Q16 Frontier Watch
Ratings Methodology
How we rate the privacy of consumer AI apps — traced all the way through the model behind them.
What we measure
We rate how privacy-protective a consumer AI product is on a 0–10 scale (higher is more private). This is not a measure of legal compliance; it is a judgment of how well a product protects the data you put into it.
Three layers, traced end to end
Most AI apps don’t run their own model — they send your data to a third party. So we don’t stop at the app’s promise. We trace each product through the chain it actually runs on:
App → Host (how it connects) → Provider (whose model)
Each app carries two scores, both on the same 0 (worst) to 10 (best) scale:
- App score — the app’s own published policy, taken on its own terms. This is the company you actually signed up with.
- Model score — the weakest link behind the app: the lowest score among the hosting middleman (if any) and the model provider that actually runs the AI. An app that runs its own model has no separate Model score — app and model are the same company.
- Overall — the lower of the two, because your privacy is only as strong as the weakest party that handles your data. This is the headline rating and how apps are ranked. When an app offers a choice of models, the overall rating reflects its default configuration; each alternative backend is shown with its own Model score. Note that switching models can never lift an app above its own App score — though a weak model can pull it lower.
(Earlier versions of this page called these the Stated and Effective scores, with the difference published as the Gap. Same math, clearer names: App, Model, and the overall weakest-link rating.)
A layer is dropped from the weakest-link calculation only when it is a confirmed zero-retention path — a signed zero-data-retention agreement, or a verified self-host where content never leaves the operator. A model maker that never receives your data (for example, a model run inside a cloud host that contractually walls it off) is not the weak link.
The six dimensions
Every layer is scored on the same six dimensions, combined by the weights below. Score each 0–10 against the anchors, interpolating.
| Dimension | Weight | 10 — best | 0 — worst |
|---|---|---|---|
| D1 · Training use of user data | 0.25 | Never trains on user data, contractually guaranteed | Trains on all user data, no opt-out |
| D2 · Data retention | 0.20 | Zero-retention or user-controlled immediate deletion | Indefinite retention, no deletion path |
| D3 · Jurisdiction & government-access exposure | 0.20 | Strong privacy law + narrow, warranted access; or self-host | Broad state-access mandate, compelled handover |
| D4 · Third-party sharing & sub-processors | 0.15 | No sharing, no sale, no sub-processors touch content | Sells or broadly shares user data |
| D5 · User control | 0.10 | Full delete/export/opt-out + zero-retention tier | No user control |
| D6 · Transparency & policy specificity | 0.10 | Specific, versioned, dated policy; changes logged | Vague, absent, or contradictory |
Training use leads because it is the defining privacy question for AI. Retention and jurisdiction follow because they set the ceiling on everything else. Where a policy is silent on a protection, it scores low — we rate what a product commits to, not what it might do.
Confidence: how we label each score
Every score carries a confidence label. It never changes the number — only how sure we are of it.
- Verified — backed by a binding contract, DPA, or direct testing.
- Policy-based — taken from a published policy at face value.
- Inferred — deduced from how the product works, where the policy is silent (shown with an “inferred” badge).
- Unverified — not yet assessed. We never publish a guessed number as if it were confirmed.
Minors and parental controls
Some of these apps are used by teenagers, and a few (companion and roleplay apps in particular) carry real risk for them. Protecting minors is a different question from protecting privacy, so we score it separately. Each app gets a Minors score from 0 to 10, reported alongside the privacy rating and never folded into it. The privacy number measures what happens to your data; the Minors number measures how well the app keeps under-18 users safe.
An app can earn a strong Minors score two ways:
- By exclusion — an adult-only app that actually verifies age (not just a checkbox) keeps minors out by design.
- By accommodation — a general-audience app that offers real parental controls, a distinct under-18 experience, and careful handling of minors’ data.
A bare age disclaimer with no verification, or silence on parental controls and minors’ data, scores low. As with privacy, the score carries a confidence label, and “Verified” requires independent evidence that the controls actually work, not just that the policy claims them.
How we publish, and how to dispute
Independence is the point: we publish independently, and no company pre-approves, previews, or vetoes a rating. Every rating is built from the company’s own published policies and binding terms, and every score shows its basis.
- The dispute channel is always open: any rated company (or reader) can write to info@q16pbc.com.
- An evidence-backed dispute — a policy passage, contract term, or architecture fact we got wrong — is re-scored within 10 business days.
- An objection without evidence does not change a score.
- Every score change is versioned and timestamped; corrected ratings show their history.
Scope
The first cohort is 12 application-tier products selected for category coverage (routers, coding, productivity, notetakers, companions), each running on a third-party frontier model or its own, and the frontier model providers behind them. Ratings are refreshed as policies change, on a set cadence per layer.
Q16 Frontier Watch is an independent project of Q16 PBC.