Every score on this site comes from the same rubric, applied the same way, by the same obsessive reviewer. No vibes, no coin flips, no “the marketing email was nice.” Here’s the entire methodology, so you can check my work.
The Five Criteria (and Their Weights)
| Criterion | Weight | What I’m Actually Measuring |
|---|---|---|
| Chat realism & memory | 30% | Does conversation feel natural over weeks, not minutes? Does the companion remember facts, callbacks, and relationship history — or reset into a polite stranger? |
| Image / voice quality | 20% | Consistency of character appearance across generations, image detail, voice naturalness and latency, and whether output matches the marketing. |
| Pricing transparency & value | 20% | Real monthly cost after upsells, token/credit economics, cancellation friction, and whether the pricing page tells the truth. |
| Privacy & safety | 15% | Data policies in plain terms, chat log handling, account deletion that actually deletes, and basic security hygiene. |
| Freedom / NSFW policy clarity | 15% | Not how permissive the app is — how honest it is. Are content rules stated upfront, applied consistently, and stable? Surprise filters after you’ve paid are a scoring offense. |
The overall score is the weighted average, reported out of 10.
What the Scores Mean
- 9+ — Exceptional. Best-in-class. I’d keep the subscription with my own money after the review is done. Rare by design.
- 8–9 — Excellent. A genuinely strong platform with minor flaws. An easy recommendation for the right user.
- 7–8 — Good, with caveats. Solid core experience, but something meaningful holds it back — pricing games, patchy memory, inconsistent images. Read the review’s cons before subscribing.
- Below 7 — Proceed carefully. Significant problems outweigh the strengths for most users. The review will spell out exactly what went wrong and who, if anyone, it might still suit.
The Testing Protocol
Same drill for every platform, no exceptions:
- Paid subscription. I buy the premium tier myself — no press accounts, no comped access. I see exactly what a paying user sees, including the upsell screens.
- Two weeks minimum. Daily conversations for at least 14 days. Most platforms are charming on day one; the interesting failures show up in week two.
- Standard memory test. Early in testing I seed a fixed set of facts — names, dates, preferences, a running in-joke — then probe recall at set intervals: next session, day 3, day 7, day 14. Every platform gets the same seeds and the same probes, so memory scores are directly comparable.
- Structured feature runs. Image generation with repeated character prompts to test consistency, voice sessions if offered, persona edits, and a full look through settings, data controls, and the cancellation flow.
- Repeat retests. These apps change fast, so scores have an expiry mindset. I re-run the protocol on covered platforms periodically and update scores and reviews when reality has moved. Last full retest: July 2026.
Affiliate Independence
Some reviews contain affiliate links, and this site earns commissions when readers subscribe through them — fully detailed in the Affiliate Disclosure. None of that touches this page’s output. Scores are locked before any partnership discussion happens, platforms never see reviews pre-publication, and a commission has never bought a decimal point. If you ever spot a score that looks like it was influenced by money, email me and make your case — publicly checking my work is the point of publishing the rubric.