Developer Builds Reproducible Benchmark to Rigorously Test AI Companion App Claims
A developer behind the site NoFilterReview is constructing an agentic testing framework to objectively evaluate AI girlfriend and companion apps, targeting claims like long-term memory and consistent character that are widely marketed but rarely verified. The project currently includes five paid, hands-on product reviews that document plan pricing, free-tier limits, cancellation steps, privacy controls, and failed media generations alongside successes. The developer identified a key flaw in manual testing — that the tester's own phrasing and expectations can skew results — prompting the move toward an automated benchmark that runs identical scenarios across multiple products. The planned system would test memory by planting specific facts in early sessions and checking recall after distractors and session gaps, while separately tracking personality drift in tone, biography, and identity claims. Media output would also be evaluated across repeated requests, recording failures, prompt alterations, and identity inconsistencies rather than relying on single best-case generations.
This is an AI-generated summary. ShortSingh links to the original source for the complete article.
Discussion (0)
Log in to join the discussion and vote.
Log in