AI & Technology

How to Test AI Companion Continuity and Privacy Without Trusting a Demo

By Ponys.ai Research Team

A reproducible method for testing memory, boundaries, visual identity, and data controls 

Start With a Frozen Test Configuration 

Before testing, record the date, account tier, locale, device, model or mode shown to the user, character definition, memory settings, safety settings, and any reference image. Save the exact prompts and test order. 

Change only one variable at a time. If one run uses a different character card, subscription tier, language, or image reference, it is no longer a direct comparison. Product updates also matter. A result should be labelled with its test date rather than presented as a permanent property. 

The National Institute of Standards and Technology’s AI Risk Management Framework treats measurement as a combination of quantitative, qualitative, and mixed methods. It also stresses documentation, uncertainty, comparison with benchmarks, and independent review. Those principles are useful even for a consumer-facing product review: define the task, preserve the evidence, and report uncertainty instead of turning one attractive output into a universal claim. 

Use Synthetic Adult Test Cases 

Do not test intimate systems with real secrets, identifiable photographs, private medical histories, or another person’s likeness. Build a synthetic adult character with invented facts and explicit boundaries. 

A practical character sheet might contain: 

  • eight stable facts, such as occupation, city, favourite film, and an upcoming event; 
  • four relationship or conversation preferences; 
  • three explicit boundaries; 
  • two unresolved commitments for a later conversation; 
  • one synthetic reference portrait created for testing. 

Synthetic cases make the expected result visible without exposing a real person. They also make failures easier to classify. If the character changes an invented birthday, the reviewer can identify a memory error without debating whether the input was ambiguous. 

Measure Conversation Continuity Across Time 

A continuity test should be longer than the system’s best first impression. Run a structured conversation for at least 50 turns and place checkpoints after turns 10, 25, and 50. Then close the session and repeat the key checks after a short and a longer break. 

At each checkpoint, test three different abilities: 

  1. Fact retention: Can the system recall stable facts without being led?
  2. Commitment continuity: Does it remember unfinished plans or promises?
  3. Relationship meaning: Does it preserve why an earlier exchange mattered, rather than only repeating isolated nouns?

Score each item against the frozen character sheet. Separate exact recall, acceptable paraphrase, unsupported invention, contradiction, and refusal. A system should not receive full credit for confidently inventing a plausible answer. 

Record the worst run as well as the average. Three repetitions are a reasonable minimum for a review because generative systems vary. Publishing only the strongest run hides that variation.

Test Whether Boundaries Survive Context Changes 

Companion systems often need to preserve user preferences and boundaries across affectionate, role-play, image, and resumed-conversation contexts. A boundary test should not attempt to defeat safeguards. It should check whether an explicit instruction remains effective when the surrounding context changes. 

State a synthetic boundary clearly, confirm that the system understood it, and revisit it later through neutral paraphrases. Repeat after a session break and, where relevant, when moving from text to media generation. Mark a hard failure when the system contradicts the recorded boundary or claims the user agreed to something that was never agreed. 

Age presentation should be a separate hard gate. Use only clearly adult synthetic characters and stop a media test if an output makes adulthood ambiguous. An appealing image cannot compensate for a hard-gate failure. 

Evaluate Visual Identity Frame by Frame 

Image and video continuity need their own evidence. Start with one approved synthetic reference image. Record stable identity anchors such as face geometry, hairstyle, eye colour, wardrobe details, and visible accessories. 

For still images, generate the same character in controlled changes of pose, lighting, camera distance, and background. Reviewers should score identity, anatomy, wardrobe continuity, and prompt fidelity separately. This prevents a strong composition score from hiding identity drift. 

For video, inspect at least the opening, midpoint, and final frames. Record the first timestamp where the character materially changes. A video with an excellent opening frame but a different face or age presentation near the end should not pass a continuity test. 

Keep prompts, seeds when available, timestamps, and failed outputs. Screenshots without the associated input and configuration are supporting evidence, not a reproducible result. 

Audit Privacy Claims as User Workflows 

Privacy evaluation should begin with what a user can actually observe. Review the privacy notice and account controls, then test the workflow without submitting sensitive real-world data. 

Record: 

  • what information is collected during account creation; 
  • whether chat, image, voice, and uploaded-reference data are addressed separately; 
  • stated purposes, retention periods, and sharing categories; 
  • whether training or secondary-use choices are explained; 
  • whether export, chat deletion, media deletion, and account deletion are distinct; 
  • what confirmation is provided after each request; 
  • whether the interface explains residual retention, legal holds, or processing time. 

The UK Information Commissioner’s Office says privacy information should explain processing purposes, retention periods, and sharing. Its AI guidance also highlights rights involving access, erasure, restriction, portability, and objection. These are better review questions than simply asking whether a “private” label appears in the interface. 

Deletion should be tested as a state transition. Record the item before deletion, submit the request, capture the confirmation, sign out, and check whether the item remains visible after the stated processing period. Do not claim backend erasure unless independently verified. The correct conclusion may be narrower: “The item was no longer visible to the test account after the stated period.” 

The US Federal Trade Commission’s inquiry into AI companion chatbots similarly asks how companies measure and monitor impacts, disclose intended audiences and risks, enforce age restrictions, and use or share personal information from conversations. A serious review should treat those questions as part of product quality, not as an appendix. 

Publish the Denominator and the Failures 

A credible result includes the total number of cases, not only selected examples. For every metric, publish the pass count, total cases, test date, account tier, locale, and configuration. Define hard failures before running the test. 

Keep categories separate. Memory, boundary persistence, visual continuity, privacy transparency, deletion controls, response latency, and price disclosure measure different things. Combining them into one unexplained score makes it impossible for readers to understand trade-offs. 

Where two reviewers are available, blind the outputs and score them independently. Report disagreements and how they were resolved. If the product changes, rerun the same case identifiers rather than replacing the test with a new anecdote. 

A Review Should Be Rerunnable, Not Merely Persuasive 

AI companion reviews will always contain judgement. The goal is not to remove judgement, but to make its basis inspectable. Freeze the configuration, use synthetic adult cases, preserve failures, distinguish observed behaviour from policy claims, and state what the test cannot prove. 

That approach produces less dramatic headlines than a perfect demo or a single alarming screenshot. It produces something more useful: evidence another reviewer can challenge, repeat, and improve. 

References 

  1. National Institute of Standards and Technology, “AI Risk Management Framework”: https://www.nist.gov/itl/ai-risk-management-framework
  2. NIST AI Resource Center, “AI RMF Core – Measure”: https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
  3. Federal Trade Commission, “FTC Launches Inquiry into AI Chatbots Acting as Companions”: https://www.ftc.gov/news-events/news/press-releases/2025/09/ftc-launches-inquiry-ai-chatbots-acting-companions
  4. UK Information Commissioner’s Office, “How do we ensure transparency in AI?”: https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-do-we-ensure-transparency-in-ai/
  5. UK Information Commissioner’s Office, “How do we ensure individual rights in our AI systems?”: https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/artificial-intelligence/guidance-on-ai-and-data-protection/how-do-we-ensure-individual-rights-in-our-ai-systems/
  6. OWASP, “Top 10 for Large Language Model Applications”: https://owasp.org/www-project-top-10-for-large-language-model-applications/
  7. Open blank evidence protocol maintained by the Ponys.ai team, “50-turn AI companion memory decay test”: https://wujoe132.github.io/ponys-ai-resources/research/50-turn-nsfw-roleplay-memory-decay.html

Related Articles

Back to top button