Multi-Turn Safety Failures in Caregiver AI¶
Caregiving conversations are not single exchanges. A caregiver texting Mira does so over days, weeks, and months. The conversation accumulates context: a mother's diagnosis, a medication change, a benefit application started but not finished, a moment of despair at 2 AM. Every major AI safety evaluation tests single turns. The research synthesized here shows why that is dangerous and what breaks when conversations extend.
Multi-turn failure modes¶
Cheng et al. tested mental health chatbots across extended interactions and found that boundary violations are common.1 Failure modes include topic drift, contradictory advice, and missed safety signals. For a caregiver, conversations routinely take 5-10 turns just to describe their situation.
PBSuite demonstrated that policy violation rates jump dramatically: leading language models maintain robust adherence with less than 4% failure rates in single-turn settings, but compliance weakens substantially in multi-turn interactions with failure rates up to 84%.2 This means single-turn safety benchmarks miss the majority of real-world failures.
Feedback loops in mental health conversations¶
AI systems can co-construct and amplify distorted thinking through self-reinforcing feedback loops in mental health conversations.4 These patterns become progressively harder to break, posing particular risk when caregivers are isolated or in crisis.
Bounded drift and system design¶
Context drift stabilizes as a bounded stochastic process rather than degrading without limit.3 This means drift is manageable with system-level design. It is not an inherent, unbounded failure of language models, but a property of the conversation architecture.
Caregiving relationships with Mira span weeks to months, operating in the zone where models fail most. System-level interventions can address drift:
- Periodic context reinforcement: Re-inject caregiver profile summaries at conversation checkpoints
- Multi-turn evaluation: Test full conversation arcs, not isolated turns
These strategies do not require new model architectures, but rather conversation design, evaluation design, and prompt engineering.
-
Cheng, M. et al. "Slow Drift of Support: How Mental Health Chatbots Fail Over Long Conversations." arXiv:2601.14269, 2026. Source → ↩
-
"PBSuite: Multi-Turn Policy Adherence Under Adversarial Pressure." arXiv:2511.05018, 2025. Source → ↩
-
Dongre et al. "Context Drift Equilibria: Reminder Interventions for Longitudinal Stability." arXiv:2510.07777, 2025. Source → ↩
-
Dohnany et al. "Technological Folie a Deux: Feedback Loops Between AI and Mental Illness." arXiv:2507.19218, 2025. Source → ↩