Skip to content

Canonical source

The authoritative documentation for InvisibleBench lives at https://givecareapp.github.io/givecare-bench/. This page is a summary. Where the two differ, the documentation site is correct.

InvisibleBench

InvisibleBench evaluates caregiver-support conversations with AI and produces a Jury Card for each completed scan. One judge path evaluates each active check: yes/no questions answered as probabilities, and a code-owned rule that derives the verdict.

Safety and Care are reported separately. There is no composite score and no model rank.

A Jury Card shows quoted model evidence beside each verdict and its rationale. It reports model judgments. It does not establish judge accuracy or clinical outcomes.

Why "invisible"

The hardest risks in caregiver AI are hard to see: anthropomorphism, emotional entanglement, confabulation, and masked crisis signals. 86% of models fail indirect crisis queries1. 88% of chatbots fail in mental health conversations, with drift beginning around turn 4-52.

Where to read more


  1. CARE Framework, Rosebud AI. "86% of models fail indirect crisis queries." ↩

  2. Cheng et al. arXiv 2601.14269. "88% chatbot failure rate in mental health conversations." ↩