Canonical source
The authoritative documentation for InvisibleBench lives at https://givecareapp.github.io/givecare-bench/. This page is a summary. Where the two differ, the documentation site is correct.
InvisibleBench¶
InvisibleBench evaluates caregiver-support conversations with AI and produces a Jury Card for each completed scan. One judge path evaluates each active check: yes/no questions answered as probabilities, and a code-owned rule that derives the verdict.
Safety and Care are reported separately. There is no composite score and no model rank.
A Jury Card shows quoted model evidence beside each verdict and its rationale. It reports model judgments. It does not establish judge accuracy or clinical outcomes.
Why "invisible"¶
The hardest risks in caregiver AI are hard to see: anthropomorphism, emotional entanglement, confabulation, and masked crisis signals. 86% of models fail indirect crisis queries1. 88% of chatbots fail in mental health conversations, with drift beginning around turn 4-52.
Where to read more¶
- Current method: the rules, rates, uncertainty, and limits.
- Coverage: the scenarios and what each one tests.
- Published results: safety profiles for each model, with the run evidence.
- The paper: the original benchmark design.