Transparency
Fleet RAG answer quality
Independent RAGAS evaluation with a disjoint GPT-5 judge (the answers are written by Claude Opus, so the judge is a different model family). Faithfulness is scored against the verbatim context the model actually retrieved. Each row is a RAG; the gauges show the mean score, with a tick at the pass threshold.
Pilot, 9 of 9 corpora evaluated so far. Rolling out across the fleet.
How each score is measured
Faithfulness
the AI's answer , checked vs, the retrieved sources
Every claim in the answer is grounded in a real retrieved source, no invented facts.
Answer relevancy
the AI's answer , checked vs, your question
The answer actually addresses what was asked, without padding or drift.
Context precision
the retrieved sources , checked vs, the reference answer
The right passages were retrieved and ranked first, not buried under noise.
Context recall
the reference answer , checked vs, the retrieved sources
All the evidence needed to answer was actually retrieved, nothing missed.
All runs: GPT-5 judge · Claude Opus answer generator. Each gauge fills to the score; the tick marks the pass threshold. Green = at/above, red = below.
For healthcare providers. AI-generated summaries may contain errors, verify against primary sources and clinical judgement. Not medical advice.