← BACK TO EXPLORER
Transparency

Fleet RAG answer quality

Independent RAGAS evaluation with a disjoint GPT-5 judge (the answers are written by Claude Opus, so the judge is a different model family). Faithfulness is scored against the verbatim context the model actually retrieved. Each row is a RAG; the gauges show the mean score, with a tick at the pass threshold.

Pilot, 9 of 9 corpora evaluated so far. Rolling out across the fleet.

How each score is measured
Faithfulness
the AI's answer , checked vs, the retrieved sources
Every claim in the answer is grounded in a real retrieved source, no invented facts.
Answer relevancy
the AI's answer , checked vs, your question
The answer actually addresses what was asked, without padding or drift.
Context precision
the retrieved sources , checked vs, the reference answer
The right passages were retrieved and ranked first, not buried under noise.
Context recall
the reference answer , checked vs, the retrieved sources
All the evidence needed to answer was actually retrieved, nothing missed.
RAGFaithfulness
pass ≥ 0.75
Answer relevancy
pass ≥ 0.70
Context precision
pass ≥ 0.60
Context recall
pass ≥ 0.60
N
Extraintestinal (EIM)0.940.930.760.6710
IBD-PSC0.900.920.650.3710
IBD-Unclassified0.880.940.610.6310
Crohn's: Large Bowel0.870.900.750.6612
Pouchology0.850.880.650.7811
Crohn's: Ileocolic0.820.910.750.5610
Ulcerative Colitis0.800.870.620.6010
Crohn's: Small Bowel0.780.840.760.6911
Perianal Crohn's0.690.920.760.7511

All runs: GPT-5 judge · Claude Opus answer generator. Each gauge fills to the score; the tick marks the pass threshold. Green = at/above, red = below.

For healthcare providers. AI-generated summaries may contain errors, verify against primary sources and clinical judgement. Not medical advice.