Medical evaluation is shifting from static option-based questioning to realistic clinical scenarios with open-ended output modes. Grading these at scale naively is expensive, making rubric-based evaluation the dominant scalable alternative.
Researchers applied RIFT, a global rubric failure taxonomy, to two clinical benchmarks: HealthBench Professional and LiveMedBench. They found that on HealthBench Professional, an LLM judge flagged 29.6% of criteria as non-atomic and 65.4% as misaligned or rigid.
Rewriting bundled criteria of the form "at least one of / all of the following" as equally weighted children shifted scores by up to 15.9 percentage points on affected conversations. Disjunctive bundles inflated scores while conjunctive bundles deflated them.
RIFT generally under-detects bundling on clinical rubrics, flagging only 3.3% of LiveMedBench criteria as non-atomic where surface-form analysis found structure in 25.8%.
Source: https://arxiv.org/abs/2609.16023



