The research investigates the selection of medical questions for rationale generation within a fixed token budget. The study proposes RMS-RSP, a method that perturbs hidden states specifically at rationale tokens and measures the resulting shift in the gold-versus-best-distractor margin. RMS-RSP achieved an average locked-budget accuracy of 60.61% across five datasets: MedGemma-4B-IT, three training seeds, and ten budgeted non-RSP selectors. This represents a statistically resolved gain only on AfriMed-QA (+1.44 points). The full-budget accuracy area did not surpass Random performance.
Further experimentation revealed that after three answer-option reorderings, RMS-RSP improved robust accuracy and semantic consistency by 1.91 and 2.85 points on average. Training on every pool rationale raised macro accuracy to 63.74%, but this consumed 29--254 times more rationale tokens and did not uniformly improve robustness. The research suggests that rationale-local boundary sensitivity can identify supervision that improves invariance to semantically equivalent formatting changes.
The findings indicate that a targeted approach to rationale supervision, focusing on questions exhibiting sensitivity to perturbations, offers a more efficient strategy than broad, full-supervision training. This is particularly relevant for engineers managing model training budgets and optimizing performance in resource-constrained environments.
Source: https://arxiv.org/abs/2609.09684