This study presented a rubric-guided large language model (LLM) solution for opioid use disorder (OUD) computable phenotyping. The framework utilized an 18-item rubric to instruct the LLM in automatically extracting critical text and supporting evidence for determining OUD flags. The system was evaluated on a dataset of 253 patients, including 68 confirmed OUD cases. The LLM-based CP demonstrated superior performance compared to existing methods, achieving an F1 score of 0.774 and an AUROC of 0.934. This represents a relative improvement of 12.8% and 44.4% over machine learning-based CP and zero-shot LLMs, respectively.
Source: https://arxiv.org/abs/2609.05682