TEFM addresses the challenges of applying large language models to domains requiring structured data analysis. The framework focuses on token efficiency through the use of Behavioral Code tokens, compressing lengthy observations. This approach dramatically reduces token consumption while minimizing information loss. TEFM also prioritizes faithfulness by employing a dual-fidelity objective. This objective simultaneously optimizes code-level reconstruction and prediction-level fidelity, identifying minimal sufficient feature subsets. Experiments across clinical and security domains demonstrate competitive classification accuracy with approximately 1% token retention and 2% token retention respectively. These results indicate a substantial reduction in resource requirements.
Source: https://arxiv.org/abs/2609.09552