The research investigates the challenges of admitting updates into continual learning systems for embodied agents. A range-based confidence gate was found to fail in certifying unchanged old-task behavior within substantial budgets. This suggests a need for a more nuanced approach to update admission. A standard paired-binomial construction reduced the burden when outcome disagreements were rare, admitting 31.6% of a common update stream across 2,000 episodes per stage. Unconditional replay demonstrated better learning in closed-loop runs. The study also included a learned-dynamics stress test to differentiate model bias from feedback-selection error.
The proposed admission-audit protocol incorporates analytical and synthetic evidence. It specifies certified historical-reference promotion and a round-level missed-opportunity metric. A constructed one-step pushing diagnostic with 32 seeds utilized this protocol, admitting 31.6% of a common update stream. This contrasts with the range-based gate, which admitted zero updates.
Physical-robot and VLA validation remain open. The research contributes an admission-audit protocol with analytical and synthetic evidence.
Source: https://arxiv.org/abs/2609.10873