NVIDIA researchers, together with MIT and Oxford, introduced Physis-Lang. This framework adds physics reasoning to video captions, making models better at understanding physical scenes.
Traditional captions only describe what happens, not why. Physis-Lang includes explanations of causes, interactions, and effects. It also creates negative prompts to avoid unlikely outcomes.
The system uses a self-evolving loop. A captioner writes descriptions, and a physics critic scores them. The prompts are then refined based on these scores. This process improves caption quality over iterations.
The framework was tested on PhysCapBench, a benchmark with thousands of videos and assertions. Caption accuracy improved from 78.64 to 87.82 points after multiple iterations. The process also increased the number of causal steps captured.
Training data includes 183,000 videos, with 71,000 filtered and 112,000 retrieved clips. Retrieval added significant points to the models' physics understanding. The approach outperformed Veo 3.1 on most benchmarks.
Physis-Lang uses a lightweight fine-tuning method called LoRA, which adjusts attention projections without changing the model architecture. It can be run locally or via API, reducing costs.
This development offers a way to make video models more physically accurate and reliable. It can help in applications needing precise scene understanding.