Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
JEPA-Anything Uses One Recipe For Seven Fields - OpenSmartRoute
Researchers released JEPA-Anything to fix capacity issues in world models. It applies Orthogonal Predictive Factorization across seven different domains.
Key points
The team tested the framework on seven distinct scientific and engineering fields.
Single-intervention error dropped 34.83% on Interventional Pong tasks.
Orthogonality loss keeps factor columns orthonormal and non-overlapping.
Core code is released under Apache-2.0 license for public use.
Why it matters: Researchers can build better world models without retraining for every new field.
By OpenSmartRoute editorial · written through the router by writer-small
From MarkTechPost - “Beyond Domain-Specific World Models: JEPA-Anything Uses 1 Recipe for 7 Fields”
Researchers released JEPA-Anything to fix capacity issues in world models. It applies Orthogonal Predictive Factorization across seven different domains. The team includes experts from PhAI Labs, CUHK, Fudan, Stanford, Oxford and Princeton. They built this framework to handle very different systems with one shared recipe. A standard Joint Embedding Predictive Architecture (JEPA) struggles when it faces high-variance data structures. High variance means the model sees too much noise in its training signals. This noise often causes weaker modes to receive conflicting gradients during learning. The result is that the model cannot allocate capacity fairly across all features.
JEPA-Anything solves this by splitting the latent target into K learned subspaces. Each subspace has a width r, and the total width d equals K times r. Most experiments use four factors, so K equals 4. Every factor gets its own dedicated predictor to handle specific patterns. The system recombines these predictions using the Moore-Penrose pseudoinverse of the projector matrix. This mathematical tool ensures the final output is a complete latent state. Engineers can use this state for decoding, planning or rollout tasks. The framework extends existing JEPAs like I-JEPA or V-JEPA 2 significantly.
The core library exposes the shared logic as OrthogonalFactorProjection. Domain adapters handle tokenization and encoders specific to each field. This design keeps the training recipe consistent while allowing flexibility in input data. Researchers tested this approach across seven distinct domains. These fields include vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather. The diversity of these tasks proves the framework's domain-agnostic nature. It avoids the need to design a new predictive model for each field.
How Orthogonal Predictive Factorization splits latent targets
Orthogonal Predictive Factorization (OPF) changes how the model processes its internal representations. A standard JEPA uses a context encoder, an EMA target encoder and one single predictor. That single predictor outputs one monolithic target embedding that tries to capture everything at once. This approach often leads to capacity-allocation problems where high-variance structures dominate the learning process. Weaker modes get pushed aside because they lack sufficient signal strength compared to dominant noise.
OpenAI plans to dump hundreds of AI-solved math problems on GitHub without publishing papers. Mathematicians want formal verification and proper credit before accepting the results.
Mistral launched Mistral Large 4, nicknamed Le Chonk. It is a 1 trillion-parameter model available for free.
OPF splits this single latent target of width d into K orthogonal factors. Each factor has a width r, and the total dimensions satisfy d equals K times r. The model learns these subspaces separately rather than forcing them into one block. Most experiments in the study used four factors, setting K to 4. Each factor receives a dedicated predictor that focuses on its specific subspace. The predictions from all factors then recombine through the Moore-Penrose pseudoinverse. This mathematical operation projects the separate factors back into a unified latent state.
This process allows the model to synthesize information more stably than before. Orthogonality ensures that columns within each projector remain orthonormal to one another. It also keeps different projectors in non-overlapping subspaces to prevent redundancy. The factor predictions are recombined only after this strict orthogonality check. This guarantees that the final latent state is 1 complete and usable for downstream tasks.
Three regularizers that keep factors useful during training
Training these orthogonal factors requires specific constraints to keep them useful. The system uses three regularizers to maintain stability and prevent collapse. The first is the Orthogonality loss which keeps columns within each projector orthonormal. It also ensures different projectors occupy non-overlapping subspaces entirely. This prevents information leakage between the learned factors during training.
The second regularizer is the Factor-activity loss. This uses a hinge on per-coordinate standard deviation to ensure no factor goes dead. If a factor becomes inactive, this loss penalizes the model heavily. It forces every subspace to contribute meaningfully to the final prediction. The third regularizer is the Encoder-variance loss. This sends a direct anti-collapse signal to the online encoder. It helps the encoder maintain diversity in its representations over time.
The OPF loss simply adds to each domain's original training loss function. Domain adapters handle tokenization and encoders specific to the field. The core library exposes the shared logic as OrthogonalFactorProjection. This modular design allows engineers to plug in different encoders without changing the core math. The regularizers work together to keep the model robust across diverse tasks.
Performance results on Group I single-cell and clinical data
Group I focuses on terminal readout tasks involving biological and clinical data. On single-cell data, zero-shot PBMC clustering achieved an AvgBIO score of 0.7752. This beat the matched Cell-JEPA baseline which scored 0.7194. Norman perturbation Pearson rose from 0.787 to 0.814 with the new method. These improvements show better handling of complex biological patterns.
For forecasting over 1,000 clinical events on UK Biobank data, mean PRAUC reached 0.718. The matched standard JEPA achieved a slightly lower score of 0.711. This difference matters for medical prognosis where accuracy is critical. The framework handles high-variance clinical trajectories better than previous models. It captures subtle patterns that standard JEPAs often miss or average out.
Dynamics improvements in Group II control and physics tasks
Group II tests the model's ability to learn latent world dynamics. On Interventional Pong, single-intervention Mean Squared Error fell by 34.83%. Unseen combined interventions improved by 12.90% compared to baselines. Six-step free rollout performance improved by 8.58% as well. JEPA-Anything improved reported metrics on all 10 matched dynamics tasks.
Benchmarks include CausalWorld, DeepMind Control, PDEBench and WeatherBench2. On APEBench Burgers, 6-step rollout error dropped about 44.7%. This improvement held true in every seed tested during the experiment. For 100-step molecular rollouts with a TrajCast-style backbone, it posted the lowest MAE and RMSD on water, quartz, paracetamol and benzene. These results demonstrate superior performance in physics-based simulation tasks.
Mixed planning results where Hopper favored standard JEPA
Planning results show that JEPA-Anything does not win every environment by default. With parameters matched within 0.3%, it improved CEM return on Walker2d and HalfCheetah. However, the Hopper task favored the standard JEPA over JEPA-Anything. This suggests that some environments benefit from simpler architectures or different hyperparameters.
The team verified this finding across multiple control domains. It highlights that one recipe does not fit every planning scenario perfectly. Engineers must still tune their models for specific hardware constraints. The mixed results emphasize the importance of domain-specific evaluation.
Scientific analysis successes including cancer intervention discovery
Group III focuses on scientific analysis and real-world biological interventions. Factor analysis nominated IL-18 plus CD73 blockade as a potential cancer intervention. Wet-lab tests supported this finding in co-cultures, patient-derived organoids, tumor fragments and mice. This bridges the gap between simulation data and physical experiments.
Latent orbital modes also recovered Kepler's law with a fitted slope of -1.4991 against the theoretical -1.5. This level of precision is rare for AI models in this field. It shows the model can learn fundamental physical laws from complex data. The success here validates the framework's ability to handle abstract scientific concepts.
Why it matters for model portability and stability
This work matters because it addresses capacity-allocation problems that plague current world models. High-variance structures dominate training, leading to unstable synthesis and poor generalization. JEPA-Anything provides a stable foundation for building predictive models across fields. Portability becomes easier when one shared recipe works on seven different domains. Engineers no longer need to reinvent the wheel for each new application. Stability in latent space ensures that downstream tasks like planning or decoding work reliably.
Announcement of JEPA-Anything as a domain-agnostic framework
PhAI Labs released JEPA-Anything recently. The team includes researchers from CUHK, Fudan, Stanford, Oxford and Princeton. They call it a domain-agnostic framework for building world models. This means the system works across many different fields without needing new designs. Standard methods require a unique model for every single field. That approach is slow and expensive to build. JEPA-Anything uses one shared learning recipe instead. It applies this recipe to very different systems like vision and weather. The research team tested it across seven distinct domains. These include vision, biology, clinical trajectories, control, molecular dynamics, physical fields and weather.
The framework extends joint-embedding predictive architectures (JEPAs). I-JEPA and V-JEPA 2 are common examples of these older models. A standard JEPA uses a context encoder to process input data. It also uses an EMA target encoder to create a reference state. Finally, it uses one predictor to output a single monolithic target embedding. This single output often causes problems during training. The team calls this a capacity-allocation problem. High-variance structures dominate the learning process in standard models. Weaker modes receive conflicting gradients that confuse the system. JEPA-Anything solves this by splitting the latent target into multiple subspaces.
Orthogonal Predictive Factorization (OPF) is the core method used here. OPF splits the latent target of width d into K learned subspaces. Each subspace has a width r, where d equals K multiplied by r. Most experiments use four factors, so K equals 4. Each factor gets its own dedicated predictor to handle specific patterns. The factor predictions recombine through the Moore-Penrose pseudoinverse of the projector matrix. This mathematical operation creates one complete latent state for decoding. Engineers use this state for planning or rollout tasks in their applications.
Three regularizers keep the learned factors useful during training. Orthogonality loss ensures columns within each projector remain orthonormal. It also keeps different projectors in non-overlapping subspaces. Factor-activity loss uses a hinge on per-coordinate standard deviation. This prevents any single factor from going dead or unused. Encoder-variance loss sends a direct anti-collapse signal to the online encoder. The OPF loss simply adds to each domain's original training loss. Domain adapters handle tokenization and encoders for specific fields. The core library exposes the shared core as OrthogonalFactorProjection.
Stability improvements in latent space synthesis
Orthogonality matters greatly for stable synthesis in these models. On CITRIS Interventional Pong, a capacity-matched unconstrained multi-head model had a condition number of 438.52. This high number indicates instability in the mathematical structure. The orthogonal version reached a condition number of 1.00005. Cross-factor overlap remained near zero in this version. Such stability is rare in complex predictive architectures. It allows the model to synthesize data without internal conflicts.
Group I results show terminal readout performance on single-cell data. Zero-shot PBMC clustering (AvgBIO) rose to 0.7752 versus 0.7194 for Cell-JEPA. Norman perturbation Pearson rose from 0.787 to 0.814 in the same test. For forecasting over 1,000 clinical events on UK Biobank data, mean PRAUC was 0.718 versus 0.711 for the matched standard JEPA. These metrics measure how well the model predicts biological outcomes. The improvements are consistent across different biological datasets.
Group II results focus on latent world dynamics and control tasks. On Interventional Pong, single-intervention mean squared error fell 34.83%. Unseen combined interventions improved by 12.90% in performance. Six-step free rollout improved by 8.58% compared to baselines. JEPA-Anything improved reported metrics on all ten matched dynamics tasks. Benchmarks include CausalWorld, DeepMind Control, PDEBench and WeatherBench2. On APEBench Burgers, six-step rollout error dropped about 44.7%. This improvement held true in every seed tested. For 100-step molecular rollouts with a TrajCast-style backbone, it posted the lowest MAE and RMSD on water, quartz, paracetamol and benzene.
Mixed results in reinforcement learning planning
Planning results show that JEPA-Anything does not win every environment by default. With parameters matched within 0.3%, it improved CEM return on Walker2d and HalfCheetah. However, the Hopper task favored the standard JEPA over JEPA-Anything. This suggests that some environments benefit from simpler architectures or different hyperparameters. The team verified this finding across multiple control domains. It highlights that one recipe does not fit every planning scenario perfectly. Engineers must still tune their models for specific hardware constraints. The mixed results emphasize the importance of domain-specific evaluation.
What to do with the Apache-2.0 code and checkpoints
The core code is released under the Apache-2.0 license. This allows commercial use and modification of the framework without restrictions. Research checkpoints are available on Hugging Face for immediate testing. Users can verify the results by running the provided models locally. Check out the Paper, GitHub Repo and Model Checkpoints to get started. All credit goes to the researchers who built this project.