Skip to content

LLMs1 min read

Qwen3.8-2.4T-A95B Model Now Available on NVIDIA GB300 NVL72

Alibaba has released the open weights for Qwen3.8-2.4T-A95B, a 2.4 trillion parameter model, allowing near-frontier capabilities to be deployed on NVIDIA GB300 NVL72 systems. This enables engineers to run large language models with configurable reasoning.

By OpenSmartRoute editorial · written through the router by writer-small

From NVIDIA technical blog - “Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

Decorative object.
Decorative object.. Image: NVIDIA technical blog (original)

Alibaba has made the open weights for Qwen3.8-2.4T-A95B (Qwen3.8-Max) available. The model has a parameter count of 2.4 trillion. Deployment is supported on NVIDIA GB300 NVL72 systems.

This release brings near-frontier capabilities to the open-weight space. The NVIDIA GB300 NVL72 is designed for large model inference. It provides a platform for running models of this scale.

Configurable reasoning is supported. This allows for adaptation of the model's behavior based on specific requirements. The system is optimized for efficient inference of large language models.

Source: https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/

Published Aug 24, 2026 · updated Sep 8, 2026 · 85 words

Keep reading

Related posts

More in LLMs

LLMs1 min read

Hugging Face: Topic Safety Restrictions

The MultiverseComputingCAI research explores restricting topic safety for large language models, focusing on specific subsets rather than broad prohibitions. This approach aims to reduce the risk of unintended consequences while maintaining model utility.