NVIDIA announced NV-Reason-CT, a new vision language model built specifically for 3D computed tomography (CT) scans. Unlike older models that treat slices independently, NV-Reason-CT processes the entire volume as a true 3D input. This approach preserves spatial relationships across all slices, which is essential for understanding complex anatomical structures.
The model combines a full 3D vision transformer encoder with a language model trained to generate chain-of-thought reasoning. It produces detailed structured reports covering many abnormalities in the chest and abdomen. The reports follow a systematic review process similar to radiologists, including differential diagnoses and confidence levels.
NV-Reason-CT can also engage in multistep conversations. Clinicians can ask follow-up questions or request clarifications, making the model a more interactive diagnostic partner. It supports clinical workflows by providing step-by-step reasoning that can be audited and trusted.
The model achieved state-of-the-art results on the CT-RATE benchmark, with a Macro-F1 score of 0.614 and an AUROC of 0.871. Radiologists validated its structured reports and reasoning traces, noting the step-by-step thinking improves trust and safety.
NV-Reason-CT is an open research foundation. It is designed for researchers and developers to fine-tune for specific CT applications. It is not a clinical product or autonomous diagnostic system. Users can adapt it for their own specialized workflows.



