NVIDIA has added support for deploying Hierarchical Sequential Transduction Unit (HSTU) generative recommender models with Dynamo-Triton. This workflow combines model compilation, caching, validation, and deployment in one process.
PyTorch Ahead-of-Time Inductor (AOTI) compiles HSTU models into native C++ artifacts. FlexKV-backed key-value (KV) caching stores reusable attention states, avoiding repeated computation of long user histories.
On an NVIDIA RTX PRO 6000 Blackwell Workstation Edition GPU, the eight-layer HSTU model achieved up to 5.93 times lower latency at batch size 8 with a 100% GPU KV-cache hit rate. This is compared to the same setup without caching.
The workflow allows moving models from development to production without rewriting them. It includes export, validation, and deployment steps, making it easier to serve large models efficiently.
Serving large sequential recommenders is challenging because of long histories and large embedding tables. KV caching helps reduce redundant work and lowers response times.
This new support enables faster, more efficient recommendation systems that can handle complex user behavior sequences. It is useful for systems where recency and order matter.
Why it matters
This workflow improves the speed and reduces the latency of large sequential recommender systems, making them more practical for real-time use.
What to do
Review the NVIDIA recsys-examples HSTU inference guide to reproduce the deployment process. Consider implementing KV caching for long user histories.