Ollama has added support for decision models based on TypeSafe’s Jev API. This API allows fast, typed decisions on local machines. The new feature is available in Ollama version 0.35 via the /v1/systemone endpoint.
Users can send text as state with questions, and the model responds with answers in one request. This setup is useful for tasks like ticket triage, model routing, and content moderation. Because the models run locally, requests are faster and do not need to go over a network.
One model, Nimble 9B, runs on a MacBook Pro M5 Max and averages 91 milliseconds per decision. This speed makes it suitable for real-time applications like gaming or content filtering. Other models include Tev1 4B and Tev1:0.8B, both from Together AI.
Bespoke Labs evaluated Nimble and Tev1 on public data sets. Nimble and Tev1 performed well, with accuracy measured across 3,880 decisions. More models are expected to be added soon, including cloud-based options.
To get started, download or upgrade Ollama to the latest version. Then, download a decision model like Nimble. You can use curl commands or the TypeSafe SDK to make requests. Future updates will improve performance on Apple Silicon and introduce new models.
Source: https://ollama.com/blog/ollama-now-supports-jev-style-decision-models