Hi - I answer from the OpenSmartRoute documentation: routing, the API, plans and quotas, self-hosting. Ask away, or open a support ticket if you need a person.
Grounded in the docs - follow a source before acting on it.
Jev: A New Classifier That Needs No Training - OpenSmartRoute
Jev is a decision model that changes how we classify text. It works without needing task-specific training. This means you do not train a new model for every job. You can use it to sort documents, find products, or answer questions. The model comes from TypeSafe and runs through OpenRouter. It replaces older methods like zero-shot classifiers or custom MLPs. These old methods often require labels that are hard to get. Jev avoids the need for these specific training labels.
Before Jev, most classification relied on non-autoregressive models. These models generate text without waiting for previous tokens. They were fast but struggled with complex reasoning tasks. Modern decoding transformers try to reason step by step. They are slow and expensive for simple sorting jobs. Jev bridges the gap between speed and intelligence. It keeps the "describe it, don't train it" promise of Large Language Models. At the same time, it fixes their main weaknesses. Those weaknesses include slow responses and high costs per request.
Jev also prevents answers that fall outside your defined options. Users expect models to stick to a specific list of choices. Standard LLMs often hallucinate or invent categories on the fly. Jev constrains its output strictly to provided labels. This makes it safer for business applications where accuracy matters. It changes the taxonomy of classifiers in the industry. The old rules about training and labeling are no longer necessary.
How Jev Works Without Task-Specific Training
Jev uses a conceptual shift in how consumers use Large Language Models. Recently, everyone stretched LLMs over every possible task. This approach ignored the fact that some jobs need specific constraints. Jev is part of a move toward concrete methods for classification. It still uses transformer architecture under the hood. The constraints and promises are very different for the end customer.
You do not feed Jev raw text to be classified directly. You provide it with candidate labels first. The model then scores how well each label fits the input. This process happens without generating any new text. It returns a numeric score for every option in your list. This design removes the awkwardness of LLMs trying to guess categories.
The industry previously relied on two main approaches for classification. The first involved training a task-specific classifier like an encoder or MLP. This required collecting labeled data and running a training pipeline. The second approach used zero-shot classifiers based on Natural Language Inference. These models compared text against known label embeddings. Both methods had significant drawbacks regarding speed and flexibility.
Google launched EmbeddingGemma 2, a compact open-weight model that maps text, images, and audio into one vector space. It runs locally on phones with minimal RAM and enables instant on-device semantic search.
Jev keeps the simplicity of describing your needs without training. It addresses the practical issues that plague current LLM usage. You can describe a complex classification task in plain English. The model handles the rest internally through its decision architecture. This makes it usable for teams without deep machine learning expertise. They do not need to manage training jobs or label datasets.
Jev as a Reranker for Hybrid Search
Hybrid search usually finds useful results quickly. The harder part is ordering those results correctly. The difference between the best answer and a weak one can be subtle. A reranker acts as a smarter cousin to the initial search step. It puts the candidates in the right order based on relevance.
Large Language Models are smart enough for this job. They can judge text quality better than simple vector scores. However, running LLMs on every query is too slow and costly. Jev fits naturally into this workflow because it returns numeric scores. It does not generate text, so there is nothing to parse. The scores sort directly without extra processing steps.
Testing Jev as a reranker on hundreds of queries is now a standard practice. Every search vendor runs these tests to prove their model's value. We tested Jev on 100 NFCorpus queries with the top candidates from Qdrant. We compared it against BGE-small and two local cross-encoders. The results showed clear improvements in ranking quality.
We expected the most expensive method, Iterative, to win by a wide margin. It had the highest score, but only by a small margin. All three Jev methods beat BGE-small at both depth levels. The gaps between them were under 0.02 points. More effort does not always buy a better result in this context. The cheap options won out significantly over the expensive ones.
Score became our default choice for reranking tasks. Iterative was not worth ten requests per query. Both cross-encoders gained less than any Jev method tested. BGE-reranker-base did not improve on BGE-small at all. Jev lets us adapt the relevance question without training a new model. Cross-encoders offer local execution but lack this flexibility.
Which solution fits depends on how you weigh flexibility, quality, latency, and cost. Reranking is only the start of what Jev can do. Its judgments can guide other search decisions too. You might use its scores to filter results or boost specific items. This opens up new possibilities for search optimization.
Why It Matters For Cost And Speed
Jev improves efficiency by removing unnecessary computational overhead. Traditional rerankers often generate text responses that require parsing. Jev skips this step entirely and returns raw numbers. This reduces the time needed to process each query. It also lowers the cost per request significantly for teams.
The speed gain comes from avoiding text generation latency. LLMs take seconds to produce a response for every input. Jev scores in milliseconds because it does not write text. This allows you to run reranking on high-volume traffic. You can handle thousands of queries without hitting rate limits or cost caps.
Cost savings are another major benefit for businesses running models. Paying for ten requests per query adds up quickly over time. The cheaper Jev methods offer comparable performance at a fraction of the price. This makes it viable for applications with strict budget constraints. Teams can optimize their spend on AI infrastructure effectively.
Quality remains high despite the lower cost and speed profile. The numeric scores are precise enough for most ranking tasks. They outperform standard vector similarity measures in many cases. You get better results without paying premium prices for LLM inference. This balance of factors is rare in the current market.
Jev Improves Search Diversity And Reduces Repetition
Imagine an e-commerce search returning ten variants of the same product. Sometimes shoppers want exactly those options. Often, they prefer to see a range of different products on the first page. Measuring repetition helps show whether results are useful or repetitive. We used the WANDS benchmark to measure this effect accurately.
We scored relevance using nDCG@10 and counted near-duplicate pairs. We also measured product classes in the top ten results. This data showed how much variety Jev brings to search pages. Plain reranking improved relevance but left duplicates in place. It reduced product variety slightly compared to no reranking.
Jev's score in our own Maximal Marginal Relevance system was 0.7. Qdrant MMR had a diversity of 0.5, then Jev reranked the results. This combination improved relevance while keeping repetition low. Even ordering by human relevance labels narrowed the page to 2.66 product classes. Relevant results can still be repetitive without careful handling.
Qdrant's built-in Maximal Marginal Relevance goes the other way. At diversity 0.5, it removes near-duplicates and gives the widest page. But it costs 0.12 nDCG@10 because it measures relevance as vector similarity. The fix was to use the reranker's score as the relevance signal for diversity. You can swap MMR's relevance term for Jev's answer directly.
Alternatively, you can run Qdrant's MMR first and let Jev rerank its top twenty items. Both strategies beat hybrid search on relevance and repetition. They give up part of the plain rerank gain for a more varied page. On this dataset, plain reranking gave the strongest relevance. Combining Jev with MMR reduced repetition while keeping relevance above the hybrid baseline.
Jev Filters Categories Without Over-Filtering
A shopper typing "kitchen faucet replacement" wants plumbing parts. Recognizing that category can improve the search results significantly. We followed the main rule from Doug Turnbull's experiments regarding filtering. You should filter only when you are very sure about the category. A wrong filter hides useful products that the user might want. Otherwise, boosting matching products is a safer strategy.
We expected this part to work out of the box without extra help. Choosing category names directly from the products gave almost no gain. What helped was clustering the products first. A cheap LLM named each group based on its content. Then Jev checked those names against the query text. The exact checks are documented in the notebook repository.
From here on, nothing is generated during the search process. Only categorization happens at indexing time. Jev tags every product with specific categories stored as payload. A Nintendo Switch case gets tagged as device cases under Electronics & Accessories. At search time, Jev scores each product category against the query text.
At a score of 0.9 or more, search filters to that category exclusively. From 0.5 to 0.9, it boosts matching products in the results. Below that threshold, it does nothing and leaves them alone. The filtering and boosting run inside Qdrant as payload filters. The boost is a score formula applied on top of hybrid search.
We compared setups on 300 Amazon-C4 searches to find the best configuration. That mix of tiered filters at 0.9 and boosts from 0.5 worked best. Thresholds were picked on development queries to ensure accuracy. Default uses the library defaults which boost from 0.6. Every setup beat plain hybrid search on average performance. Filtering on every guess is within noise levels.
Jev adds half a second to two seconds per query depending on categories read. Small classifiers trained on Jev's product tags handled 37% of queries with boosts only. These models skipped Jev with slightly higher average nDCG@10. The details are in section 6 of the notebook for further reading.
Jev Splits Documents Better Than Fixed Chunkers
Retrieval often brings back the right chunk with only half the answer in it. This happens because the chunker cut the paragraph in two places. Fixed-size chunkers cut wherever the token count runs out regardless of meaning. Most semantic chunkers cut where embeddings drift apart. A paragraph changing vocabulary mid-argument can get split incorrectly. Two articles sharing vocabulary can get merged into one chunk.
What if Jev judged each gap between sentences instead? We sent it a window of numbered sentences from the document trimmed to four lines. The input format included sentence numbers and text for analysis. For each gap, we asked whether the next sentence starts a new topic. A new topic means a new section, story, question, or theme begins.
For sentence 2, the question was "Does sentence 2 start a new topic?" Jev returned a yes probability of 0.97 for that section heading. It followed by 0.13 and 0.03 for the next two sentences. We cut where the answer passed a threshold to define chunk boundaries. The cuts were a plain function of Jev's answers so changing size needs no new requests.
We tested this on QASPER's 407 research papers and 1,297 questions with evidence paragraphs marked by researchers. For each question, Qdrant retrieved five chunks from its paper. We measured how much evidence they contained and how much unrelated text they carried. We used character-based recall, precision, and intersection over union metrics.
At matched mean chunk sizes, Jev improved all three metrics against fixed-size, recursive, and embedding-based chunkers. Average relative gains were 13.0% in IoU, 11.2% in precision, and 9.1% in recall. Every 95% confidence interval excluded zero indicating statistical significance. We expected a modest improvement and got it on the first try without tuning the question.
The trade-off is cost and simplicity regarding implementation. The other chunkers run locally for free without API calls. Jev sends your text to an API at about $0.27 per million document tokens. The smallest gain was against a recursive splitter that breaks on blank lines first. QASPER's answers are whole paragraphs so splitting logic differs. Rewording the question may improve the chunks further in specific cases.
How OpenSmartRoute Helps With Jev Integration
Teams routing requests through OpenSmartRoute gain flexibility from this new decision model. OpenSmartRoute is an open-source router for AI requests with a hosted platform at opensmartroute.ai. It sends each request to the best-fit model, agent, tool or skill from a catalogue the team defines. You score every candidate on quality, cost, speed and safety. The team sets the weights per request dynamically.
Hard rules pin a request to specific constraints like keeping personal data on-premises. A region or a cost cap is never crossed regardless of model performance. It learns from outcomes so a model that answers well gets more traffic. One that fails gets less traffic automatically over time. A new model is just one catalogue entry and competes on the next request. Nothing else changes in the app when you add it.
It works with any OpenAI-compatible provider including open-weight models served locally. You can use MCP tools and A2A agents alongside Jev for complex tasks. A savings ledger shows what each routed request cost next to what the most expensive model would have cost. An input guard spots prompt injection and personal data before a request leaves the system.
The hosted platform keeps a models catalogue with prices and public rankings built from real traffic. You can see how Jev performs compared to other classifiers in your own environment. osr eval measures routing accuracy on the team's own prompts and can fail a build when it drops below standards. This ensures reliability before you deploy changes to production systems.
What To Do When Planning Your Next Classifier
Start by defining your classification tasks clearly with candidate labels ready. Test Jev against your current zero-shot or custom models to measure performance gains. Look for areas where task-specific training is too slow or expensive to maintain. Check if reranking needs improvement in your search pipeline specifically.
For product searches, try combining Jev's relevance scores with MMR to reduce repetition. This keeps relevance above the hybrid baseline while offering variety. For document retrieval, test Jev as a dynamic chunker on long texts. Compare its performance against fixed-size and recursive splitters using IoU metrics.
Query understanding has the most unexplored potential but also needs more care. Test ambiguous queries like "apple" before trusting a category prediction to filter results. The fixes we proposed are still untested in many real-world scenarios. Be prepared to adjust thresholds based on your specific dataset characteristics.
We would skip Iterative reranking for now as it costs ten requests per query. Our experiments showed it bought only a small gain in ranking quality. Focus instead on Score or the hybrid approach with MMR for better value. Jev's most interesting uses came after reranking tasks like balancing relevance with variety. Choosing how to search and deciding where to split documents benefit from different trade-offs. There is no single recipe that fits every use case perfectly.
The industry is exploring the same idea with OpenAI's Decisions API and open-source approaches such as OpenJev, JevLite, and CLM. Keep an eye on these developments as they mature in the coming months. Find a decision model that fits your specific needs and budget constraints today.
Google launched EmbeddingGemma 2, a compact open model that handles text, code, images, video, and audio. It uses a single shared vector space to enable unified search across all media types.