Using Jev to improve search quality

Since getting access to Jev, the team's been exploring how it can improve our classification workloads. We didn't originally plan to test it as a reranker, but decided to try it out after seeing some early signs that it may have some use there. We found it outperformed off-the-shelf rerankers and came close to frontier models in quality.
Our team spends a lot of time on search. Sales data is extremely messy and is often scattered across many sources like Slack, Gong, Salesforce, Notion, and Google Drive. We break search down into two problems: extracting useful signal from noise and serving the most relevant context when a user or agent needs it. This post covers reranking, just one component of the latter.
A quick rundown of Jev
Jev is a generalized decision model from TypeSafe (a "System One" model as they call it). Unlike LLMs you are probably more familiar with, it doesn't generate text or write code. Instead, you just give it data and questions about the data, and it returns structured answers your code can use.
There are three types of outputs it supports:
- Choice selects from a set of options and returns probabilities for those options.
- Score evaluates the input against a scale you define.
- Noul returns a probability between 0 and 1 for a yes/no question.
AI agents run into surprisingly many classification and decision tasks, from routing requests to checking whether something meets an approval policy or judging a specific part of output. You'd typically hand these tasks to a small LLM like GPT-6 Luna or a purpose-built fine-tuned model, but Jev gives us a new cost-effective, intelligent option. So far, it's been cheap, consistent, and very capable.
Using it as a reranker
I tend to think of Jev as a model for classification tasks, so using it as a reranker felt a little odd at first. But it makes a ton of sense when you frame it as "does document X help answer input query Y?" With Jev, you can run Noul and get a confidence score for each document. You can then just sort the results by scores and voilà, you've done reranking.
Candidate documents can be batched into a single request with one Noul question per document. A request looks roughly like this:
{
"model": "jev-latest",
"state": {
"query": "Did Acme ask for SSO during the security review?",
"documents": {
"doc_0": "Security review notes: Acme's IT team said SSO through Okta is a hard requirement...",
"doc_1": "Q3 pipeline review: Acme moved to stage 3 after the pricing call..."
}
},
"questions": {
"doc_0": {
"type": "noul",
"instructions": "Does `documents.doc_0` help answer `query`? Prefer passages with the specific facts needed.",
"criteria": {
"true": "Contains specific information that answers or is necessary for answering the query",
"false": "Unrelated, only tangentially related, or lacks the needed facts"
}
},
"doc_1": {
"type": "noul",
"instructions": "Does `documents.doc_1` help answer `query`? Prefer passages with the specific facts needed.",
"criteria": {
"true": "Contains specific information that answers or is necessary for answering the query",
"false": "Unrelated, only tangentially related, or lacks the needed facts"
}
}
}
}Jev then answers every question with a probability and those become the ranking scores:
{
"answers": {
"doc_0": { "type": "noul", "noul": 0.94 },
"doc_1": { "type": "noul", "noul": 0.06 }
}
}In many pipelines, it's common to run searches for an initial list of candidates then pass the results to a reranker to bring the most relevant results to the top. The ordering matters as we want to pass the highest quality payload back to agents. At Opine, we use a hybrid search pipeline combining multiple techniques to tune for relevancy: semantic search, full text search (BM25), query rewriting, date decay, reciprocal rank fusion, reranking, and more.
A query fans out to two searches that run in parallel
The setup
Jev went up against several off-the-shelf cross-encoder rerankers: Cohere Rerank 3.5, Rerank 4 Fast, Rerank 4 Pro, and Voyage Rerank 3. For a reference point against higher intelligence models, GPT-6 Luna, GPT-6 Sol, and Claude Opus 5.5 were in the mix too.
These tests were conducted with windows of 60 and 200 search results. 60 is what we use in production. A window of 200 gives a model more opportunities to find useful context but also adds cost and latency. Given Jev's pricing, it's worth knowing whether expanding the initial set of results gets us better results.
R@20 (recall) was used as the main quality metric to understand what share of a query's highly relevant documents from the candidate pool make it into the first 20 results. The second metric measured was nDCG@20 (normalized discounted cumulative gain), which credits partially relevant documents and rewards putting more relevant results higher. The evaluation included 64 queries with at least one highly relevant document in the candidate pool. A small experiment, but enough for a vibe check and to understand whether we should pursue deeper testing.
An independent model (GPT-6 Astra) was used to grade each document on a relevance scale and build the final answer key.
The results
Jev not only beat every other reranker model but also came close to frontier LLM quality at a fraction of the price.
With 60 candidates, Jev improved recall@20 by 12 points. Voyage 3, the best cross-encoder, got 7 for a little less money. GPT-6 Luna was right behind Jev at 11. Sol and Opus did better by 2 to 3 points, but cost significantly more.
Jev also gets a lot more out of a bigger candidate window. Going from 60 to 200 candidates bumped its gain from 12 points to 17.
If you plot quality against cost, Jev sits right on the Pareto frontier.
Recall@20 gain over hybrid search, points
- Jev
- Frontier LLM
- Cross-encoder
- @200: reranks top 200 results
- @60: top 60
The other metric, nDCG@20, mostly backs this up. Jev scored 0.876 at 200 candidates, ahead of Luna (0.864) and well ahead of Voyage 3 (0.815). Sol and Opus led with scores around 0.92.
Cost and latency
At 1,000 searches, Jev's estimated cost is $0.80 at 60 candidates and approximately $2.19 at 200. This puts the wider run close to the cost of the cheaper rerankers at a narrower window, meaning we may be able to drive up quality at near equal cost.
Cost per 1,000 searches (log scale)
- Jev
- Frontier LLM
- Cross-encoder
- @200: reranks top 200 results
- @60: top 60
The cross-encoder rerankers were still faster though, with median response times at 0.3 to 0.7 seconds. Jev took 0.9 seconds at 60 candidates and 1.9 seconds at 200.
For an agent that's already spending several seconds on a tool call or doing work in the background, it's a small price to pay for higher quality outputs.
What's next
So far, reranking looks like a great fit for Jev! I continue to be impressed by its intelligence, speed, and cost.
These results are promising, but we still have a bit more homework to do. Before it goes into production, we need to do a deeper check of the relevance labels against human judgment, run against a larger set of queries, and run evals end to end to understand how well it holds up. Recall and nDCG only tell us each reranker found the documents the judge graded as highly relevant, but it takes a closer look to judge whether the data is actually helpful to agents.
It's also worth noting that while Jev and the frontier models reached similar recall, they weren't picking the same documents. At 200 candidates, Jev shared only 57% of its top 20 results with Opus and 58% with Sol. Roughly 8 or 9 of every 20 documents were different. Opus and Sol agreed with each other more, with 72% of the documents matching.
If you already run a reranker in production, Jev is likely worth testing on your own queries. For our customers, expect to see higher quality outputs soon!


