Technology
Where Jev probably fits: between LLMs and purpose-trained classifiers
Harry Yu & Wang Tao

Level3AI has built on the Workflow Agent architecture for more than two years. A production conversation is not one model call. It is a pipeline of stages, and classifications decide the transitions between them. Is the message relevant? Does the customer want to speak to a human? Is it urgent? Is the conversation over? A typical deployment has dozens of these decisions, each with its own label set.
When these classifiers run at high volume in production, they set the agent's cost, latency, and reliability.
The LLM default for classification
An LLM is the easiest classifier to plug in: write a prompt, list the labels, and parse the answer. But LLMs aren't built specifically for classification. They decode the request token by token, and they're priced for output generation. In addition, LLMs are not deterministic. The same input, sent ten times, doesn't necessarily return the same label. When a label determines how a message is routed, a flip can mean a different outcome for the customer.
For high-volume CX tasks, we built purpose-trained classifiers instead. They're faster, cheaper, and reproducible. However, each one needs a labelled dataset, a training pipeline, and an owner, which makes them expensive to build and maintain.
Some tasks are hard to train for at all. When the AgentOps team defines its own labels, there's no shared dataset, and the labels can change whenever the AgentOps team updates its taxonomy.
What is Jev?
TypeSafe announced Jev in early access on September 15, 2026. It's not a language model and doesn't generate text. Instead, it ingests your input once and evaluates multiple typed questions against it in parallel. Each question returns a typed value with a calibrated confidence attached: yes/no, a choice from a set, or a score on a scale.
TypeSafe claims "zero hallucinations" because Jev only returns approved answers. But an approved answer is not always the same answer to the same input. TypeSafe also describes Jev's intelligence-per-dollar as "off the charts," claiming it outperforms any other model on the market.
To test these claims in a real CX use case, we benchmarked Jev not only against GPT-5.6 Terra, but also against one of Level3AI's own purpose-trained classifiers.
Comparison of Jev, LLM and purpose-trained classifier
We used one classifier from our production pipeline and 2,000 testing samples, sending each sample ten times to each system.[1] These are early numbers, not a verdict. Accuracy is the mean across the ten rounds.[2] Inconsistent samples are those that received more than one label across the ten runs. Cost is one full pass over all 2,000 samples.[3]

Accuracy. GPT-5.6 leads. Our 0.6B model is 0.37 points behind it, and Jev is 1.38 points behind. The differences are within the margin of error given the dataset size.
Consistency. Our 0.6B model returned the same label on all ten runs for every sample. The zero is exact, not rounded. GPT-5.6 flipped on 14 samples and Jev on 19. In this test, returning a typed answer did not make Jev more reproducible than the LLM.

Cost. Jev is 13x cheaper than GPT-5.6, and our 0.6B model is 60x cheaper. GPT-5.6’s cost assumes a 75% cache hit rate; the same pass costs $6.28 with no caching. Jev’s cost and ours stay the same at any hit rate.
Latency. Jev is about 3x faster than GPT-5.6 at p50, and its tail is tighter. Jev’s p90 is only 53 ms above its p50, compared with 194 ms for GPT-5.6, and its p99 is 556 ms against 1,482 ms. For a stage that runs before every reply, the tail is what the customer feels. Our 0.6B model’s p99 is 166 ms.

Where Jev fits
Jev does not replace purpose-trained classifiers. Where we have invested in training, the purpose-trained classifier is faster, cheaper, and more consistent than the other systems in the table.
Jev fits two kinds of tasks.
The long tail. Some tasks are still running on LLMs because no one's had time to train a model for them. Moving a task like this to Jev cuts cost by an order of magnitude and p99 latency by nearly two-thirds. For low-stakes classifications, that tradeoff is fine. For the rest, Jev can serve as an interim solution until a purpose-trained classifier is developed.
User-defined labels. Jev takes the label set as part of the question, so when a customer defines or changes their taxonomy, the classifier follows immediately, with no dataset or retraining. Here, a purpose-trained classifier is slow to build and impractical to maintain. LLM will still be the preferred option.
Jev is not the best of both worlds. It probably sits in the middle, between LLMs and purpose-trained classifiers. For user-defined labels, the middle may be where it stays.
Jev is not open source, but its architecture is well documented, and we are building our own implementation. We will run both: Jev where it fits today, and our own implementation where we want Jev-style typed decisions on our infrastructure, trained on our data.
We will publish what we learn as we move more classification tasks off LLMs and continue to build out the Workflow Agent architecture for CX.
References
[1] GPT-5.6 Terra was run at temperature 0. All 14 of its inconsistent samples occurred at that setting.
[2] We used a relatively balanced test set, so the default threshold applies and accuracy is a meaningful metric. In production, most classification tasks are heavily skewed. For those, our usual architecture is a cascade: a small classifier optimized for recall, followed by a more powerful LLM optimized for precision. Jev fits the first leg: its calibrated probabilities let us lower the threshold for recall without retraining.
[3] Hosting cost for our 0.6B model is extrapolated from GPU costs on GCP.








