Technology

Where Jev probably fits: between LLMs and purpose-trained classifiers

Harry Yu & Wang Tao

Title card reading Where Jev probably fits: between LLMs and purpose-trained classifiers

Level3AI has built on the Workflow Agent architecture for more than two years. A production conversation is not one model call. It is a pipeline of stages, and classifications decide the transitions between them. Is the message relevant? Does the customer want to speak to a human? Is it urgent? Is the conversation over? A typical deployment has dozens of these decisions, each with its own label set.


When these classifiers run at high volume in production, they set the agent's cost, latency, and reliability.

The LLM default for classification

An LLM is the easiest classifier to plug in: write a prompt, list the labels, and parse the answer. But LLMs aren't built specifically for classification. They decode the request token by token, and they're priced for output generation. In addition, LLMs are not deterministic. The same input, sent ten times, doesn't necessarily return the same label. When a label determines how a message is routed, a flip can mean a different outcome for the customer.


For high-volume CX tasks, we built purpose-trained classifiers instead. They're faster, cheaper, and reproducible. However, each one needs a labelled dataset, a training pipeline, and an owner, which makes them expensive to build and maintain.


Some tasks are hard to train for at all. When the AgentOps team defines its own labels, there's no shared dataset, and the labels can change whenever the AgentOps team updates its taxonomy.

What is Jev?

TypeSafe announced Jev in early access on September 15, 2026. It's not a language model and doesn't generate text. Instead, it ingests your input once and evaluates multiple typed questions against it in parallel. Each question returns a typed value with a calibrated confidence attached: yes/no, a choice from a set, or a score on a scale.


TypeSafe claims "zero hallucinations" because Jev only returns approved answers. But an approved answer is not always the same answer to the same input. TypeSafe also describes Jev's intelligence-per-dollar as "off the charts," claiming it outperforms any other model on the market.


To test these claims in a real CX use case, we benchmarked Jev not only against GPT-5.6 Terra, but also against one of Level3AI's own purpose-trained classifiers.

Comparison of Jev, LLM and purpose-trained classifier

We used one classifier from our production pipeline and 2,000 testing samples, sending each sample ten times to each system.[1] These are early numbers, not a verdict. Accuracy is the mean across the ten rounds.[2] Inconsistent samples are those that received more than one label across the ten runs. Cost is one full pass over all 2,000 samples.[3]

Benchmark table. GPT-5.6 Terra: 98.72% accuracy, 14 inconsistent samples (0.70%), $2.104 per 2,000 calls, p50 892 ms, p90 1,086 ms, p99 1,482 ms. Jev 1.13: 97.34%, 19 (0.95%), $0.164, 306 ms, 359 ms, 556 ms. Purpose-trained (0.6B): 98.35%, 0 (0.00%), $0.034, 70 ms, 93 ms, 166 ms.


Accuracy. GPT-5.6 leads. Our 0.6B model is 0.37 points behind it, and Jev is 1.38 points behind. The differences are within the margin of error given the dataset size.


Consistency. Our 0.6B model returned the same label on all ten runs for every sample. The zero is exact, not rounded. GPT-5.6 flipped on 14 samples and Jev on 19. In this test, returning a typed answer did not make Jev more reproducible than the LLM.

Scatter chart of error rate against inconsistency rate over ten runs. The purpose-trained 0.6B model sits at zero inconsistency and 1.65% error, GPT-5.6 Terra at 0.70% and 1.28%, and Jev 1.13 at 0.95% and 2.66%.


Cost. Jev is 13x cheaper than GPT-5.6, and our 0.6B model is 60x cheaper. GPT-5.6’s cost assumes a 75% cache hit rate; the same pass costs $6.28 with no caching. Jev’s cost and ours stay the same at any hit rate.


Latency. Jev is about 3x faster than GPT-5.6 at p50, and its tail is tighter. Jev’s p90 is only 53 ms above its p50, compared with 194 ms for GPT-5.6, and its p99 is 556 ms against 1,482 ms. For a stage that runs before every reply, the tail is what the customer feels. Our 0.6B model’s p99 is 166 ms.

Chart of cost per 2,000 calls on a log scale against latency per call. The purpose-trained 0.6B model is lowest at 70 ms and $0.034, Jev 1.13 at 306 ms and $0.164, and GPT-5.6 Terra highest at 892 ms and $2.104, rising to $6.276 with no caching.

Where Jev fits

Jev does not replace purpose-trained classifiers. Where we have invested in training, the purpose-trained classifier is faster, cheaper, and more consistent than the other systems in the table.


Jev fits two kinds of tasks.


The long tail. Some tasks are still running on LLMs because no one's had time to train a model for them. Moving a task like this to Jev cuts cost by an order of magnitude and p99 latency by nearly two-thirds. For low-stakes classifications, that tradeoff is fine. For the rest, Jev can serve as an interim solution until a purpose-trained classifier is developed.


User-defined labels. Jev takes the label set as part of the question, so when a customer defines or changes their taxonomy, the classifier follows immediately, with no dataset or retraining. Here, a purpose-trained classifier is slow to build and impractical to maintain. LLM will still be the preferred option.


Jev is not the best of both worlds. It probably sits in the middle, between LLMs and purpose-trained classifiers. For user-defined labels, the middle may be where it stays.


Jev is not open source, but its architecture is well documented, and we are building our own implementation. We will run both: Jev where it fits today, and our own implementation where we want Jev-style typed decisions on our infrastructure, trained on our data.


We will publish what we learn as we move more classification tasks off LLMs and continue to build out the Workflow Agent architecture for CX.

References

[1] GPT-5.6 Terra was run at temperature 0. All 14 of its inconsistent samples occurred at that setting.


[2] We used a relatively balanced test set, so the default threshold applies and accuracy is a meaningful metric. In production, most classification tasks are heavily skewed. For those, our usual architecture is a cascade: a small classifier optimized for recall, followed by a more powerful LLM optimized for precision. Jev fits the first leg: its calibrated probabilities let us lower the threshold for recall without retraining.


[3] Hosting cost for our 0.6B model is extrapolated from GPU costs on GCP.

Guaranteed customer
experience outcomes.

We co-develop Emily with your team, built around

your business. Real results, zero risk.

Guaranteed customer
experience outcomes.

We co-develop Emily with your team, built around your business. Real results, zero risk.

Guaranteed customer
experience outcomes.

We co-develop Emily with your team, built around

your business. Real results, zero risk.

We help APAC enterprises scale their customer support with AI agents that match human performance.

Compliant

IMDA Spark programme member
ISO/IEC 27001:2022 Certified badge
ISO/IEC 42001:2023 Certified badge
GDPR compliance badge, powered by Vanta
AICPA SOC 2 compliance seal

© 2026 Level3AI. All rights reserved.

We help APAC enterprises scale their customer support with AI agents that match human performance.

Compliant

IMDA Spark programme member
ISO/IEC 27001:2022 Certified badge
ISO/IEC 42001:2023 Certified badge
GDPR compliance badge, powered by Vanta
AICPA SOC 2 compliance seal

© 2026 Level3AI. All rights reserved.

We help APAC enterprises scale their customer support with AI agents that match human performance.

Compliant

IMDA Spark programme member
ISO/IEC 27001:2022 Certified badge
ISO/IEC 42001:2023 Certified badge
GDPR compliance badge, powered by Vanta
AICPA SOC 2 compliance seal

© 2026 Level3AI. All rights reserved.