Dev.to AI ๐Ÿค– Ai ๐Ÿ‘ 0 ๐Ÿ“– 12 min read

Jev vs Local Models: Where Japanese Intent Routing Stands After 1,000 Utterances

๐Ÿ“ Originally published (in Japanese) at forge.workstyle.tech. The router that decides "which process to assign this phrase to" is crucial for the usability of an assistant. There are methods where the model generates t

๐Ÿ“ Originally published (in Japanese) at forge.workstyle.tech.

The router that decides "which process to assign this phrase to" is crucial for the usability of an assistant. There are methods where the model generates text and classifies it every time, and others where a specialized model is used for classification. Local models make it easier to keep data on hand and allow you to set your own operational conditions.

So, what can be used for routing short Japanese utterances?

I used the classification judgment sentences from the voice assistant exista and measured 1,000 Japanese utterances using Jev and local models. I didn't test 1,000 conversations. Instead, I evaluated each utterance individually, appending only a single line explaining the previous operation for follow-up utterances. This was not a test where long conversational flows were passed to the model.

What We Compared

Exista categorizes user input into the following eight categories: browser operations, screen operations, coding, commands, tool usage, app-specific operations, casual conversation or questions, and incomplete statements. The last two categories are used to classify inputs that donโ€™t correspond to specific actions.

We compared five different setups:

Setup Execution Environment
Jev TypeSafe API
Qwen3.6-35B-A3B Custom Gateway (vLLM)
qwen3:8b This Mac (Ollama)
qwen3.5:4b This Mac (Ollama)
SemIf (Qwen3.5-4B) This Mac (MLXใƒปBF16)

Qwen3.6-35B-A3B is a Mixture of Experts (MoE) model with 3B parameters per token. The "35B" in its name refers to the total model size, but the actual computation per token is closer to that of a 4B or 8B model. SemIf uses the same Qwen3.5-4B model but is configured to directly output probabilities for the selected options.

How it was measured

The 1,000 cases created for evaluation replaced 43 cases that overlapped with previous tests, with the same composition. There were 66 cases where the correct answer could not be determined or where the answers conflicted between the roles responsible for assigning the correct answers, and these were counted separately. The remaining 934 cases were used for the final comparison. The breakdown of the 934 cases is as follows: 382 single utterances, 233 continuations of previous operations, 170 switches to different tasks, 111 post-operation chats, and 38 short and ambiguous utterances.

Seven agent entities were responsible for creating the utterances, and seven different entities were responsible for assigning the correct answers without seeing the proposals from the creators. The initial correct answer assignment was not used, as the utterances were already sorted by destination and the correct answers could be written in bulk based on the numbering. The second assignment was done by shuffling all the cases and replacing the numbers with ones that were not related to the content of the utterances.

The creators' proposals and the second correct answer assignment matched in 956 out of 957 cases. This is a high consistency rate, but it is possible that the creators and the correct answer assigners, who are of the same model type, have similar judgment tendencies. This point will be revisited later.

The explanation of the judgment text options used the same classification text as the product. For continuations of utterances, only one line, "X was operated on (what was done) N seconds ago," was added. The entire conversation history was not provided. The temperature was set to 0. For Qwen-series models, the same answer format (JSON) as the product was specified, and additional specifications were added to prevent consideration. For Jev, the destination (choice) and whether it was a save request (noul) were asked in a single call.

From Jev's measurement data, it is possible to confirm the utterance, correct answer, and judgment result, as shown in the following example. q is the utterance, expect is the correct answer, got is the model's answer, confidence is the confidence level, ok indicates whether it is correct or not, save and capabilities are the probabilities of "save request" and "asking about capabilities," respectively, ms is the response time, kind is the type of utterance, and recent is the one-line description of the previous operation added to the model.

{"q": "ไปŠๅบฆใฏ --watch ไป˜ใ‘ใฆๅŒใ˜ใฎ่ตฐใ‚‰ใ›ใฆ", "expect": "command", "got": "command",
 "confidence": 0.99, "ok": true, "save": 0.04, "saveOk": true,
 "capabilities": 0.05, "capOk": true, "ms": 181, "kind": "continue",
 "recent": "command: yarn test ใ‚’ๅฎŸ่กŒใ—ใŸ"}

(The utterance means "Now run the same thing again with --watch"; recent is the one-line context "command: ran yarn test".)

When comparing the answers of Jev and Qwen3.6-35B-A3B one by one for the 934 cases, Jev alone answered correctly in 55 cases, Qwen alone answered correctly in 10 cases, and both failed in 14 cases. The McNemar test (exact test) resulted in p โ‰ˆ 1.2ร—10โปโธ. The difference in the number of correct answers between the two models is unlikely to be due to chance alone in this evaluation set. However, this test does not guarantee the design of the evaluation set or the validity of the correct answer labels.

The utterances and correct answers for reproduction are in scripts/jev-route-k1000-cases.json, and the measurement script is in scripts/eval-jev.ts. The script is sent to the model only when --run is specified. Case specification is done with --cases=k1000, and local models are specified with --local=<Ollama model name>. To send to Exista's judgment role, specify --exista.

Results: Jev Scored 97.4% on the 934-Case Set

Evaluation Range Jev Qwen3.6-35B-A3B qwen3:8b qwen3.5:4b SemIf
All 1,000 cases 95.7% (957) 90.9% (909) 80.8% (808) 81.1% (811) 59.4% (594)
Final 934 cases 97.4% (910) 92.6% (865) 82.2% (768) 82.7% (772) 62.3% (582)

In the final 934 cases, categorized by type, there was a significant difference in the continuation judgment. Jev achieved 99.6% out of 233 cases, Qwen3.6 achieved 86.3%, and both qwen3:8b and qwen3.5:4b achieved 70.4%. Examples that involve switching to another task or engaging in small talk while considering the previous operation cannot be measured by simply classifying a single utterance.

Utterance Type (Final) Number of Cases Jev Qwen3.6-35B-A3B qwen3:8b qwen3.5:4b SemIf
Continuation 233 99.6% 86.3% 70.4% 70.4% 88.8%
Single Utterance 382 97.1% 94.0% 84.3% 83.0% 69.1%
Switching to Another Task 170 98.2% 98.8% 88.2% 92.4% 55.3%
Small Talk After Operation 111 100% 96.4% 90.1% 98.2% 15.3%
Short and Ambiguous Utterance 38 76.3% 78.9% 84.2% 65.8% 0%
Ambiguous Utterance (Separate Calculation) 66 71.2% 66.7% 60.6% 59.1% 18.2%

Overall, Jev had the most correct answers, but it did not win in all categories. In the switching category, Qwen3.6 achieved 98.8%, surpassing Jev's 98.2%. For short and ambiguous utterances, qwen3:8b achieved the highest rate of 84.2%. Looking only at the overall correct answer rate for the router can hide which utterances are incorrect.

In a separate binary judgment of "Is it a request to save?", Jev achieved 98.3%, Qwen3.6 achieved 99.4%, qwen3:8b achieved 97.1%, and qwen3.5:4b achieved 99.3%. The classification judgment and the save judgment can have different results even with the same model. Therefore, if the actual product makes multiple judgments, it is desirable to evaluate each judgment separately.

Setting Median Response Time Execution Location
Jev 170ms TypeSafe API
Qwen3.6-35B-A3B 306ms Custom Gateway (vLLM)
qwen3:8b Approximately 6.2 seconds This Mac (Ollama)
qwen3.5:4b Approximately 4.5 seconds This Mac (Ollama)
SemIf Approximately 10 seconds This Mac (MLX, BF16)

Jev and Qwen3.6 were executed over the network, while the others were executed on this Mac. Since the network, execution environment, and settings are not unified, the table should be read as the measured values in each measurement environment.

Sorting by Confidence Doesn't Always Improve Results

Jev returns confidence scores along with its predictions. In the 934 production cases this time, all 758 cases with a confidence score of 0.9 or higher were correct. For scores of 0.8 or higher, 826 out of 827 were correct, and for scores of 0.6 or higher, 881 out of 890 were correct.

Confidence Threshold Matching Cases Correct Among Them
0.9 or above 758 758 (100%)
0.8 or above 827 826 (99.9%)
0.6 or above 890 881 (99.0%)

I also tried routing only low-confidence cases to Qwen3.6, but the number of correct answers did not increase. Cutting off at 0.8 resulted in 899 correct answers, and at 0.6, 903 correct answers, both falling short of Jev's standalone 910 correct answers. This is because even in low-confidence cases, Jev outperformed Qwen3.6.

Using confidence scores for fallback seems like a reasonable design. However, routing to another model doesn't guarantee improvement. It's necessary to measure, for each threshold, how many cases are handled, how many are correct, and how many are incorrect, using your own evaluation set.

Insights from SemIf

We analyzed whether the differences in SemIf could be explained solely by the arrangement of options and the format of descriptions across 299 cases. These 299 cases were extracted from the 899 production-level cases out of the initial 957 evaluations, maintaining the ratio of speech types.

Settings in 299 cases Total Continuation (77) Switch (54) Chit-chat (35) Ambiguous (11) Standalone (122)
Jev 98.0% 77 52 35 10 119
Qwen3.6-35B-A3B 92.0% 67 53 33 8 114
qwen3.5:4b (text response) 84.6% 54 49 34 7 109
SemIf A: Original input 62.2% 71 27 4 0 84
SemIf B: Shuffle option order 67.2% 72 34 11 0 84
SemIf C: B + single-line description 67.9% 62 32 26 2 81

Moving from A to B resulted in a 5-point increase. There was a bias in the order of options, making later options less likely to be chosen. However, even after shuffling the order, the result was 67.2%. Further simplifying the description to a single line in C yielded 67.9%. Despite input adjustments, the result fell short of the 84.6% achieved by the same model with text responses, lagging by approximately 17 points.

Longer descriptions also strengthened the tendency to select the same answer as the previous one. The same answer as the previous one was chosen in 124 out of 177 cases for B and 87 out of 177 cases for C, yet the overall accuracy remained nearly unchanged. At least in this test, simply standardizing the judgment format did not replicate Jev's results. Jev's strength likely lies not only in its usage but also in the model itself.

How to Utilize the Results of This Study

The judgment of continuity has practical significance in voice assistants and chat-type agents. When a user says "do the same thing again" or "go down a bit more," the system can redirect to the same operation based on the previous operation. Alternatively, when a user says "remember this" during a task, the system can determine whether to treat it as a continuation of the previous operation or route it to a different process for saving.

In this study, the results for 233 consecutive utterances were as follows: Jev achieved 99.6%, Qwen3.6 achieved 86.3%, and the standard 4B and 8B models achieved 70.4%. Although this result is limited to the data provided with short preceding contexts, it provides a reason to include consecutive utterances in the evaluation set. Evaluations that only collect single utterances may overlook important product failures.

The median response time for Jev was 170ms. While the response time cannot be directly compared to local models due to environmental differences, it provides material to consider using the model to return routing judgments with short waiting times. We also evaluated "save request" Yes/No judgments, such as ใ€Œไฟๅญ˜ใฎไพ้ ผใ‹ใ€ ("is it a save request?"). However, Qwen3.6 and qwen3.5:4b performed slightly better in this task, indicating that the best option may vary depending on the judgment task.

Judgment-Specific Models and General-Purpose Models

Jev is a judgment-specific model from TypeSafe that does not generate text, but instead returns a fixed judgment. Choice can return up to 255 options, Score can return 2-10 levels, and Noul can return a Yes probability of 0-1. The input is text only, and the weights are not publicly available. To use it, you need to go through TypeSafe's API.

On the other hand, frontier models like Claude and GPT are general-purpose and can be used not only for classification but also for generating text and various other tasks. The wide range of applications is a major advantage, but the cost is not only based on the input, but also on the output. The configuration changes depending on whether you use it only for judgment or also for generating responses after judgment.

Here, we compared Jev with local models, etc., and did not measure the accuracy of frontier models. Therefore, we cannot conclude that "Jev is more accurate than Claude". To compare accuracy, we need to prepare the same utterance, the same context, the same correct answer, and the same output conditions.

According to the publicly available specifications of Jev, the response time is 70-500ms, the input cost is $0.042 per 1 million tokens, and the output is free. The median response time of 170ms observed in the experiment is the value for this API usage, and it is not a guarantee of the same speed in all environments.

Model Input (per 1 million tokens) Output (per 1 million tokens)
Jev $0.042 Free
Claude Haiku 4.5 $1 $5
Claude Sonnet 5 $2 $10
Claude Opus 5.5 $4 $20

The prices of Claude in the price table are the official prices published by Anthropic at the time of writing. Jev is designed as a judgment-specific model with free output, but the cost changes depending on the number of input tokens and the actual request content. For example, if the input is 1,000 tokens for one judgment, the input cost of Jev would be 1,000 รท 1,000,000 ร— $0.042, which is approximately $0.000042. This is a calculation example assuming that token number, and it does not show the actual cost of each request.

There are also English evaluations in the public benchmarks. In LangWatch's test of 15 types and approximately 10,000 cases, Jev was 12-68 points higher than smaller public models. SemIf showed that the consistency rate with Jev was 84.5% in TypeSafe's public 102 cases, and it was about 5 times faster than generating JSON. Both of these are different from the Japanese evaluation in this time, and they are not direct evidence of the Japanese accuracy.

If You Evaluate This Yourself

Start by collecting hundreds of real-world utterances from your product and labeling them with ground truth. Assign different people to the data collection and labeling tasks, and shuffle the entries to prevent guessing the labels based on IDs or ordering. Don't force ambiguous examples into a single correct answer; instead, isolate them and report them separately.

Count standalone requests, continuations of previous actions, task switches, and post-action chit-chat separately. Analyzing error rates by categoryโ€”not just the overall averageโ€”makes it easier to decide which model is best suited for specific use cases. Binary decisions, such as confirming saves or answering feature-related questions, should also be measured as a distinct metric separate from general routing accuracy.

If using a setup where low-confidence cases are routed to a different model, validate this pipeline with real data before deployment. As seen in this case, the fallback model is not necessarily more accurate on low-confidence examples. Review the number of correct predictions and the total number of items delegated for each confidence threshold before moving to production.

If you can't use Jev, in our in-house setup, Qwen3.6-35B-A3B achieved 92.6% accuracy across 934 production cases. Standard 4B and 8B models performed in the high 80s (~82%), with continuation detection particularly struggling at 70.4%. Please note these results are specific to our dataset and do not generalize to all Japanese-language tasks. If you prioritize local deployment, consider MoE models in the 35B-A3B class and verify whether they meet your performance criteria for each utterance type.

Limitations

This time, the correct answer was given by Claude. Although the roles of creating and giving correct answers were separated, they are the same type of model, so there is a possibility that the judgment tendencies are consistent. The consistency rate of 99.8% may be too high, and verification with correct answers given by humans has not been done yet.

The utterances used were created ones, not those of actual end-users. The context is only the previous line, and the ability to read the flow of long conversations is not being measured. If the distribution changes with actual operational data, the results may also change.

SemIf's instruction sentences are in English format, with only the content translated into Japanese (ๆ—ฅๆœฌ่ชž) (Japanese language). If the instruction sentences are translated into Japanese, the results may improve. Additionally, SemIf's 27B version is said to have been 96% accurate in English, but it is for NVIDIA's GPU, so it was not tried in this environment.

The conditions for response time are also not consistent. Jev and Qwen3.6 are via communication, while other models are running on this Mac's Apple Silicon. The speed in the table cannot be treated as a win or loss under the same conditions.

Finally, Jev cannot be run manually. The weights are not publicly available, and it can only be used through TypeSafe's API. The accuracy and speed this time may not be reproducible as is if the model on the API side is updated or the usage environment changes.

A companion essay on the same experiment, written from the builder's side, is on note (in Japanese): here.

๐Ÿ“ฐ Read the original article on Dev.to AI

Originally published by Dev.to AI. Aggregated on AIWithGhost for educational purposes โ€” full credit and traffic to the original publisher.