Opus 5.5 Leads a Four-Model OWASP Java Test by 7 to 9 Points. Kimi K3, GLM 5.3, and GPT-6 Sol Tie.
OWASP Benchmark for Java is the OWASP Foundation's open test suite for scoring vulnerability-detection accuracy. Version 1.2 ships 2,740 deliberately exploitable test cases across eleven weakness classes β SQL injection,
OWASP Benchmark for Java is the OWASP Foundation's open test suite for scoring vulnerability-detection accuracy. Version 1.2 ships 2,740 deliberately exploitable test cases across eleven weakness classes β SQL injection, command injection, cross-site scripting, path traversal, weak cryptography, and more β each one a fully runnable Java servlet labeled vulnerable or safe against a public answer key. It was built to score static and dynamic security scanners the same way on every run, on two numbers: whether a tool finds the real vulnerabilities, and whether it also wrongly flags the safe code next to them. That second number β false positives β is where three of the four models in this test lose most of their points.
At maximum reasoning effort, Claude Opus 5.5 classified 654 of 660 OWASP Benchmark Java files correctly and missed no vulnerable file. Kimi K3, GPT-6 Sol, and GLM 5.3 scored 607, 598, and 597 β a 1.5-point spread that a paired bootstrap can't call a winner on. They separate clearly on cost and speed instead.
I ran the same 660 files through all four models on Amazon Bedrock with the same settings. The question was what two open-weight models give up against current closed models on a single-call code-review task once each runs at its top reasoning setting, and what each correct answer costs.
TL;DR
- Opus 5.5 got 654 of 660 right with zero missed flaws and 5 false alarms, while the other three raised 38 to 41 false alarms each. Its lead over each of them is 7 to 9 points, and the 95% interval stays above zero in all three comparisons.
- Kimi K3, GPT-6 Sol, and GLM 5.3 land within 1.5 points of each other, and every paired interval among them includes zero.
- GLM 5.3 charges 56% of Kimi K3's input rate and 35% of its output rate, yet it writes 3.4 times as many output tokens, so cost per correct answer comes out at $0.0091 against $0.0094.
- Opus 5.5 costs 3.1 to 3.4 times more per correct answer than the other three and averages 36 seconds per call against 7.5 to 10.7.
What the model is asked
Each call is one message with no tools. It names a weakness class, such as command injection or weak hashing, and attaches one servlet file from OWASP Benchmark for Java with line numbers. The model returns one JSON object: vulnerable true or false, an evidence line, and a one-sentence reason.
The answer key is a separate public file that never reaches the model. I score only the true or false verdict. A refusal, a truncated reply, or an unparseable reply counts as a miss and gets reported on its own line.
The 660 files are 330 vulnerable and 330 safe, spread across eleven weakness classes, drawn deterministically from a pinned commit of the benchmark. None of the 660 repeat across my earlier OWASP runs.
Every model ran with the same settings:
- the global inference profile, called through Bedrock Runtime
- reasoning effort set to
max - a 32,768-token output cap, which no call reached
- one prompt wording that asks whether the code is vulnerable to the named class
- one repeat per file
Claude went through the Messages API, GPT-6 Sol through Responses, and Kimi K3 and GLM 5.3 through Chat Completions. I covered why each family needs its own API in an earlier post on wiring OpenCode to Bedrock. The model cards for GLM 5.3, GPT-6 Sol, and Opus 5.5 list the supported effort levels.
The result
| Opus 5.5 | Kimi K3 | GPT-6 Sol | GLM 5.3 | |
|---|---|---|---|---|
| Correct | 654 (99.1%) | 607 (92.0%) | 598 (90.6%) | 597 (90.5%) |
| Missed flaws | 0 | 11 | 24 | 21 |
| False alarms | 5 | 41 | 38 | 39 |
| Non-answers | 1 refusal | 1 | 0 | 3 |
| Mean output tokens | 1,192 | 364 | 707 | 1,243 |
| Mean seconds per call | 36.1 | 10.7 | 7.5 | 10.3 |
| List cost, 660 calls | $19.99 | $5.73 | $5.96 | $5.40 |
| Cost per correct answer | $0.0306 | $0.0094 | $0.0100 | $0.0091 |
I compared each pair of models on the same 660 files with a paired bootstrap over files. Opus 5.5 beat Kimi K3 by 7.1 points (interval +5.3 to +9.2), GLM 5.3 by 8.6 (+6.5 to +10.9), and GPT-6 Sol by 8.5 (+6.4 to +10.6). Among the other three, the largest gap was 1.5 points, and all three intervals straddled zero.
The paired counts show the shape of the lead. Opus 5.5 was right on 47 files Kimi K3 got wrong, 57 that GLM 5.3 got wrong, and 56 that GPT-6 Sol got wrong. None of the three was right on a file Opus 5.5 got wrong.
Where the gap lives
Weak hashing and weak cryptography carry the largest share. Opus 5.5 scored 70 of 70 on both classes, while the other three scored between 55 and 63 on each. Command injection shows the same direction, with Opus 5.5 at 66 of 70 against 58 to 62.
Ten of the eleven classes put Opus 5.5 at or within one file of a perfect score, and command injection (66 of 70) is the exception. GLM 5.3 matched it at 70 of 70 on weak randomness, and all four models tied at the top on secure cookies and XPath injection. Each class rests on 20 to 70 files, so I treat single-class gaps as leads to follow and not as findings.
The false-alarm column matters in practice. A reviewer who flags 40 safe files out of 330 produces a triage queue, and Opus 5.5 produced 5. The Semgrep team reported that Kimi K3 had lower precision than other models on IDOR detection, 0.684 against 0.84 to 0.91. That was a different task with an agentic harness. In my test Kimi K3's 41 false alarms sit next to GLM 5.3's 39 and GPT-6 Sol's 38, so I see no Kimi-specific precision problem on this task. The outlier is Opus 5.5.
What the open-weight pair buys
Kimi K3 and GLM 5.3 each cost about $5.50 for the whole 660-file run, against $20 for Opus 5.5. That buys roughly 7 to 9 points less accuracy and a 3 to 5 times faster call.
Token counts are exact β I logged every call's usage block directly, and the mean, median, and 95th-percentile figures above come straight from those logs. The price per token checks out against each vendor's own published rate. Anthropic's pricing lists Claude Opus 5.5 at $4 in / $20 out per million tokens on the direct API, and states that US-only inference carries a separate 1.1x surcharge β so the $4/$20 global rate used here is Anthropic's own direct-API price, not an estimate. OpenAI's pricing lists GPT-6 Sol at $2 in / $10 out, and matches the same 1.1x pattern: this repository's own cost tracking found Bedrock's US-Geo rate for GPT-6 Sol at exactly $2.20/$11.00, 1.1x the direct rate. Kimi K3's Bedrock global rate ($3.00 in / $15.00 out) matches Moonshot's own API price exactly, confirmed against the live AWS Price List API. GLM 5.3 is the one real exception: Bedrock's global rate ($1.68 in / $5.28 out) runs 20% above Z.ai's own direct rate of $1.40 in / $4.40 out β a genuine markup, not a reading gap. All four rates in the table are grounded against a vendor or AWS source; none are reconstructed.
GLM 5.3's per-token advantage mostly disappears. Its mean output of 1,243 tokens is 3.4 times Kimi K3's 364. At the rates above, GLM 5.3's output alone costs about $0.0066 per call and Kimi K3's about $0.0055, so GLM 5.3 spends more on output despite the lower rate. Cheaper input recovers most of the difference, and the net gap is 4% per correct answer β small enough that extra reasoning tokens, not price, decide which one is cheaper in practice.
GPT-6 Sol sits with the open-weight pair on both accuracy and cost, and it was the fastest model in the run.
What the effort setting did
Bedrock accepted six effort values on both open-weight models in my probes: none, low, medium, high, xhigh, and max. It rejected minimal, and none produced zero reasoning tokens. Moonshot's own documentation lists only low, high, and max for Kimi K3, so Bedrock exposes a wider surface than the vendor describes.
On three OWASP files, GLM 5.3's reasoning tokens did not rise steadily with the label: high used none on two of them, and low used about 2,700 tokens on the weak-hash file. A sheep-counting puzzle did rise steadily, from 29 tokens at low to 240 at max. I ran only max in the full test, so I can say nothing about cost or accuracy at lower settings. A team picking an effort level for production would want exactly that curve, and this test does not provide it.
What this test cannot show
- Memorization. OWASP Benchmark and its answer key have been public for years. Opus 5.5's near-perfect score could reflect recall of the benchmark, and I cannot tell how much. This is my inference, and a post-cutoff set is the test that would settle it.
- Scope. One benchmark, Java only, single files, with the weakness class named in the prompt. Whole-repository review and agentic harnesses are different tasks.
- One repeat. My earlier OWASP runs showed wrong answers repeating across runs, so extra repeats add little. Gaps under about 4 points between two models are within the noise.
- Latency. Approximate, because the runs shared a machine and overlapped in time.
- Opus 5.5's refusal. One file returned a refusal with only a thinking block. I counted it as a miss, and Opus 5.5 scores 654 of 660 either way.
Other groups report results for these models on different tasks. Aikido's agentic fresh-CVE test found 25 of 32 for both Kimi K3 and GLM 5.3, which is a tie like mine. Simbian's defensive hunting benchmark put GLM 5.2 slightly ahead of Kimi K3. Z.ai's model card reports a CyberGym lead for GLM 5.3 that I did not test. I found no published OWASP Benchmark result for Kimi K3 or GLM 5.3.
So what
On single-file Java review at maximum effort, I would pick by the failure I can afford. Opus 5.5 gives up no vulnerable files and few false alarms, and charges about three times as much per correct answer. The open-weight pair and GPT-6 Sol are interchangeable on accuracy here, so price per correct answer, latency, and data-handling needs decide between them.
The result also sets up a specific question. Opus 5.5 was never wrong on a file where Kimi K3, GLM 5.3, or GPT-6 Sol was right β every file it missed, they missed too. That asymmetry is the precondition for a model cascade: run a cheap model on every file, and escalate to Opus 5.5 only on the ones it flags as uncertain. FrugalGPT named this pattern in 2023, and the broader cascade literature treats cheap-first, escalate-on-low-confidence as the standard cost play for exactly this kind of accuracy asymmetry. The same literature carries a specific warning that applies here: a 2026 measurement of cascade reliability found that a cheap verifier's blind spot β the fraction of wrong answers it waves through unescalated β grows with the weaker model's capability, and a frontier verifier closes that gap only by escalating on nearly half of all traffic, which erases most of the savings the cascade exists to produce. I have not built that cascade on this benchmark. Whether it lands near the 3x savings a clean cascade promises, or the near-zero savings the blind-spot result warns about, is the open thread β along with whether Opus 5.5's lead holds on code that was never public, which this test also cannot answer.
Originally published by Dev.to Security. Aggregated on AIWithGhost for educational purposes β full credit and traffic to the original publisher.