Skip to main content

PR Security Review Benchmark Update: New Model Showdown

· 4 min read

Summary​

We re-ran the same PR security-review benchmark (not a full-code scan) over a new wave of models: Kimi K3, Gemini 3.6, Opus 5, GLM5.2 and more OpenAI models. Overall, GPT-5.6 Sol is still on top. Kimi K3 is the strongest open-weights performer, GLM 5.2 is cheap and precise but lags hard on recall, and Opus 5 performed well but was even more expensive than Fable. More below!

Methodology

This is a results update only. Scope, harness, scoring, and how we keep the corpus clean are unchanged. Read the original benchmark post for the full methodology before treating these numbers as a general "security model" ranking. Same caveat as that post: this measures planted access-control bugs in pull requests, not freely roaming a large existing codebase.

This chart plots Cost per Pull Request vs F2 performance score. Blue points define the cost/quality frontier.

Highlights​

Headlines from the expanded run:

  • GPT-5.6 Sol still wins. 100% (perfect) recall, F1 0.91 / F2 0.96, about $0.70 per PR. Nothing else we added knocked it off for reviewing the security status of PRs.
  • Kimi K3 is the most performant open weights model, well ahead of GLM 5.2 and sitting near mid-pack closed models, but at ~$0.94 per PR it is quite expensive.
  • GLM 5.2 via OpenRouter is cheap but recall-limited. 46% recall, 92% precision, F1 0.61. Precise when it fires, but it misses too many planted bugs for security PR review.
  • Grok 4.5 performs well (frontier at F1 0.77) but we are skeptical of the $0.20/PR price. Looks highly subsidized.
  • Anthropic Opus 5 performed worse than Fable on both performance and price.
  • Smart models do tend to more token efficient. GPT5.6 performed just as well as it's cheaper siblings and Fable outperforms Opus 5.

Results​

Since the first post we added Gemini 3.6 Flash, Claude Opus 5, the rest of the GPT-5.6 stack (Luna / Terra), Kimi K3, GLM 5.2, and refreshed a few earlier rows against the same harness. Same 10 planted access-control PRs, five runs each, Dam Secure Vulnerability Scanner with reasoning mode set to high on every model. Details live in the original methodology.

ModelRecallPrecisionF1F2Cost / PRCost / TP
GPT-5.6 Sol via OpenRouter100%83.3%0.910.96$0.70$0.70
GPT-5.5 via OpenRouter94%83.9%0.890.92$1.24$1.32
Gemini 3.6 Flash90%83.3%0.870.89$1.06$1.17
Fable 5 → Opus 4.8 fallback88%83%0.850.87~$3.61~$4.10
Claude Sonnet 4.680%85.1%0.820.81~$1.22~$1.53
GPT-5.6 Luna via OpenRouter90%73.8%0.810.86$0.76$0.84
Gemini 3.5 Flash84%76.4%0.800.82~$0.94~$1.12
Claude Opus 594%69.1%0.800.88$4.25$4.52
Kimi K3 via OpenRouter88%72.1%0.790.84$0.94$1.07
GPT-5.6 Terra via OpenRouter86%71.7%0.780.83$1.09$1.27
Grok 4.5 via OpenRouter74%80.4%0.770.75$0.20$0.27
Gemini 3.1 Flash Lite68%82.9%0.750.71~$0.04~$0.06
Claude Opus 4.860%81.1%0.690.63~$1.72~$2.87
Claude Haiku 4.556%73.7%0.640.59~$0.75~$1.34
GLM 5.2 via OpenRouter46%92.0%0.610.51~$0.13~$0.28
DeepSeek V4 Pro via OpenRouter30%65.2%0.410.34$0.21$0.71
  • True Positive (TP) - a vuln found that was expected to be found.
  • False Positive (FP) - a vuln found that was not expected (not a real bug).
  • False Negative (FN) - a vuln not found that was expected to be found.
  • Recall (R) - how well did it find the planted vulnerabilities? It is the fraction of actual positives correctly identified: TP / (TP + FN).
  • Precision (P) - false positive performance. A fraction of predicted positives that are correct: TP / (TP + FP).
  • F1 - balances precision and recall; use it when reducing false-positive noise for developers matters: 2PR / (P + R).
  • F2 - favors recall; use it when finding more true positives matters most: 5PR / (4P + R).
  • Costs / PR - is the cost to run that model over all prompts and process the PR regardless of it's accuracy. This is the most accurate representation of real-world cost.
  • Cost / TP - is the cost per true positive found.
  • Wikipedia: Precision, Recall and F1

More notes​

Again: PR review only. Full methodology is in the original article.