Skip to main content

2 posts tagged with "Benchmark"

Evaluations, benchmarks, and measured comparisons of security tooling and models

View All Tags

PR Security Review Benchmark Update: New Model Showdown

· 4 min read

Summary

We re-ran the same PR security-review benchmark (not a full-code scan) over a new wave of models: Kimi K3, Gemini 3.6, Opus 5, GLM5.2 and more OpenAI models. Overall, GPT-5.6 Sol is still on top. Kimi K3 is the strongest open-weights performer, GLM 5.2 is cheap and precise but lags hard on recall, and Opus 5 performed well but was even more expensive than Fable. More below!