| GLM 5.2 beats Claude in our benchmarks(semgrep.dev) | |
| 1097 points by jms703 54 days ago | 504 comments | |
tl;dr: Semgrep benchmarked open-weight and frontier models on detecting Insecure Direct Object Reference (IDOR) vulnerabilities, and found that Zhipu AI's GLM 5.2 scored 39% F1 with just a bare prompt, beating Claude Code (32%) at roughly $0.17 per vulnerability found and ~1/6 the cost. However, both trailed Semgrep's own purpose-built multimodal pipeline (53–61% F1), reinforcing that the harness around a model matters more than the model itself. Other open-weight models (MiniMax M3, Kimi K2.7) lagged significantly, so GLM 5.2 appears to be a standout rather than evidence of open weights broadly catching up. | |
HN Discussion:
| |