arXiv:2609.32601v1 Announce Type: new Abstract: Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability.

Read the full article at arXiv cs.CR →