arXiv:2609.32601v1 Announce Type: new Abstract: Vulnerability-detection benchmarks score the verdict an agent reaches, not the evidence it gathered. A model that recalls a CVE from pretraining therefore scores the same as one that traced the data flow. We study a task where this difference matters, deciding whether a commit introduces a vulnerability.
VulContextBench: A Benchmark for Security Context Retrieval in Coding Agents
About this summary. This is a short, independently written summary of an article first published by arXiv cs.CR. Cyber Security News did not report or verify the underlying story. Read the original: https://arxiv.org/abs/2609.32601
Source attribution: headline and facts are from arXiv cs.CR (arxiv.org). Summary method: excerpt of the source description. See our source attribution policy.





