arXiv:2609.22664v1 Announce Type: new Abstract: Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or reproduces a proof of concept.

Read the full article at arXiv cs.CR →