arXiv:2609.35870v1 Announce Type: new Abstract: Large language model agents can correctly judge that an action should be blocked while still preferring to take it. We ask why this judgment-action disconnect arises, and whether explicit safety judgment causally governs subsequent action preference. Across three open-weight language models, safety-predictive information remains recoverable from action states, arguing against a simple information-loss account.
Says Block, Still Acts: Why LLM Safety Judgments Fail to Govern Action in LLM Agents
About this summary. This is a short, independently written summary of an article first published by arXiv cs.CR. Cyber Security News did not report or verify the underlying story. Read the original: https://arxiv.org/abs/2609.35870
Source attribution: headline and facts are from arXiv cs.CR (arxiv.org). Summary method: excerpt of the source description. See our source attribution policy.

