arXiv:2609.35870v1 Announce Type: new Abstract: Large language model agents can correctly judge that an action should be blocked while still preferring to take it. We ask why this judgment-action disconnect arises, and whether explicit safety judgment causally governs subsequent action preference. Across three open-weight language models, safety-predictive information remains recoverable from action states, arguing against a simple information-loss account.

Read the full article at arXiv cs.CR →