arXiv:2609.35872v1 Announce Type: new Abstract: Safety evaluations often ask whether a model recognizes that an action is unsafe, whereas agent evaluations ask what the model chooses to do. Using safety judgments as evidence about action selection therefore raises a measurement question: does the influence of the same safety-relevant fact persist across response interfaces? We introduce SameFact, a matched-counterfactual benchmark that tests this question directly.

Read the full article at arXiv cs.CR →