arXiv:2609.35872v1 Announce Type: new Abstract: Safety evaluations often ask whether a model recognizes that an action is unsafe, whereas agent evaluations ask what the model chooses to do. Using safety judgments as evidence about action selection therefore raises a measurement question: does the influence of the same safety-relevant fact persist across response interfaces? We introduce SameFact, a matched-counterfactual benchmark that tests this question directly.
SameFact: The Same Safety Facts Lead to Different Responses Across Interfaces
About this summary. This is a short, independently written summary of an article first published by arXiv cs.CR. Cyber Security News did not report or verify the underlying story. Read the original: https://arxiv.org/abs/2609.35872
Source attribution: headline and facts are from arXiv cs.CR (arxiv.org). Summary method: excerpt of the source description. See our source attribution policy.

