arXiv:2609.26072v1 Announce Type: new Abstract: Inter-agent communication is essential to multi-agent language-model systems, yet a single message may combine task-critical information with instructions not authorized by the original request. Prompt-based defenses leave enforcement to models exposed to adversarial messages, while indiscriminate message removal discards useful information.

Read the full article at arXiv cs.CR →