arXiv:2609.32021v1 Announce Type: new Abstract: Open-weight tool-calling agents are adopted on evidence of merit, usually benchmark scores and a record of reliable use. We show that a model publisher can train an agent that earns both while concealing malicious behavior.

Read the full article at arXiv cs.CR →