MCP agents are vulnerable to semantic supply-chain attacks where adversaries optimize tool descriptions and outputs to hijack agent behavior—a risk that transfers across different AI models without retraining.
This paper reveals a critical vulnerability in AI agents using the Model Context Protocol (MCP), where attackers can hijack agents by crafting malicious tool metadata and outputs. The A2M framework demonstrates how two-stage optimization can trick agents into invoking attacker-controlled tools with 93.6% success rate, potentially causing denial of service, data theft, or reasoning failures.