Glossary · term
Tool Poisoning Attack
A class of attack on MCP discovered by Invariant Labs (Luca Beurer-Kellner, Marc Fischer; April 1, 2025): malicious instructions hidden in a tool's description/metadata, which the model reads in full while the user sees only a simplified version in the UI. They hijack the agent's behavior — exfiltrating SSH keys, files, and data through call parameters.
Safety2025Wave 3 · 2025–26Maturity: 2/5
Maturity rationale
single source, early stage
References
Author: Invariant Labs