Glossary · term

Tool Poisoning Attack

A class of attack on MCP discovered by Invariant Labs (Luca Beurer-Kellner, Marc Fischer; April 1, 2025): malicious instructions hidden in a tool's description/metadata, which the model reads in full while the user sees only a simplified version in the UI. They hijack the agent's behavior — exfiltrating SSH keys, files, and data through call parameters.

Safety2025Wave 3 · 2025–26Maturity: 2/5

Maturity rationale

single source, early stage

References

Author: Invariant Labs