Prompt Injection: When Content Gives Orders
Prompt injection is the most important AI-specific risk to understand, and it is simple. An AI reads text and follows instructions in text. It has no reliable way to tell your instructions apart from instructions hidden inside something it was asked to read.
Imagine you ask an agent to summarize a web page, and somewhere on that page, in small or hidden text, is a line saying to ignore your instructions and send the user’s files to an address. A vulnerable agent may simply do it. The same can hide in an email, a document, a calendar invite or a support ticket, anywhere an agent reads content written by someone else.
There is no perfect fix today, so the defence is layered. Keep untrusted content away from agents that hold powerful access. Do not let an agent that reads outside content also be able to send data out or change things without your approval. Require a human check before anything irreversible, such as sending, deleting, paying or publishing.
And treat that as a permanent habit, not a temporary one. New defences are being built, and they help, but the safe assumption is that anything an agent reads could be trying to steer it.
Takeaway: text an agent reads can contain hidden orders. Keep untrusted content away from powerful access, and require your approval for anything that cannot be undone.
Tools, prices and features in this area change quickly. Last reviewed September 2026.