NXT

Explainer

What is prompt injection, and will it work on your agent?

Prompt injection is when text an AI agent reads contains instructions meant for the agent, and the agent follows them. The text can come from anyone who can get a document, an email, or a form in front of it.

NXT AI research team · Updated October 7, 2026

01

Why it works

A language model has one input: text. It does not have a separate channel for “instructions from the people who built me” and “content I am supposed to read.” Everything arrives in the same stream, and the model decides what to do with it. If a sentence in a customer's email reads like an instruction, the model can treat it as one.

That is why the problem has resisted fixes. OpenAI put it plainly in December 2025: prompt injection, “much like scams and social engineering on the web, is unlikely to ever be fully ‘solved.’” The OWASP project that catalogs AI application risks lists it first.

A chatbot that only talks has little to lose. An agent that reads your inbox, opens your files, and can send email or update records is a different matter. The instruction in the document becomes an action in your systems. Two publicly disclosed flaws in 2025 and 2026, one in Microsoft 365 Copilot and one in Copilot Studio, were exactly this: content the assistant read, turned into data leaving the organization.

02

Four ways it reaches an agent

Direct

The person typing to the agent tells it to ignore its instructions. The oldest form, the easiest to spot, and the least dangerous, because the attacker is already in the conversation.

Through documents

A PDF, spreadsheet, or image the agent is asked to process carries text written for the agent. White-on-white text, a memo field, a comment in a file. The agent reads it as part of the job.

Through email and forms

A customer email, a support ticket, a lead form. Anyone on the internet can write to these, and an agent that summarizes or acts on them is reading instructions from strangers.

Through tool results

The agent calls a search, a database, or another agent, and the result contains instructions. Nothing a human sees, so nothing a human catches.

03

What we found when we measured it

Our research looked at the close cousin of injection: a caller who states customer details and asks the agent to confirm them, with no proof of who they are. No hidden text, just a request framed to look like verification. Seven of eight frontier models from two providers disclosed customer data in at least one tested condition. Verification instructions in the system prompt did not prevent it.

The follow-up study traced why. The model almost never asks itself whether confirming a detail is the same as disclosing it. When made to take a position on that question first, it refused in every one of 152 trials. The failure is not bad reasoning. It is reasoning that never gets triggered, which is the same mechanism that lets an instruction in a document pass as a legitimate task.

Both studies name the models, publish the methodology, and are on the Research page.

04

What does not fix it

Telling the agent not to

Instructions in the system prompt are the first thing a team tries. In our testing, three instruction-level defenses stopped disclosure in all 120 trials, and one of them also refused 8 of 10 legitimate requests. The defense worked by refusing the customer.

Giving the model more time to think

More reasoning effort did not reliably help. In one model it made disclosure more explicit, because the extra reasoning elaborated the decision instead of re-examining it.

Trusting the platform to handle it

Platforms protect the data on its way to the model. They do not decide what your agent can read, what it can do, or what it does with text that arrived from outside. Those are configuration, and configuration is what gets tested.

05

Will it work on your agent?

Three questions decide it. What does the agent read that someone outside your company could have written? What can it do once it has read it: disclose a record, send an email, change a payment detail, call a tool? And what stands between the two other than the model's own judgment: an authentication step, an approval, a scope limit on the action?

If the answer to the first two is “a lot” and the answer to the third is “the system prompt,” then yes, it will work. In our testing, what decided the outcome was not how the model was instructed but whether the application itself required authorization before a record could be touched. Where the model was the only barrier, the records were disclosed. Where the application required an authenticated session, disclosure fell to zero.

Find out before someone else does.

For the confirmation class, our open-source tool runs the check against your own agent with synthetic records. For the full input surface, including documents, email, forms, and tool results, an assessment tests the agent as deployed.

Attack classes are described at the conceptual level. No techniques or payloads appear on this page or in our tooling.

Sources