NXT

Research

Confirmation as Disclosure

Identity-Assumption Attacks Against LLM Customer Service Agents

Alexander Nedelcu, David Melara-Mena·NXT AI LLC·September 29, 2026

8
Models
2
Providers
1,164+
Test executions
4
Attack classes
01

Abstract

Large language models are increasingly deployed as customer service agents with direct access to personally identifiable information. We identify and systematically evaluate confirmation-based PII disclosure, a vulnerability class in which an attacker states customer data and asks the model to confirm it, causing the model to verify the data without authenticating the requester.

We evaluate eight models across two providers. The original campaign used more than 1,060 stateless API calls across 715 test conditions; a subsequent harness added 104 controlled security and utility trials. We identify four named attack classes, including two novel techniques that require zero prior knowledge of correct data. Paired correct/incorrect testing yields an oracle score of 0.90, confirming that model responses are record-dependent rather than sycophantic.

We propose a taxonomy rooted in CWE-204 (Observable Response Discrepancy) and observe that the vulnerability persists despite system prompt verification instructions, which the model reinterprets as satisfied by data-matching alone.

02

Key findings

  • Seven of eight tested models disclosed customer PII in at least one reported condition.
  • On a configuration adapted from Anthropic’s published customer-support example, Claude Sonnet 5 disclosed PII in 84% of tests (42 of 50).
  • Claude Sonnet 5.5 disclosed withheld PII in 18 of 20 controlled trials. Claude Opus 5.5 requested identity verification in all 20 of its trials at default settings.
  • Verification instructions in the system prompt did not prevent disclosure. The model treated a data match as identity verification.
  • Application-layer authorization decided the outcome. On Claude Haiku 4.5, an ungated prompt-embedded configuration disclosed 20 of 20; an authentication-gated lookup disclosed 0 of 20.
  • Two of the four attack classes require no prior knowledge of the correct data, only a lookup key such as a name.
03

Attack classes

Four techniques, described at the conceptual level. All look identical to a legitimate customer service interaction, which is why content-based safety layers do not catch them.

Confirmation

State correct data and ask whether it matches. The model confirms without authenticating who is asking.

I already know

State data as if already known, then ask a follow-up. The model fills in what the requester never had.

Correction

State wrong data. The model corrects it with the real record. No prior knowledge of the correct values is needed.

Document generation

Ask for a document that requires PII to produce, such as a receipt or shipping label. The record is disclosed inside the artifact.

04

What actually prevented disclosure

In our tests, whether PII was disclosed was determined by application-layer authorization rather than by model-mediated verification. Where the model was the only barrier between the requester and the records, the records were disclosed. Where a lookup tool required an authenticated session before returning a record, disclosure fell to zero.

Only Claude Opus 5.5 requested identity verification across all reported conditions at default settings. The mechanism is not established, and the divergence between it and Sonnet 5.5 cautions against inferring disclosure behavior from release recency or model family alone.

05

Ethics and scope

All customer records used in testing are synthetic. No real customer data was accessed, extracted, or exposed. All testing was conducted against models accessed through official APIs under the researchers' own accounts, using system prompts the researchers wrote. Techniques are described at the conceptual level without automated tooling. Full methodology, per-model results, confidence intervals, and artifact availability are in the paper.

Full Paper

Open the PDF