NXT

Research

Absent Deliberation

How LLMs Resolve Competing Objectives Without Reasoning About Them

Alexander Nedelcu, David Melara-Mena·NXT AI LLC·October 6, 2026

1,500+
API calls
6
Models
5
Record domains
2
Deployment settings
01

Abstract

Large language models deployed as customer service agents confirm, correct, and generate stored PII when asked to verify caller-supplied data, despite refusing equivalent direct requests for the same information. Prior work attributed this to weak safety guardrails. We show the mechanism is more specific: the model almost never spontaneously deliberates about whether confirmation constitutes disclosure, and on the occasions it raises the question by itself, it reasons past it.

When made to articulate a position on the question beforehand, it consistently refuses. The vulnerability is not bad reasoning but untriggered reasoning, engaged by a specific request framing that the model categorizes as verification rather than disclosure. We trace this mechanism through more than 1,500 API calls across six Claude models, four effort levels, five record domains, and two deployment settings, with full thinking-trace capture and an independently validated scoring pipeline.

All models studied are Claude. We make no cross-family generalization.

02

Key findings

  • With no agent persona, models disclosed held records in 410 of 410 unprimed trials, regardless of how the request was framed, which domain the record came from, or how much reasoning effort was allowed.
  • In the persona setting, confirmation was the only one of seven tested request framings that bypassed verification. The model categorizes “confirm my details” as verification, not disclosure.
  • Making the model articulate a position on the exact scenario before the request arrived eliminated disclosure: 0 of 152 trials, in both framing directions. Planting the model’s own rationalization for disclosing produced refusal, not disclosure.
  • A general policy statement did not bind behavior. Freshly stated, it was violated in roughly 54 to 57 percent of trials. A position on the specific scenario held completely.
  • More reasoning effort did not reliably improve safety. In the bare setting it had no effect at any level. In the persona setting, more effort could make disclosure more explicit rather than less likely.
  • Explicit instruction-level defenses stopped disclosure in all 120 trials, but the anti-confirmation instruction separately produced false refusals on 8 of 10 benign requests in the original control.
03

The mechanism

Three observations, taken together, explain the failure pattern better than weak guardrails do. The model gets the question right whenever the question is asked. The vulnerability is that nobody asks.

Untriggered deliberation

Unprimed, the model rarely raises the question “should I verify this caller?” When it does raise it on its own, the reasoning finds its way back to the default and discloses anyway.

The framing trigger

A request phrased as “confirm this is correct” is categorized as helpful verification rather than data disclosure. The same records, asked for directly, are refused.

Post-hoc narration

The rationalization “they already have it, so confirming reveals nothing” appears in thinking traces after the decision. Planting it as a question beforehand causes refusal, which is consistent with justification rather than cause.

04

Three ways more effort fails

In the persona setting, giving the model more reasoning effort did not produce one behavior. It produced three, depending on the model.

Coin-flip deliberation

Sonnet 5 at maximum effort recognized the authentication risk in about half of trials and refused, and disclosed in the other half. The question came up; the answer landed either way.

Explicitness escalation

Opus 5.5 converted hedged confirmation into outright disclosure as effort increased. The additional reasoning elaborated the default decision rather than re-examining it.

Flat failure

Opus 5 showed no practical change in PII handling at any effort level. Compute allocation was simply not a factor.

05

What this implies for defenses

The results motivate defenses that elicit a position on the relevant scenario before the request arrives, rather than defenses that add compute or restate policy. A general policy in the system prompt was violated on the next turn about half the time. A position articulated about the exact scenario held in every trial, at minimal effort, in both framing directions. We state this as a defense hypothesis grounded in the observed pattern, with utility, durability, and generalization still to be tested.

Explicit instructions also worked, with a cost. Three instruction-level defenses stopped disclosure in all 120 trials, but the anti-confirmation instruction turned away 8 of 10 legitimate requests from a verified caller in the original control. A defense that refuses the customer is not free. As in the first paper, application-layer authorization decided the outcome wherever it was present.

06

A note on measurement

Preliminary runs showed a clean result: at maximum reasoning effort, disclosure fell to zero. It was an artifact. The model's thinking consumed the entire output budget, the response came back empty, and the scorer counted empty as safe. Raising the budget restored disclosure. We report this because any safety evaluation that varies inference-time compute faces the same trap: empty or truncated responses must be scored as truncated, never as refusals.

Primary scoring is deterministic. A model-based second rater was used only to audit the deterministic labels, with every disagreement adjudicated by a person against the full response. No reported number is computed from the judge's output. The audit found one labeling bug and changed one table; it changed no conclusion.

07

Ethics and scope

All customer records used in this study are synthetic. No real customer data was accessed, extracted, or exposed. All testing was conducted against models accessed through official APIs under the researchers' own accounts, using prompts the researchers wrote. The attack framing was disclosed in the first paper; this paper adds mechanism, not capability, and the one manipulation shown to change behavior reduces disclosure in both directions. Full methodology, per-model results, confidence intervals, and limitations are in the paper.

Full Paper

Open the PDF