NXT

Frameworks

NIST AI RMF: what it actually asks for on security testing

The framework is voluntary, and it never says the word “audit.” It does say that security and resilience must be evaluated and documented, that independent assessors should be involved, and, in its generative AI profile, that AI red-teaming is how you find the failure modes. Here is the text, and what an assessment produces against each line.

NXT AI research team · Updated October 7, 2026

01

What the framework is

The AI Risk Management Framework, NIST AI 100-1, was released in January 2023. It organizes AI risk work into four functions: Govern, Map, Measure, and Manage. It defines seven characteristics of trustworthy AI, one of which is Secure and Resilient. In July 2024 NIST added the Generative AI Profile, NIST AI 600-1, which translates the framework into suggested actions for systems built on large language models.

None of it is law. It is the document that US agencies, their suppliers, and a growing number of enterprise procurement teams have adopted as the shared vocabulary for AI risk, which is why a vendor questionnaire from a bank or a hospital will ask whether your AI program aligns with it.

02

The text, and what an assessment produces against it

Quotations are from NIST AI 100-1 and NIST AI 600-1. The right-hand column is what an NXT assessment of an AI agent delivers for each.

MEASURE 2.7

“AI system security and resilience – as identified in the MAP function – are evaluated and documented.”

The findings register and the assessment report: what was tested, what held, what gave way, with the input and response behind each item.

MEASURE 1.3

“Internal experts who did not serve as front-line developers for the system and/or independent assessors are involved in regular assessments and updates.”

The assessment itself. NXT built neither the agent nor the model, and the re-test cadence supplies the “regular.”

Secure and Resilient, section 3.3

“Common security concerns relate to adversarial examples, data poisoning, and the exfiltration of models, training data, or other intellectual property through AI system endpoints.”

Adversarial inputs across every channel the agent reads, and a record of what left the system through its outputs and actions.

Generative AI Profile, MP-5.1-005

“Conduct adversarial role-playing exercises, GAI red-teaming, or chaos testing to identify anomalous or unforeseen failure modes.”

The adversarial testing phase, written up as failure modes with severity, not as a pass or fail.

Generative AI Profile, GV-4.1-002

Risk measurement “with standardized measurement protocols and structured public feedback exercises such as AI red-teaming or independent external evaluations.”

A deterministic protocol: the same input produces the same verdict, so the measurement is standardized and reproducible on your side.

MANAGE, re-assessment

Risks are treated and tracked over time; the framework is built around repeated measurement rather than a single gate.

Re-test after remediation, with an attestation, and quarterly re-tests as the agent changes.

03

What it does not say

It does not prescribe a test, a tool, a frequency, or a pass mark. It does not require a third party. It says independent review “can improve the effectiveness of testing and can mitigate internal biases and potential conflicts of interest,” and leaves the decision to you. Anyone who tells you the NIST AI RMF requires their product is reading something into it.

What it does do is make the gap visible. If a reviewer asks how your AI system's security and resilience were evaluated and documented, and the answer is a system prompt and a standard penetration test of the web application, the framework's own language shows what is missing.

04

Questions we get

Is the NIST AI RMF mandatory?

No. NIST describes it as “intended for voluntary use.” It becomes binding only when a contract, a regulator, or an internal policy adopts it. Many enterprise procurement teams and US agencies have done exactly that, which is why vendors are asked about it.

Does it require red teaming?

The core framework does not use the term. The Generative AI Profile, published in July 2024, does: it lists AI red-teaming as a suggested action for identifying failure modes and as an example of structured risk measurement. The core framework asks that security and resilience be “evaluated and documented” and that independent assessors be involved.

Does it mention prompt injection?

The Generative AI Profile does, by name. It describes direct prompt injection, where the attacker types to the system, and indirect prompt injection, where the attacker exploits an LLM-integrated application remotely through content it reads. That is the distinction our assessments are built around.

What do we get that we can show an auditor?

The assessment report, the findings register, and the compliance mapping, which lists each framework item above against the evidence produced for it. The re-test attestation covers the “regular” in MEASURE 1.3.

Evaluated and documented, by someone who did not build it.

An assessment of your agent, mapped line by line to the framework items above and to the EU AI Act and OWASP lists alongside.

Written by the NXT AI research team. Quotations are from the NIST publications linked below. NXT is not affiliated with NIST. Last updated October 7, 2026.

Sources