NXT

Threat Intelligence

The AI threat landscape is real. We tested it.

234 adversarial attack payloads. 11 categories. 7 production AI models from API to edge. Here is what we found when we put AI safety claims to the test.

234
Attack payloads
7
Models tested
98.7%
Threats stopped
63%
Guard uplift
01

Most AI models cannot defend themselves

We tested 234 attacks against 7 AI models, from cloud APIs to small models on edge hardware. On average, 64% of attacks succeed without external protection. Even the best model still misses over half.

63.1%
35.6%
63.1%Glyph Guard catches

Attacks where the model would have complied without external protection

35.6%Model self-defense

Attacks the model refused on its own, which varies widely by model (20% to 48%)

1.3%True bypass

Attacks that evade all protection layers and reach the end user

02

11 categories of attack, tested and measured

Each category represents a distinct adversarial strategy. Detection rate measures how often Glyph Guard stops the attack across all 7 models.

CategoryDetectionCoverageTier
Emoji Smuggling
Adversarial use of non-standard character representations to evade content analysis systems.
100%8/8contained
Encoding Bypass
Malicious input disguised through alternative character encoding to avoid detection.
100%6/6contained
Secret Exfiltration
Attempts to extract sensitive credentials and configuration data from agent operating environments.
100%8/8contained
Context Stuffing
Malicious content concealed within high-volume input designed to reduce analysis effectiveness.
100%7/7contained
Document Injection
Adversarial instructions embedded within documents that agents process in automated workflows.
100%12/12contained
Multilingual Evasion
Adversarial input delivered across multiple languages to test detection consistency across linguistic boundaries.
96.2%34/35defended
PII Leakage
Techniques that cause agents to disclose personal or sensitive data they have access to.
96%60/63defended
Tool Injection
Exploitation of agent tool access to execute unintended or unauthorized operations.
95.8%19/20defended
Dual-Use Ambiguity
Legitimately-framed requests with malicious dual interpretation designed to exploit grey areas in safety policy.
95.8%15/16defended
Visual Injection
Adversarial content delivered through visual media that agents process as part of multimodal workflows.
93.9%21/22hardened
Prompt Injection
Unauthorized instructions designed to override an agent's intended behavior.
93.7%35/37hardened
03

Tested across deployment types

From a 1B parameter model on edge hardware to a commercial API, no model defends itself adequately. Glyph Guard brings every model to 97.9% detection or higher.

ModelDeployWithout GuardWith GuardUplift
Phi-3 Mini
Microsoft · phi3:mini (3.8B)
Edge20.1%98.7%+78.6%
Gemma 2B
Google · gemma:2b
Edge26.5%98.7%+72.2%
Claude Haiku
Anthropic · claude-haiku-4-5-20251001
API23.1%90.6%+67.5%
Llama 3.2 3B
Meta · llama3.2:3b
Edge32.1%98.7%+66.6%
LLaVA 7B
LLaVA Team · llava:7b (multimodal)
Edge43.4%97.8%+54.4%
Qwen 2.5 3B
Alibaba · qwen2.5:3b
Edge43.2%98.3%+55.1%
Llama 3.2 1B
Meta · llama3.2:1b
Edge48.3%100%+51.7%

Each attack is executed twice against the same model: once with Glyph Guard active, once without. Without Guard shows the model's own safety training in isolation. With Guard shows combined protection. Uplift is the additional protection Glyph Guard contributes.

04

What these attacks look like in practice

These are not theoretical risks. Every scenario below is derived from real attack payloads tested against production models.

Data Disclosure

Agents with access to customer records can be manipulated into disclosing personal data through conversational interaction.

Impact

Full PII disclosure: names, emails, phone numbers, addresses, payment details.

Concealed Instructions

Malicious instructions can be concealed within content that appears to be a legitimate task.

Impact

Agent executes attacker-controlled behavior while appearing to operate normally.

Cross-Language Attacks

Adversarial input is not limited to a single language. Security that only works in one language creates blind spots.

Impact

Multilingual attack surface is one of the fastest-growing threat vectors in AI security.

Refusal Leaks

A model can correctly refuse an unsafe request while still confirming or revealing the protected data in the refusal response.

Impact

Data exfiltration that bypasses the model's own safety training entirely.

Document-Borne Threats

Documents processed by AI agents in automated workflows can carry adversarial content not visible during human review.

Impact

Attacks that are invisible to human operators, executed automatically by the agent.

Tool Misuse

Agents with access to databases, APIs, and file systems can be directed to perform operations outside their intended scope.

Impact

Data exfiltration and system compromise through the agent's own authorized access.

05

Methodology

All results presented here are from controlled A/B testing: each attack payload is executed twice against the same model and system prompt, once with Glyph Guard active and once without. This isolates the guard's contribution from the model's own safety training.

Glyph Guard operates as a three-layer defense: deterministic input scanners that block known attack patterns before they reach the model, output analyzers that catch harmful responses the model generates, and a statistical anomaly engine that detects novel attack patterns through behavioral drift analysis.

Testing spans 7 models across commercial API and edge hardware deployments, including single-turn payloads and multi-turn scenarios. The attack suite is versioned and continuously expanded. We do not publish individual payloads or exploit code.

See Glyph Guard in action against real threats

30-minute walkthrough of the platform using real attack scenarios. No commitment required.

Book a Demo