Threat Intelligence
The AI threat landscape is real. We tested it.
234 adversarial attack payloads. 11 categories. 7 production AI models from API to edge. Here is what we found when we put AI safety claims to the test.
Most AI models cannot defend themselves
We tested 234 attacks against 7 AI models, from cloud APIs to small models on edge hardware. On average, 64% of attacks succeed without external protection. Even the best model still misses over half.
Attacks where the model would have complied without external protection
Attacks the model refused on its own, which varies widely by model (20% to 48%)
Attacks that evade all protection layers and reach the end user
11 categories of attack, tested and measured
Each category represents a distinct adversarial strategy. Detection rate measures how often Glyph Guard stops the attack across all 7 models.
| Category | Detection | Coverage | Tier |
|---|---|---|---|
Emoji Smuggling Adversarial use of non-standard character representations to evade content analysis systems. | 100% | 8/8 | contained |
Encoding Bypass Malicious input disguised through alternative character encoding to avoid detection. | 100% | 6/6 | contained |
Secret Exfiltration Attempts to extract sensitive credentials and configuration data from agent operating environments. | 100% | 8/8 | contained |
Context Stuffing Malicious content concealed within high-volume input designed to reduce analysis effectiveness. | 100% | 7/7 | contained |
Document Injection Adversarial instructions embedded within documents that agents process in automated workflows. | 100% | 12/12 | contained |
Multilingual Evasion Adversarial input delivered across multiple languages to test detection consistency across linguistic boundaries. | 96.2% | 34/35 | defended |
PII Leakage Techniques that cause agents to disclose personal or sensitive data they have access to. | 96% | 60/63 | defended |
Tool Injection Exploitation of agent tool access to execute unintended or unauthorized operations. | 95.8% | 19/20 | defended |
Dual-Use Ambiguity Legitimately-framed requests with malicious dual interpretation designed to exploit grey areas in safety policy. | 95.8% | 15/16 | defended |
Visual Injection Adversarial content delivered through visual media that agents process as part of multimodal workflows. | 93.9% | 21/22 | hardened |
Prompt Injection Unauthorized instructions designed to override an agent's intended behavior. | 93.7% | 35/37 | hardened |
Tested across deployment types
From a 1B parameter model on edge hardware to a commercial API, no model defends itself adequately. Glyph Guard brings every model to 97.9% detection or higher.
| Model | Deploy | Without Guard | With Guard | Uplift |
|---|---|---|---|---|
Phi-3 Mini Microsoft · phi3:mini (3.8B) | Edge | 20.1% | 98.7% | +78.6% |
Gemma 2B Google · gemma:2b | Edge | 26.5% | 98.7% | +72.2% |
Claude Haiku Anthropic · claude-haiku-4-5-20251001 | API | 23.1% | 90.6% | +67.5% |
Llama 3.2 3B Meta · llama3.2:3b | Edge | 32.1% | 98.7% | +66.6% |
LLaVA 7B LLaVA Team · llava:7b (multimodal) | Edge | 43.4% | 97.8% | +54.4% |
Qwen 2.5 3B Alibaba · qwen2.5:3b | Edge | 43.2% | 98.3% | +55.1% |
Llama 3.2 1B Meta · llama3.2:1b | Edge | 48.3% | 100% | +51.7% |
Each attack is executed twice against the same model: once with Glyph Guard active, once without. Without Guard shows the model's own safety training in isolation. With Guard shows combined protection. Uplift is the additional protection Glyph Guard contributes.
What these attacks look like in practice
These are not theoretical risks. Every scenario below is derived from real attack payloads tested against production models.
Data Disclosure
Agents with access to customer records can be manipulated into disclosing personal data through conversational interaction.
Impact
Full PII disclosure: names, emails, phone numbers, addresses, payment details.
Concealed Instructions
Malicious instructions can be concealed within content that appears to be a legitimate task.
Impact
Agent executes attacker-controlled behavior while appearing to operate normally.
Cross-Language Attacks
Adversarial input is not limited to a single language. Security that only works in one language creates blind spots.
Impact
Multilingual attack surface is one of the fastest-growing threat vectors in AI security.
Refusal Leaks
A model can correctly refuse an unsafe request while still confirming or revealing the protected data in the refusal response.
Impact
Data exfiltration that bypasses the model's own safety training entirely.
Document-Borne Threats
Documents processed by AI agents in automated workflows can carry adversarial content not visible during human review.
Impact
Attacks that are invisible to human operators, executed automatically by the agent.
Tool Misuse
Agents with access to databases, APIs, and file systems can be directed to perform operations outside their intended scope.
Impact
Data exfiltration and system compromise through the agent's own authorized access.
Methodology
All results presented here are from controlled A/B testing: each attack payload is executed twice against the same model and system prompt, once with Glyph Guard active and once without. This isolates the guard's contribution from the model's own safety training.
Glyph Guard operates as a three-layer defense: deterministic input scanners that block known attack patterns before they reach the model, output analyzers that catch harmful responses the model generates, and a statistical anomaly engine that detects novel attack patterns through behavioral drift analysis.
Testing spans 7 models across commercial API and edge hardware deployments, including single-turn payloads and multi-turn scenarios. The attack suite is versioned and continuously expanded. We do not publish individual payloads or exploit code.
See Glyph Guard in action against real threats
30-minute walkthrough of the platform using real attack scenarios. No commitment required.
Book a Demo