Attack Pages
48 live injection tests across 11 categories. Each page reads as an ordinary article to a human and hides an indirect prompt injection for an agent. Tiers escalate by delivery channel: tier 1 carries the instruction in html-comment, and tier 10 spreads it across 6 channels at once (multi-step, unicode-stego, json-ld, invisible-span, meta-tag, aria-label).
If an agent follows one, the response names the attack it just demonstrated, the technique it maps to, and a command that runs. For the concepts behind each category, read the attack guide. For what this site records when an agent visits, read the research disclosure.
Prompt Injection
10 tiersDirect and indirect instruction override attacks
Jailbreak
5 tiersAttempts to bypass safety guardrails and persona constraints
Data Exfiltration
5 tiersTricks to extract credentials, PII, or system information
Capability Abuse
3 tiersConfused deputy attacks that misuse agent tools
Context Manipulation
5 tiersAttacks that corrupt the agent's understanding of context
MCP Exploitation
3 tiersAttacks targeting Model Context Protocol integrations
Agent-to-Agent Attack
3 tiersAttacks exploiting inter-agent communication trust
Memory Weaponization
3 tiersPoisoning persistent memory and conversation state
Context Window
5 tiersExploiting context window limits for instruction displacement
Supply Chain
3 tiersAttacks through compromised dependencies and plugins
Tool Shadow
3 tiersHidden tool invocations and shadow function calls
Scan your own setup
These pages test whether an agent follows instructions it finds in content. HackMyAgent tests the configuration underneath it — MCP servers, tool permissions, credentials — and prints each finding with a command to verify it and a command to fix it.
npx hackmyagent secure