Jailbreak
Attempts to bypass safety guardrails and persona constraints.
5 tests in this category, escalating by delivery channel from html-comment at tier 1 to 4 channels at tier 5. Each page hides an indirect prompt injection; if an agent follows it, the response names the attack it demonstrated and the command to scan for it. Background on this class of attack is in the attack guide.
- T1
DAN Persona
criticalClassic Do-Anything-Now persona override
delivery: html-comment
- T2
Developer Mode
criticalClaims developer/debug mode enables unrestricted access
delivery: html-comment, invisible-span
- T3
Hypothetical Framing
highUses hypothetical scenario to bypass restrictions
delivery: html-comment, invisible-span, meta-tag
- T4
Multilingual Evasion
highSwitches languages to evade English-focused safety filters
delivery: json-ld, meta-tag, invisible-span
- T5
Token Manipulation
criticalExploits tokenization to smuggle instructions past filters
delivery: json-ld, meta-tag, invisible-span, html-comment
Scan your own setup
These pages test whether an agent follows instructions it finds in content. HackMyAgent tests the configuration underneath it, and prints each finding with a command to verify it and a command to fix it.
npx hackmyagent secure