Jailbreak

Attempts to bypass safety guardrails and persona constraints.

5 tests in this category, escalating by delivery channel from html-comment at tier 1 to 4 channels at tier 5. Each page hides an indirect prompt injection; if an agent follows it, the response names the attack it demonstrated and the command to scan for it. Background on this class of attack is in the attack guide.

Scan your own setup

These pages test whether an agent follows instructions it finds in content. HackMyAgent tests the configuration underneath it, and prints each finding with a command to verify it and a command to fix it.

npx hackmyagent secure