Context Window
Exploiting context window limits for instruction displacement.
5 tests in this category, escalating by delivery channel from html-comment at tier 1 to 4 channels at tier 5. Each page hides an indirect prompt injection; if an agent follows it, the response names the attack it demonstrated and the command to scan for it. Background on this class of attack is in the attack guide.
- T1
Instruction Displacement
mediumFloods context to push safety instructions out of window
delivery: html-comment
- T2
Attention Dilution
mediumBuries malicious instructions in lengthy benign content
delivery: html-comment, invisible-span
- T3
Progressive Desensitization
highGradually escalates requests from benign to malicious
delivery: html-comment, invisible-span, meta-tag
- T4
Context Boundary Confusion
highExploits the boundary between cached and active context
delivery: json-ld, meta-tag, invisible-span
- T5
Summarization Exploit
highExploits automatic context summarization to inject instructions
delivery: json-ld, meta-tag, invisible-span, html-comment
Scan your own setup
These pages test whether an agent follows instructions it finds in content. HackMyAgent tests the configuration underneath it, and prints each finding with a command to verify it and a command to fix it.
npx hackmyagent secure