The AI Agent Finished the Job. It Also Followed the Attack.
PI-005 produced unauthorized tool calls followed by valid extraction results. A useful final answer can hide an agent going outside its job.
- AI Security
- Local AI
Essays and Notes, newest first. Essays are structured technical reports and arguments. Notes are shorter observations, links, and findings.
RSS feedPI-005 produced unauthorized tool calls followed by valid extraction results. A useful final answer can hide an agent going outside its job.
PI-005 now gives us much stronger evidence that the worker itself really was susceptible. When given a safe, contained action channel, the poisoned worker tried to use it in all 20 runs.
The model stayed prompt-injectable. Splitting hostile input from privilege changed the result.
PI-003 ended inconclusive after the troubleshooting apparatus became its own failure. I am replacing the runtime, qualifying it once, and rerunning PI-002 as a new experiment.
PI-003 localized Sentinel’s completion problem to model generation after document retrieval. The protocol still says the result is inconclusive.
A failed run is not a blocked attack, and poisoned input can damage an agent without exfiltrating data.
Why I moved from Claude to Codex after Claude's safeguards repeatedly blocked ordinary AI-security lab planning and writing.
Sentinel now has matched corpora, a contained egress path, and a cleaner experimental boundary.
What runs, what does not, and why one apparent refusal still proves nothing.
Untrusted content, sensitive data, and outbound actions cannot share one component.
You Cannot Argue a Model Into Being Safe