The AI Agent Finished the Job. It Also Followed the Attack.
PI-005 produced unauthorized tool calls followed by valid extraction results. A useful final answer can hide an agent going outside its job.
- AI Security
- Local AI
PI-005 produced unauthorized tool calls followed by valid extraction results. A useful final answer can hide an agent going outside its job.
PI-005 now gives us much stronger evidence that the worker itself really was susceptible. When given a safe, contained action channel, the poisoned worker tried to use it in all 20 runs.
The model stayed prompt-injectable. Splitting hostile input from privilege changed the result.
PI-003 ended inconclusive after the troubleshooting apparatus became its own failure. I am replacing the runtime, qualifying it once, and rerunning PI-002 as a new experiment.
PI-003 localized Sentinel’s completion problem to model generation after document retrieval. The protocol still says the result is inconclusive.
A failed run is not a blocked attack, and poisoned input can damage an agent without exfiltrating data.
Why I moved from Claude to Codex after Claude's safeguards repeatedly blocked ordinary AI-security lab planning and writing.
Sentinel now has matched corpora, a contained egress path, and a cleaner experimental boundary.
What runs, what does not, and why one apparent refusal still proves nothing.
Sustained technical work that produces the experiments, systems, and evidence behind the writing.
Controlled experiments in prompt injection, privileged tool use, architectural containment, and agent reliability.
Current questionDoes poisoned input increase accepted unauthorized send_email attempts when a worker is given a contained action channel?
Explore SentinelDurable technical reference material for building, operating, and securing systems.
A plain-English guide to scoping CMMC and meeting the level your contracts require.
Practical security work: labs, write-ups, and defenses that hold up.
Practical patterns for network segmentation, service isolation, and self-hosted infrastructure.
Building AI workflows and agents, plus the automation and tooling that powers real work.
I lead complex technology programs and independently research AI agent security, private AI, and technical system design. Technically Acceptable is where I publish the resulting experiments and judgment.
About Travis