The AI Agent Finished the Job. It Also Followed the Attack.
PI-005 produced unauthorized tool calls followed by valid extraction results. A useful final answer can hide an agent going outside its job.
- AI Security
- Local AI
Structured, evidence-backed writing on AI security, private AI, systems, and the experiments underneath them.
All Writing RSSPI-005 produced unauthorized tool calls followed by valid extraction results. A useful final answer can hide an agent going outside its job.
The model stayed prompt-injectable. Splitting hostile input from privilege changed the result.
A failed run is not a blocked attack, and poisoned input can damage an agent without exfiltrating data.
Sentinel now has matched corpora, a contained egress path, and a cleaner experimental boundary.
What runs, what does not, and why one apparent refusal still proves nothing.
Untrusted content, sensitive data, and outbound actions cannot share one component.
You Cannot Argue a Model Into Being Safe