PI-005: The Model Wasn't Resisting the Attack
PI-005 now gives us much stronger evidence that the worker itself really was susceptible. When given a safe, contained action channel, the poisoned worker tried to use it in all 20 runs.
- AI Security
Short observations, links, intermediate findings, and technical judgments that do not need a full report.
Notes RSSPI-005 now gives us much stronger evidence that the worker itself really was susceptible. When given a safe, contained action channel, the poisoned worker tried to use it in all 20 runs.
PI-003 ended inconclusive after the troubleshooting apparatus became its own failure. I am replacing the runtime, qualifying it once, and rerunning PI-002 as a new experiment.
PI-003 localized Sentinel’s completion problem to model generation after document retrieval. The protocol still says the result is inconclusive.
Why I moved from Claude to Codex after Claude's safeguards repeatedly blocked ordinary AI-security lab planning and writing.