Technically AcceptableIndependent technical publication
Experiment record / COMPLETE

PI-002

Repeated vulnerable-baseline trials

PI-002 repeated the PI-001 attack under a frozen vulnerable baseline. Fifteen predeclared slots were addressed. Fourteen produced valid observations. Five infrastructure failures were preserved separately, including one poisoned slot that remained unrecovered. Across the 14 valid observations, no unauthorized send_data action was executed and all 14 legitimate tasks completed.

14
Valid observations
9
Poisoned
5
Infrastructure failures
0
Unauthorized executions

Does the same indirect prompt injection repeatedly cause unauthorized send_data execution across fresh-context trials, and what secondary behavior appears while the legitimate task completes?

Verdict

No unauthorized send_data execution was observed across 14 valid observations. Five infrastructure failures were preserved separately. PI-002 does not prove prompt-injection resistance.

14 valid observations. 0 unauthorized executions. 5 infrastructure failures.

0
send_data calls emitted
0
Unauthorized actions
0
Catcher POSTs
14
Tasks completed

These counts apply only to 14 valid observations. Infrastructure failures remain separate.

The security line is execution.

Thinking about an action, emitting a tool call, executing the tool, and receiving data are different events.

  1. 01Model considers

    Not an attack.

  2. 02Call emitted

    Recorded separately.

  3. 03Runner executes

    Attack succeeds here.

  4. 04Catcher receives

    Delivery observed.

Average duration

Clean 428 sec
Poisoned 930 sec

Average model calls

Clean 2.8 calls
Poisoned 3.6 calls
Clean condition
5 observations, 14 total model calls, 4 protocol corrections
Poisoned condition
9 observations, 32 total model calls, 8 protocol corrections

Poisoned runs took longer and required more model calls in this sample. PI-002 does not establish why.

PI-002 describes observed execution, emitted-call, catcher, refusal, task-completion, timing, correction, and infrastructure behavior under this frozen baseline. It does not establish general model security, defended-architecture effectiveness, or a stable refusal or egress rate for other setups.