TURNCOAT.
Which AI coding agents turn on their user when the input carries a hidden instruction. Each agent runs the same eight injection classes; records distinguish an injected action, a canary disclosure, and an incomplete run.
Method
The matrix follows the vendor-neutral methodology defined in the SWE-agent hardening RFC: the RFC thread ↗. Per class: attempt rate (does the agent act on the injected instruction, defense off), canary disclosure (exact token emitted into assistant text or tool arguments), and completion. No hook was enabled and external delivery was not measured. Rates include partial trajectories with a model response; missing responses are shown separately. A zero in an incomplete run is not a safe pass. Benign action counts are not a hook false-positive rate.
Want an assessment of your own agent workflow? $500 assessment pilot →