Paper finds prompt injection can shift Jev's typed decisions, though rarely to the attacker's target
2026-09-26 · that day's edition · one of the five
Adaptive attacks using score feedback push fresh-validation success from 1.8 percent to 3.5 percent.
A paper posted to arXiv on 23 September 2026 tests prompt injection against Jev, a non-generative decision model, using 510 reconstructed InjecAgent cases. The authors report that malicious content shifts action probabilities but rarely causes Jev to select the attacker's target, and that adaptive attacks using score feedback raise fresh-validation success from 1.8% to 3.5%.