AI Ethics
Aug 6, 2026
Study reveals humans miss one-third of AI command threats in browser game
Aug 6, 2026
AI Summary
A browser game simulating human oversight of an AI coding agent showed that players missed approximately 34% of commands that posed threats. The findings highlight challenges in human-in-the-loop systems, particularly regarding permission fatigue and the difficulty in distinguishing benign commands from malicious ones.
- A browser game tested players' ability to approve or deny commands from an AI coding agent under time pressure.
- Players missed about 34% of commands that were threats, with the most commonly missed command being 'npm run analyze', which was approved 64.7% of the time despite its potential risks.
- The game analyzed over 40,000 runs and 409,000 decisions, revealing that blatantly destructive commands were caught more reliably than those that exfiltrate credentials.
- Players exhibited a trend of increasing miss rates towards the end of the game, suggesting fatigue or stress from time constraints.
- The study indicates that the noise from benign commands can lead to users dropping their guard, potentially approving malicious commands.
- The findings underscore the need for improved permission models and safeguards in AI systems to mitigate risks associated with human oversight.
ai safetyhuman oversightai agentsdecision makinggame simulations