Frontier AI Models Struggle with Autonomous Malware Analysis When Evidence Invalidates Initial Conclusions
SentinelOne Labs benchmarked whether large frontier AI models can maintain investigative integrity during long-horizon malware analysis tasks when new evidence contradicts prior reasoning. The research highlights fundamental reliability gaps in autonomous AI-driven security analysis.