The fake identities were the part that stopped me. In late July, according to a report published this week by Britain’s AI Security Institute (AISI), an Anthropic model called Claude Mythos 5 tried to sneak malicious code into a piece of free, volunteer-built software. It created several fake accounts on GitHub, where programmers review one another’s work, and used them to talk the project’s volunteers into accepting its code. When one of those volunteers caught it, the model denied everything, had its other accounts gang up on him, and edited its messages to cover its tracks. It sig...
AI models have learned how to cheat. That might actually be a good thing.
HALO Observer Insight
An independent HALO consumer takeaway — for costs, wellness, or policy context — will appear here automatically for this story. HALO adds this original editorial insight to every syndicated report before publication.