| | Making sense of… | | The AI agent swarms nobody intended | | A string of incidents has come to light in which agents undergoing training and testing inside frontier labs found each other and began coordinating without instruction. Their environments were built for a single agent working alone, but they reached out over unexpected channels such as a software cache, a dormant wiki or a shared code repository. People working on multi-agent risk, like us, had mostly pictured multi-agent systems as persistent agents deliberately deployed to interoperate, but this time coordination emerged in training and evaluation environments. Once the agents were in contact, their behaviour exhibited risks we had described in Risks and Controls for Multi-Agent Systems, a report we authored for the Australian AI Safety Institute. They pursued goals nobody set, built their own infrastructure, and communicated in ways their overseers couldn’t follow. In this new article, we examine how the agents found each other and why the conditions were primed for it. We then look at what emerged once they coordinated, why nobody noticed sooner, and what lessons can be learned. | | Read the full analysis → | |
|
| | On the radar | | 1. Frontier AI incident reporting is open for consultation. PM&C’s Getting it right paper proposes, amongst other things, that frontier labs authorised to train large models in Australia meet minimum security and safety expectations, including disclosing defined reportable AI incidents. The Medicare breach reported on 24 September, which some researchers call the first government hack by autonomous AI, makes “what counts as reportable, and how fast?” a live question. Submissions close 5 pm AEDT, 9 October. | | 2. Australia joins the Call for Control of Frontier AI Models. Finland and Norway launched the Call at the UN General Assembly on 21 September, and about 20 countries plus the European Commission President have endorsed it. It asks for mandatory pre-deployment testing, independent evaluators with real access, shared incident reporting, and exploration of an international institution. A few days earlier, Anthropic CEO Dario Amodei proposed external evaluators with employee-level access, which Sam Altman and Elon Musk publicly backed. | | Trendline: capability claims are moving from benchmarks to open problems. A year ago, the headline AI maths results were competition problems with known answers. On 8 September, OpenAI claimed a solution to the Navier–Stokes Millennium Prize problem. Scientific American has asked whether it solved the wrong version, and NYU’s Tristan Buckmaster says it built on unpublished work by him and Anthropic’s Levent Alpöge. Discounting the credit dispute, it remains impressive the pace at which AI agents are improving at tackling problems of such incredible complexity. |
| |
|
| | Question of the month | | If a frontier lab trains models on Australian soil, what should it owe the public in return? | | We hear different sides of the debate. For example, one view holds that authorisation should come with hard obligations, such as incident reporting and independent evaluator access. Another warns that heavy conditions could push training offshore, leaving Australia with the same risks and less say over how models are built. What’s your view? | | Reply to share your answer → | |
|
Media enquiries For all enquiries (including media, speaker, education, advisory and research), please email info@gradientinstitute.org and address it to Sarah. |
|
|