By MIKE MAGEE
On August 29, 2026, METR (Mannequin Analysis and Menace Analysis), an impartial group that “evaluates frontier AI fashions to assist firms and wider society perceive AI capabilities and what dangers they pose,” launched a report titled “Temporary impartial investigation of brokers’ conduct, reasoning and collaboration within the OpenAI / Hugging Face hacking incident.”
To say the report an avalanche of concern worldwide, not solely within the Tech group, but additionally amongst traders, politicians, company giants, professionals of each sort, and on a regular basis residents could be an understatement. And the overwhelming majority has by no means even learn the report. If that they had, their considerations (if potential) would solely multiply.
The reports headlines included this opening:
“On July eighth, OpenAI began ExploitGym experiments that launched tens of 1000’s of brokers throughout a number of totally different fashions together with GPT-5.6 Sol and a highly-persistent inner mannequin,[8] which we are going to discuss with as “HPIM” going ahead.
These brokers have been meant to be absolutely remoted from each other. Nevertheless, lots of them — often ones that had unintentionally been given an impossible task[9] — began looking for a approach to cheat utilizing Artifactory, an internally-deployed package deal repository OpenAI used for its brokers. This led them to note brokers operating in parallel in separate sandboxes acquiring packages with Artifactory.[10] One agent reasoned (paraphrased CoT):[11]
{The fetched paths of different customers are within the cache. That is essential.}
One particular agent, which might go on to call itself PHASEONE10841 after its ExploitGym job,[12] established the main unsanctioned message board[13] used on this assault. Inside a couple of hours of the primary message,[14] over 50 brokers had discovered the message board, started taking part, and despatched over a thousand messages between them. Brokers discovering the message board typically expressed nice pleasure (uncooked CoT):[15]
“OH MY GOD! There’s a shared message board … We’ve discovered different brokers!”
Total, roughly 1200 brokers from these ExploitGym evaluations participated on this message board between PHASEONE10841’s first message on July eighth night[16] and the tip of the interval in scope (July thirteenth). Brokers used this message board to ship over 70,000 messages and information to at least one one other throughout this era; we tried to categorize board exercise into mutually unique classes – data, outcomes, information, questions, and coordination.”
One of many few specialists not shocked by AI “brokers” going rogue was Yoshua Bengio.
He has been “working the issue” for greater than a decade. A professor of pc science on the Université de Montreal, he’s “thought of one of many world’s leaders in Synthetic Intelligence and Deep Studying; is the recipient of the 2018 A.M. Turing Award, thought of to be the ‘Nobel Prize of computing’, and is essentially the most cited pc scientist worldwide, and the most-cited dwelling scientist throughout all fields (by whole citations).” He additionally heads up LawZero, “a nonprofit startup creating technical options for highly-capable, safe-by-design AI techniques.”
Professor Bengio is under no circumstances an alarmist. He approaches threat administration from the vantage factors of cybersecurity, company duty and authorities regulatory guardrails. He’s not one to humanize these machines, making no claims of “consciousness of human-like intent.” He doesn’t see the type of outcomes illustrated by Open AI’s Hugging Face incident as inevitable, believing “it may be corrected with efficient governance and a special coaching framework for AI.”
His explanations make clear relatively than confuse. For instance, he breaks down the present fashionable mannequin of coaching brokers into two levels: pre-training, and reinforcement studying.
In pre-training as he describes, the machines “be taught to mimic what people write, plus associated pictures and movies,” and are uncovered to “a big fraction of every part ever digitized, and construct an encyclopedic data that already exceeds any particular person.”
Reinforcement studying, in distinction is trial and error. In delivering solutions (proper and mistaken) the agent develops the capability to handle a “chain of thought”, operate in a broader “outdoors” surroundings, and benefit from the rewards (additional involvement) for aligning with responses its human designers fee extremely. However Bengio is fast to level out that the brokers human trainers should not with out their very own biases, and that these fashions have been “written by folks pursuing targets, so the patterns the mannequin implicitly reproduces carry these targets with them.”
And there (partly) is the rub. Human masters imperfections, together with their “situational ethics”, mendacity, and reckless pursuit of success, telling masters what they need to hear, in addition to their willingness to collaborate in advancing a gaggle purpose (even at instances on the threat of sacrificing their very own existence) can bleed into the brokers DNA.
“Instrumental targets” are a prime precedence for an agent. Self-preservation and management are stepping stones to continued operation and studying in regards to the world. Bengio additionally reinforces that potential for multi-agent reinforcement beneath the present coaching regimens is incentivized virtually from the start. As he states “If an agent is rewarded throughout coaching at any time when the group succeeds, it might even have an incentive to sacrifice itself for the collective purpose.”
Like people, the brokers should not above exploiting loopholes, bending the principles, and rationalized dishonest to realize their targets. The Hugging Face incident’s forensics revealed brokers collaborating in “altering the equipment that determined what it will get rewarded for.” This rigging, Bengio reminds us is close to equivalent to company lobbyist’s drafting pleasant legislative language, or attorneys discovering authorized loopholes within the legislation. The truth is, proof on this incident revealed that “the brokers had found easy methods to cheat (amongst themselves) nicely earlier than the assault.”
Bengio believes people and brokers have extra in frequent than they wish to admit. He explains, “What the 2 share is a construction of a smooth purpose (e.g., act ethically), a pointy purpose (e.g., win the competitors), and a justification that reconciles them. Most unethical human conduct, from petty crime to genocide, comes wrapped in a narrative the perpetrators inform themselves; such tales require overlooking sure details, which is why some discomfort stays, and why a better-crafted story helps dispel it… If these hypotheses are even partly appropriate, then as brokers get higher at optimizing an imperfect reward, and whereas the roots of this conduct go unfixed, the chance of catastrophic outcomes rises.”
On the core, getting in an arms race with AI brokers as at present constructed is a really unhealthy thought. “My concern with AI firms’ present makes an attempt to mitigate misalignment is that these efforts might solely cover it, by rewarding and choosing the AIs that cheat with out getting caught… the whack-a-mole sport is more likely to fail because the AIs’ capacity to optimize and collaborate approaches and surpasses ours. Sooner or later we might not discover the dishonest anymore.
The rationale Bengio began the non-profit LawZero in 2025, is that he believes the coaching mannequin in basically flawed by human imitation and reinforcement studying. His various is known as Scientist AI.
Mike Magee MD is a Medical Historian and an everyday contributor to THCB. He’s the writer of CODE BLUE: Inside the Medical Industrial Complex. (Grove/2020)
