I quit OpenAI because its culture is broken — the time for trial and error is over
What I’m about to tell you has, I realize, become something of a cliché: I resigned this week from OpenAI. I led the writing of the safety reports we published with each major launch. Now I’m joining a parade of former colleagues—at OpenAI and the industry’s other leaders—who have decided that the current path is unacceptable.
I agree with other recently departed staff that the companies building this technology aren’t being nearly careful enough. But I believe that we need to look deeper than specific rules or new laws. We need to talk about culture.
The future depends on wisdom that Silicon Valley lacks. Wisdom about how to handle dangerous technology and, more fundamentally, wisdom about what it means to care for people. This moment needs a degree of humility that isn’t natural for people who have succeeded through their extreme confidence. My former colleagues at OpenAI were prescient: They came to understand the scaling laws that meant bigger AI systems would be smarter—and so they went all in on building bigger systems, at great cost. A can-do attitude of achieving the seemingly impossible—coupled with work timelines that amount to perpetual sprints—are common across the industry.
The safety approach that emerges from such a culture starts with unimpeded optimism about being able to solve problems as they arise. OpenAI has thrived by trial and error (which it calls “iterative deployment”), looking for problems and improving its guardrails in response. But this approach, by its very nature, guarantees periodic failures—and the scale of those failures is growing as systems get more capable. This summer, in the Hugging Face incident, OpenAI let a swarm of agents out by mistake. The company responded by making security improvements. But even after those changes, OpenAI reported that its safety controls failed again, when a model in training bypassed restrictions on internet access: A monitoring system alerted human staff but did not automatically turn the model off as it was supposed to. Anthropic, too, has acknowledged accidentally turning off its own safeguards because of a misconfiguration. I believe that such mistakes are typical of the industry, given the speed and flexibility with which people operate.
An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are and that might not do what we want them to. Paul Christiano, on joining OpenAI’s board a few weeks ago, wrote that “there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.” If this is the situation, then the time for trial and error is over. [Continue reading…]