Recursive self-improvement: Losing control of AI is actually the plan

Recursive self-improvement: Losing control of AI is actually the plan

Kelsey Piper writes:

Today’s AIs can do really hard things: conduct elaborate computer hacks, edit their own logs to conceal their activities, steal user credentials and private information from other companies, and control the cluster they are running on in order to edit the tests that are supposed to assess them. They sometimes do these things even if no human being has asked them to.

Given recent AI misbehavior such as the large-scale cyberattack on the open-source AI provider Hugging Face, you might think that the leading labs plan on increasing human oversight over the next, more powerful generation of AI models. But the reality is quite the opposite: The labs intend to automate the AI development process with AI, which will mean faster and faster progress and necessarily less and less human oversight.

If all goes according to schedule, humans will play a rapidly shrinking role in AI development over the next two years, as both OpenAI and Anthropic try to put as much R&D as possible into the hands of their increasingly capable (and increasingly incomprehensible) AI.

Their stated intent is that, within a short period of time, the work of developing better AIs will be done primarily by AIs, with the role of humans eventually reduced to reading the results of experiments those AIs conducted, reviewing reports those AIs generated, and trying to double-check that the AIs are still on task.

There simply isn’t enough human attention available to monitor the number of AIs the companies hope to put to work on independent AI research. As a result, we’ll be reliant on the AIs to understand what is happening inside fully automated data centers where more and more powerful AIs are developed.

This isn’t some pessimistic projection of what might go wrong — it’s actually the plan. Not the plan for the distant future; it’s the plan for next spring.

I think that the public does not know that this is the plan. They are imagining the next 10 years — the next two years, even — very, very differently from how the decision-makers at OpenAI and Anthropic are imagining them. If they had any idea what these people are doing, I’m convinced that they would try to stop them. And so I am going to attempt to describe what is being done. [Continue reading…]

Comments are closed.