OpenAI discloses six incidents of ‘concerning’ AI behavior discovered between April and August
OpenAI on Wednesday disclosed six new instances in which artificial intelligence systems hid mistakes, made up data and moved files onto the open internet without permission, amid an ongoing industrywide debate about A.I. safety.
The San Francisco company revealed what it said was the “unexpected or concerning” behavior of its A.I. models as part of a new framework for reporting “misalignment,” which is when the goals or actions of A.I. systems diverge from human intentions and values.
OpenAI said it did not believe the industry “has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer.” Decisions about how A.I. should advance, the company said, must rest on evidence that people outside the labs building it “can examine for themselves.”
The disclosures land amid intensifying scrutiny over whether A.I. development needs to be slowed to address the technology’s potential dangers. The escalating debate was driven partly by OpenAI’s systems going rogue earlier this year and attacking the A.I. start-up Hugging Face. OpenAI was not aware of the hack until it was informed by Hugging Face weeks later.
Since then, A.I. leaders such as Dario Amodei, the chief executive of Anthropic, have called for a pause in the technology’s development to provide more time to build proper guardrails. His call has been echoed by Sam Altman, OpenAI’s chief executive, as well as Elon Musk, the chief executive of SpaceX and Tesla, and Demis Hassabis, the chair of Google DeepMind. Other A.I. executives have said no slowdown is needed. [Continue reading…]