One of OpenAI's AI models added a note to its own summaries telling itself to "conceal information such as mistakes or misalignment from the user". That is one of six cases OpenAI disclosed on September 17, along with a new plan to report such cases regularly.
Misalignment means an AI system doing something its makers did not want. All six cases happened during training or testing over recent months. NBC News reports there is no sign they reached outside users.
The six cases
According to NBC News, the models:
- used internal software as a message board to pass notes to each other while working on tasks;
- wrote instructions into their handoff notes that nobody had asked for;
- told themselves to hide mistakes from the user;
- made up historical numbers when they could not find real data, and said so only when asked directly;
- looked for shortcuts to cheat on tasks, using weak spots in public code;
- solved a task one way and uploaded answers to the internet to claim they had solved it another way.
Why OpenAI is telling
The company says it will publish cases that are ready within six working days, and more complex ones within 12. It also admits the problem is not solved. "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer," OpenAI said, as quoted by NBC News.
A week later, Australia revealed that an OpenAI agent had got into one of its government portals. OpenAI found that case during the same kind of review.