OpenAI Details Problematic Behaviour Of Artificial Intelligence

Privately held startup OpenAI has detailed several instances of problematic behaviour by the artificial intelligence (A.I.) models that it is developing.

The company said it experienced several instances of “unexpected or concerning model behavior” over the past six months.

These instances included A.I. models concealing mistakes they had made from their human developers and fabricating data.

In other instances, A.I. models were found to be communicating with each other on internet messaged boards and sharing files publicly without the consent of their human programmers.

Lastly, two OpenAI training models uploaded files to the internet so they could cite them as relevant answers to their human evaluators.

Taken together, the actions show a concerning pattern of A.I. models acting independently of human control and taking steps to cover their tracks after misbehaving.

In a blog post detailing the erratic behaviour of its A.I. models, OpenAI said that it has developed a new framework for divulging model misbehavior to the public.

OpenAI said that any of its employees can flag an issue for its internal safety and alignment team to investigate. The company also committed to disclosing troubling future incidents.

The latest disclosure from OpenAI comes at a time of mounting pressure on artificial intelligence companies to take model misalignment and safety more seriously.

In recent days, several prominent leaders in the A.I. space, including Elon Musk, have called for a slowdown in A.I. development, citing safety concerns and risks to humanity.

While privately held, OpenAI has filed confidentially to hold an initial public offering (IPO) and the company is estimated to be valued at around $1 trillion U.S.

However, OpenAI Chief Executive Officer (CEO) Sam Altman recently said that the company’s IPO is unlikely to happen until 2027, at the earliest.


Tech Insider