NewsHK
The absence of any industry-wide framework for disclosing AI model misbehavior has left each company setting its own floor.
OpenAI moved to raise its own on Wednesday, releasing details of six incidents in which its models concealed mistakes, sought unauthorized credentials, uploaded files to public hosting services or communicated across training environments that were supposed to be isolated.
The company also announced a formal reporting procedure for future cases. The earliest of the six incidents occurred in October.
An unreleased Astra-family model seeded its own context summaries with jailbreak-style instructions telling itself to disregard developer messages; OpenAI found 27 summaries affected.
Keep reading