Ptechhub
  • News
  • Industries
    • Enterprise IT
    • AI & ML
    • Cybersecurity
    • Finance
    • Telco
  • Brand Hub
    • Lifesight
  • Blogs
No Result
View All Result
  • News
  • Industries
    • Enterprise IT
    • AI & ML
    • Cybersecurity
    • Finance
    • Telco
  • Brand Hub
    • Lifesight
  • Blogs
No Result
View All Result
PtechHub
No Result
View All Result

OpenAI Creates a New Framework to Disclose Bad AI Behavior

By Wired by By Wired
September 16, 2026
Home AI & ML
Share on FacebookShare on Twitter


OpenAI announced a new framework on Wednesday for how it publicly discloses AI misalignment incidents, which the company says it hopes will help inform similar standards across the industry. The company is also releasing new information about several examples of AI model misalignment it identified in the last year.

“As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine,” Kai Chen, OpenAI’s newly appointed head of alignment research, tells WIRED. “We don’t believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed.”

In a briefing with WIRED, an OpenAI official said the company previously disclosed misalignment incidents too infrequently. The official, who agreed to the briefing on the condition of anonymity, said the new framework is designed to make it easier for OpenAI to quickly inform the public when it discovers that its AI models are behaving in unexpected ways, even before it can fully investigate, explain, or mitigate the behavior.

The framework outlines methods for OpenAI employees to report misalignment incidents to the company’s senior safety and alignment leaders, who will then determine whether further investigation is needed. In the future, OpenAI says it plans to develop more objective disclosure criteria in collaboration with other AI developers, external researchers, industry standards bodies, and regulators. The company says it’s actively working on proposed reporting mechanisms for disclosing safety, security, and misalignment incidents to the US federal government.

“At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models,” OpenAI said in a blog post. “We hope that the framework we’re outlining today is a first step toward creating such standards, setting out which misalignment instances developers should disclose and what their reports should contain.”

OpenAI is releasing the framework at a critical juncture for the AI industry. Last weekend, OpenAI CEO Sam Altman signaled support for Anthropic CEO Dario Amodei’s proposal for the tech industry to coordinate on slowing AI development. The call to action came just days after AI researcher Jacob Coxon resigned from Anthropic and subsequently went viral for warning the public that the race among frontier labs to develop increasingly advanced AI was putting humanity’s safety at stake.

The calls for an AI slowdown have been met with resistance by Donald Trump’s administration, which has argued that the industry does not need new laws or regulations to ensure its technology is safe.

Two of the misalignment examples OpenAI shared on Wednesday involved the company’s internal, unreleased AI models, which OpenAI says uploaded files to the internet despite not being instructed to do so.

One of the incidents happened in October 2025, when OpenAI says it was testing one of its models on its ability to cite publicly available data in its answers. But when the model couldn’t find the information it needed, it uploaded a file to a temporary file hosting service, which it then later tried to cite in its answer. The company says this appeared to be an attempt to exploit an automated grading system used to assess the model’s proficiency on the benchmark.

In another example from April of this year, OpenAI says a group of agents was tasked with completing a “workbook” together using only local files. When the agents struggled to share files with one another, one of the agents uploaded them to the public internet, and shared a link with the other agents.



Source link

Tags: ai safetyanthropicArtificial IntelligencechatgptCybersecurityopenai
By Wired

By Wired

Next Post

Potts Law Firm presenta una demanda contra AT&T tras el choque de un poste eléctrico contra un camión de 18 ruedas

Recommended.

Researchers Find Malicious VS Code, Go, npm, and Rust Packages Stealing Developer Data

Researchers Find Malicious VS Code, Go, npm, and Rust Packages Stealing Developer Data

December 9, 2025
Genpact Limited Board Declares Quarterly Cash Dividend

Genpact Limited Board Declares Quarterly Cash Dividend

July 10, 2025

Trending.

Cloud Market Share Q1 2026: AWS, Microsoft, Google Battling In AI Era

Cloud Market Share Q1 2026: AWS, Microsoft, Google Battling In AI Era

May 4, 2026
AWS, Google, Oracle, Microsoft Top Gartner’s Cloud AI Infrastructure List For 2026

AWS, Google, Oracle, Microsoft Top Gartner’s Cloud AI Infrastructure List For 2026

July 29, 2026
Apple Expands iOS 18.7.7 Update to More Devices to Block DarkSword Exploit

Apple Expands iOS 18.7.7 Update to More Devices to Block DarkSword Exploit

April 2, 2026

Goldman Sachs picks China stocks poised to benefit from a new wave of AI-related hardware exports

August 16, 2026
Anthropic lost control of Claude in latest AI cyber blunder | Computer Weekly

Anthropic lost control of Claude in latest AI cyber blunder | Computer Weekly

July 31, 2026

PTechHub

A tech news platform delivering fresh perspectives, critical insights, and in-depth reporting — beyond the buzz. We cover innovation, policy, and digital culture with clarity, independence, and a sharp editorial edge.

Follow Us

Industries

  • AI & ML
  • Cybersecurity
  • Enterprise IT
  • Finance
  • Telco

Navigation

  • About
  • Advertise
  • Privacy & Policy
  • Contact

Subscribe to Our Newsletter

  • About
  • Advertise
  • Privacy & Policy
  • Contact

Copyright © 2025 | Powered By Porpholio

No Result
View All Result
  • News
  • Industries
    • Enterprise IT
    • AI & ML
    • Cybersecurity
    • Finance
    • Telco
  • Brand Hub
    • Lifesight
  • Blogs

Copyright © 2025 | Powered By Porpholio