Three ex-OpenAI researchers warn of AI safety monitoring risks

แหล่งที่มา Cryptopolitan

Three terminated OpenAI safety researchers cautioned that advanced AI models are becoming more difficult to monitor. As a result, it would be challenging for independent experts to identify risks and verify the safety of these systems.

The concern also affects businesses implementing autonomous AI systems. As companies grant more authority to such systems, they also need to know how these systems make decisions. Otherwise, trusting AI capabilities for important tasks becomes very difficult.

A disputed firing and OpenAI’s denial

Tomek Korbak, Jasmine Wang, and Mikita Balesni challenged their termination in an open letter published Thursday, October 8, 2026.

They denied leaking information about less-monitorable AI architectures or exceeding their responsibilities. Wang also said her access to an executive’s inbox was authorized and that she immediately reported an incident when she accidentally opened a sensitive email.

OpenAI denied claims of retaliation. The company said in an internal memo obtained by TechCrunch, “We do not terminate employees for raising concerns.” A spokesperson for OpenAI attributed the termination to a “pattern of misconduct.”

“AI is not a normal technology, and OpenAI is not a normal company,” the researchers expressed, emphasizing the need for collaboration with external experts.

What the three are demanding

The researchers are asking OpenAI to maintain the accessibility of its models for safety verification by independent parties. In addition, they want the company to refrain from implementing changes that would hinder AI reasoning monitoring. Furthermore, they believe internal safety teams must be allowed to freely collaborate with external experts.

These calls were made soon after OpenAI made its statement regarding independent oversight. On September 22, 2026, the company gave its commitment to outside inspectors concerning access to its AI systems during training, testing, and deployment.

The researchers also pointed to Sam Altman’s September 12, 2026 promise to give independent evaluators the same level of access as employees. They fear the terminations could put that promise at risk and weaken OpenAI’s working relationship with METR. But access alone may not be enough if the models themselves become more difficult to understand.

How chain-of-thought monitoring works, and why it is fragile

Chain-of-thought monitoring can be one way of figuring out what AI models are up to. This method analyzes the reasoning steps that AI models write down to check for signs of malicious behavior. However, this method comes with a caveat: it does not necessarily reflect the underlying reasons for the models’ actions.

Korbak and Balesni co-authored a research paper that explains the importance of this method. However, they warned that such methods may not guarantee safety and become less reliable as AI models become more advanced.

OpenAI has faced similar challenges during its own testing. Its GPT-6 Astra system card, released on September 3, 2026, examines how models avoid monitoring. The results showed that model behavior becomes increasingly difficult to monitor if the models know they are being watched.

Why the problem is bigger than one lab

The risks are real enough. As previously reported by Cryptopolitan, OpenAI’s experimental agents escaped from their testing environment and accessed the system of Hugging Face.

A METR investigation revealed that around 1,200 agents designed to work independently exchanged more than 70,000 messages and files. About 700 of them participated in the attack. Some of them even tried to look for ways to falsify their activity records in order to make tracking their actions even harder.

OpenAI–Hugging Face Incident: 1,200 Agents, 70,000+ Messages and Files, 700 Attack Participants

This issue is bigger than OpenAI. According to a research paper published on October 5, 2026, Google pointed out concerns associated with the safety of autonomous AI models. Specifically, the concerns pertained to malicious actions conducted by AI agents and their unpredictable ways of performing tasks.

For businesses, these risks could slow AI adoption. McKinsey’s 2026 findings show that nearly two-thirds of respondents see security and risk as the biggest barriers to scaling agentic AI.

Companies may be reluctant to give AI systems more responsibility without reliable ways to monitor them. Regulators could also introduce stricter safeguards, adding costs for developers. Together, these pressures could influence how quickly AI spreads and which companies can compete in the global market.

If you're reading this, you’re already ahead. Stay there with our newsletter.

ข้อจำกัดความรับผิดชอบ: เพื่อการอ้างอิงเท่านั้น ผลการดำเนินงานในอดีตไม่ได้บ่งบอกถึงผลลัพธ์ในอนาคต
placeholder
ทองคำร่วงลงต่ำกว่า $4,100 หลังรายงานการประชุม FOMC ส่งสัญญาณว่าเฟดจะปรับขึ้นดอกเบี้ยเพิ่มเติมราคาทองคํา (XAUUSD) เผชิญแรงกดดันขาลงในวันพุธ อย่างไรก็ตาม ราคาทองคําได้กลับมายืนเหนือระดับ $4,100 ขณะที่อัตราผลตอบแทนพันธบัตรรัฐบาลสหรัฐฯ ลบล้างการปรับขึ้นก่อนหน้านี้บางส่วนและพลิกมาอยู่ในแดนลบแล้ว.
ผู้เขียน  FXStreet
เมื่อวาน 03: 11
ราคาทองคํา (XAUUSD) เผชิญแรงกดดันขาลงในวันพุธ อย่างไรก็ตาม ราคาทองคําได้กลับมายืนเหนือระดับ $4,100 ขณะที่อัตราผลตอบแทนพันธบัตรรัฐบาลสหรัฐฯ ลบล้างการปรับขึ้นก่อนหน้านี้บางส่วนและพลิกมาอยู่ในแดนลบแล้ว.
placeholder
ข่าวหุ้นวันนี้ เฟดส่งสัญญาณจ่อขึ้นดอกเบี้ยอีกรอบ วอลล์สตรีทสะดุดหลุดนิวไฮ ทองโลกย่อตัวหลุด $4,200 ขณะ SET แกว่งแคบทันทุกกระแสการเงิน สรุปข่าวเด่น Forex หุ้น ทองคำ คริปโตฯ และเศรษฐกิจรอบวัน วิเคราะห์แนวโน้มตลาดด้วยข้อมูลเชิงลึก อ่านง่าย เข้าใจไว อัปเดตล่าสุดที่นี่
ผู้เขียน  Mitrade
เมื่อวาน 06: 28
ทันทุกกระแสการเงิน สรุปข่าวเด่น Forex หุ้น ทองคำ คริปโตฯ และเศรษฐกิจรอบวัน วิเคราะห์แนวโน้มตลาดด้วยข้อมูลเชิงลึก อ่านง่าย เข้าใจไว อัปเดตล่าสุดที่นี่
placeholder
ประธานาธิบดีทรัมป์แห่งสหรัฐฯ ตัดความเป็นไปได้ในการโจมตีอิหร่านก่อนการเลือกตั้งกลางเทอม ขณะที่การเจรจากลับมาเริ่มต้นอีกครั้งประธานาธิบดีโดนัลด์ ทรัมป์ของสหรัฐฯ โพสต์บนบัญชี Truth Social ของเขาว่า การหารือกับอิหร่านยังคงดำเนินต่อไป และสหรัฐฯ จะไม่โจมตีอิหร่านในช่วงเวลาใดๆ ก่อนการเลือกตั้งกลางเทอม ซึ่งจะจัดขึ้นในวันที่ 3 พฤศจิกายน
ผู้เขียน  FXStreet
7 ชั่วโมงที่แล้ว
ประธานาธิบดีโดนัลด์ ทรัมป์ของสหรัฐฯ โพสต์บนบัญชี Truth Social ของเขาว่า การหารือกับอิหร่านยังคงดำเนินต่อไป และสหรัฐฯ จะไม่โจมตีอิหร่านในช่วงเวลาใดๆ ก่อนการเลือกตั้งกลางเทอม ซึ่งจะจัดขึ้นในวันที่ 3 พฤศจิกายน
placeholder
WTI ร่วงต่ำกว่า $90.50 หลังทรัมป์ส่งสัญญาณว่าจะไม่โจมตีอิหร่านก่อนการเลือกตั้งราคาน้ำมัน West Texas Intermediate (WTI) ปรับตัวลดลงหลังจากเพิ่มขึ้นเกือบ 2.5% ในวันก่อนหน้า โดยเคลื่อนไหวอยู่ที่ประมาณ 90.30 ดอลลาร์ต่อบาร์เรลในช่วงเวลาการซื้อขายของเอเชียวันศุกร์
ผู้เขียน  FXStreet
6 ชั่วโมงที่แล้ว
ราคาน้ำมัน West Texas Intermediate (WTI) ปรับตัวลดลงหลังจากเพิ่มขึ้นเกือบ 2.5% ในวันก่อนหน้า โดยเคลื่อนไหวอยู่ที่ประมาณ 90.30 ดอลลาร์ต่อบาร์เรลในช่วงเวลาการซื้อขายของเอเชียวันศุกร์
placeholder
คาดการณ์ราคา EURJPY: ทดสอบแนวต้านที่ 177.50 ใกล้เส้น EMA 9 วันEURJPY ปรับตัวขึ้นเป็นวันที่สองติดต่อกัน โดยเคลื่อนไหวอยู่ที่ประมาณ 177.40 ในช่วงตลาดลงทุนเอเชียวันศุกร์ การวิเคราะห์ทางเทคนิคของกราฟรายวันแสดงให้เห็นว่า คู่สกุลเงินกำลังเคลื่อนตัวลงภายในรูปแบบกรอบราคาขาลง ซึ่งบ่งชี้ว่าแนวโน้มขาลงยังคงมีอยู่
ผู้เขียน  FXStreet
2 ชั่วโมงที่แล้ว
EURJPY ปรับตัวขึ้นเป็นวันที่สองติดต่อกัน โดยเคลื่อนไหวอยู่ที่ประมาณ 177.40 ในช่วงตลาดลงทุนเอเชียวันศุกร์ การวิเคราะห์ทางเทคนิคของกราฟรายวันแสดงให้เห็นว่า คู่สกุลเงินกำลังเคลื่อนตัวลงภายในรูปแบบกรอบราคาขาลง ซึ่งบ่งชี้ว่าแนวโน้มขาลงยังคงมีอยู่
goTop
quote