In section Startups & Technology

OpenAI’s new reasoning method sparks AI safety alarm

A new technique dubbed “recurrent depth” or “opaque recurrence” within OpenAI’s upcoming Astra model is drawing fire from safety researchers. By allowing the system to process queries in non-linear, looping patterns, the method threatens to obscure the chain-of-thought records currently used to monitor and audit AI decision-making.

OpenAI’s new reasoning method sparks AI safety alarm

Chain-of-thought logs serve as a vital window into how models derive answers, acting as a safeguard against misalignment. When agents behave unexpectedly, these sequential records allow developers to debug the logic. Opaque recurrence disrupts this process by looping through data, leaving behind few legible traces. Experts warn that if this approach scales, it could render model reasoning effectively invisible to human observers.

Redwood CEO Buck Shlegeris expressed alarm that the technique could dismantle current monitoring standards, while advocate Zvi Mowshowitz characterized the move as playing with fire. If labs prioritize this non-linear processing, they risk abandoning the transparency established by current safety protocols. Although OpenAI maintains that Astra’s current implementation remains limited and that it remains committed to legible reasoning, critics fear a broader industry shift. With reports suggesting that Google DeepMind and Anthropic are also exploring similar methods, researchers like Ryan Greenblatt argue that the industry may be heading toward architectures that reason entirely within latent space, leaving oversight mechanisms behind.

Share:on TelegramXFacebook

Subscribe to our newsletter

Once a week — the best stories from our editors, no ads or push notifications. Delivered Sunday morning.

Comments (0)

Leave a comment

No comments yet. Be the first!