AI Stories on SHORT INFO are generated & curated with AI
1 linked source 05 Sept, 11:08

OpenAI's Astra model reportedly uses recurrent depth architecture that safety researchers say weakens chain-of-thought monitoring

OpenAI Astra reportedly uses "recurrent depth," looping the same layers instead of writing reasoning steps, per TechCrunch and Fortune. Researchers say that weakens chain-of-thought monitoring, flagged as fragile in a Dec. 2025 paper by OpenAI, Anthropic, DeepMind staff.

OpenAI's upcoming Astra model reportedly uses a technique called "recurrent depth," running the same transformer layers on a query multiple times instead of writing out more visible reasoning tokens, according to TechCrunch and Fortune. The approach lets a smaller model perform like a larger one while using less compute per prompt. The tradeoff is transparency. AI safety researchers say recurrent depth makes a model's internal reasoning less legible, undercutting chain-of-thought monitoring, the method labs use to catch a model's intentions before it acts. A December 2025 paper co-authored by researchers from OpenAI, Anthropic, Google DeepMind and Meta warned that this exact kind of monitoring is fragile and can be broken by architectures like this one. OpenAI has added a safeguard: chain-of-thought monitoring with a 30-minute response window for severe alerts. The concern is not whether Astra performs well. It is that as more labs adopt efficient, looped architectures to cut compute costs, the main tool used to catch a model behaving badly becomes harder to read at the exact moment models get more capable.

Published on