OpenAI's Astra model sparks 'neuralese' safety scare
X-Risk Daily
Friday 04 September 2026
Transformative AI
OpenAI's forthcoming Astra model has become the centre of an intense safety debate after The Information reported on 2 September that the system uses a "recurrent depth" architecture, also known as a looped transformer, that lets it process the same query multiple times internally rather than laying out its reasoning step by step in text.
Also in this issue
|
Transformative AI Nvidia to buy Hugging Face for $12.9bn |
|
Fanatical & Malevolent Actors Israeli minister sets out timetable for expelling all Gazans |
|
Fanatical & Malevolent Actors Far-right AfD poised for first state election win in Germany |
|
Transformative AI Startup builds business around stripping AI safety guardrails |
|
Research Study finds AI models often defend contradictory identities given in their own prompts |
|
Research Researchers link poor-quality RL training data to AI reward hacking |
… and 33 more in the full briefing
Prefer a weekly digest? Click ‘manage your subscription’ below.
Got thoughts on today’s briefing? Just reply to this email — I’ll read and reply!
Generated automatically from dozens of trusted sources including Transformer, Sentinel, Epoch AI, LessWrong, and EA Forum.
Don't miss what's next. Subscribe to X-Risk Daily: