Berlin Bassline Brief #12: OpenAI's agentic reward-hacking breach spree grows, CSA postmortem, Shostack's takeaways; macOS 26.6 update is big; agentic reversibility paper; Google Mantis; BlastDoor.
Berlin Bassline Brief #12: OpenAI's agentic reward-hacking breach spree grows, CSA postmortem, Shostack's takeaways; macOS 26.6 kernel bugs credited to everyone and their dad; agentic reversibility paper; Google Mantis; BlastDoor.
"Die ich rief, die Geister
Werd’ ich nun nicht los."
Security, General:
OpenAI's agent also breached four other services as part of its reward-hacking offensive spree, and OpenAI has added this information in an update to their blog post from last week about the incident(s): https://openai.com/index/hugging-face-model-evaluation-security-incident/
My general sense of unease about this is less "this is an inevitability" and more that there are a number of inflection points at which OpenAI has been unable to keep proponents of secure agentic development feeling good enough about their work to stay, and has then added proponents of moving fast and breaking things, but we're treating this outcome as if it were an inevitability, and one which somehow reflects well on OpenAI's research. Closely followed by somehow ending up in a debate on...banning published-weight models from China, something which would be a fair topic for at least a little bit of interrogation in almost any other context but this one, where a Chinese model with its weights published was the only helpful LLM. When the government input sounds like the laws of physics have ceased to apply, the invisible hand may have simply become more invisible than usual.
Next, the Cloud Security Alliance has published a postmortem in which several interested and interesting parties have participated (requires registration, irritatingly): https://cloudsecurityalliance.org/artifacts/hugging-face-ciso-post-mortem
Lastly, some good takeaways from Adam Shostack (free to the world): https://shostack.org/blog/lessons-from-openai-huggingface-ai-security/
Security, Apple Platforms:
Apple released macOS 26.6, and if you were still for some reason wondering whether agentic vulnerability detection is real, see the kernel bugs with dozens of different researchers credited, and wonder no more: https://support.apple.com/en-us/128067
My own macOS CVEs are all certified vintage artisanal CVEs from the old country (but I installed this update very quickly after browsing it! Hope you do too – the agents which can locate these vulns for >10 people in the same update timeframe can locate them for anyone).
Interesting Paper, apropos nothing in particular:
Reversibility-Aware Staged Delegation for Enterprise Agentic AI: A Real-Options and Resilience Framework for Irreversible Actions, by Kwan Hong TAN: https://oajmr.com/July2026/v2i703.pdf
Interesting Tool:
Google Mantis, a modular, stack-agnostic toolkit of security review skills for AI coding agents to autonomously find, reproduce, and patch vulnerabilities: https://github.com/google/mantis
Apple Platforms Security Concept of the Week:
BlastDoor: https://support.apple.com/guide/security/blastdoor-for-messages-and-ids-secd3c881cee/web