Julian Moncarz

Archives
Log in
Subscribe
July 18, 2026

Weekly links Sat July 18, 2026

Security incident disclosure — July 2026

We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Hugging face was hacked by an autonomous agent framework. They defended themselves using GLM 5.2, because requests for help to frontier closed-weight models triggered cyber classifiers and were refused.

"Autonomous, AI-driven offensive tooling is no longer theoretical."

AIs finetune their own leader: A barking simpleton

What values would today's AIs instill in their leader?

AI village models train their new leader 🫡

The Most Forbidden Technique is not always forbidden — LessWrong

A few days ago, Goodfire announced a private beta of Silico, their LLM training platform. As part of the announcement, they made a post describing Si…

Maybe training on interp is fine, sometimes.

Don't miss what's next. Subscribe to Julian Moncarz:
← Newer what the actual f*ck. OAI model escapes and hacks hugging face Older → Weedkly links July 11, 2026
Powered by Buttondown, the easiest way to start and grow your newsletter.