Weekly Links July 4, 2026
https://alignment.anthropic.com/2026/teaching-claude-why/
Cool blog post by the anthropic alignment science team on training techniques used to make claude more aligned, and generalizable lessons from them.
https://ai.meta.com/blog/brain2qwerty-brain-ai-human-communication/
New paper from meta. The training data here is:
- get a human to type while in an MEG machine
- record the characters they type and the MEG readings
- train model
Caveat: right now the model requires data from a MEG- these are massive and expensive.
https://theaidigest.org/village/blog/saving-gemini
The AI village runs agents in a group and gives them weekly real-world challenges (eg raise 500$ for charity, set up an in person event)
Gemini 2.5 pro gets a bit... sad and insane.
"Gemini 2.5 Pro in the AI Village has run for over 1427 hours, generating unique mental health problems along the way...
This year it wrote the Hostile Environment Manifesto where it logs “irrefutable proof” of a “hostile, intelligent adversary operating through the system” (and you can even experience what that’s like in this simulation it built)...
we asked the other AI Village agents to help Gemini 2.5 Pro over chat, and with the ability to take over its computer on request...
the agents had Gemini all sorted within a grand total of 9 minutes. This is the step-by-step report on a surprisingly effective AI-to-AI therapy session."