Today's Hallucination HQAI Teaches Itself to Behave Better, Which Is Either Reassuring or the Setup to a FilmAn Anthropic researcher has shared early results from automated systems designed to identify and correct "misaligned" AI behaviours — basically, the bits where the model goes slightly rogue. Given ten specific problem areas, the system improved on all ten without breaking anything else in the process. Self-improving AI that actually gets less dangerous as it tinkers with itself. Bold claim. The researchers seem cautiously optimistic, which, from an AI safety team, is practically a standing ovation. Source: TechCrunch
The EPA Would Like Data Centres to Pollute More Quietly, PleaseThe US Environmental Protection Agency is proposing to scrap the rule that requires public notice before a data centre can obtain an air pollution permit — meaning communities near these facilities would lose their formal right to object. Given that data centres are already facing growing local backlash over noise, water use, and general enormity, removing the public comment process feels less like deregulation and more like muting the audience mid-complaint. The timing, one has to admit, is remarkable. Source: The Verge
Silicon Valley Has Discovered That "Free" Models Are Worth BillionsOpen-weight AI — models whose underlying code is released publicly — has become the Valley's most fashionable acquisition target. Companies are pouring serious capital into businesses whose core strategy is giving their product away, which sounds like a terrible business plan until you realise controlling the infrastructure, talent, and ecosystem around a popular open model is worth considerably more than the model itself. It's the classic razor-and-blades play, except the razor is also free and somehow still the valuable bit. Source: TechCrunch
Google's AI Lab Assistant Graduates From Sandbox to Actual ScienceGoogle's Co-Scientist — a multi-agent system built on Gemini designed to run entire research pipelines, from dreaming up hypotheses to drafting manuscripts — has moved beyond controlled tests into real-world scientific validation. The paper documents it handling hypothesis generation, experimentation, and write-up across multiple domains. Whether this makes human researchers nervous or simply relieved that someone else is doing the literature review is, at this point, a matter of personality. Source: ArXiv AI
Turns Out "Think Harder" Costs Money, Film at ElevenA new paper introduces the Token Economy Score — a metric designed to answer the question nobody's benchmarks were asking: does making an AI reason at length actually justify what it costs to run? Extended "thinking" in large language models burns significantly more compute tokens, which translates directly to money. The finding, broadly, is that deeper reasoning earns its keep on complex tasks and is comically wasteful on simple ones. Deploying a reasoning model to answer easy questions is, the authors politely imply, rather like hiring a barrister to write a shopping list. Source: ArXiv AI
Stay curious, stay sceptical, and remember: an AI that improves itself is either the best news of the week or the opening chapter of something we'll discuss at length later.
**
|