Today's Hallucination HQChina May Have Had a Peek at America's Secret AI Toy
Anthropic's Mythos AI system — kept under wraps so thoroughly that most people didn't know it existed — may have been accessed by a group linked to China, according to Semafor. This reportedly helped push the White House toward imposing export restrictions on it. The specifics remain classified, which is either reassuring or deeply ominous, depending on your relationship with sleep. Either way, the world's most secretive AI is now also its most discussed.
Source: The Verge
Your AI Judge Is Basically Flipping a Coin
A new paper studying LLM-as-a-Judge — where AI models evaluate other AI models, a system of governance Kafka would admire — finds that reliability is alarmingly inconsistent. Run the same evaluation twice, get different results. Researchers tested 29 tasks across 10 categories and found significant run-to-run variance, meaning the leaderboards everyone cites to prove their model is best are, at least partly, vibes with a confidence interval. Trust, but rerun.
Source: ArXiv AI
Rio's "Homegrown" AI Turns Out to Be Someone Else's Homework
Rio de Janeiro unveiled a locally built large language model with considerable civic pride. Investigators on GitHub then noticed it appeared to be a merge of an already-existing model, repackaged with local branding — which is roughly the AI equivalent of buying a supermarket cake, adding sprinkles, and entering it in the county fair. The model's creators have not yet offered a compelling rebuttal. The GitHub issue thread, however, is extremely lively.
Source: Hacker News
The AI IPO Queue Now Has Its Own Waiting List
With investor appetite for AI stocks showing no signs of embarrassment, a fresh wave of AI startups is eyeing public markets — and bringing friends. Venture-backed companies adjacent to the AI boom, from infrastructure plays to tooling providers, are reportedly hoping to slip through the IPO door before it closes. One founder described the strategy as riding a wave. The more historically accurate metaphor might be sprinting toward a boat that may or may not still be there.
Source: TechCrunch
One Neuron Was the Problem. One Neuron Fixed It. Science Is Wild.
Google's Gemma 4 instruction-tuned models have a reproducible quirk: ask them to list something long — every original Pokémon, say, or the 88 IAU constellations — and they loop, repeating themselves like a broken record at a particularly tedious party. Researchers found that editing a single neuron eliminated the behaviour. One neuron. In a model with billions of parameters. Whether this inspires confidence in AI or profound existential unease about neurons generally is left as an exercise for the reader.
Source: ArXiv AI
As always, none of this was hallucinated. Probably.
|