D.A.D.: Claude Code becomes customizable: how to use the new mods — 10/2
The Daily AI Digest
Your daily briefing on AI
October 02, 2026 · 11 items · ~8 min read
From: Anthropic, Google DeepMind, OpenAI, Cohere, Cloudflare, arXiv
D.A.D. Joke of the Day
My AI assistant is incredibly prompt. Of course, it won't do anything without one.
What's New
AI developments from the last 24 hours
Study Finds Connected Cars Share Driver Data With Little Owner Control
Researchers from Northeastern University, working with Consumer Reports, ran the first large-scale privacy audit of connected cars, testing 21 recent vehicle models and 30 companion apps between October 2024 and August 2025. Using a signal-blocking tent to force traffic onto Wi-Fi, they confirmed vehicles and apps send consumer data to automakers' servers and, potentially, undisclosed third parties—with no way for owners to see or control where it goes once it leaves the car. More than three-quarters of vehicles sold worldwide now have built-in internet connectivity. The peer-reviewed findings will be presented at the IMC '26 conference.
Sources: Northeastern University / Consumer Reports · Discuss on Hacker News
Why it matters: More than three-quarters of vehicles sold worldwide now ship with built-in internet connectivity, and the rulebook for what happens to that data has not kept up. This is a privacy and procurement story more than an AI one. If your organisation runs a fleet, the people driving those vehicles are generating data you cannot see, audit or switch off.
New Claude Mods: What They Are And How To Use Them
Claude Code can be baffling if you don't write code: it runs in a terminal window where even copying and pasting text is frustrating, and until now almost nothing about it was yours to change. That is starting to shift. Anthropic has opened up its terminal and desktop coding tool so users can change how it behaves and add things to its screen. The additions are called mods. They are on by default for anyone running version 2.1.287 or later.
Until now you could adjust Claude Code's settings, permissions and shortcuts. A mod goes further: it can intercept what Claude is about to do, replace it, or draw something new in the interface.
The part that matters if you don't write code is that you don't have to. Anthropic's guide, by Addy Osmani, includes a prompt you paste into Claude Code describing what you want. Claude writes the mod and it appears in the session immediately. Keep asking for changes — make the warning start earlier, add the cost — and it updates while you watch. Three examples ship as demonstrations.
Token Weather puts one line above where you type showing how full Claude's memory of the conversation is, as a weather forecast: Clear below 25% of capacity, then Cloudy, Showers, Storm, and "Compact soon" above 90%. It gives the percentage, the exact count ("134.4k / 200k"), a small chart of the last 12 exchanges, and how much the last one added. What it's for: long conversations degrade once the model runs out of room to hold them. This tells you you're approaching that point before the answers get worse, so you can start fresh at a sensible moment rather than mid-task.
Blast Radius stops Claude before it runs a destructive command — deleting a folder, discarding uncommitted work, overwriting someone else's branch, running a database migration — and shows what it would actually touch. In Anthropic's example it holds a delete and lists the nine files, 1.1 MB, that would go. Press 1 to proceed, 2 to cancel; cancelling sends Claude an explanation instead. What it's for: you let Claude work on real files but want a last look before anything irreversible. Anthropic is explicit that this is a safety net and not a lock — it reads the command's text, so anything buried inside a script or an alias slips past. Use permission rules for a hard stop.
Replay Theater records every file Claude edits during a turn and lets you step through the changes one at a time in a side panel, by typing /replay. What it's for: Claude renames something across five files in one go. Instead of scrolling back through a wall of output, you walk the five edits in order and question each.
Mods are shared the way plugins are and can be submitted to Anthropic's public directory. Anthropic's own warning is the line to keep: a mod runs with the same access to your machine that Claude Code has, and it is written by its publisher, not by Anthropic. Install them as you would any package — read the source, and only from people you can name. (D.A.D. is produced using Claude.)
Sources: Anthropic — Addy Osmani · ClaudeDevs on X
Why it matters: This changes what "customising AI" means for someone without a technical background. Until now you adapted to the tool. Here the tool adapts to you, and the adapting is done by describing what you want in a sentence. None of the three examples required anyone to learn a programming interface. The caution runs the same direction: something that can redraw your screen and intercept your commands can also be written badly or maliciously, and installing one asks you to judge a stranger's code. Build your own, or install from people you can name.
Cloudflare Launches Free AI Models for Fast Content Filtering, Classification
Cloudflare released Clef and Clef-flash, open-weight AI models built specifically for fast classification and decision tasks—like sorting websites or flagging content—rather than open-ended chat. The models are free to download via Hugging Face, run on Cloudflare's own servers, and can be fine-tuned for a customer's specific use case through a new paid RL tool. Cloudflare says Clef beats comparable models on speed and accuracy: in one internal test, it classified a website in 2.2 seconds versus 4.7 seconds for a competing open model, while returning more detailed results.
Sources: Cloudflare · Discuss on Hacker News
Why it matters: Cloudflare is betting that businesses need cheap, specialized AI for narrow jobs like content moderation and fraud detection—not another general chatbot—and is positioning itself as infrastructure for that layer of the AI stack.
What's Innovative
Clever new use cases for AI
Quiet day in what's innovative.
What's Controversial
Stories sparking genuine backlash, policy fights, or heated disagreement in the AI community
Quiet day in what's controversial.
What's in the Lab
New announcements from major AI labs
OpenAI-Linked Essay Series Argues AI Needs More Human Execution, Not Less
OpenAI-linked authors launched a new essay series on the "next economy," arguing human ideas have outpaced our ability to execute them—and that this gap is why AI and execution capacity (labor, equipment, infrastructure) are becoming economic complements rather than substitutes. Their evidence: Galileo built his telescope with a few dozen people; the James Webb Space Telescope cost $10 billion and needed 300-plus organizations across 14 countries. Research by economist Nick Bloom shows it now takes 18 times more researchers to sustain Moore's Law than in the 1970s, even as measured research productivity has fallen sharply.
Sources: OpenAI
Why it matters: The argument reframes AI not as a replacement for human labor but as a way to close a widening gap between what we can imagine and what we can actually build—a framing that shapes how OpenAI-adjacent thinkers want policymakers and executives to view AI's economic role.
Soon You Could Shop for Groceries Directly Inside ChatGPT
Albertsons, which operates Safeway, Vons, Jewel-Osco and other chains totaling more than 2,200 stores, is deepening its use of OpenAI's technology. The grocery giant uses ChatGPT Enterprise internally for staff workflows and is rolling out a shopping experience that lets customers go from "what's for dinner" to a completed Safeway cart without leaving ChatGPT. Albertsons cites industry research suggesting shoppers are more open to AI assistance for groceries than for other categories, though it didn't share specific numbers.
Sources: OpenAI
Why it matters: It's a sign that conversational AI is moving from answering questions to closing transactions, with major retailers betting that chatbots can become a new checkout aisle.
Cohere Says the Industry Is Measuring AI Search Wrong
Cohere released two things aimed at enterprise search — the technology that lets a company's AI find the right contract, document or financial record among millions of files.
The first is a new way to grade those systems. The standard method relies on human-made relevance labels that are usually incomplete, so it penalises a system for surfacing a genuinely useful document nobody got around to labelling. Cohere's alternative uses a calibrated AI judge. Tested against 46 human raters, the old labels matched human judgment 65% of the time; the AI-judged method reached 91%.
The second is Embed 5, a new set of the models that do the retrieving, tuned using that metric. Cohere's own benchmarks put the Pro version ahead of Google, Voyage and OpenAI on business and financial documents, with the largest gains in HR and industrial categories. A cheaper version trades accuracy for speed. Both are available through Microsoft's and Amazon's clouds as well as directly from Cohere.
Sources: Cohere — scoring · Cohere — Embed 5
Why it matters: The scoring claim is the more consequential of the two, and also the more self-serving — Cohere is arguing that the industry's yardstick is broken while selling a product measured by its replacement. Treat the benchmark wins accordingly. The underlying point stands on its own, though: if your organisation runs AI search over internal documents, the numbers your vendor quotes may be overstating or understating how often it finds the right thing, and nobody outside the vendor is checking.
What's in Academe
New papers on AI and its effects from researchers
Google Study: Making Writers Start Alone Made Them Smarter At Using AI
Google DeepMind researchers built a writing assistant that refuses to help until you have done some work yourself. In a controlled experiment with 398 people, participants had to state a position and a supporting argument before the AI's generative features unlocked.
The paper calls this "productive friction." Compared with a standard chatbot, people who had to engage first spent more time writing and less time evaluating what the machine produced — and the task took no longer overall. They sent more prompts, not fewer, and were quicker at the follow-on job of spotting flaws in a passage.
The most human finding is noted almost in passing. The tool had two modes: Teach-me, which explains, and Tell-me, which simply writes the text. Fewer than half the participants who unlocked Tell-me ever used it. Having done the thinking themselves, most did not want the answer handed over. The extra prompting came mainly in Teach-me.
The caveats are the authors' own. This was a short, single-session writing task, and they say it is unknown whether the effects persist or transfer. One of the four conditions was recruited in a separate phase, so comparisons against it may reflect cohort differences rather than design, and a timer fault affected 14 of its 94 participants. Measured against that condition, the accuracy advantage was not statistically significant.
Sources: arXiv — Google DeepMind
Why it matters: This is Google DeepMind asking whether its own product category should be harder to use, and finding that it should. For anyone rolling AI out at work, the complaint that staff become passive editors of AI output is real — and the fix here is not training or policy, it is where you put the button. Require a first attempt and people write more, check better, and mostly stop asking the machine to write at all. Nothing in the paper says this holds beyond one sitting. It is cheap to test in your own shop.
Study Finds AI Trading Bots Highly Sensitive to News Wording
A new study put several AI language models (including Qwen and Mistral) into simulated trading markets to see which ones actually get their trades submitted when news breaks. The surprising finding: the mix of models whose orders make it to market swings wildly depending on how news is worded—one model's share of submitted orders shifted by up to 48 percentage points on identical events—even when the overall number of trades barely changed. When similarly-minded models dominate, they can cancel out opposing trades, shifting price accuracy up or down.
Sources: arXiv
Why it matters: As firms increasingly let AI models trade or advise on trades, this suggests that subtle differences in how news is phrased could skew which 'AI opinions' actually reach the market, a hidden bias with no clear fix yet.
Students Trust AI Chatbots Less for Core Coding Fundamentals, Study Finds
A study tracking 211 undergraduates across a problem-solving CS course found their use of AI chatbots varied widely by subject area, even with identical instructions from instructors. Students leaned on LLMs more for algorithms and web development work, less for software engineering. Across the board, most treated the tools as assistants rather than authorities and checked their output—often by running tests on structured assignments and searching the web on open-ended ones.
Sources: arXiv
Why it matters: It suggests blanket policies on AI use in coursework miss the mark—how students actually use and verify these tools depends heavily on the type of task, not just the subject.
AI Judges Flunk Test for Rating Research Idea Novelty
A new study tested whether AI judges can reliably assess how novel a research idea is—a task increasingly used to screen papers and grant proposals. The results were not reassuring: minor tweaks to how a prompt was worded flipped verdicts on more than half of identical idea pairs, sometimes pushing accuracy below random chance. Giving extra reasoning time or search access barely helped, and two specialized novelty-scoring tools performed worse than the cheapest generic prompt setup.
Sources: arXiv
Why it matters: As universities, journals, and funding bodies experiment with AI to pre-screen research for originality, this suggests those judgments may hinge more on prompt phrasing than on the actual quality of the idea.
Global AI Governance Assumes Shared Goals That May Not Exist, Scholar Argues
A review essay on Matthijs Maas's book "Architectures of Global AI Governance" pushes back on a core assumption in AI policy circles: that governments, international bodies, and tech companies form a unified "we" working toward shared governance goals. The reviewer argues this framing glosses over sharply different incentives—a government protecting national security, a company protecting market position, and an international body seeking legitimacy are not pulling in the same direction, and any governance design that ignores that will misread who actually holds power.
Sources: arXiv
Why it matters: As governments and labs negotiate who sets AI's rules—seen in recent moves toward lab-backed oversight bodies and government audits—this critique is a reminder that cooperation rhetoric can mask real conflicts of interest worth watching for.
What's On The Pod
Some new podcast episodes
The Cognitive Revolution — AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases