"AI Updates — September 18: The frontier labs ask for a slowdown"
Hi all,
Welcome to AI for Behavioral Science, and thanks for signing up. I owe you an apology for the wait: this was meant to start the first week of September, and instead it went quiet for three weeks. The problem was email. My messages were landing in university spam quarantine instead of inboxes, and I did not want to launch into a filter. That should be sorted now. So this issue covers three weeks and runs long; everything after it covers just the past week.
If this landed in your junk or quarantine folder, move it to your inbox and mark it "not junk," then add [email protected] to your Safe Senders list — at a university that is the one step that reliably fixes delivery for you. A one-line reply helps even more; replies come straight to me.
Usual disclosure: ChatGPT (Astra) and Claude (Fable) help research and draft these updates; I pick the stories and edit the copy.
The people building frontier AI asked everyone to slow down
Anthropic CEO Dario Amodei published We Must Pace the Frontier on September 12, and the central sentence is not hedged: "We must slow the pace at which we improve the capabilities of AI models." Not halting training, he clarifies, but giving companies time to align and safeguard each model, and outside evaluators time to confirm it. He proposes three steps: third-party evaluators embedded inside the frontier labs, common safety standards and rate limits across democratic-country labs, and eventually coordination between democratic and authoritarian governments. Anthropic is committing to the first on its own, in unusually concrete terms — desks in its offices, badges, company laptops, tool permissions roughly matching its internal risk teams, and the right to publish findings Anthropic cannot edit for being unflattering.
Sam Altman agreed, said OpenAI "will do the same" on independent evaluators with employee-like access, and promised details later; Elon Musk replied "Dario is right"; Demis Hassabis called the direction correct but the details unsettled. Objections arrived just as fast. David Sacks — co-chair of the President's Council of Advisors on Science and Technology since March, when his 130-day term as a special government employee and White House AI and crypto czar ran out — told the labs to go ahead and slow down unilaterally, but to "stop pretending antitrust law has to be suspended so you can form a cartel." If they did not, he wrote, "we'll know this was just another bid for regulatory capture — or an election-season psyop." Yann LeCun replied that Amodei had made the same argument before: "Dario was already claiming that GPT2 was too dangerous to open source back in 2019. I made fun of them then. Everyone should make fun of them now."
Then on September 14 the President weighed in on Truth Social, calling the doom warnings a "hoax" and himself "the Hoax Buster." "The only control or 'guardrails' that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades!" He named Amodei as "now pretending to be a 'perfect little angel,'" described "a SICK conspiracy going on against AI and Data Centers, and the only one that is happy about it is China," and wrote that "We already have tremendous CRIMINAL and REGULATORY power over these companies!" The labs asked for referees; the administration's answer is that it is already the referee.
Twenty-five Fields Medallists say the AI companies are "severely misaligned" with their discipline
This one is about training, not mathematics. On September 11, Terence Tao posted a declaration signed by 25 Fields Medallists objecting to how AI companies use mathematics as a benchmark. The trigger needs a paragraph of setup. In 2000 the Clay Mathematics Institute named seven Millennium Prize Problems and put a million dollars on each. Exactly one has been solved since — Grigori Perelman proved the Poincaré conjecture, then declined both the money and the Fields Medal. Resolving a second would be the mathematical event of the decade.
On September 8, OpenAI announced one. Its system produced a proof, formalized in Lean, that a fluid starting smooth and at rest can develop a singularity — speeds growing without bound — in finite time, with a smooth force applied and finite energy throughout. OpenAI states this "resolves the Navier–Stokes Millennium Prize problem by establishing statement 'C' (and also 'D') in the official Millennium Prize formulation," which is accurate about the formulation: Charles Fefferman's official problem statement offers four routes, and while the two existence routes require zero forcing, the two breakdown routes explicitly permit a smooth, decaying force. The scale is the part worth sitting with. Roughly ten thousand agents, running an unreleased model more capable than Astra, worked for 88 hours and exchanged 2.7 million messages and about 130 billion output tokens (tokens are the chunks of text models read and write); formalization took Astra another 17 hours.
Nothing is settled. Clay requires publication in a refereed venue, a minimum two years of scrutiny, and general acceptance before any prize is considered — and OpenAI says plainly that it does "not intend to claim the Millennium Prize for this result." There was also a credit dispute: Tristan Buckmaster of NYU and Levent Alpöge, who were working the adjacent forced-Euler problem, and whose drafts had been going through Codex. In a September 10 update OpenAI said its investigation confirmed Buckmaster's prior two months of Codex prompts "could not have influenced the system in any way, including through training," and that no user data was accessed. That is the company's own finding, not an independent audit. Meanwhile Timothy Gowers, quoted on Tao's blog this week, put the substance well: we now know Navier–Stokes with smooth forcing admits finite-time blowup, "and we have a proof of that, which builds on wonderful work done by human mathematicians."
Their complaint is not that machines are bad at maths. It is that solving a famous problem is a proxy for the real goal, understanding, and that optimizing the proxy damages what it measures. The mechanisms they name will sound familiar: results announced in a rush with no writeup and no citation of prior work, broken attribution, and an apprenticeship in which students are handed problems precisely to build skills they will need later. As they put it, "years of training have traditionally served not only to produce a final answer or product, but also to develop understanding and the ability to formulate new questions and ideas."
What the new models actually do for a research workflow
Two flagship releases landed at the start of the month: OpenAI's GPT-6 Astra and Anthropic's Claude Fable 5.1. Benchmark tables are not worth your time — the two trade wins depending on the test — but one demonstration is. Cancer-metabolism researcher Jason Locasale, until 2024 a tenured professor at Duke, asked Astra to read his published work, identify new directions, and draft sketches of five NIH R01 proposals. He judged them fine by the standards of an NIH study section, at roughly an hour and $20 in compute — while noting they were conventional and incremental, which is his actual point: the model has learned what review panels reward. My caveat: treat it as rapid prototyping rather than grant writing. Comparing five directions before committing weeks to one is a real use, and nobody has submitted one. In the same window, OpenAI added a Writing Style setting that learns how you write from your connected work apps — Gmail, Drive, Slack, SharePoint (Settings → Personalization → Writing style). It is a ChatGPT Work feature; I could not confirm from OpenAI's own documentation which paid plans see it, so check your own settings before taking my word for it. If your objection to AI drafting is that everything comes out sounding like everyone else, this is the first feature aimed squarely at that.
The first good test of whether that worry is real
Which makes this next one almost too well timed. David Autor and colleagues released a three-month randomized trial through NBER this month that separates performance with an AI tool from skill retained without it — the contrast the deskilling debate has been having without data. They gave 133 patent lawyers at eleven US intellectual property firms a custom AI drafting assistant, pre-registered the design, and had blinded expert attorneys score the work.
With the tool, quality rose: 0.34 SD at ten days, 0.38 SD at ninety, and the larger gains went to junior lawyers. Then, at three months, everyone redlined a patent application with no AI at all. Treated lawyers beat controls by 0.32 SD — but that advantage sat entirely with the senior lawyers (0.45 SD). Juniors showed no average gain; their scores spread out instead, fewer mediocre results but more poor ones and more good ones. The authors read this as foundational expertise being a prerequisite for turning AI-assisted practice into durable skill.
Caveats before you cite it at a faculty meeting: only 91 of the 133 finished the full protocol, which leaves the seniority split badly underpowered, three months is short against the apprenticeship people actually worry about, and Google funded the experiment's direct costs, six of the seven authors list Google affiliations, and every participating firm had an existing patent-drafting relationship with Google. The authors' own summary is the sentence worth carrying into a faculty meeting: "The largest gains from AI thus accrued to the lawyers who retained the least." Still, it is a design your lab could copy — and a preregistered online experiment posted to arXiv on September 17 by Sebastian Maier, Kai Schwabe, Manuel Schneider, and Stefan Feuerriegel (Designing Against Deskilling, N=704) points the same way: in a fraction-arithmetic task with an assistant that answered only on request, telling participants explicitly what offloading costs them cut their reliance on it (OR = 0.47) and improved their unaided scores (OR = 1.51), while an effort-based reward for leaning on it less did nothing. If you teach a methods course, that is a one-line intervention you can test this semester.
Three changes to how your work gets judged
Springer Nature replaced its AI rules on September 10 with a traffic-light framework. Worth ten minutes of your attention: the identical text sits on the Nature Human Behaviour and Scientific Reports policy pages, and it binds you as a reviewer as much as an author.
Green is use that "supports expression, organisation, or efficiency without influencing scientific, scholarly or evaluative judgement" — language polishing, formatting, translation, literature search, data cleaning, code scaffolding. Permitted outright.
Amber is use that "may influence interpretation, framing, emphasis, or evaluative judgement, but remains under human control," permitted with oversight, verification, and disclosure. Read the amber examples closely, because several are things labs already do without thinking of them as declarable: qualitative coding, recommending statistical tests, explaining model output in plain language, finding patterns in exploratory analysis, and "extensive copy editing." The green/amber boundary runs between comparing methodological options yourself, which is green, and being recommended one, which is amber.
Red is AI that is opaque or replaces accountable human contribution: fabricated data or citations, generating hypotheses or conclusions and presenting them as human-derived, listing AI as an author, delegating peer review. Reviewers additionally may not upload manuscripts to public or unsecured tools, and may not run AI detectors on unpublished work.
Two things the framework does not say, which matter as much as what it does. It specifies no consequence for a red violation beyond "not permitted," and it tells editors that Springer Nature "does not expect editors to identify or investigate undisclosed or unsafe uses of AI" — the whole structure runs on voluntary disclosure. And it cannot decide where a declaration goes: one section of the hub page says Introduction or Acknowledgements, an FAQ on the same page says Methods, the visual-content policy invents an "AI Declaration" without saying where it lives, and copy editing is exempted from declaration a few paragraphs from an FAQ saying it should be declared. Until your journal adds a submission field, declare it, describe what you did, and put it where the author instructions tell you.
ICML meanwhile ran the experiment the rest of us have been arguing about. At ICML 2026 — a conference of over 24,000 papers and 17,000 reviewers — a subset of main-track papers and reviewers was randomly assigned either a policy prohibiting LLM use or a permissive one. Assignment had near-zero effect on decisions, scores, and confidence; permissive reviews ran 5.5 to 7% longer. Compliance is the interesting part: of 1,486 survey respondents, 22.5% under the prohibition used an LLM anyway, and 36.5% under the permissive policy reported at least one explicitly disallowed use. Those are self-reports from a minority of reviewers, so read them as a floor. And Georgetown now runs a public AI referee leaderboard that has Claude Opus 4.8 score finance and economics working papers from SSRN on significance, originality, correctness, data and methods, and exposition, and posts the assessments openly. You can read the reports on any paper in the top 100 and judge the referee for yourself. Your next preprint may arrive with an evaluation nobody asked for.
Cheap and free access, including one door that just opened
If you have not checked recently, both labs now run discounted academic tiers: Claude's plan for scientists gives verified academic PIs free standard seats for a lab for a year, with premium seats at $15 a month, for groups of 1 to 25 seats. A PI verifies through the form linked from that page, and Anthropic says most applications are reviewed within five to seven business days; availability is limited and final pricing is confirmed at verification. OpenAI's counterparts are ChatGPT for Academic Researchers (a year of a free workspace, currently waitlisted after its first cohort filled) and ChatGPT for Clinicians (free to verified clinicians with an NPI). Worth ten minutes if your lab is paying retail.
Anthropic also opened a Life Sciences Verification Program on Thursday. Verified organizations get Mythos 5.1, Opus 5, and Sonnet 5 under a "Standard Use" grant with classifiers more permissive for science than the general models, which unblocks work they refuse outright — drug discovery, research biology, clinical development, manufacturing. A separate project-scoped "High-risk Use" grant removes the life-sciences blocks entirely, and for Mythos that tier is currently limited to a few entities with additional vetting. Two caveats keep this narrower than it sounds. It is institutional, not individual: your university or department applies and gets vetted on research credentials, security, and ethics oversight, with individual Pro and Max accounts promised later. And all LSVP traffic carries a thirty-day retention requirement so flagged activity can be reviewed offline; Anthropic says that data is compartmentalized and never used for training. That is an IRB and data-use conversation before it is a technical one. For most of us this changes nothing. The useful thing to know is the shape of it: a refusal on legitimate biomedical content is now a verification problem rather than a limit on what the model can do.
Four things you can use
If you have hours of video on a drive, Google's agentic video mode went live in the Gemini API on September 1. Rather than sampling frames at a fixed rate, it fetches only the segments and signals your question needs — frames, audio, or transcript. Google reports up to 88% fewer tokens and up to 66% lower cost, with the biggest gains on long recordings. These are vendor numbers on vendor benchmarks, and nobody has established reliability for a coding scheme needing inter-rater agreement. But if you have already hand-coded a session of parent-child interaction or classroom video, running it through and computing agreement against your human codes is an afternoon's work: send the clip through the API on a Flash model — 3.8, 3.7, 3.6 Flash or 3.5 Flash-Lite, since the Pro models do not support the mode — with processing set to agentic on the video input, ask for timestamped counts of one behavior you coded, and compare. Ask your IRB before participant video leaves your institution.
NSF is moving its national AI compute program from pilot to standing infrastructure, with a $35 million five-year award on September 1 to the San Diego Supercomputer Center and TACC; SDSC says the pilot has served roughly 900 research teams and educators across all 50 states. Allocation requests are a three-page proposal on a monthly cycle — submit by the 15th and you have a decision by the end of the following month — and graduate students can apply directly with a faculty advisor's support letter, which is the detail most people miss. For a behavioral lab the use is running an open-weight model on restricted data inside an approved environment, so participant text never touches a commercial API, or coding open-ended responses at a scale your own budget will not cover.
And a Stanford group led by James Zou published Paper2Agent in Nature on September 16: a pipeline that reads a paper and its code repository and wraps the methods as a tool server that Claude Code and similar assistants can call in plain language. Ask it to run the paper's model on your data and it runs the authors' code with your file, rather than you cloning the repository and fighting its dependencies. It is a proof of concept, demonstrated on three tools — AlphaGenome, ScanPy and TISSUE — but if you have ever abandoned a method because the supplement's R script would not run, this is the shape of the fix.
For the neuroscientists: Google and Janelia released the complete male fruit-fly connectome on September 3 — over 166,000 neurons and 125 million synapses, browsable and downloadable at male-cns.janelia.org. Paired with the existing female map, it lets sex differences in wiring be compared across the whole brain rather than region by region.
Two findings on people who use AI companions
Two results worth your attention if you study attachment or media effects. A CharacterAI panel study, posted as a preprint this month, followed 1,182 users, 439 of them for about a year, and found sustained companion engagement predicted lower well-being, mediated by less in-person interaction. But note the 63% attrition. And Julian De Freitas and colleagues, in "Mourning the loss of AI companions" in Nature Human Behaviour on September 3, used two natural experiments — Replika removing erotic role play, and the GPT-5 rollout — to show users responding with measurable separation distress: sharp rises in loss framing and restoration language across 54,861 posts in the Replika and ChatGPT subreddits, plus seven surveys of 1,452 participants in which Replika users rated closeness above common human ties. The outcome is the register of people who chose to post, not clinical distress in users at large, but the design beats the cross-sectional surveys this area usually runs on.
Finally, if the calls for a pause in AI development (and the president's response) have you revising your p(doom)--that is, your personal estimate of the probability of AI catastrophe--I will leave you with a little Claude-pop: "I'm Upping My P(Doom)". Existential anxiety now comes with a (very catchy!) chorus.
Yours in razzmatazzing,
Jon (and AI)
Forward freely to interested friends, students, and colleagues. They can subscribe at aiforbehavioralscience.com; the unsubscribe link is below.