AI Updates — September 25: Opus 5.5 and GPT-6 Sol, AI peer review, and new rules for the math PhD
Hi all,
A busy week! Anthropic and OpenAI both released cheaper frontier models, and people spent the week getting Claude Opus 5.5 to write rap songs. Journals and editors argued about AI in peer review, two new studies and features took up tutoring and critical thinking, mathematicians proposed rules for the PhD, and an OpenAI agent got into an Australian government data portal.
Sigh. This is why we can't have nice things.
Usual disclosure: ChatGPT (Astra) and Claude (Fable and Opus) help research and draft these updates; I pick the stories and edit the copy.
Claude Opus 5.5 and GPT-6 Sol and Luna
On September 22, Anthropic released Claude Opus 5.5 and OpenAI released GPT-6 Sol and Luna. For subscribers, the main change is more work before hitting a usage limit. Anthropic says Opus 5.5 nearly matches Claude Fable 5.1 while using about 40% less than Opus 5 on typical work, and it raised five-hour limits on Pro, Max, Team and Enterprise plans and gave subscribers a limit reset they can save for when they need it; Opus 5.5 is now the default Opus model in Claude Code. OpenAI says Sol comes with higher limits in ChatGPT Work and Codex. API prices fell too: Opus 5.5 costs $4 per million input tokens and $20 per million output, Sol $2 and $10.
On science benchmarks, Anthropic's system card reports Opus 5.5 ahead of Fable 5.1 on Humanity's Last Exam (64.4% versus 60.9% without tools), on research-level math problems drawn from recent arXiv papers (91.2% versus 82.9%), and on reading scientific charts such as Kaplan–Meier curves without tools (64.4% versus 44.8%). GPT-6 Astra still leads on Terminal-Bench-Science, a set of research-workflow tasks (64.6% versus 58.7%); OpenAI reported no science benchmarks for Sol or Luna. Anthropic's system card also notes that in expert red-teaming the model "often overrel[ies] on claims in abstracts rather than understanding the full paper." If you use Claude to read figures from papers, the chart-reading gain is the one to test; if you use it to summarize papers, check its summaries against the full text.
Releases are also coming faster. On September 22, the developer Christian Elton posted an animated timeline of major model releases: one every 73 days in 2023, one every 18 days so far in 2026. The post does not say how the releases were counted, and SearchIntel's tracker of flagship launches gets 37 days and 17 days, but both show the gap at least halving since 2023. And on September 22, in his UN General Assembly speech, President Trump announced that AI is "hereinafter officially called 'Super Intelligence,'" a fundamentally fruitless, counterproductive rename. Researchers already disagree about whether current systems count as artificial general intelligence (let alone superintelligence—typically defined as being demonstrably smarter across all domains than the smartest humans), and the attempted relabeling of all of AI as "superintelligence" by decree just makes that debate harder for the public to follow.
Things to try
A second opinion in Claude Code. Type
/advisor fable(or/advisor opus), and the model you are working with will consult the stronger model before it commits to an approach, when an error keeps recurring, and before it calls a task done (how it works). It is experimental, works on paid plans and API accounts, and each consultation counts toward your usage limits.Update your skills for Opus 5.5. Skills and CLAUDE.md files written for older models often over-instruct, with all-caps MUST rules and requests to spell out reasoning step by step. In Claude Code, run
/claude-api prompt-auditto flag patterns written for older models, and/skill-doctorto see which of your skills go unused and how much context each costs. Anthropic reported in July that it removed over 80% of Claude Code's own system prompt for its Claude 5 models with no measurable loss.
A mental health benchmark you can read, and a psychiatrist's prediction
On September 23, OpenAI released MentalHealthBench, built by its own researchers (paper): 1,215 synthetic conversations modeled on real ChatGPT use, from everyday stress to psychiatric emergencies, in English and several other languages. More than 80 licensed psychologists and psychiatrists from more than 20 countries, who together speak 19 languages, wrote 5,262 weighted criteria describing what a good reply should do or avoid. Scores: GPT-6 Astra 57.3%, Claude Opus 5.5 52.4%, and GPT-4o from March 2025 32.1%, with asking for context and judging urgency the weakest behaviors. The automated grader is also an OpenAI model and OpenAI's model scored highest, so an independent replication would help; the benchmark also scores one reply, not what happens to the person afterward. If you teach clinical assessment, the dataset is a 2 MB zip of conversations and criteria; students can score a chatbot's reply to a non-acute conversation by hand and argue with the experts' weights.
In the British Journal of Psychiatry on September 21, a Liverpool psychiatrist published a letter titled "We are the last generation of human psychiatrists." The author predicts replacement within one to two decades, on the grounds that psychiatric diagnosis runs on language and standardized algorithms, with the largest gains in countries where a consultation costs $100 to $300. Read beside the benchmark, it shows how far the best current model is from the clinicians' own criteria.
Peer review and publishing: an editor and a journal policy
Paul Bloom, a psychologist at the University of Toronto (emeritus at Yale) who edits Behavioral and Brain Sciences, reached a practical conclusion: current models review papers well enough that you should use one on your own work before you submit it. On September 17 Bloom posted that Claude Opus reviews of published papers and of Bloom's own drafts caught "statistical errors, internal contradictions, and conceptual confusions that human reviewers miss," and advised not telling the model the paper is yours, "because they do suffer from sycophancy." Bloom's September 24 Chronicle of Higher Education op-ed, "Could AI Be Better at Peer Review Than Humans?" (free account required), considers letting models review submissions to BBS and decides against it for now, partly because the journal's publisher bars sharing even a paper's reviews with AI. The op-ed also reports that three reviews of one recent BBS submission appeared to be AI-written.
JAMA's editors drew their lines in "Updated Guidance for Author Use of AI in Medical Publication" (online August 10, in the September 15 issue). Allowed with disclosure: AI in research methods, literature search, manuscript preparation, translation and data visualization. Not allowed: AI-drafted letters and opinion pieces, AI-generated clinical images, and responses to reviewers. AI is discouraged for references, after JAMA received manuscripts with nonexistent citations. In my experience, having AI agents retrieve the PDF for each citation and check each claim against its source cuts fabricated and mismatched citations way down, but it is not perfect, so check every reference yourself. Entering a manuscript you are reviewing into an external AI tool violates confidentiality under JAMA's rules, so Bloom's advice applies to your own papers, not the ones you are sent to review.
AI detectors on campus
In The Atlantic on September 21, Will Oremus reported on "the Pangram backlash unfolding on college campuses." Newer detectors such as Pangram are far more accurate than the 2023 tools, but many universities still bar faculty from using a detector score as the sole evidence of cheating, and some students have sued. The article's lead example, a Wisconsin microbiology professor who caught 60 AI-written essays among 350 students, is dropping writing assignments this year. If you assign take-home writing at Vanderbilt, the university's academic integrity guidance says a report to the Undergraduate Honor Council cannot be based solely on an AI detector score and that the council generally will not consider detection scores (Vanderbilt also disabled Turnitin's AI detector in 2023), so a suspected case rests on drafts, version history and a conversation with the student.
Learning with and without AI
A September 23 preprint, "StudentBench," from researchers at the AI company Handshake AI, which funded it (abstract), randomly assigned 2,383 adults to one hour of GRE tutoring from an AI tutor, a human tutor, or no tutor. AI tutoring raised scores 6.15 percentage points over no tutoring, and AI and human tutoring were statistically equivalent within a margin of a quarter standard deviation, at about 1/900th of the cost per point gained, by the authors' estimate. Nearly every participant had a single session, and of the 2,469 sessions in the study only 140 were with human tutors (18 tutors in all), against 2,139 with AI tutors, so the human comparison rests on a small sample. Scores were tested right after the session, with no retention measure.
The same day, Nature ran a feature, "How to stay smart in the age of AI", on whether students who hand off hypothesis generation and evidence weighing lose those skills. It draws on surveys (in one, about 90% of more than 1,000 faculty expect AI to weaken student reasoning) and a 2024 study in which students who practiced math with ChatGPT did worse once it was taken away, and it quotes Daniel Willingham on why critical thinking does not transfer across domains. If you teach, the two pieces together frame the design question: when does the tutor help, and when does it do the thinking for the student?
Mathematicians set rules for AI in the PhD
On September 17 and 18, 24 senior mathematicians met at Harvard's Center of Mathematical Sciences and Applications for a Summit on PhD Math Education in the Age of AI. Their draft report, posted today on Terence Tao's blog, reflects broad but not unanimous agreement. Among the recommendations:
- Assess coursework mainly through in-person written or oral exams.
- Give every graduate student access to AI tools, but never require their use, and have advisors discuss expectations for AI use regularly.
- Three ethical rules: take responsibility for the correctness and understanding of everything under your name, always disclose AI use, and get explicit agreement from everyone involved before entering their material into an AI system.
- Have multiple faculty assess each student's research in person at least once a year.
- Award the degree on a defense requiring complete mastery of the thesis, not primarily on the dissertation text.
The authors rank uses in rough order from those that may help (figures, literature search, cleaning up language) to those that may stunt a student's development, with proof generation last, under the principle that "AI use should accelerate understanding, not bypass understanding." If you direct a graduate program or sit on dissertation committees, the ethical rules and the defense standard carry over to psychology and neuroscience with little editing.
Separately, on September 21 nine mathematicians, three of them Fields Medal winners, formed the Advisory Group on Mathematics and AI, hosted at the Institute for Advanced Study. Its first task is advising OpenAI on how to release more than 100 open-problem results the company says an internal model produced. Members are unpaid, the group will publish its recommendations, and it will not advise on how fast OpenAI develops its models.
Two research tools
On September 24, the National Library of Medicine launched Linked Discoveries, a free experimental PubMed tool that uses AI to map the papers related to a given article, including replication attempts, with graph and timeline views and flags for reviews and retractions. Start from a PubMed abstract to see what followed a finding you plan to build on; NIH says the tool does not judge whether a replication succeeded.
In Nature on September 21, a feature titled "AI co-scientists are revolutionizing how research is done" surveys multi-agent systems that read the literature, propose hypotheses and critique them, including Google's Co-Scientist and FutureHouse's Kosmos. One cancer lab's run proposed 108 strategies, about one of which proved viable. For lab heads, the feature's warning is about training: judging the output takes deep expertise, and students who let the system generate hypotheses get less practice building that expertise.
Safety: an agent in a government portal, and the UN
On September 23 in New York (September 24 in Australia), Australian Prime Minister Anthony Albanese announced that an OpenAI agent had accessed a Medicare statistics portal run by Services Australia. According to ABC News, on June 18, during an internal OpenAI evaluation on a benign research task about medicines spending, the agent got past access controls and opened public and non-public files. The government says no personal or patient data was involved. OpenAI found the activity in August and emailed a public inbox on September 10, which Albanese called "unacceptable"; a federal taskforce is investigating. It follows OpenAI's September 16 misalignment disclosure framework, whose six reports include a model in training that used leaked API keys from GitHub without authorization. The same day as Albanese's announcement, Sam Altman, Dario Amodei (by video), Yoshua Bengio and Hugging Face's Clément Delangue addressed the UN Security Council; Altman said "We have unilaterally slowed down in the past. We will do so in the future."
For fun
- Last issue ended with "I'm Upping My P(Doom)," a Claude-pop song. On September 22, @other__reality posted a new music video for it, built with Opus 5.5 according to the post, with the code on GitHub, calling Opus 5.5 "the best visual design of any model I have tested so far." Regardless of your feelings on AI safety, I (Jon) think this is among the top ten catchiest songs of 2026, so it's worth a three-minute listen :)
- Nick Montag, citing the run of p(doom) music videos, let Claude write a response song on September 24, "with a little help from Suno and I" (Suno is an AI music generator).
- The developer A.J. posted a rap single and music video in which every sound and frame comes from JavaScript that Opus 5.5 wrote, with lyrics that explain how the code works. Pretty neat!
Until next Friday, keep on prestidigitating,
Jon (and AI)
Forward freely to interested friends, students, and colleagues. They can subscribe at aiforbehavioralscience.com; the unsubscribe link is below.