What AI is actually changing at my company
Six new models in September and the one habit worth changing, a test for anything sold as an "agent", what's shifting inside my own company, and a new homepage.
Hi there,
The last edition said ChatGPT had won me back and that Claude's Fable was still the best choice for hard thinking. Since then, Anthropic and OpenAI have released a run of new models, and one of them replaced Fable as my default. So this edition starts with what changed, and with the part of my advice that didn't change.
After that: a test for anything sold to you as an "agent", something I haven't written about much yet (what's actually shifting inside my own company), and a new front door for this site.
๐ง A month of new models, one habit worth changing
The short version of September: Claude Opus 5.5 is now the model I use for almost everything. It's as good as Fable on most work, beats it on a test of real-world tasks across 44 occupations, and costs 40% of Fable's price per token. On a subscription, the number that matters more: it uses up my allowance noticeably more slowly than Opus 5 did. Fable keeps a small edge on very big, long-running jobs, and I expect even that to shift with the next Fable release.
OpenAI answered the same day with GPT-6 Sol and Luna, at half the price of the models they replace or less. A week later it announced GPT-6.1 Sol at its developer conference, and Anthropic released Sonnet 5.5, which keeps Sonnet's lower price and lands two points behind Opus 5.5 on that same test of real-world work.
Sonnet 5.5 is a curious release: almost as good as Opus, and cheaper. I have a feeling I'll still stick with Opus in my setup and let it hand the smaller, clearly defined jobs to Sonnet working as its helper. GPT-6.1 Sol stays my alternative workhorse. It's probably a bit less capable than Opus but potentially cheaper to use, and it's a good second opinion when I want something checked by two AIs (that's what Don't Let Your AI Grade Its Own Homework is about).
So where does that leave ChatGPT, a few weeks after it won me back? Since Opus 5.5, I'm a total Opus fan again. But I work on a lot of different projects with different kinds of work, and I spread them across different setups: the right tool for each job.
I switch more often than most people should, because trying these tools is part of what this site is for. If you don't follow model names for fun, here's what I'd take from all this. On price, OpenAI and Anthropic are now close enough that the choice mostly comes down to preference. What matters far more than picking this month's winner is knowing your main tool really well: where it cuts corners, how it reacts when you push back, which tasks it quietly does badly. You only learn that by sticking with one main tool, and switching every few weeks throws it away. If your company picked Copilot or another tool for you, the same holds. Learning what you've got in depth beats envying what you don't have.
One habit carries over to every model released this year: describe the outcome and how you'd know it worked, and leave the steps to the model. Instead of this:
Read the attached survey results. Group the comments by theme. Count each theme. Summarise the top three.
try this:
I need to decide which two customer complaints to fix first this quarter. Here are the survey results. Done means: I can defend the choice to my boss in two minutes with numbers from the data, and you've told me where the data is too thin to be sure.
The first prompt gets you a model following your steps. The second lets it look for a better approach than the one you'd have written down, and the newer models are good at finding one. For when to pay for which tier, the updated Think Expensive, Execute Cheap has my current split.
๐ค Is it really an AI agent?
Every product has an agent now. I wrote a two-part Agent Field Guide for the moment someone puts one in front of you.
Part one, Is It Really an AI Agent?, comes down to one question: after the system sees the result of an action, who decides what happens next, fixed rules or the model? A tool that runs the same five steps every Monday is a workflow. That's often the better design, but it shouldn't be priced like an agent. The tip I'd keep from it: bring one awkward real case to the next vendor demo, the one where data is missing or two sources disagree. Ask what the system looked at, what it decided, what it changed, how it checked the result, and why it stopped. If they can only show you the case where everything goes right, you've watched a sales pitch.
Part two, Before You Build an AI Agent, Configure One, is for when the job really does need an agent. My advice is the opposite of how I built my own (rented server, several setups broken badly enough that only a full restore helped). Start with one weekly job in the AI tool you already have, three static files and no live access. Before anything else, write a ten-line job contract: inputs, allowed tools, what "done" means, what needs your approval, what happens on failure. If you can't fill in those lines, you're not ready to automate the job. Run it five times, keep the failures, and only then decide whether you need more.
๐ข What's actually shifting at my company
People keep asking me what actually changes when a company takes AI seriously. Here's what I see at ours, nine months in.
It stopped being my tool. Since early this year I've been building what I call our company AI operating system: a shared knowledge base about how we work, plus saved procedures ("skills") that anyone's AI can run, like turning a client call into a summary for our records. It has 17 skills now. I built nine of them. Colleagues built the other eight, all since June: interview kits for hiring, a decision memo, a summary skill for data-analysis sessions, and more. All of it runs inside our company's approved setup, under Swiss data-protection law. There was no mandate behind it. I shared some of my skills, explained how they work and how to create one, encouraged people to try, and made clear that experimenting here was low-risk, so they felt safe doing it.
Access isn't the bottleneck. Coaching is. Usage is uneven. The leadership team are the heavy users, and the talent and sales teams use it too. For everyone else, the hurdle is that people need someone to sit next to them and show them how it works on their own work, and there's never enough of that time. So I do it in smaller doses: continuous coaching, sharing results and how-tos as they come up, and now and then an "AI school" for a group. The lesson from those sessions: people don't need a feature tour, they need to see real work get done. And I send colleagues to this site. Seriously.
๐ A new front door for the site
The homepage got rebuilt. The old one dated from when this site was mainly a directory of AI tools: a headline, a row of star ratings, and then you were on your own. The new one starts with the two-minute quiz and six "sound familiar?" lines, each opening the guide I'd hand you first, and it lists every guide in the order I'd read them. If you've taken the quiz before, it welcomes you back with your level and the next skill to tick.

Claude Opus 5.5 did the whole redesign from two prompts: one to plan it, one to do the work. It proposed the page, built it in English, German and French, checked it in a browser at six screen widths, and ran a separate reviewer over its own copy that flagged claims the site couldn't back up. My corrections fit into one message: a line break, a more modest description of my talks, and one element I didn't want.
That only works because it didn't start cold. The repository already knows who this site is for, how I write, the words I never use, my exact job title and which claims need a source. It's the brilliant new hire again, except this one had been properly onboarded, which is why one prompt was enough. Have a look.
๐งช Two guides to send to a colleague
If someone on your team is just starting out, or keeps their distance from AI, these two are for them.
I rebuilt Getting Started with AI from scratch. It used to be a list of advice. Now it's six short exercises you work through in about an hour, each showing one behaviour you need to understand before you trust an AI with real work. Here's one you can do right now: open a new chat, ask for a specific fact from your own field that you can check, and tell it to answer from memory without searching. Then open another new chat and ask exactly the same question again. Compare the facts, not the wording. If they differ, at least one answer is wrong, and nothing in either answer tells you which. If they match, check anyway.
And on 1 October I published Is It Safe? What Happens to What You Type Into AI, for everyone whose first question about AI is about their data. It covers where your chats go, who can read them, whether they train the model, and what has actually gone wrong so far, followed by how I decide what to share. It ends with a 30-minute checklist: find out whether you're on a personal or a work account, make your own choice about training, read what the assistant remembers about you, and delete old shared links.
๐ Quick hits
The Orchestrator's High, updated after four more months of running agents. The new part I'd read: "Starting is free. Finishing isn't." More than once I had to ask an agent to remind me what my own project was about before I could continue.
Perplexity went from five stars to four, and my subscription is on hold. The research built into Claude, ChatGPT and Gemini is now good enough for most people. If you research for a living, the review says why it's still worth paying for.
New on the site: See What's Possible, working things AI built from a plain-English prompt, running in your browser. The first is a synthesizer you can play, with the full prompt to copy and an explanation of what the model had to work out that the prompt never mentioned.
Also new: a Speaking and teaching page, with the keynotes, guest talks, teaching and podcasts I've done so far. If you're planning an event or a team session on practical AI adoption, that's where to start.
One question this time: what's one thing AI has changed at your work in the last three months, for better or worse? Just hit reply. If enough of you answer, I'll share the patterns in a future edition, anonymised.
Best,
Fabian