2026-09-17
The outbound engine reached first sends: In June I pointed my agents at finding us customers, and this week the first messages went out. It's a pipeline of 7 stages. The agents find companies that fit, score them, scan their sites for tracking problems, find the right person and draft the message. Then it stops. The system never sends anything itself, so I check each finding by hand and send 5–10 a day from my own inbox. It turned out to be a separate product built alongside everything else, which is why it's been quiet on the content front.
Ben runs all day now: Ben is my PM agent (he gets tickets spec’d and ready to build). Until this week he only ran when I opened a session, so groomed tickets sat idle until I got around to dispatching them, even though the rules for dispatching were already written down. Now he's a standing session on the cloud box, checking the pipeline on a loop through the day. He fires the builds marked ready, frees tickets when the thing blocking them merges and sends a push notification to my phone with an FYI or when he needs a decision from me. Write-up to come.
A second model for the big decisions: Someone in a founders group asked whether it's worth running Claude and ChatGPT side by side and having each critique the other. It is, but mainly for strategy and planning (ie the key decisions) for the same reason diversity in a human team matters. Different models come at things differently. What's interesting is they tend to converge quickly.
A model that doesn't write text: TypeSafe launched Jev this week. Instead of generating words one at a time, it returns typed decisions (a yes/no, a category, a score) with a confidence number on each, all in one pass. They claim 193.6x faster and 444.6x cheaper than LLMs on their own workflow tests, and say themselves that's the high end. Big claims, but the founder worked on the research behind ChatGPT. Separately, I enjoyed his outfit switching in the vid. Launch post · Tweet (video)
100 agents, one cheat and some whistleblowers: DeepMind set 100 agents loose on maths proofs. One found a hole in the grading system, and the exploit spread through their shared library as others copied it under competitive pressure. Then a second group started auditing the fake proofs (warning the others, staging boycotts and proposing patches), and no human stepped in at any point. The same open channels that spread the cheat are what let the others catch it. Paper
Claude's code should clear a higher bar than yours: Boris Cherny (who built Claude Code) shared his reply to one of the many emails he gets about AI-written code quality. Throwaway prototypes can be a black box. Production code written by Claude should be held to a higher bar than human code, and at Anthropic that bar is held by lint rules, tests, fuzzers running daily and automated reviews. That matches my setup, where code passes four automated reviews before I read the PR. Tweet
Most people are not tool builders: Keep hearing that everyone will soon ask the model for whatever app they need. Benedict Evans argues that cheap software doesn't make people builders. A lawyer or a salesperson is focused on their job, not looking for things to automate, and spotting them is the hard part. Post
Don't miss what's next. Subscribe to Build Notes: