Collin's Thoughts logo

Collin's Thoughts

Archives
Log in
Subscribe
August 25, 2026

Prioritization is a trap

It’s a trap

We’ve all heard it: “What’s the priority?” Product teams turn into a tradeoff machine. You are balancing tech debt, live product support, and new features. Or maybe you have three products: one with a revenue opportunity, one with angry customers, and one that is starting to fall over under load.

Then comes the call. Meetings pop up on the calendar - “quick connect?” - a roadmap appears. The product director, product manager, perhaps a program manager and tech lead all get involved. Someone runs the inputs through a framework and the output is high, medium, low. Or worse, they are just spitballing.

That ritual made more sense when a team could only carry one thing at a time. After all, you only had so many hands on keyboards.

I think we should do less of it.

Something that once took a sprint to prototype can now take a few days, and sometimes a few hours. Agents and automation do not make production work free, but they do let you explore more than one path before you declare the other paths dead. Run the work in parallel. Stay in the loop when judgment is required. Build enough to learn.

Don’t hear what I’m not saying, prioritization still earns its keep. You need to choose clients, set quarterly direction, and make real tradeoffs when the same people or budget cannot do two things at once. But a lot of cross-business prioritization is really a team design problem.

Move people into durable teams that own a surface area long enough to make it better -- fewer support issues, more reliable systems, lower run cost, faster delivery. Then make the big prioritization decision through resourcing. Grow a team. Split one. Shrink one. Close a product. Those are clearer decisions than asking a room full of people to rank five things that all need doing.

If you have five good ideas, I would rather build a thin version of three than spend a month proving why only one deserves to exist. Spend more time getting faster at building. Or just spend the time building the damn thing.

Have models plateaued?

A lot of people had a “this actually works” moment around the leap in model quality last fall. The systems have continued getting better since then. I pulled the last three generations of leading models from DeepSWE, and it is hard to argue that the benchmark line has stopped moving. DeepSWE long-horizon coding-agent results, grouped by three recent model generations.

Still, the visible gain feels flatter in day-to-day work. A benchmark improvement can be real without turning into the same percentage improvement in a team’s velocity. Once a model is competent enough to finish a task, the next constraints show up: context, tool use, latency, cost, integration, and whether somebody can review the output with confidence.

I’m watching utility now more than another leaderboard point. A model that gets a little less clever but returns usable work twice as fast, or for a lot less money, can change how a team operates (side note: I am much more bullish on Gemini 3.7 flash than others).

Is coding solved?

Every few weeks, someone says coding is solved. Yes, code generation is getting very good (I have only hand written a few lines of code in the last few months) but that does not settle the rest of software engineering.

Boris Cherny put the distinction more sharply on August 20: “Coding is solved, bugs are not yet solved. Fix incoming.”

Boris Cherny on X: “Coding is solved, bugs are not yet solved. Fix incoming.”

A generated diff still needs a clear problem definition, tests that cover the behavior you care about, review from somebody who understands the system, deployment discipline, and observability once it hits production. The Claude status page is a useful reminder that the tool itself is another production dependency. Completion is not a delivery contract.

Claude status history showing 90-day uptime for claude.ai, Console, API, Claude Code, and Cowork.

The bar is not “can a model produce code that looks right?” The bar is whether a team can ship, operate, and change the thing safely. We are much closer than we were a year ago. “Solved" is still doing a lot of work in that sentence, but admitting that doesn’t get investors excited for an IPO.

Financial trend I’m watching: power is the allocation problem

The AI trade has focused on chips. I think more attention is going to move upstream to delivered power.

Gartner projects global data-center electricity consumption will reach 565 TWh in 2026, up 26% year over year. It expects AI-optimized servers to account for 31% of that power use this year and to consume more power than conventional servers in 2027.

This is not a stock tip (or is it?) but is a constraint worth following. A company can have capital and an accelerator allocation, then still wait on a grid connection, transmission, or generation. The moment compute revenue depends on when capacity can be energized, the scarce asset is no longer only the chip and I don’t expect this problem to work itself out quickly.

One I have not tested: Grok

Grok (and Grok Bot) keeps coming up in conversations about model choice. I have not tried it yet, so I do not have a take to offer. If you have used it for real work, reply and tell me where it is materially better or worse than the other tools you already use.

Worth Reading

The free field guide to agent security — OWASP
OWASP’s State of Agentic AI Security and Governance v2.01 was published in June and is free to download. If you are deploying agents with real access, your security team is going to ask whether you have read it. Read it before that meeting.

The IBM and OpenAI announcement, from the primary source — IBM
IBM lays out the three workflow areas: legacy operations, application modernization and product development, and cybersecurity with AI-risk management. It also joins OpenAI’s Elite partner tier. Read it as a preview of the consultancy pitch your board will hear soon.

OpenRouter’s Series B post, three months before the reported $7B Stripe deal — OpenRouter The May post reported 25 trillion tokens a week, more than 8 million developers, and 400-plus models. The investor list included NVentures, Snowflake, Databricks, MongoDB, and ServiceNow. A lot of people saw this coming. Stripe moved first.

The bottom-up case for Anthropic’s profit — SemiAnalysis
SemiAnalysis models the financials by SKU, tier, and customer type rather than accepting the headline number. Useful if you want to know what “profitable” means for a company that is still spending heavily to grow capacity.

Quote of the Week

“He who thinks the idea owns nothing. He who executes the idea owns everything.” — M. J. DeMarco

Forwarded this email? Subscribe here.

— Collin

Don't miss what's next. Subscribe to Collin's Thoughts:
Older → Stop calling AI subscriptions subsidized
collinwilkins.com
LinkedIn
Powered by Buttondown, the easiest way to start and grow your newsletter.