
We tried outsourcing our support to off-the-shelf software. Their pitch: generate perfect customer support experiences with AI. All their systems did was read our in-house notes and guess.
A user would write in asking if their July payment was recorded. The rented product would send back a generic article about our billing practices. The same one it sent to everyone. It didn't know who the customer was. Nor their history with our company. Nor the specifics of their billing setup. The bots weren't dumb. They just didn't know anything. Context problem.
Enterprise companies insist on shelling out millions of dollars for these AI solutions. Generic products nervously avoiding real problems. Frustrating customers and somehow creating more low-value work for employees. Most companies have yet to see real returns from AI:

When Genius Models Fail
We are way past the excuse of intelligence. These models have performance comparable to a PhD. 91% on GPQA Diamond, against a 65% human-expert baseline. They are solving open math and computer science problems. Industry-ignorant brainiacs. Brainpower can't be the bottleneck.
We improved our own support metrics, reducing resolution time by 88% and replies to resolve by 46%. All from a custom harness. We struggled for over a year to figure out how to make AI work. Solutions working 10% of the time. Answers focusing on a different question than the one asked. We rebuilt the retrieval layer four times.
An LLM is a PhD hire on their first day in the field. Brilliant. Can spin a thesis with the best of them. Pointed in the wrong direction, they will hurt your business. Not because they can't help, but because you're asking them to solve PhD-level problems with pre-school level tooling. Botching agent onboarding is an organizational failure.
Context and tooling become the new bottleneck.
I've been tracking where our tools fail. We logged over 2,000 errors across 57,000 actions since March. 73% trace to one tool calling our database wrong1. The models can't own these failures. The fault is mine, for neglecting to provide the proper context. Context problem.
A Model for Value
If context is the scarce input, the question becomes who owns it.
Each AI Lab is the alternative to the next. Capabilities converge, prices collapse, and the lead changes hands every few weeks. Your context has no alternative. Nobody else holds your customers' history, your error logs, your post-mortems.

Domain, system, and individual knowledge is boutique. Universal tools lose usefulness the farther they travel from your articles of incorporation. Real value must be secured at the company level. Responding to customers in their own vocabulary. Aggregating inputs for your company's strategy. Only you can do either.
In July a user wrote in, "your program will not let me enter the show." Our harness spent six minutes and eighty cents identifying:
a duplicate piece she couldn't delete
a validation error that wouldn't clear
A rented bot would have sent her an article about uploading files with a FAQ at the bottom. Our harness pulled her entire history and went hunting. Her logs, error history, the module she was interacting with, all of it. Sentry was silent for her account, so it knew there was a client-side error, not a crash. By the time our support agent opened her ticket, it had a code change ready to deploy and caught a data bug we would not have seen otherwise. Context resolved:

These models will not just "figure it out." Integration costs will collapse once the model crosses a threshold. Yet they will never touch zero. Your business will keep generating context only you hold, and building your harness will not be a one-time task. Scaffolding tailored to you and the terrible problems you wish to conquer. Doing it now compounds institutional knowledge and earned secrets. As the models get better, so does the harness.
Our inbox stays clear. Who each customer is, what they bought, where they got stuck: on file before anyone opens the ticket. When someone asks about their July payment, the harness pulls their account and the specific charges.
Reducing resolution time by 88% matters less than what it bought. Now, we are free to run at our own terrible problems.
Thank you to Andrew, Justin, and Harrison for contributing and reading drafts.
Footnotes
LLMs provide reams of data on what they are doing. It is how we found the 73%, in Sentry logs on our own agents: review the agent threads, fix the top offender, repeat. Basic blocking and tackling, fast feedback cycles, impressive gains. ↩
You just read issue #15 of Myles Marino. You can also browse the full archives of this newsletter.