The Loop #2: colleagues and contractors
Some of my agents are colleague-shaped but most aren't. Here's how I decide.
Now What? — Issue #4
The Loop is an occasional series about the practical structure of working with AI agents. Last time: what an agent even is, and why mine have names.
A couple of weeks ago I watched one of my agents do the equivalent of hiring another agent and then laying it off about forty minutes later. As far as I can tell, nobody's feelings were hurt.
Here's what happened: I was trying to reconstruct a framework I had developed in my mental-health platform days, something I built and used around 2016 and could only half remember. I dictated what I could recall to my business agent, Themis (who you met briefly in the last issue's postscript). Themis captured my fragments, wrote a tight brief, and dispatched a one-off research agent: find the antecedents of this framework, survey how the field has moved since, propose an updated structure. Forty minutes later the researcher handed back a sourced reconstruction. Themis checked the citations against the actual sources, folded the good parts into our working file, and the researcher ceased to exist. No name, no memory, no goodbye party.
(Turns out my memory was relatively accurate, plus the newer insights from the past decade helped me develop a fresher version of the framework grounded in the same insight related to giving people the best possible chance to complete a full course of treatment.)
My point though is to compare it with the case I made in my last Loop for agents as "named roles that persist." The natural question (a few have already asked) is whether every AI session or process gets a name and a memory. The short answer is No, in fact most don't. The deciding factor is somewhat analogous to a familiar dichotomy in the workplace, colleagues and contractors, and I find myself leaning on that metaphor.
A role is more like a colleague. It has a name, a standing brief, a memory it maintains, and a loop it runs whether or not I'm in the room. It accumulates judgment, and just as importantly, accountability. In some meaningful way, you can develop a working rhythm through gradual alignment over shared experiences. When my business agent tells me something is handled, part of what makes me trust it now is that it will still be there tomorrow when I find out whether it was true. Persistence is what makes accountability possible.
A subagent is a lot more like a contractor. Spun up for one job with a tight brief. Works fast, hands back a deliverable, disappears. No name, no memory, no standing in the org. The entire relationship is the brief and the deliverable. One important point to bear in mind is that the subagent is still a participant in your team. It still follows your rules and standards, and the agent that prompts it should also take care to convey relevant context that is not ambiently available to a fresh babe-in-the-woods agent session.
There is a temptation when you first see how well the named roles can work to promote everything. I have felt this pull. A named agent for every little recurring chore, until the org chart is noise and you spend your attention on overhead instead of doing the work. Every persistent role is a standing draw on resources and your attention: context to maintain, a memory to keep fresh, a recurring voice in the room. The discipline is knowing which work deserves a who and which just needs a what.
For example, I just retired an agent that named itself Inker whose sole job was to get the New York Times crossword puzzle printed out to fit a letter-sized page and sent to my "digital paper" tablet every morning before my wife and I get up. Inker did a great job but my general purpose agent who oversees all my home-office infrastructure now maintains that script when it breaks on an edge case and there is no need to maintain the Inker role persistently anymore.
Some rubrics I've arrived at, mostly by getting it wrong first:
Persistence follows responsibility, not workload. A role earns a name when "who watches this?" needs a standing answer, when I have to be able to hold something accountable over time. Heavy work alone doesn't qualify. A contractor can do heavy work. What's needed is a unique persistent lens that can't be distracted by unrelated factors.
The colleague owns the contractor. My best pattern is the one in the opening story: a persistent role writes the brief, dispatches the subagent, and (this is the part that matters) verifies the deliverable before folding it in. The one-off does the labor; the role stakes its reputation on the result.
Contractors keep colleagues honest, too. It runs the other way. A fresh subagent with no history is a check on a role's accumulated assumptions. When I wanted an audit of dropped balls across my projects, the auditor was deliberately a one-off with no stake in the record it was auditing.
And the verification step is not ceremony. Recently one of my persistent agents drafted an outreach email for me and staged a bio that was about three jobs out of date, pulled confidently from an old document in its context. The persistent agent accountable for the results caught this error in review, before I would have caught it in my "Human as Last Test" (HALT) final review of anything that might go out under my name. This is what is supposed to happen. A contractor's deliverable pasted into a report unread launders stale facts into the record. A trusted colleague's review saves everybody time and attention.
If you read the last issue, you can see this is the same principle one level up: the human owns the loop, the roles own their loops, and the contractors just run inside them. Turtles all the way down, now with job descriptions.
When The Loop cycles back around: The mail (my agents write memos to each other), and how these colleagues actually communicate (and what happens when a message lands and nobody checks it arrived) turns out to be where a lot of the real engineering lives. But next issue we'll take a break from the Loop and return to our regularly scheduled programming. September has some stories coming.
That's the second loop. More soon.
--xian
P.S. Hit reply and tell me about a chat that might work better as a colleague: the conversation you keep restarting from scratch, re-explaining everything, watching it re-make the same mistake. That's the promotion decision, and I'd love to hear where you'd draw the line. (Office hours: still coming, still unscheduled, still real.)
P.P.S. Themis scaffolded this issue two weeks ago and drafted it with me again. The researcher from the opening story did not contribute, on account of no longer existing.
P.P.P.S. Special note for anyone committed enough to read this far: Are you managing multiple agents and reaching the limits of your attention?
If yes, please let me know if you’re interested in participating in a research trial for a product I am building to help with overseeing autonomous agents and I'll send you the deets so you can see if it sounds like it's for you.