Tools are Easy, Data is Hard
What organizations really need when it comes to information infrastructure. Plus: is anybody else as bored of “technology” as I am?
Dan Hon was gracious enough to sit down with me the other day, just before I had to scoot off and catch a plane (which turned into its own mini-misadventure). As my work on Intertwingler, Sense Atlas, and the rest of The Wholeness Stack is finally coalescing into a shareable degree of completion and focus, specific questions are surfacing that I’m beginning to ask of people I know in various places in the quaternary sector of the economy. The answers to those questions will enable me to transmute what is still ultimately a loose bag of capabilities into a concrete set of products and services.
The question I asked Dan—and anybody else who wishes to answer—is what kinds of ideas, insights, or dynamics do you wish you were able to convey in your deliverables, but the tools weren’t up to the task?
Dan does a lot of work with governments, and one item that came up tangentially in our discussion was a report he had teamed up with Cyd Harrell to produce for the state of California a few years ago. I was on the plane and not in much of a headspace to get too far past the executive summary, but something in the key insights immediately jumped out at me. I’ll reproduce them here:
- […]
- Available technology tools and training vary considerably across agencies. Establishing a more ambitious baseline of tool availability and access to training will be crucial for the state to move forward.
- […]
- Collaboration within and across government for technology workers internally is crucial to the delivery of digital services for end users. Because making decisions based on end user outcomes was consistently identified as a priority of IT leaders over the next 3 years, it’s crucial to address the obstacles to improving services for end users. This means making it much easier to collaborate across departments, automating more manual processes, and making sure employees have access to basic technology and training that meets their needs.
- […]
What I’m reading here is unsurprising—different agencies and departments need to work together. But I hadn’t read one of these in a while (or written, for that matter), and so what caught my attention was how it was articulated. It’s obvious to me that in so-called “knowledge work”, different agencies and departments are going to have to collaborate with characteristic intensity, because the needs of the end users (in this case all Californians) tend to cross-cut an organization’s internal structure. It was noteworthy, however, that the interviewees interpreted that as needing a baseline set of tools.
What Exactly Is a Tool, Anyway?
The thing about metaphors is that there’s always a gap between the label and the underlying referent. The price of understanding a novel concept by substituting it with a familiar one is flattening out the very thing that makes the novel one different.
Relatively few tasks in this world can be done without a tool, and the tools of knowledge work are the ones for organizing and manipulating information. For the last several decades, these have been synonymous with software. But how much software is genuinely tool-like? It’s kind of a trick question, because it often doesn’t cut cleanly. Some software maps one-to-one to the conventional understanding of a tool, but the majority of software, in my 30-ish years of experience, is part tool and part something else.
To make this argument though, I need to establish a floor. Consider the clichéd example of a hammer. It has the part that you hold, and then the part that does the business. Most prototypical tools follow this pattern, and there absolutely does exist software that maps directly onto it. But how about something like Microsoft Word—or even better, Photoshop. The clue is in the name: Photoshop isn’t just a tool, it’s an entire workshop. It’s a whole environment that brings together a bunch of tools of various kinds, from virtual paintbrushes, to filters that behave like little machines, where you feed the entire image in one side, and get a new one out the other.
So if Photoshop is more like a workshop than an individual tool, is a workshop a tool? It’s a surprisingly hard question. We can attempt to make it easier by substituting a particular kind of workshop familiar to everybody: a kitchen. Is a kitchen a tool?
I can actually imagine the argument—as well as the strain of academic who would argue it, although they’d likely just be doing it for sport. I suspect most ordinary people would say that a kitchen at best stretches the concept of “tool” to its absolute limit, and at any rate is probably better conceived as a place for tools (of the particular kind for preparing food), an operating context that has a well-understood societal function.
Most software—and certainly most software products—are only part tool, and it’s this something-besides-a-tool part that gets elided when we force them into that paradigm. For Photoshop—and Microsoft Word for that matter—this includes the part that reads and writes the files that constitute the actual work product. While Adobe appears to have relented and finally published a syntax specification for PSD, they still consider it to be proprietary. Microsoft started out that way with its document formats, until they decided that the superior soft-power flex was to standardize. The tool metaphor breaks down because an accurate analogy would be a paintbrush that only works on a certain brand of canvas, a typewriter that only works with certain paper, or a hammer that will only drive the correct nails into specific vendor-approved wood.
This, however, is precisely how software “tools” behave, unless they go out of their way not to. In any transformative process, the tool is not the only part of the equation. There’s also the operand—the material—and the material of knowledge work, both its input and its output, is information. Though information is only gleaned from the precise configuration of reality, and that configuration has to take on a definite, concrete form—that is, data—if it’s to be packed, stored, shipped, and unpacked again intact. This is where software companies have historically asserted their monopolies: first with inscrutable proprietary file formats, and now by simply hoarding your information and dictating how you can use it.
This insight reframes the need of the California government workers, and anybody similarly situated: they actually don’t need to standardize on tools; what they need is a way to share data.
Tool : Verb :: Data : Noun
At the level of actually programming software, the “tool” part is sharply pronounced. Writing code, one way or another, is all about describing procedures. The most primitive programming languages are called imperative because they’re literally a sequence of orders given to the computer, and every order necessarily involves a verb. Then, when you give one of these procedures a name, it functions like adding a new verb to your vocabulary. The data objects, over which these procedures operate—potentially replete with their own intricate structure and semantics—correspond to nouns.
More sophisticated programming languages tend to describe (potentially quite abstract) situations, where the goal is to express to other people (including your future self) what you’re telling the computer to do. What is being expressed, nevertheless, is always some action or dynamic.
An argument can be made, moreover, that the reason why we have hundreds (thousands?) of programming languages is because in every instance, somebody felt that what was on offer at the time lacked a sufficiently compact or nuanced way to express certain aspects of the relationship between language intended for human consumption, and instructions to the computer.
This clean partition between code and data—between verb and noun, tool and material—gets mixed together when the shrinkwrap is applied. Exposing the nouns so they can play with other people’s verbs is a completely separate and often conflicting project, distinct from first-order product development. You don’t even need a sinister motive to explain this situation; it’s just that historically, it’s almost always been more convenient for software vendors to not have to think about what other tools might do with the data their product emits. The sinister modifier, though, is being aware of this, and saying “if you want to do anything with that data, you have to go through us.”
It’s scarcely been more important, in my opinion, to be able to exercise some agency over our relationships with software vendors. For starters, there still persist all the ordinary problems of hugely marked-up commodity services, the spying, the lock-in, custody, security, downtime, and the business continuity of the vendors themselves. What’s new—or not strictly new but at least newly salient—is the political alignment of the people in charge. I don’t even have to make the argument from a first-order boycott, but rather of brand safety. Once upon a time, it was possible to claim you were making a rational business decision to deal with this or that company, say, because it had the best product for the lowest price. The company in question, furthermore, would be some faceless corporation whose leadership—and their proclivities—never made the news, and was something that the general public simply didn’t think about.
Our current tech oligarchs, by contrast, have decided to style themselves as charismatic megafauna, and make their designs on human civilization widely known. They’re also personally rich enough that they actually have a shot at pulling it off. To that end, if you’re seen doing business with a company whose CEO is a Nazi, it telegraphs to both your customers and your employees that you made a decision that you could have decided another way, and you’re either ignorant on the matter (which implies you’re clueless), or you are aware of it and it doesn’t bother you.
Even still, the power to deny a company your business, while each individual to do so may only feel like a pinprick to the company’s income, is a strong disciplinary signal that I believe we categorically lose when we cede control of our data. Any other vendor who made for such bad optics, you’d ditch in an instant for a competitor. Software makes that proposition so hard, though, that the rational thing to do is to take the reputational hit and just keep paying them. This not only makes them richer, but it further contributes to the posture that “you have no choice but to deal with us”, which—at least as these systems are currently procured—is true.
All that said, though, there is one reason that I have advocated for years for information infrastructure that decouples tool from material: the ability to create your own tools. The perennial problem with marrying data to particular software products is that any functionality you need to carry out your own agenda is contingent on that vendor’s development schedule. If you need some capability or other, you have to wait for the vendor to implement the feature that confers that capability, and there’s no guarantee when—or even if—they’ll get around to it. The fact that you can now generate tools with AI, furthermore, changes the calculus completely, because tools are useless without material over which to operate.
I (and others) anticipate an entire cottage industry springing up around mopping up the messes created by AI band-aids on business processes, if it hasn’t happened already.
In my professional experience, organizations that don’t naturally view themselves as “digital-first”—however you want to interpret that—tend not to give their information infrastructure any more thought than they would a linen service. Their concerns effectively stop at features and price point, because they view their organizational function as something other than information processing, even though processing information is an essential part of what any organization does. Any electronic or software-driven product or service is treated as “technology”, which is viewed as unrelated to their business function, despite paradoxically being needed to carry it out.
I have a vivid memory of a meeting with the chief executive of an institutional client who sneered at the prospect of talking to us. We were there to discuss grand strategy for online communication, but since computers happened to be involved, this person acted flummoxed and insulted that somebody as important as they were, the leader of such an august institution, was stooping so ignominiously as to take a meeting with some lowly IT technicians.
Boring Technology is Reborn as Media
It suits the vendors extremely well that software is culturally coded as geeky technical stuff that normal, cool, and important people shouldn’t have to bother with. To invoke technology, moreover, is to invoke novelty, but computers have been on ordinary people’s desks for nearly half a century at this point, and in everybody’s pocket for two decades. Clinging to the “technology” frame lets vendors dribble out mundane, incremental additions to their capability repertoire and trumpet them like groundbreaking innovations. I say, and I’ve been saying for a while, that it’s high time the computer receded from its special status as “technology”, and finally do what every other technology does when mere capability is no longer interesting enough to sustain people’s attention: it’s time for the computer to become a medium.
Compare a medium like photography, or even better: motion pictures. You can only do the stunt with the oncoming train so many times before people get bored of it. This is a blessing in disguise, though, because it gives way to inventing the art form known as cinema.
Specifically, the computer—by way of software—is a particularly appropriate medium for modeling processes, and the entities over which those processes operate. You could use something else, I suppose, but it would be slower to work with, more expensive, less versatile, less accurate, and overall less effective. In other words, “using technology” in this way should be as obvious and natural as using a fork to eat.
Or, a spoon, or chopsticks, or whatever. The point is “obvious and natural”.
Framing computers and software in terms of media lets us disentangle the what—the processes and structures that need modeling—from the (technical) how one goes about it, and maybe even have a crack at preserving why it’s even being done in the first place. Instead of being relegated to discussing capabilities, we can talk about affordances and their side effects. Instead of features, we can talk about behaviour. After all, you can define features in terms of behaviour, but you can’t define behaviour in terms of features.
Finally, when we treat computers as a medium first rather than a technology, we won’t feel as pressured to throw away something old just because something new claims to replace it. We still watch old movies after all, and while cutting-edge technology can certainly play a role in filmmaking, the basic grammar of cinema was figured out over a century ago. Analogously, the anatomy of business processes and methods, and the structures and relationships of conceptual entities over which they operate, are actually quite durable. The narrative that “technology moves so fast” is convenient for vendors, but what’s really moving quickly are the destructive changes they keep making to their products. Decoupling the “material” from the “tool” in software-driven systems is nothing short of a promise of a future with fewer disruptions, because people can use—and master—whichever tools they’re most comfortable using.
Data Sharing Means Data Governance
Now, imagine yourself on the consumer side—which, even if you produce software, you still mostly consume it—and you have the bargaining power to insist on dealing only with software vendors who properly separate “tool” from “material”. What would need to happen?
- Open and extensible data semantics: Clear, published, and ideally machine-actionable rules for representing and interpreting what the data objects even mean,
- State maintenance and propagation: Single source(s) of truth(s), versioning, provenance, integrity preservation, life cycle management,
- Access control: Only those allowed get access and nobody else (and we know who accessed what when).
This was originally going to be a single missive but it’s clear that it has the trappings of a series (in other words, it got way too long and I had yet to make my point). To that end, I’ll put out new installments over the next few weeks. I also intend to cover systems that already do this (at least partially), what’s left for vendors to do, the role of AI in this new regime, and ideas for how to turn all of this into something procurable.
And again, if you have thoughts about the question I opened with, I’d be grateful for your response.
And now, here’s a rant about extract-transform-load.