Write Like Nobody's Reading logo

Write Like Nobody's Reading

Archives
About Paul
Log in
Subscribe
9 August 2026

Tools for Better Data, Now

Tools are easy now, data (has always) been hard

Bending Spoons’ acquisition of Airtable has set many a hare running. Through my work with the News Product Alliance, I know of several journalism/tech-related organisations who (at least at one point) defaulted to, and recommended the use of, Airtable to capture, structure and store all kinds of data. And now, many people are concerned by the acquisition, due to the company’s track record of, well, not always doing the best thing for the end user, let’s say.

The news sparked an exchange between Dan Hon and Dorian Taylor, on BlueSky, which caught my eye, and that I’m going to be riffing on this issue, because it is a useful jumping off point to unpack many of the things I think are fundamental to forging a better future with technology:

This ties back to the discussion around Airtable, because, as with many software-as-a-service products, although the tool/product is what is most appreciated (and is given the most attention), the thing that is almost always more valuable, is the data created and held within that tool.

This isn’t a simple rehashing of a decades’ old talking point, however. What I think Dorian is getting at, and what I think is an absolutely crucial point is - it’s the usability of the data that is most important. Having data (or rather, information, as it needs to be understandable to be useful), in a shape that it can be used, reused, flexibly, programmatically, is the main thing here.

The data can be anything - what’s important is the ability to do things with it. And not just use it as training data or fodder for data science. As I mentioned in the last issue, it’s the ‘point-at-ability’, the fact you can not only request it, but use it, that’s most interesting to me.

Information is possibility. Making that information findable, accessible, retrievable, understandable and programmatically usable is my (and I think, Dorian’s?) point.

This has always been the case, however. It’s just becoming even more apparent in the rush towards an apparent future of agents acting on people’s behalf.


Maybe it is the tools, after all

Further into Dorian’s thread, he references one of his brief video diaries, on the topic of ETL (Extract, Transform, Load - a standard practice in working with any kind of data or information in digital environments).

It’s worth unpacking these terms. Each of them has, in a multitude of systems, had code created to perform these tasks, and the craft and thought that has gone into these is worth pondering. Yet really, they should be considered activities that require human-computer interaction - and thus need as much care and attention as to the design of them as anything else.

Let’s start with Load, because to me, it’s the least interesting of the three. I think it misses the point, honestly. Yes, loading data into some form of digital system is necessary, but if we take a step back, ETL is literally just the process of communication - extracting from one mind, transforming into a media for transfer, and loading (or perhaps, receiving, interpreting and committing to memory) to someone, or something, else.

Indeed, I’d rather see more focus on what happens after the ‘L’. Maybe it should be ETP (Play), ETC (Create), or more abstractly, ETU (Use!) - not just usable by humans in the kinds of tools that Dorian and Dan were originally discussing (like Airtable), but by machines too - that’s the programmatically useful part.

Extract is usually thought of as a stage in the context of retrieving some data from an existing digital system. Although I absolutely balk at the use of this particular word in what I’m about to write, this stage of the process is as much about ‘extracting’ knowledge from our minds, as it is about requesting and retrieving data from a digital system. By which I mean to emphasise that we need to respect the impedence mismatch between the (still fuzzy, unknown) way in which our minds think, and the way computers require data to be shaped, in order to be maximally useful, and design to accomodate this.

Transform is, as Dorian says in his video, “where the artistry is”. And yet, this is where I’d argue that tools are still hard. Yes, it’s even easier (apparently) than ever to create tools that seemingly work and do the intended job, but as Dorian notes back in the BlueSky thread:

“…what still is a huge hazard, however, is the integrity of the data over which those tools operate.”

My take on this, though, is not just about the integrity of the data, but of the design of the tools that allow people to extract and capture the data itself. Ever since software such as iTunes went mainstream (and even before then, of course), making it second-nature to capture accurate, well structured (dare I say ‘meta') data has been a challenge.

This, to me, is still one of the great unsolved challenges of UX & product design, and one I simply can’t accept is one we should give up on, just because it’s hard or because we think humans are lazy. I also don’t think the answer is just more forms, or better form design.

In Rachel Coldicutt’s excellent piece on forms, she says:

No matter how much I love forms, it’s a pretty one-way relationship. A successfully completed form might give you access to something you didn’t have before, but an unsuccessfully completed form either gets you nowhere or gets you into trouble. There’s no incremental progress, just win or lose.

One reason for this is that the data schemas are rigid — even with new methods of input like conversational interfaces, the required information doesn’t flex. The database needs what the database needs.

… which brings us on to the other major thing I think we absolutely need to change if we are to progress. It’s also something that I think is especially relevant when it comes to structured data for any kind of storytelling, be that fictional, documentary, sporting or journalistic.

There is no ‘one model to rule them all’

There’s been a lot written about metadata over the years. A lot of it boils down to those who believe it’s worth putting time and effort into designing and crafting good metadata, and those who either believe it’s an unreachable utopia, or it’s a job we can hand off to our magic machines to statistically divine themselves, and therefore it’s not worth spending effort on. “Just search!” they say, and the computer will find it.

My first full time job was to maintain and update a single data model that was meant to describe everything a media company would do, from commissioning, through production and broadcast, to archiving - with the idea that this model would be something everyone could conform to (or at least translate to), in order to exchange information. But, just like the XKCD comic, now you have yet another standard that everyone has to adopt.

Cory Doctorow wrote a piece back in 2001, entitled Metacrap. In it, he eviscerates the idea of a ‘meta-utopia’, with seven reasons why it’d never be possible. Whilst he does conclude that there is value in metadata, and that it’s the idea of a world of perfectly accurate, well described and managed metadata which is the fallacy, I’ve always felt that these seven reasons are used as ammo against any attempt to spend effort on better information management at all. No immediate return on investment, and the idea of exponential benefit - that well managed information can have a multitude of uses - seems to put an end to most efforts. Be pragmatic, yes, but see these things as ecosystems, not as short-term projects.

There will never be a single, perfect way of describing something. And the whole beauty of human expression, of communication and understanding, for me, lies within that. Showing those different interpretations, laying them out, side by side, so we can compare, contrast, connect and adapt, feels like a crucial step we need to take in the forms of media we use to communicate. A ‘plural point of view’, as it were:

Wikipedia’s neutral point of view (NPOV) policy presupposes that it is possible to write from an objective perspective. We do not strive to establish a “true” account of events, explanation of practices, or definition of terms. We do not believe this exists in fandom.

Our intent with Fanlore is to create a space where fans can tell their own stories from their own perspective. The plural point of view policy asks fans to recognize the point of view from which they tell the story, and invites those with other, differing points of view to tell their own stories about the same events, places, concepts and people.

A similar fundamental idea is that of multi-model agnosticism, or the recognition that because every model we create to describe aspects of the world can never be a full recreation of the true complexity of the real thing, we should be more adept at accepting we will need to use many models during our time.

As John Higgs puts it, in his excellent book on the seminal British music artists, the KLF:

“Multiple-model agnosticism, then, is a way out of postmodernism which doesn’t lead into the belief that, out of all the billions of people in the world, you are the only one who really gets it and everyone else are idiots. The problem is, however, that our models are too damned convincing, and it is a struggle to remember that they are models and not reality… They are just models, after all, and models are significantly less detailed than what they represent. Reality itself is ablaze with infinite connections: every particle in the cosmos affects every other particle. It’s Too Much, it really is, and seeing reality in all its innate finery would be so overpowering that you’d be in no state to nip down the shops when you need a pint of milk.”

The point of it all was never the data format. It’s the philosophy - people should not only publish what they know (or want to communicate), but they should publish what they mean alongside that. Rather than everyone having to agree on a singular model (which would be lovely, if and when possible, but shouldn’t be assumed as a pre-requisite), the value comes in encouraging others to share their perspectives, effectively, and then connect and agree across ‘models’ where it makes sense. No lingua franca, no enterprise model, no service bus, it’s a networked world of information architecture.

Next time…

Interpretations, Perspectives and Provenance - or, the one where I tie it all back into my previous work in information architecture for narrative.

Not a new idea, and not contingent on the current wave of LLM-inspired AI. If anything, it relies even more on both good quality data, and trust mechanisms.

Don't miss what's next. Subscribe to Write Like Nobody's Reading:
Older → Networked Narratives
paulrissen.com
Powered by Buttondown, the easiest way to start and grow your newsletter.