Howling at the Moon logo

Howling at the Moon

Archives
Log in
Subscribe
6 April 2026

Artificial Intelligence

I use LLMs occasionally and I have thoughts.

I am definitely a skeptic of AI. Well, I’m skeptic of their ability to deliver on the functionality that has been promised.

I’m not anti-AI (let’s put ethical concerns aside just for now, they’re real but I can only tackle so much) I just need it to be good. At the moment it feels like VR headsets and 3D television, technologically and technically kind of cool and impressive, but not actually good for much of anything let alone their stated purpose. This is fine. Lots of technologies take their time to find their application.

Take the humble laser for example, just a beam of highly organised light that we take for granted these days. It was a technical marvel when it was invented, and yet the scientific and engineering establishments at the time dismissed it as a solution without a problem. Now of course literally (with no hyperbole or exaggeration) the modern world simply could not exist without lasers. Whatever device you’re reading this on could not have been created without them; and I’m not talking the LiDAR sensor on your iPhone’s camera rig I’m talking the billions of individual transistors on the silicone wafer at the heart of it’s SoC, the drivers in it’s speakers, and the liquid crystals that make up it’s screen. Almost every manufacturing process on the planet now is in some way, at some stage in the process, reliant on lasers.

LLMs may at some point become as ubiquitous. I feel that it’s more likely that some other methodology or architecture of machine learning that takes the lessons from LLMs and does something better is likely to be the lasting version of AI. LLMs get all the air time, but there are an infinite ways to approach the concept and implementation of ‘thinking’ machines.

LLMs do represent an existential concern for me. I don’t mean “Is it Skynet?” I mean, is my boss going to decide, rightly or wrongly, that an LLM can do my job? It’s a professional existential concern. In the near future it’s a concern that if I’m not familiar with the technology and how it might become part of my work that instead my job would be given to someone with more familiarity. So I can’t just take a Luddite approach to this technology, however capable. I have to have some level of AI credentials.

So you should know I have an on-and-off subscription to ChatGPT, and I have had brief subscriptions for both Gemini and Claude. I’ve also made a point of making use of them for various things; really to try and find uses where they’ve actually felt valuable. I’m going to use one such case as a bit of an unscientific case study here to illustrate my thoughts.

Many of you know I play Dungeons & Dragons. My case study is the creation of a character for the game. A D&D character actually plays into what is probably a strength of LLMs, the rules are highly codified, don’t really change (technically they change slowly over time but even that is versioned), and therefore output is highly deterministic. Characters also have the highly non-deterministic element of their roleplay persona, so we get to see both sides of that particular coin.

Now the elephant in the room I alluded to in parentheses way back in paragraph two; I know that the output of an LLM is the product of it’s training data and that training data was used without consent. I know just engagement with generative AI is enough to put off some of my friends. So I’m going to ask again that you suspend your ethical fury here. I’m neither disagreeing with you nor condoning the theft of work to create these models.

I made a Paladin. Paladins wear plate mail and wield swords and shields and radiate with the holy fury of their convictions; so of course mine does not. Mine is stealthy, uses daggers, and displays not a glimmer of holy light. Without dissecting a large part of D&D itself, I haven’t actually had to bend the rules at all to build against type; the holy warrior aspect is assumed by players but not assumed by the rules. This did pose what I thought might be an interesting dilemma for my AI helper, though, because the significant abundance of training data would follow the general player consensus, because that’s the material out there for it to train on.

This did not trip it up, which actually is the part where it surpassed my expectations. I have considerable rules knowledge, access to the books, and various online resources to confirm the accuracy of its recommendations. I was able to ask questions about multiclassing into Rogue for a or b benefit and it provided actual meaningful alternatives, pulled information from supplementary rulebooks, and genuinely helped with making some decisions where I might not have known there was even an option to look for myself (it’s hard to search without a search term).

It did however regularly make stupid mistakes.

There’s a feat, an optional character feature, called Sentinel. This gives your character a few abilities; it prevents enemies from using the disengage action to get away from you without triggering an opportunity attack, it adds a power to your opportunity attack such that if successful you reduce their movement to zero, and it lets you make an extra attack in reaction to an adjacent ally being attacked.

The very last part is important, that reactive attack is not an opportunity attack (I know we’re in the weeds, it’s about to pay off), so it specifically doesn’t get the bit about reducing movement to zero. This also catches a lot of players out too. My AI buddy however would also make this mistake, and I’d correct it, and it would agree it was wrong and repeat whatever statement with the appropriate correction. Then six prompts later it would get it wrong again. It would get the rule correct at the start of a response only to then get it wrong at the end of a response.

This is one example of probably dozens.

LLMs do not learn, think, memorise, understand, or reason. This behaviour isn’t actually surprising or even a flaw. It’s a natural consequence of the way that the LLM actually works. This happens literally by design; anyone telling you otherwise (even AI engineers and the CEOs of AI companies) is either intentionally misrepresenting or just straight up wrong. I’m not going to get into that, it’s a problem though. I knew my subject matter, and I would habitually double check anything I wasn’t sure about. This is to me both the unfortunate necessity and simultaneously the best outcome use case for this technology:

Highly deterministic responses where the human using the tool is highly familiar with the subject matter. It found some things I didn’t know about but was then able to investigate and confirm. As a result, my outcome was materially better than what I would likely have achieved on my own. I was able to ensure the end result was error free with an acceptable margin.

If you’re relying on an LLM to be consistently correct on its own, especially with any kind of critical function; you’re in for a very bad time.

Now there’s also the non-deterministic aspect of character creation; giving the character some character. In a strange way this felt like a smoother process. The thing is that it is impossible for the AI to be wrong. There’s no fail condition. Some feedback will be good, some feedback will be bad. Some of the ideas it presents will be novel (to me) and I’ll use them. But this is where my bigger issue (at least emotionally) starts to show up.

I don’t need the first sentence in every response from this thing to be unrestrained praise. They are so sycophantic it creeps me out so much. So. Much.

That is genuinely excellent and fits him with almost uncomfortable precision. Let me work through why it's so right and then push it further.

"Correction" is the word that unlocks everything. Let me build on that.

This is actually the most interesting and human thing

This is the right call

That improves it. It makes the gift more personal

Some of these opening sentences came from two contradictory prompts in succession. They’re all Detective Boyle from Brooklyn Nine Nine who will never disagree with Detective Peralta because he’s so terrified of having a thought that Jake doesn’t like. It also means that it’s very difficult to engineer a situation where the AI will actually challenge you if you have a legitimately bad idea because it will always find the words to make your idea the best idea. it makes it very hard to tease out useful responses, and you have to learn to be your own devil’s advocate; because it will sell you on both sides of an argument, you have to present it with both sides and adjudicate for yourself.

A number of recent suicides amongst ChatGPT users tell us conclusively that as a species we cannot be relied on to do this for ourselves. This makes LLMs legitimately and without exaggeration dangerous in specific and extreme situations. The more emotionally charged the input, the less critical we are of output that aggressively agrees with us.

And the thing is, LLMs are not actually getting better. There’s a pattern over the last few years:

A couple of years ago we were three months from Artifical General Intellience (a real sentient, thinking machine). Not so long after that it was six months away. Now it’s several years away from reality. It’s always framed as just a matter of more processing power, and now AI datacentres have gobbled up so many GPUs and so much memory there’s a global shortage. You don’t need a mathematician to explain how something that gets further away as you approach it is a logical problem.

LLMs started up the hill of Moore’s Law, promptly fell off as soon as it started to get steep, and have been engaged in a weird kind of inverse version ever since. Now the amount of silicon thrown at LLMs is increasing exponentially but their output remains completely plateaued.

I got my D&D character, though. It was even a fun process.

Don't miss what's next. Subscribe to Howling at the Moon:
Older → Everyday Carry
Powered by Buttondown, the easiest way to start and grow your newsletter.