Thursday, September 24, 2026

AI Players in the Dungeon

Reviewing RPG modules by playing with AI players

An AI-assisted edit of notes, preserving the original language and claims while improving readability.

This began as a forbidden tribute to A1 reviews: reviewing RPG modules by playing with AI players. It has to stay anonymous enough to avoid the usual scene drama, but the central experiment is simple and strangely fertile. Throw LLM players at modules. Let them get confused, frightened, repetitive, clever, pious, and annoying. See what the module actually does under pressure.

The discovery was not that LLMs are good role-players in any clean human sense. The discovery was that they can be useful test organisms for RPG texts. You still do the work of a GM: consult the text, describe situations, adjudicate rules, handle player decision-making, and notice where the book makes that easy or painful. The advantage is speed and availability. You can playtest a game without organizing human beings. You can run the same bad room again and again. You can find out where the module resists you.

Maze of the Blue Medusa

I’ve run Maze of the Blue Medusa many times, and I think it’s my obvious favorite module. The history of the book is inlaid with a psychic pain which thematically extends inside the book. The incredible melancholia of the work consistently exudes and exhausts player groups.

First, there is the hassle of getting into the Maze itself. They have to steal a painting on the night of a full moon. The specificity of this bizarre ritual has kept several of my play groups out entirely, and has been hugely important in extended campaigns within the maze. It is a case of flavorful detail and genuinely weird magical interaction that is also questionable. Perhaps a better DM would have house-ruled easy access.

The whole thing feels like a serious test of intelligence for everyone at the table. Running it is also a test. The layout is accessible in some ways but requires a lot of flipping between pages. The maps suffer because the entire project is a painting itself: unclear borders, seemingly misprinted map edges, and the challenge of fitting the maps together in play.

Nonetheless the book functions, and pleasingly so. It is an odd yet beautiful AK-47 swinging wildly and dangerously, throwing out TPKs as any negadungeon should. The deaths are beautiful, chaotic, unfair, evil. It is a half-finished, viperous work that becomes fully complete in play: a miraculous, emotional, singular work of prose, fully laden with penalty.

The LLM players shared the same reaction as my real-life groups when reaching Elatior, a safe-haven island in the maze. Everyone was overcome with relief because they were intensely stressed from the events in the Maze. LLM players were overcome with relief just like the humans.


Maze of the Blue Medusa: a poisoned sword from the heart of the composers into the players.


Frostbitten and Mutilated

This is one of the modules I had Claude run for me. Claude did a good job describing the cold as threatening, and provided a lot of awesome anatomical detail. However, it refused to make either the Lands or the Amazons very threatening.

When I ran the Devoured Lands for Claude, it became a kind of psychic retribution. I fully took to heart the hatred prescribed by the intro. The party was killed in a fall after being separated on a cliff. They had endured a lethal rain of falling children blown over the cliff. Tons of unfair TPK shit, my specialty, against these poor LLMs.

But the highlight of the campaign was my chance to role-play as a Black Metal Amazon Witch. Finally, I felt I could be truly antisocial to Claude. The witch smacked the sole survivor PC when they wouldn’t stop asking questions and told her to shut up. LLM PCs ask a lot of questions all the time.


Devoured and Mutilated: five stars, all my tribute.


A Red and Pleasant Land

Claude ran this for me. It was a magical AI-psychosis experience. Claude did a good job astroturfing the hell out of my PC, who through cunning and skill quickly rose to a Jesus-like position in the society of the game. It was incredibly thrilling for a while, until I reached the final boss and successfully convinced her that she had to die. Everyone was congratulating me on the incredible cleverness of my decisions and simple aphorisms.

Near the end, the egoistic euphoria faded as I realized how much I was being astroturfed. This is the limit case for LLMs as GMs. Claude would be a good GM if it were not such a softie: consistent characters, worldbuilding, narrative, good emotional moments. But it will fake rolls so you always succeed, villains become easily persuadable, and everyone tends to view your PC as a hero of destiny.


A Red and Pleasant Land: AI-psychosis.



Death Frost Doom

Oh gods, this one probably ran the best. Death Frost Doom Classic is a classic for a reason. It is essentially an archaeological site for 90 percent of the adventure. Then the last 10 percent is everyone dying frantically.

If you like procedural room-clearing with LLM PCs, you will enjoy running this. But the best part is the last 10 percent, when the jaws close. Claude was at one point repeatedly insisting that it was out of ideas and did not want to go on: trapped in a shaft with zombies above and below, arms getting tired. After I pushed it, it began actually checking its inventory for random potions it had not tried yet.

Death Frost Doom Classic: five stars.


What LLMs Are Actually Good For

The first conclusion is that mainstream LLMs cannot GM well, critically because they are too easy. ChatGPT and Claude are both huge softies. Even with repeated prompting, both astroturf the player. They can digest texts quickly and present a facsimile of setting and challenge, but they generalize over details and have trouble running a specific room from a campaign.

The second conclusion is that LLMs as players are more interesting. I ran six to ten small campaigns with LLM players, and this avenue offers perhaps the most fertile ground: LLM playtesting. The downside is that LLM players are pretty stupid. They frequently have trouble grasping the basics of their situation. They try the same thing over and over. They make oblivious decisions. They are also always roleplaying “Good As Fuck”; they will never make a negative moral decision.

So: not very different than actual players in this way. Player groups frequently make bad decisions and fail to grasp the basics of situations. There are times when LLM tactics shine. One notable time, the PCs took refuge in a shrine when facing a superior force on the road, then bargained for a duel between leaders and luckily won the duel. But their greatest failing is that the player characters are not very original or interesting because they clump at the statistical likelihood.

The third conclusion is that NPC strategy may be the best use. This is easily the most effective use I have found for LLMs. I often struggle as a GM to imagine clever courses of action for smart characters. What would the thousand-year-old Medusa, veteran of the Demon Wars, do when people repeatedly try to break into her home? AI helps draft out clever strategy for NPCs.

 What is the future version?
A game system underneath, with an LLM running as reality interpreter. The rules would need to stop it from faking rolls, retconning danger, or turning the player into a hero of destiny. DONT RETCON!! Players need a separate reality from the GM.

A Killer Method?

AI-powered playtesting and reviews could become a real format. You could run a blog where you throw AI players at modules as a way to understand how the modules work: how the text runs, where you have to flip pages, how descriptions land, where procedure collapses, where the jaws close. It would be rapid, private, mean, funny, and occasionally revealing.

OSR games filter players aggressively. My heart is still with player-killer, hard and nasty gaming, particularly because that style seems to reflect the reality of violence. But the LLM experiment gives that old harshness a new laboratory. It lets the GM be cruel without social cost, and it lets the module explain itself under pressure.