AI-DnD · architecture notes
The Engine Room
An AI Dungeon clone is a chat wrapper until four things are true. The prompt is budgeted. The numbers are refereed. The story is a tree. Someone else’s JavaScript runs in a sandbox. This page shows how each one is built, and what it cost to learn.
0.1 What the thing is
You write a scenario, then play an open-ended text adventure where a language model narrates the world. You type “I open the door”, the model writes what happens next, and it remembers what came before.
Four subsystems carry the weight, and every hard problem in the codebase belongs to one of them:
- A context engine. The model has a finite input window. The app decides, every single turn, which pieces of the story get into the prompt and which get dropped.
- A world-state engine. The scenario declares stats (
hp,trust,day). The model proposes changes each turn; Python decides what actually sticks. - A story tree. The story is not a list. Any turn can hold more than one take, and writing below a take that isn’t live starts a branch that borrows every turn above the fork rather than copying it.
- A scripting sandbox. Real AI Dungeon JavaScript imports and runs, inside an embedded QuickJS interpreter.
It runs locally against Ollama for free, or hosted against any OpenAI-compatible endpoint.
0.2 The stack, and what each part is doing
| Piece | Job in this system |
|---|---|
| FastAPI | HTTP server. Routes, plus the SSE streaming. |
| SQLAlchemy | ORM. Adventure, Action, Memory are Python classes; attribute access becomes SELECTs — which is exactly how the egress bug happened. |
| SQLite / Postgres | One file on disk locally; Neon Postgres when hosted. Same code, two dialects. |
| React + Vite | The SPA. Built files are served by FastAPI, same origin, same port. |
| httpx | Calls the model endpoint, streaming. |
| tiktoken | Counts tokens, so budgeting is arithmetic rather than a guess. |
| QuickJS | Embeddable JS engine used as the user-script sandbox. |
Part 1 — The high-level design
What the boxes are, who talks to whom, and the two or three decisions that every later decision inherits.
1.1 One process, two things outside it
Production is a single Docker web service on Render. It serves both the API and the built SPA. Only two things live off-box: the database and the model endpoint.
This shape is not an accident. It is what makes the in-process turn lock and the in-memory rate limiter honest. It is also what the Known limitations list is measured against.
1.2 The domain model: template versus instance
User
├─ Scenario (the template) ── stat_schema, prompt, memory, author's note
│ └─ StoryCard, Script
└─ Adventure (the playthrough) ── world_state, script_state,
│ head_branch_id, head_depth,
│ persona_name, persona_pronouns, persona_desc
├─ Branch (one line of it) ── parent_branch_id, fork_depth, lineage, name
├─ Action (one node) ── branch_id, depth, parent_id, live,
│ text, context_snapshot, state_after
├─ StoryCard (its own copy)
├─ Memory (text, embedding, branch_id, depth, use_count)
└─ AdventureScript
The decision that shapes everything else: a scenario declares what stats exist. An adventure holds what they are right now. Creating an adventure copies the scenario’s story cards, scripts, and plot fields into it. Editing a scenario later never mutates a game in progress. This is the same reasoning as instantiating a class: shared definition, independent state.
The opt-in escape hatch is Update from scenario. GET /adventures/{id}/refresh returns a diff: per-field old and new values, card additions, updates and removals, and world-state paths added or removed. POST to the same path applies it under the turn lock. It never touches the opening action, the adventure’s title, its summary, or player-authored cards.
Re-copying scenario text would have re-injected literal ${Hero} placeholders. The answers were consumed once at creation and thrown away. And without story_cards.source_ref (card:<id> / npc:<key>, NULL for player-authored), there was no link back. A rename read as delete-plus-add, and player cards would have been clobbered. Two migrations were the price of one feature.
1.3 The turn pipeline
Everything between “player pressed a button” and “text is on screen”. Teal steps are user-script hooks — the points where someone else’s JavaScript gets to rewrite the turn.
- onInput hookUser JS may rewrite or block the input outright.
- store the player actionWritten before generation, so a failed turn still shows what you typed.
- retrieve memoriesEmbed the last 4 actions, cosine-rank the bank, take the top K.
- build_context()The budget allocator. Fixed sections first, elastic ones into what is left.
- onModelContext hookUser JS may rewrite the entire assembled prompt.
- snapshot the exact promptStored on the action. Powers Insights. Costs ~74 KB a turn: §3.1 makes that cheap to read, §3.5 cheap to store.
- provider.generate()One streamed call. Tokens forwarded to the browser as they arrive.
- onOutput hookLast chance for user JS to touch the text.
- extract + referee the state blockParse the fenced JSON delta, clamp it, strip it from the prose.
- save the nodeStamped with
state_afterandworld_state_after— the scoreboard as this turn leaves it. - fire-and-forget: summarize + embedBackground task, own DB session, never on the shared demo key.
Two choices are visible in that list before any detail. The prompt is snapshotted, not reconstructed. That makes prompt bugs findable, and it makes the database expensive. Every node also records the state it leaves behind. Rewinding to before a turn is a read of the node in front of it. That single move is why undo, retry, and a branch switch are the same restore.
1.4 Two personalities, one codebase
One env var switches the whole app. Not two builds — the differences are gated at each site.
| Local (default) | Hosted | |
|---|---|---|
| Users | One auto-created local user | Guest on first visit, optional account |
| Auth | None — no cookies, no login | Signed session cookie |
| Rate limits | Off | On |
| Row caps | Off | On |
/docs | On | Off |
| Provider | Whatever Settings points at | User’s key, or the shared demo key |
Someone running this on their own laptop should never be throttled by their own app. They should never see a login screen, and they should get the interactive API docs. A hosted deployment needs all four of those to be the opposite.
Guests upgrade in place. A visitor gets a guest User row on first load. Registering sets email and password_hash on that same row, so every adventure played as a guest survives with no re-parenting.
Guests idle past AIDND_GUEST_RETENTION_DAYS are swept by a single Core DELETE. The FK graph is ON DELETE CASCADE the whole way down. Without it, the ORM path would have pulled every action and memory into Python just to delete them.
1.5 Why there is no agent framework
Graph-based agent frameworks earn their complexity with branching, cyclic, multi-step control flow. Their path depends on what the model decides, with loops, tool calls, and persisted state between steps. This pipeline is a fixed linear sequence with exactly one model call. There is no routing decision, no tool selection, no loop.
There is a second, more specific reason. The budgeting logic is the product. Buffer-window and summary-memory abstractions are opinionated about how to fit history into a window. Here, the Insights panel exposes each context component, its token cost, and the trigger word that pulled it in. Assembly has to be explicit and inspectable.
Part 2 — The low-level design
The arithmetic, the schemas and the invariants. Each of these is a place where the obvious implementation is wrong, and the reason it is wrong was found by measurement or by a bug.
2.1 Context assembly is a budget problem
Say the budget is 8,000 tokens. A 200-turn adventure has far more story than that. The naive fix is to send the last N turns, but that breaks in both directions. N short exchanges waste the window. N long ones overflow it. Either way, it throws away the premise and the promise you made to the innkeeper.
So the prompt is split into two kinds of sections. Fixed sections are included whatever they cost. Elastic ones fit into what is left.
Fixed is about the budget, not the position. A fixed section always goes in. Where it sits in the prompt is a separate question, and §2.3 answers it. Several sections listed as fixed below sit after the history.
reserved = every fixed section + author's note + front memory
+ length hint + refusal note + emit reminder
available = max(256, context_token_budget - reserved)
cards spend up to available * 0.4
history spends available - cards_used, filling backwards from newest
The details that are decisions
- Cards are capped at 40% of the elastic budget. Story cards trigger on keyword match. A scene naming six things could pull six lore entries and leave no room for the story. The cap turns that failure into “some lore is missing” instead of “the model has no idea what just happened”. Cards that do not fit are still reported to Insights with
included: false. - History fills newest-first and stops. Old material is not lost. It has already been summarized into memories and the running summary, both of which are fixed sections and are included whatever they cost.
- If even the newest turn is over budget, it is hard-truncated, not dropped. A prompt with no story produces nonsense. A prompt with the tail of the last turn produces something.
- The author’s note is injected 3 actions from the end (
AUTHORS_NOTE_DEPTH = 3), not at the top. It is a steering control, so it goes where steering works best: recency. - The world-state reminder takes the very last slot. The full emit rule lives hundreds of tokens up in the system block. A one-line reminder occupies the position closest to where the model starts writing.
The state block is stripped from text before storage. Replayed history then showed the model twenty of its own past turns with no state block, teaching it by example to stop emitting one. Once it missed a turn, it never recovered, and retry did not help.
The fix has two halves, both gated on the adventure actually having a schema. First, a one-line EMIT_REMINDER sits in the recency slot. Second, _history_text() reconstructs each past turn’s delta block from the stored delta and re-appends it. Action.text stays clean, so the UI, the embeddings, and card trigger-matching are unaffected. History carries only the per-turn delta, for format imitation. The full scoreboard is rendered once, up top.
2.2 The performance trap hiding inside that
Building the context needs only the newest ~6,000 tokens of story. The obvious implementation reads adventure.actions, which loads every row of the adventure and then discards 90% of it. At turn 200 that was 839 KB read to use about 70 KB, and it grew every turn.
history.py serves three shapes straight from SQL: a tail, a slice, and a count. window_covering() fetches the newest 32 actions and measures their real token count. If that falls short, it projects how many more it needs from the average it just measured, instead of blindly doubling:
average = tokens / len(actions)
projected = int(budget / average * 1.15) + 8
Each round fetches only what it does not already hold, so no row is read twice. The same turn now costs 129 KB, flat from about turn 50. The cost is bounded by the context budget, not by the length of the story.
2.3 Prompt caching is a layout problem
An endpoint caches a prompt prefix. It compares this request against the last one, and it reuses the request up to the first byte that differs. It bills every byte after that point in full.
One section that changes every turn therefore costs the whole prompt. If you place the live stat block above the story history, a single changed stat value bills the history again. The history is most of the prompt.
§2.1 lists what goes into the prompt. This section sets the order. The sections that change least often go first.
The three layout rules
- The static block does not change between turns. It holds only what you edit by hand: the narrator instructions, the guide built from the stat schema, the emit rule,
ai_instructions, the persona, and the plot essentials. None of these change during a story, so the bytes are identical every turn. Do not add a changing section tosystem_sections. Live stats, retrieved memories, and the rewritten summary belong below the history. This mistake produces no visible error. It raises the token cost of every turn. - The live sections sit below the history, ordered by how often they change. The order is summary, lore, memories, then stats. The summary changes every 15 actions. Lore changes with the scene. Memories change on most turns. Stats change on nearly every turn. In this order, a turn that changes only a stat leaves the other three sections inside the cached prefix.
- The tail sections keep the last slots. These are the front memory, the length hint, the refusal note, and
EMIT_REMINDER. §2.1 gives the reason: the model responds most strongly to the text it reads last. These sections are short, so a cache miss on them costs little. Position matters more than caching here.
reserved still counts the summary, the memories, and the live state, because the prompt still contains them. Only their position changes. world_lore is the exception: the builder charges it to available, alongside the cards.A cache lives on the machine that served the request. OpenRouter distributes requests across upstream providers, and each upstream keeps its own cache. A stable prompt that reaches a new upstream starts with an empty cache.
_apply_provider_routing() names a preferred upstream for the vendors where this matters, for example provider: {"order": ["deepseek"]}. This is a preference, not a restriction. Fallbacks stay enabled. If the preferred upstream is down, the turn runs elsewhere and misses the cache.
Endpoints other than OpenRouter receive no provider field.
The provider measures the hit rate. It reads prompt_tokens_details.cached_tokens from the endpoint's final usage chunk. That value counts the prompt tokens the endpoint served from cache instead of billing in full. A later chunk that carries no usage does not clear the value.
test_prompt_caching.py pins the layout. It changes one stat, then asserts five things: the static block does not change, the volatile sections stay below the history, the tail sections keep the last slots, the budget still charges for the live sections, and the new turn only appends to the cached prefix.
2.4 World state: the AI proposes, Python referees
Three options, and two of them lose.
- A deterministic dice engine. What a real RPG does. It loses here because the action space is unbounded. Mapping arbitrary natural language onto a fixed rules system is harder than the problem being solved.
- Let the model own the numbers. Fails immediately. Models are bad at arithmetic, worse at holding a number across twenty turns, and unable to obey their own frequency rules. Tell one “change this at most every 5 turns” and it changes it every turn.
- The model proposes, the engine disposes. Chosen. The model narrates and appends a JSON delta. Python validates and clamps it before anything is stored.
narration: "The blade catches your shoulder. Gwen shouts and drags you back."
```state
{"player.hp": -15, "npc.gwen.trust": 5, "milestones.escaped": true}
```
| Rule applied, in order | What it stops |
|---|---|
| Path must exist in the schema | Hallucinated stats |
| Value must be the right type | "a lot" instead of -15 |
| Cooldown | Changing a stat more often than the scenario allows |
| Counters can’t decrease | The in-game day going backwards |
max_delta_per_turn | Losing 90 hp to a stubbed toe |
Clamp to min/max | Negative hp, trust above 100 |
Milestones are sticky, true only | Un-completing a quest |
| Flags are two-way booleans | Nothing — deliberately unrestricted |
Everything rejected is reported, not silently swallowed: Insights shows applied, clamped and rejected paths per turn, and a chip under each narration shows what actually changed.
Word bands are the reliability mechanism
"hp": { "min": 0, "max": 100, "initial": 100,
"bands": [[0,20,"very weak"], [20,40,"hurt"], [40,60,"minor damage"],
[60,90,"healthy"], [90,100,"full health"]] }
The live state line shows the current band label, like hp 55/100 (minor damage). The model reads a word, not just a number. The stat guide also prints the whole ladder once per turn. Models reason well over semantics and badly over arithmetic. “He’s badly hurt, so a solid hit takes him to very weak” is a judgment a model can make. “55 minus 22 is 33” is one it gets wrong often enough to matter.
Two philosophies underneath
Nothing in this engine raises. A malformed delta returns {} and the turn continues. The parser strips trailing commas and leading + signs. It accepts a fence labelled state, one labelled json, or an unlabelled one. It falls back to a bare object hugging the end of the text, but only if that object parses into something delta-shaped, so prose ending in } is never eaten. This tolerance exists because the hosted demo runs on free-tier models. A stricter parser would mean good models work and free ones don’t.
One call, not two. Narrate-then-extract is more reliable per call, but it costs twice the latency and twice the rate-limit budget. On a 20 req/min free tier, that halves the playable turn rate. The tolerant parser plus the terminal reminder buys the same reliability for less cost.
2.5 Output length, decided by measurement
The hint’s whole job is protecting the state block, which is emitted last and is therefore what truncation eats. So it must shorten output. The first version lengthened it.
| Phrasing, cap 800, n=5 | Mean words |
|---|---|
| No hint | 174 |
| “Keep this turn under about 506 words.” | 246 |
| “Hard limit … must not exceed 506 words … a typical turn is much shorter.” | 170 |
A budget reads to the model as a target to fill. Every one of the five budget runs was longer than every unhinted run, 41% longer on average. That pushes output toward the wall the hint exists to avoid. Ceiling phrasing is statistically indistinguishable from no hint at loose caps, but it still works at tight ones. At cap 250, unhinted runs hit finish_reason: length 2 times in 6. Ceiling-hinted runs hit it 0 times in 6.
A one-sided ceiling turned out to be half a fix. Across other models, the same prompt gave wildly different lengths. A terse model has nothing to act on but “much shorter”, so it collapses to two paragraphs. The hint is now a band with deliberately asymmetric bounds, so neither side reads as a number to hit:
must not exceed 506 words, and it should not stop short of about 177.
Prefer the lower end of that range unless the scene genuinely needs more.
LENGTH_FLOOR_SHARE = 0.35 of the ceiling
MIN_LENGTH_FLOOR_WORDS = 60 # below this the floor is dropped and the
MAX_LENGTH_FLOOR_WORDS = 300 # tight-cap string stays byte-identical
No test was run on the band wording. Two risks remain open. A stated range may invite landing mid-range on verbose models. Truncation is still silent, since nothing in the app reads finish_reason yet. The better design is a target_length preference separate from max_output_tokens, which is currently a safety wall doubling as the length dial. It was rejected as too big for the ask.
2.6 The memory bank
Turn 4 said you promised the innkeeper you’d return. At turn 90, that is long gone from the prompt. But if you walk back into the inn, it should come back. Three layers make that happen:
| Layer | Cadence | What it is |
|---|---|---|
| Memory | every 6 actions, from 12, once 1 action sits past the block | One or two past-tense sentences of concrete fact, third person, at most 50 words. |
| Story summary | every 15 actions | A single ≤250-word overview, rewritten by folding in the new memories. |
| Retrieval | every turn | Embed the last 4 actions (≤600 tokens), cosine-rank the bank, inject the top 5. |
Retrieval is what answers the innkeeper problem. The promise is a memory. The memory has a vector. Walking into the inn produces a query vector near it.
- A memory hangs off the node whose block it ends on. That is a
(branch_id, depth)coordinate, not the adventure and not a position in a list. This makes “which memories described this turn?” an indexed lookup. It is also why memories inherit correctly across a fork: the ones above the fork point already sit on ancestors both lines read. - A block waits one action past its end before the bank summarizes it (
SETTLE_SLACK = 1). This rule saves money. It does not correct an error. Retry and take-switching accept only the newest action. A memory whose block ends on the tip is therefore the only memory you can still discard, and every retry of that turn paid to write it twice. Blocks close every 6 actions and a normal turn writes 2 actions, so this cost one turn in three. Quality does not change, because the history window still holds a newly closed block in full. - Cursors only advance on success. Every AI call here is best-effort. If summarization fails, the cursor is unchanged, and the same block is retried on a later turn. There is no retry loop, no backoff, no dead-letter queue. The cadence is the retry mechanism.
- Pinned memories count toward
top_k. Otherwise, 6 pinned plustop_k=5injects 11 and blows the budget the whole context engine exists to respect. - A dimension mismatch scores 0.0, it does not crash. Change your embedding model, and old 768-dim vectors meet a 1536-dim query.
zip()would happily truncate and score garbage silently. - Eviction is LRU-ish and non-destructive. Over capacity (default 200), the least-used unpinned memories are marked
forgotteninstead of deleted, so you can un-forget one. - Background calls never spend the shared demo key. Summarization and embedding providers are built from the user’s own settings, never from the demo config. Unmetered background calls on a server-funded key would be a bill.
What the summarizer prompt contains
The memory prompt used to contain six actions of second-person prose and nothing else. It named no protagonist, no cast, and no setting. It also gave no instruction about grammatical person. Two problems follow.
The grammatical person varies. The model chooses one per call. A single bank then holds “You entered the crypt”, “The player entered the crypt”, and “He entered the crypt” for the same kind of event.
The memory omits names. Given You push the door open. She grabs your arm., the only accurate memory is “You entered a room and she stopped you”. Retrieval injects that memory forty turns later, into a scene with three women in it. The memory then supplies no usable information.
The summary repeats both problems, because the summarizer builds it from the memory bullets.
Two changes correct this. Both apply to the summary prompt as well.
- The adventure names a protagonist. A persona carries a name, pronouns, and a description. Only you can edit it, so it does not change during a story. It therefore sits in the static block of §2.3 and costs nothing after the first turn. The builder emits it whether or not the adventure has an RPG layer, because an adventure without stats still has a protagonist.
- A cast brief precedes both prompts. It names the protagonist, then the NPCs from
stat_schema.npcs. If the adventure has no RPG layer, it adds the story cards that match the block being summarized. The ceilings are 8 members, 240 characters each, and 300 tokens of plot essentials for the setting. Keyword matching alone omitted people, so the builder tops up the roster. A block can center on a character that never triggers a keyword.
Gwen: trust 40 (wary) changes the memory prompt on every turn, which loses the caching described in §2.3. It also frames one event two ways, depending on when the summarizer reads it. That is the problem this section corrects, so the brief must not reintroduce it.The framing rule occupies more than half of the memory prompt. It states its reason as well as the instruction: write in third person, name the protagonist, and name characters instead of using bare pronouns. The reason is that retrieval injects a memory on its own, with no surrounding story.
The prompt also caps a memory at 50 words. That number comes from measurement. Given only “1-2 plain sentences”, a real model wrote 34 words for one block and 105 words for the next. Retrieval injects 5 memories per turn at the default memory_top_k, so this cap sets what the bank costs per turn.
An existing bank does not update itself, because the bank writes each memory once. tools/rewrite_memories.py replays an existing bank through the current prompt. tools/memory_ab.py is the harness that compared the two prompts.
Two full rounds of egress work ran with the memory bank effectively disabled. Retrieval needs an embedding model, and embedding providers are BYOK-only by construction. So the demo never embeds, and the stress harness had none configured. Measured later in production: 134 memories, 1536 dims, ~31 KB each as JSON text. The ranking walked adventure.memories, so the whole bank crossed the wire every turn just to pick five. On a 100-memory adventure that was 3,024 KB per turn against 129 KB for everything else, about 96% of a turn. Break-even is 4.2 memories.
This is not a repeat of the earlier fix. There is no repeating group, no denormalization. It is a format problem (JSON floats at 20 bytes where a float is 4) plus a fetch-frequency problem. Rule: any egress measurement must run with an embedding model configured.
2.7 The story is a tree
The largest structural change the project has had, and the one with the most reasoning behind it.
The story used to be a list, and a mutable one. Retry rewrote the last entry in place. Undo and delete removed entries from the middle. Everything derived from the story was indexed by position in that list: memories, the running summary, the two marks saying how far each had got. A position means something different once anything in front of it is deleted. That one fact produced a family of bugs that all looked different:
- Deleting a middle action slid a never-summarized action into the “already covered” range. A recent action silently never became a memory.
- Discarding a memory left its actions behind the mark, describing nothing.
- Retry rewrote text after the mark had passed it. The memory then described narration no longer in the story.
- The retried row stayed attached while its replacement was written. The model was shown the attempt it was meant to replace and wrote a continuation of it. That exclusion had to be threaded through four separate readers.
- Attempts lived in a JSON array on the row, with a mirrored copy of the live one in the ordinary columns. That is a repeating group and a denormalization in one.
The shape
branches(id, adventure_id, parent_branch_id, fork_depth, lineage, name)
actions (id, adventure_id, branch_id, depth, parent_id, live, text, …, state_after)
memories(…, branch_id, depth)
adventures(…, head_branch_id, head_depth)
depth is a position along a path, not a global turn number: A4 and B4 are two alternatives, not two turns. Reading branch C is one query:
SELECT * FROM actions
WHERE (branch_id = 'C')
OR (branch_id = 'B' AND depth <= 5)
OR (branch_id = 'A' AND depth <= 3)
ORDER BY depth DESC LIMIT 32 -- → A0 A1 A2 A3 B4 B5 C6 C7
Why branch_id + depth, and not parent pointers alone. Parent pointers are the obvious way to store a tree, but they are the wrong way to read one. Reading a story would need N round trips up a chain, which throws away the windowing work of §2.2. Depth replaces the old index as the ordering key, so reads keep the shape they already had. ltree was rejected because it breaks SQLite dev parity and its path column grows per row. Closure tables were rejected for being O(n²).
A branch stores no story of its own. A fork costs only an id, a parent, a fork depth, and a cached ancestry. Measured on a 40-turn story forked twenty times, against the same story flat: a page load of 31,652 B against 31,433 B, or 1.007×, about 103 bytes per branch. No migration, no vacuum, no copy.
What a player actually does
‹ 2/4 › is free, since the story below simply empties: that take has no children yet. A branch is created when you write below a take that is not the live one, never before.That rule collapses two operations into one and deletes a distinction from the UI. The first version of this screen had a chip that switched at the tip and only previewed above it, with a second button to take that line. One control's meaning depended on where the reader was standing. It shipped, was driven by hand, and was found unusable. The replacement is a pager that only ever steps, a fork button on every turn, and no tip-versus-past distinction at all.
Takes are grouped by parent, not by coordinate
This is the load-bearing detail, and it is not obvious. The natural way to find “the other takes of this turn” is by coordinate: same branch, same depth. That is wrong in both directions:
B ── C C1 C2 <- three takes, one parent (B)
│ └── D1' D2' <- two takes, parent C2
└── D1 D2 D3 <- three takes, parent C1
Standing on the C2 path at that depth must read 2/2, not 5. Coordinate grouping gets that right by accident: writing under a non-live take forks, and the two sets land on different branches. It gets C wrong. Once C is forked onto a branch of its own, it sits alone at its coordinate and reads 1/1, having lost C1 and C2 from a pager that must still say 1/3.
So a node carries parent_id, read for nothing else. The alternative was making a branch’s fork point a node rather than a depth, and it was rejected. The whole point of lineage is that a read is an OR-clause per branch instead of a walk. Re-pointing the fork at a node changes path resolution itself, dragging in the cursors, the memory depths, and both bundle formats. parent_id is one indexed lookup, never a walk.
Cursors become anchors
The two marks, how far the memory bank has got and how far the summary has got, used to be counts. A count is a position in a list. Every rule about sliding, rewinding, and translating between positions and Action.index existed to patch up the fact that the list moves.
A cursor is now an anchor: (branch_id, depth), the node up to and including which the work is done. Deleting an action does not move it. “What is not covered yet?” becomes a question about the story instead of a list index, and it answers correctly no matter what has been deleted in front of it. The branch half is what makes it survive forking. A depth alone is ambiguous once two branches both have a node 41.
position_of_index, note_action_removed, settled_story_actions, and the cursor-rewind machinery were deleted, not left unused. So was the one-turn memory holdback that existed because a retry could rewrite an action the mark had already passed. (SETTLE_SLACK later put one action of slack back, for what redoing a block costs rather than for what it could get wrong.)
A sibling group breaks every query that assumed one row per depth. _latest_narration ordered by (depth desc, id desc) and got the newest attempt rather than the live one. The index screen quoted a take the player had thrown away. Anything ranking actions by coordinate needs live in the filter.
delete_turn meant “every take at this coordinate”. Once the group spans branches, undo reached onto another line and deleted a take nobody asked about. Anything that reads a take group and then writes has to say whether it means the turn or the coordinate.
The adventure GET does not build ActionOut. It hands the window to the relationship with set_committed_value and lets Pydantic walk it. Patching every place that builds ActionOut still misses this one path, and every page load takes it.
One review finding was rejected, and that is the part worth keeping. A memory the player types lands on the head’s coordinate. A retry withdraws every memory at that coordinate, so the note disappears. That is reproduced and real, but it is the rule working: a memory anchored to a node describes that node and goes when the node goes. The root node is the one exception. The pre-tree bank was parked on depth 0, the one depth every branch can see, so withdrawing it would retire a whole bank in a click.
2.8 Undo and retry that actually rewind
Most implementations of undo delete the last message. That is wrong here, because a turn mutates three things: the text, the scripting scoreboard (script_state), and the RPG stats (world_state).
Every node carries state_after and world_state_after, deep copies of the adventure once that turn had played. Rewinding to before a turn is a read of the node in front of it, so undo, retry, and a branch switch are the same restore. The cooldown clock comes along free. It lives inside the world state at _meta.last_changed, so each line of the story carries its own without anything having to know there is one.
Nothing a retry replaces is thrown away. The old attempt stays as another take, a sibling node with live false, and the pager steps between them. Retry is not a special case. It is the tree with the branch not yet created.
- The turn being retried is excluded from its own context. Its takes are still attached to the adventure. Without
exclude_action_id, the model is shown the attempt it is replacing as established story. The exclusion had leaked into four readers: history replay, story-card trigger matching, in-scene NPC detection, and the memory-bank similarity query. Anything reading the story during generation takes the exclusion. - A retry reuses the turn’s depth, not the next one. Cooldowns are measured along the path. A new depth would advance the clock the cooldown rules run on, and it would quietly unlock stats that should still be waiting.
- If regeneration fails, the rollback is reversed.
generate_turnwraps the generator intry/finally. A provider error, an empty reply, a scriptstop, or the browser hanging up puts the previous take back in charge. Otherwise server state drifts from the text still on the user’s screen.
2.9 The turn lock
One turn at a time per adventure. The subtlety is where the check goes. A StreamingResponse does not start iterating its generator until the response begins. A check inside the generator would let two rapid requests both pass before either claims the slot. And because sync FastAPI endpoints run in a threadpool, the test-and-set needs a real threading.Lock.
def acquire_turn_lock(adventure_id): # in the REQUEST handler
with _active_turns_guard:
if adventure_id in _active_turns:
raise HTTPException(409, "A turn is already generating…")
_active_turns.add(adventure_id)
async def with_turn_lock(adventure_id, gen): # wraps the SSE generator
try:
async for event in gen: yield event
finally:
_active_turns.discard(adventure_id)
The lock is in-memory, so it is a single-process guarantee. That is honest for the deployment it targets. Two workers would need the lock in the database.
2.10 Streaming
The app uses Server-Sent Events, not WebSockets. Traffic is one-directional, and a bidirectional connection for a unidirectional problem is cost with no return. FastAPI reads the provider’s stream and yields data: {"type":"chunk","text":"…"}. The frontend reads the body with a ReadableStream reader, buffers on \n\n boundaries, and dispatches each event. Event types: player, reasoning (thinking traces, into their own collapsible panel with their own budget), chunk, stopped, error, done.
X-Accel-Buffering: no. nginx-style reverse proxies buffer by default, turning a stream into one delivery at the end.
The security-headers and body-size middlewares are written as pure ASGI rather than Starlette’s BaseHTTPMiddleware. The latter buffers the response body, which would break streaming outright.
The empty-reply case is diagnosed, not just reported as “empty”. If a reasoning model streams thinking but no story text, it spent its whole budget thinking. The error says so and names the three settings to change.
2.11 The QuickJS sandbox
Real AI Dungeon scripts define modifier(text) and call it as the last line, with globals like state, history, and storyCards. The same contract runs here, inside an embedded QuickJS interpreter. The safety properties are mostly structural. Nothing was removed, because nothing was there to begin with.
| Property | How |
|---|---|
| No filesystem, network or process access | QuickJS has none by default |
| Memory cap | 16 MB per run |
| CPU cap | 2 seconds per run |
| No shared state between runs | A fresh Context per hook execution |
| A broken script can’t break a turn | Every failure returns .error with text, state and cards unchanged |
Data crosses the boundary as JSON. Python serializes {state, text, history, storyCards, info} in, and the script’s results come back out the same way. There is no object bridge to exploit.
One bug-compatibility is deliberate: addStoryCard returns the new card’s index. The first card returns 0, which is falsy, so if (!addStoryCard(...)) misfires. That is upstream AI Dungeon’s behavior. Matching real scripts is the entire point of the feature.
Part 3 — Running it in public
A hosted demo on a free tier turns three things into engineering problems that a local app never has: bytes on the wire, a spending surface, and strangers.
3.1 The 189× egress fix
Neon’s free tier allows 5 GB of network transfer per project per billing period, and it hard-blocks at connection time. There is no read-only grace. You cannot even pg_dump your way out once it trips. This app tripped it.
The diagnosis came from the monitoring tab, not a guess. The database was ~55 MB, and had transferred 5 GB. Compute was idle most of the day, and the pooler never exceeded three connections. That is the whole database, ninety times over: a payload-per-request problem, not traffic volume and not a leak.
Action.context_snapshot holds the entire assembled prompt, about 74 KB per row and 94% of the database. Every adventure load pulled it for every action, just to read two small fields out of it. SQLAlchemy loads all columns by default.
| Step | What it does |
|---|---|
| 1 | Move the two things actually needed per action into their own column, Action.world_delta. |
| 2 | Mark the heavy columns deferred — context_snapshot, variants, reasoning — so they load only when asked for. |
| 3 | Backfill with dialect-specific server-side SQL (json_extract on SQLite, #> on Postgres) so the 39 MB never crosses the wire. |
One adventure load went from 38.5 MB to 0.20 MB. A second round found two more problems. Action.variants was bulk-fetched just to compute a count, fixed the same way plus a real variant_count column. And story_actions() walked the relationship, so a turn was O(story length), which is what §2.2 replaced. A 200-turn playthrough went from 84.5 MB to 23.0 MB. A delete went from 115 KB to 5 KB.
tests/test_egress.py hooks SQLAlchemy’s before_cursor_execute. It captures every statement the ORM emits and fails if a bulk load ever names those columns again. The regression is caught by asserting on the SQL, not on a timing. It was verified by sabotage.Query.count() wraps the entity select in a subquery, so the emitted SQL names every deferred column. No bytes come back, but the database still reads them. A SQL-grepping guard cannot tell that apart from a real bulk fetch. Use db.query(func.count(Action.id)) instead.
SQL and Python must agree exactly on what counts as a story action, or cursors point at the wrong one. trim() strips only spaces, so fold \n\r\t with replace() first.
Read table sizes from column sums, not n_live_tup. After a migration rewrites a table, that estimate goes stale in the direction that makes bloat look smaller. It said ~40 MB was reclaimable; the real figure was 79 MB. sum(octet_length(col)) is the honest number.
After a migration that rewrites actions, run one VACUUM FULL actions; on the direct endpoint, not the pooler. It takes an ACCESS EXCLUSIVE lock. And DROP COLUMN is metadata-only in Postgres, so it frees nothing by itself.
/api/health instead, which deliberately does not query the database.3.2 Migrations, hand-rolled
No Alembic. An append-only list of (version, SQL) pairs, with the current version in SQLite’s PRAGMA user_version or a one-row table on Postgres. 64 versions so far.
- A fresh database is created by
Base.metadata.create_all()— always current — and stamped at the latest version. It never replays history. - An existing database runs every migration above its stored version, in order.
Someone may run this single-file SQLite app for months. The entire requirement is “add a column, don’t lose their data”. Alembic’s autogenerate, branching, and down-migrations are machinery for a team with a staging environment. This is 250 lines, and you can read all of it.
The constraint it creates is written at the top of the file. Change models.py so fresh databases are current, and append a pair so existing ones upgrade. Migrations 2–23 predate Postgres support and use SQLite-only syntax. That is harmless, because every Postgres database starts fresh, but anything added since must run on both dialects.
Repairing duplicate action indexes uses UPDATE … FROM with a window function rather than a correlated subquery. SQLite may evaluate a correlated subquery against partially-updated rows, which would produce fresh duplicates while “repairing” them.
Separately, the test suite only ever builds fresh databases. It does not cover migrations as upgrades. Check the upgrade path by hand whenever one is added.
3.3 The spending surface, and the strangers
The demo lets people play with no signup and no API key, on a key the server pays for. That makes it the most defended code in the project.
resolve_provider_config() is the single place the BYOK-versus-demo decision is made. On the demo branch it pins two things. It pins the model to a whitelist, so no caller-supplied override can aim a server-funded key at an expensive model. It pins the endpoint to the configured demo URL, so the key cannot be redirected to a URL the user controls and harvested. A defensive __post_init__ raises if a demo config somehow carries a non-whitelisted model.
That backstop tests using_demo, not api_key == DEMO_API_KEY. Keying on the key value looks stricter, but it is wrong. The demo key is an ordinary OpenRouter key, so a user can legitimately paste that same value into their own settings as BYOK. Every resolution then raised, 500ing even GET /auth/me, which is the SPA’s bootstrap call. Nothing rendered at all.
/auth/me is a single point of failure for the whole frontend. Anything it touches must not be able to raise. And using_demo is what actually means “the server is paying”.
Secrets
Everything derives from one server-side secret. Passwords use hashlib.scrypt (N=214, r=8, p=1, per-password salt, constant-time compare); this is a stdlib function, so it needs no extra dependency. Sessions are v1.<user_id>.<HMAC-SHA256> with no expiry, because long-lived guest sessions are the point. Stored LLM keys are Fernet-encrypted at rest with an enc: prefix, so legacy plaintext rows stay recognizable. The secret auto-generates for local installs, but multi-user mode refuses to start without the env var. Hosted filesystems are ephemeral. A regenerated secret on every deploy would silently log out every user and orphan their stored API keys.
Two findings from an authorized pen-test
Rate limits were bypassable (High, confirmed live). The Dockerfile ran uvicorn with --forwarded-allow-ips "*", which trusts the leftmost X-Forwarded-For value. Render forwards rather than strips the inbound client XFF, appending the real client IP on the right. So the leftmost hop was fully attacker-controlled, and every rotation bought a fresh rate-limit bucket. Proof: a fixed IP hit 429 after 10 login attempts; rotating a spoofed XFF passed 14 of 14. That defeated the only anti-brute-force control in the app.
Fixed in two independent layers. limits._client_ip() now reads the rightmost hop, the one Render appends, which a client cannot push past. It is tunable via AIDND_TRUSTED_PROXY_HOPS. A new per-account throttle also applies: 8 failures per 15 minutes, keyed on the target email, so a botnet with many real IPs still cannot brute one account. The accepted tradeoff is that an attacker can keep a known account in a 15-minute cooldown. That is a nuisance, and it is strictly better than brute force.
SSRF via the BYOK endpoint (Medium). A hosted user could point Settings.endpoint_url at an internal or metadata address, and the connection test echoed part of the response back. netguard.endpoint_block_reason resolves the host and refuses anything where not ip.is_global. It checks at request time, so it resists DNS rebinding. It is a deliberate no-op in local mode, since a local install reaching localhost:11434 is the intended case.
What held, and is worth naming: per-object authorization was solid throughout. Every {id} route filters on user_id, and sub-resources re-check they belong to their parent. The API key is write-only in every response shape. Session tokens are unforgeable. The sandbox has no host bindings.
The guard rails, in numbers
| Guard | Limit |
|---|---|
| Turn generation | 10 / min |
| Auth attempts, per IP | 10 / 5 min |
| Guest creation, per IP (each is a DB row) | 30 / 5 min |
| Script test runs (each costs up to 2 s CPU) | 30 / min |
| Connection test (outbound HTTP to a user URL) | 10 / min |
| Adventures / scenarios / scripts per user | 100 / 200 / 200 |
| Actions per adventure | 5,000 |
| Request body, and on import endpoints | 2 MB / 20 MB |
| Demo turns per user per day | 20 |
Import endpoints check bundle list lengths against the same caps live creation enforces. Otherwise the cap is bypassed by uploading a file. That check had a bug worth remembering: a cap has to count what gets written, not what the file says. The import counted a v1 file’s turns, but each turn expands into a row per saved attempt. A file inside a 5,000-action cap could write 50,000 rows.
3.4 Counting visits without undoing §3.1
The obvious way to build analytics is to write a row per request and read rows per dashboard query. That would have undone the entire egress fix. So a visit is a write and never a read. Counts accumulate in a process-local dict and flush every 60 seconds as UPSERTs into a generic (day, metric, label) → hits table, plus one row per visitor per day for the funnel flags. Every dashboard query is a GROUP BY. It returns tens of rows no matter how much traffic sits behind it.
The counters are anonymous and the access log beside them is not, on purpose. A visitor in the counter tables is HMAC(secret, "visitor:<user id>") truncated to 32 chars. It is one-way, so those tables cannot be joined back to users, and keyed, so no client can compute one. Story content never reaches that module. The identifying half lives in a separate module and a separate table, so the anonymity of the counters is a property of the code rather than a convention.
- The funnel counts people, not clicks. A player who starts six adventures is one person who started an adventure. That is the whole reason the per-visitor-day table exists. Its flags only ever turn on.
- A failed turn is an HTTP 200 with a bad ending. Status-code middleware cannot see one. A demo whose model started refusing every request would look perfectly healthy. All five SSE error paths now go through one
turn_error()helper. - The tests run on SQLite; production is Neon. A flush that raises is caught and logged. A dialect mistake in the UPSERTs would have stayed invisible while the dashboard quietly stayed empty. One test compiles both statements against the Postgres dialect without connecting to one.
Part 4 — The scoreboard
What the work bought, and what it deliberately did not.
3.5 Compressing the prompt snapshots
§3.1 stopped the server from reading context_snapshot. It did nothing about storing it. The column holds 89% of the database, measured as 150.8 MB of JSON across 944 actions, and 232 kB on a single row of the longest adventure. The free tier allows 512 MB. Reading the column was the defect. Storing it sets the limit on how long the database fits that allowance.
Postgres compresses the column already. TOAST reduces 150.8 MB to about 89 MB, a factor of 1.7. TOAST uses pglz, which favors fast decompression on data that a query might filter on. No query filters on this column. One screen reads a prompt whole, and it reads it rarely.
zlib in the application layer reaches a factor of three to four on the same text. It costs one decompression, on a request that already makes an LLM call. Level 6 gives most of that benefit. Level 9 spends noticeably more CPU on prompt text and saves about one percent more space.
TypeDecorator, not a second column. Every call site still writes action.context_snapshot = {...} and reads back a dict. deferred=True, undefer(), and load_only() still name the same attribute. Only the storage format changes.The migration runs in three steps. It adds context_snapshot_z, backfills it, then drops the old column and renames the new one. The backfill and the drop share one transaction. If the backfill raises, the DROP rolls back with it, and the prompts remain.
Postgres does not release disk space on its own. DROP COLUMN only marks the column dropped, and the backfill leaves one dead tuple per row. The table therefore grows first, and it peaks at roughly twice its starting size while both columns exist. Autovacuum makes that space reusable, but it does not shrink the files.
Run VACUUM FULL actions; once after the deploy. On the 2026-08-17 figures, the database measures 99.6 MB before the migration, peaks near 200 MB, and settles at about 53 MB. If you skip the vacuum, nothing breaks, and the database keeps the larger files.
Measured results
| Property | Figure |
|---|---|
| Database egress per adventure load | 38.5 MB → 0.20 MB |
| Turn read cost at turn 200 | 839 KB → 129 KB, flat |
| 200-turn playthrough, total reads | 84.5 MB → 23.0 MB |
| Cost of a branch | ~103 B · 20 forks load at 1.007× |
| Prompt snapshot storage | 150.8 MB → ~53 MB, 89% of the DB |
| Length hint, budget vs ceiling phrasing | 246 vs 170 words (n=5) |
| Backend tests | 632 |
| Schema versions | 76 |
| Sandbox limits | 16 MB, 2 s, fresh context |
| Context defaults | note at depth 3, cards ≤ 40% |
| Memory cadence | memory /6 +1 slack, summary /15, top-5 |
| Memory length ceiling | 50 words, measured against 34–105 |
| Actual hosting bill | ~$0.30/mo, mostly $0 collected |
Two of those tests encode a performance property rather than a behavior: test_egress.py asserts on the SQL the ORM emits, and test_history_window.py asserts that the read cost stops growing with story length.
Known limitations
Deliberate trades for a single-user-first app that also happens to be hosted, written down so nobody has to discover them the hard way.
- Single process. The turn lock, the rate limiter and the summarization task all assume one worker. A second would need a row-level advisory lock and Redis.
- No vector index. Retrieval does cosine similarity in Python over the whole bank. Fine at the 200-memory cap; at 10,000 it wants pgvector.
- Prompt snapshots are compressed, not expired. §3.1 made them cheap to read and §3.5 made them cheap to store, but nothing deletes them. Storage still grows with every turn played, just four times more slowly.
- In-memory rate-limit windows reset on restart, so a restart grants a brief extra allowance.
- Background summarization is a fire-and-forget asyncio task, so it does not survive a restart. At real load it belongs in a queue.
- The two memory marks are one pair on the adventure, not one per branch. Switching lines makes the mark on the line being left unreadable, so that ground is summarized again. It fails in the safe direction — redo, never skip — but switching back and forth costs AI calls.
- Story cards are adventure-wide, so a card invented on one branch shows on all of them.
- Editing an already-summarized turn leaves its memory stale. Replacing a turn withdraws what was derived from it; editing one in place does not.
- The frontend has no test runner. This is the standing reason the project keeps finding UI bugs by hand — every one of the tree-UI findings above was invisible to 632 green backend tests.
- Truncation is silent. Nothing reads
finish_reasonyet, which is the obvious next step for §2.4.