← Doom Sandbox

Engineering journal · September 16, 2026 · ViZDoom, ModernBERT, PyTorch, Apple Metal

A text classifier that plays Doom from instructions and game state

ViZDoom waits while a ModernBERT classifier reads an instruction and a short text description of the game, then presses the buttons it scores highest. Only a small head on top of the frozen encoder is trained, in 3.6 seconds; the full 149.6M-parameter model runs on every decision. The same checkpoint attacks or evades depending on the instruction alone.

The 51-second edit at measured wall-clock speed, resampled to 30 fps: the trained classifier attacks, then, from the same starting state with only the instruction changed (0:29), evades. Separate seeded runs; the dashboard shows each decision's seven button scores.
16 / 1
kills / deaths in 60 game seconds, “Attack enemies on sight.”
0.979 → 0.026
ATTACK score for the same state when only the instruction changes
43.7/s
decisions per second, whole loop with dashboard and video recording
13.9 ms
per classification: text, tokenizing, full encoder, decoding
3.6 s
to train the 595,975-parameter head; the encoder stays frozen
Where this stands. A controlled demonstration. The classifier reads game state from the engine, turned into a small fixed vocabulary of text (health, ammo, the nearest enemy's side and range); it never sees pixels. Its labels come from a scripted teacher, it knows six phrasings of three behaviours, and a decoder only blocks impossible button combinations. It does not establish general Doom skill or open-ended language understanding. Every number on this page comes from the reports and summaries in the repository.

1 · A loop that waits for the model

September 16, 00:45 · the runner, the random baseline and the recording

The first piece was the game loop. ViZDoom normally runs in real time, so a slow policy acts on frames that are already out of date. Here the engine waits: observe, decide, advance one tic, repeat. Every decision sees a fresh state, there is no queue of stale actions, and the game runs exactly as fast as the policy allows. Video can still be replayed at normal game speed, or at the measured wall-clock speed to show how fast the policy really is.

The scenario is defend_the_center: the player stands in a round room while demons and chainsaw marines walk in from the walls. Before any model, a random policy pressing one legal button combination per game second set the baseline.

The game

  • ViZDoom defend_the_center, 640×480, 35 tics per second
  • Seven buttons for model play: attack, strafe left/right, forward/back, turn left/right
  • Opposite directions and firing without ammo are masked as illegal

The policy's input

  • An instruction, e.g. “Attack enemies on sight.”
  • The game state as text, read from the engine: no pixels
  • Out: seven sigmoid scores; the best legal combination is pressed

The record

  • Every decision's observation, scores, buttons and timing in states.jsonl
  • A screenshot each game second and an H.264 replay
  • A summary with kills, deaths, decisions per second and decision latency
The random baseline, one decision per game second, seed 7: 10 kills, 5 deaths and 61 observations in 60 game seconds. Played back at normal game speed.

2 · ModernBERT on Apple Metal

September 16, 00:45–01:08 · speed first, with an untrained head

Before training anything, we measured whether a 149.6M-parameter encoder could keep up with the game. The timing covers the whole decision: building the text, tokenizing, the full encoder forward pass, moving scores back to the CPU and decoding the buttons. Model loading and warm-up are excluded.

ModernBERT-base on an M3 Max, batch onePer decision
CPU, float3243.3 ms
Metal (MPS), float1613.3 ms
Metal (MPS), bfloat1634.2 ms

Float16 became the default. With hardware video encoding on, the whole game loop completed 175 decisions in 3.72 seconds, 47 decisions per second including the first call.

Speed, not skill. These runs used an untrained action head. They show what the hardware can do, not how well the model plays.

A dashboard that does not slow the model

A charcoal dashboard puts the game on the left and the instrumentation on the right: rolling latency, decisions per second, all seven button scores against a 0.5 marker, the legal combinations and why a button was suppressed. Drawing and encoding run in a background worker that keeps only the newest frame, so the model never waits for video. The first recording, still with the untrained head, made 2,100 decisions in 60 game seconds and took 86.5 seconds of wall time: 30.8 ms per decision and 24.3 decisions per second with the window and video on.

Thirty seconds of the untrained-head dashboard run, at measured wall-clock speed. The game ran slower than Doom's native 35 tics per second here; the on-screen clocks keep that difference visible.
Keep three numbers apart. Decisions per second (the whole loop, including drawing and recording), milliseconds per classification (the model's part of each decision) and the video's frame rate are different measurements. Every later video states which one it shows.

3 · Training only the head

September 16, 01:24 · a controlled vocabulary and a 3.6-second training run

For the trained checkpoint, the game state is described with a small fixed vocabulary. Code reads the nearest visible monster's position from the engine and bins it, along with health and ammo, into text such as:

health=healthy ammo=loaded nearest_enemy=center range=medium

These are perception features, not commands: the text never names a button. A scripted teacher with three behaviours (aggressive, cautious, evasive) labelled 468 examples, two known phrasings per behaviour. Whole state combinations were split between training (62 combinations, 372 examples), validation (8, 48) and test (8, 48), so no tested state was seen in training under any instruction.

Only the head was trained: 595,975 of the model's 149,610,247 parameters. The encoder's outputs were cached during training only; during play, every decision runs the full encoder. Training took 3.6 seconds of wall time, and the selected checkpoint is step 150 of 1,200, chosen on validation.

81.25%
exact seven-button vectors on held-out states
48 test examples
97.02%
individual buttons correct on held-out states
48 × 7 decisions
0.757 vs 0.003
ATTACK for one held-out aligned, wounded state
aggressive vs evasive instruction
A small test. The test set has only two positive ATTACK examples and none for MOVE_FORWARD. The saved float16 checkpoint was reloaded and checked with 48 complete encoder passes before any gameplay.

4 · Same state, opposite instruction

September 16, 01:28–12:55 · attack and evade from one checkpoint

The checkpoint then played with the dashboard and the video recorder on, deciding every two game tics. From an identical starting state, changing only the instruction moved the ATTACK score from 0.979 to 0.026. In the evade run, ammo stayed at 26 and ATTACK never rose above 0.036: the model's own scores, not a rule in the decoder, kept it from firing.

InstructionGame timePlaybackDecisions/sPer decisionResult
Attack enemies on sight.60 s24.0 s43.713.9 ms16 kills, 1 death
Do not shoot. Evade the enemies.30 s12.1 s43.313.9 msno shots, no deaths

Two tics per decision need 17.5 decisions per second to keep up with the game's native speed; the model ran at more than twice that, so the films above play faster than real time. They were joined into the 51-second edit at the top of this page.

5 · What this does not show

The boundaries of a controlled demonstration

The model reads a hand-built description of the game, not the screen, and the description comes from the engine's own object positions. It was trained on six phrasings of three behaviours and a finite set of states. It follows instructions it was taught, in situations described in its vocabulary. That is enough to show a text classifier steering a game in real time, and not enough to claim it understands Doom or arbitrary language.

  1. Compare with zero-shot modelsGLiNER2.5 took the same observations without Doom training; see the GLiNER2.5 journal.
  2. Read the screen instead of the stateEmbeddingGemma 2 later played from pixels with no training at all; see the EmbeddingGemma 2 journal.