1 · A loop that waits for the model
September 16, 00:45 · the runner, the random baseline and the recording
The first piece was the game loop. ViZDoom normally runs in real time, so a slow policy acts on frames that are already out of date. Here the engine waits: observe, decide, advance one tic, repeat. Every decision sees a fresh state, there is no queue of stale actions, and the game runs exactly as fast as the policy allows. Video can still be replayed at normal game speed, or at the measured wall-clock speed to show how fast the policy really is.
The scenario is defend_the_center: the player stands in a round room while demons and chainsaw
marines walk in from the walls. Before any model, a random policy pressing one legal button combination per game
second set the baseline.
The game
- ViZDoom
defend_the_center, 640×480, 35 tics per second - Seven buttons for model play: attack, strafe left/right, forward/back, turn left/right
- Opposite directions and firing without ammo are masked as illegal
The policy's input
- An instruction, e.g. “Attack enemies on sight.”
- The game state as text, read from the engine: no pixels
- Out: seven sigmoid scores; the best legal combination is pressed
The record
- Every decision's observation, scores, buttons and timing in
states.jsonl - A screenshot each game second and an H.264 replay
- A summary with kills, deaths, decisions per second and decision latency
2 · ModernBERT on Apple Metal
September 16, 00:45–01:08 · speed first, with an untrained head
Before training anything, we measured whether a 149.6M-parameter encoder could keep up with the game. The timing covers the whole decision: building the text, tokenizing, the full encoder forward pass, moving scores back to the CPU and decoding the buttons. Model loading and warm-up are excluded.
| ModernBERT-base on an M3 Max, batch one | Per decision |
|---|---|
| CPU, float32 | 43.3 ms |
| Metal (MPS), float16 | 13.3 ms |
| Metal (MPS), bfloat16 | 34.2 ms |
Float16 became the default. With hardware video encoding on, the whole game loop completed 175 decisions in 3.72 seconds, 47 decisions per second including the first call.
A dashboard that does not slow the model
A charcoal dashboard puts the game on the left and the instrumentation on the right: rolling latency, decisions per second, all seven button scores against a 0.5 marker, the legal combinations and why a button was suppressed. Drawing and encoding run in a background worker that keeps only the newest frame, so the model never waits for video. The first recording, still with the untrained head, made 2,100 decisions in 60 game seconds and took 86.5 seconds of wall time: 30.8 ms per decision and 24.3 decisions per second with the window and video on.
3 · Training only the head
September 16, 01:24 · a controlled vocabulary and a 3.6-second training run
For the trained checkpoint, the game state is described with a small fixed vocabulary. Code reads the nearest visible monster's position from the engine and bins it, along with health and ammo, into text such as:
health=healthy ammo=loaded nearest_enemy=center range=medium
These are perception features, not commands: the text never names a button. A scripted teacher with three behaviours (aggressive, cautious, evasive) labelled 468 examples, two known phrasings per behaviour. Whole state combinations were split between training (62 combinations, 372 examples), validation (8, 48) and test (8, 48), so no tested state was seen in training under any instruction.
Only the head was trained: 595,975 of the model's 149,610,247 parameters. The encoder's outputs were cached during training only; during play, every decision runs the full encoder. Training took 3.6 seconds of wall time, and the selected checkpoint is step 150 of 1,200, chosen on validation.
4 · Same state, opposite instruction
September 16, 01:28–12:55 · attack and evade from one checkpoint
The checkpoint then played with the dashboard and the video recorder on, deciding every two game tics. From an identical starting state, changing only the instruction moved the ATTACK score from 0.979 to 0.026. In the evade run, ammo stayed at 26 and ATTACK never rose above 0.036: the model's own scores, not a rule in the decoder, kept it from firing.
| Instruction | Game time | Playback | Decisions/s | Per decision | Result |
|---|---|---|---|---|---|
| Attack enemies on sight. | 60 s | 24.0 s | 43.7 | 13.9 ms | 16 kills, 1 death |
| Do not shoot. Evade the enemies. | 30 s | 12.1 s | 43.3 | 13.9 ms | no shots, no deaths |
Two tics per decision need 17.5 decisions per second to keep up with the game's native speed; the model ran at more than twice that, so the films above play faster than real time. They were joined into the 51-second edit at the top of this page.
5 · What this does not show
The boundaries of a controlled demonstration
The model reads a hand-built description of the game, not the screen, and the description comes from the engine's own object positions. It was trained on six phrasings of three behaviours and a finite set of states. It follows instructions it was taught, in situations described in its vocabulary. That is enough to show a text classifier steering a game in real time, and not enough to claim it understands Doom or arbitrary language.