Doom Sandbox · engineering journals · 2026

Language and embedding models playing Doom

Three experiments in ViZDoom, from a text classifier reading the game state to a multimodal embedding model reading the screen with no training at all. The game waits for every decision, so each model plays at its own speed. Each journal tells what was built, what failed and what was measured, with the recorded runs, and says which parts are learned, which are zero-shot and which are fixed rules.

Zero-shot · EmbeddingGemma 2 · RTX 4090

EmbeddingGemma 2, zero-shot from pixels

Screen crops and cosine similarity to text prompts, no training: 37 kills and 2 deaths attacking, no deaths evading, with every reading on screen.

multimodal embeddingszero-shotpixelscosine similarity

Read the journal →

Trained head · ModernBERT · Apple Metal

ModernBERT policy

A frozen ModernBERT reads an instruction and the game state as text; a head trained in 3.6 s makes it attack or evade, at about 44 decisions per second.

text classifierinstruction followingimitation labels

Read the journal →

Fine-tuned · GLiNER2.5 · RTX 4090

GLiNER2.5, zero-shot then fine-tuned

Zero-shot it fell short; fine-tuned on 12,000 teacher-labelled examples in four minutes, it matches the teacher on every held-out test, with every score on screen.

zero-shot classificationfine-tuninginstruction followingvideo

Read the journal →