← Doom Sandbox

Engineering journal · October 6, 2026 · ViZDoom, EmbeddingGemma 2, sentence-transformers, RTX 4090

A multimodal embedding model plays Doom from pixels, with no training

EmbeddingGemma 2 was released on October 6 and embeds text, images, audio and video into one 768-dimensional space. The same day, it played Doom zero-shot: every decision cuts the frame into 19 crops, embeds them, and compares each with fixed text prompts by cosine similarity. A small fixed rule set turns those matches into buttons. Nothing is trained, and the policy never reads the game state.

The 2 min 23 s showcase, recorded on an RTX 4090 and played back at normal game speed (the recorded runs made about 3 decisions per wall-clock second with the dashboard on). Each frame shows the 17 window crops coloured by their cosine margin, the bearing marker, the HEALTH and AMMO crops with their best-matching prompts, the instruction's mode cosines and the buttons pressed. An eye marks image inputs, # marks text.
37 / 2
kills / deaths in 60 game seconds, “Attack enemies on sight.”
0 deaths
in 60 game seconds, “Do not shoot. Evade the enemies.”, no shots fired
0
training steps, labels or fitted thresholds
89%
of fire decisions had a monster under the crosshair, 150 logged frames
~70/s
image embeddings per second on the RTX 4090 at 140 vision tokens
Where this stands. A same-day zero-shot proof of concept. EmbeddingGemma 2 is used frozen and unchanged; perception is cosine similarity between crop embeddings and fixed text prompts; a teacher-shaped rule set picks the buttons. The prompts, the crop band and the rules were designed for this one scenario while looking at its frames, and each game result is a single 60-second run on one seed. The policy reads pixels only; the game state appears on the dashboard for checking. Every number on this page comes from reports/embeddinggemma2/ or the experiment's README.

1 · Released this morning

October 6 · getting a day-old model running locally

EmbeddingGemma 2 has 740M parameters in three parts: a 270M text backbone, a 170M vision encoder and a 300M audio encoder, all mapping into one 768-dimensional space that can be truncated to 512, 256 or 128 dimensions. Text, images, video frames and audio can even be interleaved into a single input. The repository was not gated, and transformers 5.19.0, the first release that knows the embedding_gemma2 architecture, came out the same day.

On an Apple M3 Max it worked on the first try. The model card warns that float16 overflows its activations, so the Mac runs float32 and the RTX 4090 bfloat16. Metal float32 matched the CPU reference to a cosine of 1.00000, and bfloat16 to 0.99996.

First retrieval checks, known answers, Metal float32Result
Text query → document, including Spanish and German queries8/8, also at 128 d
Caption → photo, among 27 stock photos27/27
Spoken sentence → its paraphrase, including Spanish audio8/8
Spoken descriptions (“A zebra.”) → the right photo, among 276/6
Video → its label4/5
Text → the right two-image interleaved pair9/9
No sense of direction in video. A ball moving left to right and one moving right to left scored the same against both captions (0.739 vs 0.735). Time order did not register.

2 · No training

October 6 · cosine similarity to prompts decides everything

The first plan was the ModernBERT recipe: freeze the encoder and train a small head. The brief changed it to no training at all. Every decision must come from comparing embeddings with text prompts.

Embedding a whole frame does not work: mean-pooling averages away where things are, and asking whether the monster is left, centre, right or absent was right 26% of the time, below the 38% that always guessing the most common answer gets. Crops fix this. The view is cut into pieces and each piece is embedded on its own; the status bar's HEALTH and AMMO boxes become their own images, compared with prompts such as “HEALTH 40%” and “AMMO 12”.

100%
AMMO read exactly from its crop
150 logged frames, 43 of them empty
0.5
points of mean HEALTH error
HEALTH crop vs 21 prompts
9 / 9
instructions mapped to the right mode
three phrasings never written into the prompts

These readings come from the Apple-silicon probe against logged game state (zero-shot-probe-mps-negation-presence.json). The instruction is matched once per run, text to text, against three short behaviour descriptions.

3 · Contrast, not negation

October 6 · what a monster prompt should be compared with

A single prompt's cosine means little on its own; a crop is always somewhat similar to “a monster”. The decision needs a contrast: a positive set of prompts against a negative set, each averaged into one prototype, with the sign of the difference deciding. The natural negative, “is not a monster”, barely works. Embedding models place a phrase and its negation close together, and “is a monster” against “is not a monster” separated monster crops from empty ones with an AUC of only 0.62. Contrasting with what is actually in the scene (“an empty brown stone wall”, “a grey tiled floor”, “a dark night sky”) reached 0.96. Ensembles of three or four prompts beat single prompts in every set we tried.

One fused embedding cannot choose a command. Interleaving the instruction, the game state as text (“health 60%, ammo 14”) and the view into one input, then picking the closest of five command prompts, collapsed: every variant chose the same command for all 300 test situations, 15.7% against a 45.7% majority baseline. The shared text dominates the embedding; frame-to-frame differences barely move it. One question per crop, each against its own contrast, is what works.

👁 The crops (image)

  • 17 overlapping 128-px windows, 32 px apart, across the band where monsters stand (y 120–245)
  • The HEALTH and the AMMO boxes of the status bar
  • All 19 embedded in one batch per decision

# The prompts (text)

  • “a monster”, “a pink demon”, “a zombie soldier” vs wall, floor and sky
  • “HEALTH 0%” … “HEALTH 100%” and “AMMO 0” … “AMMO 50”
  • Three behaviour descriptions for the instruction

The decisions

  • A monster is visible if any window's margin is above zero
  • The best window, refined by a parabola through its neighbours, gives the bearing
  • Fixed rules: turn toward it, fire within ±0.12 of the crosshair, back off when hurt
No fitted thresholds. Every decision is the sign of a cosine difference or an argmax across crops. The crop band came from the logged monster boxes: their tops cluster near y 191 and their bottoms near y 233.

4 · Firing at its own gun

October 6 · the first runs on the RTX 4090

The first 60-second attack run on the GPU, which still decided presence with the “not a monster” contrast, scored 12 kills and 7 deaths. No monster was in view for 868 of its 1,050 decisions, yet the policy saw one in 860 of them and held ATTACK 96% of the time. Each shot raises the pistol, and its muzzle flash reaches into the centre crops, where it matches “a monster”: firing kept the policy firing, and monsters walked up from the sides.

Two frames of an empty room with the pistol raised and a bright muzzle flash in the centre
Game seconds 4 and 12 of that run: an empty room by the game's own state, and ATTACK pressed in both. The recoiling pistol and its flash sit in the middle of the crop band.
The first 20 seconds of that run at normal game speed, on the generic dashboard; it locks onto the centre and keeps shooting.

Two changes fixed it. The weapon is no longer rendered (set_render_weapon(False)), a display setting that gives the policy no extra information. And presence now comes from the scene contrast itself: a monster is visible when some window is closer to the monster prompts than to wall, floor and sky.

Perception on the same 150 logged frames, RTX 4090“Not a monster” ruleScene rule
Monster visible or not, correct (always answering yes: 75.3%)75.3%80.7%
Empty views correctly rejected2.7%94.6%
Fire decisions with a monster under the crosshair52.6%89.3%
Bearing error to the closest monster, median0.0260.025

Both probes read AMMO exactly in 99.3% of frames, with a mean HEALTH error of 1.0. The scene rule misses about a quarter of the frames that contain a monster; the policy then turns to search, which usually brings the monster back into view. The next 60-second attack run scored 37 kills and 2 deaths.

5 · Evade is its own policy

October 6 · new contrasts on the same crops

With the scene rule, an evade policy that was simply the attack rules with firing turned off died 5 times in 60 seconds: when it saw nothing it turned to search, and monsters reached it in the meantime. Evade became its own policy, with its own contrasts on the same window crops.

ContrastPositive promptsNegative promptsDecides
danger“a monster right in front of you, up close”, …“a small monster far away”, …close: face it, back off, strafe away
escape“an empty open floor”, “a clear path …”, …the monster promptsfar: turn toward open space

It never fires, keeps circling backwards while it scans, and the dashboard shows the danger reading and an escape marker. The showcase's evade run survived all 60 seconds without a shot.

37 / 2
attack: kills / deaths
60 game seconds, seed 7
0 deaths
evade, its own policy
attack rules without firing: 5 deaths
16 / 1
the trained ModernBERT, for reference
reads game state as text, not pixels
Cautious adds nothing here. A cautious run (“retreat when health falls below 40”) matched an attack run exactly, 37 kills and 2 deaths: with this seed its 40-health rule never changed a decision.

6 · How many embeddings per second

October 6 · RTX 4090, bfloat16, through sentence-transformers

InputOne at a timeBatched
Short text, 8 tokens36/s10,179/s at 1,024
Passage, 130 tokens36/s1,004/s at 64
Image, 70 vision tokens21/s134/s at 19
Image, 140 vision tokens (the Doom crops)18/s71/s at 8
Image, 280 vision tokens (the model default)18/s33/s at 64
Image + short text, 140 vision tokens21/s73/s at 8

Images cost far more than text: each vision token is built from nine image patches, so a 140-token crop sends about 1,260 patches through the vision encoder. Adding text to an image costs almost nothing. With 19 crops per decision at about 70 images per second, the policy decides about three and a half times per second before the dashboard and recording. Rates include CPU preprocessing; nothing was compiled or tuned (throughput.json).

7 · What this does not show

The boundaries of a same-day proof of concept

Nothing was trained, but the prompts, the crop band and the rules were designed for one scenario while looking at its frames, and each result is one 60-second run on one seed. The rules decide the strategy; the embedding model supplies perception and reads the instruction. Strategy, distance and numbers beyond the HUD's own digits are not tested, and neither is any other map.

  1. Lower image detailAt 70 vision tokens the encoder runs about twice as fast; whether perception holds up is untested.
  2. More seeds and scenariosRepeat the attack and evade runs across seeds and other ViZDoom maps before reading anything into small differences.
  3. CompareThe text-only approaches are in the ModernBERT and GLiNER2.5 journals.