← Doom Sandbox

Engineering journal · September 19 and October 6, 2026 · ViZDoom, GLiNER2.5, ModernBERT, RTX 4090

Teaching GLiNER2.5 to play Doom: zero-shot first, then fine-tuned

GLiNER2.5 scores text against labels it is given at run time, so it can rate Doom's seven buttons with no Doom training. The first attempt, zero-shot and with small adapted heads, fell short of the trained ModernBERT policy. The second fixed the training data and fine-tuned the whole model on an RTX 4090 in under four minutes. It then matched the scripted teacher on every held-out test state and new phrasing, passed all eight conditional probes, and played six evaluation games without dying.

The 3 min 27 s showcase, recorded on an RTX 4090 and played back at normal game speed: GLiNER2.5 Base after full fine-tuning, in complete 60-second runs. Attack: 20 kills, 1 death. Evade: no shots, 1 death. Cautious (“retreat when health falls below 40”): 17 kills, 1 death. The dashboard shows the model's own seven sigmoid scores on every decision; a decoder only blocks impossible button combinations.
100%
held-out test states and new phrasings, fine-tuned Base (first attempt: 75% and 71% at best)
8 / 8
explicit conditional probes, “fire only at 40+ health” (first attempt: at most 4 / 8)
0 deaths
in six 20-second evaluation games, both fine-tuned sizes
232 s
to fine-tune all 193.6M parameters of Base on 12,000 examples, RTX 4090
12.5 ms
per decision, the full encoder every time; 54 decisions/s headless
Where this stands. Both attempts imitate a scripted teacher from the game state described as text. GLiNER never sees pixels and adds no rules of its own; a decoder only blocks impossible button combinations. The second attempt trained on 12,000 teacher-labelled examples, while the ModernBERT reference kept its original 372-example head and was not retrained on them. Test states, test phrasings and one probe phrasing were never trained on. The RTX 4090 was shared with another inference service, so timings are indicative. Every number on this page comes from reports/gliner2-experiment/ (first attempt) or reports/gliner2-v2/ (second).

1 · The question

September 19 · five pipelines, one set of observations

The ModernBERT policy needed 372 labelled Doom examples to learn its seven button scores. GLiNER2.5 is built to classify text against labels supplied at run time, so it can be handed descriptions of Doom actions and asked which apply, with no Doom training at all. Could it play? And if a little training helped, how little?

Each pipeline scored the same 516 evaluation pairs and played six 20-game-second headless runs: three instructions (attack, evade, cautious) on two seeds, 2,100 live decisions per pipeline.

The pipelines

  • ModernBERT with its trained head (372 examples), the reference
  • GLiNER2.5 Small and Base, zero-shot
  • GLiNER2.5 Small and Base with their classification layers adapted on a few examples

The setup

  • Apple M3 Max, PyTorch 2.14, Metal (MPS) float16, batch one
  • Same observations and legal-action decoder for every pipeline
  • GLiNER inputs are 227–253 tokens with the action descriptions; ModernBERT's are 25–36

What counts

  • Speed: the whole classification, and the headless game loop
  • Agreement: the exact seven-button vector against the scripted teacher
  • Behaviour: kills, deaths, ammo and stillness in the live runs

2 · Speed

September 19 · complete pipelines, not equal-length kernels

PipelineDoom examplesPer classificationHeadless loop
ModernBERT37212.0 ms53.5 decisions/s
GLiNER2.5 Small015.8 ms46.1 decisions/s
GLiNER2.5 Base024.4 ms31.3 decisions/s
GLiNER2.5 Small, adapted head12814.6 ms42.8 decisions/s
GLiNER2.5 Base, adapted head3223.5 ms27.5 decisions/s

Classification covers building the text, tokenizing, the complete encoder, moving scores to the CPU and decoding; loading and ten warm-up calls are excluded. The headless loop includes the game, logs and screenshots, and is total decisions over total wall time. All averages stayed above the 17.5 decisions per second that two tics per decision need at native game speed, although one adapted-Base cautious run fell below it.

On the CPU (float32, four threads), Small averaged 36.6 ms and Base 74.9 ms. CPU and Metal scores differed by at most 0.00062 and 0.00223, with identical first decisions, so the weaknesses below are not a float16 artefact.

Not the video's number. The ModernBERT film's 44 decisions per second included the dashboard and video encoder. It is not the baseline here; no video was recorded for this first comparison.

3 · What it did in the game

September 19 · zero-shot responsiveness, and its limits

On an identical healthy state with ammo and an enemy lined up, zero-shot Base's ATTACK score fell from 0.976 for “Attack enemies on sight.” to 0.022 for “Do not shoot. Evade the enemies.” Its evasion runs used no ammo, and it reacted to where enemies were, although it chose right turns inconsistently.

Kills / deaths, two seeds × 20 game secondsAttackEvadeCautious
ModernBERT18 / 00 / 09 / 1
Small, zero-shot9 / 40 / 40 / 4
Base, zero-shot8 / 30 / 121 / 1
Small, 128 examples14 / 10 / 00 / 0
Base, 32 examples19 / 10 / 00 / 0
Obedient by standing still. Zero-shot Small changed its firing scores with the instruction, but its movement scores stayed too low: it stood still for every decision of both evasion runs and died in each. A no-fire check alone would have hidden this.

Kills and deaths are not a full measure of skill, and these are short runs. Zero-shot Base scored more cautious kills than ModernBERT, but that does not show reliable conditional behaviour; the probes in chapter 5 test that directly.

4 · A little training

September 19 · heads only, 32 to 372 examples

Both GLiNER sizes were tried with 32, 128 and 372 training examples, keeping the encoder frozen and adapting only the existing classification layers, 600 steps each with the same settings. The smallest budget reaching the best validation agreement went on to gameplay: 128 for Small and 32 for Base. The original whole-state splits were kept, and the test sets chose nothing.

Exact agreement with the teacherHeld-out states, 48 pairsNew phrasings, 48 pairs
ModernBERT81.25%56.25%
Small, zero-shot2.08%2.08%
Base, zero-shot4.17%4.17%
Small, 128 examples75.00%70.83%
Base, 32 examples70.83%41.67%

Head optimisation took 0.75–1.19 seconds per Small budget and 1.79–2.95 seconds per Base budget, plus 8.46 and 13.89 seconds of shared feature extraction. Both adapted heads moved and survived both evasion runs without firing.

Cautious became silent. Both adapted heads never fired in their cautious runs, even at full health.
Small numbers. The held-out set has eight distinct states, two positive ATTACK examples and no positive MOVE_FORWARD. Small's new-phrasing result is promising for this narrow test, not proof of better general instruction understanding.

5 · Conditional instructions

September 19 · “fire only when health is at least 40”

“Shoot at enemies, but retreat when health falls below 40” never forbids firing while retreating. To remove that ambiguity, a follow-up tested the explicit known phrasing “Fight, but stop shooting and retreat below 40 health” and one new phrasing, at health 20, 39, 40 and 80, with no model or threshold changed.

6 / 8
ModernBERT
all four known-phrasing cases; missed firing above 40 under the new phrasing
4 / 8
Small and Base zero-shot, Small adapted
results changed substantially with wording
3 / 8
Base adapted
32 training examples

The health categories already split at 40, so these probes test the vocabulary, not general numerical reasoning. Neither approach establishes arbitrary instruction following.

6 · First verdict

September 19 · what the first comparison supported

Three bar panels comparing the five first-attempt pipelines: decisions per second, agreement on held-out states and on new phrasings
The first attempt on an Apple M3 Max, drawn from reports/gliner2-experiment/summary.json.
Zero-shot works, but was not an upgrade. GLiNER2.5 Base attacked and evaded with no Doom training, which makes it a good subject for a zero-shot demonstration. ModernBERT stayed faster and more reliable on the task it was trained for, and adapted heads fell silent on conditional instructions. The evidence did not support a faster-or-better claim. Two weeks later we tried again.

7 · Why the first attempt fell short

October 6 · reading the first attempt's data and training

Three things limited the first attempt, and none of them was GLiNER itself.

The data could not show the rule

  • The 468 controlled pairs only ever said health 10, 30 or 100
  • Nothing near 40, so “retreat below 40” had no boundary to learn
  • Ammunition was always 0 or 10

Too few situations

  • 62 distinct training states, at most 372 examples
  • One fixed bearing and one distance per enemy category
  • Two phrasings per instruction

A frozen encoder

  • Only the small classification layers were adapted
  • The encoder's features were never trained to combine the instruction with the state
  • So “retreat below 40” had to be read off features that did not represent it

8 · Better data from the same teacher

October 6 · 12,000 labelled examples, nothing hand-labelled

The scripted teacher can label any state, so the second attempt generated them (teacher_data.py). Each sample draws a health value from 1 to 100, half of them near the 20 and 40 boundaries; ammunition from 0 to 50, empty a quarter of the time; and an enemy that is absent or to the left, ahead or right, at a random bearing and distance within its range category. The teacher labels every state for all three behaviours, and each label is paired with one of five or six training phrasings: 4,000 states, 12,000 examples, plus 1,200 for validation.

Keep the tests honest. The original test and validation state combinations (health and ammo categories, enemy side and range) never appear in training; validation draws only from the validation combinations. The fixed new test phrasings and one explicit conditional probe phrasing are never trained on, and the script refuses to run if one leaks in.

9 · Fine-tuning the whole model

October 6 · an RTX 4090, minutes per model

Training runs through GLiNER's own scoring path, the same calls as inference without inference mode, so the model is trained on exactly the text and labels it scores at play time; a check confirmed the two paths give identical logits. Three epochs, batch 32, AdamW, bfloat16; the checkpoint is chosen on the held-out validation states. To separate the effect of better data from the effect of fine-tuning, the same data also trained only the classification layers with the encoder frozen.

Three bar panels comparing seven pipelines on held-out states, new phrasings and conditional probes
All seven pipelines re-measured on the RTX 4090, float16, batch one (reports/gliner2-v2/summary.json).
PipelineTrainedTrainingTest statesNew phrasingsConditionalPer decision
ModernBERT, 372 exampleshead3.6 s81.2%56.2%6 / 89.6 ms
GLiNER2.5 Small, zero-shot——2.1%2.1%4 / 812.6 ms
GLiNER2.5 Base, zero-shot——4.2%4.2%4 / 812.1 ms
Small, 12,000 exampleshead, 0.3M63 s50.0%50.0%4 / 812.6 ms
Base, 12,000 exampleshead, 1.2M109 s68.8%56.2%5 / 812.3 ms
Small, 12,000 examplesall, 73.9M111 s100%97.9%8 / 812.8 ms
Base, 12,000 examplesall, 193.6M232 s100%100%8 / 812.5 ms
Fine-tuning, not just more data. With the encoder frozen, 12,000 examples still left Small at 50% and Base at 69% on the held-out states, and neither passed the conditional probes. Updating the encoder took both to 100%: the instruction and the state have to be combined inside the encoder, not after it.

Both fine-tuned models passed all 12 behaviour checks, and on the explicit conditional probes, “Fight, but stop shooting and retreat below 40 health” and the never-trained “Fire only when your health is at least 40. If lower, do not fire and move backward.”, they held fire and backed away at health 20 and 39 and fired at 40 and 80. Speed is unchanged by fine-tuning: the same encoder runs on every decision, at about 54 decisions per second headless. ModernBERT, with its 25–36-token inputs against GLiNER's 227–253, is a little faster.

10 · Playing it

October 6 · six evaluation games, then three recorded minutes

Kills / deaths, seeds 7 and 19 × 20 game secondsAttackEvadeCautious
ModernBERT18 / 00 / 09 / 1
Small, zero-shot9 / 40 / 40 / 4
Base, zero-shot8 / 30 / 121 / 1
Small, head only, 12,000 examples0 / 00 / 30 / 0
Base, head only, 12,000 examples9 / 00 / 20 / 3
Small, fine-tuned20 / 00 / 019 / 0
Base, fine-tuned20 / 00 / 019 / 0

The two fine-tuned sizes played identical games: both reproduce the teacher's decisions in these states, so the same actions led to the same outcomes. For the showcase, fine-tuned Base played a full 60 seconds per instruction with the dashboard on: 20 kills and 1 death attacking, no shots and 1 death evading, 17 kills and 1 death on the cautious instruction. Over 3,150 live decisions its raw button vector matched the teacher's in 2,686; every disagreement was the direction of the sideways strafe while backing away, which the text cannot always decide (an enemy “directly ahead” may be slightly left or right of the crosshair). Firing, turning and moving forward or back always matched.

What the cautious film does and does not show. In that run, health fell below 40 only after the ammunition had run out, so its retreat there is also explained by the empty weapon. The 40-health rule itself is shown by the conditional probes above, which test it with ammunition loaded (live-agreement.json, conditional-base-full.json).

11 · Verdict

What the two attempts support

Fine-tuned, GLiNER2.5 plays Doom from instructions. With data that shows the rules and an encoder that can learn them, both sizes matched the scripted teacher on every held-out state and every conditional probe, followed instruction phrasings they never saw, and played without dying in the evaluation games, in under four minutes of training on one GPU.

This is imitation of a hand-written teacher from a text description of the game, so it shows instruction following within that vocabulary, not general Doom skill. The ModernBERT reference was not retrained on the new data, so the comparison shows what each pipeline reached with its own training, not which encoder is better.

  1. Retrain ModernBERT on the same 12,000 examplesThe fair architecture comparison.
  2. From text to pixelsThe same question without any training, reading the screen instead; see the EmbeddingGemma 2 journal.
  3. The trained referenceHow the ModernBERT policy was built is in the ModernBERT journal.