← World Sandbox

Engineering journal · October 8, 2026 · AMB3R-SLAM · Depth Anything 3 · Seedance 2.5 · RTX 4090

From six seconds of generated video to a 3D map, one stage at a time

A text prompt became a 6 s video of a small robot rolling through a junkyard. AMB3R-SLAM turned its 145 frames, plain RGB from one uncalibrated camera, into a camera path and a 3.33 million-point coloured map in 20.6 s on one GPU. We opened the pipeline into eight stages, recorded what each one produced, and built a viewer that replays them all in 3D.

The SLAM camera (yellow frustum, showing the frame it sees) and its path (orange) as the map grows, with the generated input video in the corner, at 1×; then an orbit of the finished map. Rendered on the CPU from the recorded run; each point appears when the camera passes the frame it came from.
145 frames
6 s of generated 1280 × 720 video at 24 fps: the only input
20.6 s
the whole SLAM run on one RTX 4090, 7.0 frames per second
3.33 M
points in the map, each kept pixel of 37 frames
4 submaps
from 6 feed-forward passes over 11 to 35 views each
7.8 GB
peak GPU memory
What this is. The video is generated from a text prompt by Seedance 2.5; no robot or camera recorded it. AMB3R-SLAM (Wang and Agapito, 2026) runs from a pinned submodule: the authors' commit 5418465 plus two small fixes of ours to its front-end (chapter 4), in a fork; our tracer wraps its methods at run time to keep every intermediate result. The images, figures, film and viewer all come from one recorded run (the figure of the fix compares it with the run before); every number on this page is in docs/synthetic_slam/results/. Distances are in the model's own units: one camera cannot measure absolute scale.

1 · From pixels to a map

What the pipeline does, before the details

SLAM (simultaneous localisation and mapping) answers two questions from a moving camera at once: where was the camera at every moment, and what does the world around it look like? Classical systems such as ORB-SLAM answer them by matching keypoints between frames and refining poses and 3D points together in an iterative optimisation called bundle adjustment.

AMB3R-SLAM takes a different route. Its core is a feed-forward 3D model, Depth Anything 3: hand it a batch of images and, in a single forward pass with no iterations, it returns where each camera was, how far away every pixel is, and how confident it is. The catch is that one pass only fits a few dozen frames. The system's contribution is everything around the model: a light front-end that tracks every frame, and a hierarchical backend that stitches hundreds of passes into one consistent map. These are its eight stages.

1 Frames 1280 × 720 video frames, resizedto 518 × 294: 37 × 21 patches EACH FRAME 2 Front-end tracking DA3-Small on 4 views: keyframe,2 recent, new → a pose per frame EACH FRAME 3 Submaps DA3 Giant, one pass over ~32 views:poses, depth, confidence, intrinsics EVERY 32 FRAMES 4 Stitching Sim(3): scale, rotation, shift,fitted on the frames they share EACH NEW SUBMAP 5 Long-context links one pass over 24 views from up to8 submaps; link only what is covisible EVERY 4 SUBMAPS 6 Loop closure DBoW2 asks "been here before?";a joint pass verifies each answer EVERY 5TH FRAME 7 Pose graph one node per submap; optimiseall Sim(3) links together AT THE END 8 Map drop sky, edges, low confidence;place every pixel with its submap AT THE END
Stages 1 and 2 run for every frame, so the robot always has a pose. Stages 3 to 6 run as frames accumulate. Stages 7 and 8 run once the video ends; the upstream default optimises the pose graph offline.

Input

  • One camera, plain RGB: no depth sensor, no IMU, no LiDAR
  • Uncalibrated: the model estimates the focal length
  • Any frame rate; below 10 Hz the run is told the rate

Models

  • Front-end: DA3-Small, every frame
  • Backend: DA3-Nested-Giant-Large-1.1, encoder weights in bf16 (2.68 GB)
  • Both are feed-forward: one pass, no iterations, no training on this scene

Output

  • A camera pose for every frame
  • A dense coloured point cloud, every kept pixel of the mapped frames
  • Here, every stage's intermediate results too, from our tracer

2 · Six seconds that never happened

A text prompt, one generation, one take

We wanted footage from a small robot's point of view in an industrial setting, and no robot to film it with. Seedance 2.5 generated it from a prompt: a camera 30 cm above the ground rolling down a dirt lane between stacked wrecks, curving left past tyres and a pickup truck, then past a car with its hood open. One generation at 720p, 6 s long, took 178 s on the provider's side, and we used the first result.

The video is 1280 × 720 at 24 fps: 145 frames. Nothing else goes into SLAM: no camera calibration, no depth, no motion sensor.

Eight frames of the generated junkyard video, from the dirt lane between stacked cars to the pickup truck and an open-hood car
Eight of the 145 frames. The prompt and the generation record are in docs/synthetic_slam/.

3 · Stage 1: frames

What the model actually sees

Depth Anything 3 is a vision transformer: it cuts each image into 14 × 14-pixel patches and reasons about those. AMB3R-SLAM resizes every frame so the long side is 518 pixels, here 518 × 294, a grid of 37 × 21 patches. Fine detail below a patch is blurred into its neighbourhood; this is why the map is soft at small scales.

The frame rate sets how many frames each submap spans. At 10 Hz or more, the backend takes about one frame in four; a 24 fps clip is well above that.

A generated frame at full size beside the 518 by 294 model input with its patch grid
Frame 72 at full size, and as the model receives it. The orange square is one 14-pixel patch.

4 · Stage 2: front-end tracking

A pose for every frame, from four views

For each new frame, the small model DA3-Small sees four images together: the current keyframe, the two most recent frames and the new one. It returns their relative poses; the front-end fits a Sim(3) transform, with its scale taken from the keyframe's depth, and that gives the new frame's pose. It ran 144 times, once per frame after the first, for 3.99 s of model time.

A keyframe is the reference the tracker leans on. When less than 35% of the keyframe is still visible in the new frame, or 32 frames have passed, the new frame becomes the keyframe; that happened at frames 0, 17, 47, 61 and 78. And every 32 frames the backend hands the front-end a fresh keyframe: its own newest frame, with its depth.

Four frames of one front-end pass; the overlap curve with keyframe promotions; the front-end's scale estimate over time; the final path from above
Top: the four views of the pass that tracked frame 58. Bottom: how much of the keyframe stays in view, falling until a new keyframe is chosen; the front-end's own scale estimate (log scale), with the frames where the backend hands it new poses; and the final camera path.
A fix along the way. Tracing every stage showed the front-end's live estimate running away from the final path: 21.5 times too long in this clip's first run. Two causes: the backend's hand-off mixed two coordinate systems, and the front-end took its scale from frames it had just placed itself, so errors fed on themselves. We fixed both in our fork of AMB3R-SLAM (rozgo/amb3r-slam, commit 552e17f): the front-end now takes its scale from its keyframe's depth and restarts from the backend's poses at each hand-off. Its live path now stays within 1% of the final one here (RMS), and on the real 13-minute Oxford walk of chapter 12 its error against the laser ground truth fell from 39.7 m to 3.9 m.
The front-end's live scale and live step length against the final path, before and after the fix, and the live and final paths from above after it
The front-end's live estimate before and after the fix. Left: its scale, with the six hand-offs. Middle: how far it moves per frame, against the final path (1 = they agree). Right: after the fix, the live path (dashed) on the final one. Numbers: junkyard_v1_reanchor_fix.json and spires_christ_church_05_live.json in docs/synthetic_slam/results/.

5 · Stage 3: submaps, the feed-forward pass

32 views in, poses and depth for every pixel out

This is where the 3D comes from. Every 32 frames the backend takes the last 96 frames, keeps the front-end's most confident frame in each group of four, adds frames so no gap exceeds four, and hands those 33 to 35 views to DA3-Nested-Giant-Large in one pass. Out come, for every view, a camera pose, a depth for every pixel, a confidence for every pixel, the camera intrinsics and a sky mask. That set, in its own coordinate frame and units, is a submap.

Each pass took 1.79 to 1.92 s on the RTX 4090. Two short warm-up passes (11 and 21 views, at frames 31 and 63) gave the robot a map before the first full window; both were replaced. The model estimated a focal length of about 248 × 239 pixels at the 518 × 294 input, a horizontal field of view of about 93°, which suits a low robot camera.

Six views of submap 2 with their depth maps, confidence maps and sky masks
Six of submap 2's 34 views. Depth: warm is near. Confidence: brighter is surer; the ground is the surest, sky and distant stacks the least. Teal marks the pixels the model called sky.
Submap 2's points in 3D with its 32 camera frustums along the path
Submap 2 alone, from one pass: its views' depths placed with its own poses, with every second camera drawn as a teal frustum.

6 · Stage 4: stitching submaps

Sim(3) on shared frames

Every submap has its own origin and its own scale: a single camera cannot tell a small near scene from a large far one. But consecutive submaps overlap, and the frames they share appear in both. Fitting their two reconstructions of those frames gives a Sim(3) transform (a scale, a rotation and a translation) that places the new submap in the old one's frame. Scale comes from depth ratios; rotation and translation from the shared cameras.

Each new submap is aligned to its predecessor, and to the one before that when they still share frames. The five alignments here shared 18 to 33 frames and fitted with relative residuals of 0.54% to 1.51%, in 0.03 s in total.

A timeline of the six submap passes, their chosen frames, and the alignments between overlapping submaps
Each row is one model pass; dots are the frames it used. Arrows are alignments, with the frames shared and the fit residual. The two hatched passes are the warm-up submaps that were replaced.
The final map coloured by the submap each point came from, with the camera path in white
The map coloured by submap. The four reconstructions overlap and line up where they share frames.

7 · Stage 5: long-context links

Checking far-apart submaps against each other

Alignment between neighbours accumulates small errors, like a chain. Long-context mapping adds links between submaps that are not neighbours: one extra pass over 24 frames spread across a window of up to eight submaps. If two distant submaps both fit that pass well, the pass links them directly. A link is only kept if the two submaps actually see the same place, which ALIKED keypoints test: at least 20% of their frame pairs must share enough matched keypoints.

Here the window covered all four submaps in one 24-view pass (1.24 s). The only far-apart pair, submaps 0 and 3, failed the test: 1 of 72 frame pairs (1.4%). That is right: the robot's first frames and its last frames look at different parts of the yard.

The 24 frames of the long-context window and the closest frame pair between submap 0 and submap 3
Top: the window's 24 frames, bordered by the submap each was compared for (grey: shared by both). Bottom: the most similar pair across the two submaps, frames 28 and 96.

8 · Stage 6: loop closure

"have I been here before?"

On long routes, the strongest correction comes from recognising a place seen earlier. Every fifth frame the backend adds the image to a DBoW2 bag-of-words database (ORB features against a fixed vocabulary) and asks for the most similar earlier frame at least 100 frames back. A candidate is then verified by reconstructing both places together in one pass, and checking that they overlap and fit their own submaps before it becomes a loop edge.

This clip has 29 probes and no revisits: the robot never comes back, so no candidate was proposed. On the 13-minute Oxford walk in chapter 12, the same stage closed one loop.

9 · Stage 7: the pose graph

One optimisation over every link

Each submap becomes a node; each alignment, long-context link and loop becomes an edge that says where one node should sit relative to another. A robust Sim(3) optimisation then moves the nodes until the edges agree as well as they can. With loops, this is where accumulated drift is corrected.

Here the graph has 4 nodes and 5 edges and no loops, and the chain of alignments already agreed: the optimisation moved no submap by more than 0.07% of the path length, or changed its scale by more than 0.78%, in 0.12 s.

The pose graph: four submap nodes with adjacent and span edges and a rejected long-context candidate; bars of how much each node moved
Left: the graph, with each edge's fit residual. Right: what the final optimisation changed.

10 · Stage 8: from depth to a map

Which pixels become points

Each submap pass predicted depth for every pixel. To build the map, the recorder takes one frame in four, prefers the submap where that frame sits nearest the middle, and drops three kinds of pixel: those below the 20th percentile of the frame's confidence, those on a depth edge (a jump of more than 3% within 3 × 3 pixels, where foreground and background blur together), and sky. The rest are placed in the world with their submap's final Sim(3).

From 37 frames that gives 3,330,429 points, and 212,818 when sampling every fourth pixel; exporting took 0.18 s.

One frame and the same frame with dropped pixels coloured: sky, depth edges and low confidence
Frame 80: 57% of its pixels kept; 29% sky and 15% depth edges dropped. The confidence cut removes nothing in this frame: at least a fifth of its pixels share the lowest confidence value, and the cut keeps ties.
The final map from behind the robot's starting point, with keyframe frustums along the path
The finished map from behind the start, with the front-end's 5 keyframes as yellow frustums.
The final map from above and to the side
The same map from above: the lane, the car stacks, the tyre piles and the pickup.

11 · Explore it in 3D

Every stage as an overlay, synced with the video

The viewer replays the run. By default it follows behind the SLAM camera: the yellow frustum is the pose SLAM computed, with the frame it was looking at on its image plane. Every recorded stage is a layer you can switch on: the front-end's four-view window, keyframes, the latest submap pass, the pose graph, loop probes, and the front-end's live estimate. Colour the points by photo, by submap, by confidence or by time. Reveal them as the camera films them, or only once the backend has committed their submap. Ego view looks through the computed camera, with a slider to blend the real frame over the map.

1,000,000 of the map's points, the 145 poses and every stage's record, in three.js. Space plays and pauses; the arrow keys step one frame. Open the viewer full screen.

12 · Speed and scale

Where the time goes, and a real 13-minute walk

The whole run took 20.62 s for 145 frames, 7.0 frames per second, after 11.5 s to load the models. The two models' passes took 14.3 s of it. Throughput, not the video's frame rate, decides whether this is real time: at about 6 to 7 frames per second it keeps up with a camera sampled at 5 Hz, and falls behind a 24 fps clip by 3.4×.

StageSecondsWhat ran
2 · Front-end4.94144 DA3-Small passes over 2 to 4 views (3.99 s in the model)
3 · Submaps9.286 DA3 Giant passes over 11 to 35 views
4 · Stitching0.035 Sim(3) alignments, and 2 for the warm-up
5 · Long context1.621 pass over 24 views, 1 covisibility check
6 · Loops0.3229 retrieval probes
7 · Pose graph0.124 nodes, 5 edges
Everything else4.32image handling, re-anchoring, bookkeeping
8 · Map export, after the run0.183,330,429 points

Before this project we ran the authors' code on a real recording, for scale: the 13-minute Christ Church walk from the Oxford Spires dataset (2024), one camera of a handheld rig at 5 Hz, 816 m on foot, with laser-scanned ground truth. The tiles are the authors' code; the live estimate on the same walk, before and after our fix, is in chapter 4.

3,961frames
800 s of camera at 4.95 Hz, one camera
processed in 693.6 s: 1.15× faster than the recording
1.86m
trajectory error over an 816 m walk
after a Sim(3) fit; the unaligned map is 1.43× off in scale
9.1GB
peak GPU memory for 3,961 frames
7.8 GB for this clip's 145: memory stays bounded
1loop
closed on the Oxford walk
0 here: the junkyard robot never returns

13 · When to use it

What it is good at, and where something else fits better

A good fit

  • Ordinary video from one camera, even uncalibrated: a phone, a dashcam, a drone, a generated clip
  • You want a dense coloured map, not sparse feature points
  • Long routes: memory stays bounded and loops correct drift
  • Scenes with moving people or cars: there is no bundle adjustment that assumes a static world
  • A GPU with about 10 GB or more

A poor fit

  • Tracking at 30 Hz or more, for control: about 6 to 7 frames per second here
  • Running on a robot's CPU or an embedded board: classical visual-inertial SLAM is far lighter
  • True metric scale from one camera: use its RGB-D, stereo or LiDAR modes, or an IMU

14 · Swapping a stage

Where each stage lives, and what could replace it

The tracer (src/world_sandbox/slam/trace.py) wraps the methods below without editing them, and saves what each returns; the two fixes we made to the code itself (chapter 4) are in our fork. That record is the contract for replacing a stage: a new front-end, model or retrieval method can be run on the same video and compared stage by stage, in the same figures and the same viewer.

StageUpstream codeWhat the trace keepsCould be swapped for
1 · Framesdatasets/demo.pythe 145 input images at 518 × 294another resolution or undistortion
2 · Front-endFrontEnd.trackpose, keyframe, views, overlap and scale per framea stronger small model, IMU fusion, classical odometry
3 · Submaps_new_submap → model.DA3poses, depth, confidence, sky and intrinsics per viewVGGT-Ω (--model_name omega, built in), other feed-forward models
4 · Stitching_align_pair, tools/align.pyeach Sim(3), frames shared, residualother robust estimators
5 · Long context_long_context_window, tools/covis.pywindows, covisibility scores, edgesanother covisibility test
6 · Loops_loop_query, _verify_loopprobes, scores, candidates, verdictsMegaLoc or SALAD retrieval (built in)
7 · Pose graph_optimise, tools/pose_graph.pysubmap poses before and afteranother optimiser, online mode
8 · MapMapRecorder (our subclass)points with source frame, submap and confidencemeshes, Gaussian splats
  1. The front-end, furtherKeep the live pose smooth where it snaps back at a hand-off (up to 3.2 m on the Oxford walk), and confirm on more sequences that the fix leaves the map as it was (on Oxford the final path moved from 1.85 m to 2.29 m).
  2. A longer generated loopA clip that returns to its start, to exercise loop closure and the pose graph on synthetic input.
  3. VGGT-Ω against Depth Anything 3The same video and tracer with the other backend model, compared stage by stage.
  4. The viewer on a real captureThe tracer and viewer work on any folder of frames.