← MuJoCo Sandbox

Engineering journal · June–September 2026 · Rust/Candle CUDA · MuJoCo 3.13 · RTX 4090

Teaching JEPA world models to fly a drone

A world model learns how a quadrotor responds to its rotor commands, and a sampling planner uses those predictions to choose what the motors do next. The work began as a CUDA-native port of LeWM, became SkyJEPA, a drone model rebuilt from a paper that makes a hand-written controller fly more accurately, and grew into the JEPA Gym, a MuJoCo test bench where a new JEPA architecture can be trained and flown against matched baselines in one evening.

The JEPA Gym, September 18. One X-frame drone flown by two matched JEPA models, the plain baseline (left) and the OPF variant adapted from JEPA-Anything (right): first with the physics readout, then the learned readout, then turning in place, and finally the five-seed results. Real-time replay of recorded MuJoCo poses and model predictions. The test domain and training seed were fixed before evaluation, and propeller animation is decorative.
756 / 756
SkyJEPA tracking runs passed across three training seeds and four test conditions, with no ground contact
26–34%
lower worst-case tracking error when SkyJEPA guides the hand-written controller (mean 4.4–6.7% lower)
1,344
learned-model flights across the JEPA Gym's two studies, all with zero contact steps
3.3 ms
p95 prediction and scoring with the learned readout on an RTX 4090, versus about 14 ms with the physics readout
3 h 48 min
from the first look at JEPA-Anything to the verified five-seed JEPA Gym study
Where this stands. Simulation only, in two simulators: the project's native rotor simulator for SkyJEPA and MuJoCo 3.13 for the JEPA Gym. The world models are learned; the geometric flight controller is written by hand. In every flight that controller proposes motor commands and MPPI uses the learned predictions to choose among variations of them, so the models steer the drone rather than fly it alone. The JEPA-Anything models are this project's adaptations of its OPF method, trained here, not the authors' checkpoints. Nominal-physics MPPI and the geometric controller remain the references to beat. Every number on this page comes from files in the repository.

1 · A world-model runtime in Rust and CUDA

June 2–15 · porting LeWM, then pointing it at a drone

The project started as a port. LeWM is a JEPA world model from the open stable-worldmodel project: it encodes an observation into a latent state, predicts how that state changes under candidate actions, and scores the predictions against a goal. We rebuilt its runtime in Rust on Candle CUDA tensors so that decoding, encoding, rollout, scoring and action selection all stay on the GPU. CEM, MPPI and iCEM planners run there too.

The port was checked against the official Python implementation on the same RTX 4090 before any speed claim. Over a fixed PushT batch, the largest difference in predicted latents was 7.3 × 10⁻⁴, and the cost ranking picked the same best action.

Runtime diagram: observations decoded on the GPU, a LeWM checkpoint, the Candle CUDA runtime and CUDA planners
The runtime: observations are decoded and preprocessed on the GPU, a LeWM checkpoint is loaded into Candle CUDA, and the planners score candidate action batches without leaving the device.
Stage, image PushT planning (iCEM, 1,024 samples, 5 iterations)Rust/Candle vs Python/PyTorch
Media decode and preprocessing (nvJPEG and CUDA kernels)3–4× faster
Image encoding1.37–1.51× faster
iCEM planning1.13× faster
Selected-score evaluation1.66× faster

From June 12 the same runtime learned a drone. A vector-observation LeWM trained on recorded racing-drone flights, a Bevy simulator replayed them, and a gate-loop test asked the planner to fly the course. Fitting the simulated plant to the recordings let the recorded expert actions complete all four gates, up from two. Plan traces then showed where the remaining work was: the model's objective often preferred its own action sequence to the expert's while the drone drifted off the recorded path. The next model would need to be built for control from the start.

Measure the plant apart from the planner. Replaying expert actions through the simulator tests the simulator alone. Once that replay flew the full lap, any remaining failure belonged to the model and the planner, and that is where the project went next.

2 · SkyJEPA, rebuilt from a paper

September 3 · commits from 17:16 to 22:10, from the paper to a trained controller

The SkyJEPA paper describes a JEPA world model for quadrotors. Its authors' model, training code and data were not available, so the implementation was reconstructed from the paper, keeping the paper's interfaces and sizes separate from the choices made here. It runs natively in the Rust/Candle runtime with a dedicated 200 Hz rotor-physics simulator and a 20 Hz controller.

Learned latent dynamics

  • Ten past states of 18 values each (position, velocity, attitude, angular rate) pass through a causal TCN into a 24-value latent
  • Windows of four rotor forces pass through a second TCN
  • A GRU rolls the latent forward 20 steps; loss is multi-step latent error plus SIGReg

Physics-inspired prober

  • Trained second, on the frozen latent model
  • Maps each predicted latent to motion through differentiable SO(3) integration
  • Turns latent predictions into positions the controller can score

Control

  • MPPI samples 512 rotor-command sequences over 15 steps (0.75 s)
  • A hand-written geometric controller supplies the starting sequence
  • The learned predictions decide how to adjust it

Two problems were solved the same day. The first dataset was rejected because it did not excite the rotors differently enough to learn from, and was replaced by 2,000 ten-second flights across 100 randomized vehicles. Sampling raw rotor forces also proved a poor way to start each plan, so a trim-aware geometric controller now proposes the first sequence and MPPI improves it. That hybrid design stayed for the rest of the project.

The first trained controller passed all 63 test flights, against 62 for the geometric controller alone. Its worst trajectory error was 0.41 m, against 0.89 m, at a planning p95 of 8.69 ms.

The native SkyJEPA simulator: a drone tracing a yellow figure-eight trail
September 3: the first trained SkyJEPA controller flying a figure eight in the native rotor simulator. The yellow trail is the executed path.
Let the hand-written controller stabilize; let the model improve it. Keeping the drone in the air and predicting its dynamics are different jobs. Giving each its own part made the learned model's contribution measurable against the same controller flying alone.

3 · Results that hold up

September 4–5 · an audit, a fixed protocol and three training seeds

Before claiming more, the pipeline was audited and hardened. Test vehicles now come from physical configurations never seen in training. Every configuration contains both hover and moving flights. Checkpoints are bound to their normalization, timestep and training history, and interrupted training resumes exactly. The evaluation protocol, with three training seeds, the test populations and the controller settings, was fixed before any final test ran.

Each model trained in about 34 minutes on a shared RTX 4090, from 1,600 ten-second flights. Every controller then flew the same 63 hover, circle and figure-eight cases on unseen vehicles.

SkyJEPA seed 7 flying a randomized drone on a figure eight, 20 s at normal speed. Yellow: executed path. Cyan: reference. Magenta: predicted path. Green bars: commanded rotor forces. On-screen timing includes rendering; the headless benchmarks are the timing reference.
756/ 756
tracking runs passed, no ground contact
Three seeds × four conditions, including ±10% hover-calibration error and heavier, slower-motor drones
26–34%
lower worst-case tracking error than the controller alone
Mean error 4.4–6.7% lower
≈⅔
less tracking error than the same planner with random weights
Training clearly teaches useful dynamics
0
missed 50 ms control deadlines
741 of 756 runs also met the stricter 10 ms p95 target
Controller, 63 test flightsPassesMean RMSEWorst RMSEPlanning p95
Random-weight model + MPPI51 / 630.604 m1.864 m8.93 ms
Hand-written controller alone63 / 630.215 m0.571 m0.0015 ms
Trained SkyJEPA + controller, three seeds63 / 63 each0.201–0.206 m0.378–0.423 m8.82–8.94 ms
Nominal-physics MPPI + controller63 / 630.161 m0.304 m3.22 ms
The reference still ahead. Nominal-physics MPPI plans with the drone's equations of motion and nominal parameters. In this clean simulator it tracked best and planned fastest, which makes it the target for every learned model that follows.

This is the work behind the short film Self-Improving Agent Trains a UAV with JEPA: an agent read the paper, wrote the code and trained a drone controller that flies more accurately than the hand-written one.

4 · Trying JEPA-Anything

September 18, 18:11–18:54 · a new architecture, compared in under an hour

On September 17 the JEPA-Anything paper introduced Orthogonal Predictive Factorization (OPF). OPF splits the latent state into complementary groups of coordinates, predicts each group, and recombines them into the next state, with the aim of longer, more stable predictions. The release turned out to be larger than its GitHub page suggested: a Hugging Face repository holds 36 checkpoint files. Its core library and three released locomotion models ran on our RTX 4090.

The goal was to compare it with SkyJEPA, not replace it, and to film the comparison. OPF went into SkyJEPA's own structure: the same encoders, the same 24-value recurrent state and four factor groups of six coordinates recombined through a learned orthonormal basis. Its latent model trained in PyTorch in 24 seconds, then a new importer brought it into the Rust runtime, matching the Python rollout to within 9 × 10⁻⁸ over 20 steps. The prober and all flying ran natively.

SkyJEPA (left) and the OPF adaptation (right) flying the same drone from the same recorded data with the same planner. A circle and a figure eight, each 20 simulated seconds at 1× speed, then eight seconds of aggregate results. The video seed was fixed before either model was evaluated.
Model, 63 flights per conditionUnseen drones, mean RMSEHeavier, slower motorsPlanning p95
SkyJEPA0.198 m0.206 m5.5 / 5.7 ms
JEPA-Anything OPF adaptation0.236 m0.233 m6.3 / 6.2 ms
Nominal-physics MPPI0.170 m0.166 m0.7 / 0.7 ms

All three passed 126 of 126 tracking cases with no ground contact and no missed deadlines. SkyJEPA held its lead over this first OPF adaptation in both conditions. Open-loop prediction over 26,400 held-out windows tells the same story, and shows what a better architecture would need to fix:

Prediction horizonSkyJEPAOPF adaptationConstant velocity
0.25 s0.031 m0.081 m0.054 m
0.75 s, the planning horizon0.286 m0.414 m0.453 m
1.00 s0.511 m0.669 m0.767 m
3.00 s8.16 m8.82 m3.32 m
Results screen: SkyJEPA, OPF adaptation and nominal physics tracking errors on unseen and heavier drones
The film's results screen, drawn directly from the benchmark files.
A fixed protocol turns a new paper into an evening's experiment. Because the data, splits, planner budget and baselines were already fixed, trying OPF meant writing an adapter and an importer, not a new evaluation. The first comparison video was finished 41 minutes after the session began.

5 · Building the JEPA Gym

September 18, 19:37–21:15 · a charcoal engineering gym for robots, in MuJoCo

To make comparisons like this routine, the work moved to MuJoCo 3.13. The brief was a gym for robots: abstract, with training aids and overlays, but good to look at. It became the JEPA Gym: a charcoal floor grid, landing pads, reference hoops, lane markings and live telemetry. MuJoCo simulates rigid-body motion and ground contact at 400 Hz. Each motor has first-order lag, thrust gain and reaction torque, the airframe has linear drag, and control runs at 20 Hz. The training aids have no collision response.

The JEPA Gym: a charcoal drone hovering among reference hoops on a dark grid
The charcoal flight lab: the drone, reference hoops and the target path (teal). Rendered from recorded MuJoCo states.

Looking at the drone shaped the gym as much as the tests did. Watching it fly showed that it never turned to face its path, so heading control was added: face the direction of travel, or turn in place independently. A closer look showed landing skids that did not follow the nose and rotors in a plus layout. The skids were realigned, and the drone became an X frame: diagonal arms, motors 0.24 m from the centre at 45° to the body axes, the nose between the front motors, and a motor mixer that matches those positions.

Airframe review of the first plus-frame drone from four views
Before: the first drone, a plus frame with one rotor at the nose.
Airframe review of the X-frame drone from four views
After: the X frame, with diagonal arms and the nose between the front motors.
Heading control on the X frame: facing the path on a figure eight, then turning in place, 10 s in real time. This feature demonstration is flown by the geometric controller, not a learned model. On the first drone, the same checks covered a 514° yaw sweep and a 20 s figure eight with 5.9 cm position RMSE, both without contact.

The gym briefly supported several airframes, which made it more complex than intended. It was cut back to one X-frame drone and one motor mixer. The simplified code reproduced the recorded flights exactly, with zero state difference. On a Mac, an interactive viewer opens the same gym, with live telemetry and replay of recorded learned-model flights.

Look at the robot. The skids and the rotor layout passed every numerical test; they were caught by watching the drone. Visual review sits alongside the measurements in this gym, not after them.

6 · The first gym study

September 18, by 20:35 · three seeds, the first drone, a promising signal

The gym compares two matched models: a plain LE-WM JEPA baseline and the same architecture with OPF. Both are trained on the same data, predict 0.75 s ahead, and fly through the same MPPI controller and geometric action prior. Each representation is then frozen and given two readouts. The physics readout learns corrections inside supplied equations of motion. The learned readout predicts future states directly, without the physics integrator.

The first study, with three training seeds on the original plus-frame drone, gave OPF a promising signal:

8.6–9.9%
lower prediction error with OPF and the learned readout
Held-out heavy and slow drones, and ordinary test drones
7.9%
lower action regret with OPF on ordinary test drones
Learned readout. Per seed on held-out drones, regret fell for seeds 7 and 17 and rose for seed 29
384/ 384
matched learned-model flights without ground contact
Both models, both readouts
32/ 32
flights by the original SkyJEPA, transferred unchanged
5.6 cm tracking on a new simulator and rotor geometry; the geometric controller had 5.0 cm
Side-by-side flights of the JEPA baseline and OPF variant with the learned readout
Learned readout, no physics integrator: baseline (left) and OPF (right) on a held-out heavy, slow-motor drone.
Results table of the three-seed study
The three-seed results screen.
Worth running properly. Three seeds, one drone that was about to change, and a seed-to-seed spread in action regret: enough to justify a larger study on the corrected X frame, with heading changes in the training data.

7 · Five seeds, one X-frame drone

September 18, 20:59–21:59 · a fixed protocol, 1,008 flights, verified

The protocol was fixed before training: seeds 7, 17, 29, 43 and 59; 864 ten-second episodes across 72 physical configurations, balanced across fixed, path-facing and turning-in-place heading; 48 configurations for training and eight each for validation, ordinary testing and a held-out region of heavier drones with slower motors. Both models have 40,928 trainable parameters and read ten past states and nine motor commands. Checkpoints were chosen on validation data only, and no test result tuned anything. The full run of data, training and evaluation took 50 minutes 32 seconds on the RTX 4090.

Physics readout flights, baseline and OPF
Physics readout, path-facing figure eight.
Learned readout flights, baseline and OPF
Learned readout, same flight, no physics integrator.
Learned readout, turning in place
Learned readout, turning in place.
960/ 960
learned-model flights with zero contact steps
5 seeds × 2 models × 2 readouts × 48 flights; plus 48 reference flights
2.3%
lower action regret with OPF, physics readout
0.8% with the learned readout; held-out heavy, slow drones
3.3ms
p95 prediction and scoring, learned readout
About 14 ms with the physics readout; excludes simulation and prior construction
3,072
MuJoCo rollouts to score the models' choices
96 saved states × 32 candidate action sequences
Held-out heavy and slow dronesAction regretPrediction RMSETracking RMSEHeading RMSE
Physics readout · JEPA baseline0.2410.119 m3.13 cm0.97°
Physics readout · OPF variant0.2350.143 m3.48 cm0.97°
Learned readout · JEPA baseline0.5280.502 m6.67 cm1.00°
Learned readout · OPF variant0.5240.599 m7.60 cm0.99°
Geometric controller alone, reference––2.25 cm–

Action regret measures how close a model's top-ranked action comes to the best one actually available, scored by MuJoCo itself; lower is better. OPF picked slightly better actions with both readouts. The learned readout flew every flight to completion without the supplied physics equations, more than four times faster to score than the physics readout.

Results screen of the five-seed study
The film's closing screen: all five seeds, not only the flights shown.
What we hoped for. We hoped OPF's better action ranking would also mean tighter tracking. In this study it did not: OPF's tracking error was 11–14% higher and its prediction error about 20% higher, and the geometric controller alone, at 2.25 cm, remains the number to beat on these clean simulated flights. One likely reason is that MPPI flies a weighted blend of many candidates, so a better top pick does not necessarily change the command it sends.

8 · What we learned

Across June to September

The pipeline is the asset. Fixed data, splits, controllers, readouts and baselines mean a new JEPA architecture can be trained, flown and filmed against the same references in an evening. For JEPA-Anything that took 3 hours 48 minutes, from first look to a verified five-seed study.
Learned models help most in the hard cases. SkyJEPA's largest gain over the hand-written controller was in the worst flights, 26–34%, rather than the average. Edge cases are where learned dynamics earn their place.
Better ranking is not automatically better flying. OPF ranked actions better and tracked worse. The planner's averaging sits between the model and the motors, so a model has to be measured in the closed loop, not only by its predictions.
Supplied physics is valuable, and so is speed. The physics readout predicts four times more accurately than the learned readout; the learned readout scores four times faster. Both flew every flight, which leaves room to combine them.
Strong references keep the claims honest. Nominal-physics MPPI and the geometric controller are hard to beat in a clean simulator because they know its equations. That is also why learned dynamics should be tested where those equations stop being right.
Models transfer. SkyJEPA, trained on another simulator with another rotor layout, flew 32 of 32 MuJoCo flights unchanged, within 0.7 cm of the geometric controller.

9 · Next: new JEPA architectures

The gym is ready for the next idea

  1. Try each new JEPA architecture as it appearsNew JEPA variants are appearing quickly. The gym, the protocol and the baselines are ready; each new one needs an adapter and an evening, and its results sit next to SkyJEPA's, OPF's and the references'.
  2. Fly where the equations breakGusting wind, slung payloads, a damaged motor or a large unknown mass: conditions where the hand-written controller and nominal physics are wrong, and learned dynamics have room to win.
  3. Predict further aheadAt 3 s, a constant-velocity guess still beat both SkyJEPA and the first OPF adaptation. Architectures designed for long, stable rollouts, the promise behind OPF, are aimed squarely at that gap.
  4. Give the model less helpWeaken or remove the geometric action prior and measure how much of the steering the world model can take over on its own.
  5. More seeds, more motionTrain with heading changes throughout, add seeds, and confirm on a fresh, untouched sample of drones.
  6. Bring winners homePort the best variant into the Rust/Candle CUDA runtime for real-time control, then test it on logs from real flights.

Everything here is simulated. The result is a working path from a new world-model idea to measured flights, and a set of baselines that any future JEPA has to beat.