1 · A world-model runtime in Rust and CUDA
June 2–15 · porting LeWM, then pointing it at a drone
The project started as a port. LeWM is a JEPA world model from the open stable-worldmodel project: it encodes an observation into a latent state, predicts how that state changes under candidate actions, and scores the predictions against a goal. We rebuilt its runtime in Rust on Candle CUDA tensors so that decoding, encoding, rollout, scoring and action selection all stay on the GPU. CEM, MPPI and iCEM planners run there too.
The port was checked against the official Python implementation on the same RTX 4090 before any speed claim. Over a fixed PushT batch, the largest difference in predicted latents was 7.3 × 10⁻⁴, and the cost ranking picked the same best action.
| Stage, image PushT planning (iCEM, 1,024 samples, 5 iterations) | Rust/Candle vs Python/PyTorch |
|---|---|
| Media decode and preprocessing (nvJPEG and CUDA kernels) | 3–4× faster |
| Image encoding | 1.37–1.51× faster |
| iCEM planning | 1.13× faster |
| Selected-score evaluation | 1.66× faster |
From June 12 the same runtime learned a drone. A vector-observation LeWM trained on recorded racing-drone flights, a Bevy simulator replayed them, and a gate-loop test asked the planner to fly the course. Fitting the simulated plant to the recordings let the recorded expert actions complete all four gates, up from two. Plan traces then showed where the remaining work was: the model's objective often preferred its own action sequence to the expert's while the drone drifted off the recorded path. The next model would need to be built for control from the start.
2 · SkyJEPA, rebuilt from a paper
September 3 · commits from 17:16 to 22:10, from the paper to a trained controller
The SkyJEPA paper describes a JEPA world model for quadrotors. Its authors' model, training code and data were not available, so the implementation was reconstructed from the paper, keeping the paper's interfaces and sizes separate from the choices made here. It runs natively in the Rust/Candle runtime with a dedicated 200 Hz rotor-physics simulator and a 20 Hz controller.
Learned latent dynamics
- Ten past states of 18 values each (position, velocity, attitude, angular rate) pass through a causal TCN into a 24-value latent
- Windows of four rotor forces pass through a second TCN
- A GRU rolls the latent forward 20 steps; loss is multi-step latent error plus SIGReg
Physics-inspired prober
- Trained second, on the frozen latent model
- Maps each predicted latent to motion through differentiable SO(3) integration
- Turns latent predictions into positions the controller can score
Control
- MPPI samples 512 rotor-command sequences over 15 steps (0.75 s)
- A hand-written geometric controller supplies the starting sequence
- The learned predictions decide how to adjust it
Two problems were solved the same day. The first dataset was rejected because it did not excite the rotors differently enough to learn from, and was replaced by 2,000 ten-second flights across 100 randomized vehicles. Sampling raw rotor forces also proved a poor way to start each plan, so a trim-aware geometric controller now proposes the first sequence and MPPI improves it. That hybrid design stayed for the rest of the project.
The first trained controller passed all 63 test flights, against 62 for the geometric controller alone. Its worst trajectory error was 0.41 m, against 0.89 m, at a planning p95 of 8.69 ms.
3 · Results that hold up
September 4–5 · an audit, a fixed protocol and three training seeds
Before claiming more, the pipeline was audited and hardened. Test vehicles now come from physical configurations never seen in training. Every configuration contains both hover and moving flights. Checkpoints are bound to their normalization, timestep and training history, and interrupted training resumes exactly. The evaluation protocol, with three training seeds, the test populations and the controller settings, was fixed before any final test ran.
Each model trained in about 34 minutes on a shared RTX 4090, from 1,600 ten-second flights. Every controller then flew the same 63 hover, circle and figure-eight cases on unseen vehicles.
| Controller, 63 test flights | Passes | Mean RMSE | Worst RMSE | Planning p95 |
|---|---|---|---|---|
| Random-weight model + MPPI | 51 / 63 | 0.604 m | 1.864 m | 8.93 ms |
| Hand-written controller alone | 63 / 63 | 0.215 m | 0.571 m | 0.0015 ms |
| Trained SkyJEPA + controller, three seeds | 63 / 63 each | 0.201–0.206 m | 0.378–0.423 m | 8.82–8.94 ms |
| Nominal-physics MPPI + controller | 63 / 63 | 0.161 m | 0.304 m | 3.22 ms |
This is the work behind the short film Self-Improving Agent Trains a UAV with JEPA: an agent read the paper, wrote the code and trained a drone controller that flies more accurately than the hand-written one.
4 · Trying JEPA-Anything
September 18, 18:11–18:54 · a new architecture, compared in under an hour
On September 17 the JEPA-Anything paper introduced Orthogonal Predictive Factorization (OPF). OPF splits the latent state into complementary groups of coordinates, predicts each group, and recombines them into the next state, with the aim of longer, more stable predictions. The release turned out to be larger than its GitHub page suggested: a Hugging Face repository holds 36 checkpoint files. Its core library and three released locomotion models ran on our RTX 4090.
The goal was to compare it with SkyJEPA, not replace it, and to film the comparison. OPF went into SkyJEPA's own structure: the same encoders, the same 24-value recurrent state and four factor groups of six coordinates recombined through a learned orthonormal basis. Its latent model trained in PyTorch in 24 seconds, then a new importer brought it into the Rust runtime, matching the Python rollout to within 9 × 10⁻⁸ over 20 steps. The prober and all flying ran natively.
| Model, 63 flights per condition | Unseen drones, mean RMSE | Heavier, slower motors | Planning p95 |
|---|---|---|---|
| SkyJEPA | 0.198 m | 0.206 m | 5.5 / 5.7 ms |
| JEPA-Anything OPF adaptation | 0.236 m | 0.233 m | 6.3 / 6.2 ms |
| Nominal-physics MPPI | 0.170 m | 0.166 m | 0.7 / 0.7 ms |
All three passed 126 of 126 tracking cases with no ground contact and no missed deadlines. SkyJEPA held its lead over this first OPF adaptation in both conditions. Open-loop prediction over 26,400 held-out windows tells the same story, and shows what a better architecture would need to fix:
| Prediction horizon | SkyJEPA | OPF adaptation | Constant velocity |
|---|---|---|---|
| 0.25 s | 0.031 m | 0.081 m | 0.054 m |
| 0.75 s, the planning horizon | 0.286 m | 0.414 m | 0.453 m |
| 1.00 s | 0.511 m | 0.669 m | 0.767 m |
| 3.00 s | 8.16 m | 8.82 m | 3.32 m |
5 · Building the JEPA Gym
September 18, 19:37–21:15 · a charcoal engineering gym for robots, in MuJoCo
To make comparisons like this routine, the work moved to MuJoCo 3.13. The brief was a gym for robots: abstract, with training aids and overlays, but good to look at. It became the JEPA Gym: a charcoal floor grid, landing pads, reference hoops, lane markings and live telemetry. MuJoCo simulates rigid-body motion and ground contact at 400 Hz. Each motor has first-order lag, thrust gain and reaction torque, the airframe has linear drag, and control runs at 20 Hz. The training aids have no collision response.
Looking at the drone shaped the gym as much as the tests did. Watching it fly showed that it never turned to face its path, so heading control was added: face the direction of travel, or turn in place independently. A closer look showed landing skids that did not follow the nose and rotors in a plus layout. The skids were realigned, and the drone became an X frame: diagonal arms, motors 0.24 m from the centre at 45° to the body axes, the nose between the front motors, and a motor mixer that matches those positions.
The gym briefly supported several airframes, which made it more complex than intended. It was cut back to one X-frame drone and one motor mixer. The simplified code reproduced the recorded flights exactly, with zero state difference. On a Mac, an interactive viewer opens the same gym, with live telemetry and replay of recorded learned-model flights.
6 · The first gym study
September 18, by 20:35 · three seeds, the first drone, a promising signal
The gym compares two matched models: a plain LE-WM JEPA baseline and the same architecture with OPF. Both are trained on the same data, predict 0.75 s ahead, and fly through the same MPPI controller and geometric action prior. Each representation is then frozen and given two readouts. The physics readout learns corrections inside supplied equations of motion. The learned readout predicts future states directly, without the physics integrator.
The first study, with three training seeds on the original plus-frame drone, gave OPF a promising signal:
7 · Five seeds, one X-frame drone
September 18, 20:59–21:59 · a fixed protocol, 1,008 flights, verified
The protocol was fixed before training: seeds 7, 17, 29, 43 and 59; 864 ten-second episodes across 72 physical configurations, balanced across fixed, path-facing and turning-in-place heading; 48 configurations for training and eight each for validation, ordinary testing and a held-out region of heavier drones with slower motors. Both models have 40,928 trainable parameters and read ten past states and nine motor commands. Checkpoints were chosen on validation data only, and no test result tuned anything. The full run of data, training and evaluation took 50 minutes 32 seconds on the RTX 4090.



| Held-out heavy and slow drones | Action regret | Prediction RMSE | Tracking RMSE | Heading RMSE |
|---|---|---|---|---|
| Physics readout · JEPA baseline | 0.241 | 0.119 m | 3.13 cm | 0.97° |
| Physics readout · OPF variant | 0.235 | 0.143 m | 3.48 cm | 0.97° |
| Learned readout · JEPA baseline | 0.528 | 0.502 m | 6.67 cm | 1.00° |
| Learned readout · OPF variant | 0.524 | 0.599 m | 7.60 cm | 0.99° |
| Geometric controller alone, reference | – | – | 2.25 cm | – |
Action regret measures how close a model's top-ranked action comes to the best one actually available, scored by MuJoCo itself; lower is better. OPF picked slightly better actions with both readouts. The learned readout flew every flight to completion without the supplied physics equations, more than four times faster to score than the physics readout.
8 · What we learned
Across June to September
9 · Next: new JEPA architectures
The gym is ready for the next idea
Everything here is simulated. The result is a working path from a new world-model idea to measured flights, and a set of baselines that any future JEPA has to beat.