← MuJoCo Sandbox

Engineering journal · September 9, 2026 · MuJoCo 3.12 · NeuralOperator FNO and PINO · PyTorch 2.14

Flying a slung parcel through simulated wind with learned forecasts

A 2.0 kg quadrotor carries a 0.20 kg parcel on a 0.65 m cable that can go slack, through two gates, while an independent Navier–Stokes solver blows gusting crosswinds across its path. The same programmed predictive controller flies every run; only its 2 s wind forecast changes: held frozen, or advanced by an FNO or PINO neural operator trained for this project. The forecasts are the only learned part.

Version 2, seed 400: the same programmed controller with a frozen-wind forecast (left) and the learned PINO forecast (right), in identical 1.5× simulated wind. The full 31 s flight plays at real time, then a 10 s results card. Rendered from saved states of the physical MuJoCo runs; the forecast panel compares PINO with the solver's future wind for evaluation only.
24 / 24
standard deliveries: six held-out winds × four forecasts
4.28 cm
mean parcel tracking error with the FNO forecast, vs 10.89 cm with frozen wind (PINO 4.42 cm)
4.34 cm
the same in version 2's 1.5× wind, vs 19.59 cm frozen (PINO 4.51 cm), five complete paired windows
5 / 6
version 2 missions per method; seed 404 failed for all four and stays on record
73.4 s · 82.3 s
FNO and PINO training, 64 epochs each on one RTX 4090
Where this stands. Finished on September 9, 2026, in two versions, both kept. MuJoCo simulates the aircraft, the cable, the rotors and every contact; an independent 2D Navier–Stokes solver supplies the wind, applied as drag forces; FNO and PINO networks trained here forecast that wind 2 s ahead. The predictive controller and the delivery sequence are programmed, not learned, and they read exact simulator state and the full current wind field, not camera images or wind sensors. The wind is two-dimensional and one-way coupled: the aircraft does not disturb the air. Every number on this page comes from files in the repository.

1 · A static scene first

September 9 · a research brief, then a scene that does not move

The project began with a neural-operator research brief and one question: what could we build in simulation that would make a strong video? The answer we accepted was a quadrotor carrying a parcel on a cable through changing crosswinds, flown side by side with and without a learned wind forecast. MuJoCo keeps the rigid bodies, the cable, the rotors and the contacts. The neural operator only forecasts the air. The brief was a starting point. This project does not reproduce any published benchmark.

Before anything moved we built the static scene and checked it from several cameras: a generic quadrotor made of geometric primitives, a foam parcel, pickup and delivery platforms 5.4 m apart, two gates, and onboard and external cameras.

The course: pickup platform with quadrotor and parcel, two gates, delivery platform
The course before any motion: the quadrotor and parcel on the pickup platform, two gates, and the delivery platform 5.4 m away.
Quadrotor hovering over the parcel on its cable
The aircraft, the cable and the 26 × 26 × 20 cm parcel. Rotor disks are visual only. Thrust comes from four site actuators.
View through the gates toward the delivery platform
Looking down the route through both gates.
Downward camera looking at the parcel on the platform
The downward payload camera. Cameras are for the viewer. The controller never sees their images.
PartValue and reason
Aircraft2.0 kg; diagonal inertia 0.035, 0.035, 0.055 kg m²; a compact inspection drone, not a vendor model
Parcel0.20 kg, 26 × 26 × 20 cm; a light foam box with a large area facing the wind
RotorsFour sites at ±0.24 m; 0–10 N each, so 40 N against a 21.58 N loaded weight; 0.03 s first-order lag
CableUnilateral spatial tendon, 0.65 m between attachments; it can go slack and never pushes
ContactFriction 0.8; compliant MuJoCo contact; the parcel is never welded to anything
IntegrationStandard CPU MuJoCo, 2 ms step (500 Hz), implicitfast, gravity 9.81 m/s²
RouteLift, a 5.4 m crossing at 1.90 m cruise height through two gates, then lower and release: a 43 s flight
Release has to be earned. The controller releases the cable only after the parcel touches the delivery platform and is moving slower than 0.15 m/s. The tendon then goes slack. No flight code writes positions or velocities to make a delivery happen. The videos replay saved poses from the physical runs.

2 · The wind is its own simulation

September 9 · an independent Navier–Stokes solver, applied as drag

Every flight flies through the output of an independent numerical solver, including the flights that use neural forecasts. No flight ever flies through a network's prediction. The solver handles the forced, incompressible 2D vorticity equation on a periodic 12 × 12 m domain, with viscosity 0.04 m²/s, Fourier derivatives, 2/3 dealiasing and third-order SSP Runge–Kutta. Its internal step is at most 0.01 s at 64², 0.005 s at 128² and 0.0025 s at 256². Tests check exact Taylor–Green vortex decay and the divergence without involving any learned model.

Each seed draws its own mix of 12 low Fourier modes, with an initial vorticity RMS of 2 s⁻¹, random phases and steady forcing. These modes sit on top of a mean wind of 1.4–2.2 m/s across the route and −0.4 to 0.4 m/s along it. The evolving vortices add gusts that change in space and time.

Drag at five points

  • F = c ‖u − v‖ (u − v) at each rotor and at the parcel's centre
  • Aircraft c = 0.14 kg/m, split over four rotors; parcel c = 0.045 kg/m
  • v is the actual velocity of each point, so swinging and tilting change the drag
  • mj_applyFT adds each force and its moment, rebuilt from zero every 2 ms step

One-way coupling

  • Air moves horizontally; the aircraft and gates do not change it
  • Drag dissipates energy relative to the air, but moving air can still do work on the aircraft
  • The same wind drives every method within a seed

Not modelled

  • Rotor wash, obstacle wakes and pressure integration
  • Vertical turbulence; cable mass and cable drag
  • Calibration against a real aircraft or real wind
Keep the subsystems apart. MuJoCo supplies rigid-body dynamics and contact. The solver supplies the air. The networks only predict the air. Because of that separation, a forecast can be judged against the solver's actual wind, and a flight can be judged against MuJoCo's physics.

3 · One controller, four forecasts

September 9 · the controller is written by hand; only the forecast changes

The controller is programmed, not learned. At 10 Hz it plans 2 s ahead with a linearized model of the aircraft and its swinging parcel. The model tracks position, speed, swing angle, swing rate and the lag of horizontal rotor force, and a quadratic cost over the horizon penalizes parcel error, aircraft error, swing, effort and slew. The controller solves the unconstrained problem and then clips horizontal force to ±6 N per axis. A geometric attitude loop at 100 Hz turns that force into an orientation and four rotor commands. This is an approximate predictive controller, not a constrained nonlinear MPC, and the route is fixed rather than planned.

Wind enters the plan as predicted drag along the reference path. The four methods differ only in where those 2 s of future wind come from:

MethodWind over the 2 s horizonDeployable
Frozen fieldThe last observed field, held fixedYes
FNOThe observation advanced by the data-only neural operatorYes
PINOThe observation advanced by the operator that was also trained on a PDE residualYes
True futureThe solver's actual future windNo; a diagnostic, not an optimum
Side-by-side frozen and PINO flights entering the crosswinds, with forecast panels
Standard flight, seed 300, at 6.0 s: frozen field (left) and PINO (right). Bottom row: the one-second PINO forecast beside the solver's future field (shown for evaluation only), the downward payload camera, and tracking error.
Matched by construction. Within each seed, the reference path, controller matrices, gains, actuator limits, initial conditions and solver wind are identical. Every controller sees the same full current wind field every 0.25 s, plus exact aircraft and parcel state. That is privileged simulator state, not perception. Forecasts are cached before each flight, and each entry depends only on the observation at its timestamp, the known forcing and the model weights. A test corrupts the future reference arrays and confirms the forecasts do not change. Checkpoint hashes are stored in every flight report.

4 · Training the forecasters

September 9 · 48 training winds, one RTX 4090, 64 epochs per model

Both predictors are the official NeuralOperator FNO: four Fourier layers, width 32, 16 × 16 retained modes, and 603,041 parameter elements counting complex spectral weights. Each takes the current vorticity, the steady forcing and the mean wind, and predicts a mean-preserving update for the next 0.25 s. Chained eight times, that covers the controller's 2 s horizon. The same weights also run on other grid sizes.

PINO here is an FNO trained with one extra loss term. The two runs share the initialization, data, minibatch order and learning-rate schedule for 64 epochs. From epoch 33, every fourth PINO minibatch adds a Navier–Stokes residual at weight 0.03, computed on 128² inputs that are Fourier-interpolated from the coarse data. There are no 128² labels. The residual uses a trapezoidal time step, so it acts as a soft regularizer and does not guarantee physical consistency.

PartitionSeedsUse
Training0–4748 trajectories × 64 steps: 3,072 supervised pairs on 64²
Checkpoint selection100–107One-step error on independent trajectories
Flow evaluation200–20364² and 128² from 2, 12 and 28 s; 256² from 12 s; leads 0.25, 1 and 2 s
Flight development200Controller tuning before any held-out flight
Flight evaluation300–305All four methods on all six winds (version 2: 400–405)
73.4s
FNO training, 64 epochs
RTX 4090; excludes setup, data generation and validation
82.3s
PINO training, 64 epochs
Same data and schedule, plus the residual
0.096
FNO one-step error relative to the frozen field (PINO 0.097)
Vorticity, 0.25 s ahead, selection seeds; both chose epoch 64
0.064m/s
FNO velocity RMSE at a 2 s lead (PINO 0.065)
Held-out seeds 200–203, three start times, 64²
Grid and start timesFNO, 2 s velocity RMSEPINO
64², starts at 2, 12 and 28 s0.0643 m/s0.0650 m/s
128², the same starts0.0643 m/s0.0650 m/s
64² or 128², 12 s start only0.0413 m/s0.0434 m/s
256², 12 s start, unseen by both losses0.0413 m/s0.0434 m/s

At a 2 s lead, the forecast velocity error is 2–5% of the frozen field's error, depending on the start time. The 256² row looks better only because it uses the 12 s start alone, and at that same start the coarser grids give the identical error. A direct 128² PINO rollout has a relative vorticity error of 0.057133. Fourier interpolation of the 64² rollout gives the same 0.057133.

Training ran on an RTX 4090. Evaluation and every flight ran on an Apple M3 Max, with forecasts on Metal and MuJoCo on the CPU. Optimization can follow different paths on different devices, so the checkpoints are committed and evaluation does not depend on retraining.

A finer grid is not finer accuracy. The same weights run at 128² and 256² and match the coarse result. That shows the operator transfers across grids. It does not show that the operator recovers detail it never learned.

5 · CUDA and Metal disagreed

September 9 · a portability bug, a pinned upstream fix, then retraining

Same weights, different answers. With the released NeuralOperator 2.0.0, the same model produced different 128² inverse-FFT outputs on CUDA than on Metal and the CPU. Training had finished without error and every shape matched, so neither check revealed the problem. The repository records the discrepancy and the fix, but not the size of the original disagreement.

An FNO layer multiplies Fourier modes by learned complex weights and returns to a real field through a real inverse FFT. That transform assumes the spectrum is Hermitian-symmetric, as the spectrum of any real field is. The transform leaves any part that breaks the symmetry undefined, and different FFT backends handled that part differently. NeuralOperator fixed this upstream: commit 00b7d86 explicitly enforces Hermitian symmetry before the real inverse FFT. We pinned the dependency to that commit in uv.lock, turned off TF32 so CUDA arithmetic stays comparable with the CPU and Metal, and retrained both models. The committed weights are the retrained ones.

ModelGridCPU vs CUDACPU vs Metal
FNO64²4.8 × 10⁻⁶5.1 × 10⁻⁶
FNO128²6.8 × 10⁻⁶6.0 × 10⁻⁶
FNO256²7.0 × 10⁻⁶9.1 × 10⁻⁶
PINO64²5.1 × 10⁻⁶5.4 × 10⁻⁶
PINO128²6.2 × 10⁻⁶6.9 × 10⁻⁶
PINO256²8.2 × 10⁻⁶9.5 × 10⁻⁶

Values are the largest absolute vorticity difference, in s⁻¹, after eight chained 0.25 s steps (2 s), with the fixed models. For scale, the initial vorticity RMS is 2 s⁻¹. When the repository was later set up on the GPU machine, both committed models passed the CPU/CUDA check again, with a largest difference of 8.23 × 10⁻⁶ s⁻¹.

A clean training run is not a cross-device check. The comparison is now a script (scripts/check_wind_backend.py). A regression test checks the Hermitian setting and that gradients stay finite. The repository's guidelines keep the pinned revision until a replacement passes the same checks.

6 · Standard flights: 24 of 24

September 9 · six held-out winds, four forecasts, one controller

Standard profile, seed 300, the first evaluation seed, chosen before the final results: frozen field (left) and PINO (right) in identical wind. The film has a 4 s opening hold, the full 43 s flight at real time, a 5 s hold on the final frame, then the results card, 62 s in all. It is rendered from saved states of the physical runs. Tracers follow the solver's wind, arrows show applied forces, and coloured paths show controller predictions.
24/ 24
deliveries across all methods and winds
Every release after platform contact; no gate collisions, no MuJoCo warnings
10.89cm
frozen field: mean crossing tracking RMSE
Horizontal parcel error, t = 6–32 s; peak swing 28.81°
4.28cm
FNO forecast
60.7% lower; peak swing 25.36°
4.42cm
PINO forecast
59.4% lower; true-future diagnostic 4.30 cm
Wind seedFrozen fieldFNOPINOTrue future
3009.90 cm3.88 cm3.88 cm3.65 cm
30113.33 cm6.02 cm6.39 cm6.31 cm
3028.26 cm3.03 cm2.96 cm2.99 cm
3039.36 cm2.70 cm2.76 cm2.43 cm
30413.43 cm3.71 cm3.73 cm3.72 cm
30511.07 cm6.34 cm6.79 cm6.72 cm
Mean10.89 cm4.28 cm4.42 cm4.30 cm

The forecast is what helps. Both learned forecasts cut the error by about 60% compared with holding the wind fixed, and they land within 0.14 cm of each other and of the true-future diagnostic. FNO is lower than PINO on four of the six winds and PINO on one, and they tie on seed 300. On three winds (301, 304 and 305), FNO even beats the true future. The controller's model is approximate, so a perfect forecast does not guarantee the lowest error.

PINO parcel delivered while the frozen-field parcel is still lowering
36.4 s: the PINO side has released after contact (at 36.32 s). The frozen-field release on this seed came at 36.36 s. All 24 releases fell between 36.30 and 36.47 s.
Results card: mean tracking RMSE for the four forecasts
The film's results card. The bars average all six held-out winds, not only the seed shown.
PINO did not beat FNO. The physics residual is a soft constraint added to the same network, and on these six winds it neither helped nor hurt measurably. The finding is the value of a forecast, not of the physics loss.
Physical checks. Every run stayed under the 10 N rotor limit; the highest was 9.734 N, in a frozen-field run. The largest cable constraint extension was 0.400 mm. There were no gate collisions and no MuJoCo numerical warnings. The whole goal, from brief to committed video, took 54 min 21 s of elapsed time. That covers GPU setup, the FFT fix, training, 27 passing tests and rendering. The two final training runs took 2 min 36 s of it.

7 · Version 2: stronger wind, faster crossing

September 9 · a follow-up, kept separate from the original

After the first version we were asked for more aggressive dynamics, and then to keep the original intact. Version 2 is a separate aggressive preset with fresh seeds 400–405 and its own output folders. The original video, checkpoints and standard preset are unchanged. Rerunning the standard trajectory reproduced it exactly.

To make the wind 1.5× stronger while keeping the same physics, version 2 uses the Navier–Stokes similarity transform unew(x, t) = s·ubase(x, s·t) with s = 1.5. Viscosity becomes 0.06 m²/s, vorticity scales by s and forcing by s². Wind speed and the rate of gust evolution both rise 1.5×, the Reynolds number stays the same, and the aircraft's mass, gravity and actuator time constants are unchanged. The trained models needed no retraining. They forecast 3 native seconds for each 2 s physical horizon, and the output is rescaled. The observation clock scales too, to 6 Hz.

ParameterOriginalVersion 2
Wind speed and evolution rate1×1.5×
5.4 m crossing24 s12 s
Peak commanded cruise speed0.3375 m/s0.675 m/s
Crossing acceleration scale1×4×
Tracking window6–32 s6–20 s
Complete flight43 s31 s
Wind seeds300–305400–405
Masses, cable, contact, controller gains, 0–10 N rotorsUnchanged
Test the identity, not the integrator. The first scaling test compared runs that used unmatched time steps. It failed on ordinary Runge–Kutta discretization error, not on the physics. Once the dimensionless steps were matched, the test isolated the scaling identity, and the scaled solver now agrees with the transformed original to below 10⁻⁵ in vorticity.
Version 2 at 18 s: both parcels swinging past the second gate, with motion trails
18.0 s, seed 400: trails show each parcel's last three seconds of actual motion. The white cross marks the horizontal target.
Version 2 at 24.4 s: PINO parcel delivered, frozen-field parcel still swinging
24.4 s: PINO released at 24.32 s, while the frozen-field parcel is still swinging onto the platform. It released at 24.62 s.

The film at the top of the page is seed 400, the first new evaluation seed, chosen before the final results. Both flights deliver. The frozen-field parcel strays up to 47.7 cm from its target, against 11.4 cm for PINO, and the crossing RMSE is 20.17 cm against 5.99 cm. Peak sling angles are close, 46.2° against 43.4°, and peak aircraft tilt is 21.2° against 16.9°. The forecast keeps the parcel on its path. It does not stop the wind from swinging it.

The first version 2 cut ran 50 s, with a 4 s frozen opening and a 5 s frozen end. Both holds were removed the same evening, leaving the 31 s flight at real time and the 10 s results card: 41 s and 1,025 frames, all of which decoded without error.

8 · Seed 404, kept on record

September 9 · one wind beyond this controller and these rotors

Seed 404 failed for all four methods. All four runs hit the 10 N rotor limit. These were physical and control failures, and MuJoCo reported no numerical warnings. The reports keep each termination reason and partial trajectory, and the comparison command exits with an error status when any mission fails.
MethodWhat happenedRun ended
Frozen fieldStruck a gate and dropped below the flight envelope15.67 s
FNODrifted past the lateral envelope during the lift, before the crossing began5.58 s
PINOThe same lateral drift during the lift5.54 s
True futureFlew all 31 s and set the parcel on the platform, but contacted a gate, so the mission failed31.00 s

The seed 404 runs met winds of 8.09–8.33 m/s at the aircraft and parcel, among the strongest of the six seeds. Wind speed alone does not explain the failure, though. On seed 403 the wind reached 8.13 m/s and every method also hit the rotor limit, yet all four delivered.

The failure also changes how the error is measured. The failed FNO and PINO runs ended before the crossing began, so their reports hold no tracking samples. Counting them would have entered zero error for the two worst flights. The RMSE comparison therefore uses only the five windows that every method completed (seeds 400, 401, 402, 403 and 405).

5/ 6
missions completed, for every method
The same success rate; seed 404 failed for all four
19.59cm
frozen field: mean tracking RMSE
Five shared complete windows, t = 6–20 s
4.34cm
FNO forecast
77.8% lower on those five windows
4.51cm
PINO forecast
77.0% lower; true-future diagnostic 4.47 cm
Version 2 results card with failures retained
Version 2's results card names the failed wind and leaves incomplete windows out of the bars.
What the 77% means. The 77% applies only to the five completed windows. It is not a better success rate, since every method completed 5 of 6. FNO was again lower than PINO on four of the five winds, and PINO was lower on seed 405, by 0.01 cm. Versions 1 and 2 are different experiments, so their percentages are not a controlled comparison. The follow-up took 24 min 29 s and ended with 32 passing tests. Removing the holds took another 2 min 43 s.

9 · Next steps

What would make the claim stronger

  1. Forecast live, inside the loopForecasts are now cached before each flight from causal inputs. Running the operator inside the 10 Hz loop, and timing it there, would turn cached playback into a deployable measurement.
  2. Partial, noisy observationsEvery controller sees the full current field every 0.25 s. A real aircraft would sense the wind at a few points, late and with noise, and forecasting from that is the harder, more useful problem.
  3. Find the edge, not one caseSeed 404 shows where this controller and 10 N rotors stop coping. A predictive controller that respects rotor limits, plus more seeds at graded wind strengths, would map that boundary.
  4. Test the physics loss where it could matterWith plenty of data from one flow family, PINO matched FNO. Fewer labels, longer horizons or winds outside the training family are where a PDE residual might make a measurable difference, measured with several training seeds per model.
  5. Richer airTwo-way coupling, rotor wash, obstacle wakes, vertical gusts and cable aerodynamics, with flow data beyond this synthetic 2D family, and then validation against a real system.

Nothing here is calibrated to a real aircraft or real wind. The result is the value of forecasting the wind within this simulated test family, flown by a programmed controller.