Five learners. One shared world.
Ground vehicles, boats, quadcopters, fixed wings and submarines learn local navigation together. A* supplies the global routes. Each vehicle acts on its own scheduled sensors and remembers its own trajectory; movement families share policy weights.
Loading recorded evaluation…
| Family | Deployed PPO | Three-seed range | Route tracker | Random actions | Contacts: PPO / tracker |
|---|
Full mission arrivals, including unavailable requests in the denominator. Contact counts are blocked decisions, not unique impacts. All methods use the same global routes and low-level tracker; PPO and random actions choose local overrides. The reference has no learned avoidance. Wing arrivals are fly-throughs.
What learned together
Every rollout advances all five learning policies in the same physical worlds. Each family has its own PPO optimizer and value function. Clear-route training progresses to opposing traffic and full missions. These policies learn individual navigation objectives; team tactics and combat are not part of this task.
Command the fleet
Click a vehicle in Map Lab, select a family, and right-click a destination. Air and submarine units support explicit height/depth targets. Sensor equipment changes what the policy observes; range overlays only change what you see. Fixed-wing routes include heading and physical turn/climb limits.