← MuJoCo Sandbox

Engineering journal · October 5–6, 2026 · MuJoCo 3.12 · PufferLib 5.0

Building a surgical thread-insertion robot in simulation

A gantry robot with a needle and a thread tube, a 40 µm thread and a tissue phantom with vessels. This is the record of getting the physics right, teaching the robot to align its needle within 10 µm, and then teaching it to insert threads, site after site, while the workcell, the sensors and the tissue all disturb it.

Replay
Controller
Controller
Disturbance
Evaluation run
Camera
Overlays
Playback
Speed
Speed
t 0.00 s phase – site – thread depth – mm from target – µm needle force – mN lateral – µm vertical – µm target – level – latency – ms loading…

Replays of recorded MuJoCo states; the page does not simulate. Insertion: the learned policy (insert_v2, final, deterministic) or the scripted yardstick placing threads at sites 0, 5 and 1, on the first evaluation seeds of the baseline at each disturbance level; frames every 2 ms during the needle's work and every 8 ms between. The thread is lime, and also drawn as a line where it is thinner than a pixel; the Micro camera shows the tissue translucent. Sticking, release and new threads in the tube are stand-in constraints (chapter 15). Alignment: the first 30 of the 200 predetermined evaluation seeds for each policy and level, recorded at 50 Hz. Overlays show recorded disturbance values, drawn larger so they read at any zoom: force arrows at 1 mN per 7 screen pixels, and micrometre offsets (sensing errors, tissue motion) magnified by the factor shown in the legend. Drag to orbit, right-drag or Shift-drag to pan, scroll or pinch to zoom.

60 / 60
threads placed by the learned policy, 21 full-simulation runs
8.0 µm
its median placement at nominal disturbances (approved cycle: 23 µm)
84%
within 10 µm on 200 held-out training-environment episodes
0
contacts with the tissue or a placed thread, full simulation
1 h 39 min
final insertion training run on one RTX 4090
Where this stands. A learned policy now drives the insertion: it steers the needle over each site, compensates for where the thread's end actually waits, and decides when to strike. The programmed stroke then sticks the needle to the thread, inserts it 2 mm, releases it and moves on, three sites per run, under workcell, sensing and tissue disturbances. It was trained in a fast C environment and checked in the full thread simulation, where it places every thread and lands a median 8 µm from the target at nominal disturbances. Sticking, release and thread feeding are stand-in constraints, not a modelled cartridge. Every number on this page comes from recorded runs in the repository.

How this was built

A harness, a brief, and two coding agents working in a loop

Two things made this possible: a harness and a brief. The harness was in place before this project began. It is a repository set up for simulation and learning work: MuJoCo for physics, PufferLib 5.0 for reinforcement learning, a locked Python environment, Git LFS for assets, a commit-based workflow between a laptop and a remote RTX 4090 workstation, earlier MuJoCo projects to build on, and written engineering rules for how simulations are modelled, tested, recorded and documented here. The brief states what to build and the gates the work must meet.

Everything specific to this robot was written by two coding agents working together inside that harness from the brief, one running GPT-6 Astra and one running Claude Opus 5.5. That includes the workcell and needle model, the unit rescaling, the C++ thread model and contact kernel, the C task core shared by training, tests, evaluation and replay, the disturbance layer, the PufferLib environment, the evaluation tools, and this page with its WebAssembly replay. The project's first commit was on October 5.

  1. Brief and gateswhat to build, how it is judged
  2. Planthe next measurable step
  3. Buildcode, models, configuration
  4. Runsimulation or training
  5. Measureagainst the gates, on held-out seeds
  6. Recordresults and failures, committed
  7. Change one thingthen run again

We call this improvement loop AutoResearch. The agents plan a step, implement it, run it, measure the result against the brief's gates, record what failed, and change one variable before the next attempt. The training runs in this journal show the pattern: a diverged run exposed a reward bug, two runs that blew up led to lower learning rates, and a stable but slow run led to a longer training budget. A person sets the brief, reviews the results and corrects course. Their corrections were decisive: train on the GPU workstation rather than fall back to a CPU, make the gates serve the mission rather than the reverse, and pursue robustness under disturbances. Their checks of this page in a browser found four display bugs.

"Self-improving" means the work improves, not the agents. The models behind them do not change. The code, the physics, the policies and the process get better through measured iterations, and every mistake stays on the record.

1 · A static workcell first

October 5 · model, inspect, then move

We built the robot as a static scene before anything moved: an XYZ gantry, a downward insertion slide with a needle, a mounted microscope, and a domed tissue phantom with six targets and vessel markings. Clearance and limits were checked at 120 sampled poses before any dynamics. The first tool also had a thread retainer slide and a cassette of thread samples; the alignment experiments of chapters 7–14 use it. The insertion tool replaces them with a thread tube beside the needle (chapter 15).

The gantry's tool over the phantom: needle and thread tube
The insertion tool over a 658 × 448 mm deck.
A lime thread in a clear tube, its end under the needle point
The needle and the thread tube: the thread's end waits on the needle's path.
AxisTravel (mm)Moving mass (kg)Force limit (N)
X−175…1752.4±100
Y−75…751.3±80
Z−8…500.65±40
Insertion (down)0…180.025±2
Retainer (down), first tool only0…60.004±0.3

The workcell, part by part

Each view highlights one part against the faded machine, in the insertion workcell at a recorded moment: the thread stuck to the needle as it enters the tissue. Select a part to enlarge it; arrow keys step through.

2 · A 40 µm thread at SI scale will not compile

October 5 · unit rescaling

A 44 mm thread of 40 µm diameter, split into 32 segments, has rotational inertias near 10−19 kg m². MuJoCo rejects them at compilation. Rather than inflate the mass or radius, we rescaled the engine's units: millimetres and grams, with seconds unchanged. Every dimensional quantity converts at the model boundary, and results are reported back in SI.

@dataclass(frozen=True)
class Units:
    name: str
    length: float = 1.0  # engine length units per metre
    mass: float = 1.0  # engine mass units per kilogram

    @property
    def force(self):
        return self.mass * self.length

    @property
    def torque(self):
        return self.mass * self.length**2

MM_G = Units("mm_g_s", 1000, 1000)

src/sixlegs/neural_insertion/units.py: forces scale by M·L, inertias and torques by M·L², moduli by M/L.

The physical mass and dimensions are unchanged, and bending then matched the analytical cantilever within 2.4%. The same conversion applies to every scene we build, so the robot and the thread share one unit system.

3 · Porting a published rod model, and correcting it

October 5 · discrete elastic rods inside MuJoCo

We ported the published discrete-elastic-rod (DER) plugin for MuJoCo to version 3.12 as a separate library, leaving the engine untouched. An energy-gradient audit found that the published force mapping produced zero restoring torque for a purely twisted rod despite positive twist energy. We kept the original variant and added a corrected one that projects the same Cartesian forces through MuJoCo's own Jacobians, plus the material-frame end moments. Its forces match the energy gradient to 2.85 × 10−9.

// Project the same Cartesian elastic forces through MuJoCo's Jacobians,
// and include material-frame end moments.
mju_copy(d->qfrc_passive, before.data(), m->nv);
const mjtNum zero[3] = {0, 0, 0};
for (int i=0; i<r->nv+2; ++i) {
  int body = r->i0 + std::min(i, r->nv);
  mj_applyFT(m, d, r->nodes[i].force.data(), zero, r->nodes[i].pos.data(),
             body, d->qfrc_passive);
}
double moment = r->beta_bar*(r->edges[r->nv].theta-r->edges[0].theta)/r->bigL_bar;
Eigen::Vector3d first = moment*r->edges[0].e.normalized();
Eigen::Vector3d last = -moment*r->edges[r->nv].e.normalized();
mj_applyFT(m, d, zero, first.data(), r->nodes[0].pos.data(), r->i0, d->qfrc_passive);
mj_applyFT(m, d, zero, last.data(), r->nodes[r->nv+1].pos.data(),
           r->i0+r->nv, d->qfrc_passive);

src/sixlegs/neural_insertion/der_bridge.cc, the corrected direct projection.

Bending stayed within 2.4% of the analytical cantilever, an undamped bent rod relaxed with energy drift falling from 0.075% to 0.038% as the step halved, and forces were identical after copy, restore and reset. For sim-to-real this matters more than any single benchmark: a rod whose twist cannot push back would mislead every later thread-handling result.

4 · The contact model

October 5 · a continuous contact force and a fourth-order integrator

MuJoCo's soft contact computes a reference acceleration a_ref = −B·v − K·d(r)·pos with an impedance d(r). With d held constant, as in our first law, the damping force switches on as a step at zero penetration, so a grazing contact receives a finite impulse. A penetration ramp makes the force continuous in the state. And because the default integrator is first order through a contact lasting a few tens of microseconds, we step the thread with MuJoCo's RK4, which evaluates collision and the constraint solve at four stages. Both are standard MuJoCo settings; the engine is unmodified.

Before relying on it we checked the engine against an independent calculation that never calls MuJoCo: same-state contact forces match a closed-form solution to 10−11, and agree across unit systems to 10−9.

def law_for(tau):
    return ContactLaw(f"ramp_{tau*1e6:g}us", d0=mujoco.mjMINIMP, power=1., time_constant_s=tau)

law = law_for(2e-5)  # thread contact in the task regime: 20 µs, stepped at 5 µs with RK4
# solimp = (0.0001, 0.9999, 1 µm width, 0.5, 1): impedance rises linearly over 1 µm

src/sixlegs/neural_insertion/task_fixtures.py with contact_laws.ContactLaw.

if (rk) {                 // same sequence as mj_step for RK4
  mj_checkPos(m, d);
  mj_checkVel(m, d);
  mj_forward(m, d);
  mj_checkAcc(m, d);
} else {
  mj_step1(m, d);
}
...
mj_RungeKutta(m, d, 4);

src/sixlegs/neural_insertion/contact_kernel.cc: our batch stepping kernel with per-step contact telemetry; bit-identical to mj_step.

5 · Thread contact in the task regime

October 6 · gates derived from what the robot will actually do

Each phase of the mission touches the thread in its own way, at millimetres per second. We built one fixture per phase, with a rigid 150 µm probe driven only by force-limited actuators, and required the thread's centreline to agree within 1 µm when the timestep is halved and within 0.01 µm across unit systems, with penetration checked at every step over every contact.

Task phaseFixtureTimestep, 5 → 2.5 µsUnits
Transport and alignSettle on the support0.13 µm≤ 1.5e−10 µm
Insert to depthNeedle drags across thread (it rolls)≤ 5.1e−7 µm≤ 6e−7 µm
Pick and grasp100 µN press, force balance ≤ 6.8e−5≤ 2.4e−7 µm≤ 9e−11 µm
ReleaseHook slides out, 150 µm fall0.19 µm≤ 7.3e−7 µm

Every fixture passes at a 5 µs RK4 step with a 20 µs contact time constant, for both material presets. Release failed at 10 µs (1.58 µm), because its short fall lands faster than settling does; we refined the step rather than relax the gate, and kept the failure on record.

Sim-to-real gaps we know about. Thread materials are presets (100 MPa illustrative, 2.5 GPa polyimide), not measurements. Contact compliance under load, about 1 µm per contact at 100 µN, is a property of MuJoCo's soft contact on a 2.4 µg segment, not a calibrated stiffness. The phantom's collision mesh has three rings, so its facets sit up to about 190 µm below the analytic dome. At this stage the phantom could not be punctured (chapter 15 adds a puncture model), and observations were exact simulator state rather than sensor readings.
Task regime plots
Largest passing timestep, penetration against contact time constant, fixture agreement over time and the press force balance.

6 · The robot's first motion

October 6 · programmed servo, gates fixed before the run

A programmed servo adds inertial, gravity, viscous and Coulomb feedforward and a bounded integral term to each force-limited position actuator, by offsetting its command. Two design facts surfaced on the first run: Z reaches only 8 mm down, so the needle slide reaches the surface; and both fine slides hang against gravity at their retracted stops, so they are carried 0.5 mm out.

ctrl = q_ref + (kv*v_ref + M*a_ref + qfrc_bias + damping*v_ref + coulomb + I) / kp
# applied force = kp*(q_ref - q) + kv*(v_ref - qdot) + feedforward + I, clipped by the actuator force range

src/sixlegs/neural_insertion/motion.py

Six-target tour, real time. Hover error ≤ 0.50 µm lateral and ≤ 0.09 µm vertical; no contacts; 16.8% peak force.

7 · The learning task, exactly

October 6 · trained task: needle alignment (robot only)

One C core (native/surgical_core.h) holds physics stepping, servo, observations and rewards. It is compiled into the PufferLib 5.0 adapter for training and into a local library for tests and evaluation, so they cannot drift apart. Physics runs at 1 ms; the policy acts at 50 Hz (20 physics steps per action).

Episode

Observation (16 values, exact simulator state)

IndexQuantityScalingClip
0–2Tip minus goal, x y z÷ 5 mm±5
3–5Tip velocity, x y z÷ 40 mm/s±2
6–8Previous applied action—±1
9–11Tip minus goal, fine scale÷ 100 µm±1
12–13Offset to the nearest vessel centreline, x y÷ 5 mm±1
14Tip height above the dome÷ 10 mm−1…2
15Remaining time fraction1 − step/2500…1
for (int i = 0; i < 3; i++) {
    double e = tip[i] - c->goal[i];
    obs[i] = (float)sa_clip(e / 0.005, -5, 5);
    obs[3 + i] = (float)sa_clip(vel[i] / SA_VMAX_XY, -2, 2);
    obs[6 + i] = c->prev_action[i];
    obs[9 + i] = (float)sa_clip(e / 1e-4, -1, 1);
}
double vx, vy;
sa_vessel(c, tip[0], tip[1], &vx, &vy);
obs[12] = (float)sa_clip(vx / 0.005, -1, 1);
obs[13] = (float)sa_clip(vy / 0.005, -1, 1);
obs[14] = (float)sa_clip((tip[2] - sa_surface(c, tip[0], tip[1])) / 0.01, -1, 2);
obs[15] = (float)(1.0 - (double)c->tick / SA_HORIZON);

Action (3 continuous values)

Each output is clipped to [−1, 1] and cubed, then scaled to needle-tip velocity: 40 mm/s in X and Y, 20 mm/s vertically. The cube gives 40 µm/s resolution at 0.1. The command drives the X and Y stages and the insertion slide; a 10 ms filter smooths the reference, which the programmed servo tracks.

Reward per policy step

TermValue
Progress (potential shaping)Φ(d′) − Φ(d), with Φ(d) = −log₁₀((d + 2 µm) / 1 mm) and d the tip-to-goal distance
Action smoothness−0.002 · Σ (aᵢ − aᵢ,prev)² on the clipped actions
Vessel proximity−0.05 when the tip is under 3 mm above the dome and within the vessel radius + 0.3 mm of a centreline
Success+2, episode ends
Collision−2, episode ends

The logarithmic potential makes going from 10 mm to 1 mm worth the same as going from 10 µm to 1 µm, so the policy gets useful signal at every scale. There is no curriculum: every episode is the full task.

double potential = sa_potential(distance);
double r = potential - c->potential;
c->potential = potential;
double smooth = 0;
for (int i = 0; i < SA_ACT; i++) {
    double applied = sa_clip(action[i], -1, 1);
    double da = applied - c->prev_action[i];
    smooth += da * da;
    c->prev_action[i] = (float)applied;
}
r -= 0.002 * smooth;

src/sixlegs/neural_insertion/native/surgical_core.h, after the fix described in chapter 9.

Baselines on the evaluation seeds

PolicySuccessCollisionMean time
Scripted reference (programmed)100%0%2.06 s
Zero action0%0%—
Uniform random0%5%—
Caught before training: the scripted reference first failed twice. Descending early cut through the dome, and a 4 s deadline at 20 mm/s could not cover the longest traverse. Both were task-design errors, fixed before any policy was trained, so the learned policy is judged on a task known to be solvable.

8 · The network we train

PufferLib 5.0 default recurrent policy, 128 hidden units, 2 MinGRU layers

Read from PufferLib's own code and the checkpoint's layout: 100,867 parameters, no biases. The same weights run on the GPU during training and on the CPU (deterministic mean action) for evaluation and this replay.

Observation 16 values error ×2 scales, velocity, action, vessel, height, time Encoder Linear 16 → 128 2,048 weights MinGRU × 2 layers (memory per world) each layer: Linear 128 → 384 = [ h | z | w ] 49,152 weights per layer g(h) = h + ½ if h ≥ 0, otherwise σ(h) st = st−1 + σ(z) ⊙ (g(h) − st−1) y = σ(w) ⊙ st + (1 − σ(w)) ⊙ x (highway) x: layer input · s: state, zeroed at episode start recurrent state s, 2 × 128 Decoder Linear 128 → 4 512 weights Action μ ∈ ℝ³ + learned log σ (3) train: a ~ N(μ, σ) evaluate: a = μ Value V(s) PPO critic, training only a → clip[−1,1] → a³ → tip velocity (40, 40, 20 mm/s) → 10 ms filter → programmed servo → MuJoCo, 20 × 1 ms

Training configuration

SettingValue
AlgorithmPufferLib 5.0 PPO (pinned revision 6ffa5b1), clip 0.2, value coefficient 2.0
OptimizerMuon, learning rate 0.003 cosine-annealed to 0
Discount, GAE λ0.995, 0.95
Entropy coefficient5 × 10−4
Batch2048 worlds × 64-step horizon, minibatch 16,384
HardwareRTX 4090 for the network; MuJoCo physics on 28 CPU threads
Budget100 M steps (99.9 M), 29 minutes, about 59,000 steps/s
void puf_step(Env* env) {
    SAEpisode e;
    float reward;
    int done = sa_step(&env->core, env->agents[0].actions, env->agents[0].observations, &reward, &e);
    env->agents[0].rewards[0] = reward;
    env->agents[0].terminals[0] = (float)done;
    if (done) {
        env->log.perf += e.success;
        ...
        sa_reset(&env->core, env->agents[0].observations);
    }
}

experiments/neural_insertion/puffer/surgical_align/surgical_align.h: the PufferLib adapter is a thin wrapper over the shared core.

9 · Training on the RTX 4090: two failures, then success

October 6 · three runs, one change at a time

Final lateral error at episode end (training rollouts)

Success rate (training rollouts, noisy actions)

Policy update size (KL divergence)

Collision rate

Run 1 diverged, and it was our bug. Episode returns reached −46,557 from terms that should never exceed a few units. The smoothness penalty used the policy's raw, unbounded Gaussian samples instead of the clipped actions the robot receives; once action means drifted past ±1 the penalty grew without limit.
-        double da = action[i] - c->prev_action[i];
+        double applied = sa_clip(action[i], -1, 1);
+        double da = applied - c->prev_action[i];
Run 2 learned, then collapsed. For 12.7 M steps it improved steadily; then, as action noise shrank, each update at the default learning rate moved the policy too far. KL rose from 0.16 to about 300.
+learning_rate = 0.003
 gamma = 0.995
Run 3 changed only the learning rate. It stopped colliding within 10 M steps, closed in through hundreds and then tens of microns, first succeeded at about 52 M steps and reached 100% in training by 69 M.

10 · Results on 200 held-out episodes

Deterministic actions, evaluated on CPU with PufferLib's own network code

PolicySuccessCollisionsMean timeWorst final errorLow steps over vessels
Learned, final (99.9 M)100%0%1.40 s8.6 µm lat., 4.4 µm vert.0.25
Learned, 65.5 M100%0%1.98 s10.0 µm lat., 7.4 µm vert.0.39
Scripted reference100%0%2.06 s3.4 µm lat.1.68

The learned policy is a third faster and passes low over vessels less often, but settles closer to the tolerance than the scripted controller: an episode ends as soon as the 0.3 s hold is met. Compare both in the replay at the top.

Four evaluation episodes chosen before evaluation (the first seeds covering four targets), real time.

11 · A perfect world is too easy

October 6 · the brief is rewritten around disturbances

With exact state, no delay and a still target, both the scripted controller and the learned policy succeed every time. A real stage gets none of that. We rewrote the brief: every mission phase is learned under one shared, documented disturbance layer; the programmed servo stays only as the motor drive; scripted controllers become yardsticks. Results are reported on held-out seeds at stated disturbance levels, including full strength.

The layer lives in the same C core used for training, evaluation and replay. One number, the level from 0 to 1, scales every magnitude, and in training each episode draws its own level uniformly from that range. The magnitudes are illustrative until replaced by measurements of the real stage, sensors and tissue.

DisturbanceModelFull strength (level 1)Enters through
Slide force noiseLow-pass random force, 50 ms correlation timestd 100, 100, 50, 10, 2 mN on X, Y, Z, insertion, retainerApplied joint force
Extra frictionCoulomb, smoothed below 0.1 mm/sup to 100, 100, 50, 5, 1 mN, drawn per episodeApplied joint force
Table vibrationThree sinusoids per axis, 5–60 Hzup to 10 mm/s² per componentInertial force on each carriage
Tip sensingDelay, then Gaussian noise0–10 ms latency; 0.5 µm and 0.05 mm/s noiseObservation only
Target estimateBias, slow drift and per-step noise1 µm bias, 1 µm/√s drift, 3 µm noise per axisObservation only
Tissue motionBreathing 0.2–0.4 Hz, pulse 1–2 Hzup to 30 + 10 µm vertical, 10 + 3 µm lateral amplitudeTrue target position and velocity
// Every 1 ms physics step: rebuild the external force on each slide from scratch.
for (int j = 0; j < SA_NJ; j++) {
    double sigma = SA_FORCE_SIGMA[j] * q->level;        // low-pass force noise
    q->force[j] += -q->force[j] * dt / SA_FORCE_TAU
                 + sigma * sqrt(2 * dt / SA_FORCE_TAU) * sa_normal(&c->drng);
    double v = d->qvel[c->dof[j]];
    double inertial = -c->mass[j] * dot(a_base, SA_AXIS[j]);   // table vibration
    double friction = -q->friction[j] * tanh(v / 1e-4);       // extra Coulomb friction
    d->qfrc_applied[c->dof[j]] = q->force[j] + friction + inertial;
}
Scale used later. Training runs 4–6 used this first scale. With the user, the robot group was then rescaled to a quality workcell (scale v2): level 1 is now a precision stage on an isolation table, a tenth of the robot magnitudes above (force noise 10, 10, 5, 1, 0.2 mN; extra friction up to 20, 20, 10, 1, 0.2 mN; vibration 1 mm/s²), and level 2 is the stress setting, twice that. Sensing and tissue are unchanged. The evaluation in chapter 13, the replays at the top and the insertion runs use scale v2.
Level 0 changes nothing. Disturbance parameters come from their own random stream, so at level 0 the new core reproduces all 200 recorded evaluation episodes of the earlier policy and of the scripted controller exactly. Rewards and success use true state, and success now also requires the tip to match the moving target's velocity within 0.2 mm/s.

The replay at the top draws every component in the scene. Each carriage carries force arrows at 5 cm per 0.1 N: amber for noise, cyan for friction, lilac for vibration. The table shows its acceleration and recent trace. The Micro camera shows, at true scale, the moving target with its 10 µm tolerance cylinder, the noisy target estimate and the late tip measurement. Values sit in panels beside each object. Airflow is not modelled yet; it matters for the thread, which this task does not carry.

12 · Training under disturbances

October 6 · runs 4 and 5, one change between them

Final lateral error at episode end (training rollouts)

Policy update size (KL divergence)

Success rate (training rollouts, noisy actions)

Collision rate

13 · Robustness on 200 held-out episodes

Deterministic actions, the same 200 predetermined seeds at every level

Success against disturbance level

14 · What the training used

One workstation: RTX 4090 and a 16-core Ryzen 9 5950X

This workload is CPU-bound. PufferLib steps 28 copies of the MuJoCo world on CPU threads, about 29 of the machine's 32 hardware threads, while the GPU runs only the 101k-parameter network and its PPO update. It idles near 10% utilisation and about 59 W of its 450 W limit, with under 1 GB of memory for the trainer. Runs 1–3 predate the resource logger, so their figures come from PufferLib's own dashboard, which reports device-wide GPU memory including the desktop. From run 4, a logger samples GPU power, temperature, the trainer's GPU memory and host CPU load once a minute.

15 · Thread insertion: a thread tube and stand-in handling

October 6 · three mechanisms, then a decision about what to prove

Picking up a 40 µm thread is the hard mechanical part of this machine. We tried it three ways: a slotted needle with a latch over a 44 mm free thread (about 600 times slower than real time, and the eyelet rattled loose on landing); a needle with a ledge and a rotary pincher; then a cartridge with a cannula and a spring latch under the thread's loop. Each attempt was watched on video, and each taught something, but none of it was what this project sets out to prove.

After reviewing photos of a current insertion robot, the user set the goal: prove the robot's movements and abstract the thread handling. The thread waits in a tube on an arm beside the needle, angled 15° from vertical, its end on the needle's path. On its way down the needle sticks to that end, carries the thread into the tissue, and lets go at depth. The rest is physical: the robot's dynamics, the thread as an elastic rod, the tube's walls, the puncture and the tissue's grip.

The thread's end waits under the needle point
Waiting: the thread's plain end sits 1 mm below the needle point and 1 mm above the tissue.
A lime thread through the translucent tissue surface
Placed: the thread's end 2 mm deep, drawn with the tissue translucent.
Stand-in (a MuJoCo constraint)Switched onSwitched offStands for
Tube hold on the thread's back endthread ready in the tubewhen the needle has the threada feed brake
Needle bond to the thread's endneedle point within 30 µm of the endat deptha chemical attachment and release
Tissue grip on each segment that entersbelow the surface, near the puncturepulled out of the tissuetissue holding the thread; slips above 16.7 µN per 0.46 mm segment
New thread in the tubebetween sites (a spare, straight and at rest)feeding the next thread
What this proves, and what it does not. These runs test positioning, alignment under disturbances, insertion to depth, release, and moving between sites. They do not test how a real cartridge picks, holds or releases a thread, or whether the thread survives it.

16 · A fast thread

October 6 · coarser settings, as a game engine treats a rope

The first thread ran at a 5 µs step: light, stiff contacts between the thread, the needle and the holding parts settle in microseconds. With the bond replacing those contacts, we coarsened the thread the way rope simulations do: a 50 µs RK4 step, contacts that settle in 200 µs, a looser solver, and added joint inertia and damping that slow the thread's fastest wiggles without changing its shape or sag.

SettingFirst threadFast thread
Step5 µs50 µs
Segments32 × 1.4 mm (44 mm), later 6 × 1.4 mm18 × 0.46 mm (8.25 mm)
Contact time constant20 µs200 µs
Joint armature, dampingnone3 × 10−17 kg m², 1 × 10−12 N m s
Tissue gripexplicit spring-dampersoft constraint with slip
Wall time per simulated secondabout 600 s (44 mm thread)about 10 s (one thread), 20 s (three)
for b, eq in grips:                       # each segment that has entered the tissue
    point = data.xipos[b]
    if self.surface(point) - point[2] <= 0:  # pulled out: let go
        data.eq_active[eq] = 0
        continue
    force = norm(data.efc_force[rows_of(eq)])
    if force > p.f_retain:                   # slip: move the pin toward the segment
        anchor = tissue_origin + m.eq_data[eq, 0:3]
        m.eq_data[eq, 0:3] = point + (anchor - point) * p.f_retain / force - tissue_origin

src/sixlegs/neural_insertion/tube_cycle.py, TubeTissue: the tissue's grip as constraints MuJoCo solves implicitly.

Each failure was watched live in the viewer. A thread built on the tube's centreline settled 80 µm onto the bore and the needle missed its end. At 45°, long rigid segments jammed in the tube's mouth; with shorter ones the thread wrapped the needle at insertion speed, so the tube went to 15°. An explicit tissue spring on 0.8 µg segments diverged once the bond let go, and an undamped thread whipped after release.

17 · Three sites in a row

October 6 · a new thread for each site, placed threads stay simulated

A run places threads at sites 0, 5 and 1. After each site the tool lifts 8 mm, clear of the 6 mm of thread left standing; a spare thread appears in the tube; the tool moves over the next site, measures the target, makes one correction from the measured needle tip, and inserts. Placed threads stay simulated and gripped by the tissue, and the run checks that the tool's moves do not disturb them.

Three sites without disturbances. The needle's work plays at 0.02× real time, the moves between sites at 0.2×.
Three lime threads standing in three target rings
After a run: each thread stands in its target ring, its end 2 mm deep.

18 · Disturbances in the full cycle

October 6 · the same layer as alignment, and the tissue really moves

The insertion runs use the disturbance layer of chapter 11 at scale v2: force noise, extra friction and table vibration on every slide; a late, noisy needle-tip measurement; a target estimate with noise, bias and drift; and breathing and pulse. In alignment only the target point moved. Here the phantom itself is moved, and its surface, the targets, the puncture model and the grip on placed threads move with it. The replay at the top draws all of it: force arrows on the slides, the moving target's path, the estimate and the measurement while the robot measures.

Three sites at level 2, the stress setting.

19 · The insertion baseline

The scripted yardstick on evaluation seeds fixed before any training

Ten evaluation seeds (1001–1010) at levels 1 and 2, and one undisturbed run, three sites each. Placement is the distance from the thread's end to the target ring's centre. A learned insertion policy will be compared on these same seeds and levels.

Placement of every thread, by site and level

Where the misses come from. At site 0 the tool starts right over the target and the thread has not moved, so its end waits on the needle's axis and lands within a few micrometres. Before every later site the tool lifts and travels, and the thread settles against the tube's bore: its end then waits about 18 µm off the needle's axis (60 runs, 180 sites). The yardstick aims the needle, not the thread's end, so it carries that offset into the tissue; letting go adds a sideways jump (median 3 µm, 90th percentile 16 µm). Sixty runs with random site orders measured both, and they became the thread model of the training environment in chapter 20.

20 · A fast environment for learning the insertion

October 6 · millions of episodes need a faster world than the full thread

The full simulation is the reference, but at about 20 s of computing per simulated second it cannot supply the hundreds of millions of steps that PPO needs. The training environment is a C core, insert_core.h, built on the alignment core: the same robot, servo, disturbance layer and needle–tissue force model, at a 1 ms step, with 2,048 copies stepped on 28 CPU threads. It does not simulate the thread. It models the two things about the thread that decide where it lands, both measured in the full simulation: where the thread's end waits relative to the needle's axis, and how far it jumps sideways when the needle lets go.

Every episode is new

drawn from the episode's seed

  • The phantom turned by a random angle; its vessels turn with it
  • 0 to 5 threads already standing at earlier sites
  • A random site, clear of vessels and of those threads
  • The tool starting anywhere over the field, 4–10 mm up
  • Disturbances from none to stress, and the thread end's offset and release jump

What the policy sees

24 numbers, 50 times a second, delayed and noisy

  • Needle point to the hover pose, coarse and fine, and its velocity
  • The thread end's offset from the needle's axis, as a camera would measure it
  • The nearest vessel, the height above the dome
  • The two nearest standing threads
  • Its previous actions, phase and time left

What it does

4 actions

  • X, Y and Z stage velocities
  • Start the stroke. The stroke is the approved programmed motion: 40 ms onto the thread, 25 ms in, 30 ms at depth, release, snap back, with the stages held still

What it is rewarded for

placement, safety, smoothness

  • Progress toward the pose that puts the thread's end over the target
  • 2 × exp(−placement / 10 µm) when the thread is released
  • −2 for touching tissue or a standing thread, dragging the waiting thread, or missing its end
  • Small costs for jerky motion and passing low over vessels
Checked against the full simulation before training. A scripted controller that compensates for the measured thread-end offset, given the same 24 inputs built from the full simulation's own sensing, placed three threads at 7.8, 6.2 and 6.3 µm (level 1, seed 1001), close to the 5.6 µm median the C model predicts for it. The same harness, policy_cycle.py, later runs every learned checkpoint in the full simulation.

21 · Learning to insert

October 6 · two runs on the RTX 4090, about two hours in all

Run 1 learned within minutes to approach before striking, and showed a gap in the environment: after the stroke started, the stages still took velocity commands, which on the real machine would drag the needle sideways through tissue. Run 2 holds the stages still from the start of the stroke and times the stroke exactly as the full cycle does. It ran to 300 M steps, and its success on training episodes climbed from zero to 72%.

Threads within 10 µm (training rollouts, with exploration noise)

Mean placement error of placed threads (training rollouts)

Checkpoints on 200 held-out episodes

Within 10 µm, deterministic policy, training environment

The moment to strike is learned too. The fourth action starts the stroke when it exceeds 0.5. In run 1 the policy's average output stayed just below that, and its exploration noise fired the stroke: it had learned where to be, with the output rising from −2.6 far away to +0.17 within 10 µm of the hover pose, but the exact moment was left to chance. In run 2 the noise-free policy strikes on its own from the first checkpoint. Results are reported both ways.

22 · Results in two simulations

Seeds fixed before training; the same seeds and sites for every controller

60/ 60
threads placed in 21 full-simulation runs
no tissue or thread contact, no failed run
8.0µm
median placement, nominal disturbances
approved cycle 23 µm on the same seeds
20/ 30
within 10 µm, nominal, full simulation
approved cycle 11, compensating script 19
84%
within 10 µm, 200 held-out episodes, nominal
compensating script 80.5%, needle-aiming 3%

Two yardsticks frame the result. The approved cycle aims the needle at the target, as chapter 19 did. A second scripted controller, written for this comparison, reads the measured thread-end offset and aims the thread's end instead; it is the strongest hand-written controller we have. At nominal disturbances the learned policy places nearly twice as many threads within 10 µm as the approved cycle (20 against 11 of 30) and cuts the median from 23 to 8 µm; at the stress level the medians are 12 and 25 µm. Against the compensating script it is level in the full simulation at nominal disturbances, ahead in the training environment, and behind in the full simulation at the stress level.

Threads within 10 µm, full simulation (bar labels: medians)

Every placed thread, full simulation

The final policy, deterministic, at nominal disturbances (seed 1001): approach, strike, three threads at 1.8, 9.7 and 30.8 µm. The needle's work plays at 0.02× real time, the moves between sites at 0.2×. The replay at the top of the page has all seven runs of each controller.

Full simulation, 30 threads per level

Training environment, 200 held-out episodes per level

23 · Where the remaining micrometres come from

The final policy's 63 threads in the full simulation, followed through the stroke

Following each thread's end from the moment the policy strikes separates what the policy controls from what it does not. The policy's part is the aim: when the stroke starts, the thread's end is a median 6 µm from the target. The programmed stroke carries it into the tissue and moves it about 3 µm. The release usually adds another 3 µm, but now and then the end jumps: 8 of the 63 threads moved more than 15 µm sideways when the needle let go, and those set the 90th percentile.

Thread end to target at three moments

The release is the next lever. While the thread is held at depth, the needle's contact with the thread it carries works against the bond, and the gripped segments jitter. With that contact off, the jitter fell about eightfold in a test. Calmer threads should mean smaller release jumps for every controller.

24 · What comes next

Each step is measured on the same seeds as everything above

  1. Calm the thread at depthTurn off the needle's contact with the thread while the needle carries it, re-measure the release jumps over the same 60 runs, refit the training environment's thread model, and retrain. Test: jitter at depth fell from up to 515 rad/s to 63 rad/s with that contact off.
  2. Close the stress-level gap between the simulationsThe policy is ahead of the compensating script in the training environment at level 2 and behind it in the full simulation. Finding what the C model misses under stronger disturbances will make training transfer better. Level 2: 69% vs 64.5% in C; 12 vs 18 of 30 in the full simulation.
  3. Harder conditions, more seedsRaise the disturbances to where hand-written compensation breaks down, cover sites across the whole dome, and train three independent seeds per policy, as the brief asks.
  4. Learn more of the cycleThe stroke and the release are programmed today. Learning the stroke's timing, depth and release moment is the natural next phase for the policy.
  5. Then the real handlingA cartridge mechanism instead of stand-in constraints, a calibrated tissue model with a dura layer, threads as thin ribbons, and camera images instead of simulator measurements.

Materials and tissue parameters are presets, not measurements, and the controllers read simulator measurements rather than camera images.