How this was built
A harness, a brief, and two coding agents working in a loop
Two things made this possible: a harness and a brief. The harness was in place before this project began. It is a repository set up for simulation and learning work: MuJoCo for physics, PufferLib 5.0 for reinforcement learning, a locked Python environment, Git LFS for assets, a commit-based workflow between a laptop and a remote RTX 4090 workstation, earlier MuJoCo projects to build on, and written engineering rules for how simulations are modelled, tested, recorded and documented here. The brief states what to build and the gates the work must meet.
Everything specific to this robot was written by two coding agents working together inside that harness from the brief, one running GPT-6 Astra and one running Claude Opus 5.5. That includes the workcell and needle model, the unit rescaling, the C++ thread model and contact kernel, the C task core shared by training, tests, evaluation and replay, the disturbance layer, the PufferLib environment, the evaluation tools, and this page with its WebAssembly replay. The project's first commit was on October 5.
- Brief and gateswhat to build, how it is judged
- Planthe next measurable step
- Buildcode, models, configuration
- Runsimulation or training
- Measureagainst the gates, on held-out seeds
- Recordresults and failures, committed
- Change one thingthen run again
We call this improvement loop AutoResearch. The agents plan a step, implement it, run it, measure the result against the brief's gates, record what failed, and change one variable before the next attempt. The training runs in this journal show the pattern: a diverged run exposed a reward bug, two runs that blew up led to lower learning rates, and a stable but slow run led to a longer training budget. A person sets the brief, reviews the results and corrects course. Their corrections were decisive: train on the GPU workstation rather than fall back to a CPU, make the gates serve the mission rather than the reverse, and pursue robustness under disturbances. Their checks of this page in a browser found four display bugs.
"Self-improving" means the work improves, not the agents. The models behind them do not change. The code, the physics, the policies and the process get better through measured iterations, and every mistake stays on the record.
1 · A static workcell first
October 5 · model, inspect, then move
We built the robot as a static scene before anything moved: an XYZ gantry, a downward insertion slide with a needle, a mounted microscope, and a domed tissue phantom with six targets and vessel markings. Clearance and limits were checked at 120 sampled poses before any dynamics. The first tool also had a thread retainer slide and a cassette of thread samples; the alignment experiments of chapters 7–14 use it. The insertion tool replaces them with a thread tube beside the needle (chapter 15).
| Axis | Travel (mm) | Moving mass (kg) | Force limit (N) |
|---|---|---|---|
| X | −175…175 | 2.4 | ±100 |
| Y | −75…75 | 1.3 | ±80 |
| Z | −8…50 | 0.65 | ±40 |
| Insertion (down) | 0…18 | 0.025 | ±2 |
| Retainer (down), first tool only | 0…6 | 0.004 | ±0.3 |
The workcell, part by part
Each view highlights one part against the faded machine, in the insertion workcell at a recorded moment: the thread stuck to the needle as it enters the tissue. Select a part to enlarge it; arrow keys step through.
2 · A 40 µm thread at SI scale will not compile
October 5 · unit rescaling
A 44 mm thread of 40 µm diameter, split into 32 segments, has rotational inertias near 10−19 kg m². MuJoCo rejects them at compilation. Rather than inflate the mass or radius, we rescaled the engine's units: millimetres and grams, with seconds unchanged. Every dimensional quantity converts at the model boundary, and results are reported back in SI.
@dataclass(frozen=True)
class Units:
name: str
length: float = 1.0 # engine length units per metre
mass: float = 1.0 # engine mass units per kilogram
@property
def force(self):
return self.mass * self.length
@property
def torque(self):
return self.mass * self.length**2
MM_G = Units("mm_g_s", 1000, 1000)
src/sixlegs/neural_insertion/units.py: forces scale by M·L, inertias and torques by M·L², moduli by M/L.
The physical mass and dimensions are unchanged, and bending then matched the analytical cantilever within 2.4%. The same conversion applies to every scene we build, so the robot and the thread share one unit system.
3 · Porting a published rod model, and correcting it
October 5 · discrete elastic rods inside MuJoCo
We ported the published discrete-elastic-rod (DER) plugin for MuJoCo to version 3.12 as a separate library, leaving the engine untouched. An energy-gradient audit found that the published force mapping produced zero restoring torque for a purely twisted rod despite positive twist energy. We kept the original variant and added a corrected one that projects the same Cartesian forces through MuJoCo's own Jacobians, plus the material-frame end moments. Its forces match the energy gradient to 2.85 × 10−9.
// Project the same Cartesian elastic forces through MuJoCo's Jacobians,
// and include material-frame end moments.
mju_copy(d->qfrc_passive, before.data(), m->nv);
const mjtNum zero[3] = {0, 0, 0};
for (int i=0; i<r->nv+2; ++i) {
int body = r->i0 + std::min(i, r->nv);
mj_applyFT(m, d, r->nodes[i].force.data(), zero, r->nodes[i].pos.data(),
body, d->qfrc_passive);
}
double moment = r->beta_bar*(r->edges[r->nv].theta-r->edges[0].theta)/r->bigL_bar;
Eigen::Vector3d first = moment*r->edges[0].e.normalized();
Eigen::Vector3d last = -moment*r->edges[r->nv].e.normalized();
mj_applyFT(m, d, zero, first.data(), r->nodes[0].pos.data(), r->i0, d->qfrc_passive);
mj_applyFT(m, d, zero, last.data(), r->nodes[r->nv+1].pos.data(),
r->i0+r->nv, d->qfrc_passive);
src/sixlegs/neural_insertion/der_bridge.cc, the corrected direct projection.
Bending stayed within 2.4% of the analytical cantilever, an undamped bent rod relaxed with energy drift falling from 0.075% to 0.038% as the step halved, and forces were identical after copy, restore and reset. For sim-to-real this matters more than any single benchmark: a rod whose twist cannot push back would mislead every later thread-handling result.
4 · The contact model
October 5 · a continuous contact force and a fourth-order integrator
MuJoCo's soft contact computes a reference acceleration a_ref = −B·v − K·d(r)·pos with an
impedance d(r). With d held constant, as in our first law, the damping force switches on
as a step at zero penetration, so a grazing contact receives a finite impulse. A penetration ramp makes the
force continuous in the state. And because the default integrator is first order through a contact lasting a few
tens of microseconds, we step the thread with MuJoCo's RK4, which evaluates collision and the constraint solve at
four stages. Both are standard MuJoCo settings; the engine is unmodified.
Before relying on it we checked the engine against an independent calculation that never calls MuJoCo: same-state contact forces match a closed-form solution to 10−11, and agree across unit systems to 10−9.
def law_for(tau):
return ContactLaw(f"ramp_{tau*1e6:g}us", d0=mujoco.mjMINIMP, power=1., time_constant_s=tau)
law = law_for(2e-5) # thread contact in the task regime: 20 µs, stepped at 5 µs with RK4
# solimp = (0.0001, 0.9999, 1 µm width, 0.5, 1): impedance rises linearly over 1 µm
src/sixlegs/neural_insertion/task_fixtures.py with contact_laws.ContactLaw.
if (rk) { // same sequence as mj_step for RK4
mj_checkPos(m, d);
mj_checkVel(m, d);
mj_forward(m, d);
mj_checkAcc(m, d);
} else {
mj_step1(m, d);
}
...
mj_RungeKutta(m, d, 4);
src/sixlegs/neural_insertion/contact_kernel.cc: our batch stepping kernel with per-step contact telemetry; bit-identical to mj_step.
5 · Thread contact in the task regime
October 6 · gates derived from what the robot will actually do
Each phase of the mission touches the thread in its own way, at millimetres per second. We built one fixture per phase, with a rigid 150 µm probe driven only by force-limited actuators, and required the thread's centreline to agree within 1 µm when the timestep is halved and within 0.01 µm across unit systems, with penetration checked at every step over every contact.
| Task phase | Fixture | Timestep, 5 → 2.5 µs | Units |
|---|---|---|---|
| Transport and align | Settle on the support | 0.13 µm | ≤ 1.5e−10 µm |
| Insert to depth | Needle drags across thread (it rolls) | ≤ 5.1e−7 µm | ≤ 6e−7 µm |
| Pick and grasp | 100 µN press, force balance ≤ 6.8e−5 | ≤ 2.4e−7 µm | ≤ 9e−11 µm |
| Release | Hook slides out, 150 µm fall | 0.19 µm | ≤ 7.3e−7 µm |
Every fixture passes at a 5 µs RK4 step with a 20 µs contact time constant, for both material presets. Release failed at 10 µs (1.58 µm), because its short fall lands faster than settling does; we refined the step rather than relax the gate, and kept the failure on record.
6 · The robot's first motion
October 6 · programmed servo, gates fixed before the run
A programmed servo adds inertial, gravity, viscous and Coulomb feedforward and a bounded integral term to each force-limited position actuator, by offsetting its command. Two design facts surfaced on the first run: Z reaches only 8 mm down, so the needle slide reaches the surface; and both fine slides hang against gravity at their retracted stops, so they are carried 0.5 mm out.
ctrl = q_ref + (kv*v_ref + M*a_ref + qfrc_bias + damping*v_ref + coulomb + I) / kp
# applied force = kp*(q_ref - q) + kv*(v_ref - qdot) + feedforward + I, clipped by the actuator force range
src/sixlegs/neural_insertion/motion.py
7 · The learning task, exactly
October 6 · trained task: needle alignment (robot only)
One C core (native/surgical_core.h) holds physics stepping, servo, observations and rewards. It is
compiled into the PufferLib 5.0 adapter for training and into a local library for tests and evaluation, so they
cannot drift apart. Physics runs at 1 ms; the policy acts at 50 Hz (20 physics steps per action).
Episode
- Start: gantry at rest, X uniform in ±40 mm and Y in ±30 mm about the field centre; needle tip 3–8 mm above the dome; Z stage at −6 mm, retainer at 0.5 mm.
- Goal: one of six targets, uniformly; the hover point 1 mm above it.
- Success: lateral and vertical error under 10 µm and tip speed under 0.2 mm/s for 15 consecutive policy steps (0.3 s).
- Failure: any contact between robot and environment, or a nonfinite state.
- Deadline: 250 policy steps (5 s). Success, failure and deadline all end the episode.
- Seeds: training base seed 1 per world; evaluation seeds 1,000,000–1,000,199, never used in training or tuning.
Observation (16 values, exact simulator state)
| Index | Quantity | Scaling | Clip |
|---|---|---|---|
| 0–2 | Tip minus goal, x y z | ÷ 5 mm | ±5 |
| 3–5 | Tip velocity, x y z | ÷ 40 mm/s | ±2 |
| 6–8 | Previous applied action | — | ±1 |
| 9–11 | Tip minus goal, fine scale | ÷ 100 µm | ±1 |
| 12–13 | Offset to the nearest vessel centreline, x y | ÷ 5 mm | ±1 |
| 14 | Tip height above the dome | ÷ 10 mm | −1…2 |
| 15 | Remaining time fraction | 1 − step/250 | 0…1 |
for (int i = 0; i < 3; i++) {
double e = tip[i] - c->goal[i];
obs[i] = (float)sa_clip(e / 0.005, -5, 5);
obs[3 + i] = (float)sa_clip(vel[i] / SA_VMAX_XY, -2, 2);
obs[6 + i] = c->prev_action[i];
obs[9 + i] = (float)sa_clip(e / 1e-4, -1, 1);
}
double vx, vy;
sa_vessel(c, tip[0], tip[1], &vx, &vy);
obs[12] = (float)sa_clip(vx / 0.005, -1, 1);
obs[13] = (float)sa_clip(vy / 0.005, -1, 1);
obs[14] = (float)sa_clip((tip[2] - sa_surface(c, tip[0], tip[1])) / 0.01, -1, 2);
obs[15] = (float)(1.0 - (double)c->tick / SA_HORIZON);
Action (3 continuous values)
Each output is clipped to [−1, 1] and cubed, then scaled to needle-tip velocity: 40 mm/s in X and Y, 20 mm/s vertically. The cube gives 40 µm/s resolution at 0.1. The command drives the X and Y stages and the insertion slide; a 10 ms filter smooths the reference, which the programmed servo tracks.
Reward per policy step
| Term | Value |
|---|---|
| Progress (potential shaping) | Φ(d′) − Φ(d), with Φ(d) = −log₁₀((d + 2 µm) / 1 mm) and d the tip-to-goal distance |
| Action smoothness | −0.002 · Σ (aᵢ − aᵢ,prev)² on the clipped actions |
| Vessel proximity | −0.05 when the tip is under 3 mm above the dome and within the vessel radius + 0.3 mm of a centreline |
| Success | +2, episode ends |
| Collision | −2, episode ends |
The logarithmic potential makes going from 10 mm to 1 mm worth the same as going from 10 µm to 1 µm, so the policy gets useful signal at every scale. There is no curriculum: every episode is the full task.
double potential = sa_potential(distance);
double r = potential - c->potential;
c->potential = potential;
double smooth = 0;
for (int i = 0; i < SA_ACT; i++) {
double applied = sa_clip(action[i], -1, 1);
double da = applied - c->prev_action[i];
smooth += da * da;
c->prev_action[i] = (float)applied;
}
r -= 0.002 * smooth;
src/sixlegs/neural_insertion/native/surgical_core.h, after the fix described in chapter 9.
Baselines on the evaluation seeds
| Policy | Success | Collision | Mean time |
|---|---|---|---|
| Scripted reference (programmed) | 100% | 0% | 2.06 s |
| Zero action | 0% | 0% | — |
| Uniform random | 0% | 5% | — |
8 · The network we train
PufferLib 5.0 default recurrent policy, 128 hidden units, 2 MinGRU layers
Read from PufferLib's own code and the checkpoint's layout: 100,867 parameters, no biases. The same weights run on the GPU during training and on the CPU (deterministic mean action) for evaluation and this replay.
Training configuration
| Setting | Value |
|---|---|
| Algorithm | PufferLib 5.0 PPO (pinned revision 6ffa5b1), clip 0.2, value coefficient 2.0 |
| Optimizer | Muon, learning rate 0.003 cosine-annealed to 0 |
| Discount, GAE λ | 0.995, 0.95 |
| Entropy coefficient | 5 × 10−4 |
| Batch | 2048 worlds × 64-step horizon, minibatch 16,384 |
| Hardware | RTX 4090 for the network; MuJoCo physics on 28 CPU threads |
| Budget | 100 M steps (99.9 M), 29 minutes, about 59,000 steps/s |
void puf_step(Env* env) {
SAEpisode e;
float reward;
int done = sa_step(&env->core, env->agents[0].actions, env->agents[0].observations, &reward, &e);
env->agents[0].rewards[0] = reward;
env->agents[0].terminals[0] = (float)done;
if (done) {
env->log.perf += e.success;
...
sa_reset(&env->core, env->agents[0].observations);
}
}
experiments/neural_insertion/puffer/surgical_align/surgical_align.h: the PufferLib adapter is a thin wrapper over the shared core.
9 · Training on the RTX 4090: two failures, then success
October 6 · three runs, one change at a time
Final lateral error at episode end (training rollouts)
Success rate (training rollouts, noisy actions)
Policy update size (KL divergence)
Collision rate
- double da = action[i] - c->prev_action[i];
+ double applied = sa_clip(action[i], -1, 1);
+ double da = applied - c->prev_action[i];
+learning_rate = 0.003
gamma = 0.995
10 · Results on 200 held-out episodes
Deterministic actions, evaluated on CPU with PufferLib's own network code
| Policy | Success | Collisions | Mean time | Worst final error | Low steps over vessels |
|---|---|---|---|---|---|
| Learned, final (99.9 M) | 100% | 0% | 1.40 s | 8.6 µm lat., 4.4 µm vert. | 0.25 |
| Learned, 65.5 M | 100% | 0% | 1.98 s | 10.0 µm lat., 7.4 µm vert. | 0.39 |
| Scripted reference | 100% | 0% | 2.06 s | 3.4 µm lat. | 1.68 |
The learned policy is a third faster and passes low over vessels less often, but settles closer to the tolerance than the scripted controller: an episode ends as soon as the 0.3 s hold is met. Compare both in the replay at the top.
11 · A perfect world is too easy
October 6 · the brief is rewritten around disturbances
With exact state, no delay and a still target, both the scripted controller and the learned policy succeed every time. A real stage gets none of that. We rewrote the brief: every mission phase is learned under one shared, documented disturbance layer; the programmed servo stays only as the motor drive; scripted controllers become yardsticks. Results are reported on held-out seeds at stated disturbance levels, including full strength.
The layer lives in the same C core used for training, evaluation and replay. One number, the level from 0 to 1, scales every magnitude, and in training each episode draws its own level uniformly from that range. The magnitudes are illustrative until replaced by measurements of the real stage, sensors and tissue.
| Disturbance | Model | Full strength (level 1) | Enters through |
|---|---|---|---|
| Slide force noise | Low-pass random force, 50 ms correlation time | std 100, 100, 50, 10, 2 mN on X, Y, Z, insertion, retainer | Applied joint force |
| Extra friction | Coulomb, smoothed below 0.1 mm/s | up to 100, 100, 50, 5, 1 mN, drawn per episode | Applied joint force |
| Table vibration | Three sinusoids per axis, 5–60 Hz | up to 10 mm/s² per component | Inertial force on each carriage |
| Tip sensing | Delay, then Gaussian noise | 0–10 ms latency; 0.5 µm and 0.05 mm/s noise | Observation only |
| Target estimate | Bias, slow drift and per-step noise | 1 µm bias, 1 µm/√s drift, 3 µm noise per axis | Observation only |
| Tissue motion | Breathing 0.2–0.4 Hz, pulse 1–2 Hz | up to 30 + 10 µm vertical, 10 + 3 µm lateral amplitude | True target position and velocity |
// Every 1 ms physics step: rebuild the external force on each slide from scratch.
for (int j = 0; j < SA_NJ; j++) {
double sigma = SA_FORCE_SIGMA[j] * q->level; // low-pass force noise
q->force[j] += -q->force[j] * dt / SA_FORCE_TAU
+ sigma * sqrt(2 * dt / SA_FORCE_TAU) * sa_normal(&c->drng);
double v = d->qvel[c->dof[j]];
double inertial = -c->mass[j] * dot(a_base, SA_AXIS[j]); // table vibration
double friction = -q->friction[j] * tanh(v / 1e-4); // extra Coulomb friction
d->qfrc_applied[c->dof[j]] = q->force[j] + friction + inertial;
}
The replay at the top draws every component in the scene. Each carriage carries force arrows at 5 cm per 0.1 N: amber for noise, cyan for friction, lilac for vibration. The table shows its acceleration and recent trace. The Micro camera shows, at true scale, the moving target with its 10 µm tolerance cylinder, the noisy target estimate and the late tip measurement. Values sit in panels beside each object. Airflow is not modelled yet; it matters for the thread, which this task does not carry.
12 · Training under disturbances
October 6 · runs 4 and 5, one change between them
Final lateral error at episode end (training rollouts)
Policy update size (KL divergence)
Success rate (training rollouts, noisy actions)
Collision rate
13 · Robustness on 200 held-out episodes
Deterministic actions, the same 200 predetermined seeds at every level
Success against disturbance level
14 · What the training used
One workstation: RTX 4090 and a 16-core Ryzen 9 5950X
This workload is CPU-bound. PufferLib steps 28 copies of the MuJoCo world on CPU threads, about 29 of the machine's 32 hardware threads, while the GPU runs only the 101k-parameter network and its PPO update. It idles near 10% utilisation and about 59 W of its 450 W limit, with under 1 GB of memory for the trainer. Runs 1–3 predate the resource logger, so their figures come from PufferLib's own dashboard, which reports device-wide GPU memory including the desktop. From run 4, a logger samples GPU power, temperature, the trainer's GPU memory and host CPU load once a minute.
15 · Thread insertion: a thread tube and stand-in handling
October 6 · three mechanisms, then a decision about what to prove
Picking up a 40 µm thread is the hard mechanical part of this machine. We tried it three ways: a slotted needle with a latch over a 44 mm free thread (about 600 times slower than real time, and the eyelet rattled loose on landing); a needle with a ledge and a rotary pincher; then a cartridge with a cannula and a spring latch under the thread's loop. Each attempt was watched on video, and each taught something, but none of it was what this project sets out to prove.
After reviewing photos of a current insertion robot, the user set the goal: prove the robot's movements and abstract the thread handling. The thread waits in a tube on an arm beside the needle, angled 15° from vertical, its end on the needle's path. On its way down the needle sticks to that end, carries the thread into the tissue, and lets go at depth. The rest is physical: the robot's dynamics, the thread as an elastic rod, the tube's walls, the puncture and the tissue's grip.
| Stand-in (a MuJoCo constraint) | Switched on | Switched off | Stands for |
|---|---|---|---|
| Tube hold on the thread's back end | thread ready in the tube | when the needle has the thread | a feed brake |
| Needle bond to the thread's end | needle point within 30 µm of the end | at depth | a chemical attachment and release |
| Tissue grip on each segment that enters | below the surface, near the puncture | pulled out of the tissue | tissue holding the thread; slips above 16.7 µN per 0.46 mm segment |
| New thread in the tube | between sites (a spare, straight and at rest) | feeding the next thread | |
16 · A fast thread
October 6 · coarser settings, as a game engine treats a rope
The first thread ran at a 5 µs step: light, stiff contacts between the thread, the needle and the holding parts settle in microseconds. With the bond replacing those contacts, we coarsened the thread the way rope simulations do: a 50 µs RK4 step, contacts that settle in 200 µs, a looser solver, and added joint inertia and damping that slow the thread's fastest wiggles without changing its shape or sag.
| Setting | First thread | Fast thread |
|---|---|---|
| Step | 5 µs | 50 µs |
| Segments | 32 × 1.4 mm (44 mm), later 6 × 1.4 mm | 18 × 0.46 mm (8.25 mm) |
| Contact time constant | 20 µs | 200 µs |
| Joint armature, damping | none | 3 × 10−17 kg m², 1 × 10−12 N m s |
| Tissue grip | explicit spring-damper | soft constraint with slip |
| Wall time per simulated second | about 600 s (44 mm thread) | about 10 s (one thread), 20 s (three) |
for b, eq in grips: # each segment that has entered the tissue
point = data.xipos[b]
if self.surface(point) - point[2] <= 0: # pulled out: let go
data.eq_active[eq] = 0
continue
force = norm(data.efc_force[rows_of(eq)])
if force > p.f_retain: # slip: move the pin toward the segment
anchor = tissue_origin + m.eq_data[eq, 0:3]
m.eq_data[eq, 0:3] = point + (anchor - point) * p.f_retain / force - tissue_origin
src/sixlegs/neural_insertion/tube_cycle.py, TubeTissue: the tissue's grip as constraints MuJoCo solves implicitly.
17 · Three sites in a row
October 6 · a new thread for each site, placed threads stay simulated
A run places threads at sites 0, 5 and 1. After each site the tool lifts 8 mm, clear of the 6 mm of thread left standing; a spare thread appears in the tube; the tool moves over the next site, measures the target, makes one correction from the measured needle tip, and inserts. Placed threads stay simulated and gripped by the tissue, and the run checks that the tool's moves do not disturb them.
18 · Disturbances in the full cycle
October 6 · the same layer as alignment, and the tissue really moves
The insertion runs use the disturbance layer of chapter 11 at scale v2: force noise, extra friction and table vibration on every slide; a late, noisy needle-tip measurement; a target estimate with noise, bias and drift; and breathing and pulse. In alignment only the target point moved. Here the phantom itself is moved, and its surface, the targets, the puncture model and the grip on placed threads move with it. The replay at the top draws all of it: force arrows on the slides, the moving target's path, the estimate and the measurement while the robot measures.
19 · The insertion baseline
The scripted yardstick on evaluation seeds fixed before any training
Ten evaluation seeds (1001–1010) at levels 1 and 2, and one undisturbed run, three sites each. Placement is the distance from the thread's end to the target ring's centre. A learned insertion policy will be compared on these same seeds and levels.
Placement of every thread, by site and level
20 · A fast environment for learning the insertion
October 6 · millions of episodes need a faster world than the full thread
The full simulation is the reference, but at about 20 s of computing per simulated second it cannot supply the
hundreds of millions of steps that PPO needs. The training environment is a C core, insert_core.h,
built on the alignment core: the same robot, servo, disturbance layer and needle–tissue force model, at a 1 ms
step, with 2,048 copies stepped on 28 CPU threads. It does not simulate the thread. It models the two things about the thread that decide
where it lands, both measured in the full simulation: where the thread's end waits relative to the needle's axis,
and how far it jumps sideways when the needle lets go.
Every episode is new
drawn from the episode's seed
- The phantom turned by a random angle; its vessels turn with it
- 0 to 5 threads already standing at earlier sites
- A random site, clear of vessels and of those threads
- The tool starting anywhere over the field, 4–10 mm up
- Disturbances from none to stress, and the thread end's offset and release jump
What the policy sees
24 numbers, 50 times a second, delayed and noisy
- Needle point to the hover pose, coarse and fine, and its velocity
- The thread end's offset from the needle's axis, as a camera would measure it
- The nearest vessel, the height above the dome
- The two nearest standing threads
- Its previous actions, phase and time left
What it does
4 actions
- X, Y and Z stage velocities
- Start the stroke. The stroke is the approved programmed motion: 40 ms onto the thread, 25 ms in, 30 ms at depth, release, snap back, with the stages held still
What it is rewarded for
placement, safety, smoothness
- Progress toward the pose that puts the thread's end over the target
- 2 × exp(−placement / 10 µm) when the thread is released
- −2 for touching tissue or a standing thread, dragging the waiting thread, or missing its end
- Small costs for jerky motion and passing low over vessels
policy_cycle.py, later runs every learned checkpoint in the full simulation.21 · Learning to insert
October 6 · two runs on the RTX 4090, about two hours in all
Run 1 learned within minutes to approach before striking, and showed a gap in the environment: after the stroke started, the stages still took velocity commands, which on the real machine would drag the needle sideways through tissue. Run 2 holds the stages still from the start of the stroke and times the stroke exactly as the full cycle does. It ran to 300 M steps, and its success on training episodes climbed from zero to 72%.
Threads within 10 µm (training rollouts, with exploration noise)
Mean placement error of placed threads (training rollouts)
Checkpoints on 200 held-out episodes
Within 10 µm, deterministic policy, training environment
22 · Results in two simulations
Seeds fixed before training; the same seeds and sites for every controller
Two yardsticks frame the result. The approved cycle aims the needle at the target, as chapter 19 did. A second scripted controller, written for this comparison, reads the measured thread-end offset and aims the thread's end instead; it is the strongest hand-written controller we have. At nominal disturbances the learned policy places nearly twice as many threads within 10 µm as the approved cycle (20 against 11 of 30) and cuts the median from 23 to 8 µm; at the stress level the medians are 12 and 25 µm. Against the compensating script it is level in the full simulation at nominal disturbances, ahead in the training environment, and behind in the full simulation at the stress level.
Threads within 10 µm, full simulation (bar labels: medians)
Every placed thread, full simulation
Full simulation, 30 threads per level
Training environment, 200 held-out episodes per level
23 · Where the remaining micrometres come from
The final policy's 63 threads in the full simulation, followed through the stroke
Following each thread's end from the moment the policy strikes separates what the policy controls from what it does not. The policy's part is the aim: when the stroke starts, the thread's end is a median 6 µm from the target. The programmed stroke carries it into the tissue and moves it about 3 µm. The release usually adds another 3 µm, but now and then the end jumps: 8 of the 63 threads moved more than 15 µm sideways when the needle let go, and those set the 90th percentile.
Thread end to target at three moments
24 · What comes next
Each step is measured on the same seeds as everything above
Materials and tissue parameters are presets, not measurements, and the controllers read simulator measurements rather than camera images.