experiment · evergreen

Training

SmolVLA fine-tuned for 20,000 steps on a Colab A100, then three checks on the Mac before the arm.

A policy is the function that turns what a robot sees into what it does next. SmolVLA is a small one that takes camera images and a sentence. I did not train it from scratch. I started from lerobot/smolvla_base and fine-tuned it on my 45 episodes with LeRobot's default setting: only the part that produces actions is trained, the vision-language model stays frozen.

The numbers: batch size 64, 20,000 steps, about 3 hours 45 minutes on a Colab A100. The loss went from 0.479 at step 100 down to 0.077. The notebook does all of it and uploads a checkpoint every 5,000 steps, so a dropped connection does not cost the whole run.

Before the model touched the arm I checked three things on the Mac:

  • It loads onto the Mac's GPU and computes 50 steps of motion in 0.35 seconds.
  • On recorded frames its prediction differs from my recorded motion by about 0.4–1° on average. Those are training frames, so this only says it learned its data.
  • On the same frame, swapping the sentence for another cell changes the prediction by up to 50°. The model listens to the sentence.

None of this says the arm will succeed. It rules out the cheap mistakes before the motors are on.

#smolvla #colab #mps

See this note on the whiteboard →