Teaching a Robot Arm Tic-Tac-Toe with 45 Demonstrations
An SO-101 arm, a SmolVLA policy fine-tuned on Colab and Claude picking the cell: the path from zero, two surprises on the arm, and how small the evidence still is.
I taught an SO-101 robot arm to play tic-tac-toe, or XOX as we call it in Turkish. The arm plays X and I play O. Claude looks at the board through a camera and picks a cell. The moving is done by SmolVLA, a small neural network that I fine-tuned on 45 moves I showed the arm by hand.
My first idea was simple: show the arm each piece going to each cell once, and train at the end. It was wrong. A model like this does not learn "pick" and "place" as ideas. It imitates what it saw, and it needs to see each situation many times. So I started with a pilot whose goal was not success, but to see the whole chain work once, from recording to training to the arm.
This post is the short version. The step-by-step one, with the file that does each step, is in the garden, and the code is on GitHub. I did the recording and the playing. The code and the log analysis were written together with Claude Code.
The path
Fix the scene. The grid, the piece slots, the arm base and the camera above the table sit on tape marks, and the light stays the same. The model is lost in a scene it has not seen.
Record. I moved a second arm by hand and the robot copied it. One move per episode: leave the rest pose, pick the X from one fixed slot, put it in the cell, come back. Five episodes for each of the nine cells, on boards that look like a real game, from empty to nearly full. Then I checked every episode and recorded one again. The result is 45 episodes, about 16 minutes, public on Hugging Face.
Train. On a Colab A100 I fine-tuned lerobot/smolvla_base for 20,000
steps, which took about 3 hours 45 minutes. Before the
model
touched the arm I checked it on the Mac. When I swap the sentence for another
cell on the same camera frame, its prediction changes. So it listens to the
sentence.
Two surprises on the arm
On the first run the arm stayed in its rest pose. On the second it waited 6.8 seconds, picked the X and ran out of time above the cell. The third worked, but in jerks.
From the second run on I logged every command, and the logs showed why. The
model outputs motion in chunks of 50 steps, about 1.7 seconds. An inference
setting I had copied from a recipe, without measuring it here, computed each
new chunk from an old observation and appended it to the end of the previous
one. Every 50 steps the arm was told to jump back: in the second and third
runs, 17 of the 18 jumps above 6° sit exactly on a chunk boundary. The fix
was one flag, --inference.type=sync. On the next run the arm started at
second 1.3 and no jump was above 6°.
There was a second cause, and it was me. In the recordings I waited 2.4 seconds on average before I moved, and the model learned to wait. No flag fixes that. It is in the data.
Giving it a brain
The model knows nine sentences, like put the red X in the top left cell.
"Put it in the corner" means nothing to it. So Claude gets one tool with one
argument, the cell, and code writes the exact training sentence. Claude never
drives the arm.
Claude also reads the board from the camera above the table. On eight frames from real runs, each read twice, all 16 reads were correct. In the automatic mode nothing is typed: the camera sees the X the arm placed and the O I placed.
What works, and how little I know
Ten moves ran to their end on the real arm, and all ten put the X in the right cell. Six of the nine cells have been tried. All of it happened on one evening, on the table the data was recorded on, and neither game was played to the end. That is a pilot. It is not a success rate.
I did not expect even this much. LeRobot's documentation suggests about 50 episodes for a single task, and here 45 are split over nine.
What is next
To make it better, the first thing I will do is add data: 30 to 40 episodes per cell, with full boards, and this time I start moving the moment recording starts. Then ten trials per cell on boards the model has not seen, a minimax check on Claude's choice, and the O pieces.
If you want to try it yourself, start with the garden. It follows the order I did things in.