Teaching a Robot Arm Tic-Tac-Toe with 45 Demonstrations

An SO-101 arm, a SmolVLA policy fine-tuned on Colab and Claude picking the cell: the path from zero, two surprises on the arm, and how small the evidence still is.

I taught an SO-101 robot arm to play tic-tac-toe, or XOX as we call it in Turkish. The arm plays X and I play O. Claude looks at the board through a camera and picks a cell. The moving is done by SmolVLA, a small neural network that I fine-tuned on 45 moves I showed the arm by hand.

My first idea was simple: show the arm each piece going to each cell once, and train at the end. It was wrong. A model like this does not learn "pick" and "place" as ideas. It imitates what it saw, and it needs to see each situation many times. So I started with a pilot whose goal was not success, but to see the whole chain work once, from recording to training to the arm.

This post is the short version. The step-by-step one, with the file that does each step, is in the garden, and the code is on GitHub. I did the recording and the playing. The code and the log analysis were written together with Claude Code.

The path

Fix the scene. The grid, the piece slots, the arm base and the camera above the table sit on tape marks, and the light stays the same. The model is lost in a scene it has not seen.

Record. I moved a second arm by hand and the robot copied it. One move per episode: leave the rest pose, pick the X from one fixed slot, put it in the cell, come back. Five episodes for each of the nine cells, on boards that look like a real game, from empty to nearly full. Then I checked every episode and recorded one again. The result is 45 episodes, about 16 minutes, public on Hugging Face.

Train. On a Colab A100 I fine-tuned lerobot/smolvla_base for 20,000 steps, which took about 3 hours 45 minutes. Before the model touched the arm I checked it on the Mac. When I swap the sentence for another cell on the same camera frame, its prediction changes. So it listens to the sentence.

Two surprises on the arm

On the first run the arm stayed in its rest pose. On the second it waited 6.8 seconds, picked the X and ran out of time above the cell. The third worked, but in jerks.

From the second run on I logged every command, and the logs showed why. The model outputs motion in chunks of 50 steps, about 1.7 seconds. An inference setting I had copied from a recipe, without measuring it here, computed each new chunk from an old observation and appended it to the end of the previous one. Every 50 steps the arm was told to jump back: in the second and third runs, 17 of the 18 jumps above 6° sit exactly on a chunk boundary. The fix was one flag, --inference.type=sync. On the next run the arm started at second 1.3 and no jump was above 6°.

There was a second cause, and it was me. In the recordings I waited 2.4 seconds on average before I moved, and the model learned to wait. No flag fixes that. It is in the data.

Giving it a brain

The model knows nine sentences, like put the red X in the top left cell. "Put it in the corner" means nothing to it. So Claude gets one tool with one argument, the cell, and code writes the exact training sentence. Claude never drives the arm.

Claude also reads the board from the camera above the table. On eight frames from real runs, each read twice, all 16 reads were correct. In the automatic mode nothing is typed: the camera sees the X the arm placed and the O I placed.

What works, and how little I know

Ten moves ran to their end on the real arm, and all ten put the X in the right cell. Six of the nine cells have been tried. All of it happened on one evening, on the table the data was recorded on, and neither game was played to the end. That is a pilot. It is not a success rate.

I did not expect even this much. LeRobot's documentation suggests about 50 episodes for a single task, and here 45 are split over nine.

What is next

To make it better, the first thing I will do is add data: 30 to 40 episodes per cell, with full boards, and this time I start moving the moment recording starts. Then ten trials per cell on boards the model has not seen, a minimax check on Claude's choice, and the O pieces.

If you want to try it yourself, start with the garden. It follows the order I did things in.