🧪 New tutorials are being published — build from robot arms to sensors step by step
Skip to content

Stage 6: Model Deployment (Windows) ​

This stage loads the trained policy so the robot can execute tasks autonomously, and records evaluation videos to verify the results. This is the finale of the whole workflow and the key test of the training outcome.


Prerequisites ​

  • Stage 5: Model Training completed

  • Training produced outputs/train/soarm_amazing_hand_pick/checkpoints/last/pretrained_model/

  • The camera indices are recorded


Step 1: Confirm the Model Files ​

PowerShell
# Confirm the model directory exists
dir outputs\train\soarm_amazing_hand_pick\checkpoints\last\pretrained_model

It should contain model files such as model.safetensors.

⚠️ Note (model path): --policy.path must point to the pretrained_model directory (containing the config + weights), not the checkpoint root directory.


Step 2: Deploy and Evaluate ​

PowerShell
lerobot-rollout `
  --strategy.type=episodic `
  --robot.type=so101_amazing_hand `
  --robot.port=<follower_arm_com> `
  --robot.hand_port=<hand_com> `
  --robot.id=amazing_hand_follower `
  --robot.cameras='{
    wrist: {type: opencv, index_or_path: 0, width: 640, height: 480, fps: 30, fourcc: "MJPG"},
    top: {type: opencv, index_or_path: 1, width: 640, height: 480, fps: 30, fourcc: "MJPG"}
  }' `
  --policy.path=outputs\train\soarm_amazing_hand_pick\checkpoints\last\pretrained_model `
  --dataset.repo_id=rollout_soarm_amazing_hand_pick_eval `
  --dataset.root=D:\lerobot_data `
  --dataset.push_to_hub=false `
  --dataset.num_episodes=10 `
  --dataset.single_task="Pick up the cube with the dexterous hand" `
  --display_data=true

Replace <follower_arm_com> / <hand_com> with the actual COM numbers; replace the camera index_or_path with your camera indices.

💡 Notes: Use lerobot-rollout but do not add ****--teleop.type; the policy then controls the robot autonomously (replacing manual teleoperation). The data is saved as an evaluation set. --dataset.root / --dataset.push_to_hub=false are the same as in Stage 4; purely local saving requires no HF login.


Evaluation Procedure ​

  1. Return the robot + hand to the starting position

  2. Press Enter to start: the policy executes the task autonomously

  3. Observe whether the grasp succeeds (press Enter to continue after each episode)

  4. Repeat for num_episodes episodes

Evaluation metric: success rate = successful episodes / total episodes

⚠️ Note 1 (reset consistency): Start every episode from the same starting position, otherwise the policy fails to generalize and the success rate will be artificially low.

⚠️ Note 2 (safety): On the first autonomous run, it is recommended to keep a hand on the robot / run slowly and observe, to confirm the policy's motions are reasonable. The policy may make unexpected movements.

⚠️ Note 3 (expected success rate): ACT typically achieves a 50-80% success rate with 20 episodes of data. If it is lower than expected, go back and record more data or adjust the training step count.


Iterative Optimization ​

If the evaluation success rate is unsatisfactory, adjust in order of priority:

PriorityOptimizationAction
1Record more high-quality dataGo back to Stage 4 and record an additional 20-30 more consistent episodes
2Increase the training step countGo back to Stage 5, --steps=100000
3Check starting-position consistencyStrictly reset before every evaluation episode
4Adjust the task descriptionMake sure single_task matches the task

This completes the full closed loop of SO-ARM101 + AmazingHand: calibration → teleoperation → collection → training → deployment.


Troubleshooting ​

SymptomCauseSolution
Model fails to loadWrong/incomplete pathConfirm --policy.path points to the pretrained_model directory
Policy does not moveCamera/observation errorConfirm the camera indices match those at training time; check the --display_data view
Policy moves erraticallyInconsistent starting position / poor dataReset strictly; record more data
Behavior differs from trainingEnvironment differencesConfirm the cameras, lighting, and object positions match those at recording time