A benchmark of spatial-constraint puzzles for robot manipulation

Two claws, hooked through each other.

Pulling only jams them. They come apart when turned in the right order.

Here two robot arms do it in simulation, replaying the motion they planned.

RoboTangle

Can your robot get untangled?

Thirteen physical puzzles where the pieces are hooked, threaded, knotted or caged, and only the right sequence of moves gets them apart.

Scroll

RoboTangle is a benchmark of physical puzzles for robots: interlocked claws, cast-metal brainteasers, a caged hedgehog, a lock, a buzz wire, knots, threaded rope, fibers and a blanket.

Every puzzle comes in three tiers, easy, medium and hard, and every tier is a frozen simulation instance that always starts from the same state. An agent gets six hours and drives a simulated robot, a two-arm ALOHA or a Franka arm, by writing code. An attempt counts only if its recorded motion solves the puzzle again when it is replayed in a fresh simulator.

13
puzzle families
39
frozen instances, three tiers each
2
evaluation modes
6 h
per attempt
01

Puzzles

Each card plays the reference solution: the oracle's trajectory, replayed by the verifier in a fresh simulator, sped up. Switch tiers on a card to see how the puzzle tightens.

02

Try one yourself

Each rigid puzzle comes two ways. Play it: move or turn the highlighted piece with the handle, and it stops wherever it would pass through another piece. Or watch the oracle: its whole run in 3D, every piece where the simulator had it. The ropes are simulated live and shown solving themselves.

moves 0 time 0:00 blocked · the pieces touch
lockedapart
03

Two modes

Every instance is posed twice. In both, the agent's recorded motion is replayed in a fresh simulator, and only a replay that solves the puzzle counts.

privileged

Code with the simulator open

  • Object poses, contacts and kinematics can be read at any time.
  • Resets and replays are unlimited within the six hours.
  • The agent submits one trajectory; a separate verifier replays it twice.
from rcb_topogym import Session
s = Session("interlocked_claws", instance="/app/instance.json")
s.reset()                      # back to the frozen start, any time
s.objects()["piece_1"]          # poses, bounds, velocity
s.record_start()
s.step(action)                 # joints + grippers, one 1/40 s step
actions = s.record_stop()
s.save_trajectory(actions, "/app/output/trajectory.npz")
standard

The robot as a remote service

  • Camera images (RGB, depth on request) and the robot's joint readings, nothing else.
  • No object poses and no resets: like real hardware.
  • The service records the episode, which is then replayed and graded.
Overhead camera view of the claw puzzle
overhead
Left wrist camera view
left wrist
Right wrist camera view
right wrist
04

Results

Each row is one model working inside one harness, the agent program around it, such as Codex CLI or Claude Code. Pick a harness to compare the models run in it. VLA policies, which output robot actions directly, are ranked on their own.

Leaderboard

Harness

Agent rollouts

replays of graded agent trajectories, sped up