RoboDojo#

RoboDojo connects Isaac Sim / IsaacLab, dual ARX-X5 arms and an RLinf Pi0.5 policy to RPent’s planners, tools and memory. The integration is experimental; simulator compatibility and task success require GPU validation. Follow the upstream RoboDojo and XPolicyLab instructions for simulator, CUDA and policy dependencies, and download the required assets and checkpoints. For shared backend interfaces, see Add a Robot or Simulator.

Python environment#

Install the runtime and agent dependencies from the RPent root in Python 3.11. The robodojo-sim extra installs rlinf-robodojo-runtime and the RLinf environment adapter; the runtime owns the Isaac Sim / IsaacLab pins. robodojo also installs SAM3 and the openpi policy runtime. Use one robot extra per environment. Run uv from this directory so it reads the root project’s cuRobo build dependencies, and pass this repository’s simulator overrides explicitly:

uv pip install -e ".[robodojo]" --extra-index-url https://pypi.nvidia.com \
  --override requirements/robodojo-override.txt

On Blackwell GPUs, append --torch-backend=cu128 to select a PyTorch build with sm_120 support. The default CUDA 12.6 build of PyTorch 2.7.0 does not support these GPUs. Use the same option when reinstalling this extra. Note that uv treats an already-installed 2.7.0+cu126 as satisfying the pinned torch==2.7.0 and keeps it, so switching an existing environment also needs --reinstall-package torch --reinstall-package torchvision; rebuilding the environment is the simpler alternative.

requirements/robodojo-override.txt carries the eight pins where the simulator stack disagrees with the agent stack or with rpent-openpi; passing them with --override keeps every other robot resolving the versions it was validated against. They allow dependency resolution; they do not establish simulation task success. The RLinf integration branch supplies both the environment adapter and pi05_robodojo_arx_x5. The RoboDojo preset selects the OpenPI eval loader with a 50-step horizon, a padded 32-dimensional model action, and 14-dimensional environment actions. Real-weight RPC inference has returned finite actions of shape (1, 50, 14); this is not a task-success guarantee.

The simulation bridge requires runtime version 0.3.0 or later. The Git references are temporary until versioned releases are published. Fresh dependency installation and GPU rollout must be validated separately; offline bridge tests do not establish simulator compatibility. IsaacLab’s upstream non-editable packaging can omit its extension configuration. If that defect affects the installed revision, use a separately prepared editable IsaacLab installation; the runtime wheel does not repair IsaacLab. The pinned afca7b09 revision fails with a missing config/extension.toml after a wheel install. Check out the IsaacLab fork and revision named by the runtime dependency (yuechen0614/IsaacLab) into a separate writable directory, then, from the RPent root and in the same environment, run:

uv pip install --no-deps \
  -e /path/to/IsaacLab/source/isaaclab \
  -e /path/to/IsaacLab/source/isaaclab_assets \
  -e /path/to/IsaacLab/source/isaaclab_tasks

Run this step once after installation; retain the checkout while using the environment.

Sources and assets#

The runtime wheel includes the validated RoboDojo code tree, so a RoboDojo checkout is not required. By default, the environment uses robodojo_runtime.source_root(); --source-root can override it with a local checkout. Scene data remains separate from Python dependencies. Set ROBODOJO_ASSETS_ROOT to the directory containing Assets/:

  • RoboDojo robot/object/material/layout assets: download only Assets/** from RoboDojo’s dataset. For example, with the Hugging Face CLI available:

    hf download RoboDojo-Benchmark/RoboDojo --repo-type dataset \
      --include 'Assets/**' --local-dir /data/robodojo
    export ROBODOJO_ASSETS_ROOT=/data/robodojo
    

    The packaged runtime reads ROBODOJO_ASSETS_ROOT/Assets; no symlink into the installed code is needed. Check that Robots, Object, Material and Eval_Layout contain real files, not LFS pointers.

  • NVIDIA USD/material assets referenced by IsaacLab: these are separate from isaaclab_assets. Obtain the matching asset pack from NVIDIA’s asset download instructions. RoboDojo’s utils/ensure_usd_path.py rewrites URLs under Assets/Isaac/5.0: Isaac Sim 5.1 keeps referencing that 5.0 tree, so those references need the 5.0 asset pack, while scenes or extensions that name 5.1 paths need the 5.1 asset pack. The two packs contain different trees and are not interchangeable. For the 5.0 references, preserve /data/nvidia/Assets/Isaac/5.0 and export:

    export ROBODOJO_USD_ASSET_PREFIX=/data/nvidia
    

    This existing upstream variable applies only to that URL prefix, not every IsaacLab asset or a 5.1 asset tree. Other IsaacLab references use Kit’s /persistent/isaac/asset_root/cloud setting; configure that separately for a matching local pack when offline. A 5.1 pack does not cover the 5.0 paths.

The two similarly named dependency entries are not two scene-data bundles: isaacsim[all,extscache] supplies simulator binaries and extension caches (the Linux x86-64 / CPython 3.11 5.1.0 Kit and Kit-SDK cache wheels alone are 3,021,340,845 and 1,345,115,764 bytes). Keep this runtime installation; a data environment variable cannot replace it. isaaclab_assets at afca7b09 contains 119,143 bytes of source/configuration and no USD files, so retain it too. Git tree sizes are about 53.6 MB for IsaacLab and 129.9 MB for cuRobo (including robot meshes); these are uncompressed tree totals, not measured clone sizes. cuRobo’s packaged meshes remain part of its runtime. Thus uv still downloads a large simulator runtime, but not the separate scene datasets or policy checkpoints. Download checkpoints separately and use PI05_CHECKPOINT_PATH and SAM3_CHECKPOINT_PATH as described below.

Policy and perception checkpoints#

RoboDojo releases its Pi_05 weights as an orbax/JAX checkpoint, while the RLinf openpi loader reads the PyTorch build, so convert them once after downloading. Perception uses SAM 3, which needs its own checkpoint.

# 1. Download the released Pi_05 checkpoint (about 7 GB). Use the same
#    --local-dir root as the assets: hf download keeps the repository path,
#    so /data/robodojo/ckpt would nest a second ckpt/ level.
hf download RoboDojo-Benchmark/RoboDojo --repo-type dataset \
  --include 'ckpt/RoboDojo/Pi_05/**' --local-dir /data/robodojo

# 2. Convert it. The release directory is named Pi_05 and the converter picks
#    the pi05 branch by looking for the lowercase string "pi05" in the path,
#    so convert through a link that contains it.
ln -s /data/robodojo/ckpt/RoboDojo/Pi_05/RoboDojo-sim-arx_x5-joint-0 \
  /data/robodojo/pi05_robodojo_arx_x5
python -m rlinf.utils.ckpt_convertor.convert_openpi_jax_to_python \
  --checkpoint-dir /data/robodojo/pi05_robodojo_arx_x5/59999 \
  --config-name pi05_aloha \
  --output-path /data/robodojo/pi05_robodojo_arx_x5_torch

# 3. Point the run at the converted checkpoint and at SAM 3.
export PI05_CHECKPOINT_PATH=/data/robodojo/pi05_robodojo_arx_x5_torch
export SAM3_CHECKPOINT_PATH=/data/sam3/sam3.pt

The converter belongs to RLinf and ships inside the rlinf package that the robodojo-sim extra installs, so call it as a module: an RPent checkout has no rlinf/utils/ tree. Pass the 59999 step directory, because the converter restores <checkpoint-dir>/params. The released weights are a Pi_05 aloha model (32-D actions, horizon 50), which is what the stock pi05_aloha config describes in the openpi registry installed with the extra. XPolicyLab’s bundled openpi carries an ..._arx-x5_seed_0 name for the same architecture, but that registry is not the one the converter imports.

Keep the normalization statistics that ship with the release next to the converted weights; the client loads them together. --inspect_only prints the orbax parameter keys without converting, which is the quickest way to check a download before converting it.

RPent configuration#

The default --policy-backend rlinf requires a policy interpreter that provides the pi05_robodojo_arx_x5 config and its openpi dependencies; RLinf main does not include that config yet. Set PI05_CHECKPOINT_PATH to a compatible RLinf checkpoint and its normalization statistics. The robodojo extra selects an integration branch that includes the required preset. Use --policy-backend xpolicylab for an independently prepared XPolicyLab runtime.

Configure SAM3’s checkpoint using SAM3_CHECKPOINT_PATH, and export the placement settling budget. The default leaves objects unstable in official mode; the variable is read by the RoboDojo checkout, not by RPent, and the CLI passes it on to the child services it starts:

export ROBODOJO_PLACEMENT_SETTLE_STEPS=1000

Start with the packaged code:

rpent --robot robodojo --task put_bottles_into_dustbin --layout 0

XPolicyLab is not included in the runtime wheel. --xpolicylab-root defaults to SOURCE_ROOT/XPolicyLab; set it for a separate checkout. In a single environment, services default to the current interpreter; use --sim-python or --pi05-python to override it when needed. The CLI constructs child import paths without reading a workspace’s config/runtime.env or changing the parent environment. Child processes inherit shell environment variables. Omitting --cuda-device preserves CUDA_VISIBLE_DEVICES, including an unset value; passing it explicitly selects the GPU for locally started services. For XPolicyLab, an omitted GPU flag starts its Python policy entry point with --pi05-python directly; an explicit flag uses its shell launcher.

Use --env-endpoint, --vla-endpoint, and --sam3-endpoint to attach to already running services. A borrowed service requires no local source or Python path for that component. The CLI starts the shared rpent.robots.components.pi05_vla_server --embodiment robodojo by default. With --policy-backend xpolicylab, it starts xpolicylab_vla_server with --policy-root pointing to XPolicyLab/policy/Pi_05 instead. Select the matching backend when borrowing a VLA endpoint. Changing backend does not convert checkpoints; the RLinf client encodes native observations into openpi’s wire format.

Every owned service logs and writes into the run’s output directory: the CLI passes it as --save-dir to the environment server and as --output-dir to the optional XPolicyLab entry point. Its inner policy log is vla_server.log. Direct service launches default to the current directory; use separate output directories for concurrent runs.

Verify the installation#

Run one bounded development episode and confirm the services come up before the planner takes over:

export ROBODOJO_PLACEMENT_SETTLE_STEPS=1000
export OMNI_KIT_ACCEPT_EULA=YES
rpent --robot robodojo --task put_bottles_into_dustbin --layout 0 \
  --planner codex --model <planner-model> --max-turns 1 \
  --sim-python /path/to/sim-env/bin/python \
  --pi05-python /path/to/pi05-env/bin/python \
  --output-dir /path/to/run-output

OMNI_KIT_ACCEPT_EULA=YES answers Kit’s license prompt; without it a non-interactive start stops at that prompt and exits.

Expected behaviour:

  • /path/to/run-output contains robodojo_env_server.log, sam3_server.log, robodojo_vla_server.log and, once the policy server is spawned with XPolicyLab, vla_server.log.

  • The environment server reports ready, and the first observation carries cam_head, cam_left_wrist and cam_right_wrist with intrinsics and extrinsics, plus joint and gripper state.

  • The run writes one MP4 per camera under /path/to/run-output/videos.

  • Shutdown leaves no owned child process behind and the GPUs return to idle.

A service that exits during startup is the usual failure mode; read its log in the run output directory first. Isaac Sim start-up takes tens of seconds and the first run also compiles shaders.

Key modules#

  • robots/robodojo/rlinf_env.py — agent recording, camera metadata and episode diagnostics over RLinf’s RoboDojoEnv.

  • robodojo_runtime/bridge.py — simulator creation, reset, observations and control execution; imported lazily by RLinf.

  • robots/robodojo/env_server.py — Isaac Sim RPC server (main-thread rendering; head + dual-wrist RGB-D with intrinsics/extrinsics; joint/ee actions; per-camera video recording).

  • robots/robodojo/env_client.py — rpent-side client inheriting BaseEnvClient.

  • Shared rpent/robots/components/pi05_vla_server.py and rpent/robots/components/pi05_vla_client.py — default RLinf openpi policy with the robodojo embodiment.

  • Optional rpent/robots/components/xpolicylab_vla_server.py and rpent/robots/components/xpolicylab_vla_client.py — Pi_05 policy service (XPolicyLab WebSocket) adapted to the shared BaseVLAFacade / BaseVLAClient protocol.

  • robots/robodojo/toolkit.py / tools.py — primitives: view_env_state, back_project, segment, move_to, set_gripper, pi0_pick, stabilize, place_in_bin.

  • robots/robodojo/robot_spec.py — RobotSpec factory (CLI, run config, runtime orchestration).

  • robots/robodojo/tasks.py — task inventory from the configured source checkout.

Implementation details#

The environment server initializes Isaac Sim before simulator imports and serializes simulator requests on its main thread for camera rendering. Each process owns one simulator application. Reset returns an observation dictionary; step returns (obs, reward, done, info). Chunk stepping is unsupported, so primitives issue individual steps through the environment action path, retaining its bounds and counters.

The RLinf client maps head, left-wrist and right-wrist RGB to main_images, wrist_images and extra_view_images. states contains left arm (6), right arm (6), left gripper (1), right gripper (1), with observed gripper values unchanged (1=open, 0=closed); task_descriptions carries the instruction.

The XPolicyLab adapter passes observations/actions through unchanged and serializes update_obs/get_action with reset; it does not isolate sessions. RoboDojo requires three-camera inputs and 14-DoF joint actions; neither backend converts checkpoints or joint actions to end-effector actions.

Tools and information access#

The planner supplies tools directly; call their listed names. Dustbin placement and bottle recovery guidance are included only in the put_bottles_into_dustbin task context, not the generic system prompt. Recorded-state reading and calibrated depth projection are local to robots/robodojo/tools.py. Projection uses Isaac’s negative optical Z axis; segmentation uses the shared SAM3 client. view_env_state, back_project and segment are read-only: they do not advance the environment or trigger post-action state capture.

Every tool declared in robots.robodojo.tools is registered for every task; the toolkit does not filter tools by task name. No planner tool exposes reward details or ground-truth safety alarms, so planners judge progress from observations only. The underlying env.get_reward_details and env.get_safety_status RPCs remain available on dev servers for external evaluation and diagnostics, not as planner tools. Reward and official success belong to the evaluation path after the planner’s action channel is closed. finalize_run records the runner-provided result; it does not call these RPCs itself.

Tool visibility is not an isolation mechanism: the schemas, handlers, automatic post-action state, raw observation fields, logs, memory, and common file tools that the shared toolkit exposes are unchanged. Only --planner flash (eval-fair) narrows the robot tools down to the replay set.

Development and frozen replay#

Normal planners retain the development tool set without scoring or safety diagnostics. --planner flash selects eval-fair: RPent’s native Flash planner invokes RobotSpec.run_flash without an LLM. --explore remains unsupported for RoboDojo; development here means the normal planner loop, not that CLI mode.

Development records actions and observations in the shared states.json manifest. Read-only perception results are attached to the next action for Flash export. Existing flash_trace.json exports remain readable. To record a transferable waypoint, call segment on cam_head, then back_project at its returned centroid_rc (row, col), then the action. The pixel is the floored mean of foreground mask coordinates, not the box center. Recording and replay share this derivation; recorded mask area and coordinate sums allow export to reject a changed centroid. Do not batch different objects’ segmentations ahead of depth calls. Repeat that perception pair before each action. Export the trace before evaluation:

python -m robots.robodojo.flash.generate \
  --trace /path/to/dev-run/states.json \
  --task put_bottles_into_dustbin \
  --destination /path/to/memory/robodojo/flash

The version-2 JSON contains a task, symbolic SAM3 queries, the explicit sam3_mask_centroid_floor_v1 derivation method, optional wrist-camera refinement, and ordered actions with arguments and three-dimensional offsets. It contains no reference object coordinates, score or predicate results. Export accepts move_to, set_gripper and pi0_pick; unsupported actions (including place_in_bin and stabilize), failed calls and unanchored moves are rejected. Record those operations as supported basic actions instead. Existing plans are never overwritten. Review and freeze the plan before eval; no candidate selection or plan writing occurs during replay. Version-1 box-center plans and traces without mask moments must be recorded again; they are not silently converted. The shared XPolicyLab facade owns only the policy processes it spawns and stops them on close, startup failure, SIGTERM and normal interpreter exit; borrowed policy services remain running.

Use the usual runtime flags together with:

rpent --robot robodojo --task put_bottles_into_dustbin --layout 1 \
  --planner flash --memory-profile local --memory-dir /path/to/memory/robodojo \
  --sim-python /path/to/sim-env/bin/python \
  --pi05-python /path/to/pi05-env/bin/python

The plan is flash/<task>_plan.json under the selected robot memory. HF mode uses the native robodojo/flash/** sync filter; no RoboDojo plans are bundled or guaranteed to be published there. Missing or invalid plans fail closed.

Replay first grounds all anchors in the opening head frame. Each waypoint is the live anchor plus its recorded offset (maximum offset length 0.5 m). Export with --refine-camera cam_left_wrist or cam_right_wrist to require a live wrist reading before moves; it must agree within 5 cm, otherwise replay stops. This does not move the wrist to create visibility: record an approach that makes the selected view useful, or leave refinement disabled. pi0_pick is unchanged; acceptance additionally requires a closed gripper and wrist-localized object within 12 cm of either EEF. A failed hold retries the preceding approach with fresh head grounding, at most three pick attempts. Errors, lost localization, unreachable waypoints and step-budget exhaustion stop replay without reset. These checks are not a collision-safety guarantee.

Eval-fair additionally excludes common file/memory tools and task-specific helpers. The service advertises the mode in its metadata, rejects reset and diagnostic RPCs, returns zero reward and no task-success feedback, and exposes only RGB-D, camera calibration, instruction and arm proprioception. Ground-truth bottle alarms are disabled. Existing dev endpoints are rejected for eval-fair; boot creates the episode and the eval client does not reset it again. Automatic state logs therefore contain public observations and action diagnostics only. The reduced eval prompts contain no scoring instructions and are unused by Flash.

Flash done and its planner completion status mean the frozen sequence completed, not that the official task predicate passed. Official scoring must remain outside the replay context. GPU, real policy and simulator validation are required before claiming benchmark compatibility or success.

Task language#

Task-language RPC and public observations use RoboDojo’s initialized description manager. Empty language or unresolved template markers raise an error. The official instruction stays public in eval-fair. Omitting pi0_pick.prompt uses this resolved official language. Explicit overrides remain supported for contact segments but must identify the intended object; unresolved markers are rejected before inference. Official multi-object tasks do not specify a grasp order, and descriptive overrides do not guarantee that a checkpoint can select arbitrary instances. Verify the actual held target.

Capability scope and limitations#

RoboDojo provides dual-arm motion and gripper primitives, three-camera RGB-D, SAM3 perception, RLinf Pi0.5 (or optional XPolicyLab) and frozen Flash replay. Task names come from the configured checkout; examples include put_bottles_into_dustbin, fill_pen_holder and stack_bowls_random, not a validated success suite. place_in_bin is registered only for put_bottles_into_dustbin. Handover is not implemented. Low-Z tabletop and lateral scripted IK motions have reachability limits: inspect reached and dist_to_target rather than assuming the commanded pose was achieved. See Flash Mode for the shared evaluation-only planner; replay executes actions, it is not a read-only robot operation.

Bounded smoke runs and shutdown diagnostics#

For fill_pen_holder, use an explicit smoke budget of --planner-timeout-s 1500 --max-turns 40 with an outer timeout --signal=INT --kill-after=20s 1700s. These are validation overrides, not regular defaults. The outer limit leaves time for startup and cleanup and keeps the run below 30 minutes. Stop after a timeout or failed gate; inspect the last completed tools and provider latency before scheduling another attempt.

The CLI records planner errors in transcript_<cell>.json (error). The run result is written by the robot’s own finalizer — RoboDojo writes result.json into the run output directory — and a finalizer that raises is logged and reflected in the process exit code. Read both in dev and Flash runs; a zero process status is not proof of official task success.

Verify all three videos by full decoding and check every owned server’s exit status. Between [robodojo-env] shutdown begin and process exit, require no [Error], traceback, or Fatal Python error. Headless GLFW warnings are expected noise, not a reason to ignore shutdown errors. The environment releases writers, camera annotators/render products, and syntheticdata graph handles before stopping Replicator and closing the stage/app on the main thread.