RoboDojo#
RoboDojo connects Isaac Sim / IsaacLab, dual ARX-X5 arms and an RLinf Pi0.5 policy to RPent’s planners, tools and memory. The integration is experimental; simulator compatibility and task success require GPU validation. Follow the upstream RoboDojo and XPolicyLab instructions for simulator, CUDA and policy dependencies, and download the required assets and checkpoints. For shared backend interfaces, see Add a Robot or Simulator.
Python environment#
Install the runtime and agent dependencies from the RPent root in Python 3.11.
The robodojo-sim extra installs rlinf-robodojo-runtime and the
RLinf environment adapter; the runtime owns the Isaac Sim / IsaacLab pins.
robodojo also installs SAM3 and the openpi policy runtime.
Use one robot extra per environment. Run uv from this directory so it reads the
root project’s cuRobo build dependencies, and pass this repository’s simulator
overrides explicitly:
uv pip install -e ".[robodojo]" --extra-index-url https://pypi.nvidia.com \
--override requirements/robodojo-override.txt
On Blackwell GPUs, append --torch-backend=cu128 to select a PyTorch
build with sm_120 support. The default CUDA 12.6 build of PyTorch 2.7.0
does not support these GPUs. Use the same option when reinstalling this extra.
Note that uv treats an already-installed 2.7.0+cu126 as satisfying the
pinned torch==2.7.0 and keeps it, so switching an existing environment also
needs --reinstall-package torch --reinstall-package torchvision; rebuilding
the environment is the simpler alternative.
requirements/robodojo-override.txt carries the eight pins where the simulator
stack disagrees with the agent stack or with rpent-openpi; passing them with
--override keeps every other robot resolving the versions it was validated
against. They allow dependency resolution; they do not establish simulation task
success. The RLinf integration branch supplies both the environment
adapter and pi05_robodojo_arx_x5. The RoboDojo preset selects the OpenPI
eval loader with a 50-step horizon, a padded 32-dimensional model action,
and 14-dimensional environment actions. Real-weight RPC inference has returned
finite actions of shape (1, 50, 14); this is not a task-success guarantee.
The simulation bridge requires runtime version 0.3.0 or later.
The Git references are temporary until versioned releases are published.
Fresh dependency installation and GPU rollout must be validated separately;
offline bridge tests do not establish simulator compatibility.
IsaacLab’s upstream non-editable packaging can omit its extension configuration.
If that defect affects the installed revision, use a separately prepared
editable IsaacLab installation; the runtime wheel does not repair IsaacLab.
The pinned afca7b09 revision fails with a missing
config/extension.toml after a wheel install. Check out the IsaacLab fork
and revision named by the runtime dependency
(yuechen0614/IsaacLab) into a
separate writable directory, then, from the RPent root and in the same
environment, run:
uv pip install --no-deps \
-e /path/to/IsaacLab/source/isaaclab \
-e /path/to/IsaacLab/source/isaaclab_assets \
-e /path/to/IsaacLab/source/isaaclab_tasks
Run this step once after installation; retain the checkout while using the environment.
Sources and assets#
The runtime wheel includes the validated RoboDojo code tree, so a RoboDojo
checkout is not required. By default, the environment uses
robodojo_runtime.source_root(); --source-root can override it
with a local checkout. Scene data remains separate from Python dependencies.
Set ROBODOJO_ASSETS_ROOT to the directory containing Assets/:
RoboDojo robot/object/material/layout assets: download only
Assets/**from RoboDojo’s dataset. For example, with the Hugging Face CLI available:hf download RoboDojo-Benchmark/RoboDojo --repo-type dataset \ --include 'Assets/**' --local-dir /data/robodojo export ROBODOJO_ASSETS_ROOT=/data/robodojo
The packaged runtime reads
ROBODOJO_ASSETS_ROOT/Assets; no symlink into the installed code is needed. Check thatRobots,Object,MaterialandEval_Layoutcontain real files, not LFS pointers.NVIDIA USD/material assets referenced by IsaacLab: these are separate from
isaaclab_assets. Obtain the matching asset pack from NVIDIA’s asset download instructions. RoboDojo’sutils/ensure_usd_path.pyrewrites URLs underAssets/Isaac/5.0: Isaac Sim 5.1 keeps referencing that 5.0 tree, so those references need the 5.0 asset pack, while scenes or extensions that name 5.1 paths need the 5.1 asset pack. The two packs contain different trees and are not interchangeable. For the 5.0 references, preserve/data/nvidia/Assets/Isaac/5.0and export:export ROBODOJO_USD_ASSET_PREFIX=/data/nvidia
This existing upstream variable applies only to that URL prefix, not every IsaacLab asset or a 5.1 asset tree. Other IsaacLab references use Kit’s
/persistent/isaac/asset_root/cloudsetting; configure that separately for a matching local pack when offline. A 5.1 pack does not cover the 5.0 paths.
The two similarly named dependency entries are not two scene-data bundles:
isaacsim[all,extscache] supplies simulator binaries and extension caches
(the Linux x86-64 / CPython 3.11 5.1.0 Kit and Kit-SDK cache wheels alone are
3,021,340,845 and 1,345,115,764 bytes). Keep this runtime installation; a data
environment variable cannot replace it. isaaclab_assets at afca7b09
contains 119,143 bytes of source/configuration and no USD files, so retain it
too. Git tree sizes are about 53.6 MB for IsaacLab and 129.9 MB for cuRobo
(including robot meshes); these are uncompressed tree totals, not measured
clone sizes. cuRobo’s packaged meshes remain part of its runtime.
Thus uv still downloads a large simulator runtime, but not the separate scene
datasets or policy checkpoints. Download checkpoints separately and use
PI05_CHECKPOINT_PATH and SAM3_CHECKPOINT_PATH as described below.
Policy and perception checkpoints#
RoboDojo releases its Pi_05 weights as an orbax/JAX checkpoint, while the RLinf openpi loader reads the PyTorch build, so convert them once after downloading. Perception uses SAM 3, which needs its own checkpoint.
# 1. Download the released Pi_05 checkpoint (about 7 GB). Use the same
# --local-dir root as the assets: hf download keeps the repository path,
# so /data/robodojo/ckpt would nest a second ckpt/ level.
hf download RoboDojo-Benchmark/RoboDojo --repo-type dataset \
--include 'ckpt/RoboDojo/Pi_05/**' --local-dir /data/robodojo
# 2. Convert it. The release directory is named Pi_05 and the converter picks
# the pi05 branch by looking for the lowercase string "pi05" in the path,
# so convert through a link that contains it.
ln -s /data/robodojo/ckpt/RoboDojo/Pi_05/RoboDojo-sim-arx_x5-joint-0 \
/data/robodojo/pi05_robodojo_arx_x5
python -m rlinf.utils.ckpt_convertor.convert_openpi_jax_to_python \
--checkpoint-dir /data/robodojo/pi05_robodojo_arx_x5/59999 \
--config-name pi05_aloha \
--output-path /data/robodojo/pi05_robodojo_arx_x5_torch
# 3. Point the run at the converted checkpoint and at SAM 3.
export PI05_CHECKPOINT_PATH=/data/robodojo/pi05_robodojo_arx_x5_torch
export SAM3_CHECKPOINT_PATH=/data/sam3/sam3.pt
The converter belongs to RLinf and ships inside the rlinf package that the
robodojo-sim extra installs, so call it as a module: an RPent checkout has
no rlinf/utils/ tree. Pass the 59999 step directory, because the
converter restores <checkpoint-dir>/params. The released weights are a
Pi_05 aloha model (32-D actions, horizon 50), which is what the stock
pi05_aloha config describes in the openpi registry installed with the
extra. XPolicyLab’s bundled openpi carries an ..._arx-x5_seed_0 name for the
same architecture, but that registry is not the one the converter imports.
Keep the normalization statistics that ship with the release next to the
converted weights; the client loads them together. --inspect_only prints the
orbax parameter keys without converting, which is the quickest way to check a
download before converting it.
RPent configuration#
The default --policy-backend rlinf requires a policy interpreter that
provides the pi05_robodojo_arx_x5 config and its openpi dependencies;
RLinf main does not include that config yet. Set PI05_CHECKPOINT_PATH to a
compatible RLinf checkpoint and its normalization statistics. The robodojo
extra selects an integration branch that includes the required preset.
Use --policy-backend xpolicylab for an independently prepared XPolicyLab runtime.
Configure SAM3’s checkpoint using SAM3_CHECKPOINT_PATH, and export the
placement settling budget. The default leaves objects unstable in official
mode; the variable is read by the RoboDojo checkout, not by RPent, and the CLI
passes it on to the child services it starts:
export ROBODOJO_PLACEMENT_SETTLE_STEPS=1000
Start with the packaged code:
rpent --robot robodojo --task put_bottles_into_dustbin --layout 0
XPolicyLab is not included in the runtime wheel. --xpolicylab-root defaults
to SOURCE_ROOT/XPolicyLab; set it for
a separate checkout. In a single environment, services default to the current
interpreter; use --sim-python or --pi05-python to override it when needed.
The CLI constructs
child import paths without reading a workspace’s config/runtime.env or
changing the parent environment. Child processes inherit shell environment
variables. Omitting
--cuda-device preserves CUDA_VISIBLE_DEVICES, including an unset value;
passing it explicitly selects the GPU for locally started services.
For XPolicyLab, an omitted GPU flag starts its Python policy entry point with
--pi05-python directly; an explicit flag uses its shell launcher.
Use --env-endpoint, --vla-endpoint, and --sam3-endpoint to attach
to already running services. A borrowed service requires no local source or
Python path for that component. The CLI starts the shared
rpent.robots.components.pi05_vla_server --embodiment robodojo by default.
With --policy-backend xpolicylab, it starts xpolicylab_vla_server with
--policy-root pointing to XPolicyLab/policy/Pi_05 instead.
Select the matching backend when borrowing a VLA endpoint. Changing backend
does not convert checkpoints; the RLinf client encodes native observations
into openpi’s wire format.
Every owned service logs and writes into the run’s output directory: the CLI
passes it as --save-dir to the environment server and as --output-dir
to the optional XPolicyLab entry point. Its inner policy log is vla_server.log.
Direct service launches default to the current directory; use separate output
directories for concurrent runs.
Verify the installation#
Run one bounded development episode and confirm the services come up before the planner takes over:
export ROBODOJO_PLACEMENT_SETTLE_STEPS=1000
export OMNI_KIT_ACCEPT_EULA=YES
rpent --robot robodojo --task put_bottles_into_dustbin --layout 0 \
--planner codex --model <planner-model> --max-turns 1 \
--sim-python /path/to/sim-env/bin/python \
--pi05-python /path/to/pi05-env/bin/python \
--output-dir /path/to/run-output
OMNI_KIT_ACCEPT_EULA=YES answers Kit’s license prompt; without it a
non-interactive start stops at that prompt and exits.
Expected behaviour:
/path/to/run-outputcontainsrobodojo_env_server.log,sam3_server.log,robodojo_vla_server.logand, once the policy server is spawned with XPolicyLab,vla_server.log.The environment server reports ready, and the first observation carries
cam_head,cam_left_wristandcam_right_wristwith intrinsics and extrinsics, plus joint and gripper state.The run writes one MP4 per camera under
/path/to/run-output/videos.Shutdown leaves no owned child process behind and the GPUs return to idle.
A service that exits during startup is the usual failure mode; read its log in the run output directory first. Isaac Sim start-up takes tens of seconds and the first run also compiles shaders.
Key modules#
robots/robodojo/rlinf_env.py— agent recording, camera metadata and episode diagnostics over RLinf’sRoboDojoEnv.robodojo_runtime/bridge.py— simulator creation, reset, observations and control execution; imported lazily by RLinf.robots/robodojo/env_server.py— Isaac Sim RPC server (main-thread rendering; head + dual-wrist RGB-D with intrinsics/extrinsics; joint/ee actions; per-camera video recording).robots/robodojo/env_client.py— rpent-side client inheritingBaseEnvClient.Shared
rpent/robots/components/pi05_vla_server.pyandrpent/robots/components/pi05_vla_client.py— default RLinf openpi policy with therobodojoembodiment.Optional
rpent/robots/components/xpolicylab_vla_server.pyandrpent/robots/components/xpolicylab_vla_client.py— Pi_05 policy service (XPolicyLab WebSocket) adapted to the sharedBaseVLAFacade/BaseVLAClientprotocol.robots/robodojo/toolkit.py/tools.py— primitives:view_env_state,back_project,segment,move_to,set_gripper,pi0_pick,stabilize,place_in_bin.robots/robodojo/robot_spec.py—RobotSpecfactory (CLI, run config, runtime orchestration).robots/robodojo/tasks.py— task inventory from the configured source checkout.
Implementation details#
The environment server initializes Isaac Sim before simulator imports and
serializes simulator requests on its main thread for camera rendering. Each
process owns one simulator application. Reset returns an observation dictionary;
step returns (obs, reward, done, info). Chunk stepping is unsupported, so
primitives issue individual steps through the environment action path, retaining
its bounds and counters.
The RLinf client maps head, left-wrist and right-wrist RGB to main_images,
wrist_images and extra_view_images. states contains left arm (6),
right arm (6), left gripper (1), right gripper (1), with observed gripper
values unchanged (1=open, 0=closed); task_descriptions carries the instruction.
The XPolicyLab adapter passes observations/actions through unchanged and
serializes update_obs/get_action with reset; it does not isolate
sessions. RoboDojo requires three-camera inputs and 14-DoF joint actions;
neither backend converts checkpoints or joint actions to end-effector actions.
Tools and information access#
The planner supplies tools directly; call their listed names. Dustbin placement
and bottle recovery guidance are included only in the
put_bottles_into_dustbin task context, not the generic system prompt.
Recorded-state reading and calibrated depth projection are local to
robots/robodojo/tools.py. Projection uses Isaac’s negative optical Z axis;
segmentation uses the shared SAM3 client. view_env_state, back_project and
segment are read-only:
they do not advance the environment or trigger post-action state capture.
Every tool declared in robots.robodojo.tools is registered for every task;
the toolkit does not filter tools by task name.
No planner tool exposes reward details or ground-truth safety alarms, so
planners judge progress from observations only. The underlying
env.get_reward_details and env.get_safety_status RPCs remain available
on dev servers for external evaluation and diagnostics, not as planner tools.
Reward and official success belong to the evaluation path after the planner’s
action channel is closed. finalize_run records the runner-provided result;
it does not call these RPCs itself.
Tool visibility is not an isolation mechanism: the schemas, handlers,
automatic post-action state, raw observation fields, logs, memory, and common
file tools that the shared toolkit exposes are unchanged. Only
--planner flash (eval-fair) narrows the robot tools down to the replay set.
Development and frozen replay#
Normal planners retain the development tool set without scoring or safety diagnostics.
--planner flash selects eval-fair: RPent’s native Flash planner invokes
RobotSpec.run_flash without an LLM. --explore remains unsupported for
RoboDojo; development here means the normal planner loop, not that CLI mode.
Development records actions and observations in the shared states.json
manifest. Read-only perception results are attached to the next action for
Flash export. Existing flash_trace.json exports remain readable. To record
a transferable waypoint, call segment on cam_head, then back_project
at its returned centroid_rc (row, col), then the action. The pixel is the
floored mean of foreground mask coordinates, not the box center. Recording
and replay share this derivation; recorded mask area and coordinate sums
allow export to reject a changed centroid. Do not batch different objects’
segmentations ahead of depth calls. Repeat that perception pair before
each action. Export the trace before evaluation:
python -m robots.robodojo.flash.generate \
--trace /path/to/dev-run/states.json \
--task put_bottles_into_dustbin \
--destination /path/to/memory/robodojo/flash
The version-2 JSON contains a task, symbolic SAM3 queries, the explicit
sam3_mask_centroid_floor_v1 derivation method, optional wrist-camera
refinement, and ordered actions with arguments and three-dimensional offsets.
It contains no reference object coordinates, score or predicate results.
Export accepts move_to, set_gripper and pi0_pick; unsupported actions
(including place_in_bin and stabilize), failed calls and unanchored moves
are rejected. Record those operations as supported basic actions instead.
Existing plans are never overwritten. Review and freeze the plan before eval;
no candidate selection or plan writing occurs during replay.
Version-1 box-center plans and traces without mask moments must be recorded
again; they are not silently converted. The shared XPolicyLab facade owns only
the policy processes it spawns and stops them on close, startup failure,
SIGTERM and normal interpreter exit; borrowed policy services remain running.
Use the usual runtime flags together with:
rpent --robot robodojo --task put_bottles_into_dustbin --layout 1 \
--planner flash --memory-profile local --memory-dir /path/to/memory/robodojo \
--sim-python /path/to/sim-env/bin/python \
--pi05-python /path/to/pi05-env/bin/python
The plan is flash/<task>_plan.json under the selected robot memory. HF mode
uses the native robodojo/flash/** sync filter; no RoboDojo plans are bundled
or guaranteed to be published there. Missing or invalid plans fail closed.
Replay first grounds all anchors in the opening head frame. Each waypoint is
the live anchor plus its recorded offset (maximum offset length 0.5 m).
Export with --refine-camera cam_left_wrist or cam_right_wrist to require
a live wrist reading before moves; it must agree within 5 cm, otherwise replay
stops. This does not move the wrist to create visibility: record an approach
that makes the selected view useful, or leave refinement disabled.
pi0_pick is unchanged; acceptance additionally requires a closed gripper
and wrist-localized object within 12 cm of either EEF. A failed hold retries
the preceding approach with fresh head grounding, at most three pick attempts.
Errors, lost localization, unreachable waypoints and step-budget exhaustion
stop replay without reset. These checks are not a collision-safety guarantee.
Eval-fair additionally excludes common file/memory tools and task-specific helpers. The service advertises the mode in its metadata, rejects reset and diagnostic RPCs, returns zero reward and no task-success feedback, and exposes only RGB-D, camera calibration, instruction and arm proprioception. Ground-truth bottle alarms are disabled. Existing dev endpoints are rejected for eval-fair; boot creates the episode and the eval client does not reset it again. Automatic state logs therefore contain public observations and action diagnostics only. The reduced eval prompts contain no scoring instructions and are unused by Flash.
Flash done and its planner completion status mean the frozen sequence
completed, not that the official task predicate passed. Official scoring must
remain outside the replay context. GPU, real policy and simulator validation
are required before claiming benchmark compatibility or success.
Task language#
Task-language RPC and public observations use RoboDojo’s initialized description
manager. Empty language or unresolved template markers raise an error. The
official instruction stays public in eval-fair.
Omitting pi0_pick.prompt uses this resolved official language. Explicit
overrides remain supported for contact segments but must identify the intended
object; unresolved markers are rejected before inference. Official multi-object
tasks do not specify a grasp order, and descriptive overrides do not guarantee
that a checkpoint can select arbitrary instances. Verify the actual held target.
Capability scope and limitations#
RoboDojo provides dual-arm motion and gripper primitives, three-camera RGB-D,
SAM3 perception, RLinf Pi0.5 (or optional XPolicyLab) and frozen Flash replay. Task names come from
the configured checkout; examples include put_bottles_into_dustbin,
fill_pen_holder and stack_bowls_random, not a validated success suite.
place_in_bin is registered only for put_bottles_into_dustbin.
Handover is not implemented. Low-Z tabletop and lateral scripted IK motions
have reachability limits: inspect reached and dist_to_target rather
than assuming the commanded pose was achieved. See Flash Mode for the
shared evaluation-only planner; replay executes actions, it is not a
read-only robot operation.
Bounded smoke runs and shutdown diagnostics#
For fill_pen_holder, use an explicit smoke budget of
--planner-timeout-s 1500 --max-turns 40 with an outer
timeout --signal=INT --kill-after=20s 1700s. These are validation overrides,
not regular defaults. The outer limit leaves time for startup and cleanup and
keeps the run below 30 minutes. Stop after a timeout or failed gate; inspect the
last completed tools and provider latency before scheduling another attempt.
The CLI records planner errors in transcript_<cell>.json (error). The
run result is written by the robot’s own finalizer — RoboDojo writes
result.json into the run output directory — and a finalizer that raises is
logged and reflected in the process exit code. Read both in dev and Flash runs;
a zero process status is not proof of official task success.
Verify all three videos by full decoding and check every owned server’s exit
status. Between [robodojo-env] shutdown begin and process exit, require no
[Error], traceback, or Fatal Python error. Headless GLFW warnings are
expected noise, not a reason to ignore shutdown errors. The environment
releases writers, camera annotators/render products, and syntheticdata graph
handles before stopping Replicator and closing the stage/app on the main thread.