Remote Training Infrastructure (Nebius Serverless)
This skill runs the Positronic convert → train → serve pipeline on Nebius Serverless: Jobs for batch work (dataset conversion, training) and Endpoints for HTTP inference servers. Compute is provisioned per job/endpoint and released automatically when it finishes, so there is no idle compute cost.
All operations go through the wrapper scripts in workflows/nebius/. Run them from the repo root. The full reference is workflows/nebius/README.md.
Vendor tokens
Every script takes a <vendor> positional that selects the container image and uv extra. Supported: lerobot03_3 (ACT), lerobot (SmolVLA), openpi, gr00t.
| Vendor |
Image |
Train/serve hardware |
lerobot03_3 (ACT) |
positro/positronic:latest |
H100 |
lerobot (SmolVLA) |
positro/positronic:latest |
H100 |
openpi |
positro/openpi:latest |
H100 |
gr00t |
positro/gr00t:latest |
H100 |
Conversion always runs on CPU (cpu-e2, 8vcpu-32gb); train/serve on gpu-h100-sxm (1gpu-16vcpu-200gb). openpi/gr00t re-use the lerobot03_3 converter with their own codec namespace.
S3 Convention
s3://interim/{dataset}/{vendor}/{codec}/ — converted LeRobot datasets
s3://checkpoints/{dataset}/{vendor}/{codec_or_experiment}/ — training output
s3://inference/{dataset}/{date_or_exp}/{vendor}/ — inference eval results
Every run writes a runmetadata.yaml into its output directory capturing the full CLI command plus a snapshot of the code state (.py/*.toml). See [Analysing a past run](#analysing-a-past-run) to reconstruct what produced a given checkpoint.
Where current paths live
Concrete dataset/checkpoint S3 paths are intentionally not listed here — they rotate as new runs land and any list goes stale. The source of truth is the config in the codebase:
- Datasets:
positronic/cfg/phail/ and positronic/cfg/ds/ (e.g. @positronic.cfg.ds.sim.simstackcubes,
@positronic.cfg.phail.v10.teleopunified).
- Checkpoints: the per-vendor server presets in
positronic/vendors/<vendor>/server.py — the named configs (e.g. phail, simstack) set checkpointsdir= to the path currently in use. Read those for the live values rather than relying on a checkpoint path memorized anywhere.
To discover what physically exists, aws s3 ls under the convention above; to learn what produced a given checkpoint, read its runmetadata*.yaml (see [Analysing a past run](#analysing-a-past-run)).
Docker Images
Serverless jobs/endpoints pull positro/<vendor>:${NEBIUSIMAGETAG:-latest} from the registry — they do not mount local source, so any code change needs a rebuild + push before it takes effect remotely.
Tag gotcha: locally make push-* pushes :<branch> and :<sha> but not :latest (that only happens under CI). So after a code change, either:
cd docker && CI=1 make push-<x> — updates :latest (what the workflow pulls by default), or
cd docker && make push-<x> IMAGE_TAG=<branch> then run the workflow with
NEBIUSIMAGETAG=<branch> — tests a branch build without clobbering :latest.
Plain make push-<x> with no CI/IMAGETAG/NEBIUSIMAGE_TAG leaves serverless running the old :latest image — the change silently won't take effect.
| Image |
Source |
Used For |
positro/positronic |
positronic/docker/ |
Conversion, lerobot / SmolVLA train+serve |
positro/openpi |
positronic/docker/ (depends on positro/openpi-base) |
OpenPI train+serve, openpi stats |
positro/gr00t |
positronic/docker/ (depends on positro/gr00t-base) |
GR00T train+serve |
cd docker
make push-training # positro/positronic
make push-openpi # positro/openpi (rebuild positro/openpi-base first if ../openpi changed)
make push-groot # positro/gr00t (rebuild positro/gr00t-base first if ../gr00t changed)
make push # all images
Cross-repo base rebuilds: cd ../openpi/docker && make push (or ../gr00t/docker), then cd ../positronic/docker && make push-openpi. See docker/CONTEXTS.md.
One-time setup
The pipeline reads credentials from Nebius MysteryBox secrets and uses a shared filesystem for uv/HF/openpi caches. This is already provisioned for the Positronic-internal project. To (re)create it for a different project, follow "One-time setup" in workflows/nebius/README.md (five MysteryBox secrets + one network_ssd filesystem).
Defaults point at the Positronic-internal project; override via env when needed:
| Variable |
Default |
Purpose |
NEBIUSPARENTID |
project-e00f38wexevrr52b8j |
Project to create jobs/endpoints in |
NEBIUSSUBNETID |
vpcsubnet-e00pk1j1x6hjmr4m92 |
VPC subnet |
WANDB_SECRET |
positronic-serverless-wandb-api-key |
MysteryBox name for WandB key. Set empty to disable wandb. |
NEBIUSAUTHTOKEN_SECRET |
positronic-serverless-inference-token |
MysteryBox name (payload key AUTH_TOKEN) for the token gating served endpoints. No open-endpoint mode. |
NEBIUSCACHEFS |
computefilesystem-e00f6jyfr5wkawyrab |
Shared cache filesystem ID (mounted RW at /cache) |
Pipeline
1. Convert Dataset
convert.sh runs the right converter + codec for the vendor as a CPU Job. For openpi it blocks until convert finishes, then chains a stats job and prints the --stats_path to use for training.
bash workflows/nebius/convert.sh lerobot_0_3_3 \
[email protected]_stack_cubes \
[email protected]_0_3_3.codecs.ee \
--output_dir=s3://interim/sim_stack/lerobot/ee/
bash workflows/nebius/convert.sh openpi \
[email protected]_stack_cubes \
[email protected] \
--output_dir=s3://interim/sim_stack/openpi/ee/
# → also submits openpi-stats-* ; note the printed --stats_path=<...>/stats/assets/
Default codecs: gr00t eerot6d, lerobot ee, openpi ee, lerobot033 ee.
2. Train
train.sh runs python -m positronic.vendors.<vendor>.train as an H100 Job. The dataset bucket is mounted read-only via Mountpoint-S3 at /mnt/input for lerobot033; other vendors stream via pos3 from the s3:// path directly. --outputdir / --output_path stays an s3:// URL (handled by pos3).
Read each vendor's positronic/vendors/<vendor>/train.py docstring for its exact flags. --resume=true resumes an interrupted run.
# ACT / SmolVLA
bash workflows/nebius/train.sh lerobot_0_3_3 \
--input_path=s3://interim/sim_stack/lerobot/ee/ \
--exp_name=act_sim_stack_v1 \
--output_dir=s3://checkpoints/sim_stack/lerobot/ \
--num_train_steps=50000 --save_freq=10000
# OpenPI — needs --stats_path from the convert step's chained stats job
bash workflows/nebius/train.sh openpi \
--input_path=s3://interim/sim_stack/openpi/ee/ \
--stats_path=s3://interim/sim_stack/openpi/stats/assets/ \
--output_path=s3://checkpoints/sim_stack/openpi/ \
--exp_name=pi_sim_stack_v1 \
--num_train_steps=30000
# openpi checkpoint lands at <output_path>/pi05_positronic_lowmem/<exp_name>/
The first job after a dependency change pays the full uv/HF cold-download (~10 min); later jobs reuse /cache and start faster.
3. Serve a Checkpoint
serve.sh <vendor> <unique-endpoint-name> [server args...] creates an Endpoint on H100 port 8000, blocks until Nebius allocates its managed https:// URL, and prints a banner containing Endpoint URL: https://<managed-url>, the endpoint ID/name, and the teardown command. The container then takes ~10–15 min more to uv sync and load the model. There is no public IP and no open mode: the server rejects anything without Authorization: Bearer $AUTH_TOKEN.
# Named preset — checkpoint path comes from the vendor's server.py config
# (e.g. `demo`, `sim_stack`, `phail`). These are the source of truth; prefer them.
bash workflows/nebius/serve.sh lerobot_0_3_3 my-act-demo demo
# Explicit checkpoint dir (the subcommand is the pipeline name, one per vendor codec).
# Get <ckpt-dir> from the vendor's server.py preset or `aws s3 ls` under the S3
# convention — not memorized.
bash workflows/nebius/serve.sh lerobot_0_3_3 act-server ee \
--pipeline.source.checkpoints_dir=<ckpt-dir>
# openpi's ee pipeline also needs the EE frame the checkpoint speaks; None means the rig's default.
bash workflows/nebius/serve.sh openpi pi-server ee \
--pipeline.source.checkpoints_dir=<ckpt-dir> \
--pipeline.ee_frame=None
bash workflows/nebius/serve.sh gr00t groot-server ee_rot6d \
--pipeline.source.checkpoints_dir=<ckpt-dir>
Load the token once per shell, then sanity-check the endpoint once warm:
source workflows/nebius/common.sh # the secret serve.sh injected, whatever NEBIUS_AUTH_TOKEN_SECRET selects
SECRET_ID=$(nebius mysterybox secret get-by-name --parent-id "$PARENT_ID" \
--name "$AUTH_TOKEN_SECRET" --format json | jq -r '.metadata.id')
export AUTH_TOKEN=$(nebius mysterybox payload get-by-key \
--secret-id "$SECRET_ID" --key "$AUTH_TOKEN_KEY" --format json | jq -r '.data.string_value')
curl -H "Authorization: Bearer $AUTH_TOKEN" \
https://<managed-url>/api/v1/models # → {"models": ["<step>"]}
Tear down (releases compute, retires the managed URL):
bash workflows/nebius/stop.sh my-act-demo
To pause and keep the URL: nebius ai endpoint stop <id> (start resumes).
4. Run Inference Client
serve.sh prints the managed URL in its banner. To read it again later (by endpoint name):
nebius ai endpoint list --parent-id "$NEBIUS_PARENT_ID" --format json \
| jq -r --arg n "<endpoint-name>" \
'.items[] | select(.metadata.name==$n)
| .status.public_endpoints[] | select(startswith("https://"))'
Point positronic eval run at it; .authedremote sends AUTHTOKEN as the bearer token and fails fast if the variable is unset:
uv run --locked positronic eval run --eval=.sim.positronic.stack_cubes \
--policy=.authed_remote --policy.url=https://<managed-url> \
--output_dir=s3://inference/sim_stack_validation/<run_name>/<vendor>/
View results locally (top-level dir compares multiple runs):
uv run --locked python -m positronic.cfg.analysis sim \
--dataset.base.path=s3://inference/sim_stack_validation/<run_name> --reset_cache --https
# http://localhost:5001
Serving locally (desktop / notebook)
Nebius Serverless is for H100-class work. For local inference on a consumer GPU (LeRobot/ACT/SmolVLA, or GR00T inference), serve via Docker contexts instead — no Nebius, no per-hour compute cost. The contexts and services still live in the repo; docker/CONTEXTS.md + docker/docker-compose.yml are the source of truth for which machine/GPU/service to use (desktop = RTX 3060 12GB, notebook = RTX 4060 8GB). OpenPI/DreamZero and GR00T training still need H100 (use the Nebius pipeline above).
Run from docker/. Set CACHEROOT=/home/<user> when targeting a remote context from a Mac (the ${HOME} volume path differs). --service-ports exposes the WebSocket API on port 8000. Servers take a subcommand: a pipeline name (ee, eerot6d, …) with a custom --pipeline.source.checkpointsdir, or a named preset (phail, simstack, …) that already has one bound — check the vendor's server.py for both lists.
# Named preset (desktop)
CACHE_ROOT=/home/<user> docker --context desktop compose run --rm --pull always \
--service-ports lerobot-0_3_3-server sim_stack
# Custom checkpoint
CACHE_ROOT=/home/<user> docker --context desktop compose run --rm --pull always \
--service-ports lerobot-server ee --pipeline.source.checkpoints_dir=<ckpt-dir>
# GR00T inference — codec subcommand required
CACHE_ROOT=/home/<user> docker --context notebook compose run --rm --pull always \
--service-ports groot-server ee_rot6d --pipeline.source.checkpoints_dir=<ckpt-dir>
Run detached with -d for a background server; docker --context <ctx> ps / logs <id> / stop <id> to manage it. Point the client at the context's hostname:
uv run --locked positronic eval run --eval=.sim.positronic.stack_cubes \
--policy=.remote --policy.url=desktop:8000 \
--output_dir=<...>
Gotchas: each GR00T server uses ~6GB, so only one at a time on a 12GB GPU; on a port conflict, docker --context <ctx> ps -a | grep -E "server" then stop the stale container.
End-to-End Validation
e2e.sh runs the whole pipeline for one vendor (convert → train 200 steps → serve → /api/v1/models smoke → teardown), polling Nebius and printing a per-stage status line. ~$2–5 per run. Use it to verify a vendor still works after an image/dependency/script change.
bash workflows/nebius/e2e.sh openpi
# All four — seed the cache with one first, then fan out warm
bash workflows/nebius/e2e.sh lerobot_0_3_3
for v in lerobot openpi gr00t; do bash workflows/nebius/e2e.sh "$v" & done; wait
Override via env: E2ES3BASE, E2EEXPNAME, E2ELOGROOT, E2E_DATASET.
Analysing a past run
Each convert/train/serve run writes runmetadata.yaml into its S3 output directory. It records the exact CLI command that produced the artifact and a snapshot of the relevant source files (.py/*.toml), so a checkpoint or dataset can be traced back to the code and arguments that made it — without guessing.
# List the metadata files under a checkpoint/dataset output dir
aws s3 ls s3://checkpoints/sim_stack/openpi/ee/pi05_positronic_lowmem/<exp>/ \
--recursive | grep run_metadata_
# Read one (full command + code snapshot)
aws s3 cp s3://checkpoints/.../run_metadata_YYYYMMDD_HHMMSS.yaml - | less
To reproduce a run, copy the command recorded in runmetadata.yaml and resubmit it via the matching workflows/nebius/.sh wrapper (set NEBIUSIMAGETAG if the run used a non-latest image).
For inference runs, each episode also has a static.json alongside the recorded data; the eval viewer in [4. Run Inference Client](#4-run-inference-client) (positronic.cfg.analysis sim --dataset.base.path=…) renders these for inspection and side-by-side comparison of multiple runs.
Monitoring Jobs & Endpoints
# Jobs
nebius ai job get <aijob-id> # state: PROVISIONING/STARTING/RUNNING/COMPLETED/FAILED
nebius ai job logs <aijob-id> --follow
nebius ai job list --parent-id "$NEBIUS_PARENT_ID" --format json | jq '.items[].metadata.name'
# Endpoints
nebius ai endpoint get <endpoint-id>
nebius ai endpoint logs <endpoint-id> --follow # wait for "INFO Started server process"
nebius ai endpoint list --parent-id "$NEBIUS_PARENT_ID" --format json
The create call streams the job ID and ready-to-paste follow-up commands.
Common Issues
- Job stuck in PROVISIONING/STARTING: normal — image pull +
uv resolve.
First run on a cold cache is ~10 min; check nebius ai job logs <id> --follow.
- OpenPI train can't find stats: pass
--statspath=<outputdir-sibling>/stats/assets/
exactly as printed by convert.sh openpi. Stats must be a sibling of the dataset dir (pos3 forbids upload-inside-download).
- gr00t with read-only input mount fails: only
lerobot03_3 uses the RO
Mountpoint-S3 mount; gr00t writes back into the dataset dir, so it streams via pos3 from the s3:// path instead (handled automatically by train.sh).
- Endpoint name collision: names must be unique in the project. Pick a fresh
name or stop.sh the old one.
- Cold cache + parallel fan-out: 4 jobs racing to populate
/cache thrash;
seed with one vendor, then fan out the rest warm (see e2e.sh header).
- Wipe the shared cache: throwaway
busybox job mounting the FS — see
"Cleaning the shared cache" in workflows/nebius/README.md.
Nebius Auth (Headless)
nebius --no-browser --auth-timeout 5m iam whoami 2>&1 — extract auth URL
- User clicks URL, browser redirects to
http://127.0.0.1:PORT/?code=XXX&state=YYY
curl -s "http://127.0.0.1:PORT/?code=XXX&state=YYY" on the machine running nebius
- Auth completes, scripts work