Fiatlux: A Long-Horizon Benchmark
for Humanoid Ladder Climbing
and Light-Bulb Replacement

Pavel Bushuyeu*1, Yujin Chen*1, Anton Nikolaev*2, Brian Shu*1,3, Igor Molybog1

1HawAII, University of Hawaiʻi at Mānoa    2Independent Researcher    3Purdue University

*The first four authors contributed equally; their order is alphabetical by surname.

Abstract

Existing benchmarks evaluate tabletop manipulation, flat-floor household activity, or humanoid locomotion and manipulation as separate task groups; none scores vertical mobility and dexterous work on a fragile payload in one long-horizon episode.

We present Fiatlux, a light-bulb replacement benchmark built on NVIDIA Isaac Lab. In one episode, a Unitree G1 humanoid positions a step ladder under a ceiling or wall fixture, climbs it, exchanges a spent bulb in a socket for a fresh one, and leaves the spent one in a disposal crate. We decompose the episode into twelve subtask environments scored on difficulty-weighted gates. The goal is a successful replacement, with the fresh bulb seated, the spent one disposed of, neither dropped, and a fragility bound not crossed. Runs that fall short can earn partial credit.

Observations are split into a standard mode (signals a physical robot could sense or estimate) and a privileged mode (exact simulator state). We specify the evaluation protocol and provide reference baseline implementations spanning RSL-RL PPO, zero-shot NVIDIA GR00T N1.7 Vision-Language-Action (VLA) models, and whole-body controllers. Additionally, we provide the teleoperated recordings used to specify and check the success gates.

Contributions

What the benchmark provides

  1. Benchmark for climb-required maintenance Formalizes overhead bulb replacement as a task family. The fixture sits above the robot's standing reach, so the episode cannot be completed without climbing.
  2. Modular subtask architecture and unified scene hierarchy Decomposes multi-stage execution into 12 atomic subtasks across Flat Navigation, Manipulation and Climbing control modes, sharing one family scene hierarchy with per-task preset layouts. Each subtask is a registered environment with a start pose approximating the state its predecessor ends in, and its own success gate.
  3. Automated asset generation Converts object images into textured USD assets through a staged image-to-mesh pipeline with optional articulation prediction and a structural verifier, supporting benchmark extension through explicit calibration rather than hand-modeled meshes.
  4. Domain randomization Randomizes the scene layout per layout draw, grip friction at startup, and lighting, material colors and the robot start state per reset.
  5. Teleoperated demonstration adapters A teleoperation adapter for every subtask, in which an operator in a VR headset drives the arms through inverse kinematics while the GEAR-SONIC controller walks and balances the legs.
  6. Empirical baseline evaluation Establishes zero-shot evaluation protocols for Vision-Language-Action architectures, tested specifically on NVIDIA GR00T N1.7 paired with whole-body controllers.
12subtasksone environment each
53actuated DoF29 body joints, 12 per hand
50 Nfragility boundon a 35 g bulb; a closing grip stalls near 12 N
8 / 12subtasks shown achievable107 teleoperated takes, 80 at full credit
Task suite

Twelve subtasks, difficulty-weighted

The subtasks are not of equal difficulty, so their success rates are not averaged. Each carries a difficulty weight, and each gate is declared as a list of conjuncts, so an episode that satisfies part of a gate is distinguishable from one that achieved nothing.

Achievability is established by driving the robot through each subtask under teleoperation to its success gate. A take counts only if it satisfies every condition of the gate, sustained for the required window where one applies, and takes are recorded across multiple layout seeds so a success cannot depend on one favorable spawn.

Subtask difficulty weights and teleoperated validation takes. Eight of the twelve subtasks have at least 10 clean full-credit takes. The four climbing subtasks are dashed: those entries are outstanding, not zero. A take can satisfy its gate and still score 0.00 once the broken or dropped penalty applies, which is why full-credit counts trail take counts.
ID Subtask Control mode Weight Takes Full credit
S01Move LadderFlat navigation2.251110
S02Climb LadderClimbing3.00——
S03Remove Old BulbManipulation4.201210
S04Descend with BulbClimbing4.50——
S05Carry Bulb to DisposalFlat navigation1.501010
S06Dispose BulbManipulation1.801210
S07Approach New BulbFlat navigation1.001010
S08Grab New BulbManipulation1.401010
S09Carry Bulb to LadderFlat navigation1.501010
S10Climb with BulbClimbing4.50——
S11Screw in BulbManipulation5.403210
S12Climb DownClimbing3.00——
Validation

Teleoperated takes

One frame per subtask. Eight come from a teleoperated take that fires that subtask's gate. The four climbing subtasks have no such take — those entries are outstanding, not zero.

On-ladder stance

The robot is placed on the top tread by inverse kinematics, with no leg joints in the action space. These show that it holds the on-ladder stance with and without the bulb. The transitions to and from that stance — the climbs themselves — are not demonstrated, and neither frame is an achievability take for S02 or S10.

Results

No released baseline completes a subtask

Each policy runs one episode per subtask at each of four layout seeds. Rollout trajectories are recorded as bags and scored offline, so the same recording can be re-scored under different rules without re-running the simulator. groot uses the base checkpoint without fine-tuning and a per-subtask language instruction; zero holds joint position commands at zero offset and random samples each action uniformly from [−1, 1].

Per-subtask score: zero-shot GR00T N1.7 against the floor baselines and teleoperation. Each cell is mean ± sample standard deviation over the policy's runs. GR00T, zero and random are averaged across 4 domain-randomization seeds; teleoperation across its recorded takes. *Teleoperation's weighted score covers only the 8 subtasks with an entry (19.05 of 34.05 total weight); the 4 dashed subtasks are excluded from the sum, not scored zero.
ID Subtask GR00TN1.7 zero-shot Zerono action Randomuniform action Teleoperationhuman expert
S01Move Ladder0.000.000.001.00
S02Climb Ladder0.000.000.00—
S03Remove Old Bulb0.06 ± 0.120.000.06 ± 0.120.90 ± 0.25
S04Descend with Bulb0.250.12 ± 0.140.00—
S05Carry Bulb to Disposal0.250.250.001.00
S06Dispose Bulb0.000.12 ± 0.080.330.85 ± 0.34
S07Approach New Bulb0.000.06 ± 0.120.06 ± 0.121.00
S08Grab New Bulb0.000.000.001.00
S09Carry Bulb to Ladder0.250.12 ± 0.140.001.00
S10Climb with Bulb0.19 ± 0.120.19 ± 0.120.00—
S11Screw in Bulb0.000.12 ± 0.140.19 ± 0.120.57 ± 0.45
S12Climb Down0.000.000.00—
Weighted total 0.0876 ± 0.0261 0.0861 ± 0.0322 0.0569 ± 0.0293 0.8419*

The benchmark is unsolved

Every number in the table is gate progress, not a completion. The scorer keeps the best moment of each episode. A policy that never finishes still earns partial credit for every gate condition it holds. The 0.06, 0.19 and 0.25 entries are that partial credit. The success rate counts only the gate firing, and it is 0.00 for all three released baselines, on every subtask, at every seed. We read that result from every recorded trajectory bag.

The weighted-total row shows the same thing. Zero-shot GR00T N1.7 scores 0.0876 ± 0.0261. The zero baseline commands no action, and it scores 0.0861 ± 0.0322. The gap is smaller than either standard deviation, so the two policies do not separate. Teleoperation scores 0.8419, but that number does not compare to the other three, because it covers only the eight subtasks that have takes. Those eight subtasks are achievable but unsolved, and no take shows a climb.

How they fail

All three fail differently. With random, ten of the twelve subtasks end in a fall, none averaging more than 50 control steps. With zero, six do, and mean episode length runs from 38 to 6000 control steps against random's 19 to 46 — the robot stays upright long enough for another termination to fire first. Where the bulb is carried on or at the ladder, it is crushed or dropped before the robot falls. Elsewhere the robot neither falls nor succeeds, and runs out the time budget. For teleoperation, most failures are bulb crushes.

Protocol

Scoring and observation modes

Partial credit

Each success gate is a list of conjuncts. With n conjuncts, cmax the most held simultaneously at any step and c0 the number held at reset, gate progress is the clamped ratio (cmax − c0) / (n − c0), combined with the subtask's difficulty weight into the weighted score.

The fragility bound

An episode is broken if peak contact force over both hands, across all steps and all 34 hand and wrist bodies, exceeds 50 N. Both hands are read, since scoring one arm would record a left-handed crush as clean. Finger torque saturates at 0.5 N·m, stalling a closing finger at roughly 12 N of grip, so the 35 g bulb can be held and released well inside the bound.

Clean success is the headline metric

Success and the violation rates are computed independently, so a policy that seats the bulb by crushing it registers a success and a broken episode at once. Clean success rate — fresh bulb seated, spent one disposed of, neither dropped, fragility bound never crossed — is therefore the primary number.

Two observation modes

Standard carries only what a physical robot could sense or estimate: joint positions and velocities across 53 actuated DoF, base angular velocity from the IMU, projected gravity, estimated base height and velocity, contact forces, RGB and lidar range. The G1 has no foot sensor, so foot contact is estimated from ankle joint torques — in simulation as on hardware, with the policy consuming the same signal in both.

Privileged carries exact simulator state for critic training and privileged baselines: absolute 6D poses of robot, ladder, fixture, both bulbs and the crate, plus the score-relevant distances and seating flags. Every baseline declares which mode it consumes, so sim-to-real claims are checkable against the observation group.

Code

Getting started

Requires uv, an NVIDIA GPU with a CUDA 12.8-capable driver, and gsutil for the assets. Isaac Sim 5.1 and Isaac Lab 2.3.2 are pinned in uv.lock — no manual Isaac Lab install.

# 1. Build the environment (first run pulls ~10 GB).
uv sync

# 2. Pull the USD assets (G1, bulb/socket, ladder).
./assets/download_assets.sh

# 3. Sanity-check registration.
uv run python scripts/list_envs.py
uv run python scripts/verify_scene.py --headless

# 4. Record a run, then score it offline (no simulator needed to score).
uv run python scripts/record_run.py --task FIATLUX-Replace-v0 --policy random \
    --episodes 2 --record bag --headless --enable_cameras --out logs/runs/random0
uv run python scripts/score.py logs/runs/random0

Full setup, teleoperation and the GR00T baseline server are documented in getting_started.md.

Citation

BibTeX

@misc{bushuyeu2026fiatlux,
  title        = {Fiatlux: A Long-Horizon Benchmark for Humanoid
                  Ladder Climbing and Light-Bulb Replacement},
  author       = {Bushuyeu, Pavel and Chen, Yujin and Nikolaev, Anton
                  and Shu, Brian and Molybog, Igor},
  year         = {2026},
  howpublished = {\url{https://fiatlux-bench.github.io}}
}
Acknowledgements

This work was supported by the National Science Foundation NRT-AI 2244574 and through allocation number CIS240027 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, which is supported by National Science Foundation grants #2138259, #2138286, #2138307, #2137603 and #2138296. The technical support and advanced computing resources from University of Hawaii Information Technology Services — Research Cyberinfrastructure, funded in part by the National Science Foundation CC* awards #2201428 and #2232862, are gratefully acknowledged. This work was supported by computational resources provided by NPC Labs through the B3IQ infrastructure platform.