Christian Minich
Back to Projects

Go2 Navigation & Inspection with RL, VLA & World Models

Master's thesis training and benchmarking Reinforcement Learning, Vision-Language-Action models, and World Models for Unitree Go2 navigation and inspection in NVIDIA Isaac Sim and Isaac Lab.

Tech Stack

roboticsunitree-go2reinforcement-learningvlaworld-modelsisaac-simisaac-labpytorchjiraconfluence

Go2 Navigation & Inspection with RL, VLA & World Models

Master's Thesis | UniversitΓ€t OsnabrΓΌck | Ongoing | Deadline: 30 November 2026

What are we trying to do?

Train and evaluate three different AI approaches to control a Unitree Go2 for autonomous robot navigation and inspection inside a simulated industrial environment.

                         ISAAC SIM / ISAAC LAB

             Camera   Depth   LiDAR   IMU   Robot State
                \       |       |      |       /
                 \      |       |      |      /
                  β””β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”˜
                              β”‚
                              β–Ό
                    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                    β”‚    UNITREE GO2    β”‚
                    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β”‚
              β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
              β–Ό               β–Ό               β–Ό
             RL              VLA         WORLD MODEL
             PPO           OmniVLA      V-JEPA / LeWM
              β”‚               β”‚               β”‚
              β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                              β–Ό
                  NAVIGATION + INSPECTION
                              β”‚
                              β–Ό
               Compare Performance & Robustness

The robot should learn to move through an environment, avoid obstacles, locate a target, position itself correctly and support inspection tasks around industrial objects, machines or infrastructure.

The initial benchmark focuses on a controlled navigation task such as:

START
  β”‚
  β–Ό
Explore Environment
  β”‚
  β–Ό
Find Target
  β”‚
  β–Ό
Avoid Obstacles
  β”‚
  β–Ό
Navigate to Target
  β”‚
  β–Ό
Position Go2
  β”‚
  β–Ό
Inspect
  β”‚
  β–Ό
SUCCESS

Inputs

All approaches operate inside the same simulated environment.

InputPurpose
πŸ“· RGB CameraVisual environment
🌐 Depth CameraDistance and geometry
πŸ“‘ LiDARSpatial surroundings
🧭 IMUOrientation and movement
🦿 Joint StatesGo2 body configuration
βš™οΈ Actuator StatesRobot movement state
πŸ“ Robot PosePosition for training and evaluation
πŸ’¬ Language GoalSemantic target for VLA
🎯 Target StateGoal definition and evaluation

The benchmark records these streams synchronously so that experiments can be reproduced and compared.

The Three Approaches

Reinforcement LearningVLAWorld Model
ApproachRLVision-Language-ActionWorld Model
Model directionPPOOmniVLA / VLA adaptationV-JEPA 2 / LeWM
Main inputSensors + robot stateVision + instructionObservation sequence
OutputNavigation actionGoal-conditioned actionPredicted representation / action support
TrainingEnvironment interactionAdaptation / fine-tuningRecorded trajectories
Main questionCan it learn robust navigation?Can it understand what to navigate to?Can it predict what happens next?

The exact VLA and World Model implementations remain subject to feasibility testing during development.

Shared Task

The important part of the thesis is that the approaches are not tested in completely different setups.

                    SAME ROBOT
                   Unitree Go2
                        β”‚
                        β–Ό
                 SAME ENVIRONMENT
               Isaac Sim / Isaac Lab
                        β”‚
                        β–Ό
                   SAME TASK
             Navigate + Inspect
                        β”‚
                        β–Ό
                 SAME SCENARIOS
                        β”‚
                        β–Ό
              SAME EVALUATION PIPELINE
                  /      |      \
                 /       |       \
                β–Ό        β–Ό        β–Ό
               RL       VLA       WM

This makes it possible to investigate where each approach performs well and where its limitations appear.

Simulation

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              NVIDIA ISAAC SIM                β”‚
β”‚                                              β”‚
β”‚  Environment    Physics    Sensors    Go2    β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
                       β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚              NVIDIA ISAAC LAB                β”‚
β”‚                                              β”‚
β”‚  Training    Reset Logic    Rewards    RL     β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                       β”‚
                       β–Ό
               TRAINING + BENCHMARK

The environment can be changed systematically to test whether a model really learned useful behavior.

Room Layout        β–‘β–‘β–‘ β†’ β–“β–“β–“
Lighting           β˜€ β†’ ◐ β†’ β—Ό
Obstacle Position  A β†’ B β†’ C
Target Position    1 β†’ 2 β†’ 3
Sensor Noise       Low β†’ Medium β†’ High
Target Appearance  Known β†’ Modified β†’ Unseen

Hardware

The project uses two main GPU environments.

Local WorkstationRemote Compute
GPURTX 5090A100
Isaac Simβœ“
Isaac Labβœ“
Renderingβœ“
Sensor Generationβœ“
RL Trainingβœ“βœ“
VLA Trainingβœ“
World Model Trainingβœ“
Large Batchesβœ“
RTX 5090
   β”‚
   β”œβ”€β”€ Isaac Sim
   β”œβ”€β”€ Isaac Lab
   β”œβ”€β”€ Go2 Simulation
   β”œβ”€β”€ Sensors
   └── Benchmark Execution


A100
   β”‚
   β”œβ”€β”€ VLA
   β”œβ”€β”€ World Models
   β”œβ”€β”€ Neural Training
   └── Large Batch Experiments

Algorithms & Frameworks

AreaTechnology
RobotUnitree Go2
SimulatorNVIDIA Isaac Sim
Robot LearningNVIDIA Isaac Lab
RLPPO / RSL-RL
VLAOmniVLA / VLA adaptation
World ModelsV-JEPA 2 / LeWM direction
Deep LearningPyTorch
Fast SimulationMuJoCo / MJX
Local GPURTX 5090
Training GPUA100
Project ManagementJira
DocumentationConfluence

Benchmark Pipeline

SCENARIO
   β”‚
   β–Ό
ISAAC SIM / LAB
   β”‚
   β–Ό
GO2 + SENSORS
   β”‚
   β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚   RL    β”‚   VLA   β”‚   WM    β”‚
β””β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”˜
     β”‚         β”‚         β”‚
     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
               β–Ό
        ROBOT BEHAVIOR
               β”‚
               β–Ό
         EPISODE LOG
               β”‚
               β–Ό
          EVALUATION
               β”‚
               β–Ό
         FINAL RESULTS

Each run stores the scenario, sensor observations, robot state, actions, target, model output, timing, collisions and final result.

What Do We Compare?

NavigationRobustnessTrainingSystem
Success RateNew LayoutsTraining DataInference Time
CollisionsSensor NoiseTraining TimeGPU Usage
Time to GoalNew TargetsSample EfficiencyModel Size
Path EfficiencyObstaclesGPU HoursImplementation Effort
Final PositionLightingGeneralizationReproducibility

The result is a capability comparison, not one artificial overall score.

Project Management

The thesis is managed as a technical project using Jira + Confluence.

                    MASTER THESIS
                         β”‚
            β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
            β–Ό                         β–Ό
          JIRA                    CONFLUENCE
            β”‚                         β”‚
      Work Packages              Requirements
      Tasks                      Architecture
      Deadlines                  Research
      Progress                   Decisions
      Acceptance Criteria        Experiments
      Implementation             Documentation
            β”‚                         β”‚
            β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                         β–Ό
                 TRACEABLE PROJECT

Jira

WP1  Setup
 β”‚
WP2  Research
 β”‚
WP3  Simulation
 β”‚
WP4  Benchmark + Data
 β”‚
WP5  Pilot Experiments
 β”‚
WP6  RL + VLA + WM
 β”‚
WP7  Final Benchmark
 β”‚
WP8  Thesis
 β”‚
 β–Ό
30 NOVEMBER 2026

The Jira workflow follows:

TO DO β†’ IN PROGRESS β†’ TESTING β†’ DONE

Confluence

PROJECT
β”‚
β”œβ”€β”€ Charter & Scope
β”œβ”€β”€ Requirements
β”œβ”€β”€ Research
β”œβ”€β”€ System Architecture
β”œβ”€β”€ Experimental Design
β”œβ”€β”€ Approach Specifications
β”œβ”€β”€ Risk Register
β”œβ”€β”€ Decision Log
β”œβ”€β”€ Project Roadmap
└── Reproducibility

Jira manages what needs to be done.

Confluence documents what was designed, why decisions were made and how experiments are performed.

Current Status

JULY                 AUGUST               SEPTEMBER
Setup ──────────► Simulation ─────────► Training
                       β–²
                       β”‚
                    WE ARE HERE


OCTOBER                               NOVEMBER
Training ─────► Benchmark ─────► Results ─────► Thesis
                                              β”‚
                                              β–Ό
                                        30.11.2026

Current focus:

Go2 Environment
       +
Sensors
       +
RL / PPO
       +
Reward Design
       +
Episode Logging
       ↓
FIRST COMPLETE TRAINING PIPELINE

The next stage expands this pipeline toward the VLA and World Model implementations and then freezes the benchmark configuration for the final comparative experiments.

Final Goal

                 UNITREE GO2
                     β”‚
        β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
        β–Ό            β–Ό            β–Ό
       PPO          VLA           WM
        β”‚            β”‚            β”‚
        β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”Όβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                     β–Ό
             SAME ROBOT TASK
                     β–Ό
         NAVIGATION + INSPECTION
                     β–Ό
           CONTROLLED BENCHMARK
                     β–Ό
              COMPARE RESULTS
                     β–Ό
           MASTER'S THESIS
                     β–Ό
              30.11.2026

The final deliverable is a working simulation demonstrator and reproducible benchmark showing how RL, VLA and World Model approaches can be trained and evaluated for autonomous robot navigation and inspection with a Unitree Go2.