Academic RL Demonstration

AdaptiveRL Autonomous 3D Drone Navigation

An educational reinforcement-learning system where a PPO agent learns to navigate a simulated 3D drone through an obstacle-filled environment toward a target.

ā„¹ļø
THIS IS A SIMULATION. Academic computer-science project running a simplified 3-DOF kinematic point-mass simulation with aerodynamic drag. It is not an autopilot, PX4 controller, or 6-DOF quadrotor aerodynamics simulator.
Explore the System View on GitHub
29-D Observation Vector
3-D Continuous Acceleration
16-Ray Spherical LiDAR
49 / 49 Unit Tests Passing
READY
Interactive 3D Arena (30m Ɨ 30m Ɨ 15m)
Step: Step 0 / 200
Pos: (5.0, 5.0, 3.0)
Speed: 0.0 m/s
Goal Range: 28.3 m
Return: 0.0
SCRUB
Quick Overview

ARL in 2 Minutes

How reinforcement learning trains an autonomous agent to navigate a 3D simulated arena in 5 fundamental stages.

01 — ENVIRONMENT

3D Arena Simulation

The drone operates inside a bounded 30m Ɨ 30m Ɨ 15m arena containing procedurally spawned spherical obstacles and aerodynamic drag damping.

02 — OBSERVATION

29-D State Vector

At each 0.1s step, the agent observes normalized position, linear velocity, target vector, goal distance, and 16 spherical LiDAR range rays.

03 — DECISION

PPO Continuous Action

The neural actor network outputs continuous translational acceleration commands in the X, Y, and Z axes bounded within [-4.0, +4.0] m/s².

04 — REWARD

Objective Shaping

Dense progress reward (+10·Δd) guides flight toward the beacon, step penalty (-0.05) limits time, and collision penalties (-100.0) discourage impacts.

05 — LEARNING

Clipped Policy Updates

Stable-Baselines3 collects trajectory rollouts and computes clipped surrogate objective updates, incrementally stabilizing collision-free flight paths.

Concept & Motivation

What Is AdaptiveRL?

Traditional geometric motion planners like A* or RRT* require exhaustive environmental maps upfront. ARL explores model-free reinforcement learning: training an agent to navigate reactively using local distance sensors.

šŸŽÆ Problem 3D Reach
→
šŸ“¦ Environment 30Ɨ30Ɨ15m
→
šŸ“” Observation 29 Values
→
🧠 PPO Agent Actor-Critic
→
⚔ Action 3D Accel
→
🚁 Drone Kinematics
→
šŸ† Reward Ī”d āˆ’ Penalties
↺

The drone is simulated as a 3-DOF kinematic point-mass in continuous 3D space with linear drag damping:

dv/dt = a - (c_d * v)
dp/dt = v
Ī”t = 0.1 seconds,  c_d = 0.05,  v_max = 8.0 m/s,  a_max = 4.0 m/s²

At each discrete timestep, the commanded acceleration vector a is integrated via semi-implicit Euler integration, clamping velocity magnitude to 8.0 m/s and detecting boundary penetrations against the arena walls.

Modular Engineering

System Architecture

Four decoupled components designed for rigorous reproducibility, automated testing, and clear demonstration.

🌐 1. Environment

  • Farama Gymnasium API standard
  • 30m Ɨ 30m Ɨ 15m bounded arena
  • Procedural spherical obstacles
  • Continuous translational drag
  • Boundary & collision detection

šŸ“” 2. Perception

  • 16-Ray Spherical LiDAR
  • Analytical ray-sphere intersections
  • Normalized relative target vector
  • Linear drone velocity tracking
  • 29 continuous dimensions total

🧠 3. RL Agent

  • Stable-Baselines3 PPO
  • Actor-Critic MLP architecture
  • Clipped surrogate loss (ε = 0.2)
  • Generalized Advantage Estimation
  • Continuous action Gaussian heads

šŸ“Š 4. Evaluation & GUI

  • Random Baseline comparator
  • Obstacle density stress tests (4, 6, 8)
  • 5-Tab Streamlit flight deck
  • Deterministic multi-episode seeds
  • 49 unit tests across Python 3.10–3.12
Gymnasium Specification

Drone Environment & State Space

Precise parameters and observation specifications verified directly against src/adaptive_rl/environments/drone.py.

Observation Space

29 Dimensions (Continuous)

The agent receives a 29-element normalized Box(-1.0, 1.0) observation vector:

Components Dims Description
norm_pos 3 Drone position (X, Y, Z) normalized to [0, 1]
norm_vel 3 Linear velocity normalized by max speed (8 m/s)
norm_goal 3 Target position (X, Y, Z) normalized to [0, 1]
rel_goal 3 Relative vector from drone to target
norm_dist 1 Normalized Euclidean distance to goal
lidar_rays 16 16 spherical rangefinder readings (0 to 20m)
Total Vector 29 Full state observation space

Action & Physics

Box(-1.0, 1.0, shape=(3,))

Continuous translational acceleration commands in Cartesian 3D space:

Parameter Value Unit
Action Dimensions 3 [a_x, a_y, a_z] in [-1.0, 1.0]
Max Acceleration 4.0 m/s²
Max Speed 8.0 m/s
Time Step (dt) 0.1 seconds per step
Drag Coefficient (cd) 0.05 linear aerodynamic damping
Target Radius 1.5 meters (success threshold)
Collision Radius 0.8 meters (drone bounding sphere)

Reward Formulation

Dense Distance-Shaped

Episodic flight guidance reward calculated at every time step t:

R_t = 10 * (d_{t-1} - d_t)      # Progress reward
    - 0.05                      # Timestep penalty
    - 0.01 * ||a_t||²           # Actuator effort penalty
    + 100.0 (if d_t <= 1.5m)    # Goal arrival bonus
    - 100.0 (on collision)      # Obstacle/wall penalty
    - 50.0  (on boundary crash) # Arena boundary penalty

Encourages rapid approach toward the target coordinate while heavily penalizing collisions with spherical obstacles or arena walls.

Reinforcement Learning Core

Proximal Policy Optimization (PPO)

Why PPO was selected for continuous 3D drone navigation, and how the actor-critic policy optimizes flight paths.

Clipped Surrogate Objective

Standard policy gradient methods suffer from destructively large updates that collapse training. PPO stabilizes updates by clipping the probability ratio:

L^CLIP(θ) = E_t [ min( r_t(θ) * A_t,  clip(r_t(θ), 1-ε, 1+ε) * A_t ) ]

where:
  r_t(θ) = π_θ(a_t | s_t) / π_θ_old(a_t | s_t)
  A_t    = Generalized Advantage Estimator (GAE)
  ε      = 0.2 (clipping boundary parameter)

This bound prevents the updated policy from straying too far from previous experience, ensuring smooth, monotonically improving flight behaviors.

Hyperparameters (configs/drone_ppo_demo.yaml)

Framework Stable-Baselines3 + PyTorch
Policy Network MlpPolicy (Actor-Critic)
Learning Rate 3e-4 (Adam optimizer)
Discount Factor (γ) 0.99
GAE Parameter (Ī») 0.95
Clip Range (ε) 0.2
Batch Size 64
Demo Budget 25,000 steps (~25s on CPU)
Perception System

Interactive 16-Ray LiDAR Inspector

Test how the simulated spherical LiDAR rangefinder senses obstacles in real-time. Drag the drone to observe raycast distances dynamically.

Live Sensor Telemetry

CLEAR TRAJECTORY
-- Nearest Obstacle Clearance
-- Distance to Target

16 Spherical LiDAR Rays (Distance in Meters)

šŸ’” Rays change color dynamically: Cyan (>9m clearance) → Yellow (4–9m) → Red (<4m proximity hazard).

Empirical Validation

Verified Benchmark Results

Measured experimental results generated through genuine Python CPU execution under seed 42. No fabricated numbers.

PPO vs Random Policy Baseline

20 Test Episodes (Seed 42)
Policy Success Collision Mean Return Outcome
Random Baseline 0.0% 100.0% -101.24 Collided every run
Trained PPO Agent 5.0% 35.0% -3.34 65% survival rate

Key Finding: An untrained random policy crashes in 100% of test flights. PPO slashes collisions to 35%, proving the neural network learned meaningful obstacle avoidance and directional flight.

Obstacle-Density Stress Test

10 Episodes / Setting (Seed 42)
Condition Obstacles Collision Rate Mean Return Flight Steps
Low Density 4 20.0% +5.00 169.0
Medium Density 6 40.0% -15.20 136.2
High Density 8 70.0% -32.43 88.2

Key Finding: Collision rates climb monotonically as obstacle density increases from 4 to 8, validating the hypothesis that spatial clutter non-linearly increases 3D navigation difficulty.

Streamlit Flight Deck

The Actual ARL Application

The repository includes a 5-tab browser flight deck built with Streamlit and Plotly for live demonstration and defense.

šŸŽ® Tab 1: Live 3D Demo

Interactive Plotly 3D flight arena, trajectory trail, obstacle spheres, target beacon, and step-by-step playback slider.

šŸ“ˆ Tab 2: Train PPO

In-browser CPU training loop with real-time status notifications, progress indicator, and live reward curves.

šŸ“Š Tab 3: Evaluation

Side-by-side empirical benchmark comparing PPO against the random baseline with comparative bar charts.

šŸŽÆ Tab 4: Density Test

Stress-test policy robustness across 4, 6, and 8 obstacles with real-time performance degradation graphs.

šŸ“˜ Tab 5: Architecture

Professor viva walkthrough, 3-DOF kinematic formulas, 29-D observation vector breakdown, and honest limitations.

Launch the Live Flight Deck Locally

Run either command from your activated virtual environment to open the Streamlit interface:

bash
adaptive-rl gui
# or
streamlit run app.py
Video Showcase

Demonstration Video

Watch the trained PPO agent navigate complex obstacle configurations in real-time.

šŸŽ¬

Demo Video Coming Soon

A full video walkthrough demonstrating Tab 1 (3D flight playback) and Tab 3 (PPO vs baseline evaluation) is being recorded for the repository.

Run ARL Locally Instead
Reproducibility

Run ARL Locally (Linux Guide)

Step-by-step instructions to clone, install, verify, and run AdaptiveRL from a clean Linux terminal.

1

Install System Prerequisites

Ensure Python >= 3.10, pip, and git are available on your Linux distribution:

Ubuntu / Debian / Linux Mint
sudo apt update && sudo apt install -y python3 python3-pip python3-venv git
2

Clone Repository & Create Virtual Environment

Clone the project and create an isolated Python virtual environment:

bash
git clone https://github.com/StellarResearch/ARL.git
cd ARL
python3 -m venv .venv
source .venv/bin/activate
3

Install Dependencies & Verify Tests

Install AdaptiveRL in editable mode and run the 49 unit tests:

bash
python -m pip install --upgrade pip setuptools wheel
python -m pip install -e ".[all]"
python -m pytest tests/
4

Train & Evaluate Policy

Train the PPO demonstration agent (~25s on CPU) and run the baseline comparison:

bash
adaptive-rl train --config configs/drone_ppo_demo.yaml
adaptive-rl evaluate --config configs/drone_ppo_demo.yaml --compare-random
5

Launch 3D GUI Flight Deck

Launch the interactive browser GUI:

bash
adaptive-rl gui
Academic Boundaries

Technical Limitations

āš ļø
Kinematic Point-Mass Physics: Models translational 3-DOF acceleration with drag damping; does not simulate 6-DOF attitude quaternions, rotor gyroscopic precession, or blade wash.
āš ļø
Analytical Geometric LiDAR: Raycasting computes exact mathematical sphere intersections without modeling laser bounce noise, atmospheric dust, or surface reflectance.
āš ļø
Simulation Only: This model is not deployable directly onto real quadrotor flight controllers (like PX4 or ArduPilot) without a cascaded low-level PID attitude controller.
Future Research

Planned Research Directions

šŸš€
6-DOF Rigid Body Dynamics: Extending the simulation to full quadrotor attitude dynamics, motor RPM lag, and torque coupling.
šŸŒŖļø
Domain Randomization: Injecting randomized turbulent wind gusts, payload mass variations, and LiDAR sensor dropout noise.
šŸ¤–
Gazebo / PX4 SIL Integration: Connecting the trained SB3 policy to standard robotics middleware via ROS 2 or PX4 Software-in-the-Loop.
šŸ‘„
Multi-Agent Swarm Navigation: Investigating cooperative collision avoidance among multiple learning drones sharing the same 3D airspace.
Open Science Stack

Technology Stack & Engineering Rigor

Built with battle-tested open-source scientific computing libraries and verified through continuous integration.

Python >= 3.10
Gymnasium Farama Env
Stable-Baselines3 PPO Algorithm
PyTorch Tensor Compute
NumPy Vector Math
Streamlit Dashboard GUI
Plotly 3D Visualization
Typer & Rich CLI Tools
Pytest 49 Tests Passing
Ruff Code Linting
Mypy Static Type Analysis
GitHub Actions Multi-Python CI

CI Pipeline Verification

Every commit is tested on clean Ubuntu runners across Python 3.10, 3.11, and 3.12.

āœ“ 49 Tests Passed āœ“ Typecheck Passed āœ“ Lint Passed
Open Source Codebase

Explore the Source Repository

AdaptiveRL is released under the permissive MIT License for academic, educational, and research use.

View on GitHub Read College Demo Viva Guide Read Formal Experiment Docs