A Deep Q-Network (DQN) implementation for training an agent to land a spacecraft safely in the LunarLander-v3 environment from Gymnasium.
This project implements a complete reinforcement learning pipeline for the LunarLander-v3 environment, including:
- DQN Agent: A deep neural network-based Q-learning agent with experience replay
- Environment Integration: Full integration with Gymnasium's LunarLander-v3 environment
- Visualization & Analysis: Tools to visualize episode data and analyze agent behavior
- Demo System: Ready-to-run demos to test the agent and environment
Train an agent to safely land a lunar module between two landing pad flags with minimal fuel consumption and stable balance.
The agent observes:
- x position: Horizontal position relative to landing pad
- y position: Vertical position (decreases as it falls)
- x velocity: Horizontal speed
- y velocity: Vertical speed (negative when falling)
- angle: Tilt angle of the lander
- angular velocity: Rate of rotation
- left leg contact: Binary flag for left leg touching ground
- right leg contact: Binary flag for right leg touching ground
The agent can take one of four actions each timestep:
- 0: Do nothing
- 1: Fire left orientation engine
- 2: Fire main engine
- 3: Fire right orientation engine
The agent receives rewards for:
- โ Moving closer to the landing pad
- โ Landing safely (legs touching ground)
- โ Staying upright (low angle)
- โ Reducing speed
- โ Penalties for crashing, tilting badly, wasting fuel, and inefficient movement
gdsc-ml2-1/
โโโ main.py # Entry point with three modes
โโโ src/
โ โโโ dqn_agent.py # DQN agent implementation
โ โโโ dqn_lunarlander.pth # Trained model weights
โ โโโ env/
โ โ โโโ env_info.py # Environment constants & info printer
โ โ โโโ environment_guide.md # Detailed environment documentation
โ โโโ visualization/
โ โโโ plot_states.py # Data collection & plotting utilities
โโโ demos/
โ โโโ demo_env.py # Demo runner for the environment
โ โโโ requirements.txt # Python dependencies
โโโ README.md # This file
Implements the Deep Q-Network algorithm with:
QNetwork: Neural network that maps state (8 dims) โ Q-values (4 actions)
- Input layer: 8 neurons (state)
- Hidden layer 1: 128 neurons + ReLU activation
- Hidden layer 2: 128 neurons + ReLU activation
- Output layer: 4 neurons (one Q-value per action)
ReplayBuffer: Stores and samples past experiences to break temporal correlations
- Capacity: 50,000 experiences
- Sampling: Random batches for training stability
DQNAgent: Main training logic
- Epsilon-greedy exploration: Starts 100% random, decays to 1% random over 10,000 steps
- Two networks: Online (trained) and Target (frozen, updated every 500 steps)
- Discount factor (ฮณ): 0.99 (future rewards matter)
- Learning rate: 0.001 (Adam optimizer)
- Batch size: 64 experiences per training step
env_info.py: Easy access to environment constants
STATE_NAMES: List of all state variable namesACTION_MEANINGS: Mapping of action indices to descriptionsprint_env_info(): Prints formatted environment information
collect_episode_data(): Runs an episode and records all state/reward data plot_main_states(): Visualizes position and velocity over time plot_angle_states(): Visualizes angle and angular velocity plot_leg_contacts(): Visualizes leg contact with ground plot_reward_curve(): Plots reward per step with moving average smoothing moving_average(): Utility for smoothing noisy signals
Runs 3 episodes of the environment with random actions (baseline behavior).
Python 3.11+ with a virtual environment set up containing all dependencies from demos/requirements.txt:
- gymnasium (RL environment)
- torch (neural network framework)
- numpy (numerical computing)
- matplotlib (visualization)
- pygame (rendering)
- Box2D (physics simulation)
The main.py script provides three modes:
Watch 3 random episodes with visualization:
python main.py --mode demoShows the agent taking random actions. Total reward should be negative (untrained agent).
Run one episode and visualize the data:
python main.py --mode plotGenerates four plots:
- Position & Velocity vs. Time
- Angle & Angular Velocity vs. Time
- Leg Contact vs. Time
- Reward per Step (with moving average smoothing)
Print environment information:
python main.py --mode infoDisplays:
- All 8 state variable names and indices
- All 4 action meanings
Currently, the agent uses random actions (untrained). The trained model weights are stored in src/dqn_lunarlander.pth.
- Total Reward: Negative (agent crashes/wastes fuel)
- Angle: Increases over time โ lander tilts and loses balance
- Angular Velocity: Builds up โ uncontrolled spinning
- Vertical Velocity: Becomes increasingly negative โ crashes into ground
- Leg Contacts: Usually don't occur before crash (unstable descent)
- Controlled Descent: Balance vertical velocity to avoid crashing
- Sideways Stability: Manage x velocity and angle to stay upright
- Fuel Efficiency: Use engines sparingly
- Landing: Time the final descent to land with legs touching
All dependencies are listed in demos/requirements.txt:
gymnasium==1.2.3 # RL environments
torch==2.11.0 # Neural networks
numpy==2.4.4 # Numerical computing
matplotlib==3.10.8 # Plotting
pygame==2.6.1 # Rendering
Box2D==2.3.10 # Physics simulation
cloudpickle==3.1.2 # Serialization
... (and others)
The agent learns to estimate Q-values (expected future reward for each action):
Where:
-
$s$ = current state -
$a$ = action taken -
$r$ = immediate reward -
$\gamma$ = discount factor (0.99) -
$s'$ = next state -
$a'$ = best action in next state
- Online Network: Updated every step with gradient descent
- Target Network: Frozen copy updated every 500 steps
- Purpose: Provides stable training targets, reduces oscillations
- Stores up to 50,000 transitions (state, action, reward, next_state, done)
- Samples random batches of 64 during training
- Benefits: Breaks temporal correlations, improves sample efficiency
- ฮต-start: 1.0 (100% random exploration)
- ฮต-end: 0.01 (1% exploration, 99% greedy)
- ฮต-decay: Linear over 10,000 steps
Episode 1
Total Reward: -245.32
Episode 2
Total Reward: -312.18
Episode 3
Total Reward: -289.45
=== LunarLander Environment Info ===
State Variables:
0: x_position
1: y_position
2: x_velocity
3: y_velocity
4: angle
5: angular_velocity
6: left_leg_contact
7: right_leg_contact
Actions:
0: do nothing
1: fire left orientation engine
2: fire main engine
3: fire right orientation engine
- Gymnasium Docs: https://gymnasium.farama.org/
- DQN Paper: "Human-level control through deep reinforcement learning" (Mnih et al., 2015)
- LunarLander Details: https://gymnasium.farama.org/environments/box2d/lunar_lander/
- The project uses PyTorch on GPU if available, otherwise CPU (auto-detected)
- Pygame is used for rendering environments
- Box2D provides realistic physics simulation for the lunar lander
- Currently set up for single-episode data collection; can be extended for training loops