This project explores deep reinforcement learning in a classic grid-maze game: Pac-Man. The task is deceptively simple—collect dots, avoid enemies—but it couples long-horizon planning, partial observability, and tight navigation constraints imposed by walls and corridors. Vanilla single-frame DQN often struggles here because a single observation cannot reveal motion (e.g., whether an enemy is approaching) or short-term trends (e.g., whether Pac-Man is reversing). To address this, we build a frame-stacked DQN agent that augments each observation with a short history window, giving the network access to implicit velocity and intent without the complexity of recurrent models.
Beyond the network input, reward design is pivotal. Sparse “win/lose/eat-dot” signals alone lead to slow or unstable learning, especially in mazes with dead ends. We adopt potential-based reward shaping that is policy-invariant: a small shaping term encourages moving toward the nearest dot and away from the nearest enemy, where distances are computed on the walkable grid via BFS (not Manhattan), so the agent is never misled by walls. This is combined with clear terminal rewards (clear-all, death, timeout), a light living cost, and gentle penalties for wall bumps and instantaneous reversals to reduce oscillation.
Learning stability is further improved with Double DQN targets and Prioritized Experience Replay (PER), so the agent focuses on transitions with large TD error while correcting sampling bias via importance weights. We also use a target network for bootstrapping, ε-greedy exploration with decay, and optional gradient clipping.
For clarity and reproducibility, the environment features fixed spawns (Pac-Man near the center, enemies near the four corners), a compact state encoder (flattened grid + direction one-hot), and a renderer with an optional Q-value heatmap for diagnostics. The codebase includes training, testing, checkpointing, and replay-memory persistence to resume long runs. Taken together, these design choices make the project a practical, self-contained case study of modern DQN techniques on a nontrivial control problem.
A quick map of the repository and what each piece does.
.
├── agent.py # DQN Agent (policy/target nets, epsilon policy, frame stacking, step/reset, save/load)
├── config.py # All hyperparameters & constants (GRID_SIZE, INPUT_SIZE, HIDDEN_SIZE, LR, GAMMA, etc.)
├── environment.py # Maze world: state encoding, reward, movement/step(), fixed spawns, helpers (BFS distance)
├── learning.py # Model & learning utilities: DQN, train_dqn (Double DQN + PER), select_action helper
├── replayMemory.py # Prioritized replay buffer (push/sample/update_priorities)
├── renderer.py # Pygame renderer & heatmap diagnostics
├── train.py # Training entry point (loop, logging, checkpointing, memory save)
├── test_pacman.py # Greedy evaluation with rendering (loads a checkpoint)
└── checkpoints/ # Saved artifacts
├── latest_model.pth
├── memory.pkl
└── meta.json
-
config.py- Central place for constants (e.g.,
K_FRAMES,INPUT_SIZE,OUTPUT_SIZE,HIDDEN_SIZE,BATCH_SIZE,EPSILON_*,TARGET_UPDATE_FREQ,CHECKPOINT_DIR). - Changing stack size or maze size should only require updating config (and the derived
INPUT_SIZE).
- Central place for constants (e.g.,
-
environment.py- Maze & spawns: parses
MAZE, computesgrid_w/h, pixel dims, fixed player/enemy spawns. - Step API:
next_state, reward, done = env.step(action). - State encoder:
get_state() -> torch.FloatTensor[1, W*H + 4](flattened grid + direction one-hot). - Reward: potential-based shaping on BFS grid distance (toward dots, away from enemies) + terminal/living terms.
- Helpers:
_is_walkable,_grid_distance_to_set(multi-source BFS), coordinate conversions.
- Maze & spawns: parses
-
agent.py- Wraps the RL logic around the environment.
- Networks:
policy_net,target_net(bothDQNfromlearning.py). - Frame stacking: maintains a deque of the last
K_FRAMESsingle-frame states; exposes_stacked(). - Acting:
select_action(state=None)(ε-greedy) andstep(env, action)→ returns(stacked_t, a, r, stacked_t1, done). - Training hooks:
optimize_model(),update_target(), optional soft update if used. - I/O:
save(model_path, memory_path), optionaltry_load_memory/save_memoryif enabled. - Diagnostics:
compute_heatmap(env)(batch inference + K-frame approximation for visualization).
-
learning.pyDQN: MLP with LeakyReLU, sized forINPUT_SIZE(frame-stacked).train_dqn: one training step (Double DQN target, PER importance weights, SmoothL1/Huber loss, optimizer step, priority update, optional grad clip).select_action: stateless helper for ε-greedy evaluation.
-
replayMemory.py-
ReplayMemorywith Prioritized Experience Replay:push(state, action, reward, next_state, done)sample(batch_size, beta)→(batch, indices, weights)update_priorities(indices, td_errors)
-
-
renderer.py- Pygame window management, grid/dots/enemies rendering.
draw_heatmap()converts a Q-value grid into a colored overlay.
-
train.py- Orchestrates episodes: reset, step loop, epsilon schedule, target updates, logging.
- Saves model (
latest_model.pth), meta (meta.json), and optionally replay memory (memory.pkl) at intervals.
-
test_pacman.py- Loads a checkpoint, runs greedy (ε=0) episodes with rendering.
- Uses the same frame-stacked
Agentinterface (reset_episode,select_action,step).
- Actions: discrete {0: left, 1: right, 2: up, 3: down}.
- Single-frame state:
[1, BASE_FEAT_DIM]whereBASE_FEAT_DIM = W*H + 4. - Stacked state:
[1, INPUT_SIZE]whereINPUT_SIZE = BASE_FEAT_DIM * K_FRAMES. - Replay item:
(state, action, reward, next_state, done)withstate/next_state= stacked tensors. - Batches in training:
states, next_states → [B, INPUT_SIZE],actions/rewards/dones → [B].
This section explains the learning choices in theory, but all statements reflect what the current implementation actually does.
We learn an action-value function
with the Huber (Smooth L1).
We use Double DQN targets to reduce over-estimation bias:
where
A single Pac-Man frame does not reveal velocities or short-term trends (e.g., who is approaching? did we just reverse?). To mitigate partial observability without the complexity of RNNs, we feed the network a stack of the last
This “sliding window” supplies implicit velocity/intent cues through finite differences across frames. It typically stabilizes learning (fewer oscillations) and improves reaction quality near intersections. In replay, each transition is done=True,
Each single frame
- a flattened occupancy grid of the maze (walls, dots, enemies, Pac-Man) and
- a one-hot for Pac-Man’s current movement direction.
The stacked input is simply the concatenation of
Sparse signals (“win/lose/eat-dot”) alone are slow to learn in mazes. We therefore add potential-based shaping, which is policy-invariant: it does not change the set of optimal policies when written as
Our potentials are:
-
Toward dots:
$\Phi_{\text{dot}}(s)=-d_{\text{dot}}(s)$ -
Away from enemies:
$\Phi_{\text{enm}}(s)=+d_{\text{enm}}(s)$
Crucially, distances
- Terminal: large positive for clearing all dots; large negative for death; small negative for timeout/cap.
- Instantaneous: small living cost; positive for eating a dot; small penalties for wall bumps (no movement) and immediate reversals (to suppress oscillation).
-
Shaping: weighted sum of the two potentials above with the same
$\gamma$ as learning.
Intuition: the base events set the task, the shaping acts as a gentle compass (not a new objective).
Uniform replay wastes updates on uninformative transitions. We sample transitions with probability
where
To correct the induced bias we use importance sampling (IS) weights:
which anneal with
-
$\varepsilon$ -greedy with decay: start exploratory, then gradually exploit as$Q$ stabilizes. - Target network updates: periodical hard copy (or optional soft updates) to keep bootstrapping stable.
-
Optimization details: Adam with a small weight decay; optional gradient clipping;
no_gradaround targets to avoid graph pollution;eval()/train()toggles where needed to keep deterministic targets. -
Episode caps as terminals: time/step caps are treated as
done=Truein replay so the learner sees terminal samples regularly, anchoring value scales.
This summarizes how training actually runs in this repo—the schedules, caps, buffers, and what the numbers mean. Values shown in parentheses reflect current defaults in config.py or module defaults.
- Episode caps. Each episode ends on terminal (death or all dots cleared) or on caps:
step cap (
MAX_STEPS_PER_EPISODE=1000) or time cap (MAX_EPISODE_TIME=30s). Caps are treated as terminal in replay so the learner regularly seesdone=Truetransitions. - Act–learn cadence. On every environment step we store one stacked transition and do one gradient update.
-
State. The DQN sees K-frame stacked observations
$S_t=[s_{t-K+1},\dots,s_t]$ . Single-frame features are (flattened grid + direction one-hot). Effective input width isINPUT_SIZE = BASE_FEAT_DIM * K_FRAMES. -
Batching. Mini-batch TD learning with batch size B (
BATCH_SIZE=50).
-
Loss. Smooth L1 (Huber) on the TD error.
-
Discount.
$\gamma = 0.99$ (GAMMA=0.99). -
Optimizer. Adam with learning rate 3e-4 (
LR=0.0003). (A small weight decay may be applied in the agent to stabilize fitting.) -
Double DQN target. Action selection by the online net, evaluation by the target net: $$ a^{} = \operatorname{arg,max}{a} Q{\text{online}}(S_{t+1}, a), \qquad y_t = \begin{cases} r_t, & \text{if terminal},\[2pt] r_t + \gamma, Q_{\text{target}}(S_{t+1}, a^{*}), & \text{otherwise}. \end{cases} $$
-
Target updates. Hard copy every N episodes (
TARGET_UPDATE_FREQ=10).
- Start fully exploratory
ε_start=1.0, then decay per episode byε_decay=0.9995down toε_min=0.05. Greedy evaluation usesε=0.0.
-
Buffer. Capacity
MEMORY_CAPACITY=10000, ring buffer on CPU tensors. -
Prioritization. PER with exponent
$\alpha$ (default0.6in the buffer). New samples receive current max priority so they’re sampled at least once. -
Sampling. Draw a batch with importance sampling weights
$w_i$ using$\beta$ . -
Terminals in batch. The buffer can enforce a minimum number of terminal samples per batch
(e.g.,
min_terminal_samples=10) so targets get regularly “anchored” bydone=Truecases. -
Priority updates. After backprop, priorities are updated from fresh
$|\delta|$ .
- Terminal: big positive (clear all), big negative (death), small negative (timeout/cap).
- Per-step: light living cost; + for dot; small penalties for wall-bump and instant reversal.
-
Shaping: potential-based terms w.r.t. BFS shortest paths (toward dots, away from enemies) with the same
$\gamma$ . Shaping nudges behavior without changing optimal policies.
-
Artifacts.
checkpoints/latest_model.pth— current policy weightscheckpoints/meta.json— episode/epsilon metadatacheckpoints/memory.pkl— pickled replay buffer (optional, for long runs)
-
Snapshots. A full snapshot is written periodically (e.g., every 100 episodes).
-
Resume. On startup you can load the model and (optionally) the replay memory to continue training without a cold buffer.
- Console logs periodically print: moving loss, mean Q, and #terminal samples in the last batch.
- Heatmap (optional). A diagnostic Q-value heatmap renders the spatial preference of the current policy over the maze; for visualization it approximates stacking by repeating the single frame K times.
- Engine. Python, PyTorch, NumPy (2.0-compatible), and Pygame for rendering.
- Device. GPU if available; otherwise CPU runs (slower but functional).
- Seeds. For strict reproducibility, fix Python/NumPy/PyTorch seeds and disable nondeterministic kernels; the project runs deterministically enough for comparison even without strict seeding.
# Test (greedy[train.py](train.py) play with rendering)
python test_pacman.py --model checkpoints/latest_model.pth --fps 30Due to hardware constraints, we have not run extended training or large-scale evaluations yet. The implementation has been validated to run end-to-end (training loop, PER sampling, Double DQN targets, frame stacking, rendering/heatmap, checkpoint & replay-memory persistence). Early smoke tests show stable loss reduction, but quantitative results are intentionally omitted until longer runs are feasible.
When resources are available, we plan to report:
- Learning curves: episode return, loss, mean/max Q.
- Ablations: K=1 vs K>1 (frame stacking), with/without shaping, with/without PER.
- Qualitative behavior: survival time, collision rate, wall-bump rate, heatmap snapshots.
- Reproducibility: seed-averaged curves (≥3 seeds), same configs.