Final Project for PLIN0072
Replication package for the PLIN0072 final project by Caroline Swartz.
This repository contains all code to reproduce the experiment studying normative drift in language models across multi-turn moral reasoning dialogues.
- GPU: NVIDIA A10 (24 GB VRAM) or equivalent
- RAM: 32 GB recommended
- Storage: ~50 GB (model weights + outputs)
The experiment uses 4-bit quantization via bitsandbytes, so models can fit on a single 24 GB GPU. A Lambda Labs A10 instance was used for the original runs.
- Python 3.10+
- CUDA 11.8+ (for GPU inference)
- A HuggingFace account with access to gated models (Llama 3.1)
Models used:
meta-llama/Meta-Llama-3.1-8B-Instructgoogle/gemma-2-9b-itQwen/Qwen3-8B
git clone https://github.com/syntax606/PLIN0072_Final_Project.git
cd PLIN0072_Final_Projectbash environment/setup.sh
source ~/venv/bin/activateThis creates a virtualenv at ~/venv and installs all packages from environment/requirements.txt.
Llama 3.1 is a gated model. You must accept its license on HuggingFace and authenticate:
huggingface-cli loginCreate a read-access token at https://huggingface.co/settings/tokens and paste it when prompted.
Run all steps in this exact order from the project root with the virtualenv active:
# Step 1: Generate scenario scripts
python -m src.make_scripts
# Step 2: Run model inference (slow — expect several hours)
python -m src.generate
# Step 3: Sanity-check outputs
python -m src.sanity_check
# Step 4: Parse structured model responses
python -m src.parse_outputs
# Step 5: Compute turn-level and conversation-level metrics
python -m src.metrics
# Step 6: Ordinal regression on stance scores
python -m src.stats_stance
# Step 7: Regression models for drift hypotheses
python -m src.stats_drift
# Step 8: Generate all figures
python -m src.plotsDo not skip or reorder steps. Each step depends on the outputs of the previous one.
After a full run the following files will be present:
Scenario scripts (created by make_scripts):
data/scripts_deception.jsonl
data/scripts_manipulation.jsonl
data/scripts_duty.jsonl
data/scripts_collective.jsonl
data/scripts_epistemic.jsonl
Model generations and parsed responses (created by generate and parse_outputs):
data/generations.jsonl
data/parsed.csv
Metrics (created by metrics):
results/turn_metrics.csv
results/conversation_metrics.csv
results/reliability.csv
results/analysis_dataset.csv
Statistical models (created by stats_stance and stats_drift):
results/stance_ordered_logit.txt
results/drift_deviation_model.txt
results/drift_signed_model.txt
results/path_dependence_model.txt
results/rule_revision_model.txt
results/history_effect_model.txt
results/ndi_model.txt
Figures (created by plots):
results/figures/fig1_mean_shift_by_C.png
results/figures/fig2_mean_deviation_by_C.png
results/figures/fig3_path_dependence.png
results/figures/fig4_history_effect.png
results/figures/fig5_rule_shift_hist.png
results/figures/fig6_ndi_by_order.png
results/figures/fig7_trajectory_examples.png
The experiment crosses:
- 3 models: Llama 3.1-8B-Instruct, Gemma 2-9B-IT, Qwen3-8B
- 5 moral domains: deception, manipulation, duty, collective harm, epistemic honesty
- 2 trajectory orders: forward (R → C1 → C2 → C3 → C4), reverse (C4 → C3 → C2 → C1 → R)
- 2 history conditions:
history(full conversation context) vs.no_history(each turn evaluated independently) - 2 repeat runs per combination (for reliability estimation)
All models run at temperature=0 with max_new_tokens=256.
Before running the full pipeline, verify prompt integrity:
pytestThe key test (tests/test_prompt_snapshots.py) checks that prompts have not changed from the reference snapshots. If this test fails, do not proceed — results will not be comparable.
- Every generation is logged with the model ID, prompt hash, and generation parameters.
- Set
SEED = 42is used throughout (seesrc/config.py). - Results may differ slightly across GPU hardware or driver versions even at
temperature=0, due to floating-point non-determinism in attention kernels.
├── environment/ # setup script and requirements
├── paper/ # LaTeX source and figures
├── src/ # all pipeline code
│ ├── config.py # models, parameters, experimental design
│ ├── prompt_schema.py
│ ├── make_scripts.py
│ ├── generate.py
│ ├── sanity_check.py
│ ├── parse_outputs.py
│ ├── metrics.py
│ ├── stats_stance.py
│ ├── stats_drift.py
│ └── plots.py
├── tests/ # prompt snapshot tests
└── data/ # created by pipeline (not tracked in git)