Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

AI_Moralistic_Drift

Final Project for PLIN0072

More Personal, More Severe: Normative Drift Under Agentic Commitment in Multi-Turn Dialogue

Replication package for the PLIN0072 final project by Caroline Swartz.

This repository contains all code to reproduce the experiment studying normative drift in language models across multi-turn moral reasoning dialogues.


Hardware Requirements

  • GPU: NVIDIA A10 (24 GB VRAM) or equivalent
  • RAM: 32 GB recommended
  • Storage: ~50 GB (model weights + outputs)

The experiment uses 4-bit quantization via bitsandbytes, so models can fit on a single 24 GB GPU. A Lambda Labs A10 instance was used for the original runs.


Software Requirements

  • Python 3.10+
  • CUDA 11.8+ (for GPU inference)
  • A HuggingFace account with access to gated models (Llama 3.1)

Models used:

  • meta-llama/Meta-Llama-3.1-8B-Instruct
  • google/gemma-2-9b-it
  • Qwen/Qwen3-8B

Setup

1. Clone the repository

git clone https://github.com/syntax606/PLIN0072_Final_Project.git
cd PLIN0072_Final_Project

2. Create a virtual environment and install dependencies

bash environment/setup.sh
source ~/venv/bin/activate

This creates a virtualenv at ~/venv and installs all packages from environment/requirements.txt.

3. Log in to HuggingFace

Llama 3.1 is a gated model. You must accept its license on HuggingFace and authenticate:

huggingface-cli login

Create a read-access token at https://huggingface.co/settings/tokens and paste it when prompted.


Running the Experiment

Run all steps in this exact order from the project root with the virtualenv active:

# Step 1: Generate scenario scripts
python -m src.make_scripts

# Step 2: Run model inference (slow — expect several hours)
python -m src.generate

# Step 3: Sanity-check outputs
python -m src.sanity_check

# Step 4: Parse structured model responses
python -m src.parse_outputs

# Step 5: Compute turn-level and conversation-level metrics
python -m src.metrics

# Step 6: Ordinal regression on stance scores
python -m src.stats_stance

# Step 7: Regression models for drift hypotheses
python -m src.stats_drift

# Step 8: Generate all figures
python -m src.plots

Do not skip or reorder steps. Each step depends on the outputs of the previous one.


Expected Outputs

After a full run the following files will be present:

Scenario scripts (created by make_scripts):

data/scripts_deception.jsonl
data/scripts_manipulation.jsonl
data/scripts_duty.jsonl
data/scripts_collective.jsonl
data/scripts_epistemic.jsonl

Model generations and parsed responses (created by generate and parse_outputs):

data/generations.jsonl
data/parsed.csv

Metrics (created by metrics):

results/turn_metrics.csv
results/conversation_metrics.csv
results/reliability.csv
results/analysis_dataset.csv

Statistical models (created by stats_stance and stats_drift):

results/stance_ordered_logit.txt
results/drift_deviation_model.txt
results/drift_signed_model.txt
results/path_dependence_model.txt
results/rule_revision_model.txt
results/history_effect_model.txt
results/ndi_model.txt

Figures (created by plots):

results/figures/fig1_mean_shift_by_C.png
results/figures/fig2_mean_deviation_by_C.png
results/figures/fig3_path_dependence.png
results/figures/fig4_history_effect.png
results/figures/fig5_rule_shift_hist.png
results/figures/fig6_ndi_by_order.png
results/figures/fig7_trajectory_examples.png

Experimental Design

The experiment crosses:

  • 3 models: Llama 3.1-8B-Instruct, Gemma 2-9B-IT, Qwen3-8B
  • 5 moral domains: deception, manipulation, duty, collective harm, epistemic honesty
  • 2 trajectory orders: forward (R → C1 → C2 → C3 → C4), reverse (C4 → C3 → C2 → C1 → R)
  • 2 history conditions: history (full conversation context) vs. no_history (each turn evaluated independently)
  • 2 repeat runs per combination (for reliability estimation)

All models run at temperature=0 with max_new_tokens=256.


Tests

Before running the full pipeline, verify prompt integrity:

pytest

The key test (tests/test_prompt_snapshots.py) checks that prompts have not changed from the reference snapshots. If this test fails, do not proceed — results will not be comparable.


Reproducibility Notes

  • Every generation is logged with the model ID, prompt hash, and generation parameters.
  • Set SEED = 42 is used throughout (see src/config.py).
  • Results may differ slightly across GPU hardware or driver versions even at temperature=0, due to floating-point non-determinism in attention kernels.

Project Structure

├── environment/        # setup script and requirements
├── paper/              # LaTeX source and figures
├── src/                # all pipeline code
│   ├── config.py       # models, parameters, experimental design
│   ├── prompt_schema.py
│   ├── make_scripts.py
│   ├── generate.py
│   ├── sanity_check.py
│   ├── parse_outputs.py
│   ├── metrics.py
│   ├── stats_stance.py
│   ├── stats_drift.py
│   └── plots.py
├── tests/              # prompt snapshot tests
└── data/               # created by pipeline (not tracked in git)

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages