Skip to content

Latest commit

 

History

3 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Mini LLM Detector Benchmark — Starter

This is a tiny, runnable starter to build a 20–120 paragraph benchmark to test simple AI-text detectors.

What you’ll do

  1. Fill data/data.csv with paragraphs (human + AI).
  2. Install deps: pip install -r requirements.txt
  3. Run: python src/eval.py
  4. See: results/metrics.csv and results/summary.txt
  5. Write your 1‑pager using report/one_pager_template.md

Data format (CSV)

Columns (keep these headers exactly): id,text,label,domain,paraphrase_level,language,source,notes

  • label: human or ai
  • domain: stem, policy, narrative (use these three to start)
  • paraphrase_level: none, light, heavy
  • language: en or es->en (if you machine-translate to Spanish and back)
  • source: e.g., self, ku_website, wikipedia, chatgpt
  • notes: short free text (e.g., prompt name, link slug)
  • Wrap the text field in double quotes. If your text contains quotes, double them.

Quick test

A few sample rows are already in data/data.csv. Replace them with your own.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages