How benchmark files become CSVs for routing experiments. Main scripts: multidata_unify.py (one merged table) and construct_router_data.py (per-model responses and scores).
download_*.py → raw JSON/Parquet under data/<task>/
│
├─► multidata_unify.py → single CSV (optional; no splits)
│
└─► generate_dataset_splits.py → <data_root>/data/<TASK>/{train,val,test}.csv
│
▼
construct_router_data.py → <TASK>/<model_name>.csv
- Raw data — Run
data_processing/download_*.pyor place files wheregenerate_dataset_splits.py/multidata_unify.pyexpect them (see script configs). - Splits — Run
generate_dataset_splits.py; setdata_dirin__main__to your root. You needtrain.csv,val.csv,test.csvper task (columns includequery,ground_truth,metric). Required forconstruct_router_data.py. - API — Set keys in
configs/models.yaml(or env). MatchProviderTypeinconstruct_router_data.py(DataBuilder, currently OpenRouter-style). - Env —
export PYTHONPATH=<project_root>and run from project root. - Build model CSVs — Edit
construct_router_data.pymain():dataset_names,data_dirs,models_to_test. Runpython data_processing/construct_router_data.py. Output:{model_name}.csvper task (response, tokens,response_time,effect, …).use_saved_flag=Trueskips redoing API calls when split outputs already exist.
Optional: multidata_unify.py only exports one flat CSV (no train/val/test); it does not feed construct_router_data.py directly.
Merges tasks into one CSV: task_id, sub_task (MMLU subjects only), query, ground_truth, metric, task_description.
| task_id | Source | Notes |
|---|---|---|
| alpaca_data, GSM8K, multi_news | data/.../*.json |
Concat instruction+input; GT from output/answer |
| SQUAD | data/SQUAD/SQUAD.parquet |
question; first answers.text |
| MBPP | data/mbpp/mbpp_all.json |
Templated prompt + tests; GT code |
| mmlu_redux | data/mmlu_redux/*.json |
MCQ string; GT (A)–(D) |
generate_unified_qa_dataset(output_path, sample_size, mmlu_sample_size) — cap rows per task / per MMLU subject. Run: python data_processing/multidata_unify.py (tune __main__).
Loads three splits, calls the LLM per row, evaluates with LLMProvider.eval(metric=...), writes {model_name}.csv. Uses gpt2 tokenizer for usage fallback and truncates multi_news to 3000 tokens. Retries: 3× with backoff.
Depends on llm_engine.py, configs/models.yaml, transformers, pandas, pyyaml.
| File | Role |
|---|---|
generate_dataset_splits.py |
Writes per-task train/val/test.csv |
generate_train_val_test.py |
Alternative: unified_qa_data_{split}.csv in one folder |
utils.py, llm_engine.py |
I/O and API + metric eval |
Keep paths and column names in sync with DataBuilder.process_split if you change them.