Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

HEAR

Run 23 speech language models on your own audio task, through one interface.

EMNLP 2026 arXiv Project page Dataset Model License

Code for HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding (EMNLP 2026 Main Conference).

The harness it is built on is not specific to HEAR. Every speech LM has its own prompt format, audio placeholder token, and loading incantation; this hides all of that behind one JSONL contract, so running twenty models on your audio task is a for-loop rather than a month.


Quickstart

git clone https://github.com/dwsmart32/HEAR && cd HEAR
pip install -r requirements.txt
pip install vllm            # or transformers, depending on the model

One line per query:

{"id": "q1", "instruction": "Who speaks second?\n\nOptions:\n(A) ...\n(B) ...", "audio_path": "clip.wav", "ground_truth": "B"}

Required fields are id, instruction, audio_path. Add ground_truth only if you want the bundled scorers; relative paths resolve against the JSONL.

python inference.py --model qwen2_5_omni-7b --input my_task.jsonl --output_dir runs/mine

Results are JSON with each model response attached to its input record, so any scorer can read them. Sweeping models is a shell loop:

for m in qwen2_5_omni-7b voxtral-mini-3b phi-4-mm-6b; do
  python inference.py --model "$m" --input my_task.jsonl --output_dir runs/mine
done

--validate-input-only checks every audio path before you spend GPU time. examples/run_custom_task.py builds the JSONL from a folder of clips and runs the sweep for you.


Models

23 models, 26 registry entries, all public Hugging Face ids or API model names. The left column is what you pass to --model. Start with qwen2_5_omni-7b on a GPU, or gemini-3-flash without one.

backend --model
vLLM qwen2_5_omni-3b qwen2_5_omni-7b qwen3-omni-instruct-30b qwen3-omni-thinking-30b qwen2-audio-7b voxtral-mini-3b voxtral-small-24b phi-4-mm-6b midashenglm-7b fun-audio-chat-8b a2r-30b-a3b
transformers minicpm-o-2_6-8b minicpm-o-4_5-9b minicpm-o-4_5-9b-thinking kimi-audio-7b-instruct gemma4-e2b-it gemma4-e4b-it audio-flamingo3
API gpt-4o-audio gemini-3-pro gemini-3-flash gemini-2.5-pro
external step-audio-2-mini step-audio-r1 (bridge not bundled, disabled)

requirements.txt installs the shared core only; add vllm or transformers as needed. Gemma-4 and Audio-Flamingo 3 want transformers>=5.0, Kimi-Audio additionally needs kimia_infer from source. API models read keys from .env (cp .env.example .env).

Adding a model: copy the closest entry in registry.yaml and change path. For vLLM models, also add the key to MODEL_FAMILY_MAP in src/backends/vllm_handlers/__init__.py. Qwen-family fine-tunes need nothing special: the handler reads config.json to pick the prompt format, so a checkpoint under any name is prompted like its base. Override per entry with system_prompt.


Layout

inference.py        run any model over a JSONL task
registry.yaml       the 26 entries
src/                the harness: runner, four backends, per-family prompt adapters
examples/           run_custom_task.py, a folder of clips to a model sweep
hear/               the bundled benchmark: run_hear.py, the scorers, its prompt audio

Nothing under src/ or in inference.py knows what HEAR is.


The HEAR benchmark

HEAR asks who is speaking, not only what is said: 2,395 questions over 887 multi-party clips. Half come in counterfactual pairs, where one utterance is re-voiced as another speaker's voice. Same words, same timing, different speaker, and the answer flips, so a model reading only the transcript is caught.

The dataset is gated. Once per account: accept the agreement on the dataset page, then hf auth login.

python hear/run_hear.py --model qwen2_5_omni-7b        # download, run, score
python hear/eval_hear_vera.py --results runs/hear/results/<model>__hear_input_results.json

hear/eval_hear.py gives accuracy overall, per axis, per sub-dimension and per category. hear/eval_hear_vera.py gives Ver-A, the paper's headline metric, which counts a counterfactual pair as one unit that scores only when both members are correct: 1,235 items + 580 pairs = 1,815 units. Transcript-centric models lose 30 to 50 points crossing from per-item accuracy to Ver-A.

Both read the answer from the last <answer> span when the response has one, and from the whole string otherwise. Models that reason before answering quote option letters while they reason, so scanning the whole string picks up the wrong letter -- on IR, whose options are long verbatim quotes the model enumerates as it works, that alone moves A2R by more than 40 points.

Prompt style

Two prompts. Head, question and options are identical; only the closing instruction differs.

--prompt closing instruction
plain Respond in JSON format, e.g. {"Answer": "A"}
transcript_first transcribe the conversation into <transcript> first, reason over that transcript, then answer inside <answer>

Every baseline row in the paper was produced with plain, and A2R's with transcript_first. The default here reproduces that pairing: registry.yaml gives a2r-30b-a3b a prompt_style: transcript_first, and every other entry falls back to plain.

The two are not interchangeable, and not in the same direction for every model. Scored over the 2,395 items the Hub dataset ships, as overall accuracy / Ver-A:

plain transcript_first
A2R-30B-A3B 62.7 / 54.8 68.8 / 61.1
Qwen3-Omni-30B-Instruct 36.9 / 25.7 35.5 / 23.5

Transcribing first is worth +6.1 overall and +6.3 Ver-A to A2R, and moves its base model the other way. It is what A2R was trained to do, not a better prompt in general -- which is why the default is per-model rather than global. --prompt overrides it for any model:

python hear/run_hear.py --model a2r-30b-a3b                          # transcript_first
python hear/run_hear.py --model a2r-30b-a3b --prompt plain           # the comparison above
python hear/run_hear.py --model qwen2_5_omni-7b --prompt transcript_first

The style is part of the input file name, so a plain run and a transcript_first run never share a cached input. Under transcript_first the model is asked for {"Answer": "(A) option text"} rather than a bare letter; the scorers read both.

Audio is pre-assembled, so the audio column already holds everything the model must hear: the clip alone for VC/VCD/IR/QR/TR, clip plus a reference voice for VL/VCA, clip plus five announced voices for CVA.


Also in the release

🤗 HEAR 2,395 questions over 887 clips
🤗 CASH-60K/ in the same repo 59,762 training queries with counterfactual hard negatives
🤗 ood_benchmark/ in the same repo three held-out benchmarks, 1,286 questions
🧠 A2R-30B-A3B the merged 30B checkpoint

Citation

@misc{lee2026hearsaidwhatunlocking,
  title         = {HEAR Who Said What: Unlocking Speaker-Attributed Reasoning
                   via Counterfactual Voice Grounding},
  author        = {Dongwook Lee and Sangkwon Park and Eunwoo Song and Che Hyun Lee
                   and Youngho Cho and Junho Kim and June Young Yi and Heeseung Kim
                   and Sungroh Yoon},
  year          = {2026},
  eprint        = {2608.29120},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  url           = {https://arxiv.org/abs/2608.29120}
}

License

Code under MIT. HEAR and CASH-60K are CC BY-NC 4.0; the source corpora (AMI, ICSI, VoxMM) keep their own terms.

About

Run 23 speech language models on your own audio task through one interface: vLLM, transformers, and API backends behind a single JSONL contract. Ships HEAR, the speaker-attribution benchmark from our EMNLP 2026 paper.

Topics

Resources

Stars

7 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages