Fine-tune a Laya choice model for your own single-label text classification task. Supply labeled training and validation CSV files, an instruction, and the allowed labels in a JSON config. One run can train one classifier or share a model across several tasks. The repository also retains the Banking77, ArBanking77, and CLINC150 benchmark setup as a reproducible preset. It is designed for text with one correct label per example, not multi-label, image, or audio classification.
The default source is the pinned convaiinnovations/laya-multilingual checkpoint. Training uses supervised cross-entropy with shuffled option order. It does not reproduce upstream RLCD training. Outputs are PyTorch checkpoints; Core ML export is optional.
The published Banking77, ArBanking77, and CLINC150 fine-tuned model includes both PyTorch weights and a Core ML export. Those weights are an example output, not a requirement for training your own classifier.
- Python 3.11–3.13 and
uv - Apple Silicon with MPS or an NVIDIA GPU with CUDA for practical training
- Enough disk space for the source model, environment, and checkpoints
CPU execution is supported but can be slow. --device auto selects CUDA, then MPS, then CPU.
uv sync --extra devSee the runnable customer-support example. Each CSV needs a text column and a label column. Keep a separate test split outside the training config and use it only after selecting a checkpoint. Every label must appear in both training and validation, and exact text overlap between those splits is rejected.
text,label
I need my invoice,billing
The app will not open,technicalPlace your data in a private directory such as data/my-classifier/ and create a config beside it:
{
"version": 1,
"tasks": [
{
"name": "customer_support",
"instruction": "Choose the team that should handle this customer request.",
"labels": ["billing", "technical", "sales"],
"train": "train.csv",
"validation": "validation.csv",
"selection_metric": "macro_accuracy"
}
]
}Paths are relative to the config file. Use text_column and label_column if your CSV headers differ. Add more entries to tasks to fine-tune one shared model on several independent option sets. Task names must be unique. macro_accuracy is the default checkpoint metric; accuracy is also available.
For a label such as other or out of scope that needs special sampling and validation weight, use selection_metric: "focus_weighted_accuracy" and add:
"focus": {
"label": "other",
"sampling_rate": 0.1,
"validation_weight": 0.2
}This example samples the focus label in 10% of that task's training examples and gives it 20% of that task's checkpoint score. Choose these values from the needs of your application, not from test results.
Validate the config and inspect the task before downloading the source checkpoint:
uv run laya-ft prepare --config data/my-classifier/config.json
uv run laya-ft inspect --config data/my-classifier/config.jsonRun a one-update pilot. On the first run, the pinned source checkpoint is downloaded. Subsequent runs can use --offline.
uv run laya-ft pilot \
--config data/my-classifier/config.json \
--batch-size 4 \
--accumulation 1Then train. --epochs 1 computes enough optimizer updates to sample approximately one pass through the largest task; it is a starting point, not a guarantee of optimal accuracy. You may specify --updates instead. By default, validation runs about ten times and evaluates the complete validation split. The best checkpoint is selected by the mean of each task's configured metric.
uv run laya-ft train \
--config data/my-classifier/config.json \
--epochs 1 \
--batch-size 4 \
--accumulation 1 \
--offlineOn a Mac, prefix the training command with caffeinate -i and keep the lid open. Use --device cuda, --device mps, or --device cpu to override automatic selection. --source-checkpoint PATH selects a local compatible Laya checkpoint instead of the pinned default.
The saved model is checkpoints/<timestamp>/best. Its validation.json records the selected step and per-task metrics. inference.json stores task names, labels, and instructions so the saved model can be used without the training CSV files. run.json records the source checkpoint hash, dataset file hashes, labels, instructions, sampling settings, and training parameters. Dataset files, logs, models, and checkpoints are excluded from Git. Report results from an untouched test split only after checkpoint selection. Probabilities use neutral temperature and are not calibrated.
uv run laya-ft predict \
--checkpoint checkpoints/<timestamp>/best \
--text "The app crashes when I log in"With several tasks, add --task TASK_NAME. New checkpoints carry their task configuration in inference.json, so the training CSV files are not needed for prediction. For an older checkpoint without this file, run uv run laya-ft attach-inference-config --checkpoint checkpoints/<timestamp>/best --config data/my-classifier/config.json while the original data is still available. Omitting --text reads from standard input, which is preferable for sensitive text that should not appear in shell history. Predictions run locally; no remote inference API is called.
For Core ML deployment, set --max-options to at least the largest task's label count. Exports with fewer than 32 options caused a runtime failure in our local test, so use at least 32 even for small classifiers:
uv run --with 'laya-coreml[convert]==0.1.0' laya-coreml convert \
checkpoints/<timestamp>/best \
models/my-classifier \
--max-length 1024 \
--max-options 32
uv run --with 'laya-coreml==0.1.0' python -m laya_ft verify \
--config data/my-classifier/config.json \
--checkpoint checkpoints/<timestamp>/best \
--coreml models/my-classifierIncrease --max-options when a task has more than 32 labels. The verification command compares PyTorch and Core ML choices on validation data. Long instructions or many labels may exceed the fixed 896-token question-and-options budget; the encoder rejects truncated option markers instead of silently training on fewer choices.
Omit --config to use the pinned Banking77, ArBanking77, and CLINC150 training preset. This is separate from the zero-shot Laya–Jev benchmark. A fine-tuned result must not be called zero-shot.
uv run laya-ft prepare
uv run laya-ft inspect
uv run laya-ft pilot --updates 1 --batch-size 4 --accumulation 3
caffeinate -i uv run laya-ft train \
--updates 5400 \
--batch-size 4 \
--accumulation 3 \
--validation-interval 500 \
--offlineThe preset uses official training and validation data only. Banking77 has a deterministic, label-stratified 10% validation split from its official training file. ArBanking77 uses the official MSA and Palestinian training and validation files. CLINC150 uses train plus oos_train, and val plus oos_val. Three CLINC150 training rows matching validation text are removed. Ten exact test-text overlaps are excluded by committed SHA-256 hashes; test text itself is not loaded by the trainer. CLINC150's out of scope label has a 10% training sampling rate and contributes 20% to that task's checkpoint score. Other tasks use accuracy; all three tasks have equal weight in checkpoint selection.
The preset uses a 1,024-token context and an 896-token question-and-options budget. The earlier zero-shot benchmark used a 256-token question-and-options budget, so differences from its results reflect both fine-tuning and this input-budget change. The multilingual source model is pinned to revision 052592a15d198d9ad47da779604259b10b47b7aa.
The HWU64 test set has 64 English assistant intents under a new label taxonomy. It tests transfer to a new label set, not the share of training responsible for the earlier Banking77, ArBanking77, or CLINC150 gains. The source is pinned to a Git revision and both downloaded files are checksum-verified. The generated CSV is excluded from Git.
uv run laya-ft prepare-hwu64
uv run laya-ft compare \
--test-csv data/external/hwu64/test.csv \
--checkpoint checkpoints/<run-id>/best \
--limit 300 \
--offline \
--check-preset-overlap \
--output results/hwu64-300.jsonReplace <run-id> with your selected checkpoint directory. Use --all instead of --limit 300 for the complete test set. Duplicate texts and texts with conflicting labels are removed. Exact text matches with the preset training or validation splits are excluded and counted in the result; this check requires the preset datasets to be available locally. Both models run through the same PyTorch encoder, choice labels, 1,024-token context, 896-token question-and-options budget, and selected examples. The command reports paired accuracy, the accuracy difference, a paired bootstrap interval, and an exact McNemar p-value. For another dataset, supply a UTF-8 CSV with text and label columns; each distinct label becomes an option. The default test instruction is suitable for intent classification and can be changed with --instruction.
HWU64 originates from the NLU Evaluation Data corpus, released under CC BY 4.0. This preset uses the test split distributed by Few-Shot-Intent-Detection.
For an independent banking-domain evaluation, the MInDS-14 preset downloads only text and intent labels for the en-US, en-GB, and en-AU varieties. It uses the dataset viewer API, verifies the repository revision and generated CSV checksum, and does not download audio. MInDS-14 has no official test split for these varieties; its published train split is an evaluation-only corpus here because it was not used in this project's training or checkpoint selection. Its license is CC BY 4.0.
uv run laya-ft prepare-minds14
uv run laya-ft compare \
--test-csv data/external/minds14/english.csv \
--checkpoint checkpoints/<run-id>/best \
--instruction 'Choose the single banking intent that best matches the customer request.' \
--all \
--offline \
--check-preset-overlap \
--output results/minds14-english-full.jsonuv run ruff check .
uv run ruff format --check .
uv run python -m unittest discover -s testsThe code is MIT-licensed. The model weights and datasets retain their own licenses and are not included in this repository: Laya, Laya-CoreML, Banking77, ArBanking77, and CLINC150.