A practical CLI for converting Hugging Face models to MLX on Apple Silicon.
It accepts either a Hugging Face URL or namespace/repo, detects what is in the repo, runs the safest available conversion path, and can optionally upload the MLX output back to Hugging Face.
- Converts standard Hugging Face checkpoints (
.safetensors) to MLX. - Accepts GGUF repos and tries to resolve the original base model automatically.
- Supports a dry-run mode so you can verify the route before downloading large files.
- Optionally publishes the output as a new Hugging Face model repo.
- Detect source format from repo files.
- If
.safetensorsexists: convert directly. - If only
.ggufexists: trybase_modelresolution from model card/tags/README. - If unresolved: require
--base-repoor--experimental-gguf-direct.
python -m venv .venv
source .venv/bin/activate
pip install -e '.[dev,mlx,gguf]'Dry-run first (recommended):
hf-mlx --source Qwen/Qwen2.5-0.5B-Instruct --dry-runConvert to local MLX output:
hf-mlx \
--source Qwen/Qwen2.5-0.5B-Instruct \
--hf-token "$HF_TOKEN" \
--output ./artifacts/qwen2.5-0.5bConvert and upload:
hf-mlx \
--source Qwen/Qwen2.5-0.5B-Instruct \
--hf-token "$HF_TOKEN" \
--output ./artifacts/qwen2.5-0.5b \
--upload \
--upload-repo your-user/Qwen2.5-0.5B-Instruct-MLX- Run with
--dry-run. - Run conversion without upload and validate local inference.
- Upload only after local verification.
Prompt used during local smoke test:
Answer in one sentence: Istanbul is in which country?
Observed model response:
Istanbul is located in Turkey.
Published output from this repo:
Default working cache:
.hf_mlx_work/download/<namespace--repo>
Default MLX output:
./artifacts/mlx_model(default)- If you pass
--output /path/to/out, output is written to/path/to/out/mlx_model.
Both paths are configurable with --working-dir and --output.
- This tool does not write your token to project files.
- Direct GGUF to MLX conversion is still experimental.
- Production path is base-model safetensors fallback.
Run tests:
python -m pytest -q