Magenta RealTime 2 comprises three components:
- MusicCoCa: a text / audio style embedding model
- SpectroStream: a codec model for audio encoding and decoding
- Depthformer: a transformer-based model that generates SpectroStream tokens
MRT2 offers two model sizes:
mrt2_small(230M parameters) β runs real-time on any Apple Silicon Mac, including Air models.mrt2_base(2.4B parameters) β higher quality; requires a Pro/Max chip for real-time streaming.
The table below shows which devices support real-time streaming (generating audio faster than playback):
| Device | mrt2_small (230M) |
mrt2_base (2.4B) |
|---|---|---|
| M5 Max | β | β |
| M3 Max | β | β |
| M2 Max | β | β |
| M4 Pro | β | β |
| M2 Pro | β | β |
| M1 Pro | β | β |
| M4 Air | β | β |
| M3 Air | β | β |
| M1 Air | β | β |
Use the mrt models CLI to fetch models (automatically saved in ~/Documents/Magenta/magenta-rt-v2/)
# Download resource models:
# MusicCoCa and SpectroStream
mrt models init
# Download Depthformer models
# exported in mlxfn format
mrt models downloadExpected directory layout:
~/Documents/Magenta/magenta-rt-v2/
βββ resources/
β βββ musiccoca/
β β βββ audio_preprocessor.tflite
β β βββ music_encoder.tflite
β β βββ pretrained_vector_quantizer.tflite
β β βββ spm.model
β β βββ text_encoder.tflite
β βββ spectrostream/
β βββ spectrostream_encoder.mlxfn
β βββ decoder.safetensors
β βββ encoder.safetensors
β βββ quantizer.safetensors
βββ models/
β βββ <model_name>/
β βββ <model_name>.mlxfn
β βββ <model_name>_state.safetensors
βββ checkpoints/
βββ <model_name>.safetensors
One may find raw model checkpoints (safetensors before exporting to mlxfn) useful for research purposes. Thus we offer them via mrt checkpoints download CLI.