Models
The models each module uses, their sizes and licenses, and how downloads work
TopLocal Studio does not bundle any models. You download only the models you want, from inside the app. Sizes below are approximate.
Image
| Model | Size | Good for | License |
|---|---|---|---|
| Z-Image-Turbo | 5.5 GB | Fast text-to-image | Apache-2.0 |
| FLUX.2 klein 4B | 4.3 GB | Text-to-image and editing on machines with less memory | Apache-2.0 |
| FLUX.2 klein 9B | 8.9 GB | Text-to-image and editing with more detail | FLUX Non-Commercial |
Video
| Model | Size | Good for | License |
|---|---|---|---|
| LTX-2.5 Distilled | about 30 GB | Text or image to short video with sound (3/5/8 s, 480p/720p). Needs 24 GB+ memory. | LTX-2.x Community License |
Music
| Model | Size | Good for | License |
|---|---|---|---|
| ACE-Step 1.5 | 6.2 GB | Full songs from a description, with your own lyrics or as instrumentals | MIT |
| YuE2 | 4.6 GB | Songs with Chinese vocals | CC BY-NC 4.0 (non-commercial) |
| Qwen3-4B (lyric writer) | 2.5 GB | Drafting lyrics for your songs | Apache-2.0 |
Speech
| Model | Size | Good for | License |
|---|---|---|---|
| Qwen3-ASR 0.6B | 1.2 GB | Transcription to TXT and SRT, smaller and lighter | Apache-2.0 |
| Qwen3-ASR 1.7B | 2.5 GB | Transcription to TXT and SRT, larger model | Apache-2.0 |
| SenseVoice | 0.25 GB | Transcription with a very small download | FunASR Model License |
| Kokoro | 0.2 GB | Text-to-speech with a very small download | Apache-2.0 |
| Qwen3-TTS | 2 GB | Text-to-speech, including voice cloning | Apache-2.0 |
Licenses and your results
The app's code is licensed under Apache-2.0. Each model has its own license, and that license applies to what you create with the model. The app shows each model's license before you download it.
Some licenses do not allow commercial use, including FLUX Non-Commercial (FLUX.2 klein 9B) and CC BY-NC 4.0 (YuE2). If you plan to use your results commercially, choose models whose licenses allow it and read the full license text.
How downloads work
- Models download from Hugging Face, or from the mirror if you turn it on.
- Downloads can be resumed if they are interrupted.
- Every downloaded file is checked with sha256 before it is used.
- Models are stored in the app's data folder. See Getting Started for the location.
Hugging Face token for the video model
Only the LTX-2.5 video model needs a Hugging Face access token. All other models download without one.
To get a token, sign in at huggingface.co, create an access token in your account settings, and enter it in TopLocal Studio before downloading the video model.
China mirror
If you are in mainland China, downloads from Hugging Face may be slow or fail. Turn on the China mirror in the app's settings to download models from hf-mirror.com instead.