Qwen3.8-Flash-Next
Hugging Face (open weights) Open weights
The catch: The card calls it an experimental preview of the Qwen4 architecture, and anyone running an API business or an AI coding/office product needs a separate licence from Qwen for commercial use, with no revenue threshold.
| Access | Open weights |
|---|---|
| Free limits | Free to download and self-host; you pay only for your own compute. 125B total / 6B active (MoE) plus a 51B n-gram embedding, so it needs multi-GPU. 262,144-token native context, extensible to 1M with YaRN. Released 24 Aug 2026 as a preview of the Qwen4 architecture. |
| Modality | text, code, image, video |
| Credit card | Not required |
| Commercial use | Allowed |
| Trains on your data | – |
| Licence or terms | Qwen Community License 1.0 |
| Context window | 262,144 tokens |
| Model id | Qwen/Qwen3.8-Flash-Next |
| Base URL | – |
| Last checked | 8 Oct 2026 |
How to use Qwen3.8-Flash-Next
Download the weights and serve them yourself; the quickstart below shows one way. Licence: Qwen Community License 1.0.
vllm serve Qwen/Qwen3.8-Flash-Next # then: curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"Qwen/Qwen3.8-Flash-Next","messages":[{"role":"user","content":"Hello"}]}'I want to run the open-weight model "Qwen3.8-Flash-Next" (Qwen/Qwen3.8-Flash-Next) for my project. Please: 1. Check the model card and tell me the hardware it needs and the simplest way to serve it (vLLM, SGLang, Ollama or a hosted endpoint). 2. Give me a minimal working setup and a test request. 3. Tell me the licence terms for commercial use. 4. If a cheaper or smaller variant would do the job, mention it.
Before you ship on it
- The card calls it an experimental preview of the Qwen4 architecture, and anyone running an API business or an AI coding/office product needs a separate licence from Qwen for commercial use, with no revenue threshold.
- Qwen Community License 1.0 (LICENSE in the HF repo): MIT-style grant including sell, host and fine-tune. Clause 1: products over 100,000,000 MAU or US$20,000,000 monthly revenue must prominently show the model name.…
- Free to run means you pay for the GPUs, the serving stack and the on-call.
Questions
Is Qwen3.8-Flash-Next free?
Yes. Open weights offers it as open weights, with these limits: Self-host free · Qwen Community License 1.0. No credit card is needed.
Can I use Qwen3.8-Flash-Next commercially?
Yes, under Qwen Community License 1.0. Check the current terms before you ship.
How do I start with Qwen3.8-Flash-Next?
Download the weights and serve them yourself; the quickstart below shows one way. Licence: Qwen Community License 1.0.
Source: huggingface.co/Qwen/Qwen3.8-Flash-Next, checked 8 Oct 2026.