Skip to content

Qwen3.8-Flash-Next

Hugging Face (open weights) Open weights

The catch: The card calls it an experimental preview of the Qwen4 architecture, and anyone running an API business or an AI coding/office product needs a separate licence from Qwen for commercial use, with no revenue threshold.

AccessOpen weights
Free limitsFree to download and self-host; you pay only for your own compute. 125B total / 6B active (MoE) plus a 51B n-gram embedding, so it needs multi-GPU. 262,144-token native context, extensible to 1M with YaRN. Released 24 Aug 2026 as a preview of the Qwen4 architecture.
Modalitytext, code, image, video
Credit cardNot required
Commercial useAllowed
Trains on your data–
Licence or termsQwen Community License 1.0
Context window262,144 tokens
Model idQwen/Qwen3.8-Flash-Next
Base URL–
Last checked8 Oct 2026

How to use Qwen3.8-Flash-Next

Download the weights and serve them yourself; the quickstart below shows one way. Licence: Qwen Community License 1.0.

Quickstart
vllm serve Qwen/Qwen3.8-Flash-Next  # then: curl http://localhost:8000/v1/chat/completions -H "Content-Type: application/json" -d '{"model":"Qwen/Qwen3.8-Flash-Next","messages":[{"role":"user","content":"Hello"}]}'
Paste into Claude Code
I want to run the open-weight model "Qwen3.8-Flash-Next" (Qwen/Qwen3.8-Flash-Next) for my project.
Please:
1. Check the model card and tell me the hardware it needs and the simplest way to serve it (vLLM, SGLang, Ollama or a hosted endpoint).
2. Give me a minimal working setup and a test request.
3. Tell me the licence terms for commercial use.
4. If a cheaper or smaller variant would do the job, mention it.

Before you ship on it

  • The card calls it an experimental preview of the Qwen4 architecture, and anyone running an API business or an AI coding/office product needs a separate licence from Qwen for commercial use, with no revenue threshold.
  • Qwen Community License 1.0 (LICENSE in the HF repo): MIT-style grant including sell, host and fine-tune. Clause 1: products over 100,000,000 MAU or US$20,000,000 monthly revenue must prominently show the model name.…
  • Free to run means you pay for the GPUs, the serving stack and the on-call.

Questions

Is Qwen3.8-Flash-Next free?

Yes. Open weights offers it as open weights, with these limits: Self-host free · Qwen Community License 1.0. No credit card is needed.

Can I use Qwen3.8-Flash-Next commercially?

Yes, under Qwen Community License 1.0. Check the current terms before you ship.

How do I start with Qwen3.8-Flash-Next?

Download the weights and serve them yourself; the quickstart below shows one way. Licence: Qwen Community License 1.0.

Source: huggingface.co/Qwen/Qwen3.8-Flash-Next, checked 8 Oct 2026.