
Qwen3.8-Flash-Next - strixw (for strixite on AMD Strix Halo)
| on one AMD Strix Halo, 128 GB, with strixite | |
|---|---|
| Decode, real agent use | ~47-50 tokens/s, flat from 0 to 492k context |
| Context | 512k tokens |
| Prefill | ~1,370 tokens/s over a real 490k-token conversation |
| Agent tasks | 19 / 19 on terminal-bench-mini's core suite |
These are the weights of Qwen/Qwen3.8-Flash-Next, converted to strixw, the weight format of strixite - an inference engine I wrote from scratch in HIP for one chip, AMD's Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S, gfx1151). They are exactly the files I serve every day.
Only strixite can load these files. strixw is not GGUF or safetensors: every tensor is already in the layout of the kernel that reads it. Transformers, llama.cpp, vLLM and others can't use them - for those, start from the original checkpoint.
strixite's source code is coming soon to github.com/shawnshekari/strixite.
Platform: strixite runs on Linux with an AMD Strix Halo (gfx1151) and 128 GB of memory, tested on Fedora 43. It's built and tested by one person on one machine, so that's the only setup I can promise works; Windows and macOS aren't supported, and I don't plan to add them.
What's in it
The download is ~115 GiB instead of the original checkpoint's ~360 GB, and needs no conversion step.
| file | size | what |
|---|---|---|
converted/Qwen3.8-Flash-Next.U-gdn_in-g128/weights.strixw |
65.9 GiB | the model: 4-bit routed experts, 8-bit dense projections, MTP head included |
converted/Qwen3.8-Flash-Next.ngram-q8/ngram.table |
48.3 GiB | the per-layer n-gram embedding table (int8 rows); read from SSD on demand, not loaded into memory |
Qwen3.8-Flash-Next/tokenizer.json, generation_config.json, config.json |
small | copied unchanged from the original checkpoint |
Qwen3.8-Flash-Next/LICENSE, LICENSE |
small | the Qwen Community License 1.0 |
sha256.txt |
small | checksums of every file above |
The folder layout matches strixite's default model directory (~/models/strix-infer - the engine's working name),
so downloading there needs no configuration:
hf download wemoh/Qwen3.8-Flash-Next-strixw --local-dir ~/models/strix-infer
cd ~/models/strix-infer && sha256sum -c sha256.txt
Quantization
Asymmetric integer quantization per group along the input dimension (a BF16 scale and a BF16 minimum per group), chosen class by class from a quantization study against the FP32 reference outputs:
| weights | format |
|---|---|
| routed experts (gate / up / down), GDN input projections | 4-bit, groups of 128 |
| every other dense projection: attention q / k / v / o, GDN output, shared expert, hyper-connection mixes, PLE key / value, LM head | 8-bit, groups of 64 |
| MTP head's input projection | 4-bit, groups of 64 |
| router | BF16 (its logits are computed in FP32) |
| token embedding | BF16 |
| n-gram table | int8 per row, with a per-row scale |
| norms, gates, conv kernels, small vectors | FP32 / BF16 as in the checkpoint |
Most of the model's parameters are the 512 routed experts, of which each token reads only a few; the dense projections are read for every token. Keeping the dense path at 8 bits and the experts at 4 is what keeps accuracy close to the original at ~4.25 bits per routed weight.
The vision tower isn't included: strixite is text only.
Accuracy and speed
All measured on my machine (Strix Halo, 128 GB) with strixite, 2026-09-28 .. 2026-10-03.
- Accuracy vs the unquantized model: over 242 positions of three reference prompts, mean KL divergence from the FP32 outputs 0.028, and the top-1 token agrees at 226 of 242 positions.
- Agent tasks: terminal-bench-mini's core19 suite, 19 of 19 tasks passed.
- Decode: ~33 tokens/s plain and 45-55 tokens/s with the built-in multi-token prediction (MTP) at 2k-131k context; in real agentic use it stays at ~47-50 tokens/s out to 492k tokens of context.
- Prefill: ~1,370 tokens/s over a real 489,566-token agent conversation.
- Context: up to 512k tokens (YaRN factor 2 in the engine; the model was trained to 262,144).
Memory and storage
- ~66 GiB of weights resident in GPU memory (Strix Halo's unified memory - GTT, no VRAM carve-out needed), plus up to ~13 GiB of KV cache at 512k context.
- The n-gram table (48.3 GiB) stays on disk and is read row by row as tokens need it, through a cache in RAM: put it on a fast NVMe drive.
How it was made
- Source: Qwen/Qwen3.8-Flash-Next at revision
de4b8e4d43b917e7706784d8bb445c9af86a3540(BF16 safetensors). weights.strixw: strixite's converter (convert_qwen4exp, commiteb57cdbof my private development repo), strixw format version 4, layoutexp_gu=q4g128,exp_down=q4g128,sh_gu=q8g64,sh_down=q8g64,gdn_in=q4g128,gdn_out=q8g64,attn_qkv=q8g64,attn_o=q8g64,hc_down=q8g64,hc_up=q8g64,ple_kv=q8g64,lm_head=q8g64,mtp_fc=q4g64,embed=bf16.ngram.table: strixite'sconvert_ngram_table(the checkpoint's 128 table shards into one BF16 file), thentranscode_ngram_table --dtype q8(int8 rows with a BF16 per-row scale; median row error 0.66% of the BF16 row).- The converters will be part of the strixite release, so anyone can rebuild these files from the original checkpoint and check them.
strixite started as "how hard can it be?" It turned out an engine that holds its own against a closed-source one - faster at decode, on its home turf - is within reach of one person, with a good model, a good coding assistant, and a lot of patience. I hope it encourages you to grab your favorite LLM and framework and build something from scratch.
Limitations
- Only for strixite, and only for Strix Halo (gfx1151) with 128 GB.
- Quantization moves the outputs slightly away from the original model (numbers above); with sampling on, text differs from what the BF16 model would write.
- Text only - the original model's vision input isn't supported.
- Everything the original model card says about the model's own behaviour, limits and intended use applies here too.
License
These weights are a derivative of Qwen3.8-Flash-Next and are distributed under the
Qwen Community License 1.0, the same license as the original - its full text is in LICENSE. Note its
terms for commercial use: a "Model as a Service" or "AI Work Assistant" business needs a separate license from
Qwen, and products above 100 million monthly active users or US$20 million monthly revenue must display the model
name. See the license text for the exact wording.
Copyright for the original model belongs to Qwen. The conversion is mine; the strixite code is a separate project under its own license (AGPL-3.0).
Model tree for wemoh/Qwen3.8-Flash-Next-strixw
Base model
Qwen/Qwen3.8-Flash-Next