Instructions to use AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm") model = AutoModelForMultimodalLM.from_pretrained("AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm
- SGLang
How to use AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm with Docker Model Runner:
docker model run hf.co/AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm
- Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm
- At a glance
- What this is
- Target hardware
- Quickstart
- Step 0: Which setup do you have? (1 minute)
- Common setup (every recipe)
- 2x R9700: pick a recipe
- Verify (every recipe)
- Recipe A: 2x R9700 with the AEON DFlash2 drafter (primary)
- Recipe A-MTP: 2x R9700 with MTP (alternative)
- Recipe A-plain: 2x R9700 without speculation (exact baseline)
- Recipe Z: 2x R9700 as two independent servers (never-fails fallback)
- Tester matrix: plain vs MTP vs DFlash2
- Recipe A-Max: 2x R9700, BF16-resident MLP (experimental, superseded)
- Speed experiment: radiance (experimental, unlicensed, untested on this checkpoint)
- TensorFold
- Recipe K: 4x Radeon AI PRO R9700 with the AEON DFlash2 drafter
- Recipe B: 1x Radeon AI PRO R9700
- Recipe C: 2x 16 GB RDNA 4 (RX 9070 XT, RX 9070, RX 9060 XT)
- Recipe D: 4x 16 GB cards (RDNA 4 or RDNA 3)
- Recipe E: 2x Radeon RX 7900 XTX
- Recipe M: 2x Radeon RX 7900 XT
- Recipe L: 2x 16 GB RDNA 3 (RX 7900 GRE, RX 7800 XT, RX 7700, PRO W7700)
- Recipe F: Radeon PRO W7900 (48 GB)
- Recipe Q: 2x Radeon PRO W7900 with the AEON DFlash2 drafter
- Recipe G: Radeon PRO W7800 (32 GB)
- Recipe P: Radeon PRO V710
- Recipe H: Strix Halo (Ryzen AI Max+ 395), 64 GB
- Recipe H-128: Strix Halo (Ryzen AI Max+ 395), 128 GB
- Recipe I: Instinct MI300X / MI325X
- Recipe O: Instinct MI300A
- Recipe J: Instinct MI350X / MI355X
- Recipe N: Instinct MI210 / MI250 / MI250X
- Not supported
- Vision (Instinct and RDNA 3 only)
- Troubleshooting
- Expected performance (estimates, not measurements)
- Expected behavior and limits
- Quality and validation status
- Quantization recipe
- Model family
- User responsibility
- License and attribution
- At a glance
Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm
The AEON Ultimate uncensored 27B, quantized to OCP Microscaling (MX) formats for AMD GPUs.
The MLP runs in MXFP4, the Gated DeltaNet projections in MXFP6, and everything that carries the model's skills stays in BF16.
The checkpoint is 23.7 GB on disk and loads as about 22.8 GB of weights. Two 32 GB Radeon AI PRO R9700 cards hold it with the full 262,144-token context; one R9700 holds it with a shorter context. MXFP4 is also the format Instinct MI350 / MI355 execute natively.
The Hub's file panel labels this repo "8-bit" only because the MX weights are stored packed in U8 tensors; the real mix is MXFP4 MLP + MXFP6 Gated DeltaNet projections + BF16 for the rest, about 6.8 bits per weight on average: a 23.7 GB download that needs at least 32 GB of total VRAM.
Recommended on 2x R9700: vLLM with tensor parallelism across both cards, this checkpoint unchanged, and the bundled AEON DFlash2 drafter for lossless speculative decoding (Recipe A).
Experimental: early access. Nobody has published runtime results for this build yet. The first Radeon test pass (2x R9700) is under way, and quality and speed numbers will appear in this card as they come in. Requests are reviewed by hand, usually within a few hours; include the community access word in the form if you were given one.
At a glance
Why it's great
- Quality first. Only two parts of the model are quantized: the MLP (MXFP4, W4A4) and the Gated DeltaNet projections (MXFP6, W6A6). Full attention, the recurrence, embeddings,
lm_head, the vision tower and the MTP head all stay in BF16. That is about 6.8 bits per weight: 23.7 GB, against 55.6 GB for the master. - Full context on two 32 GB cards. 2x Radeon AI PRO R9700 hold it with the full 262,144-token context. Instinct MI350X / MI355X can also run the MXFP4 MLP natively, as an opt-in.
- Speculative decoding ships with it. The repo includes the AEON DFlash2 drafter (
dflash2/, BF16) alongside the model's own MTP head, and both are lossless. In the reference run, DFlash2's mean acceptance length was 3.73 at T=0.6, and it ran 2.7x faster than plain decoding on the same 4 prompts (details). - Turnkey across the ROCm lineup. There is a launch script per setup in
scripts/, anddetect_setup.pypicks the right one for your GPUs. The MXFP6 dequant-once patch gives the same output as stock vLLM with a much faster forward pass. - AEON Ultimate, uncensored. It is the full 27B hybrid (thinking, tool calling, and vision on Instinct and RDNA 3: see Vision), and it writes what the base model refuses. Read User responsibility.
Target systems (all Untested on AMD hardware so far)
| Setup | Recipe |
|---|---|
| 2x Radeon AI PRO R9700 (featured) | A (DFlash2) → A-MTP → A-plain → Z |
| 4x / 1x R9700 (R9700S, R9600D) | K (DFlash2) · B |
| RDNA 4 16 GB (RX 9070 XT / 9070 / 9060 XT), 2x or 4x | C · D |
| RDNA 3: 2x RX 7900 XTX, 2x RX 7900 XT, 2x 16 GB | E · M · L |
| Radeon PRO W7900 / W7800 48 GB, 2x W7900, W7800 32 GB, V710 | F · Q (DFlash2) · G · P |
| Strix Halo (Ryzen AI Max+ 395) 64 / 128 GB | H · H-128 |
| Instinct MI300X / MI325X, MI300A, MI350X / MI355X, MI210 / MI250 | I · O · J · N |
Recommended configuration
2x R9700: Recipe A. Launch it with
bash "$AEON_DIR/scripts/serve_2x_r9700_dflash2.sh". It runs:- vLLM
v0.31.0(ROCm image) at TP=2, with the bundled drafter at K=9 and the lattice[[1,2,9],[3,8,7],[9,12,6],[13,16,4]]; TRITON_ATTN, including inside--speculative-config, withVLLM_USE_V2_MODEL_RUNNER=1and prefix caching off.
The estimate is about 35 tok/s single-stream, ±50% until measured.
- vLLM
Long multi-turn agent loops: A-MTP keeps prefix caching on.
Exact baseline, or the most concurrent requests: A-plain.
TP=2 hangs: Z.
Any other setup: run
python3 "$AEON_DIR/scripts/detect_setup.py".
Using the bundled DFlash2 drafter with other engines
dflash2/ is a standard DFlash2 checkpoint. It has the same 81 tensors (names, shapes and dtypes) as z-lab/Qwen3.8-27B-DFlash2. Its config differs only in block_size 10, causal: false and draft_vocab_size. By format, any engine that runs z-lab's drafter loads it if you point its drafter path at dflash2/ (only vLLM has been run with an AEON target, and not yet on AMD). Set the draft width to 9: num_speculative_tokens 9 in vLLM, --speculative-num-draft-tokens 10 in SGLang (the block size), or --spec-draft-n-max 9 in llama.cpp. For this mixed checkpoint, vLLM is the only engine that loads the target unchanged (see TensorFold). Full guide: dflash2/README.md.
What this is
Base model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 is AEON-7's uncensored build of Qwen/Qwen3.8-27B. It is a hybrid of 48 Gated DeltaNet (linear-attention) layers and 16 full-attention layers, with a vision tower, an MTP head and a native context of 262,144 tokens.
Format: OCP Microscaling. Every block of 32 values shares one power-of-two (E8M0) scale. We used AMD Quark 0.12.post1 and exported real packed weights, not fake-quant.
- MLP (
gate_proj,up_proj,down_proj) is MXFP4, W4A4: FP4 E2M1 weights, with activations quantized to MXFP4 on the fly. - Gated DeltaNet projections (
in_proj_qkv,in_proj_z,out_proj) are MXFP6 E2M3, W6A6. - The rest stays BF16: full attention, embeddings,
lm_head, the vision tower, the MTP head, the GDN recurrence (conv1d,in_proj_a/b,A_log,dt_bias) and the norms. Precision lattice (Quark ROCm): MLP MXFP4 - GDN writers MXFP6 - attention / vision / MTP / embed / lm_head / GDN recurrence BF16.
- MLP (
Unchanged from the BF16 master: the tokenizer, chat template and vision processor. Thinking is on by default.
Why MX on AMD? MX is the open block-scaled format that AMD's newest accelerators run in hardware. On Radeon the gain is footprint: the BF16 master is 55.6 GB, while this build is 23.7 GB. The precision is mixed on purpose. The MLP and GDN projections take 4-bit and 6-bit weights, and the attention path, the recurrence, the embeddings and the vision tower stay in full BF16.
Target hardware
Every AMD GPU configuration that runs ROCm compute and can hold this checkpoint has its own labelled quickstart below. The featured setup is 2x Radeon AI PRO R9700 (Recipe A and its fallbacks).
- Native MXFP4: Instinct MI350X / MI355X (gfx950) can run the MLP on AITER's W4A4 GEMM. That path uses the same MX grid but is not bit-identical to the other recipes, so Recipe J defaults to the exact emulated path and offers native MXFP4 as a labelled opt-in. MXFP6 is emulated on every GPU, because vLLM has no MXFP6 kernel yet.
- Emulated: everything else, including Radeon RDNA 4 (R9700, RX 9070 / 9060 XT), RDNA 3 (RX 7900 / 7800 / 7700, W7900, W7800, W7700, V710), Strix Halo and Instinct MI300X / MI325X / MI300A / MI250 / MI210. vLLM keeps the packed MX weights in memory, dequantizes each layer to BF16 on every forward pass, quantizes and dequantizes the activations to the same MX grid, and runs the matmul in BF16. The numerics match W4A4 and W6A6 and memory stays near the packed size, but decoding is slower than a native kernel would be (vLLM 0.29 and 0.31 source:
QuarkOCP_MX,EmulationMxfp4LinearKernel,EmulationMxfp6LinearKernel; native MX needssupports_mx(), which is true only on gfx95x and gfx1250). Wherever memory allows, the recipes remove most of that cost without changing a bit of the output: the MXFP6 weights are dequantized once at load (a small vLLM patch inscripts/patches/) and, on the largest GPUs, the MXFP4 MLP is kept in BF16 after load (VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD).
| Other platform | Runs? | Notes |
|---|---|---|
| NVIDIA GPUs | Not supported by this build | Use NVFP4-MIXED. |
| Apple Silicon | Not with this checkpoint | An MLX conversion is being tested; results will be added here. |
Quickstart
Deploying with an AI agent (Claude Code, Codex, Cursor)? Point it at
AGENTS.md. It has the preflight, every recipe with exact commands, verification steps and an exhaustive troubleshooting table.Testing before an event? Run
bash "$AEON_DIR/scripts/preflight_smoke.sh"the day before. It checks the host, starts Recipe A (falling back to A-MTP, A-plain, then Z; a TP=2 start that hangs goes straight to Z), smoke-tests it in a few minutes and bundles the results to send back.mkdir -p ~/aeon-test && cd ~/aeon-test && bash "$AEON_DIR/scripts/run_matrix.sh"then compares plain, MTP and DFlash2 at 1–16 concurrent requests and checks TP=2 against a one-GPU reference.
Step 0: Which setup do you have? (1 minute)
Run this on the Linux host. It reads the amdgpu driver directly, so it needs no ROCm tools, no Docker and no root:
# one line per AMD GPU: gfx target, VRAM, render node, PCI name
for p in /sys/class/kfd/kfd/topology/nodes/*/properties; do
v=$(awk '$1=="gfx_target_version"{print $2}' "$p"); [ "${v:-0}" -gt 0 ] || continue
r=$(awk '$1=="drm_render_minor"{print $2}' "$p"); d=/sys/class/drm/renderD$r/device
printf 'gfx%d%x%x %5.1f GiB /dev/dri/renderD%s %s\n' $((v/10000)) $((v/100%100)) $((v%100)) \
"$(awk '{print $1/2^30}' "$d/mem_info_vram_total")" "$r" \
"$(lspci -s "$(basename "$(readlink -f "$d")")" 2>/dev/null | cut -d' ' -f2-)"
done
With ROCm tools on the host, amd-smi static --asic --vram or rocm-smi --showproductname --showmeminfo vram show the same thing. After the download, python3 "$AEON_DIR/scripts/detect_setup.py" does this lookup and prints the matching recipe and launch script, and bash "$AEON_DIR/scripts/preflight.sh" checks the rest of the host (driver, groups, PCIe slots, IOMMU, Resizable BAR, disk, ports).
Match the output to the table below. Count only the discrete GPUs.
- Ignore an integrated GPU (gfx1036, gfx1035, gfx1103, gfx1150 and similar, with a few GiB of VRAM), unless it gets passed into the container. vLLM reads the GPU architecture from the first GPU that
amdsmilists, andHIP_VISIBLE_DEVICESdoesn't change that. If the iGPU comes first, the RDNA 4 paths switch off. - To hide it, either disable the iGPU in the BIOS, or pass only the discrete GPUs' nodes into the container: replace
--device /dev/driwith--device /dev/dri/renderD129 --device /dev/dri/card1 ...for each discrete GPU. The launch scripts do this withAEON_DEVICES=auto, and they refuse to start if vLLM would see the wrong architecture.
Pick your system
Status: Verified = run with this checkpoint on that hardware (none yet). Untested = not yet run on that hardware; every recipe stays Untested until the first AMD runs come back. "Expected to work" means the vLLM code path for that architecture and GPU count is confirmed from source, and other Qwen3.x checkpoints are reported running in vLLM on the same GPU family with the same settings. Not supported = it doesn't fit, or vLLM can't run there.
MX path: once = MXFP6 dequantized once at load; at load = MXFP4 MLP also kept in BF16 after load; stock = both dequantized on every forward pass (no room for the BF16 copies). All three give identical output; they differ in speed and memory.
| Your GPUs | VRAM | gfx | MX path | Status | Recipe |
|---|---|---|---|---|---|
| 2x Radeon AI PRO R9700 (featured) | 2 × 32 GB | gfx1201 | once | Untested | A (DFlash2) → A-MTP → A-plain; fallback Z. Which one |
| 4x Radeon AI PRO R9700 / R9700S / R9600D | 4 × 32 GB | gfx1201 | once + at load | Untested | K (DFlash2) |
| 1x Radeon AI PRO R9700 / R9700S / R9600D | 32 GB | gfx1201 | stock (slow) | Untested (expected to work) | B |
| 2x Radeon RX 9070 XT / RX 9070 / RX 9060 XT 16 GB | 2 × 16 GB | gfx1201 / gfx1200 | stock (slow) | Untested (expected to work; tight) | C |
| 4x 16 GB: RX 9070 XT / 9070 / 9060 XT, RX 7900 GRE, RX 7800 XT, RX 7700, Radeon PRO W7700 | 4 × 16 GB | gfx1201 / gfx1200 / gfx1100 / gfx1101 | once | Untested | D |
| 2x Radeon RX 7900 XTX | 2 × 24 GB | gfx1100 | once | Untested | E |
| 2x Radeon RX 7900 XT | 2 × 20 GB | gfx1100 | stock | Untested | M |
| 2x 16 GB RDNA 3: RX 7900 GRE, RX 7800 XT, RX 7700, Radeon PRO W7700 | 2 × 16 GB | gfx1100 / gfx1101 | stock (slow) | Untested (tight) | L |
| Radeon PRO W7900 / W7900 Dual Slot / W7800 48 GB | 48 GB | gfx1100 | once | Untested | F |
| 2x (or 4x) Radeon PRO W7900 / W7800 48 GB; also 2x W7800 32 GB | 2 × 48 GB | gfx1100 | once + at load | Untested | Q (DFlash2) |
| Radeon PRO W7800 32 GB | 32 GB | gfx1100 | stock (slow) | Untested | G |
| Radeon PRO V710 (full GPU) | 28 GB | gfx1101 | stock (slow) | Untested | P |
| Ryzen AI Max+ 395 / Max 390 / Max 385 (Strix Halo), 64 GB | unified | gfx1151 | once | Untested | H |
| Ryzen AI Max+ 395 / Max 390 / Max 385 (Strix Halo), 128 GB | unified | gfx1151 | once + at load | Untested | H-128 |
| Instinct MI300X / MI308X / MI325X | 192 / 256 GB | gfx942 | once + at load | Untested | I |
| Instinct MI300A | 128 GB unified | gfx942 | once + at load | Untested | O |
| Instinct MI350X / MI355X | 288 GB | gfx950 | once + at load (exact); native MXFP4 opt-in | Untested | J |
| Instinct MI210 / MI250 / MI250X | 64 GB per GCD | gfx90a | once | Untested | N |
| Any single GPU under 28 GB, 8 GB cards, RDNA 2, MI100, RX 7600 series, NVIDIA | Not supported | why |
Common setup (every recipe)
You need Linux on bare metal (not a VM with GPU passthrough for multi-GPU recipes: RCCL is reported to hang there), a recent amdgpu driver for your GPU, Docker, and about 100 GB of free disk (model 23.7 GB, drafter 3.85 GB, the image's 11.5 GB download and more once unpacked, kernel caches; about 12 GB more per extra image tag).
Every recipe uses the official vLLM ROCm image vllm/vllm-openai-rocm:v0.31.0 (ROCm 7.2.3; built for gfx90a, gfx942, gfx950, gfx1100, gfx1101, gfx1150, gfx1151, gfx1200 and gfx1201; includes amd-quark).
- Don't use v0.28, v0.29 or v0.30, except through the
AEON_STRIP_ALGO_CONFIG=1escape hatch (AGENTS.md section 5.6). They crash while loading this checkpoint withAttributeError: 'dict' object has no attribute 'endswith', because their Quark loader can't handle the list-valuedalgo_configinconfig.json. v0.31.0 fixes it. - Don't use the
nightly/rocm100images on Radeon. They ship no gfx1201 kernels (ROCm/aiter#5229). - Pin the digest. As of 2026-10-04
vllm/vllm-openai-rocm:v0.31.0issha256:749f6f3f944f12af49966ac523c1f8573e4b229594b541954bd5f879b4496b1f. After the pull,docker image inspect vllm/vllm-openai-rocm:v0.31.0 --format '{{index .RepoDigests 0}}'should show it; if it shows another digest, the tag was re-pushed: tell us before relying on it. Then useAEON_IMAGE=vllm/vllm-openai-rocm@sha256:....
# 1) After you accept the access terms on the model page (log in with your own account and token)
python3 -m venv ~/.venvs/hf && ~/.venvs/hf/bin/pip install -U huggingface_hub && export PATH="$HOME/.venvs/hf/bin:$PATH" # Ubuntu may need: sudo apt install python3-venv
hf auth login
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm \
--local-dir ~/models/aeon-mxfp4-rocm
# the download includes dflash2/ (3.85 GB): the DFlash2 drafter that Recipes A, K and Q use. Skip it with
# --exclude "dflash2/*" only if you will not use DFlash2.
# 2) Variables the commands below use (re-run in every new terminal)
export AEON_DIR="$HOME/models/aeon-mxfp4-rocm"
export AEON_CACHE="$HOME/.cache/aeon-mxfp4-rocm" # compiled kernels persist here
export RENDER_GID="$(getent group render | cut -d: -f3)"; [ -n "$RENDER_GID" ] || RENDER_GID=video
mkdir -p "$AEON_CACHE" ~/aeon-test && cd ~/aeon-test # logs and results land in the current folder
# 3) Check the host, and see which recipe fits
bash "$AEON_DIR/scripts/preflight.sh"
python3 "$AEON_DIR/scripts/detect_setup.py"
Each recipe below has a one-line launch script (in scripts/, shipped with the model) and the plain docker run it executes. The scripts also check which GPUs and which vLLM version the container sees before starting, refuse known-bad combinations, wait for the server, and save the startup log. Every value can be changed through an environment variable, for example AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=4 bash "$AEON_DIR/scripts/serve_2x_r9700.sh"; AEON_DRY_RUN=1 prints the command without starting anything. The header of each script lists its target system, status and memory estimate. Check every start with Verify.
Applies to every recipe:
- Don't pass
--quantization. vLLM reads the Quark config fromconfig.json. - The first start is slow (allow 10–30 minutes). vLLM loads the weights, then Triton compiles and autotunes the Gated DeltaNet and attention kernels, and Quark JIT-compiles its MX emulation kernel with hipcc. The cache mount keeps all of that for later starts.
PYTORCH_ROCM_ARCHmakes Quark compile for your one GPU target instead of the nine the image lists. The scripts set it to the architecture the container detects. In a plaindocker run, use your own target (for examplegfx1200for an RX 9060 XT,gfx1101for an RX 7800 XT, W7700 or V710): a kernel built for a sibling target fails withinvalid device function.- BF16 KV cache everywhere (
--kv-cache-dtype auto). The checkpoint ships no calibrated KV scales, so FP8 KV would run with a scale of 1.0 and change the outputs. FP8 KV is an opt-in (AEON_KV_DTYPE=fp8) for 16 GB cards only. VLLM_ROCM_USE_AITER=0on every recipe. On Radeon, AITER's unified attention overflows RDNA 4's 64 KB LDS. On MI350/MI355 the opt-in native MXFP4 GEMM is picked without it.--attention-backend TRITON_ATTNeverywhere, and inside--speculative-configtoo. The draft model doesn't inherit the target's backend; vLLM's automatic pick on ROCm (ROCM_ATTN) collapses speculative acceptance under concurrency.- MXFP6 dequant-once patch wherever it fits (
AEON_MXFP6_ONCE=1; Recipes A, A-MTP, A-plain, D, E, F, H, H-128, I, J, K, N, O, Q): the same output as stock vLLM, an estimated 5x faster forward pass, at the cost of a BF16 copy of 10.3 GiB in total, divided by the number of GPUs. MXFP4 dequant-at-load (AEON_DEQUANT_AT_LOAD=1; H-128, I, J, K, O, Q): bit-identical, +23.4 GiB in total, divided by the number of GPUs. --override-generation-configis required. The shippedgeneration_config.jsonsays temperature 1.0; the override sets the model card's 0.6 / 0.95 / 20.- Multi-GPU Radeon uses
NCCL_PROTO=Simple(from the community R9700 TP=2 / TP=4 guide) andNCCL_P2P_DISABLE=1(kept for robustness againsthipIpcerrors; that guide found it unnecessary). The scripts set both whenever TP > 1 on a Radeon or Ryzen GPU. The same guide boots withamd_iommu=on iommu=pt pcie_aspm=off. - Ready when the log prints
Application startup complete.(docker logs -f aeon-mxfp4). Stop withdocker rm -f aeon-mxfp4.
2x R9700: pick a recipe
All four recipes serve this checkpoint byte for byte with its own numerics: W4A4 MXFP4 MLP, W6A6 MXFP6 Gated DeltaNet projections, BF16 everywhere else, BF16 KV cache. A, A-MTP and A-plain differ only in speculative decoding, which is lossless in distribution (vLLM's standard rejection sampler keeps the target's output distribution). Nothing here has run on RDNA 4 hardware yet: the acceptance figures are measured Ï„ (mean acceptance length) for this exact checkpoint, drafter and speculative config, and the R9700 speeds are estimates.
| Recipe | Script | Speculation | Prefix caching | Mean acceptance length (measured, T=0.6) | Single-stream speed on 2x R9700 (estimate) | Concurrent 16K-token requests (estimate) | Status |
|---|---|---|---|---|---|---|---|
| A (primary) | serve_2x_r9700_dflash2.sh |
AEON DFlash2 drafter, K=9 | off | 3.73 (4.40 at T=0) | ~35 tok/s | ~5–6 | Untested |
| A-MTP (alternative) | serve_2x_r9700_mtp.sh |
the model's own MTP head, K=3 | on | 2.76 | ~24 tok/s | ~10 | Untested |
| A-plain (exact baseline) | serve_2x_r9700.sh |
none | on | n/a | ~10 tok/s | ~15 | Untested (expected to work) |
| Z (last resort) | serve_2x_r9700_replicas.sh |
none | on | n/a | ~1 tok/s per stream | ~2 per card | Untested (expected to work) |
- Which one. Start with A. It has the most tokens per forward pass, and under emulation every pass is expensive, so that is where speed comes from. For long multi-turn agent loops (each turn resends a growing conversation), also try A-MTP: DFlash2 needs prefix caching off on this hybrid model in vLLM 0.31 (see Recipe A), so every turn re-prefills the whole conversation, while A-MTP reuses the cached prefix. If many long-context agents share one server, A-plain and A-MTP hold more of them at once (table).
scripts/run_matrix.shmeasures all three on your machine (see Tester matrix). - Fall back in this order: A → A-MTP → A-plain → A-plain on the v0.28.0 image (only after a vLLM-level error, not an RCCL hang: AGENTS.md section 5.6) → Z. A-plain is also the reference the other two are compared against.
- The MXFP6 dequant-once patch (
scripts/patches/, on in A, A-MTP and A-plain). Stock vLLM re-dequantizes the MXFP6 Gated DeltaNet weights on every forward pass. That one step was measured at about 2.1 s per pass, 48x the same layers in BF16, which would cap the R9700 at roughly 2 tok/s single-stream (about 7 tok/s with DFlash2). The patch calls the same dequant function once at load and reuses the result: output identical to stock (checked withtorch.equal), at the cost of about 5.2 GiB per GPU. The launcher mounts it read-only over the one vLLM file it replaces, only after checking that file is byte-identical to the v0.31.0 original.AEON_MXFP6_ONCE=0turns it off. - How the speed estimates are made. Measured (one request, T=0.6): on 4 identical prompts plain 8.1, MTP 13.5 and DFlash2 22.0 tok/s (2.7x); on 13 identical prompts MTP τ 2.76 vs DFlash2 τ 4.20; DFlash2 over all 26 prompts τ 3.73, 23.2 tok/s. Scaled to 2x R9700 by memory bandwidth (TP=2 halves the bytes per GPU), plus PCIe all-reduce: about 94 ms per forward pass with the patch. Speedup = mean acceptance length ÷ the relative cost of a drafting-plus-verify step (about 1.14 for DFlash2). Treat every R9700 tok/s figure as ±50% until someone measures it.
- Concurrency. The weight traffic is per forward pass, not per sequence, so aggregate throughput rises with concurrent requests. DFlash2 delivered 4.9x its single-stream throughput at 8 concurrent requests (113 vs 23 tok/s, measured). c=16 is unmeasured. How many requests fit at once is set by the KV and Gated DeltaNet state pool (last column; the startup log's
GPU KV cache sizeis the real number).
Verify (every recipe)
Run this after any launch, in the same terminal. For Recipe Z use NAME=aeon-mxfp4-0 (the proxy is on port 8080).
NAME=aeon-mxfp4; PORT=8000
AUTH=(); [ -n "${VLLM_API_KEY:-}" ] && AUTH=(-H "Authorization: Bearer $VLLM_API_KEY") # an array, so the header stays one argument
# 1) wait until ready (first start: 10-30 minutes); stops and prints the log if the container died
until curl -sf "${AUTH[@]}" http://127.0.0.1:$PORT/health >/dev/null; do
docker inspect -f '{{.State.Running}}' $NAME 2>/dev/null | grep -q true || { docker logs --tail 80 $NAME; break; }
sleep 15
done
# 2) smoke: expect READY
curl -s "${AUTH[@]}" http://127.0.0.1:$PORT/v1/chat/completions -H "Content-Type: application/json" -d '{
"model": "aeon", "max_tokens": 64,
"messages": [{"role": "user", "content": "Reply with exactly: READY"}],
"chat_template_kwargs": {"enable_thinking": false}}' | python3 -c 'import json,sys; print(json.load(sys.stdin)["choices"][0]["message"]["content"])'
# 3) speculative recipes (A, A-MTP, K, Q, any AEON_MTP / AEON_DFLASH run), after some traffic:
curl -s "${AUTH[@]}" http://127.0.0.1:$PORT/metrics | python3 -c '
import re, sys
v = {}
for line in sys.stdin:
m = re.match(r"^(vllm:spec_decode_num_(drafts|draft_tokens|accepted_tokens))(?:_total)?(?:\{[^}]*\})? ([0-9.e+]+)$", line)
if m: v[m.group(2)] = v.get(m.group(2), 0) + float(m.group(3))
d, a = v.get("drafts", 0), v.get("accepted_tokens", 0)
print("mean acceptance length", round(1 + a / d, 2) if d else "no drafts yet")'
Healthy: DFlash2 about 3 or more at one request, MTP about 2 or more. About 1 with several requests in flight means the draft is not on TRITON_ATTN. AGENTS.md section 6 has the full checklist (log lines, tool calls, reasoning field).
Recipe A: 2x R9700 with the AEON DFlash2 drafter (primary)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once at load; drafter in BF16 · Status: Untested on RDNA 4 (measured τ 3.73 at T=0.6 for this checkpoint, drafter and config) · Script:
scripts/serve_2x_r9700_dflash2.sh· Needs: thedflash2/folder (AEON-DFlash2-Qwen3.8-27B-BF16, 3.85 GB)
bash "$AEON_DIR/scripts/serve_2x_r9700_dflash2.sh" # AEON_DFLASH=<dir> if the drafter lives elsewhere
Plain docker run for Recipe A
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/dflash2":/draft:ro \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 -e VLLM_USE_V2_MODEL_RUNNER=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
--max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --no-enable-prefix-caching --language-model-only \
--speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic","rejection_sample_method":"standard"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
The patch path assumes vLLM lives at /usr/local/lib/python3.12/dist-packages/vllm in the image, as it does in v0.31.0. Check with docker run --rm --entrypoint python3 vllm/vllm-openai-rocm:v0.31.0 -c 'import vllm,os; print(os.path.dirname(vllm.__file__))'. The script finds the path itself and also checks the file's hash.
Why these settings:
- The drafter. AEON-DFlash2-Qwen3.8-27B-BF16 is a DFlash 2 block-diffusion drafter fine-tuned for this model from z-lab's Qwen3.8-27B drafter (Apache-2.0). It reads hidden states from five target layers and proposes 9 tokens per step; the target checks them in one forward pass. It is BF16, so it runs on ROCm as is.
"num_speculative_tokens":9matches the drafter's training block of 10 (vLLM drafts 1 + K positions). The per-batch-size lattice[[1,2,9],[3,8,7],[9,12,6],[13,16,4]]lowers K as more requests share a step; it is the schedule the acceptance was measured with.AEON_DFLASH_LATTICE=(empty) keeps K=9 at every batch size."attention_backend":"TRITON_ATTN"inside--speculative-config. The draft model doesn't inherit--attention-backend. vLLM's automatic pick on ROCm is ROCM_ATTN, which collapses DFlash2 acceptance once more than one request is in flight (vllm#53323: acceptance length 1.47 instead of 4.99 at 4 requests).VLLM_USE_V2_MODEL_RUNNER=1. DFlash2 exists only in vLLM's V2 model runner. With the variable set to 0, vLLM quietly loads the drafter as DFlash 1."draft_sample_method":"probabilistic","rejection_sample_method":"standard". Both are lossless. Probabilistic drafting is what was measured;AEON_DFLASH_SAMPLE=greedyis the alternative. Never use"synthetic": it is not lossless.--no-enable-prefix-caching. In vLLM 0.29–0.31, a hybrid Gated DeltaNet model with a DFlash2 drafter crashes or drops to 0% acceptance after a prefix-cache hit (vllm#58894, fix proposed in vllm#55601). Every request therefore prefills its whole prompt. For one-shot work that costs nothing extra; for long multi-turn loops, compare against A-MTP.AEON_DFLASH_PREFIX_FIX=1mounts the proposed one-line fix and keeps prefix caching on: experimental and untested.--dtype bfloat16avoids a drafter dtype mismatch reported on ROCm (vllm#42588) and fp16's 0% acceptance (vllm#55250).- TP=2, NCCL settings,
--language-model-only, BF16 KV: as in Recipe A-plain. - Memory (estimate) per GPU: about 10.2 GiB of weights, 5.2 GiB for the MXFP6 BF16 copy, 1.8 GiB for the drafter (it shares the target's embeddings and
lm_head), about 3 GiB for activations and graphs, and about 8 GiB left for KV cache and Gated DeltaNet state. vLLM pads the target's 16 attention layers to 20 to group them with the drafter's 5 layers (the log warnsAdd 4 padding layers, may waste at most 25.00% KV cache memory: expected), so KV costs about 40 KiB per token per GPU, and each running request also holds about 0.77 GiB of Gated DeltaNet state and about 0.1 GiB of drafter KV. Realistic capacity: about 5–6 concurrent 16K-token requests, about 4 at 32K. ExpectGPU KV cache sizearound 150–200K tokens andMaximum concurrency for 131,072 tokens per requestaround 1.2–1.5x. More requests wait in the queue; with many long-context agents,AEON_MAX_NUM_SEQS=8avoids preemption churn, or use A-plain / A-MTP, which hold more. - Pass check. The log shows
Resolved architecture: DFlash2DraftModel,MXFP6 weights dequantized once at load, and under loadSpecDecoding metrics: Mean acceptance length: ...of about 3 or more. About 1 with several requests in flight means the draft is not on TRITON_ATTN. - If it fails to load, or acceptance stays near 1, go to A-MTP and send us the log.
Recipe A-MTP: 2x R9700 with MTP (alternative)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once at load; the MTP head runs in BF16 · Status: Untested on RDNA 4 (measured τ 2.76 at T=0.6 for this checkpoint's MTP head) · Script:
scripts/serve_2x_r9700_mtp.sh
bash "$AEON_DIR/scripts/serve_2x_r9700_mtp.sh"
Plain docker run for Recipe A-MTP
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
--max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --language-model-only \
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- What it is. The model's own one-layer MTP head drafts 3 tokens; the main model checks them in one pass.
config.jsonlists the MTP head's layers as excluded from quantization, so vLLM builds them in BF16 to match the weights on disk. Measured: mean acceptance length 2.76 at T=0.6 (DFlash2: 4.20 on the same 13 prompts). - Why keep it. MTP keeps prefix caching on, so a new turn of a long conversation only prefills the new tokens. Under emulation prefill is expensive, so for long multi-turn agent loops A-MTP can finish turns sooner than A even though it decodes slower. The tester matrix measures both.
- Same as A-plain except speculation: identical environment and flags plus
--speculative-config, so the matrix compares speculation only. - Pass check.
Resolved architecture: Qwen3_5MTPin the log (it followsResolved architecture: Qwen3_5ForConditionalGeneration; vLLM 0.31 runs this model on its V2 runner, which does not print the V1 runner'sDetected MTP modelline), and under loadSpecDecoding metrics: Mean acceptance lengthof about 2 or more. - Memory: about 0.45 GiB of Gated DeltaNet state per running request per GPU; about 10 concurrent 16K-token requests (estimate).
- Tuning:
AEON_MTP=2orAEON_MTP=4.
Recipe A-plain: 2x R9700 without speculation (exact baseline)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested (expected to work) · Script:
scripts/serve_2x_r9700.sh
bash "$AEON_DIR/scripts/serve_2x_r9700.sh"
Plain docker run for Recipe A-plain
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
--max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Why these settings (shared by A and A-MTP):
--tensor-parallel-size 2splits every layer across both cards. Each card holds about 10.2 GiB of weights and dequantizes only half of them on every forward pass. All shard boundaries land on whole 32-element MX blocks (checked against the real shapes), and both cards carry 12 attention heads, 2 KV heads, 8 GDN key heads and 24 value heads.NCCL_PROTO=Simpleand--attention-backend TRITON_ATTNare what the community R9700 TP=2 / TP=4 setups use (Level1Techs guide); a TP=2 hang on dual R9700 is still reported open (vllm#40980), which is why Recipe Z exists.NCCL_P2P_DISABLE=1is kept for robustness againsthipIpcerrors (the guide found it unnecessary). vLLM's custom all-reduce is MI300/MI350-only, so Radeon cards talk through RCCL over PCIe.- No speculation, prefix caching on. This recipe avoids every speculative-decoding issue, so it is the most likely TP=2 recipe to work on the first try, and it is the reference for the tester matrix.
--max-model-len 131072 --max-num-seqs 16. With the MXFP6 BF16 copy, an estimated 10 GiB per card is left for cache: BF16 KV costs 32 KiB per token per card, so roughly 300K tokens in total, plus about 0.22 GiB of Gated DeltaNet state per running request per card (about 15 concurrent 16K-token requests). For the full 262,144-token context useAEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=4. The startup lineGPU KV cache size: N tokensis the real number.AEON_MXFP6_ONCE=0gives back about 5.2 GiB per card, at roughly a fifth of the speed.--language-model-onlyskips the vision tower (it fails to load on RDNA 4 today).--gpu-memory-utilization 0.90: use 0.85–0.88 if a card drives a display.
Recipe Z: 2x R9700 as two independent servers (never-fails fallback)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB (one full copy per card) · Arch: gfx1201 (RDNA 4) · MX path: emulated, stock (no patch) · Status: Untested (expected to work; fewest moving parts) · Script:
scripts/serve_2x_r9700_replicas.sh(stop with... replicas.sh stop)
Use this when every TP=2 recipe fails twice (hipIpcGetMemHandle error, crash at graph capture), or right away after a TP=2 hang (a serve script exits with code 3, "PROBABLY HUNG"; RCCL, vllm#40980, ROCm/ROCm#6148): a hang doesn't depend on speculative decoding, so A-MTP and A-plain would hang the same way. It runs one complete copy of the model per card, with no inter-GPU communication, no speculative decoding, no graph capture (--enforce-eager) and no AITER: aeon-mxfp4-0 on port 8000, aeon-mxfp4-1 on port 8001, and a small least-connections proxy (scripts/aeon_lb.py, or scripts/nginx_aeon_replicas.conf) on port 8080. The MXFP6 dequant-once patch doesn't fit next to a full copy of the weights on one 32 GB card, so Z runs stock emulation: it is the slowest option, a last resort.
bash "$AEON_DIR/scripts/serve_2x_r9700_replicas.sh"
Plain docker run for Recipe Z
for i in 0 1; do
docker run -d --name aeon-mxfp4-$i \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:$((8000+i)):8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e HIP_VISIBLE_DEVICES=$i -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
--max-num-seqs 4 --gpu-memory-utilization 0.88 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
--attention-backend TRITON_ATTN --enforce-eager --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
[ "$i" = 0 ] && until docker logs aeon-mxfp4-0 2>&1 | grep -q "Application startup complete"; do
docker inspect -f '{{.State.Running}}' aeon-mxfp4-0 2>/dev/null | grep -q true || break; sleep 15
done # GPU 1 reuses the kernel cache GPU 0 built; stops waiting if replica 0 died
done
python3 "$AEON_DIR/scripts/aeon_lb.py" --listen 127.0.0.1:8080 127.0.0.1:8000 127.0.0.1:8001 &
- Cost (estimate): each card holds about 20.4 GiB of weights and about 6 GiB of cache: 32,768-token context, about 4 sequences per card (about 2 concurrent 16K-token requests). Speed: about 1 tok/s per stream (stock MXFP6 emulation, TP=1, eager) and a few tok/s per card at 4 streams.
- The script starts the second server only after the first is ready, so the one-time kernel build isn't run twice at the same time. Each container gets only its own GPU's device nodes.
- First speed step once it works: drop
--enforce-eager(AEON_EAGER=0). - Don't use
--data-parallel-sizefor this: it still sets up inter-GPU process groups.
Tester matrix: plain vs MTP vs DFlash2
mkdir -p ~/aeon-test && cd ~/aeon-test && bash "$AEON_DIR/scripts/run_matrix.sh"
It starts A-plain, A-MTP and A one after another (all with 16 sequences), runs the same checks on each, then checks TP=2 against a one-GPU reference, and writes results/matrix_compare.md, results/tpcheck.md and the archive to send back:
| Check | Detail |
|---|---|
| Speed | decode tok/s (aggregate and per stream), time to first token and speculative acceptance at 1, 4, 8 and 16 concurrent requests (T=0.6, 256-token outputs), plus prefill speed on a ~4K-token prompt |
| Quality | GSM8K (first 50, greedy, exact match), IFEval (20 prompts, strict), 10 tool-calling prompts, 7 smoke checks |
| TP=1 reference | all three arms are TP=2 and share the same MX weight-sharding code, so a sharding bug would look the same on each. scripts/tp_check.py records prefill logprobs of 4 fixed prompts on A-plain, then on Recipe Z's replica 0 alone (TP=1, one GPU): expect more than 99% top-1 agreement |
| Memory | peak VRAM per GPU, GPU KV cache size from the startup log |
The three should score the same on the quality checks within noise: speculative decoding doesn't change the output distribution. It is not bit-reproducible, though: at T=0 the outputs usually diverge after a few dozen tokens, because verify batches change the floating-point reduction order (measured: 1 of 20 GSM8K outputs identical to plain decoding, mean common prefix 16%), while accuracy stayed statistically the same (plain GSM8K 20/20 and tools 10/10, DFlash2 19/20 and 9/10: two discordant items, McNemar p = 0.5). Plain decoding at different concurrency isn't bit-reproducible either. Expect 1–3 hours for all three arms after the first compile; AEON_MATRIX_ARMS="dflash2 mtp" runs a subset. An arm still compiling after the extra wait stops the matrix instead of being killed mid-build.
Recipe A-Max: 2x R9700, BF16-resident MLP (experimental, superseded)
Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: MXFP4 MLP dequantized once at load; MXFP6 on stock per-pass emulation; MTP in BF16 · Status: Untested (experimental) · Script:
scripts/serve_2x_r9700_max.sh
vLLM 0.31's VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 stores the MXFP4 MLP as BF16 (output identical; +11.7 GiB per GPU). It keeps the MXFP6 projections on stock emulation, which costs about 6x more per pass (measured) than the MXFP4 emulation this removes, and both BF16 copies together don't fit on 32 GB cards. Expect it to be slower than A-MTP. It stays in the kit as an A/B data point only (65,536 context, 4 sequences, --gpu-memory-utilization 0.92); if it runs out of memory, step AEON_MAX_MODEL_LEN down to 49152, then 32768.
Speed experiment: radiance (experimental, unlicensed, untested on this checkpoint)
radiance is a community overlay for vLLM with native RDNA 4 kernels: MXFP4 weights with FP8 activations (W4A8) and E6 (MXFP6) kernels. Users report 130–210 tok/s single-stream with DFlash2 on 2x R9700 with other checkpoints. It is not part of this kit, and we have not run it with this checkpoint.
- It can't serve this checkpoint as shipped. Its MXFP6 kernels are wired only to its own
paroquant_mxfp6format. Serving our Quark MXFP6 layers needs a new kernel class in vLLM's MXFP6 kernel list. Our weights would fold exactly: every MXFP6 row has an E8M0 scale spread of at most 6, which its fold requires (measured on all 1,032,192 rows). Its Quark MXFP4 patch targets vLLM 0.29, which can't load this checkpoint withoutAEON_STRIP_ALGO_CONFIG=1. - It changes the numerics. Activations become per-token FP8 instead of the checkpoint's MXFP4 / MXFP6, so the output is not identical to Recipes A, A-MTP and A-plain (radiance's own MXFP4 perplexity is slightly better with W4A8, but it is a deviation). Its defaults also include lossy options that must be turned off for a fidelity comparison: quantized all-reduce, FP8 KV cache and a quantized output head.
- License. The original repository has no license file. Private experiments only; don't redistribute it or code copied from it.
- If you try it, compare it against A-plain with
tests/run_radeon_eval.py --compare, and send us both reports.
TensorFold
TensorFold (Apache-2.0) drafts with DFlash2 and is bit-exact against its own serial decoding. Its ROCm support is a community fork (millaguie/TensorFold, branch rocm-r9700).
- It can't load this mixed checkpoint, so it is not a path for this model today. The ROCm fork serves Qwen3.8-27B MLX 4-bit checkpoints, GGUF, or a Quark export where every projection is MXFP4, one GPU per instance (no TP=2). Our checkpoint is detected as Quark MXFP4 and then fails at the first MXFP6 or BF16 projection; it would need an MXFP6 path, mixed projection groups and BF16 linears on ROCm. Pairing the drafter with a stock Qwen3.8-27B checkpoint would serve a different model. On 2x R9700, use vLLM Recipe A.
- The drafter is shipped for TensorFold so it can be used as soon as TensorFold supports this mixed Quark export:
tensorfold serve <AEON target> --drafter "$AEON_DIR/dflash2". Its config and tensors load in TensorFold's DFlash2 loader (checked against the source and the fork's reader). TensorFold's CUDA/ROCm engine ignores--drafter-bitsand re-packs the drafter's linears to affine 4-bit (group 64) at load, as it does for z-lab's drafter, which can lower acceptance but not output fidelity. See the drafter card.
Recipe K: 4x Radeon AI PRO R9700 with the AEON DFlash2 drafter
Target: 4x AMD Radeon AI PRO R9700 / R9700S / R9600D · VRAM: 4 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical); drafter in BF16 · Status: Untested · Script:
scripts/serve_4x_r9700.sh
bash "$AEON_DIR/scripts/serve_4x_r9700.sh"
Plain docker run for Recipe K
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/dflash2":/draft:ro \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 -e VLLM_USE_V2_MODEL_RUNNER=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 4 --max-model-len 131072 \
--max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --no-enable-prefix-caching --language-model-only \
--speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic","rejection_sample_method":"standard"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate) per GPU: weights 5.1 + MXFP6 BF16 copy 2.6 + BF16 MLP 5.9 + drafter 0.9 = 14.4 GiB, about 3 GiB of activations, about 11 GiB of pool (about 560K tokens at about 20 KiB per token per GPU: 16 KiB, padded to 20 for the drafter's layers; before the GDN state of about 0.37 GiB per running request): about 4 concurrent 128K-token requests.
- Speed (estimate): about 25–30 tok/s plain and 60–90 with DFlash2 single-stream (±50%).
- Head counts divide: 1 KV head and 6 attention heads per GPU for the target, 2 KV heads for the drafter. TP=4 with
NCCL_PROTO=Simpleis reported working on 4x R9700 with other checkpoints. - Variants: K-MTP for long multi-turn agent loops (prefix caching on):
AEON_DFLASH= AEON_MTP=3 AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=8. K-plain:AEON_DFLASH=. - Fallbacks: two Recipe A servers, one per pair of cards (
AEON_DEVICES,AEON_NAME,AEON_PORT8000 / 8001,AEON_REPLACE_OTHERS=0for the second), then Z withAEON_REPLICAS=4.
Recipe B: 1x Radeon AI PRO R9700
Target: 1x AMD Radeon AI PRO R9700 / R9700S / R9600D · VRAM: 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, stock · Status: Untested (expected to work; slow) · Script:
scripts/serve_1x_r9700.sh
bash "$AEON_DIR/scripts/serve_1x_r9700.sh"
Plain docker run for Recipe B
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
--max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- The MXFP6 dequant-once patch doesn't fit (20.4 + 10.3 GiB > 28.7 GiB usable), so this runs stock emulation: about 1 tok/s single-stream (estimate, ±50%). For real use, two or more cards.
- Memory (estimate): about 20.4 GiB of weights leaves about 6 GiB for cache: BF16 KV at 64 KiB per token plus about 0.45 GiB of Gated DeltaNet state per running sequence (prefix caching on).
- Longer context:
AEON_MAX_MODEL_LEN=65536 AEON_MAX_NUM_SEQS=2. - MTP:
AEON_MTP=3 AEON_MAX_NUM_SEQS=2(about 2.3x, estimate). - Most robust:
AEON_EAGER=1 AEON_MAX_MODEL_LEN=16384 AEON_MAX_NUM_SEQS=2 AEON_GPU_UTIL=0.88. - Dequant-at-load and DFlash2 don't fit on one 32 GB card.
Recipe C: 2x 16 GB RDNA 4 (RX 9070 XT, RX 9070, RX 9060 XT)
Target: 2x AMD Radeon RX 9070 XT or RX 9070 (gfx1201), or 2x RX 9060 XT 16 GB (gfx1200, about half the bandwidth) · VRAM: 2 × 16 GB · Arch: gfx1201 / gfx1200 (RDNA 4) · MX path: emulated, stock · Status: Untested (expected to work; tight memory) · Script:
scripts/serve_2x_rx9070xt.sh
bash "$AEON_DIR/scripts/serve_2x_rx9070xt.sh"
Plain docker run for Recipe C
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 16384 \
--max-num-seqs 2 --gpu-memory-utilization 0.92 --kv-cache-dtype auto --max-num-batched-tokens 2048 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Same code path as Recipe A-plain with half the memory: about 10.2 GiB of weights per card leaves about 2.6 GiB per card for cache, and the MXFP6 dequant-once patch doesn't fit. About 2 tok/s single-stream on RX 9070 XT, about 1 on RX 9060 XT (estimates). On RX 9060 XT cards use PYTORCH_ROCM_ARCH=gfx1200 in a plain docker run (the script detects it). If a card drives a display, use --gpu-memory-utilization 0.88. MTP (AEON_MTP=2) only if the GPU KV cache size line shows room. FP8 KV (AEON_KV_DTYPE=fp8) doubles the context but changes outputs (uncalibrated scales); it is opt-in. There is no single-card fallback: the model doesn't fit on 16 GB.
Recipe D: 4x 16 GB cards (RDNA 4 or RDNA 3)
Target: 4x AMD Radeon RX 9070 XT / RX 9070 (gfx1201) or RX 9060 XT 16 GB (gfx1200); or 4x RX 7900 GRE (gfx1100), RX 7800 XT / RX 7700 / Radeon PRO W7700 (gfx1101) · VRAM: 4 × 16 GB · Arch: gfx1201 / gfx1200 / gfx1100 / gfx1101 · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script:
scripts/serve_4x_rx9070xt.sh
bash "$AEON_DIR/scripts/serve_4x_rx9070xt.sh"
Plain docker run for Recipe D
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 4 --max-model-len 65536 \
--max-num-seqs 8 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate) per GPU: weights 5.1 + MXFP6 BF16 copy 2.6 GiB, about 2 GiB of activations, about 4.6 GiB of pool (about 300K tokens at 16 KiB per token per GPU): about 4 concurrent 64K-token requests. Longer context:
AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=4. - Speed (estimate): about 18 tok/s single-stream on RX 9070 XT, about 40 with MTP (
AEON_MTP=3), about 50 with DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=4, prefix caching off). - TP=4 shards check out (6 attention heads and 1 KV head per card; every MX shard on a whole block), and TP=4 with
NCCL_PROTO=Simpleis reported working on 4x R9700. All-reduce over consumer PCIe likely limits speed. If startup hangs at graph capture,AEON_EAGER=1. - 4x 12 GB (RX 9070 GRE, RX 7700 XT):
AEON_MXFP6_ONCE=0 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=4 AEON_GPU_UTIL=0.92(untested). - Don't use 8 cards with TP=8: 4 KV heads can't be split 8 ways.
Recipe E: 2x Radeon RX 7900 XTX
Target: 2x AMD Radeon RX 7900 XTX · VRAM: 2 × 24 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script:
scripts/serve_2x_rx7900xtx.sh
bash "$AEON_DIR/scripts/serve_2x_rx7900xtx.sh"
Plain docker run for Recipe E
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 32768 \
--max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate) per GPU: weights 10.2 + MXFP6 BF16 copy 5.2 GiB, about 2 GiB of activations, about 4.2 GiB of pool (about 138K tokens; RDNA 3 has no FP8, so KV stays BF16 at 32 KiB per token per card): about 3 concurrent 32K-token requests.
- Speed (estimate): about 15 tok/s single-stream, about 35 with MTP (
AEON_MTP=3). - Variants:
AEON_MAX_MODEL_LEN=65536 AEON_MAX_NUM_SEQS=2. Long context on stock emulation (about 5x slower):AEON_MXFP6_ONCE=0 AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=2. DFlash2 and dequant-at-load don't fit next to the patch. - 2x RX 7900 XT (20 GB): use Recipe M.
Recipe M: 2x Radeon RX 7900 XT
Target: 2x AMD Radeon RX 7900 XT · VRAM: 2 × 20 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, stock · Status: Untested · Script:
scripts/serve_2x_rx7900xt.sh
bash "$AEON_DIR/scripts/serve_2x_rx7900xt.sh"
Plain docker run for Recipe M
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 32768 \
--max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate) per GPU: about 18 GiB usable, 10.2 GiB of weights, about 2 GiB of activations, about 5.8 GiB of pool: about 4–5 concurrent 32K-token requests.
- Speed (estimate): about 2.3 tok/s single-stream (stock emulation); MTP
AEON_MTP=3. - Opt-in MXFP6 dequant-once (about 12 tok/s, estimate; tight, untested, no display on the cards):
AEON_MXFP6_ONCE=1 AEON_GPU_UTIL=0.93 AEON_MAX_MODEL_LEN=16384 AEON_MAX_NUM_SEQS=2(about 1.5 GiB of pool per GPU).
Recipe L: 2x 16 GB RDNA 3 (RX 7900 GRE, RX 7800 XT, RX 7700, PRO W7700)
Target: 2x AMD Radeon RX 7900 GRE (gfx1100), or 2x RX 7800 XT / RX 7700 / Radeon PRO W7700 (gfx1101) · VRAM: 2 × 16 GB · Arch: gfx1100 / gfx1101 (RDNA 3) · MX path: emulated, stock · Status: Untested (tight memory) · Script:
scripts/serve_2x_16gb_rdna3.sh
bash "$AEON_DIR/scripts/serve_2x_16gb_rdna3.sh"
Plain docker run for Recipe L
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 16384 \
--max-num-seqs 2 --gpu-memory-utilization 0.92 --kv-cache-dtype auto --max-num-batched-tokens 2048 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Like Recipe C on RDNA 3: about 2.6 GiB of pool per card (a 16K-token request takes about 0.5 GiB of KV plus 0.22 GiB of state per card), stock emulation, about 1.7 tok/s single-stream (estimate). On gfx1101 cards use PYTORCH_ROCM_ARCH=gfx1101 in a plain docker run (the script detects it). Use 0.88 for --gpu-memory-utilization if a card drives a display. MTP (AEON_MTP=2) only if the GPU KV cache size line shows at least 1.5 GiB. Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=8192 AEON_MAX_NUM_SEQS=1. No single-card fallback.
Recipe F: Radeon PRO W7900 (48 GB)
Target: 1x AMD Radeon PRO W7900, W7900 Dual Slot or W7800 48 GB · VRAM: 48 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script:
scripts/serve_w7900.sh
bash "$AEON_DIR/scripts/serve_w7900.sh"
Plain docker run for Recipe F
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 65536 \
--max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate): weights 20.4 + MXFP6 BF16 copy 10.3 GiB, about 3 GiB of activations, about 9.5 GiB of pool (about 155K tokens at 64 KiB per token): about 2 concurrent 64K-token requests. Dequant-at-load doesn't fit (20.4 + 23.4 GiB > 43.2 usable).
- Speed (estimate): about 8 tok/s single-stream, about 18 with MTP (
AEON_MTP=3), about 25 with DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=2, prefix caching off). - Variants:
AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=2. The full 262,144-token context on stock emulation (about 5x slower):AEON_MXFP6_ONCE=0 AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=1. Vision:AEON_VISION=1. - Two or four of these cards: Recipe Q.
Recipe Q: 2x Radeon PRO W7900 with the AEON DFlash2 drafter
Target: 2x AMD Radeon PRO W7900, W7900 Dual Slot or W7800 48 GB (variants for 4x W7900 and 2x W7800 32 GB) · VRAM: 2 × 48 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical); drafter in BF16 · Status: Untested · Script:
scripts/serve_2x_w7900.sh
bash "$AEON_DIR/scripts/serve_2x_w7900.sh"
Plain docker run for Recipe Q
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/dflash2":/draft:ro \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
-e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 -e VLLM_USE_V2_MODEL_RUNNER=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
--max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --no-enable-prefix-caching --language-model-only \
--speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic","rejection_sample_method":"standard"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate) per GPU: 10.2 + 5.2 (MXFP6 BF16) + 11.7 (BF16 MLP) + 1.8 (drafter) = 28.9 GiB, about 3.5 GiB of activations, about 10.8 GiB of pool.
- Speed (estimate): about 24 tok/s plain and about 70 with DFlash2 single-stream.
- Variants: Q-MTP for long multi-turn agent loops (prefix caching on):
AEON_DFLASH= AEON_MTP=3 AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=8. 2x W7800 32 GB:AEON_DEQUANT_AT_LOAD=0 AEON_DFLASH=(like A-plain; about 10 GiB of pool per card). 4x W7900:AEON_TP=4. - Fallback: one Recipe F server per card (
AEON_GPU_IDS=0/1,AEON_PORT8000 / 8001,AEON_NAME,AEON_REPLACE_OTHERS=0for the second).
Recipe G: Radeon PRO W7800 (32 GB)
Target: 1x AMD Radeon PRO W7800 32 GB · VRAM: 32 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, stock · Status: Untested (slow) · Script:
scripts/serve_w7800_32gb.sh
bash "$AEON_DIR/scripts/serve_w7800_32gb.sh"
Plain docker run for Recipe G
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
--max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
Like Recipe B on RDNA 3: about 6 GiB left for BF16 KV and Gated DeltaNet state; the MXFP6 dequant-once patch doesn't fit, so about 0.9 tok/s single-stream (estimate). MTP: AEON_MTP=3 AEON_MAX_NUM_SEQS=2. Text-only. For real use, two cards (Recipe Q's 2x W7800 32 GB variant).
Recipe P: Radeon PRO V710
Target: 1x AMD Radeon PRO V710 (the full GPU, not a fractional partition; usually a cloud VM) · VRAM: 28 GB · Arch: gfx1101 (RDNA 3) · MX path: emulated, stock · Status: Untested (slow) · Script:
scripts/serve_v710.sh
bash "$AEON_DIR/scripts/serve_v710.sh"
Plain docker run for Recipe P
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1101 -e VLLM_ROCM_USE_AITER=0 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 16384 \
--max-num-seqs 2 --gpu-memory-utilization 0.92 --kv-cache-dtype auto --max-num-batched-tokens 2048 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
About 25.7 GiB usable, 20.4 GiB of weights, about 3.1 GiB of pool (a 16K-token request takes about 1.0 GiB of KV plus 0.43 GiB of state). About 0.7 tok/s single-stream (estimate). No room for speculation. TP=1 needs no RCCL, so the preflight's virtual-machine warning doesn't apply here. Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=8192 AEON_MAX_NUM_SEQS=1.
Recipe H: Strix Halo (Ryzen AI Max+ 395), 64 GB
Target: AMD Ryzen AI Max+ 395 / Max 390 / Max 385 APU (Radeon 8060S / 8050S) with 64 GB · Memory: 64 GB unified (at least 45 GiB GPU-addressable) · Arch: gfx1151 (RDNA 3.5) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script:
scripts/serve_strix_halo.sh
Before you start, the GPU must be able to allocate at least about 45 GiB. vLLM sizes its memory from what HIP reports, so raise the TTM/GTT limit the way AMD documents it for Strix Halo (amd-ttm, or the ttm pages_limit module parameter; amdgpu.gttsize is deprecated), or raise the UMA frame buffer in the BIOS. scripts/detect_setup.py prints both values. Don't set HSA_OVERRIDE_GFX_VERSION (older Strix Halo guides do): gfx1151 is a native target of the image, and the launcher refuses the override.
bash "$AEON_DIR/scripts/serve_strix_halo.sh"
Plain docker run for Recipe H
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1151 -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
--max-num-seqs 2 --gpu-memory-utilization 0.85 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate): of about 48 GiB addressable, about 40.8 usable: weights 20.4 + MXFP6 BF16 copy 10.3 GiB, about 7.9 GiB of pool. BF16 KV (no FP8 on gfx1151).
- Speed (estimate): about 2.3 tok/s single-stream (about 0.4 without the patch), about 5 with MTP (
AEON_MTP=3). - Precision caveat: AMD lists only FP16 as officially validated on Ryzen APUs. This recipe runs BF16 (the reference numerics; fp16 also breaks the DFlash2 drafter): an untested risk.
- Most robust:
AEON_EAGER=1 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=1. If vLLM's free-memory check fails on unified memory,AEON_EXTRA_ARGS="--kv-cache-memory-bytes <bytes>"is an untested alternative.
Recipe H-128: Strix Halo (Ryzen AI Max+ 395), 128 GB
Target: AMD Ryzen AI Max+ 395 / Max 390 / Max 385 APU with 128 GB · Memory: 128 GB unified (at least 90 GiB GPU-addressable) · Arch: gfx1151 (RDNA 3.5) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical); MTP head in BF16 · Status: Untested · Script:
scripts/serve_strix_halo_128gb.sh
Same preparation as Recipe H, with at least 90 GiB GPU-addressable.
bash "$AEON_DIR/scripts/serve_strix_halo_128gb.sh"
Plain docker run for Recipe H-128
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx1151 -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 131072 \
--max-num-seqs 4 --gpu-memory-utilization 0.80 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --language-model-only \
--speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}' \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate): of about 96 GiB addressable, about 76.8 usable: weights 20.4 + 10.3 + 23.4 + MTP 0.8 GiB, about 3.5 GiB of activations, about 18 GiB of pool: about 2 concurrent 128K-token requests.
- Speed (estimate): about 4.5 tok/s plain and about 10 with MTP (the default) single-stream.
- DFlash2 instead of MTP:
AEON_MTP=0 AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=65536(prefix caching off). Plain:AEON_MTP=0. Most robust:AEON_EAGER=1 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=1.
Recipe I: Instinct MI300X / MI325X
Target: AMD Instinct MI300X, MI308X or MI325X · VRAM: 192 / 256 GB · Arch: gfx942 (CDNA 3) · MX path: emulated (no MX hardware), MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical) · Status: Untested · Script:
scripts/serve_mi300x.sh
bash "$AEON_DIR/scripts/serve_mi300x.sh"
Plain docker run for Recipe I
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx942 -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 262144 \
--max-num-seqs 32 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 16384 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate), MI300X: about 172 GiB usable, weights 54.1 GiB (20.4 + 10.3 + 23.4), about 114 GiB of pool (about 1.8M BF16 KV tokens): about 7 concurrent 262K-token requests. Memory is not a constraint, so use one GPU per copy and scale out with data parallelism (
AEON_DP=8) instead of tensor parallelism. - Faster:
AEON_MTP=3(best for agentic multi-turn: prefix caching stays on). For short-context throughput,AEON_DFLASH="$AEON_DIR/dflash2"(prefix caching off; withAEON_DP>1vLLM drops the batch-size lattice, so addAEON_DFLASH_K=6). - Images:
AEON_VISION=1(4 images per prompt). - An FP8 or BF16 build is the better everyday choice on MI300, since memory isn't the limit there.
Recipe O: Instinct MI300A
Target: AMD Instinct MI300A (APU) · Memory: 128 GB unified HBM3, shared with the host OS · Arch: gfx942 (CDNA 3) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical) · Status: Untested · Script:
scripts/serve_mi300a.sh
bash "$AEON_DIR/scripts/serve_mi300a.sh"
Plain docker run for Recipe O
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx942 -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 131072 \
--max-num-seqs 16 --gpu-memory-utilization 0.70 --kv-cache-dtype auto --max-num-batched-tokens 16384 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate):
--gpu-memory-utilization 0.70leaves room for the OS: about 89 GiB usable, weights 54.1 GiB, about 32 GiB of pool: about 3–4 concurrent 128K-token requests. Size the utilization from the GiB the launcher prints for the GPU; if it reports much less than 128 GiB, the GPU-allocatable limit has to be raised per AMD's MI300A guidance (ask your administrator). - Faster:
AEON_MTP=3, or DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2", prefix caching off). Several APUs:AEON_DP=<number of APUs>.
Recipe J: Instinct MI350X / MI355X
Target: AMD Instinct MI350X or MI355X · VRAM: 288 GB · Arch: gfx950 (CDNA 4) · MX path: default exact: MXFP4 emulated with the MLP kept in BF16 after load, MXFP6 dequantized once (bit-identical to every other recipe); native MXFP4 (AITER W4A4) opt-in · Status: Untested · Script:
scripts/serve_mi355x.sh
bash "$AEON_DIR/scripts/serve_mi355x.sh"
Plain docker run for Recipe J
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx950 -e VLLM_ROCM_USE_AITER=0 -e VLLM_DISABLED_KERNELS=AiterMxfp4LinearKernel -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 \
-e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 262144 \
--max-num-seqs 32 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 16384 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Why the exact path is the default. Native MXFP4 runs the MLP on AITER's W4A4 GEMM: the same MX grid and E8M0 scale rule, but a different kernel and summation order, so it is not bit-identical to the other recipes. This kit puts fidelity first, so native is a labelled opt-in:
AEON_MXFP4_EMULATE=0 AEON_DEQUANT_AT_LOAD=0(faster at high concurrency; log lineUsing AiterMxfp4LinearKernel for MXFP4 GEMM). If the native path fails with aModuleNotFoundErrorforaiter.ops.triton.gemm_afp4wfp4(newer AITER builds moved that module), stay on the default. - Check the log for
Using EmulationMxfp4LinearKernel for MXFP4 GEMM(default) andUsing EmulationMxfp6LinearKernel for MXFP6 GEMM. - Memory (estimate): about 259 GiB usable, weights 54.1 GiB (30.7 with native MXFP4), about 200 GiB of pool.
- Faster: DFlash2 (
AEON_DFLASH="$AEON_DIR/dflash2") or MTP (AEON_MTP=3). Scaling out:AEON_DP=N(with DFlash2, addAEON_DFLASH_K=6).
Recipe N: Instinct MI210 / MI250 / MI250X
Target: AMD Instinct MI210 (64 GB), MI250 / MI250X (two 64 GB GCDs, each shown as its own GPU) · VRAM: 64 GB per GCD · Arch: gfx90a (CDNA 2) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script:
scripts/serve_mi250.sh
bash "$AEON_DIR/scripts/serve_mi250.sh"
Plain docker run for Recipe N
docker run -d --name aeon-mxfp4 \
--device /dev/kfd --device /dev/dri \
--group-add video --group-add "$RENDER_GID" \
--cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
-p 127.0.0.1:8000:8000 \
-v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
-v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
-e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
-e PYTORCH_ROCM_ARCH=gfx90a -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
vllm/vllm-openai-rocm:v0.31.0 \
/model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 131072 \
--max-num-seqs 8 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
--attention-backend TRITON_ATTN --language-model-only \
--reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
--override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
- Memory (estimate) per GCD: about 57.6 GiB usable, weights 20.4 + MXFP6 BF16 copy 10.3 GiB, about 3 GiB of activations, about 24 GiB of pool: about 2–3 concurrent 128K-token requests. MXFP4 dequant-at-load doesn't fit next to it on one GCD.
- Speed (estimate): about 14 tok/s per GCD single-stream (likely launch-bound), about 33 with MTP (
AEON_MTP=3). - Scale out: one replica per GCD with
AEON_DP=<number of GCDs>. TP=2 across the two GCDs of one MI250 / MI250X:AEON_TP=2 AEON_DEQUANT_AT_LOAD=1 AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=16(xGMI; no NCCL overrides). - AITER stays off (vLLM's AITER needs CDNA 3 or newer).
Not supported
- Single GPUs under 28 GB: RX 7900 XTX (24 GB), RX 7900 XT (20 GB), every 16 GB card (RX 9070 XT / 9070 / 9060 XT 16 GB, RX 7900 GRE, RX 7800 XT, RX 7700, Radeon PRO W7700) and 12 GB cards (RX 9070 GRE, RX 7700 XT). About 20.4 GiB of text-only weights leaves no usable KV cache at 90% utilization. Smallest working combinations: 2x 24 GB (E), 2x 20 GB (M), 2x 16 GB (C for RDNA 4, L for RDNA 3), 4x 12 GB (D variant).
- 8 GB cards (RX 9060 XT 8 GB, RX 7600 and similar): even four of them leave no room for activations and cache.
- Architectures ROCm supports but the vLLM ROCm image isn't built for: Instinct MI100 (gfx908) and RDNA 2 (gfx1030: Radeon PRO W6800 / V620, RX 6800–6950 XT). vLLM's ROCm build targets gfx90a, gfx942, gfx950, gfx1100, gfx1101, gfx1150, gfx1151, gfx1200 and gfx1201. Also gfx1102 (RX 7600 / 7600 XT, Radeon PRO W7600 / W7500), which is in neither AMD's ROCm Linux list nor the image, and the MI50 / MI60 (gfx906), which current ROCm no longer supports.
- gfx1250: vLLM ships a separate image for it; these recipes don't cover it yet.
- Integrated GPUs other than Strix Halo (Ryzen AI 300 / gfx1150, Radeon 780M / gfx1103 and older): too little memory and bandwidth.
- TP=3 and TP=8: the model has 4 KV heads, which must split evenly. Use TP=1, 2 or 4, or replicas with data parallelism.
- Virtual machines with GPU passthrough for multi-GPU Radeon recipes: RCCL initialization has been reported to hang. Single-GPU recipes (B, F, G, P) don't use RCCL; for two cards, use bare metal or Recipe Z.
- Windows and WSL: not covered; vLLM's ROCm images are Linux-only.
- NVIDIA GPUs: use NVFP4-MIXED.
- llama.cpp, Ollama, LM Studio: this is not a GGUF.
Vision (Instinct and RDNA 3 only)
Every recipe starts text-only, which leaves the most memory for context. On RDNA 4 the vision tower currently fails to load in vLLM (vllm#49851), so keep --language-model-only there. Elsewhere, to serve images:
- Set
AEON_VISION=1with the script, or edit thedocker run:- remove
--language-model-only; - add
--limit-mm-per-prompt '{"image":1,"video":0}' --mm-processor-kwargs '{"max_pixels":1048576}'; - on Radeon / Strix Halo, add
-e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUEbefore the image name.
- remove
- The vision tower adds about 0.9 GB, split across the cards.
The image-size cap matters. Without it, vLLM sizes the encoder's startup memory test for this processor's maximum of about 16.7 megapixels (about 16K tokens per image), which can run out of memory. With the cap it plans for about 1,024 tokens per image. Vision is untested with this checkpoint.
Troubleshooting
The full symptom → cause → fix table is in AGENTS.md. The most common ones:
| Symptom | Likely cause | Fix |
|---|---|---|
AttributeError: 'dict' object has no attribute 'endswith' in quark.py |
vLLM image v0.28–v0.30 | Use vllm/vllm-openai-rocm:v0.31.0 |
The launch script stops with vLLM detects 'gfx1036' (or another iGPU target) |
An integrated GPU is listed first, so vLLM takes its architecture | AEON_DEVICES=auto, or pass only the discrete /dev/dri/renderD* and card* nodes, or disable the iGPU in the BIOS. HIP_VISIBLE_DEVICES alone doesn't help |
invalid device function / hipErrorNoBinaryForGpu |
Kernels built for a sibling target (for example gfx1201 on a gfx1200 card) or HSA_OVERRIDE_GFX_VERSION set |
Use the scripts (they set PYTORCH_ROCM_ARCH to the detected GPU), or set it to your target in a plain docker run; never set HSA_OVERRIDE_GFX_VERSION |
| Startup or the first request hangs with both GPUs at 100%, or responses are empty (TP ≥ 2); a serve script exits with code 3 ("PROBABLY HUNG": no log line, compiler or kernel-cache write for 20 minutes) | RCCL over PCIe (vllm#40980, ROCm/ROCm#6148) | During an event: save the log, docker rm -f aeon-mxfp4, start Recipe Z directly (skip A-MTP / A-plain; the hang doesn't depend on speculation). Later, with the host owner: check NCCL_PROTO=Simple and NCCL_P2P_DISABLE=1 are set; both cards in equal-width CPU slots; IOMMU on with iommu=pt (the community guide also uses pcie_aspm=off); bare metal. Exit code 2 means still compiling: wait (AGENTS.md section 0, hang procedure) |
hipIpcGetMemHandle ... invalid argument |
P2P IPC between the cards | NCCL_P2P_DISABLE=1, then Recipe Z |
| Stuck in compilation for over 30 minutes, one CPU core busy | Triton autotuning the Gated DeltaNet kernels on first start | Wait, then retry with AEON_EXTRA_ENV="VLLM_TRITON_FORCE_FIRST_CONFIG=1" |
out of resource: shared memory, Required: N, Hardware limit: 65536 |
A Triton kernel config exceeds RDNA's 64 KB LDS (AITER, if it was turned on; otherwise note the kernel name in the traceback) | Confirm VLLM_ROCM_USE_AITER=0 (the scripts set it). Then AEON_ATTN=default on a recipe without speculation (untested), else Recipe Z; send the log |
... larger than the available KV cache memory or No available memory for the cache blocks |
Not enough VRAM left for KV | Lower --max-model-len, then --max-num-seqs; keep --language-model-only; close other GPU apps |
| Speculative acceptance near 1 with several requests in flight | Draft model on the ROCM_ATTN backend | "attention_backend":"TRITON_ATTN" inside --speculative-config |
| DFlash2 crashes or acceptance drops to 0 after a repeated prompt | Prefix caching with this hybrid model | --no-enable-prefix-caching (Recipe A sets it; don't combine DFlash2 with prefix caching unless you are testing AEON_DFLASH_PREFIX_FIX=1) |
Launch script: ... emulation.py is not the v0.31.0 file the MXFP6 dequant-once patch was made from |
An image whose file differs from v0.31.0's | Use vllm/vllm-openai-rocm:v0.31.0, or AEON_MXFP6_ONCE=0 (same output, much slower) |
DFlash2DraftModel missing from the log, or acceptance like DFlash 1 |
Drafter folder missing, or VLLM_USE_V2_MODEL_RUNNER=0 |
ls $AEON_DIR/dflash2 (config.json + model.safetensors); keep VLLM_USE_V2_MODEL_RUNNER=1 |
| Endless thinking or repetition | Client sends temperature: 0 |
Send 0.6 or omit it; cap max_tokens; enable_thinking: false for tool loops |
| Decode speed about 20% lower than a previous run | R9700 decode speed can land lower for the life of a process (legacy-rocm-build#6347) | Restart the server and measure again |
| Tool calls come back as plain text | Tool parser not active | --enable-auto-tool-choice --tool-call-parser qwen3_coder; the client must send tools |
Expected performance (estimates, not measurements)
Nobody has measured this checkpoint on AMD hardware yet. The figures below combine measurements of this exact checkpoint in vLLM (MXFP6 emulated; an equivalent W4A4 MLP path with the same MXFP4 weight grid and the same activation quantize-dequantize, which differs from vLLM's emulation only in GEMM accumulation order, so it is not bit-identical; one request, 384-token outputs; see Quality and validation status) with a memory-bandwidth scaling to each AMD system. Treat every figure as roughly ±50%.
| 2x R9700 recipe | Measured (reference run) | Estimate, 2x R9700, MXFP6 patch on (default) | Estimate, stock emulation (AEON_MXFP6_ONCE=0) |
|---|---|---|---|
| A-plain | 8.1 tok/s (4 prompts) | ~10 tok/s | ~2 tok/s |
| A-MTP (K=3) | mean acceptance length 2.76 at T=0.6 (13 prompts); 13.5 tok/s on the same 4 prompts as plain | ~24 tok/s | ~5 tok/s |
| A (DFlash2, K=9) | 3.73 at T=0.6 (26 prompts), 23.2 tok/s (22.0 on the same 4 prompts as plain); 4.40 at T=0, 27.3 tok/s | ~35 tok/s (T=0.6), ~41 (T=0) | ~7 tok/s |
| A, 8 concurrent requests | 113 tok/s aggregate (4.9x its single-stream rate) | several times the single-stream rate; unmeasured | n/a |
| Z (per stream) | n/a | n/a | ~1 tok/s |
How the estimates are worked out:
- Per forward pass. In the reference run, stock MXFP6 emulation took about 2.1 s per pass (12–20 ms per Gated DeltaNet projection call, 144 calls), MXFP4 emulation about 0.32 s, and the BF16 rest about 0.04 s, for about 2.47 s in total. With the MXFP6 weights dequantized once, the MXFP6 part drops to the 0.044 s the same layers cost in BF16 (about 0.40 s per pass); with the MXFP4 MLP also in BF16 after load, about 0.21 s. Each AMD estimate scales that cost by memory bandwidth, divides by the TP size, and adds about 8 ms of PCIe all-reduce at TP=2 (about 12 ms at TP=4). On 2x R9700 that gives about 94 ms per pass (about 10.6 tok/s) with the patch and about 0.53 s (about 1.9 tok/s) without it.
- Speculation. Speedup = mean acceptance length ÷ (cost of one draft-plus-verify step ÷ cost of one plain step). Under emulation the verify step costs little more than a plain one, because the weight dequantization dominates and doesn't depend on how many tokens are checked; the step ratio is about 1.14 for DFlash2 and about 1.2 for MTP. So DFlash2 ≈ 3.73 ÷ 1.14 ≈ 3.3x plain, MTP ≈ 2.3x. When a pass gets cheap (both BF16 copies, many GPUs) the drafter's own cost weighs more and the gain shrinks.
- Acceptance on this target. The drafter accepts about 5–7% less on this MXFP4/MXFP6 checkpoint than on the NVFP4 build it was fine-tuned against (3.73 vs 4.01 at T=0.6, same prompts); at 4 and 8 concurrent requests the two were on par.
- Concurrency. Weight traffic is per forward pass, not per request, so aggregate throughput rises with concurrency until compute or memory runs out. 16 concurrent requests are unmeasured; the tester matrix measures them.
- Not modelled: RDNA 4 kernel efficiency, HIP-graph capture of the Quark ops, and RCCL behaviour with
NCCL_P2P_DISABLE=1. Read the A-plain time per token first; it calibrates everything else.
Other systems (single stream, default recipe path; estimates ±50%):
| Recipe | GPUs | Default MX path | Plain (est.) | With MTP (est.) | With DFlash2 (est.) |
|---|---|---|---|---|---|
| K | 4x R9700 | once + at load | ~25–30 tok/s | ~60 | ~60–90 (default) |
| B | 1x R9700 | stock | ~1 | ~2 | doesn't fit |
| C | 2x RX 9070 XT (2x RX 9060 XT) | stock | not advised | doesn't fit | |
| D | 4x RX 9070 XT | once | ~18 (stock ~3.6) | ~40 | ~50 |
| E | 2x RX 7900 XTX | once | ~15 (stock ~2.8) | ~35 | doesn't fit |
| M | 2x RX 7900 XT | stock | ~2.3 (opt-in once ~12) | ~5 | doesn't fit |
| L | 2x 16 GB RDNA 3 | stock | ~1.7 | not advised | doesn't fit |
| F | W7900 | once | ~8 (stock ~1.3) | ~18 | ~25 |
| Q | 2x W7900 | once + at load | ~24 | ~55 | ~70 (default) |
| G | W7800 32 GB | stock | ~0.9 | ~2 | doesn't fit |
| P | V710 | stock | ~0.7 | doesn't fit | doesn't fit |
| H | Strix Halo 64 GB | once | ~2.3 (stock ~0.4) | ~5 | tight |
| H-128 | Strix Halo 128 GB | once + at load | ~4.5 | ~10 (default) | ~14 |
| N | MI210 / MI250 (per GCD) | once | ~14 | ~33 | ~45 |
| I, O, J | MI300X / MI325X, MI300A, MI350X / MI355X | once + at load | launch-bound: no estimate |
At Instinct bandwidths (about 5.3 TB/s for MI300X / MI300A, 8 TB/s for MI355X) kernel-launch overhead, not memory bandwidth, likely sets the single-stream limit, so no figure is given; MI210 / MI250 figures are optimistic for the same reason.
For scale only: Puget Systems reports 15.9 tok/s stock and 62.8 tok/s tuned with MTP at concurrency 1 on 2x R9700 with Qwen3.6-27B in FP8 (article). That is a different model and a different format with no emulation.
Testers: the step-by-step tester kit ships with the files, in TESTER_GUIDE.md, tests/run_radeon_eval.py and scripts/ (start with scripts/preflight_smoke.sh). Agents: AGENTS.md.
Expected behavior and limits
- Speed. Under emulation, decode speed is limited by how many bytes each forward pass moves, not by the 23.7 GB of packed weights. See Expected performance. Tensor parallelism, speculative decoding (DFlash2, MTP) and concurrent requests all spread that cost, and the MXFP6 dequant-once patch (plus MXFP4 dequant-at-load on the largest GPUs) removes most of it wherever it fits, without changing the output. Prefill is affected much less. On MI350/MI355 the MLP can also run natively (opt-in, not bit-identical).
- Memory.
- About 22.8 GB of weights load with the vision tower, or about 21.9 GB (20.4 GiB) with
--language-model-only. The MTP head (0.85 GB) loads only when you turn on MTP. - The KV cache costs 64 KiB per token at BF16 (16 full-attention layers × 4 KV heads × 256 × 2 × 2 bytes), divided by the TP size. Every recipe uses BF16 KV, the reference numerics; FP8 KV (32 KiB) is an opt-in that changes outputs.
- Each running sequence also holds Gated DeltaNet state: about 147 MiB per state slot, FP32 as the model config sets it, divided by the TP size. A running sequence holds about 1 + K slots without prefix caching (K = draft tokens; 1 slot without speculation), and about 2 + K + 1 with prefix caching on (vLLM's default for this model, "align" mode). At TP=2 that is about 0.22 GiB per GPU for A-plain and about 0.45 GiB with MTP K=3 (see AGENTS.md section 12.2).
- The
GPU KV cache sizeline in the startup log is the real capacity.
- About 22.8 GB of weights load with the vision tower, or about 21.9 GB (20.4 GiB) with
- Speculative decoding is untested on AMD (measured Ï„ for this checkpoint: see Quality and validation status). DFlash2: see Recipe A. MTP: see Recipe A-MTP; use
{"method":"mtp",...}(the older nameqwen3_5_mtpstill works but logs a deprecation warning). Always add"attention_backend":"TRITON_ATTN"to the speculative config. - Thinking.
- Thinking is on by default, and the chat template's default reasoning effort is
xhigh. - For quick answers, send
"chat_template_kwargs": {"enable_thinking": false}. - For short thinking, send
{"enable_thinking": true, "reasoning_effort": "low"}. - Set both through
chat_template_kwargs, not through a top-level request field.
- Thinking is on by default, and the chat template's default reasoning effort is
- Inherited behavior. The BF16 master is an Early Access Draft, and very long answers can fall into loops. 4-bit MLP weights, including the residual-writing
down_proj, may make that more likely. The tester kit checks for loops explicitly. - Uncensored. This model writes what the base model refuses. Read User responsibility.
- Runtimes.
- vLLM is the only supported path.
- This is not a GGUF, so llama.cpp, Ollama and LM Studio can't load it.
- Transformers with
amd-quarkmay load it in emulation, but that is untested. - Apple needs an MLX conversion.
- Known ROCm issues.
- Decode speed on the R9700 can land about 20% lower for the life of a process (legacy-rocm-build#6347); restart and re-measure.
- A TP=2 hang on dual R9700 is still reported open (vllm#40980); Recipes A, A-MTP and A-plain use the settings of the working community setups (
NCCL_PROTO=Simple, plusNCCL_P2P_DISABLE=1), and Recipe Z avoids inter-GPU communication entirely. - vLLM v0.31.0 ships ROCm 7.2.3, which has no published TP=2 reports on R9700 yet; v0.28.0 (the tag reported working at TP=2 on R9700) uses the same ROCm 7.2.3 base and is the escape hatch for vLLM-level errors.
- vLLM v0.28–v0.30 can't load this checkpoint as shipped (
algo_configcrash in the Quark loader); use v0.31.0, or theAEON_STRIP_ALGO_CONFIG=1escape hatch. - With this hybrid model, DFlash2 speculative decoding needs prefix caching off on vLLM 0.29–0.31 (vllm#55601, vllm#58894).
- The first-start Gated DeltaNet compile hang on RDNA 4 (vllm#45929) is fixed in the images these recipes use.
Quality and validation status
Acceptance and quality were measured on an NVIDIA reference run of this exact checkpoint; AMD numbers are pending the first AMD runs.
| Check | Where | Status |
|---|---|---|
| Export integrity: 336 MX layers written; all 15 MTP tensors and 333 vision tensors carried over in BF16 | Quantization run | Done |
vLLM code path: Quark OCP MX loader, native or emulated kernel selection on each platform, TP=2/TP=4 MX sharding, MTP exclusion, algo_config handling (crashes on ≤ 0.30, fixed in 0.31) |
vLLM 0.29.0 and 0.31.0 source review | Done (a code review, not a runtime test) |
| This checkpoint in vLLM with emulated MXFP6 and an equivalent W4A4 MLP path (same MXFP4 weight grid and activation quantize-dequantize as vLLM's emulation; differs only in GEMM accumulation order, so not bit-identical): plain, MTP (K=3) and DFlash2 (K=9) | Reference run (TP=1, FP8 KV cache, vLLM 0.29-based) | Done. Mean acceptance length: DFlash2 3.73 at T=0.6 (26 prompts) and 4.40 at T=0; MTP 2.76 vs DFlash2 4.20 on the same 13 prompts. Single request, same 4 prompts: plain 8.1 tok/s, MTP 13.5, DFlash2 22.0 (2.7x). GSM8K-20 / tools-10: plain 20/20 and 10/10, DFlash2 19/20 and 9/10 (two discordant items, McNemar p = 0.5). Speculative outputs are not bit-identical to plain decoding at T=0 (1 of 20 GSM8K outputs identical; verify batches change the reduction order). No garbage output, no refusals |
| MXFP6 dequant-once patch: output identical to stock emulation | Reference run | Done (torch.equal on the layer outputs) |
| Apple Silicon, via an MLX conversion | Mac | Pending |
| W4A4 / W6A6 quality against the BF16 master, using emulated MX numerics | GPU, vLLM emulation | Pending |
| 2x Radeon AI PRO R9700, tester matrix (A-plain, A-MTP, A with DFlash2): smoke tests, GSM8K-50, IFEval-20, 10 tool calls, throughput and acceptance at c=1/4/8/16, TTFT, VRAM per GPU | ROCm 7.2.3, vLLM 0.31.0 | Pending (first tester run in progress) |
TP=2 vs TP=1 numerical cross-check (MX weight sharding): prefill logprobs of A-plain vs one Recipe Z replica, scripts/tp_check.py |
2x R9700 | Pending (part of the tester matrix) |
| Other Radeon and Ryzen systems (Recipes B–H, H-128, K–Q) | n/a | Not yet tested |
| Instinct MI300X / MI325X / MI300A / MI250 / MI210 (emulated) and MI350X / MI355X | n/a | Not yet tested |
Quantization recipe
| Module group | Linear layers | Format | Weights | Activations | Approx. size |
|---|---|---|---|---|---|
MLP gate_proj / up_proj / down_proj |
192 (64 layers × 3) | MXFP4 | FP4 E2M1, 32-element blocks, E8M0 scale | MXFP4, dynamic per 32-element block | 9.1 GB |
GDN in_proj_qkv / in_proj_z / out_proj |
144 (48 layers × 3) | MXFP6 E2M3 | FP6 E2M3, 32-element blocks, E8M0 scale | MXFP6 E2M3, dynamic | 4.3 GB |
Full attention q/k/v/o_proj, plus q_norm/k_norm |
64 (16 layers × 4) | BF16 | n/a | n/a | 3.4 GB |
GDN recurrence: in_proj_a/b, conv1d, A_log, dt_bias, norm |
96 linears + params | BF16 | n/a | n/a | 0.05 GB |
embed_tokens and lm_head (untied) |
n/a | BF16 | n/a | n/a | 5.1 GB |
| Vision tower (27 blocks + merger) | n/a | BF16 | n/a | n/a | 0.9 GB |
MTP head (1 layer + fc) |
n/a | BF16 | n/a | n/a | 0.85 GB |
- Tool: AMD Quark 0.12.post1 in eager mode, with Transformers 5.14.1 and PyTorch 2.13. The export is
real_quantized: packeduint8weights plusuint8E8M0 scales. The Quark MX settings arescale_calculation_mode: evenandround_method: half_even. - Algorithm: AutoSmoothQuant with an MSE scale search, applied to the MLP only, along the edges
post_attention_layernorm → gate/upandup_proj → down_proj. The smoothing factors are folded into those norms and weights. The GDN projections have no safe smoothing edge, so they get none. - Calibration: 1,024 chat, code and math samples (UltraChat, Open-Platypus, CodeAlpaca, GSM8K), each truncated to 64 tokens, for 65,536 tokens in total. To fit in memory, the AutoSmoothQuant scale search used a 64-token subsample per layer. MX weight scales come from the weights themselves and activations are quantized at runtime, so calibration data only affects the MLP smoothing factors.
- Exact config: see
quantization_configinconfig.jsonandBAKE_STAMP_Q2.json.
Model family
| Variant | Format | Size | Best hardware | Status | Link |
|---|---|---|---|---|---|
| BF16 master | BF16, with vision and MTP | 55.6 GB | H200, multi-GPU, RTX PRO 6000 | Public | AEON-7/…-BF16 |
| NVFP4-MIXED | NVFP4 MLP (layers 0–55); FP8 attention, GDN writers and MLP layers 56–63; BF16 for the rest (ModelOpt) | 24.7 GB | DGX Spark (GB10), RTX 5090, RTX PRO 6000 | Public | AEON-7/…-NVFP4-MIXED |
| MXFP4-MXFP6-ROCm (this repo) | MXFP4 MLP and MXFP6 GDN projections; BF16 for the rest (AMD Quark) | 23.7 GB | Instinct MI350/MI355 (native MXFP4); Radeon AI PRO R9700 32 GB (emulated) | Early Release · Gated Access | AEON-7/…-MXFP4-MXFP6-ROCm |
AEON DFlash2 drafter (bundled in dflash2/ of this repo) |
DFlash 2 block-diffusion drafter, BF16 (AEON fine-tune of z-lab/Qwen3.8-27B-DFlash2) | 3.85 GB | Ships with this repo for its DFlash2 recipes (featured: 2x Radeon AI PRO R9700). The bundle is the earlier lk3 BF16 build; the newer v1.0 (lk4) is in AEON-7/AEON-DFlash2-Qwen3.8-27B (bf16/ folder) |
Early Release · Gated Access; untested on AMD | dflash2/ (lk3) · v1.0 (lk4) |
| AEON DFlash2 drafter v1.0 (lk4) | DFlash 2 drafter: NVFP4 W4A16 pack + exact BF16 copy (bf16/) |
1.9 GB / 3.85 GB | DGX Spark, RTX 50-series (vLLM) | Early Release · Gated Access | AEON-7/AEON-DFlash2-Qwen3.8-27B |
| NVFP4-GDNFP8 | NVFP4 MLP with FP8 GDN projections (LLM Compressor) | TBA | NVIDIA Blackwell | Early access (Patreon); not public | n/a |
| FP8-MIXED | FP8 dynamic MLP; BF16 for the rest (LLM Compressor) | 38.5 GB | FP8-capable GPUs with 48 GB or more | Internal: built, not yet validated | n/a |
| MLX (Apple Silicon) | MLX conversion | TBA | Apple Silicon Macs | Planned | n/a |
User responsibility
By accessing, downloading or running this model, you agree to the following:
- You are responsible for its use. You alone are responsible for every prompt, every response, every downstream action, and any harm that results.
- No warranty. The model is provided "AS IS", without warranty of any kind.
- Follow the law. You must comply with all applicable laws and policies in every jurisdiction you operate in.
- Add safety layers in production. Use input validation, output filtering, access controls, and human review for high-risk workflows.
- The duty of care is yours. An uncensored model doesn't refuse on your behalf, so if you are unsure about a request, don't make it.
- No endorsement. The authors do not endorse any particular output.
- Arbitration. Disputes go to binding individual arbitration (under the AAA Consumer Rules if no other body is agreed), waiving jury trials and class actions.
- Indemnification. You indemnify the authors, contributors and publishers against claims arising from your use.
- Severability. If a provision is invalid, the closest enforceable equivalent replaces it.
- Acceptance. Using the model means you accept these terms. If you don't accept them, don't use the model.
License and attribution
- License: Apache-2.0, inherited from Qwen/Qwen3.8-27B.
- Qwen team (Alibaba): the Qwen3.8-27B base model.
- AEON-7: the uncensored fine-tune (the BF16 master) and this quantization. The upstream tools behind the fine-tune are credited on the BF16 card.
- AMD Quark: the quantization toolkit used to produce the OCP MX export.
- Downloads last month
- 4
Model tree for AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm
Base model
Qwen/Qwen3.8-27B