Request early access: AEON Ultimate MXFP4/MXFP6 for ROCm

Experimental early-access build, not yet tested on AMD hardware. Requests are reviewed by hand every few hours; include the community access word if you were given one.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm

AEON MIXED quant lattice - MXFP4 MLP, MXFP6 GDN writers, BF16 attention/vision/MTP

The AEON Ultimate uncensored 27B, quantized to OCP Microscaling (MX) formats for AMD GPUs. The MLP runs in MXFP4, the Gated DeltaNet projections in MXFP6, and everything that carries the model's skills stays in BF16. The checkpoint is 23.7 GB on disk and loads as about 22.8 GB of weights. Two 32 GB Radeon AI PRO R9700 cards hold it with the full 262,144-token context; one R9700 holds it with a shorter context. MXFP4 is also the format Instinct MI350 / MI355 execute natively. The Hub's file panel labels this repo "8-bit" only because the MX weights are stored packed in U8 tensors; the real mix is MXFP4 MLP + MXFP6 Gated DeltaNet projections + BF16 for the rest, about 6.8 bits per weight on average: a 23.7 GB download that needs at least 32 GB of total VRAM.

Recommended on 2x R9700: vLLM with tensor parallelism across both cards, this checkpoint unchanged, and the bundled AEON DFlash2 drafter for lossless speculative decoding (Recipe A).

Experimental: early access. Nobody has published runtime results for this build yet. The first Radeon test pass (2x R9700) is under way, and quality and speed numbers will appear in this card as they come in. Requests are reviewed by hand, usually within a few hours; include the community access word in the form if you were given one.

At a glance

Why it's great

  • Quality first. Only two parts of the model are quantized: the MLP (MXFP4, W4A4) and the Gated DeltaNet projections (MXFP6, W6A6). Full attention, the recurrence, embeddings, lm_head, the vision tower and the MTP head all stay in BF16. That is about 6.8 bits per weight: 23.7 GB, against 55.6 GB for the master.
  • Full context on two 32 GB cards. 2x Radeon AI PRO R9700 hold it with the full 262,144-token context. Instinct MI350X / MI355X can also run the MXFP4 MLP natively, as an opt-in.
  • Speculative decoding ships with it. The repo includes the AEON DFlash2 drafter (dflash2/, BF16) alongside the model's own MTP head, and both are lossless. In the reference run, DFlash2's mean acceptance length was 3.73 at T=0.6, and it ran 2.7x faster than plain decoding on the same 4 prompts (details).
  • Turnkey across the ROCm lineup. There is a launch script per setup in scripts/, and detect_setup.py picks the right one for your GPUs. The MXFP6 dequant-once patch gives the same output as stock vLLM with a much faster forward pass.
  • AEON Ultimate, uncensored. It is the full 27B hybrid (thinking, tool calling, and vision on Instinct and RDNA 3: see Vision), and it writes what the base model refuses. Read User responsibility.

Target systems (all Untested on AMD hardware so far)

Setup Recipe
2x Radeon AI PRO R9700 (featured) A (DFlash2) → A-MTP → A-plain → Z
4x / 1x R9700 (R9700S, R9600D) K (DFlash2) · B
RDNA 4 16 GB (RX 9070 XT / 9070 / 9060 XT), 2x or 4x C · D
RDNA 3: 2x RX 7900 XTX, 2x RX 7900 XT, 2x 16 GB E · M · L
Radeon PRO W7900 / W7800 48 GB, 2x W7900, W7800 32 GB, V710 F · Q (DFlash2) · G · P
Strix Halo (Ryzen AI Max+ 395) 64 / 128 GB H · H-128
Instinct MI300X / MI325X, MI300A, MI350X / MI355X, MI210 / MI250 I · O · J · N

Recommended configuration

  • 2x R9700: Recipe A. Launch it with bash "$AEON_DIR/scripts/serve_2x_r9700_dflash2.sh". It runs:

    • vLLM v0.31.0 (ROCm image) at TP=2, with the bundled drafter at K=9 and the lattice [[1,2,9],[3,8,7],[9,12,6],[13,16,4]];
    • TRITON_ATTN, including inside --speculative-config, with VLLM_USE_V2_MODEL_RUNNER=1 and prefix caching off.

    The estimate is about 35 tok/s single-stream, ±50% until measured.

  • Long multi-turn agent loops: A-MTP keeps prefix caching on.

  • Exact baseline, or the most concurrent requests: A-plain.

  • TP=2 hangs: Z.

  • Any other setup: run python3 "$AEON_DIR/scripts/detect_setup.py".

Using the bundled DFlash2 drafter with other engines

dflash2/ is a standard DFlash2 checkpoint. It has the same 81 tensors (names, shapes and dtypes) as z-lab/Qwen3.8-27B-DFlash2. Its config differs only in block_size 10, causal: false and draft_vocab_size. By format, any engine that runs z-lab's drafter loads it if you point its drafter path at dflash2/ (only vLLM has been run with an AEON target, and not yet on AMD). Set the draft width to 9: num_speculative_tokens 9 in vLLM, --speculative-num-draft-tokens 10 in SGLang (the block size), or --spec-draft-n-max 9 in llama.cpp. For this mixed checkpoint, vLLM is the only engine that loads the target unchanged (see TensorFold). Full guide: dflash2/README.md.

What this is

  • Base model: AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16 is AEON-7's uncensored build of Qwen/Qwen3.8-27B. It is a hybrid of 48 Gated DeltaNet (linear-attention) layers and 16 full-attention layers, with a vision tower, an MTP head and a native context of 262,144 tokens.

  • Format: OCP Microscaling. Every block of 32 values shares one power-of-two (E8M0) scale. We used AMD Quark 0.12.post1 and exported real packed weights, not fake-quant.

    • MLP (gate_proj, up_proj, down_proj) is MXFP4, W4A4: FP4 E2M1 weights, with activations quantized to MXFP4 on the fly.
    • Gated DeltaNet projections (in_proj_qkv, in_proj_z, out_proj) are MXFP6 E2M3, W6A6.
    • The rest stays BF16: full attention, embeddings, lm_head, the vision tower, the MTP head, the GDN recurrence (conv1d, in_proj_a/b, A_log, dt_bias) and the norms. Precision lattice (Quark ROCm): MLP MXFP4 - GDN writers MXFP6 - attention / vision / MTP / embed / lm_head / GDN recurrence BF16.
  • Unchanged from the BF16 master: the tokenizer, chat template and vision processor. Thinking is on by default.

Why MX on AMD? MX is the open block-scaled format that AMD's newest accelerators run in hardware. On Radeon the gain is footprint: the BF16 master is 55.6 GB, while this build is 23.7 GB. The precision is mixed on purpose. The MLP and GDN projections take 4-bit and 6-bit weights, and the attention path, the recurrence, the embeddings and the vision tower stay in full BF16.

Target hardware

Every AMD GPU configuration that runs ROCm compute and can hold this checkpoint has its own labelled quickstart below. The featured setup is 2x Radeon AI PRO R9700 (Recipe A and its fallbacks).

  • Native MXFP4: Instinct MI350X / MI355X (gfx950) can run the MLP on AITER's W4A4 GEMM. That path uses the same MX grid but is not bit-identical to the other recipes, so Recipe J defaults to the exact emulated path and offers native MXFP4 as a labelled opt-in. MXFP6 is emulated on every GPU, because vLLM has no MXFP6 kernel yet.
  • Emulated: everything else, including Radeon RDNA 4 (R9700, RX 9070 / 9060 XT), RDNA 3 (RX 7900 / 7800 / 7700, W7900, W7800, W7700, V710), Strix Halo and Instinct MI300X / MI325X / MI300A / MI250 / MI210. vLLM keeps the packed MX weights in memory, dequantizes each layer to BF16 on every forward pass, quantizes and dequantizes the activations to the same MX grid, and runs the matmul in BF16. The numerics match W4A4 and W6A6 and memory stays near the packed size, but decoding is slower than a native kernel would be (vLLM 0.29 and 0.31 source: QuarkOCP_MX, EmulationMxfp4LinearKernel, EmulationMxfp6LinearKernel; native MX needs supports_mx(), which is true only on gfx95x and gfx1250). Wherever memory allows, the recipes remove most of that cost without changing a bit of the output: the MXFP6 weights are dequantized once at load (a small vLLM patch in scripts/patches/) and, on the largest GPUs, the MXFP4 MLP is kept in BF16 after load (VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD).
Other platform Runs? Notes
NVIDIA GPUs Not supported by this build Use NVFP4-MIXED.
Apple Silicon Not with this checkpoint An MLX conversion is being tested; results will be added here.

Quickstart

Deploying with an AI agent (Claude Code, Codex, Cursor)? Point it at AGENTS.md. It has the preflight, every recipe with exact commands, verification steps and an exhaustive troubleshooting table.

Testing before an event? Run bash "$AEON_DIR/scripts/preflight_smoke.sh" the day before. It checks the host, starts Recipe A (falling back to A-MTP, A-plain, then Z; a TP=2 start that hangs goes straight to Z), smoke-tests it in a few minutes and bundles the results to send back. mkdir -p ~/aeon-test && cd ~/aeon-test && bash "$AEON_DIR/scripts/run_matrix.sh" then compares plain, MTP and DFlash2 at 1–16 concurrent requests and checks TP=2 against a one-GPU reference.

Step 0: Which setup do you have? (1 minute)

Run this on the Linux host. It reads the amdgpu driver directly, so it needs no ROCm tools, no Docker and no root:

# one line per AMD GPU: gfx target, VRAM, render node, PCI name
for p in /sys/class/kfd/kfd/topology/nodes/*/properties; do
  v=$(awk '$1=="gfx_target_version"{print $2}' "$p"); [ "${v:-0}" -gt 0 ] || continue
  r=$(awk '$1=="drm_render_minor"{print $2}' "$p"); d=/sys/class/drm/renderD$r/device
  printf 'gfx%d%x%x  %5.1f GiB  /dev/dri/renderD%s  %s\n' $((v/10000)) $((v/100%100)) $((v%100)) \
    "$(awk '{print $1/2^30}' "$d/mem_info_vram_total")" "$r" \
    "$(lspci -s "$(basename "$(readlink -f "$d")")" 2>/dev/null | cut -d' ' -f2-)"
done

With ROCm tools on the host, amd-smi static --asic --vram or rocm-smi --showproductname --showmeminfo vram show the same thing. After the download, python3 "$AEON_DIR/scripts/detect_setup.py" does this lookup and prints the matching recipe and launch script, and bash "$AEON_DIR/scripts/preflight.sh" checks the rest of the host (driver, groups, PCIe slots, IOMMU, Resizable BAR, disk, ports).

Match the output to the table below. Count only the discrete GPUs.

  • Ignore an integrated GPU (gfx1036, gfx1035, gfx1103, gfx1150 and similar, with a few GiB of VRAM), unless it gets passed into the container. vLLM reads the GPU architecture from the first GPU that amdsmi lists, and HIP_VISIBLE_DEVICES doesn't change that. If the iGPU comes first, the RDNA 4 paths switch off.
  • To hide it, either disable the iGPU in the BIOS, or pass only the discrete GPUs' nodes into the container: replace --device /dev/dri with --device /dev/dri/renderD129 --device /dev/dri/card1 ... for each discrete GPU. The launch scripts do this with AEON_DEVICES=auto, and they refuse to start if vLLM would see the wrong architecture.

Pick your system

Status: Verified = run with this checkpoint on that hardware (none yet). Untested = not yet run on that hardware; every recipe stays Untested until the first AMD runs come back. "Expected to work" means the vLLM code path for that architecture and GPU count is confirmed from source, and other Qwen3.x checkpoints are reported running in vLLM on the same GPU family with the same settings. Not supported = it doesn't fit, or vLLM can't run there.

MX path: once = MXFP6 dequantized once at load; at load = MXFP4 MLP also kept in BF16 after load; stock = both dequantized on every forward pass (no room for the BF16 copies). All three give identical output; they differ in speed and memory.

Your GPUs VRAM gfx MX path Status Recipe
2x Radeon AI PRO R9700 (featured) 2 × 32 GB gfx1201 once Untested A (DFlash2) → A-MTP → A-plain; fallback Z. Which one
4x Radeon AI PRO R9700 / R9700S / R9600D 4 × 32 GB gfx1201 once + at load Untested K (DFlash2)
1x Radeon AI PRO R9700 / R9700S / R9600D 32 GB gfx1201 stock (slow) Untested (expected to work) B
2x Radeon RX 9070 XT / RX 9070 / RX 9060 XT 16 GB 2 × 16 GB gfx1201 / gfx1200 stock (slow) Untested (expected to work; tight) C
4x 16 GB: RX 9070 XT / 9070 / 9060 XT, RX 7900 GRE, RX 7800 XT, RX 7700, Radeon PRO W7700 4 × 16 GB gfx1201 / gfx1200 / gfx1100 / gfx1101 once Untested D
2x Radeon RX 7900 XTX 2 × 24 GB gfx1100 once Untested E
2x Radeon RX 7900 XT 2 × 20 GB gfx1100 stock Untested M
2x 16 GB RDNA 3: RX 7900 GRE, RX 7800 XT, RX 7700, Radeon PRO W7700 2 × 16 GB gfx1100 / gfx1101 stock (slow) Untested (tight) L
Radeon PRO W7900 / W7900 Dual Slot / W7800 48 GB 48 GB gfx1100 once Untested F
2x (or 4x) Radeon PRO W7900 / W7800 48 GB; also 2x W7800 32 GB 2 × 48 GB gfx1100 once + at load Untested Q (DFlash2)
Radeon PRO W7800 32 GB 32 GB gfx1100 stock (slow) Untested G
Radeon PRO V710 (full GPU) 28 GB gfx1101 stock (slow) Untested P
Ryzen AI Max+ 395 / Max 390 / Max 385 (Strix Halo), 64 GB unified gfx1151 once Untested H
Ryzen AI Max+ 395 / Max 390 / Max 385 (Strix Halo), 128 GB unified gfx1151 once + at load Untested H-128
Instinct MI300X / MI308X / MI325X 192 / 256 GB gfx942 once + at load Untested I
Instinct MI300A 128 GB unified gfx942 once + at load Untested O
Instinct MI350X / MI355X 288 GB gfx950 once + at load (exact); native MXFP4 opt-in Untested J
Instinct MI210 / MI250 / MI250X 64 GB per GCD gfx90a once Untested N
Any single GPU under 28 GB, 8 GB cards, RDNA 2, MI100, RX 7600 series, NVIDIA Not supported why

Common setup (every recipe)

You need Linux on bare metal (not a VM with GPU passthrough for multi-GPU recipes: RCCL is reported to hang there), a recent amdgpu driver for your GPU, Docker, and about 100 GB of free disk (model 23.7 GB, drafter 3.85 GB, the image's 11.5 GB download and more once unpacked, kernel caches; about 12 GB more per extra image tag).

Every recipe uses the official vLLM ROCm image vllm/vllm-openai-rocm:v0.31.0 (ROCm 7.2.3; built for gfx90a, gfx942, gfx950, gfx1100, gfx1101, gfx1150, gfx1151, gfx1200 and gfx1201; includes amd-quark).

  • Don't use v0.28, v0.29 or v0.30, except through the AEON_STRIP_ALGO_CONFIG=1 escape hatch (AGENTS.md section 5.6). They crash while loading this checkpoint with AttributeError: 'dict' object has no attribute 'endswith', because their Quark loader can't handle the list-valued algo_config in config.json. v0.31.0 fixes it.
  • Don't use the nightly / rocm100 images on Radeon. They ship no gfx1201 kernels (ROCm/aiter#5229).
  • Pin the digest. As of 2026-10-04 vllm/vllm-openai-rocm:v0.31.0 is sha256:749f6f3f944f12af49966ac523c1f8573e4b229594b541954bd5f879b4496b1f. After the pull, docker image inspect vllm/vllm-openai-rocm:v0.31.0 --format '{{index .RepoDigests 0}}' should show it; if it shows another digest, the tag was re-pushed: tell us before relying on it. Then use AEON_IMAGE=vllm/vllm-openai-rocm@sha256:....
# 1) After you accept the access terms on the model page (log in with your own account and token)
python3 -m venv ~/.venvs/hf && ~/.venvs/hf/bin/pip install -U huggingface_hub && export PATH="$HOME/.venvs/hf/bin:$PATH"   # Ubuntu may need: sudo apt install python3-venv
hf auth login
hf download AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm \
  --local-dir ~/models/aeon-mxfp4-rocm
#   the download includes dflash2/ (3.85 GB): the DFlash2 drafter that Recipes A, K and Q use. Skip it with
#   --exclude "dflash2/*" only if you will not use DFlash2.

# 2) Variables the commands below use (re-run in every new terminal)
export AEON_DIR="$HOME/models/aeon-mxfp4-rocm"
export AEON_CACHE="$HOME/.cache/aeon-mxfp4-rocm"     # compiled kernels persist here
export RENDER_GID="$(getent group render | cut -d: -f3)"; [ -n "$RENDER_GID" ] || RENDER_GID=video
mkdir -p "$AEON_CACHE" ~/aeon-test && cd ~/aeon-test    # logs and results land in the current folder

# 3) Check the host, and see which recipe fits
bash "$AEON_DIR/scripts/preflight.sh"
python3 "$AEON_DIR/scripts/detect_setup.py"

Each recipe below has a one-line launch script (in scripts/, shipped with the model) and the plain docker run it executes. The scripts also check which GPUs and which vLLM version the container sees before starting, refuse known-bad combinations, wait for the server, and save the startup log. Every value can be changed through an environment variable, for example AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=4 bash "$AEON_DIR/scripts/serve_2x_r9700.sh"; AEON_DRY_RUN=1 prints the command without starting anything. The header of each script lists its target system, status and memory estimate. Check every start with Verify.

Applies to every recipe:

  • Don't pass --quantization. vLLM reads the Quark config from config.json.
  • The first start is slow (allow 10–30 minutes). vLLM loads the weights, then Triton compiles and autotunes the Gated DeltaNet and attention kernels, and Quark JIT-compiles its MX emulation kernel with hipcc. The cache mount keeps all of that for later starts.
  • PYTORCH_ROCM_ARCH makes Quark compile for your one GPU target instead of the nine the image lists. The scripts set it to the architecture the container detects. In a plain docker run, use your own target (for example gfx1200 for an RX 9060 XT, gfx1101 for an RX 7800 XT, W7700 or V710): a kernel built for a sibling target fails with invalid device function.
  • BF16 KV cache everywhere (--kv-cache-dtype auto). The checkpoint ships no calibrated KV scales, so FP8 KV would run with a scale of 1.0 and change the outputs. FP8 KV is an opt-in (AEON_KV_DTYPE=fp8) for 16 GB cards only.
  • VLLM_ROCM_USE_AITER=0 on every recipe. On Radeon, AITER's unified attention overflows RDNA 4's 64 KB LDS. On MI350/MI355 the opt-in native MXFP4 GEMM is picked without it.
  • --attention-backend TRITON_ATTN everywhere, and inside --speculative-config too. The draft model doesn't inherit the target's backend; vLLM's automatic pick on ROCm (ROCM_ATTN) collapses speculative acceptance under concurrency.
  • MXFP6 dequant-once patch wherever it fits (AEON_MXFP6_ONCE=1; Recipes A, A-MTP, A-plain, D, E, F, H, H-128, I, J, K, N, O, Q): the same output as stock vLLM, an estimated 5x faster forward pass, at the cost of a BF16 copy of 10.3 GiB in total, divided by the number of GPUs. MXFP4 dequant-at-load (AEON_DEQUANT_AT_LOAD=1; H-128, I, J, K, O, Q): bit-identical, +23.4 GiB in total, divided by the number of GPUs.
  • --override-generation-config is required. The shipped generation_config.json says temperature 1.0; the override sets the model card's 0.6 / 0.95 / 20.
  • Multi-GPU Radeon uses NCCL_PROTO=Simple (from the community R9700 TP=2 / TP=4 guide) and NCCL_P2P_DISABLE=1 (kept for robustness against hipIpc errors; that guide found it unnecessary). The scripts set both whenever TP > 1 on a Radeon or Ryzen GPU. The same guide boots with amd_iommu=on iommu=pt pcie_aspm=off.
  • Ready when the log prints Application startup complete. (docker logs -f aeon-mxfp4). Stop with docker rm -f aeon-mxfp4.

2x R9700: pick a recipe

All four recipes serve this checkpoint byte for byte with its own numerics: W4A4 MXFP4 MLP, W6A6 MXFP6 Gated DeltaNet projections, BF16 everywhere else, BF16 KV cache. A, A-MTP and A-plain differ only in speculative decoding, which is lossless in distribution (vLLM's standard rejection sampler keeps the target's output distribution). Nothing here has run on RDNA 4 hardware yet: the acceptance figures are measured Ï„ (mean acceptance length) for this exact checkpoint, drafter and speculative config, and the R9700 speeds are estimates.

Recipe Script Speculation Prefix caching Mean acceptance length (measured, T=0.6) Single-stream speed on 2x R9700 (estimate) Concurrent 16K-token requests (estimate) Status
A (primary) serve_2x_r9700_dflash2.sh AEON DFlash2 drafter, K=9 off 3.73 (4.40 at T=0) ~35 tok/s ~5–6 Untested
A-MTP (alternative) serve_2x_r9700_mtp.sh the model's own MTP head, K=3 on 2.76 ~24 tok/s ~10 Untested
A-plain (exact baseline) serve_2x_r9700.sh none on n/a ~10 tok/s ~15 Untested (expected to work)
Z (last resort) serve_2x_r9700_replicas.sh none on n/a ~1 tok/s per stream ~2 per card Untested (expected to work)
  • Which one. Start with A. It has the most tokens per forward pass, and under emulation every pass is expensive, so that is where speed comes from. For long multi-turn agent loops (each turn resends a growing conversation), also try A-MTP: DFlash2 needs prefix caching off on this hybrid model in vLLM 0.31 (see Recipe A), so every turn re-prefills the whole conversation, while A-MTP reuses the cached prefix. If many long-context agents share one server, A-plain and A-MTP hold more of them at once (table). scripts/run_matrix.sh measures all three on your machine (see Tester matrix).
  • Fall back in this order: A → A-MTP → A-plain → A-plain on the v0.28.0 image (only after a vLLM-level error, not an RCCL hang: AGENTS.md section 5.6) → Z. A-plain is also the reference the other two are compared against.
  • The MXFP6 dequant-once patch (scripts/patches/, on in A, A-MTP and A-plain). Stock vLLM re-dequantizes the MXFP6 Gated DeltaNet weights on every forward pass. That one step was measured at about 2.1 s per pass, 48x the same layers in BF16, which would cap the R9700 at roughly 2 tok/s single-stream (about 7 tok/s with DFlash2). The patch calls the same dequant function once at load and reuses the result: output identical to stock (checked with torch.equal), at the cost of about 5.2 GiB per GPU. The launcher mounts it read-only over the one vLLM file it replaces, only after checking that file is byte-identical to the v0.31.0 original. AEON_MXFP6_ONCE=0 turns it off.
  • How the speed estimates are made. Measured (one request, T=0.6): on 4 identical prompts plain 8.1, MTP 13.5 and DFlash2 22.0 tok/s (2.7x); on 13 identical prompts MTP Ï„ 2.76 vs DFlash2 Ï„ 4.20; DFlash2 over all 26 prompts Ï„ 3.73, 23.2 tok/s. Scaled to 2x R9700 by memory bandwidth (TP=2 halves the bytes per GPU), plus PCIe all-reduce: about 94 ms per forward pass with the patch. Speedup = mean acceptance length ÷ the relative cost of a drafting-plus-verify step (about 1.14 for DFlash2). Treat every R9700 tok/s figure as ±50% until someone measures it.
  • Concurrency. The weight traffic is per forward pass, not per sequence, so aggregate throughput rises with concurrent requests. DFlash2 delivered 4.9x its single-stream throughput at 8 concurrent requests (113 vs 23 tok/s, measured). c=16 is unmeasured. How many requests fit at once is set by the KV and Gated DeltaNet state pool (last column; the startup log's GPU KV cache size is the real number).

Verify (every recipe)

Run this after any launch, in the same terminal. For Recipe Z use NAME=aeon-mxfp4-0 (the proxy is on port 8080).

NAME=aeon-mxfp4; PORT=8000
AUTH=(); [ -n "${VLLM_API_KEY:-}" ] && AUTH=(-H "Authorization: Bearer $VLLM_API_KEY")   # an array, so the header stays one argument
# 1) wait until ready (first start: 10-30 minutes); stops and prints the log if the container died
until curl -sf "${AUTH[@]}" http://127.0.0.1:$PORT/health >/dev/null; do
  docker inspect -f '{{.State.Running}}' $NAME 2>/dev/null | grep -q true || { docker logs --tail 80 $NAME; break; }
  sleep 15
done
# 2) smoke: expect READY
curl -s "${AUTH[@]}" http://127.0.0.1:$PORT/v1/chat/completions -H "Content-Type: application/json" -d '{
  "model": "aeon", "max_tokens": 64,
  "messages": [{"role": "user", "content": "Reply with exactly: READY"}],
  "chat_template_kwargs": {"enable_thinking": false}}' | python3 -c 'import json,sys; print(json.load(sys.stdin)["choices"][0]["message"]["content"])'
# 3) speculative recipes (A, A-MTP, K, Q, any AEON_MTP / AEON_DFLASH run), after some traffic:
curl -s "${AUTH[@]}" http://127.0.0.1:$PORT/metrics | python3 -c '
import re, sys
v = {}
for line in sys.stdin:
    m = re.match(r"^(vllm:spec_decode_num_(drafts|draft_tokens|accepted_tokens))(?:_total)?(?:\{[^}]*\})? ([0-9.e+]+)$", line)
    if m: v[m.group(2)] = v.get(m.group(2), 0) + float(m.group(3))
d, a = v.get("drafts", 0), v.get("accepted_tokens", 0)
print("mean acceptance length", round(1 + a / d, 2) if d else "no drafts yet")'

Healthy: DFlash2 about 3 or more at one request, MTP about 2 or more. About 1 with several requests in flight means the draft is not on TRITON_ATTN. AGENTS.md section 6 has the full checklist (log lines, tool calls, reasoning field).

Recipe A: 2x R9700 with the AEON DFlash2 drafter (primary)

Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once at load; drafter in BF16 · Status: Untested on RDNA 4 (measured τ 3.73 at T=0.6 for this checkpoint, drafter and config) · Script: scripts/serve_2x_r9700_dflash2.sh · Needs: the dflash2/ folder (AEON-DFlash2-Qwen3.8-27B-BF16, 3.85 GB)

bash "$AEON_DIR/scripts/serve_2x_r9700_dflash2.sh"        # AEON_DFLASH=<dir> if the drafter lives elsewhere
Plain docker run for Recipe A
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/dflash2":/draft:ro \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 -e VLLM_USE_V2_MODEL_RUNNER=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
  --max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --no-enable-prefix-caching --language-model-only \
  --speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic","rejection_sample_method":"standard"}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

The patch path assumes vLLM lives at /usr/local/lib/python3.12/dist-packages/vllm in the image, as it does in v0.31.0. Check with docker run --rm --entrypoint python3 vllm/vllm-openai-rocm:v0.31.0 -c 'import vllm,os; print(os.path.dirname(vllm.__file__))'. The script finds the path itself and also checks the file's hash.

Why these settings:

  • The drafter. AEON-DFlash2-Qwen3.8-27B-BF16 is a DFlash 2 block-diffusion drafter fine-tuned for this model from z-lab's Qwen3.8-27B drafter (Apache-2.0). It reads hidden states from five target layers and proposes 9 tokens per step; the target checks them in one forward pass. It is BF16, so it runs on ROCm as is.
  • "num_speculative_tokens":9 matches the drafter's training block of 10 (vLLM drafts 1 + K positions). The per-batch-size lattice [[1,2,9],[3,8,7],[9,12,6],[13,16,4]] lowers K as more requests share a step; it is the schedule the acceptance was measured with. AEON_DFLASH_LATTICE= (empty) keeps K=9 at every batch size.
  • "attention_backend":"TRITON_ATTN" inside --speculative-config. The draft model doesn't inherit --attention-backend. vLLM's automatic pick on ROCm is ROCM_ATTN, which collapses DFlash2 acceptance once more than one request is in flight (vllm#53323: acceptance length 1.47 instead of 4.99 at 4 requests).
  • VLLM_USE_V2_MODEL_RUNNER=1. DFlash2 exists only in vLLM's V2 model runner. With the variable set to 0, vLLM quietly loads the drafter as DFlash 1.
  • "draft_sample_method":"probabilistic", "rejection_sample_method":"standard". Both are lossless. Probabilistic drafting is what was measured; AEON_DFLASH_SAMPLE=greedy is the alternative. Never use "synthetic": it is not lossless.
  • --no-enable-prefix-caching. In vLLM 0.29–0.31, a hybrid Gated DeltaNet model with a DFlash2 drafter crashes or drops to 0% acceptance after a prefix-cache hit (vllm#58894, fix proposed in vllm#55601). Every request therefore prefills its whole prompt. For one-shot work that costs nothing extra; for long multi-turn loops, compare against A-MTP. AEON_DFLASH_PREFIX_FIX=1 mounts the proposed one-line fix and keeps prefix caching on: experimental and untested.
  • --dtype bfloat16 avoids a drafter dtype mismatch reported on ROCm (vllm#42588) and fp16's 0% acceptance (vllm#55250).
  • TP=2, NCCL settings, --language-model-only, BF16 KV: as in Recipe A-plain.
  • Memory (estimate) per GPU: about 10.2 GiB of weights, 5.2 GiB for the MXFP6 BF16 copy, 1.8 GiB for the drafter (it shares the target's embeddings and lm_head), about 3 GiB for activations and graphs, and about 8 GiB left for KV cache and Gated DeltaNet state. vLLM pads the target's 16 attention layers to 20 to group them with the drafter's 5 layers (the log warns Add 4 padding layers, may waste at most 25.00% KV cache memory: expected), so KV costs about 40 KiB per token per GPU, and each running request also holds about 0.77 GiB of Gated DeltaNet state and about 0.1 GiB of drafter KV. Realistic capacity: about 5–6 concurrent 16K-token requests, about 4 at 32K. Expect GPU KV cache size around 150–200K tokens and Maximum concurrency for 131,072 tokens per request around 1.2–1.5x. More requests wait in the queue; with many long-context agents, AEON_MAX_NUM_SEQS=8 avoids preemption churn, or use A-plain / A-MTP, which hold more.
  • Pass check. The log shows Resolved architecture: DFlash2DraftModel, MXFP6 weights dequantized once at load, and under load SpecDecoding metrics: Mean acceptance length: ... of about 3 or more. About 1 with several requests in flight means the draft is not on TRITON_ATTN.
  • If it fails to load, or acceptance stays near 1, go to A-MTP and send us the log.

Recipe A-MTP: 2x R9700 with MTP (alternative)

Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once at load; the MTP head runs in BF16 · Status: Untested on RDNA 4 (measured τ 2.76 at T=0.6 for this checkpoint's MTP head) · Script: scripts/serve_2x_r9700_mtp.sh

bash "$AEON_DIR/scripts/serve_2x_r9700_mtp.sh"
Plain docker run for Recipe A-MTP
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
  --max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --language-model-only \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • What it is. The model's own one-layer MTP head drafts 3 tokens; the main model checks them in one pass. config.json lists the MTP head's layers as excluded from quantization, so vLLM builds them in BF16 to match the weights on disk. Measured: mean acceptance length 2.76 at T=0.6 (DFlash2: 4.20 on the same 13 prompts).
  • Why keep it. MTP keeps prefix caching on, so a new turn of a long conversation only prefills the new tokens. Under emulation prefill is expensive, so for long multi-turn agent loops A-MTP can finish turns sooner than A even though it decodes slower. The tester matrix measures both.
  • Same as A-plain except speculation: identical environment and flags plus --speculative-config, so the matrix compares speculation only.
  • Pass check. Resolved architecture: Qwen3_5MTP in the log (it follows Resolved architecture: Qwen3_5ForConditionalGeneration; vLLM 0.31 runs this model on its V2 runner, which does not print the V1 runner's Detected MTP model line), and under load SpecDecoding metrics: Mean acceptance length of about 2 or more.
  • Memory: about 0.45 GiB of Gated DeltaNet state per running request per GPU; about 10 concurrent 16K-token requests (estimate).
  • Tuning: AEON_MTP=2 or AEON_MTP=4.

Recipe A-plain: 2x R9700 without speculation (exact baseline)

Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested (expected to work) · Script: scripts/serve_2x_r9700.sh

bash "$AEON_DIR/scripts/serve_2x_r9700.sh"
Plain docker run for Recipe A-plain
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
  --max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

Why these settings (shared by A and A-MTP):

  • --tensor-parallel-size 2 splits every layer across both cards. Each card holds about 10.2 GiB of weights and dequantizes only half of them on every forward pass. All shard boundaries land on whole 32-element MX blocks (checked against the real shapes), and both cards carry 12 attention heads, 2 KV heads, 8 GDN key heads and 24 value heads.
  • NCCL_PROTO=Simple and --attention-backend TRITON_ATTN are what the community R9700 TP=2 / TP=4 setups use (Level1Techs guide); a TP=2 hang on dual R9700 is still reported open (vllm#40980), which is why Recipe Z exists. NCCL_P2P_DISABLE=1 is kept for robustness against hipIpc errors (the guide found it unnecessary). vLLM's custom all-reduce is MI300/MI350-only, so Radeon cards talk through RCCL over PCIe.
  • No speculation, prefix caching on. This recipe avoids every speculative-decoding issue, so it is the most likely TP=2 recipe to work on the first try, and it is the reference for the tester matrix.
  • --max-model-len 131072 --max-num-seqs 16. With the MXFP6 BF16 copy, an estimated 10 GiB per card is left for cache: BF16 KV costs 32 KiB per token per card, so roughly 300K tokens in total, plus about 0.22 GiB of Gated DeltaNet state per running request per card (about 15 concurrent 16K-token requests). For the full 262,144-token context use AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=4. The startup line GPU KV cache size: N tokens is the real number. AEON_MXFP6_ONCE=0 gives back about 5.2 GiB per card, at roughly a fifth of the speed.
  • --language-model-only skips the vision tower (it fails to load on RDNA 4 today).
  • --gpu-memory-utilization 0.90: use 0.85–0.88 if a card drives a display.

Recipe Z: 2x R9700 as two independent servers (never-fails fallback)

Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB (one full copy per card) · Arch: gfx1201 (RDNA 4) · MX path: emulated, stock (no patch) · Status: Untested (expected to work; fewest moving parts) · Script: scripts/serve_2x_r9700_replicas.sh (stop with ... replicas.sh stop)

Use this when every TP=2 recipe fails twice (hipIpcGetMemHandle error, crash at graph capture), or right away after a TP=2 hang (a serve script exits with code 3, "PROBABLY HUNG"; RCCL, vllm#40980, ROCm/ROCm#6148): a hang doesn't depend on speculative decoding, so A-MTP and A-plain would hang the same way. It runs one complete copy of the model per card, with no inter-GPU communication, no speculative decoding, no graph capture (--enforce-eager) and no AITER: aeon-mxfp4-0 on port 8000, aeon-mxfp4-1 on port 8001, and a small least-connections proxy (scripts/aeon_lb.py, or scripts/nginx_aeon_replicas.conf) on port 8080. The MXFP6 dequant-once patch doesn't fit next to a full copy of the weights on one 32 GB card, so Z runs stock emulation: it is the slowest option, a last resort.

bash "$AEON_DIR/scripts/serve_2x_r9700_replicas.sh"
Plain docker run for Recipe Z
for i in 0 1; do
docker run -d --name aeon-mxfp4-$i \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:$((8000+i)):8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e HIP_VISIBLE_DEVICES=$i -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
  --max-num-seqs 4 --gpu-memory-utilization 0.88 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
  --attention-backend TRITON_ATTN --enforce-eager --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
[ "$i" = 0 ] && until docker logs aeon-mxfp4-0 2>&1 | grep -q "Application startup complete"; do
  docker inspect -f '{{.State.Running}}' aeon-mxfp4-0 2>/dev/null | grep -q true || break; sleep 15
done   # GPU 1 reuses the kernel cache GPU 0 built; stops waiting if replica 0 died
done
python3 "$AEON_DIR/scripts/aeon_lb.py" --listen 127.0.0.1:8080 127.0.0.1:8000 127.0.0.1:8001 &
  • Cost (estimate): each card holds about 20.4 GiB of weights and about 6 GiB of cache: 32,768-token context, about 4 sequences per card (about 2 concurrent 16K-token requests). Speed: about 1 tok/s per stream (stock MXFP6 emulation, TP=1, eager) and a few tok/s per card at 4 streams.
  • The script starts the second server only after the first is ready, so the one-time kernel build isn't run twice at the same time. Each container gets only its own GPU's device nodes.
  • First speed step once it works: drop --enforce-eager (AEON_EAGER=0).
  • Don't use --data-parallel-size for this: it still sets up inter-GPU process groups.

Tester matrix: plain vs MTP vs DFlash2

mkdir -p ~/aeon-test && cd ~/aeon-test && bash "$AEON_DIR/scripts/run_matrix.sh"

It starts A-plain, A-MTP and A one after another (all with 16 sequences), runs the same checks on each, then checks TP=2 against a one-GPU reference, and writes results/matrix_compare.md, results/tpcheck.md and the archive to send back:

Check Detail
Speed decode tok/s (aggregate and per stream), time to first token and speculative acceptance at 1, 4, 8 and 16 concurrent requests (T=0.6, 256-token outputs), plus prefill speed on a ~4K-token prompt
Quality GSM8K (first 50, greedy, exact match), IFEval (20 prompts, strict), 10 tool-calling prompts, 7 smoke checks
TP=1 reference all three arms are TP=2 and share the same MX weight-sharding code, so a sharding bug would look the same on each. scripts/tp_check.py records prefill logprobs of 4 fixed prompts on A-plain, then on Recipe Z's replica 0 alone (TP=1, one GPU): expect more than 99% top-1 agreement
Memory peak VRAM per GPU, GPU KV cache size from the startup log

The three should score the same on the quality checks within noise: speculative decoding doesn't change the output distribution. It is not bit-reproducible, though: at T=0 the outputs usually diverge after a few dozen tokens, because verify batches change the floating-point reduction order (measured: 1 of 20 GSM8K outputs identical to plain decoding, mean common prefix 16%), while accuracy stayed statistically the same (plain GSM8K 20/20 and tools 10/10, DFlash2 19/20 and 9/10: two discordant items, McNemar p = 0.5). Plain decoding at different concurrency isn't bit-reproducible either. Expect 1–3 hours for all three arms after the first compile; AEON_MATRIX_ARMS="dflash2 mtp" runs a subset. An arm still compiling after the extra wait stops the matrix instead of being killed mid-build.

Recipe A-Max: 2x R9700, BF16-resident MLP (experimental, superseded)

Target: 2x AMD Radeon AI PRO R9700 · VRAM: 2 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: MXFP4 MLP dequantized once at load; MXFP6 on stock per-pass emulation; MTP in BF16 · Status: Untested (experimental) · Script: scripts/serve_2x_r9700_max.sh

vLLM 0.31's VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 stores the MXFP4 MLP as BF16 (output identical; +11.7 GiB per GPU). It keeps the MXFP6 projections on stock emulation, which costs about 6x more per pass (measured) than the MXFP4 emulation this removes, and both BF16 copies together don't fit on 32 GB cards. Expect it to be slower than A-MTP. It stays in the kit as an A/B data point only (65,536 context, 4 sequences, --gpu-memory-utilization 0.92); if it runs out of memory, step AEON_MAX_MODEL_LEN down to 49152, then 32768.

Speed experiment: radiance (experimental, unlicensed, untested on this checkpoint)

radiance is a community overlay for vLLM with native RDNA 4 kernels: MXFP4 weights with FP8 activations (W4A8) and E6 (MXFP6) kernels. Users report 130–210 tok/s single-stream with DFlash2 on 2x R9700 with other checkpoints. It is not part of this kit, and we have not run it with this checkpoint.

  • It can't serve this checkpoint as shipped. Its MXFP6 kernels are wired only to its own paroquant_mxfp6 format. Serving our Quark MXFP6 layers needs a new kernel class in vLLM's MXFP6 kernel list. Our weights would fold exactly: every MXFP6 row has an E8M0 scale spread of at most 6, which its fold requires (measured on all 1,032,192 rows). Its Quark MXFP4 patch targets vLLM 0.29, which can't load this checkpoint without AEON_STRIP_ALGO_CONFIG=1.
  • It changes the numerics. Activations become per-token FP8 instead of the checkpoint's MXFP4 / MXFP6, so the output is not identical to Recipes A, A-MTP and A-plain (radiance's own MXFP4 perplexity is slightly better with W4A8, but it is a deviation). Its defaults also include lossy options that must be turned off for a fidelity comparison: quantized all-reduce, FP8 KV cache and a quantized output head.
  • License. The original repository has no license file. Private experiments only; don't redistribute it or code copied from it.
  • If you try it, compare it against A-plain with tests/run_radeon_eval.py --compare, and send us both reports.

TensorFold

TensorFold (Apache-2.0) drafts with DFlash2 and is bit-exact against its own serial decoding. Its ROCm support is a community fork (millaguie/TensorFold, branch rocm-r9700).

  • It can't load this mixed checkpoint, so it is not a path for this model today. The ROCm fork serves Qwen3.8-27B MLX 4-bit checkpoints, GGUF, or a Quark export where every projection is MXFP4, one GPU per instance (no TP=2). Our checkpoint is detected as Quark MXFP4 and then fails at the first MXFP6 or BF16 projection; it would need an MXFP6 path, mixed projection groups and BF16 linears on ROCm. Pairing the drafter with a stock Qwen3.8-27B checkpoint would serve a different model. On 2x R9700, use vLLM Recipe A.
  • The drafter is shipped for TensorFold so it can be used as soon as TensorFold supports this mixed Quark export: tensorfold serve <AEON target> --drafter "$AEON_DIR/dflash2". Its config and tensors load in TensorFold's DFlash2 loader (checked against the source and the fork's reader). TensorFold's CUDA/ROCm engine ignores --drafter-bits and re-packs the drafter's linears to affine 4-bit (group 64) at load, as it does for z-lab's drafter, which can lower acceptance but not output fidelity. See the drafter card.

Recipe K: 4x Radeon AI PRO R9700 with the AEON DFlash2 drafter

Target: 4x AMD Radeon AI PRO R9700 / R9700S / R9600D · VRAM: 4 × 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical); drafter in BF16 · Status: Untested · Script: scripts/serve_4x_r9700.sh

bash "$AEON_DIR/scripts/serve_4x_r9700.sh"
Plain docker run for Recipe K
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/dflash2":/draft:ro \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 -e VLLM_USE_V2_MODEL_RUNNER=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 4 --max-model-len 131072 \
  --max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --no-enable-prefix-caching --language-model-only \
  --speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic","rejection_sample_method":"standard"}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate) per GPU: weights 5.1 + MXFP6 BF16 copy 2.6 + BF16 MLP 5.9 + drafter 0.9 = 14.4 GiB, about 3 GiB of activations, about 11 GiB of pool (about 560K tokens at about 20 KiB per token per GPU: 16 KiB, padded to 20 for the drafter's layers; before the GDN state of about 0.37 GiB per running request): about 4 concurrent 128K-token requests.
  • Speed (estimate): about 25–30 tok/s plain and 60–90 with DFlash2 single-stream (±50%).
  • Head counts divide: 1 KV head and 6 attention heads per GPU for the target, 2 KV heads for the drafter. TP=4 with NCCL_PROTO=Simple is reported working on 4x R9700 with other checkpoints.
  • Variants: K-MTP for long multi-turn agent loops (prefix caching on): AEON_DFLASH= AEON_MTP=3 AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=8. K-plain: AEON_DFLASH=.
  • Fallbacks: two Recipe A servers, one per pair of cards (AEON_DEVICES, AEON_NAME, AEON_PORT 8000 / 8001, AEON_REPLACE_OTHERS=0 for the second), then Z with AEON_REPLICAS=4.

Recipe B: 1x Radeon AI PRO R9700

Target: 1x AMD Radeon AI PRO R9700 / R9700S / R9600D · VRAM: 32 GB · Arch: gfx1201 (RDNA 4) · MX path: emulated, stock · Status: Untested (expected to work; slow) · Script: scripts/serve_1x_r9700.sh

bash "$AEON_DIR/scripts/serve_1x_r9700.sh"
Plain docker run for Recipe B
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
  --max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • The MXFP6 dequant-once patch doesn't fit (20.4 + 10.3 GiB > 28.7 GiB usable), so this runs stock emulation: about 1 tok/s single-stream (estimate, ±50%). For real use, two or more cards.
  • Memory (estimate): about 20.4 GiB of weights leaves about 6 GiB for cache: BF16 KV at 64 KiB per token plus about 0.45 GiB of Gated DeltaNet state per running sequence (prefix caching on).
  • Longer context: AEON_MAX_MODEL_LEN=65536 AEON_MAX_NUM_SEQS=2.
  • MTP: AEON_MTP=3 AEON_MAX_NUM_SEQS=2 (about 2.3x, estimate).
  • Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=16384 AEON_MAX_NUM_SEQS=2 AEON_GPU_UTIL=0.88.
  • Dequant-at-load and DFlash2 don't fit on one 32 GB card.

Recipe C: 2x 16 GB RDNA 4 (RX 9070 XT, RX 9070, RX 9060 XT)

Target: 2x AMD Radeon RX 9070 XT or RX 9070 (gfx1201), or 2x RX 9060 XT 16 GB (gfx1200, about half the bandwidth) · VRAM: 2 × 16 GB · Arch: gfx1201 / gfx1200 (RDNA 4) · MX path: emulated, stock · Status: Untested (expected to work; tight memory) · Script: scripts/serve_2x_rx9070xt.sh

bash "$AEON_DIR/scripts/serve_2x_rx9070xt.sh"
Plain docker run for Recipe C
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 16384 \
  --max-num-seqs 2 --gpu-memory-utilization 0.92 --kv-cache-dtype auto --max-num-batched-tokens 2048 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

Same code path as Recipe A-plain with half the memory: about 10.2 GiB of weights per card leaves about 2.6 GiB per card for cache, and the MXFP6 dequant-once patch doesn't fit. About 2 tok/s single-stream on RX 9070 XT, about 1 on RX 9060 XT (estimates). On RX 9060 XT cards use PYTORCH_ROCM_ARCH=gfx1200 in a plain docker run (the script detects it). If a card drives a display, use --gpu-memory-utilization 0.88. MTP (AEON_MTP=2) only if the GPU KV cache size line shows room. FP8 KV (AEON_KV_DTYPE=fp8) doubles the context but changes outputs (uncalibrated scales); it is opt-in. There is no single-card fallback: the model doesn't fit on 16 GB.

Recipe D: 4x 16 GB cards (RDNA 4 or RDNA 3)

Target: 4x AMD Radeon RX 9070 XT / RX 9070 (gfx1201) or RX 9060 XT 16 GB (gfx1200); or 4x RX 7900 GRE (gfx1100), RX 7800 XT / RX 7700 / Radeon PRO W7700 (gfx1101) · VRAM: 4 × 16 GB · Arch: gfx1201 / gfx1200 / gfx1100 / gfx1101 · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script: scripts/serve_4x_rx9070xt.sh

bash "$AEON_DIR/scripts/serve_4x_rx9070xt.sh"
Plain docker run for Recipe D
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1201 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 4 --max-model-len 65536 \
  --max-num-seqs 8 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate) per GPU: weights 5.1 + MXFP6 BF16 copy 2.6 GiB, about 2 GiB of activations, about 4.6 GiB of pool (about 300K tokens at 16 KiB per token per GPU): about 4 concurrent 64K-token requests. Longer context: AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=4.
  • Speed (estimate): about 18 tok/s single-stream on RX 9070 XT, about 40 with MTP (AEON_MTP=3), about 50 with DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=4, prefix caching off).
  • TP=4 shards check out (6 attention heads and 1 KV head per card; every MX shard on a whole block), and TP=4 with NCCL_PROTO=Simple is reported working on 4x R9700. All-reduce over consumer PCIe likely limits speed. If startup hangs at graph capture, AEON_EAGER=1.
  • 4x 12 GB (RX 9070 GRE, RX 7700 XT): AEON_MXFP6_ONCE=0 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=4 AEON_GPU_UTIL=0.92 (untested).
  • Don't use 8 cards with TP=8: 4 KV heads can't be split 8 ways.

Recipe E: 2x Radeon RX 7900 XTX

Target: 2x AMD Radeon RX 7900 XTX · VRAM: 2 × 24 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script: scripts/serve_2x_rx7900xtx.sh

bash "$AEON_DIR/scripts/serve_2x_rx7900xtx.sh"
Plain docker run for Recipe E
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 32768 \
  --max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate) per GPU: weights 10.2 + MXFP6 BF16 copy 5.2 GiB, about 2 GiB of activations, about 4.2 GiB of pool (about 138K tokens; RDNA 3 has no FP8, so KV stays BF16 at 32 KiB per token per card): about 3 concurrent 32K-token requests.
  • Speed (estimate): about 15 tok/s single-stream, about 35 with MTP (AEON_MTP=3).
  • Variants: AEON_MAX_MODEL_LEN=65536 AEON_MAX_NUM_SEQS=2. Long context on stock emulation (about 5x slower): AEON_MXFP6_ONCE=0 AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=2. DFlash2 and dequant-at-load don't fit next to the patch.
  • 2x RX 7900 XT (20 GB): use Recipe M.

Recipe M: 2x Radeon RX 7900 XT

Target: 2x AMD Radeon RX 7900 XT · VRAM: 2 × 20 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, stock · Status: Untested · Script: scripts/serve_2x_rx7900xt.sh

bash "$AEON_DIR/scripts/serve_2x_rx7900xt.sh"
Plain docker run for Recipe M
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 32768 \
  --max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate) per GPU: about 18 GiB usable, 10.2 GiB of weights, about 2 GiB of activations, about 5.8 GiB of pool: about 4–5 concurrent 32K-token requests.
  • Speed (estimate): about 2.3 tok/s single-stream (stock emulation); MTP AEON_MTP=3.
  • Opt-in MXFP6 dequant-once (about 12 tok/s, estimate; tight, untested, no display on the cards): AEON_MXFP6_ONCE=1 AEON_GPU_UTIL=0.93 AEON_MAX_MODEL_LEN=16384 AEON_MAX_NUM_SEQS=2 (about 1.5 GiB of pool per GPU).

Recipe L: 2x 16 GB RDNA 3 (RX 7900 GRE, RX 7800 XT, RX 7700, PRO W7700)

Target: 2x AMD Radeon RX 7900 GRE (gfx1100), or 2x RX 7800 XT / RX 7700 / Radeon PRO W7700 (gfx1101) · VRAM: 2 × 16 GB · Arch: gfx1100 / gfx1101 (RDNA 3) · MX path: emulated, stock · Status: Untested (tight memory) · Script: scripts/serve_2x_16gb_rdna3.sh

bash "$AEON_DIR/scripts/serve_2x_16gb_rdna3.sh"
Plain docker run for Recipe L
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 16384 \
  --max-num-seqs 2 --gpu-memory-utilization 0.92 --kv-cache-dtype auto --max-num-batched-tokens 2048 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

Like Recipe C on RDNA 3: about 2.6 GiB of pool per card (a 16K-token request takes about 0.5 GiB of KV plus 0.22 GiB of state per card), stock emulation, about 1.7 tok/s single-stream (estimate). On gfx1101 cards use PYTORCH_ROCM_ARCH=gfx1101 in a plain docker run (the script detects it). Use 0.88 for --gpu-memory-utilization if a card drives a display. MTP (AEON_MTP=2) only if the GPU KV cache size line shows at least 1.5 GiB. Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=8192 AEON_MAX_NUM_SEQS=1. No single-card fallback.

Recipe F: Radeon PRO W7900 (48 GB)

Target: 1x AMD Radeon PRO W7900, W7900 Dual Slot or W7800 48 GB · VRAM: 48 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script: scripts/serve_w7900.sh

bash "$AEON_DIR/scripts/serve_w7900.sh"
Plain docker run for Recipe F
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 65536 \
  --max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate): weights 20.4 + MXFP6 BF16 copy 10.3 GiB, about 3 GiB of activations, about 9.5 GiB of pool (about 155K tokens at 64 KiB per token): about 2 concurrent 64K-token requests. Dequant-at-load doesn't fit (20.4 + 23.4 GiB > 43.2 usable).
  • Speed (estimate): about 8 tok/s single-stream, about 18 with MTP (AEON_MTP=3), about 25 with DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=2, prefix caching off).
  • Variants: AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=2. The full 262,144-token context on stock emulation (about 5x slower): AEON_MXFP6_ONCE=0 AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=1. Vision: AEON_VISION=1.
  • Two or four of these cards: Recipe Q.

Recipe Q: 2x Radeon PRO W7900 with the AEON DFlash2 drafter

Target: 2x AMD Radeon PRO W7900, W7900 Dual Slot or W7800 48 GB (variants for 4x W7900 and 2x W7800 32 GB) · VRAM: 2 × 48 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical); drafter in BF16 · Status: Untested · Script: scripts/serve_2x_w7900.sh

bash "$AEON_DIR/scripts/serve_2x_w7900.sh"
Plain docker run for Recipe Q
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/dflash2":/draft:ro \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1100 -e NCCL_PROTO=Simple -e NCCL_P2P_DISABLE=1 -e VLLM_EXECUTE_MODEL_TIMEOUT_SECONDS=1800 \
  -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 -e VLLM_USE_V2_MODEL_RUNNER=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 2 --max-model-len 131072 \
  --max-num-seqs 16 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --no-enable-prefix-caching --language-model-only \
  --speculative-config '{"method":"dflash","model":"/draft","num_speculative_tokens":9,"num_speculative_tokens_per_batch_size":[[1,2,9],[3,8,7],[9,12,6],[13,16,4]],"attention_backend":"TRITON_ATTN","draft_sample_method":"probabilistic","rejection_sample_method":"standard"}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate) per GPU: 10.2 + 5.2 (MXFP6 BF16) + 11.7 (BF16 MLP) + 1.8 (drafter) = 28.9 GiB, about 3.5 GiB of activations, about 10.8 GiB of pool.
  • Speed (estimate): about 24 tok/s plain and about 70 with DFlash2 single-stream.
  • Variants: Q-MTP for long multi-turn agent loops (prefix caching on): AEON_DFLASH= AEON_MTP=3 AEON_MAX_MODEL_LEN=262144 AEON_MAX_NUM_SEQS=8. 2x W7800 32 GB: AEON_DEQUANT_AT_LOAD=0 AEON_DFLASH= (like A-plain; about 10 GiB of pool per card). 4x W7900: AEON_TP=4.
  • Fallback: one Recipe F server per card (AEON_GPU_IDS=0 / 1, AEON_PORT 8000 / 8001, AEON_NAME, AEON_REPLACE_OTHERS=0 for the second).

Recipe G: Radeon PRO W7800 (32 GB)

Target: 1x AMD Radeon PRO W7800 32 GB · VRAM: 32 GB · Arch: gfx1100 (RDNA 3) · MX path: emulated, stock · Status: Untested (slow) · Script: scripts/serve_w7800_32gb.sh

bash "$AEON_DIR/scripts/serve_w7800_32gb.sh"
Plain docker run for Recipe G
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1100 -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
  --max-num-seqs 4 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

Like Recipe B on RDNA 3: about 6 GiB left for BF16 KV and Gated DeltaNet state; the MXFP6 dequant-once patch doesn't fit, so about 0.9 tok/s single-stream (estimate). MTP: AEON_MTP=3 AEON_MAX_NUM_SEQS=2. Text-only. For real use, two cards (Recipe Q's 2x W7800 32 GB variant).

Recipe P: Radeon PRO V710

Target: 1x AMD Radeon PRO V710 (the full GPU, not a fractional partition; usually a cloud VM) · VRAM: 28 GB · Arch: gfx1101 (RDNA 3) · MX path: emulated, stock · Status: Untested (slow) · Script: scripts/serve_v710.sh

bash "$AEON_DIR/scripts/serve_v710.sh"
Plain docker run for Recipe P
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1101 -e VLLM_ROCM_USE_AITER=0 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 16384 \
  --max-num-seqs 2 --gpu-memory-utilization 0.92 --kv-cache-dtype auto --max-num-batched-tokens 2048 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'

About 25.7 GiB usable, 20.4 GiB of weights, about 3.1 GiB of pool (a 16K-token request takes about 1.0 GiB of KV plus 0.43 GiB of state). About 0.7 tok/s single-stream (estimate). No room for speculation. TP=1 needs no RCCL, so the preflight's virtual-machine warning doesn't apply here. Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=8192 AEON_MAX_NUM_SEQS=1.

Recipe H: Strix Halo (Ryzen AI Max+ 395), 64 GB

Target: AMD Ryzen AI Max+ 395 / Max 390 / Max 385 APU (Radeon 8060S / 8050S) with 64 GB · Memory: 64 GB unified (at least 45 GiB GPU-addressable) · Arch: gfx1151 (RDNA 3.5) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script: scripts/serve_strix_halo.sh

Before you start, the GPU must be able to allocate at least about 45 GiB. vLLM sizes its memory from what HIP reports, so raise the TTM/GTT limit the way AMD documents it for Strix Halo (amd-ttm, or the ttm pages_limit module parameter; amdgpu.gttsize is deprecated), or raise the UMA frame buffer in the BIOS. scripts/detect_setup.py prints both values. Don't set HSA_OVERRIDE_GFX_VERSION (older Strix Halo guides do): gfx1151 is a native target of the image, and the launcher refuses the override.

bash "$AEON_DIR/scripts/serve_strix_halo.sh"
Plain docker run for Recipe H
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1151 -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 32768 \
  --max-num-seqs 2 --gpu-memory-utilization 0.85 --kv-cache-dtype auto --max-num-batched-tokens 4096 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate): of about 48 GiB addressable, about 40.8 usable: weights 20.4 + MXFP6 BF16 copy 10.3 GiB, about 7.9 GiB of pool. BF16 KV (no FP8 on gfx1151).
  • Speed (estimate): about 2.3 tok/s single-stream (about 0.4 without the patch), about 5 with MTP (AEON_MTP=3).
  • Precision caveat: AMD lists only FP16 as officially validated on Ryzen APUs. This recipe runs BF16 (the reference numerics; fp16 also breaks the DFlash2 drafter): an untested risk.
  • Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=1. If vLLM's free-memory check fails on unified memory, AEON_EXTRA_ARGS="--kv-cache-memory-bytes <bytes>" is an untested alternative.

Recipe H-128: Strix Halo (Ryzen AI Max+ 395), 128 GB

Target: AMD Ryzen AI Max+ 395 / Max 390 / Max 385 APU with 128 GB · Memory: 128 GB unified (at least 90 GiB GPU-addressable) · Arch: gfx1151 (RDNA 3.5) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical); MTP head in BF16 · Status: Untested · Script: scripts/serve_strix_halo_128gb.sh

Same preparation as Recipe H, with at least 90 GiB GPU-addressable.

bash "$AEON_DIR/scripts/serve_strix_halo_128gb.sh"
Plain docker run for Recipe H-128
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx1151 -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 131072 \
  --max-num-seqs 4 --gpu-memory-utilization 0.80 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --language-model-only \
  --speculative-config '{"method":"mtp","num_speculative_tokens":3,"attention_backend":"TRITON_ATTN"}' \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate): of about 96 GiB addressable, about 76.8 usable: weights 20.4 + 10.3 + 23.4 + MTP 0.8 GiB, about 3.5 GiB of activations, about 18 GiB of pool: about 2 concurrent 128K-token requests.
  • Speed (estimate): about 4.5 tok/s plain and about 10 with MTP (the default) single-stream.
  • DFlash2 instead of MTP: AEON_MTP=0 AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=65536 (prefix caching off). Plain: AEON_MTP=0. Most robust: AEON_EAGER=1 AEON_MAX_MODEL_LEN=32768 AEON_MAX_NUM_SEQS=1.

Recipe I: Instinct MI300X / MI325X

Target: AMD Instinct MI300X, MI308X or MI325X · VRAM: 192 / 256 GB · Arch: gfx942 (CDNA 3) · MX path: emulated (no MX hardware), MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical) · Status: Untested · Script: scripts/serve_mi300x.sh

bash "$AEON_DIR/scripts/serve_mi300x.sh"
Plain docker run for Recipe I
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx942 -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 262144 \
  --max-num-seqs 32 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 16384 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate), MI300X: about 172 GiB usable, weights 54.1 GiB (20.4 + 10.3 + 23.4), about 114 GiB of pool (about 1.8M BF16 KV tokens): about 7 concurrent 262K-token requests. Memory is not a constraint, so use one GPU per copy and scale out with data parallelism (AEON_DP=8) instead of tensor parallelism.
  • Faster: AEON_MTP=3 (best for agentic multi-turn: prefix caching stays on). For short-context throughput, AEON_DFLASH="$AEON_DIR/dflash2" (prefix caching off; with AEON_DP>1 vLLM drops the batch-size lattice, so add AEON_DFLASH_K=6).
  • Images: AEON_VISION=1 (4 images per prompt).
  • An FP8 or BF16 build is the better everyday choice on MI300, since memory isn't the limit there.

Recipe O: Instinct MI300A

Target: AMD Instinct MI300A (APU) · Memory: 128 GB unified HBM3, shared with the host OS · Arch: gfx942 (CDNA 3) · MX path: emulated, MXFP6 dequantized once and MXFP4 MLP kept in BF16 after load (bit-identical) · Status: Untested · Script: scripts/serve_mi300a.sh

bash "$AEON_DIR/scripts/serve_mi300a.sh"
Plain docker run for Recipe O
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx942 -e VLLM_ROCM_USE_AITER=0 -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 131072 \
  --max-num-seqs 16 --gpu-memory-utilization 0.70 --kv-cache-dtype auto --max-num-batched-tokens 16384 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate): --gpu-memory-utilization 0.70 leaves room for the OS: about 89 GiB usable, weights 54.1 GiB, about 32 GiB of pool: about 3–4 concurrent 128K-token requests. Size the utilization from the GiB the launcher prints for the GPU; if it reports much less than 128 GiB, the GPU-allocatable limit has to be raised per AMD's MI300A guidance (ask your administrator).
  • Faster: AEON_MTP=3, or DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2", prefix caching off). Several APUs: AEON_DP=<number of APUs>.

Recipe J: Instinct MI350X / MI355X

Target: AMD Instinct MI350X or MI355X · VRAM: 288 GB · Arch: gfx950 (CDNA 4) · MX path: default exact: MXFP4 emulated with the MLP kept in BF16 after load, MXFP6 dequantized once (bit-identical to every other recipe); native MXFP4 (AITER W4A4) opt-in · Status: Untested · Script: scripts/serve_mi355x.sh

bash "$AEON_DIR/scripts/serve_mi355x.sh"
Plain docker run for Recipe J
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx950 -e VLLM_ROCM_USE_AITER=0 -e VLLM_DISABLED_KERNELS=AiterMxfp4LinearKernel -e VLLM_MXFP4_EMULATION_DEQUANT_AT_LOAD=1 \
  -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 262144 \
  --max-num-seqs 32 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 16384 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Why the exact path is the default. Native MXFP4 runs the MLP on AITER's W4A4 GEMM: the same MX grid and E8M0 scale rule, but a different kernel and summation order, so it is not bit-identical to the other recipes. This kit puts fidelity first, so native is a labelled opt-in: AEON_MXFP4_EMULATE=0 AEON_DEQUANT_AT_LOAD=0 (faster at high concurrency; log line Using AiterMxfp4LinearKernel for MXFP4 GEMM). If the native path fails with a ModuleNotFoundError for aiter.ops.triton.gemm_afp4wfp4 (newer AITER builds moved that module), stay on the default.
  • Check the log for Using EmulationMxfp4LinearKernel for MXFP4 GEMM (default) and Using EmulationMxfp6LinearKernel for MXFP6 GEMM.
  • Memory (estimate): about 259 GiB usable, weights 54.1 GiB (30.7 with native MXFP4), about 200 GiB of pool.
  • Faster: DFlash2 (AEON_DFLASH="$AEON_DIR/dflash2") or MTP (AEON_MTP=3). Scaling out: AEON_DP=N (with DFlash2, add AEON_DFLASH_K=6).

Recipe N: Instinct MI210 / MI250 / MI250X

Target: AMD Instinct MI210 (64 GB), MI250 / MI250X (two 64 GB GCDs, each shown as its own GPU) · VRAM: 64 GB per GCD · Arch: gfx90a (CDNA 2) · MX path: emulated, MXFP6 dequantized once at load · Status: Untested · Script: scripts/serve_mi250.sh

bash "$AEON_DIR/scripts/serve_mi250.sh"
Plain docker run for Recipe N
docker run -d --name aeon-mxfp4 \
  --device /dev/kfd --device /dev/dri \
  --group-add video --group-add "$RENDER_GID" \
  --cap-add SYS_PTRACE --security-opt seccomp=unconfined --ipc=host \
  -p 127.0.0.1:8000:8000 \
  -v "$AEON_DIR":/model:ro -v "$AEON_CACHE":/root/.cache \
  -v "$AEON_DIR/scripts/patches/vllm-0.31.0/mxfp6_emulation.py":/usr/local/lib/python3.12/dist-packages/vllm/model_executor/kernels/linear/mxfp6/emulation.py:ro \
  -e TRITON_CACHE_DIR=/root/.cache/triton -e TORCH_EXTENSIONS_DIR=/root/.cache/torch_extensions \
  -e PYTORCH_ROCM_ARCH=gfx90a -e VLLM_ROCM_USE_AITER=0 -e AEON_MXFP6_DEQUANT_AT_LOAD=1 \
  vllm/vllm-openai-rocm:v0.31.0 \
  /model --served-model-name aeon --dtype bfloat16 --tensor-parallel-size 1 --max-model-len 131072 \
  --max-num-seqs 8 --gpu-memory-utilization 0.90 --kv-cache-dtype auto --max-num-batched-tokens 8192 \
  --attention-backend TRITON_ATTN --language-model-only \
  --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_coder \
  --override-generation-config '{"temperature":0.6,"top_p":0.95,"top_k":20,"min_p":0.0,"presence_penalty":0.0,"repetition_penalty":1.0}'
  • Memory (estimate) per GCD: about 57.6 GiB usable, weights 20.4 + MXFP6 BF16 copy 10.3 GiB, about 3 GiB of activations, about 24 GiB of pool: about 2–3 concurrent 128K-token requests. MXFP4 dequant-at-load doesn't fit next to it on one GCD.
  • Speed (estimate): about 14 tok/s per GCD single-stream (likely launch-bound), about 33 with MTP (AEON_MTP=3).
  • Scale out: one replica per GCD with AEON_DP=<number of GCDs>. TP=2 across the two GCDs of one MI250 / MI250X: AEON_TP=2 AEON_DEQUANT_AT_LOAD=1 AEON_DFLASH="$AEON_DIR/dflash2" AEON_MAX_MODEL_LEN=131072 AEON_MAX_NUM_SEQS=16 (xGMI; no NCCL overrides).
  • AITER stays off (vLLM's AITER needs CDNA 3 or newer).

Not supported

  • Single GPUs under 28 GB: RX 7900 XTX (24 GB), RX 7900 XT (20 GB), every 16 GB card (RX 9070 XT / 9070 / 9060 XT 16 GB, RX 7900 GRE, RX 7800 XT, RX 7700, Radeon PRO W7700) and 12 GB cards (RX 9070 GRE, RX 7700 XT). About 20.4 GiB of text-only weights leaves no usable KV cache at 90% utilization. Smallest working combinations: 2x 24 GB (E), 2x 20 GB (M), 2x 16 GB (C for RDNA 4, L for RDNA 3), 4x 12 GB (D variant).
  • 8 GB cards (RX 9060 XT 8 GB, RX 7600 and similar): even four of them leave no room for activations and cache.
  • Architectures ROCm supports but the vLLM ROCm image isn't built for: Instinct MI100 (gfx908) and RDNA 2 (gfx1030: Radeon PRO W6800 / V620, RX 6800–6950 XT). vLLM's ROCm build targets gfx90a, gfx942, gfx950, gfx1100, gfx1101, gfx1150, gfx1151, gfx1200 and gfx1201. Also gfx1102 (RX 7600 / 7600 XT, Radeon PRO W7600 / W7500), which is in neither AMD's ROCm Linux list nor the image, and the MI50 / MI60 (gfx906), which current ROCm no longer supports.
  • gfx1250: vLLM ships a separate image for it; these recipes don't cover it yet.
  • Integrated GPUs other than Strix Halo (Ryzen AI 300 / gfx1150, Radeon 780M / gfx1103 and older): too little memory and bandwidth.
  • TP=3 and TP=8: the model has 4 KV heads, which must split evenly. Use TP=1, 2 or 4, or replicas with data parallelism.
  • Virtual machines with GPU passthrough for multi-GPU Radeon recipes: RCCL initialization has been reported to hang. Single-GPU recipes (B, F, G, P) don't use RCCL; for two cards, use bare metal or Recipe Z.
  • Windows and WSL: not covered; vLLM's ROCm images are Linux-only.
  • NVIDIA GPUs: use NVFP4-MIXED.
  • llama.cpp, Ollama, LM Studio: this is not a GGUF.

Vision (Instinct and RDNA 3 only)

Every recipe starts text-only, which leaves the most memory for context. On RDNA 4 the vision tower currently fails to load in vLLM (vllm#49851), so keep --language-model-only there. Elsewhere, to serve images:

  1. Set AEON_VISION=1 with the script, or edit the docker run:
    • remove --language-model-only;
    • add --limit-mm-per-prompt '{"image":1,"video":0}' --mm-processor-kwargs '{"max_pixels":1048576}';
    • on Radeon / Strix Halo, add -e FLASH_ATTENTION_TRITON_AMD_ENABLE=TRUE before the image name.
  2. The vision tower adds about 0.9 GB, split across the cards.

The image-size cap matters. Without it, vLLM sizes the encoder's startup memory test for this processor's maximum of about 16.7 megapixels (about 16K tokens per image), which can run out of memory. With the cap it plans for about 1,024 tokens per image. Vision is untested with this checkpoint.

Troubleshooting

The full symptom → cause → fix table is in AGENTS.md. The most common ones:

Symptom Likely cause Fix
AttributeError: 'dict' object has no attribute 'endswith' in quark.py vLLM image v0.28–v0.30 Use vllm/vllm-openai-rocm:v0.31.0
The launch script stops with vLLM detects 'gfx1036' (or another iGPU target) An integrated GPU is listed first, so vLLM takes its architecture AEON_DEVICES=auto, or pass only the discrete /dev/dri/renderD* and card* nodes, or disable the iGPU in the BIOS. HIP_VISIBLE_DEVICES alone doesn't help
invalid device function / hipErrorNoBinaryForGpu Kernels built for a sibling target (for example gfx1201 on a gfx1200 card) or HSA_OVERRIDE_GFX_VERSION set Use the scripts (they set PYTORCH_ROCM_ARCH to the detected GPU), or set it to your target in a plain docker run; never set HSA_OVERRIDE_GFX_VERSION
Startup or the first request hangs with both GPUs at 100%, or responses are empty (TP ≥ 2); a serve script exits with code 3 ("PROBABLY HUNG": no log line, compiler or kernel-cache write for 20 minutes) RCCL over PCIe (vllm#40980, ROCm/ROCm#6148) During an event: save the log, docker rm -f aeon-mxfp4, start Recipe Z directly (skip A-MTP / A-plain; the hang doesn't depend on speculation). Later, with the host owner: check NCCL_PROTO=Simple and NCCL_P2P_DISABLE=1 are set; both cards in equal-width CPU slots; IOMMU on with iommu=pt (the community guide also uses pcie_aspm=off); bare metal. Exit code 2 means still compiling: wait (AGENTS.md section 0, hang procedure)
hipIpcGetMemHandle ... invalid argument P2P IPC between the cards NCCL_P2P_DISABLE=1, then Recipe Z
Stuck in compilation for over 30 minutes, one CPU core busy Triton autotuning the Gated DeltaNet kernels on first start Wait, then retry with AEON_EXTRA_ENV="VLLM_TRITON_FORCE_FIRST_CONFIG=1"
out of resource: shared memory, Required: N, Hardware limit: 65536 A Triton kernel config exceeds RDNA's 64 KB LDS (AITER, if it was turned on; otherwise note the kernel name in the traceback) Confirm VLLM_ROCM_USE_AITER=0 (the scripts set it). Then AEON_ATTN=default on a recipe without speculation (untested), else Recipe Z; send the log
... larger than the available KV cache memory or No available memory for the cache blocks Not enough VRAM left for KV Lower --max-model-len, then --max-num-seqs; keep --language-model-only; close other GPU apps
Speculative acceptance near 1 with several requests in flight Draft model on the ROCM_ATTN backend "attention_backend":"TRITON_ATTN" inside --speculative-config
DFlash2 crashes or acceptance drops to 0 after a repeated prompt Prefix caching with this hybrid model --no-enable-prefix-caching (Recipe A sets it; don't combine DFlash2 with prefix caching unless you are testing AEON_DFLASH_PREFIX_FIX=1)
Launch script: ... emulation.py is not the v0.31.0 file the MXFP6 dequant-once patch was made from An image whose file differs from v0.31.0's Use vllm/vllm-openai-rocm:v0.31.0, or AEON_MXFP6_ONCE=0 (same output, much slower)
DFlash2DraftModel missing from the log, or acceptance like DFlash 1 Drafter folder missing, or VLLM_USE_V2_MODEL_RUNNER=0 ls $AEON_DIR/dflash2 (config.json + model.safetensors); keep VLLM_USE_V2_MODEL_RUNNER=1
Endless thinking or repetition Client sends temperature: 0 Send 0.6 or omit it; cap max_tokens; enable_thinking: false for tool loops
Decode speed about 20% lower than a previous run R9700 decode speed can land lower for the life of a process (legacy-rocm-build#6347) Restart the server and measure again
Tool calls come back as plain text Tool parser not active --enable-auto-tool-choice --tool-call-parser qwen3_coder; the client must send tools

Expected performance (estimates, not measurements)

Nobody has measured this checkpoint on AMD hardware yet. The figures below combine measurements of this exact checkpoint in vLLM (MXFP6 emulated; an equivalent W4A4 MLP path with the same MXFP4 weight grid and the same activation quantize-dequantize, which differs from vLLM's emulation only in GEMM accumulation order, so it is not bit-identical; one request, 384-token outputs; see Quality and validation status) with a memory-bandwidth scaling to each AMD system. Treat every figure as roughly ±50%.

2x R9700 recipe Measured (reference run) Estimate, 2x R9700, MXFP6 patch on (default) Estimate, stock emulation (AEON_MXFP6_ONCE=0)
A-plain 8.1 tok/s (4 prompts) ~10 tok/s ~2 tok/s
A-MTP (K=3) mean acceptance length 2.76 at T=0.6 (13 prompts); 13.5 tok/s on the same 4 prompts as plain ~24 tok/s ~5 tok/s
A (DFlash2, K=9) 3.73 at T=0.6 (26 prompts), 23.2 tok/s (22.0 on the same 4 prompts as plain); 4.40 at T=0, 27.3 tok/s ~35 tok/s (T=0.6), ~41 (T=0) ~7 tok/s
A, 8 concurrent requests 113 tok/s aggregate (4.9x its single-stream rate) several times the single-stream rate; unmeasured n/a
Z (per stream) n/a n/a ~1 tok/s

How the estimates are worked out:

  • Per forward pass. In the reference run, stock MXFP6 emulation took about 2.1 s per pass (12–20 ms per Gated DeltaNet projection call, 144 calls), MXFP4 emulation about 0.32 s, and the BF16 rest about 0.04 s, for about 2.47 s in total. With the MXFP6 weights dequantized once, the MXFP6 part drops to the 0.044 s the same layers cost in BF16 (about 0.40 s per pass); with the MXFP4 MLP also in BF16 after load, about 0.21 s. Each AMD estimate scales that cost by memory bandwidth, divides by the TP size, and adds about 8 ms of PCIe all-reduce at TP=2 (about 12 ms at TP=4). On 2x R9700 that gives about 94 ms per pass (about 10.6 tok/s) with the patch and about 0.53 s (about 1.9 tok/s) without it.
  • Speculation. Speedup = mean acceptance length ÷ (cost of one draft-plus-verify step ÷ cost of one plain step). Under emulation the verify step costs little more than a plain one, because the weight dequantization dominates and doesn't depend on how many tokens are checked; the step ratio is about 1.14 for DFlash2 and about 1.2 for MTP. So DFlash2 ≈ 3.73 ÷ 1.14 ≈ 3.3x plain, MTP ≈ 2.3x. When a pass gets cheap (both BF16 copies, many GPUs) the drafter's own cost weighs more and the gain shrinks.
  • Acceptance on this target. The drafter accepts about 5–7% less on this MXFP4/MXFP6 checkpoint than on the NVFP4 build it was fine-tuned against (3.73 vs 4.01 at T=0.6, same prompts); at 4 and 8 concurrent requests the two were on par.
  • Concurrency. Weight traffic is per forward pass, not per request, so aggregate throughput rises with concurrency until compute or memory runs out. 16 concurrent requests are unmeasured; the tester matrix measures them.
  • Not modelled: RDNA 4 kernel efficiency, HIP-graph capture of the Quark ops, and RCCL behaviour with NCCL_P2P_DISABLE=1. Read the A-plain time per token first; it calibrates everything else.

Other systems (single stream, default recipe path; estimates ±50%):

Recipe GPUs Default MX path Plain (est.) With MTP (est.) With DFlash2 (est.)
K 4x R9700 once + at load ~25–30 tok/s ~60 ~60–90 (default)
B 1x R9700 stock ~1 ~2 doesn't fit
C 2x RX 9070 XT (2x RX 9060 XT) stock 1.9 (0.9) not advised doesn't fit
D 4x RX 9070 XT once ~18 (stock ~3.6) ~40 ~50
E 2x RX 7900 XTX once ~15 (stock ~2.8) ~35 doesn't fit
M 2x RX 7900 XT stock ~2.3 (opt-in once ~12) ~5 doesn't fit
L 2x 16 GB RDNA 3 stock ~1.7 not advised doesn't fit
F W7900 once ~8 (stock ~1.3) ~18 ~25
Q 2x W7900 once + at load ~24 ~55 ~70 (default)
G W7800 32 GB stock ~0.9 ~2 doesn't fit
P V710 stock ~0.7 doesn't fit doesn't fit
H Strix Halo 64 GB once ~2.3 (stock ~0.4) ~5 tight
H-128 Strix Halo 128 GB once + at load ~4.5 ~10 (default) ~14
N MI210 / MI250 (per GCD) once ~14 ~33 ~45
I, O, J MI300X / MI325X, MI300A, MI350X / MI355X once + at load launch-bound: no estimate

At Instinct bandwidths (about 5.3 TB/s for MI300X / MI300A, 8 TB/s for MI355X) kernel-launch overhead, not memory bandwidth, likely sets the single-stream limit, so no figure is given; MI210 / MI250 figures are optimistic for the same reason.

For scale only: Puget Systems reports 15.9 tok/s stock and 62.8 tok/s tuned with MTP at concurrency 1 on 2x R9700 with Qwen3.6-27B in FP8 (article). That is a different model and a different format with no emulation.

Testers: the step-by-step tester kit ships with the files, in TESTER_GUIDE.md, tests/run_radeon_eval.py and scripts/ (start with scripts/preflight_smoke.sh). Agents: AGENTS.md.

Expected behavior and limits

  • Speed. Under emulation, decode speed is limited by how many bytes each forward pass moves, not by the 23.7 GB of packed weights. See Expected performance. Tensor parallelism, speculative decoding (DFlash2, MTP) and concurrent requests all spread that cost, and the MXFP6 dequant-once patch (plus MXFP4 dequant-at-load on the largest GPUs) removes most of it wherever it fits, without changing the output. Prefill is affected much less. On MI350/MI355 the MLP can also run natively (opt-in, not bit-identical).
  • Memory.
    • About 22.8 GB of weights load with the vision tower, or about 21.9 GB (20.4 GiB) with --language-model-only. The MTP head (0.85 GB) loads only when you turn on MTP.
    • The KV cache costs 64 KiB per token at BF16 (16 full-attention layers × 4 KV heads × 256 × 2 × 2 bytes), divided by the TP size. Every recipe uses BF16 KV, the reference numerics; FP8 KV (32 KiB) is an opt-in that changes outputs.
    • Each running sequence also holds Gated DeltaNet state: about 147 MiB per state slot, FP32 as the model config sets it, divided by the TP size. A running sequence holds about 1 + K slots without prefix caching (K = draft tokens; 1 slot without speculation), and about 2 + K + 1 with prefix caching on (vLLM's default for this model, "align" mode). At TP=2 that is about 0.22 GiB per GPU for A-plain and about 0.45 GiB with MTP K=3 (see AGENTS.md section 12.2).
    • The GPU KV cache size line in the startup log is the real capacity.
  • Speculative decoding is untested on AMD (measured Ï„ for this checkpoint: see Quality and validation status). DFlash2: see Recipe A. MTP: see Recipe A-MTP; use {"method":"mtp",...} (the older name qwen3_5_mtp still works but logs a deprecation warning). Always add "attention_backend":"TRITON_ATTN" to the speculative config.
  • Thinking.
    • Thinking is on by default, and the chat template's default reasoning effort is xhigh.
    • For quick answers, send "chat_template_kwargs": {"enable_thinking": false}.
    • For short thinking, send {"enable_thinking": true, "reasoning_effort": "low"}.
    • Set both through chat_template_kwargs, not through a top-level request field.
  • Inherited behavior. The BF16 master is an Early Access Draft, and very long answers can fall into loops. 4-bit MLP weights, including the residual-writing down_proj, may make that more likely. The tester kit checks for loops explicitly.
  • Uncensored. This model writes what the base model refuses. Read User responsibility.
  • Runtimes.
    • vLLM is the only supported path.
    • This is not a GGUF, so llama.cpp, Ollama and LM Studio can't load it.
    • Transformers with amd-quark may load it in emulation, but that is untested.
    • Apple needs an MLX conversion.
  • Known ROCm issues.
    • Decode speed on the R9700 can land about 20% lower for the life of a process (legacy-rocm-build#6347); restart and re-measure.
    • A TP=2 hang on dual R9700 is still reported open (vllm#40980); Recipes A, A-MTP and A-plain use the settings of the working community setups (NCCL_PROTO=Simple, plus NCCL_P2P_DISABLE=1), and Recipe Z avoids inter-GPU communication entirely.
    • vLLM v0.31.0 ships ROCm 7.2.3, which has no published TP=2 reports on R9700 yet; v0.28.0 (the tag reported working at TP=2 on R9700) uses the same ROCm 7.2.3 base and is the escape hatch for vLLM-level errors.
    • vLLM v0.28–v0.30 can't load this checkpoint as shipped (algo_config crash in the Quark loader); use v0.31.0, or the AEON_STRIP_ALGO_CONFIG=1 escape hatch.
    • With this hybrid model, DFlash2 speculative decoding needs prefix caching off on vLLM 0.29–0.31 (vllm#55601, vllm#58894).
    • The first-start Gated DeltaNet compile hang on RDNA 4 (vllm#45929) is fixed in the images these recipes use.

Quality and validation status

Acceptance and quality were measured on an NVIDIA reference run of this exact checkpoint; AMD numbers are pending the first AMD runs.

Check Where Status
Export integrity: 336 MX layers written; all 15 MTP tensors and 333 vision tensors carried over in BF16 Quantization run Done
vLLM code path: Quark OCP MX loader, native or emulated kernel selection on each platform, TP=2/TP=4 MX sharding, MTP exclusion, algo_config handling (crashes on ≤ 0.30, fixed in 0.31) vLLM 0.29.0 and 0.31.0 source review Done (a code review, not a runtime test)
This checkpoint in vLLM with emulated MXFP6 and an equivalent W4A4 MLP path (same MXFP4 weight grid and activation quantize-dequantize as vLLM's emulation; differs only in GEMM accumulation order, so not bit-identical): plain, MTP (K=3) and DFlash2 (K=9) Reference run (TP=1, FP8 KV cache, vLLM 0.29-based) Done. Mean acceptance length: DFlash2 3.73 at T=0.6 (26 prompts) and 4.40 at T=0; MTP 2.76 vs DFlash2 4.20 on the same 13 prompts. Single request, same 4 prompts: plain 8.1 tok/s, MTP 13.5, DFlash2 22.0 (2.7x). GSM8K-20 / tools-10: plain 20/20 and 10/10, DFlash2 19/20 and 9/10 (two discordant items, McNemar p = 0.5). Speculative outputs are not bit-identical to plain decoding at T=0 (1 of 20 GSM8K outputs identical; verify batches change the reduction order). No garbage output, no refusals
MXFP6 dequant-once patch: output identical to stock emulation Reference run Done (torch.equal on the layer outputs)
Apple Silicon, via an MLX conversion Mac Pending
W4A4 / W6A6 quality against the BF16 master, using emulated MX numerics GPU, vLLM emulation Pending
2x Radeon AI PRO R9700, tester matrix (A-plain, A-MTP, A with DFlash2): smoke tests, GSM8K-50, IFEval-20, 10 tool calls, throughput and acceptance at c=1/4/8/16, TTFT, VRAM per GPU ROCm 7.2.3, vLLM 0.31.0 Pending (first tester run in progress)
TP=2 vs TP=1 numerical cross-check (MX weight sharding): prefill logprobs of A-plain vs one Recipe Z replica, scripts/tp_check.py 2x R9700 Pending (part of the tester matrix)
Other Radeon and Ryzen systems (Recipes B–H, H-128, K–Q) n/a Not yet tested
Instinct MI300X / MI325X / MI300A / MI250 / MI210 (emulated) and MI350X / MI355X n/a Not yet tested

Quantization recipe

Module group Linear layers Format Weights Activations Approx. size
MLP gate_proj / up_proj / down_proj 192 (64 layers × 3) MXFP4 FP4 E2M1, 32-element blocks, E8M0 scale MXFP4, dynamic per 32-element block 9.1 GB
GDN in_proj_qkv / in_proj_z / out_proj 144 (48 layers × 3) MXFP6 E2M3 FP6 E2M3, 32-element blocks, E8M0 scale MXFP6 E2M3, dynamic 4.3 GB
Full attention q/k/v/o_proj, plus q_norm/k_norm 64 (16 layers × 4) BF16 n/a n/a 3.4 GB
GDN recurrence: in_proj_a/b, conv1d, A_log, dt_bias, norm 96 linears + params BF16 n/a n/a 0.05 GB
embed_tokens and lm_head (untied) n/a BF16 n/a n/a 5.1 GB
Vision tower (27 blocks + merger) n/a BF16 n/a n/a 0.9 GB
MTP head (1 layer + fc) n/a BF16 n/a n/a 0.85 GB
  • Tool: AMD Quark 0.12.post1 in eager mode, with Transformers 5.14.1 and PyTorch 2.13. The export is real_quantized: packed uint8 weights plus uint8 E8M0 scales. The Quark MX settings are scale_calculation_mode: even and round_method: half_even.
  • Algorithm: AutoSmoothQuant with an MSE scale search, applied to the MLP only, along the edges post_attention_layernorm → gate/up and up_proj → down_proj. The smoothing factors are folded into those norms and weights. The GDN projections have no safe smoothing edge, so they get none.
  • Calibration: 1,024 chat, code and math samples (UltraChat, Open-Platypus, CodeAlpaca, GSM8K), each truncated to 64 tokens, for 65,536 tokens in total. To fit in memory, the AutoSmoothQuant scale search used a 64-token subsample per layer. MX weight scales come from the weights themselves and activations are quantized at runtime, so calibration data only affects the MLP smoothing factors.
  • Exact config: see quantization_config in config.json and BAKE_STAMP_Q2.json.

Model family

Variant Format Size Best hardware Status Link
BF16 master BF16, with vision and MTP 55.6 GB H200, multi-GPU, RTX PRO 6000 Public AEON-7/…-BF16
NVFP4-MIXED NVFP4 MLP (layers 0–55); FP8 attention, GDN writers and MLP layers 56–63; BF16 for the rest (ModelOpt) 24.7 GB DGX Spark (GB10), RTX 5090, RTX PRO 6000 Public AEON-7/…-NVFP4-MIXED
MXFP4-MXFP6-ROCm (this repo) MXFP4 MLP and MXFP6 GDN projections; BF16 for the rest (AMD Quark) 23.7 GB Instinct MI350/MI355 (native MXFP4); Radeon AI PRO R9700 32 GB (emulated) Early Release · Gated Access AEON-7/…-MXFP4-MXFP6-ROCm
AEON DFlash2 drafter (bundled in dflash2/ of this repo) DFlash 2 block-diffusion drafter, BF16 (AEON fine-tune of z-lab/Qwen3.8-27B-DFlash2) 3.85 GB Ships with this repo for its DFlash2 recipes (featured: 2x Radeon AI PRO R9700). The bundle is the earlier lk3 BF16 build; the newer v1.0 (lk4) is in AEON-7/AEON-DFlash2-Qwen3.8-27B (bf16/ folder) Early Release · Gated Access; untested on AMD dflash2/ (lk3) · v1.0 (lk4)
AEON DFlash2 drafter v1.0 (lk4) DFlash 2 drafter: NVFP4 W4A16 pack + exact BF16 copy (bf16/) 1.9 GB / 3.85 GB DGX Spark, RTX 50-series (vLLM) Early Release · Gated Access AEON-7/AEON-DFlash2-Qwen3.8-27B
NVFP4-GDNFP8 NVFP4 MLP with FP8 GDN projections (LLM Compressor) TBA NVIDIA Blackwell Early access (Patreon); not public n/a
FP8-MIXED FP8 dynamic MLP; BF16 for the rest (LLM Compressor) 38.5 GB FP8-capable GPUs with 48 GB or more Internal: built, not yet validated n/a
MLX (Apple Silicon) MLX conversion TBA Apple Silicon Macs Planned n/a

User responsibility

By accessing, downloading or running this model, you agree to the following:

  1. You are responsible for its use. You alone are responsible for every prompt, every response, every downstream action, and any harm that results.
  2. No warranty. The model is provided "AS IS", without warranty of any kind.
  3. Follow the law. You must comply with all applicable laws and policies in every jurisdiction you operate in.
  4. Add safety layers in production. Use input validation, output filtering, access controls, and human review for high-risk workflows.
  5. The duty of care is yours. An uncensored model doesn't refuse on your behalf, so if you are unsure about a request, don't make it.
  6. No endorsement. The authors do not endorse any particular output.
  7. Arbitration. Disputes go to binding individual arbitration (under the AAA Consumer Rules if no other body is agreed), waiving jury trials and class actions.
  8. Indemnification. You indemnify the authors, contributors and publishers against claims arising from your use.
  9. Severability. If a provision is invalid, the closest enforceable equivalent replaces it.
  10. Acceptance. Using the model means you accept these terms. If you don't accept them, don't use the model.

License and attribution

  • License: Apache-2.0, inherited from Qwen/Qwen3.8-27B.
  • Qwen team (Alibaba): the Qwen3.8-27B base model.
  • AEON-7: the uncensored fine-tune (the BF16 master) and this quantization. The upstream tools behind the fine-tune are credited on the BF16 card.
  • AMD Quark: the quantization toolkit used to produce the OCP MX export.
Downloads last month
4
Safetensors
Model size
18B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm

Base model

Qwen/Qwen3.8-27B
Quantized
(33)
this model

Collection including AEON-7/Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-MXFP4-MXFP6-ROCm