YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

HunyuanImage-3.0-Instruct-Distil · MXFP4 mixed-precision with tuning (auto-round 0.16.0 / vLLM INC path)

Same MXFP4-for-routed-experts / MXFP8-for-everything-else coverage as -MXFP4-ar, but produced by a tuning run (iters=50, dataset=coco2014, 128 calibration forwards per layer) instead of a --model_free run, and exported by AutoRound 0.16.0 instead of 0.15.0.

This is the first artifact in this family that loads on vLLM/vLLM-Omni with no metadata edit at all — see Read this first.

AR+DiT full-pipeline output

Generated output, not a ground-truth reference. AR+DiT on 2 GPUs (AR TP1 on cuda:0 + DiT TP1 on cuda:1), prompt A cute cat, seed=42, 8 steps, guidance_scale=5.0. Captured 2026-09-22 07:53.


⚠️ Read this first

1. No patch, no normalizer — unlike a stock 0.15.0 export

The 0.15.0 builds (-MXFP4-ar) ship --layer_config into extra_config verbatim as the key ".mlp.experts.", which vLLM's INC parser cannot match, so the 4-bit override is silently dropped and the checkpoint dies later with KeyError: '…w2_weight_packed'. That build therefore needs tools/normalize_experts_extra_config_key.py.

This build does not. AutoRound 0.16.0 writes the pattern as a regex — .*mlp\.experts.* — which contains metacharacters, so the INC parser takes its regex branch and matches the container name model.layers.N.mlp.experts. The three spellings actually present on this machine:

producer extra_config key INC matches it?
0.15.0 stock export (-MXFP4-ar/config.json.pristine) .mlp.experts. ❌ all four branches miss → falls back to bits:8
0.15.0 + local normalizer (-MXFP4-ar/config.json) .*mlp\.experts ✅
0.16.0 stock export (this build) .*mlp\.experts.* ✅ — no edit needed

The .*….* wrapping is AutoRound's own normalizer (auto_round/utils/common.py:838-844: escape the literal dots, then regex = f".*{pattern}.*"), so on 0.16.0 the key that reaches vLLM is already a real regex. vLLM's side is unchanged — only the metadata carrier differs.

Resolution verified by asking the real installed vLLM code, not by reading it (python quantization_analysis/scripts/check_autoround_experts_key.py <this dir>, exit code 0):

quant_method = 'auto-round'   packing_format = 'auto_round:llm_compressor'
override_quantization_method → 'inc'

  routed experts 容器           bits=4   INCMxfp4Scheme      → MarlinExperts (weight-only FP4)
    └ 探测名 gate_proj           bits=4   INCMxfp4Scheme
  layer31 experts              bits=4   INCMxfp4Scheme
  shared_mlp                   bits=8   INCMxfp8Scheme      → MarlinMxfp8LinearKernel
  self_attn.qkv_proj           bits=8   INCMxfp8Scheme
  MoE router gate.wg          bits=16   (未量化 → Unquantized*)
  ViT fc2                     bits=16   (未量化 → Unquantized*)

Confirmed at runtime in both run modes (log lines in quantization_analysis/latest/runpack/logs/run_run_mxfp4ar_tuning_{dit,full}.log):

Using MarlinMxfp8LinearKernel for MXFP8 GEMM
Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE

2. On Hopper (SM90: H100/H200) this is a memory saving, not a speed saving

There are no FP4 tensor cores on SM90, so the experts run through Marlin as weight-only FP4 (effectively W4A16). extra_config declares act_bits: 4 for the experts (a W4A4 intent), and that declaration is inert here — nothing on this GPU consumes it. Expect a 4-bit memory footprint with 16-bit activation compute. vLLM says so itself, in every run:

WARNING [inc_mxfp4_moe.py:204] This device lacks native FP4 compute; using weight-only FP4 via the
         Marlin kernel, which may reduce performance for compute-heavy workloads.

Overview

Field Value
Base model tencent/HunyuanImage-3.0-Instruct-Distil (cfg_distilled=true, use_meanflow=true) — numerically confirmed: max|Δ| = 0, relL2 = 0.0000 for the un-quantized router tensors vs the local BF16 copy
MoE geometry 32 layers × 64 routed experts, moe_topk=8, 1 shared expert/layer, hidden 4096, moe_intermediate=3072
Scheme Mixed: routed experts → MXFP4 (E2M1, group 32, E8M0); all other quantized Linear → MXFP8 (E4M3, group 32, E8M0)
Export format --format auto_round → quant_method="auto-round", packing_format="auto_round:llm_compressor"
Loader path vLLM INC (INCConfig.override_quantization_method() maps "auto-round" → "inc")
Quantization tool auto-round 0.16.0, with tuning (iters=50, dataset=coco2014, nsamples=16, enable_quanted_input=false)
Disk size 52.61 GB (49.00 GiB) in 11 shards — BF16 base 158 GB ⇒ 0.33×
Parameters 83.045 B total: routed experts 77.31 B @4-bit (93.1 %), self_attn 1.49 B + shared_mlp 1.21 B @8-bit, 3.24 B BF16/FP32
Tensor count 9 393 — identical to -MXFP4-ar
Kept unquantized ViT (vision_model, BF16), vision_aligner (BF16), VAE (FP32, as in the base), wte/lm_head, guidance_emb/timestep_emb/timestep_r_emb/patch_embed/final_layer/time_embed*, layernorms, and the MoE router mlp.gate.wg of all 32 blocks
⚠ Router dtype stored float32 here (0.15.0 wrote bfloat16). Values are identical to the base — this is a serialization difference, not a retrained router (+16.8 MB on disk)

quantization_config from config.json (verbatim, extra_config collapsed):

{
  "quant_method": "auto-round",
  "packing_format": "auto_round:llm_compressor",
  "bits": 8, "group_size": 32, "sym": true, "data_type": "mx_fp",
  "act_bits": 8, "act_data_type": "mx_fp", "act_dynamic": true, "act_group_size": 32, "act_sym": true,
  "static_attention_granularity": "tensor", "static_kv_granularity": "tensor",
  "iters": 50, "enable_quanted_input": false,
  "autoround_version": "0.16.0",
  "block_name_to_quantize": ["model.layers.0", "…", "model.layers.31"],
  "extra_config": {
    ".*mlp\\.experts.*":                 {"bits": 4,  "act_bits": 4},       // ← the matchable key
    "model.layers.N.mlp.experts.M.*":   {"bits": 4,  "act_bits": 4},       // 4096 literal expert entries
    "model.layers.N.mlp.gate.wg":       {"bits": 16, "data_type": "float"}, // 32 literal entries
    ".*model\\.layers\\.N\\.mlp\\.gate.*": {"bits": 16, "data_type": "float"},
    ".*vae.* / .*vit.* / .*vit_aligner.*":  {"bits": 16, "data_type": "float"}
  }
}

extra_config holds 4 164 entries: 4 097 at bits:4 (= 4 096 concrete expert projections

  • 1 regex) and 67 at bits:16 (= 32 literal mlp.gate.wg + 32 .*…mlp\.gate.* regexes + .*vae.*, .*vit.*, .*vit_aligner.*). There is no ignore list in this format — 16-bit exceptions are extra_config entries — and everything outside model.layers.0…31 is untouched by construction because block_name_to_quantize is scoped to those 32 blocks.

act_* fields describe the declared scheme. What actually runs on SM90: dynamic MXFP8 activation quantization for dense Linear, no activation quantization for the experts (Marlin W4A16).


Coverage vs. the sibling -MXFP4-ar build (measured, not asserted)

Full safetensors-header census of both builds (9 393 tensors each, 83.045 B parameters each):

Module family this build -MXFP4-ar dtypes (this / -ar)
routed experts 77.309 B 77.309 B U8×8192 / U8×8192
self_attn 1.486 B 1.486 B BF16×280, F8_E4M3×64, U8×64 / same
VAE 1.261 B 1.261 B F32×280 / same
shared_mlp 1.208 B 1.208 B F8_E4M3×64, U8×64 / same
wte/lm_head 1.091 B 1.091 B BF16×2 / same
other (embeddings, DiT head, …) 0.376 B 0.376 B BF16×115 / same
ViT 0.306 B 0.306 B BF16×236 / same
MoE router mlp.gate.wg 0.008 B 0.008 B F32×32 / BF16×32 ← only difference

⇒ Which modules are quantized, and at what width, is identical. The values are not — see next section.


Quantization command

Reconstructed from calibration_recipe.json, which is shipped in this directory and is AutoRound's own dump of the run (every field below is quoted from it; only --model_name/--output_dir are the host paths recorded there):

auto-round \
  --model_name /models/HunyuanImage-3-Instruct-Distil \
  --scheme MXFP8 \
  --layer_config '{mlp.experts:{scheme:MXFP4}}' \
  --iters 50 \
  --dataset coco2014 \
  --nsamples 16 \
  --format auto_round \
  --device cuda:0 \
  --output_dir /models/changwa1/HunyuanImage-3-Instruct-Distil-Mixed-AutoRound

Recipe fields (verbatim): scheme=MXFP8, layer_config={mlp.experts: {scheme: MXFP4}}, iters=50, nsamples=16, dataset=coco2014, num_inference_steps=8, calib_num_inference_steps=8, image_size=1024x1024, guidance_scale=5.0, seed=42, format=auto_round, bot_task=image, calib_cache_dir=/models/changwa1/hunyuan-calib-cache, device=0. Per-layer calibration record: forwards = 128 for each of model.layers.0…31, plus 16 calibration_schedules (seeds 42–57) — i.e. the tuner saw real 8-step diffusion trajectories, unlike the --model_free sibling which saw none.

Self-check that the export is loadable before serving it:

python /path/to/quantization_analysis/scripts/check_autoround_experts_key.py <this model dir>
# exit 0 = experts resolve to 4-bit;  exit 1 = 4-bit override lost, metadata needs normalizing

Inference environment

Component Version
vLLM 0.29.0
vLLM-Omni latest main (validated at 1c7476ec19899f1e61838e23e9fad3505403b9b9, 0.29.0rc2.dev161+g1c7476ec1, editable install)
PyTorch 2.13.0+cu132
FlashInfer 0.6.18
Python 3.12.3
GPU NVIDIA H200 141 GB — AR-only 1×, DiT-only 2×, AR+DiT 2×
python -m venv ~/.venv-omini-latest && source ~/.venv-omini-latest/bin/activate
pip install vllm==0.29.0 && pip install -e /path/to/vllm-omni
python -c "import vllm, vllm_omni, os; print(vllm.__version__, os.path.dirname(vllm_omni.__file__))"

# host-specific, drop if unneeded here
export NCCL_NVLS_ENABLE=0
export VLLM_USE_FLASHINFER_SAMPLER=0
cd /tmp            # cwd must be neutral: children re-import vllm_omni from cwd

A full reproducible environment builder (constraints file + patch set + verifier) lives in quantization_analysis/envkit/.


Validated run modes

Mode GPUs Result Evidence from logs
DiT-only (TP2) 2 (cards 0,1) ✅ image Stage 0 … mapping: 0->0, 1->1; Model loading took 25.3899 GiB per card; image_pixels=1048576; denoise_step_latency_ms=1845.98; peak_memory_mb=39482; Saved generated image to …/run_mxfp4ar_tuning_dit.png
AR+DiT (AR TP1 + DiT TP1) 2 (cards 0,1) ✅ image 2 stages (0->0, 1->1); DiT Model loading took 46.6297 GiB, AR 46.69 GiB; AR ratio_idx=19, target size=1216x832; peak_memory_mb=61190; Saved generated image to …/run_mxfp4ar_tuning_full.png
AR-only (TP1) 1 (card 2) ✅ text Model loading took 46.69 GiB memory and 55.15 seconds; Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE; [run_ar] EXIT=0; output text It's a typical Photography (27 bytes)

⚠️ AR-only text is evidence that the AR path loads and generates, nothing more. Across this workspace the AR-only captures range from 7 to 1198 bytes and x_to_text.py does not echo the request sampling parameters, so token-level agreement with other builds cannot be claimed. In the full pipeline this build's AR emitted 8 tokens → ratio_idx=19 → target size=1216x832, whereas BF16/MXFP8 runs emit 6 tokens / ratio_idx=14 and the 0.15.0 MXFP4 builds were seen at 10 tokens / ratio_idx=14 and at 295 tokens / ratio_idx=16 on repeated runs of the same command. The AR branch is the dominant source of full-pipeline image divergence here — see Fidelity before attributing any of it to tuning.

All output PNGs are 1024×1024 in every mode (17/17 full-pipeline logs in this workspace checked against their own saved files). The target size=… line in the AR→DiT handoff is the AR's predicted aspect — ar2diffusion computes it from reso_group[ratio_idx] and logs it (stage_input_processors/hunyuan_image3.py:139-191), but the DiT takes its canvas from sampling.height or 1024 (pipeline_hunyuan_image3.py:1912), which the offline example already fills with --height/--width defaults of 1024 (text_to_image.py:125-126). ⇒ ratio_idx changes the AR text and the log line, not the saved canvas, so cross-image PSNR/SSIM here involves no mismatched-canvas resampling, and any full-pipeline divergence must come from the conditioning text, not the aspect.

# ── DiT-only, 2 GPUs ─────────────────────────────────────────────
export CUDA_VISIBLE_DEVICES=0,1
cd /tmp
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
  --model                /path/to/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound \
  --deploy-config        ./hunyuan_image3_dit_tp2.yaml \
  --prompt               "A cute cat" \
  --num-inference-steps  8 \
  --guidance-scale       5.0 \
  --seed                 42 \
  --init-timeout         1800 --stage-init-timeout 1500 \
  --output               ./dit_only.png

# ── AR+DiT, 2 GPUs ───────────────────────────────────────────────
export CUDA_VISIBLE_DEVICES=0,1
cd /tmp
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
  --model                /path/to/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound \
  --deploy-config        ./hunyuan_image_3_moe_2gpu_tp1.yaml \
  --prompt               "A cute cat" \
  --num-inference-steps  8 --guidance-scale 5.0 --seed 42 \
  --init-timeout         1800 --stage-init-timeout 1500 \
  --output               ./ar_dit.png

# ── AR-only, 1 GPU (different example script; it does NOT accept --init-timeout) ──
export CUDA_VISIBLE_DEVICES=2
cd /tmp
python /path/to/vllm-omni/examples/offline_inference/x_to_text/x_to_text.py \
  --model          /path/to/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound \
  --deploy-config  ./hunyuan_image3_ar_tp1.yaml \
  --trust-remote-code \
  --prompt         "A cute cat" --max-tokens 256 \
  --output         ./ar_only.txt

--deploy-config files and the wrapper scripts (run_dit.sh / run_ar.sh) with fully expanded commands, timeout rationale and per-mode log expectations are in quantization_analysis/latest/runpack/README.md.

Parameter notes — --num-inference-steps 8 (distilled checkpoint: do not use 50); --guidance-scale is a real input (cfg_distilled=true; vLLM-Omni feeds 1000 × guidance_scale as the guidance embedding; the images above used 5.0, upstream's Distil e2e reference uses 2.5); --prompt reaches only the AR stage.


Fidelity

Weight side — tuning moves the weights away from the base, on purpose

Dequantizing this build and -MXFP4-ar (same E2M1/MXFP4 nibble + E8M0 group-32 and E4M3 + E8M0 group-32 codewords) and comparing against the local BF16 base:

Module -MXFP4-ar (model_free / RTN) this build (iters=50 tuned)
self_attn.qkv_proj L0 0.0267 0.0572
mlp.shared_mlp.gate_and_up_proj L0 0.0267 0.0476
self_attn.o_proj L31 0.0267 0.0640
mlp.shared_mlp.down_proj L31 0.0267 0.0465
mlp.experts.3.gate_and_up_proj L0 0.1123 0.1285
mlp.experts.40.down_proj L0 0.1125 0.1260
mlp.experts.11.gate_and_up_proj L15 0.1113 0.1239
mlp.experts.63.down_proj L31 0.1124 0.1267
mean MXFP8 dense 0.0267 0.0538 (2.01×)
mean MXFP4 experts 0.1121 0.1263 (+12.7 %)
mlp.gate.wg router (L0, L31) 0.0000 0.0000 (max|Δ| = 0)

Read this correctly:

  • The -MXFP4-ar column reproduces the format floor exactly (0.0267 / ≈0.1122 = ideal per-group E8M0 RTN), which is what validates the measurement harness.
  • The tuned weights deviate more from the base, not less. That is expected: iters=50 optimizes layer output error on 128 real diffusion forwards per block, so it deliberately moves weights (and rounds differently) instead of minimizing weight-space distance. Weight rel-L2 vs the base is no longer a valid proxy for quality on a tuned checkpoint — do not compare a tuned build against an RTN build with it.
  • Router is bit-exact with the base in both builds ⇒ the router was never tuned, only re-serialized as FP32.

Image side — ⚠️ cross-session numbers, read the calibration row first

Pixel-level agreement in this workspace is bounded by session, not by quantization. Re-shooting the same build with the same command in a different session gives:

Situation PSNR vs the old capture SSIM@1/8
same build, same session re-run 41.35 – 44.51 dB (BF16: 74.2 dB, `max Δ
same build, different session 12.2 – 12.5 dB 0.223 – 0.233

⇒ Any pair measured in different sessions sits at ≈12 dB / ≈0.23 by drift alone. The DiT-only table below mixes 09-17, 09-18 and 09-22 captures, so treat the two quantized rows as "≈ drift level", not as a ranking.

DiT-only, TP2 on cards 0,1 for all cells (reference = local BF16, 09-17 22:01):

Candidate captured PSNR SSIM@1 SSIM@1/8 mean|Δ| std
MXFP8-ar (-MXFP8, auto-round/INC) 09-18 00:38 31.79 0.9811 0.9842 3.56 53.74
MXFP8-ct (-MXFP8-ct, compressed-tensors) 09-18 31.59 0.9788 0.9823 3.62 53.62
MXFP4-ar (-MXFP4-ar, model_free) 09-22 13.00 0.6720 0.2704 44.20 49.09
this build (MXFP4 + tuning) 09-22 07:15 12.88 0.6640 0.2568 45.31 45.70

Figure A — DiT-only: BF16 / MXFP8-ar / this build.

DiT-only montage

Direct pairings of this build against other captures (DiT-only): vs MXFP8-ar 12.69 dB / SSIM@1/8 0.2561; vs MXFP4-ar 22.06 dB / SSIM@1/8 0.7904 — i.e. the tuned MXFP4 build is much closer to the untuned MXFP4 build than either is to BF16 or to MXFP8, which is the signature you expect when the dominant difference is the experts' 4-bit codebook rather than the 50 tuning iterations.

AR+DiT (reference = BF16 on 4 GPUs, TP2+TP2 — cross-topology, and the AR stage re-draws its CoT every run, so this mode is the least reproducible of the three):

Candidate topology PSNR SSIM@1 SSIM@1/8 std
MXFP8-ar 2 GPUs TP1+TP1 26.85 0.9227 0.8715 49.88
MXFP4-ar 2 GPUs TP1+TP1 12.89 0.6934 0.1708 46.44
this build 2 GPUs TP1+TP1 13.54 0.6423 0.1766 43.02

Figure B — AR+DiT: BF16 / MXFP8-ar / this build.

AR+DiT montage

Individual cells, uncropped:

BF16 (DiT TP2) MXFP8-ar (DiT TP2) this build (DiT TP2)
a b c
BF16 (AR+DiT, 4 GPU) MXFP8-ar (AR+DiT, 2 GPU) this build (AR+DiT, 2 GPU)
d e f

All four cells in the bottom row show a coherent, un-degraded cute-cat image; the MXFP4 cells differ from BF16/MXFP8 in composition (SSIM falls as the scale gets coarser ⇒ structure, not texture), which is the same qualitative finding as the -MXFP4-ar build. agent_smoke_dit_only.png is this agent's own independent DiT-only reproduction of the tuned checkpoint (std 46.06), kept for cross-checking.

Not measured

No task-level benchmark (GenEval / DPG-Bench / CVTG-2K / DrawBench / WISE) has been run for this checkpoint. "Composition diverges from BF16" is measured; "quality is better or worse than -MXFP4-ar because of tuning" is not established — and cannot be from one prompt.

Reproduce every number in this card

PY=/home/kaokaolv/.venv-omini-latest/bin/python
QA=/home/kaokaolv/quantization_analysis
TUNE=/home/kaokaolv/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound
AR=/home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar

# 0) does the checkpoint resolve to 4-bit experts on the installed vLLM?  (exit 0 = yes)
$PY $QA/scripts/check_autoround_experts_key.py $TUNE

# 1) coverage census ("Coverage vs. the sibling -MXFP4-ar build")
$PY $QA/scripts/census_family_compare.py $AR $TUNE

# 2) weight-space error ("Weight side")
$PY $QA/scripts/dequant_error_vs_base.py $AR $TUNE /home/kaokaolv/models/HunyuanImage-3.0-Instruct-Distil

# 3) image metrics (any pair; --scales 1,8 = fine + coarse SSIM)
$PY $QA/bench/eval_scripts/compare_images.py \
    --ref $QA/latest/runpack/out/bf16_dit_tp2.png \
    --cands $QA/latest/runpack/out/pr_ar_dit.png $TUNE/images/user_dit_only.png --scales 1,8

# 4) the two montages (Figure A / Figure B)
$PY $QA/envkit/make_card_grid.py --out $QA/latest/runpack/out/grid_tuning_ditonly.png \
    --ref "① BF16 base|DiT-only TP2|captured 09-17 22:01"=$QA/latest/runpack/out/bf16_dit_tp2.png \
    --cell "② MXFP8-ar (auto-round/INC)|DiT-only TP2|captured 09-18 00:38"=$QA/latest/runpack/out/pr_ar_dit.png \
    --cell "③ MXFP4-mixed +tuning (0.16.0)|DiT-only TP2|captured 09-22 07:15"=$TUNE/images/user_dit_only.png \
    --scales 1,8 --width 460

Known limitations

  1. Memory-only on Hopper — experts run W4A16 through Marlin; the declared act_bits: 4 is inert.
  2. Cross-session pixel drift (~12 dB) is larger than the tuning effect being measured here, so the image tables are illustrative, not quantitative. Same-session re-shoots are the only way to rank iters=50 against iters=0.
  3. The full-pipeline images are not comparable at the pixel level at all — the AR stage regenerates its CoT/ratio tokens per run and those tokens are the DiT conditioning.
  4. Composition diverges from BF16/MXFP8 (SSIM@1/8 ≈ 0.26 in DiT-only) — the MXFP4 experts, not the tuning, are the dominant cause.
  5. The ViT stays BF16 (vision_model.encoder.*.mlp.fc2 input dim 4304 is not divisible by the 32-element MX group) ⇒ not a whole-model MXFP4 build.
  6. Router tensors are FP32 here vs BF16 in the 0.15.0 builds — harmless for vLLM, but it breaks naive byte-hash comparisons across builds (a real trap: an md5-based "different base model" conclusion drawn from this was wrong; the values are equal).
  7. block_name_to_quantize = [model.layers.0 … model.layers.31] scopes quantization to the 32 shared backbone layers; AR and DiT load the same tensors through different stacks.
  8. images/ holds verification snapshots only; it is not read when loading weights.

Contents

config.json                                 quant_method=auto-round, extra_config (4164 entries) — stock export,
                                            nothing edited; loads as-is
quantization_config.json                    flat secondary discovery carrier
calibration_recipe.json                     AutoRound's own dump of the tuning run (the source of the
                                            command above)
model-000NN-of-00011.safetensors     52.61 GB total, 9393 tensors
model.safetensors.index.json
images/
  user_ar_dit.png                           AR+DiT, 2 GPUs (hero)
  user_dit_only.png                         DiT-only TP2, this build
  agent_smoke_dit_only.png                  independent DiT-only re-run by the validating agent
  ar_only_output.txt                        AR-only output text (`It's a typical Photography`, 27 B)
  bf16_dit_only_reference.png               BF16 base, DiT-only TP2
  mxfp8_ar_dit_only.png                     -MXFP8 (auto_round) DiT-only TP2
  bf16_ar_dit_4gpu.png                      BF16 base, AR+DiT on 4 GPUs
  mxfp8_ar_ar_dit.png                       -MXFP8 (auto_round), AR+DiT on 2 GPUs
  grid_tuning_ditonly.png                   Figure A (BF16 / MXFP8-ar / this build, DiT-only)
  grid_tuning_ardit.png                     Figure B (same three, AR+DiT)
*.py / tokenizer* / utils/                  inherited from the base model
Downloads last month
14
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support