YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
HunyuanImage-3.0-Instruct-Distil · MXFP4 mixed-precision with tuning (auto-round 0.16.0 / vLLM INC path)
Same MXFP4-for-routed-experts / MXFP8-for-everything-else coverage as
-MXFP4-ar, but produced by a tuning run
(iters=50, dataset=coco2014, 128 calibration forwards per layer) instead of a --model_free run, and
exported by AutoRound 0.16.0 instead of 0.15.0.
This is the first artifact in this family that loads on vLLM/vLLM-Omni with no metadata edit at all — see Read this first.
Generated output, not a ground-truth reference. AR+DiT on 2 GPUs (AR TP1 on cuda:0 + DiT TP1 on cuda:1), prompt
A cute cat,seed=42, 8 steps,guidance_scale=5.0. Captured 2026-09-22 07:53.
⚠️ Read this first
1. No patch, no normalizer — unlike a stock 0.15.0 export
The 0.15.0 builds (-MXFP4-ar) ship --layer_config into extra_config verbatim as the key
".mlp.experts.", which vLLM's INC parser cannot match, so the 4-bit override is silently dropped and the
checkpoint dies later with KeyError: '…w2_weight_packed'. That build therefore needs
tools/normalize_experts_extra_config_key.py.
This build does not. AutoRound 0.16.0 writes the pattern as a regex — .*mlp\.experts.* — which
contains metacharacters, so the INC parser takes its regex branch and matches the container name
model.layers.N.mlp.experts. The three spellings actually present on this machine:
| producer | extra_config key |
INC matches it? |
|---|---|---|
0.15.0 stock export (-MXFP4-ar/config.json.pristine) |
.mlp.experts. |
❌ all four branches miss → falls back to bits:8 |
0.15.0 + local normalizer (-MXFP4-ar/config.json) |
.*mlp\.experts |
✅ |
| 0.16.0 stock export (this build) | .*mlp\.experts.* |
✅ — no edit needed |
The .*….* wrapping is AutoRound's own normalizer (auto_round/utils/common.py:838-844: escape the
literal dots, then regex = f".*{pattern}.*"), so on 0.16.0 the key that reaches vLLM is already a real
regex. vLLM's side is unchanged — only the metadata carrier differs.
Resolution verified by asking the real installed vLLM code, not by reading it
(python quantization_analysis/scripts/check_autoround_experts_key.py <this dir>, exit code 0):
quant_method = 'auto-round' packing_format = 'auto_round:llm_compressor'
override_quantization_method → 'inc'
routed experts 容器 bits=4 INCMxfp4Scheme → MarlinExperts (weight-only FP4)
└ 探测名 gate_proj bits=4 INCMxfp4Scheme
layer31 experts bits=4 INCMxfp4Scheme
shared_mlp bits=8 INCMxfp8Scheme → MarlinMxfp8LinearKernel
self_attn.qkv_proj bits=8 INCMxfp8Scheme
MoE router gate.wg bits=16 (未量化 → Unquantized*)
ViT fc2 bits=16 (未量化 → Unquantized*)
Confirmed at runtime in both run modes (log lines in
quantization_analysis/latest/runpack/logs/run_run_mxfp4ar_tuning_{dit,full}.log):
Using MarlinMxfp8LinearKernel for MXFP8 GEMM
Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE
2. On Hopper (SM90: H100/H200) this is a memory saving, not a speed saving
There are no FP4 tensor cores on SM90, so the experts run through Marlin as weight-only FP4
(effectively W4A16). extra_config declares act_bits: 4 for the experts (a W4A4 intent), and that
declaration is inert here — nothing on this GPU consumes it. Expect a 4-bit memory footprint with
16-bit activation compute. vLLM says so itself, in every run:
WARNING [inc_mxfp4_moe.py:204] This device lacks native FP4 compute; using weight-only FP4 via the
Marlin kernel, which may reduce performance for compute-heavy workloads.
Overview
| Field | Value |
|---|---|
| Base model | tencent/HunyuanImage-3.0-Instruct-Distil (cfg_distilled=true, use_meanflow=true) — numerically confirmed: max|Δ| = 0, relL2 = 0.0000 for the un-quantized router tensors vs the local BF16 copy |
| MoE geometry | 32 layers × 64 routed experts, moe_topk=8, 1 shared expert/layer, hidden 4096, moe_intermediate=3072 |
| Scheme | Mixed: routed experts → MXFP4 (E2M1, group 32, E8M0); all other quantized Linear → MXFP8 (E4M3, group 32, E8M0) |
| Export format | --format auto_round → quant_method="auto-round", packing_format="auto_round:llm_compressor" |
| Loader path | vLLM INC (INCConfig.override_quantization_method() maps "auto-round" → "inc") |
| Quantization tool | auto-round 0.16.0, with tuning (iters=50, dataset=coco2014, nsamples=16, enable_quanted_input=false) |
| Disk size | 52.61 GB (49.00 GiB) in 11 shards — BF16 base 158 GB ⇒ 0.33× |
| Parameters | 83.045 B total: routed experts 77.31 B @4-bit (93.1 %), self_attn 1.49 B + shared_mlp 1.21 B @8-bit, 3.24 B BF16/FP32 |
| Tensor count | 9 393 — identical to -MXFP4-ar |
| Kept unquantized | ViT (vision_model, BF16), vision_aligner (BF16), VAE (FP32, as in the base), wte/lm_head, guidance_emb/timestep_emb/timestep_r_emb/patch_embed/final_layer/time_embed*, layernorms, and the MoE router mlp.gate.wg of all 32 blocks |
| ⚠ Router dtype | stored float32 here (0.15.0 wrote bfloat16). Values are identical to the base — this is a serialization difference, not a retrained router (+16.8 MB on disk) |
quantization_config from config.json (verbatim, extra_config collapsed):
{
"quant_method": "auto-round",
"packing_format": "auto_round:llm_compressor",
"bits": 8, "group_size": 32, "sym": true, "data_type": "mx_fp",
"act_bits": 8, "act_data_type": "mx_fp", "act_dynamic": true, "act_group_size": 32, "act_sym": true,
"static_attention_granularity": "tensor", "static_kv_granularity": "tensor",
"iters": 50, "enable_quanted_input": false,
"autoround_version": "0.16.0",
"block_name_to_quantize": ["model.layers.0", "…", "model.layers.31"],
"extra_config": {
".*mlp\\.experts.*": {"bits": 4, "act_bits": 4}, // ← the matchable key
"model.layers.N.mlp.experts.M.*": {"bits": 4, "act_bits": 4}, // 4096 literal expert entries
"model.layers.N.mlp.gate.wg": {"bits": 16, "data_type": "float"}, // 32 literal entries
".*model\\.layers\\.N\\.mlp\\.gate.*": {"bits": 16, "data_type": "float"},
".*vae.* / .*vit.* / .*vit_aligner.*": {"bits": 16, "data_type": "float"}
}
}
extra_config holds 4 164 entries: 4 097 at bits:4 (= 4 096 concrete expert projections
- 1 regex) and 67 at
bits:16(= 32 literalmlp.gate.wg+ 32.*…mlp\.gate.*regexes +.*vae.*,.*vit.*,.*vit_aligner.*). There is noignorelist in this format — 16-bit exceptions areextra_configentries — and everything outsidemodel.layers.0…31is untouched by construction becauseblock_name_to_quantizeis scoped to those 32 blocks.
act_*fields describe the declared scheme. What actually runs on SM90: dynamic MXFP8 activation quantization for denseLinear, no activation quantization for the experts (Marlin W4A16).
Coverage vs. the sibling -MXFP4-ar build (measured, not asserted)
Full safetensors-header census of both builds (9 393 tensors each, 83.045 B parameters each):
| Module family | this build | -MXFP4-ar |
dtypes (this / -ar) |
|---|---|---|---|
| routed experts | 77.309 B | 77.309 B | U8×8192 / U8×8192 |
self_attn |
1.486 B | 1.486 B | BF16×280, F8_E4M3×64, U8×64 / same |
| VAE | 1.261 B | 1.261 B | F32×280 / same |
shared_mlp |
1.208 B | 1.208 B | F8_E4M3×64, U8×64 / same |
wte/lm_head |
1.091 B | 1.091 B | BF16×2 / same |
| other (embeddings, DiT head, …) | 0.376 B | 0.376 B | BF16×115 / same |
| ViT | 0.306 B | 0.306 B | BF16×236 / same |
MoE router mlp.gate.wg |
0.008 B | 0.008 B | F32×32 / BF16×32 ← only difference |
⇒ Which modules are quantized, and at what width, is identical. The values are not — see next section.
Quantization command
Reconstructed from calibration_recipe.json, which is shipped in this directory and is AutoRound's own
dump of the run (every field below is quoted from it; only --model_name/--output_dir are the host
paths recorded there):
auto-round \
--model_name /models/HunyuanImage-3-Instruct-Distil \
--scheme MXFP8 \
--layer_config '{mlp.experts:{scheme:MXFP4}}' \
--iters 50 \
--dataset coco2014 \
--nsamples 16 \
--format auto_round \
--device cuda:0 \
--output_dir /models/changwa1/HunyuanImage-3-Instruct-Distil-Mixed-AutoRound
Recipe fields (verbatim): scheme=MXFP8, layer_config={mlp.experts: {scheme: MXFP4}}, iters=50,
nsamples=16, dataset=coco2014, num_inference_steps=8, calib_num_inference_steps=8,
image_size=1024x1024, guidance_scale=5.0, seed=42, format=auto_round, bot_task=image,
calib_cache_dir=/models/changwa1/hunyuan-calib-cache, device=0.
Per-layer calibration record: forwards = 128 for each of model.layers.0…31, plus 16
calibration_schedules (seeds 42–57) — i.e. the tuner saw real 8-step diffusion trajectories, unlike the
--model_free sibling which saw none.
Self-check that the export is loadable before serving it:
python /path/to/quantization_analysis/scripts/check_autoround_experts_key.py <this model dir>
# exit 0 = experts resolve to 4-bit; exit 1 = 4-bit override lost, metadata needs normalizing
Inference environment
| Component | Version |
|---|---|
| vLLM | 0.29.0 |
| vLLM-Omni | latest main (validated at 1c7476ec19899f1e61838e23e9fad3505403b9b9, 0.29.0rc2.dev161+g1c7476ec1, editable install) |
| PyTorch | 2.13.0+cu132 |
| FlashInfer | 0.6.18 |
| Python | 3.12.3 |
| GPU | NVIDIA H200 141 GB — AR-only 1×, DiT-only 2×, AR+DiT 2× |
python -m venv ~/.venv-omini-latest && source ~/.venv-omini-latest/bin/activate
pip install vllm==0.29.0 && pip install -e /path/to/vllm-omni
python -c "import vllm, vllm_omni, os; print(vllm.__version__, os.path.dirname(vllm_omni.__file__))"
# host-specific, drop if unneeded here
export NCCL_NVLS_ENABLE=0
export VLLM_USE_FLASHINFER_SAMPLER=0
cd /tmp # cwd must be neutral: children re-import vllm_omni from cwd
A full reproducible environment builder (constraints file + patch set + verifier) lives in
quantization_analysis/envkit/.
Validated run modes
| Mode | GPUs | Result | Evidence from logs |
|---|---|---|---|
| DiT-only (TP2) | 2 (cards 0,1) | ✅ image | Stage 0 … mapping: 0->0, 1->1; Model loading took 25.3899 GiB per card; image_pixels=1048576; denoise_step_latency_ms=1845.98; peak_memory_mb=39482; Saved generated image to …/run_mxfp4ar_tuning_dit.png |
| AR+DiT (AR TP1 + DiT TP1) | 2 (cards 0,1) | ✅ image | 2 stages (0->0, 1->1); DiT Model loading took 46.6297 GiB, AR 46.69 GiB; AR ratio_idx=19, target size=1216x832; peak_memory_mb=61190; Saved generated image to …/run_mxfp4ar_tuning_full.png |
| AR-only (TP1) | 1 (card 2) | ✅ text | Model loading took 46.69 GiB memory and 55.15 seconds; Using MarlinExperts (weight-only FP4) for AutoRound MXFP4 MoE; [run_ar] EXIT=0; output text It's a typical Photography (27 bytes) |
⚠️ AR-only text is evidence that the AR path loads and generates, nothing more. Across this workspace the AR-only captures range from 7 to 1198 bytes and
x_to_text.pydoes not echo the request sampling parameters, so token-level agreement with other builds cannot be claimed. In the full pipeline this build's AR emitted 8 tokens →ratio_idx=19→target size=1216x832, whereas BF16/MXFP8 runs emit 6 tokens /ratio_idx=14and the 0.15.0 MXFP4 builds were seen at 10 tokens /ratio_idx=14and at 295 tokens /ratio_idx=16on repeated runs of the same command. The AR branch is the dominant source of full-pipeline image divergence here — see Fidelity before attributing any of it to tuning.
All output PNGs are 1024×1024 in every mode (17/17 full-pipeline logs in this workspace checked
against their own saved files). The target size=… line in the AR→DiT handoff is the AR's predicted
aspect — ar2diffusion computes it from reso_group[ratio_idx] and logs it
(stage_input_processors/hunyuan_image3.py:139-191), but the DiT takes its canvas from
sampling.height or 1024 (pipeline_hunyuan_image3.py:1912), which the offline example already fills
with --height/--width defaults of 1024 (text_to_image.py:125-126). ⇒ ratio_idx changes the AR text
and the log line, not the saved canvas, so cross-image PSNR/SSIM here involves no mismatched-canvas
resampling, and any full-pipeline divergence must come from the conditioning text, not the aspect.
# ── DiT-only, 2 GPUs ─────────────────────────────────────────────
export CUDA_VISIBLE_DEVICES=0,1
cd /tmp
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
--model /path/to/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound \
--deploy-config ./hunyuan_image3_dit_tp2.yaml \
--prompt "A cute cat" \
--num-inference-steps 8 \
--guidance-scale 5.0 \
--seed 42 \
--init-timeout 1800 --stage-init-timeout 1500 \
--output ./dit_only.png
# ── AR+DiT, 2 GPUs ───────────────────────────────────────────────
export CUDA_VISIBLE_DEVICES=0,1
cd /tmp
python /path/to/vllm-omni/examples/offline_inference/text_to_image/text_to_image.py \
--model /path/to/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound \
--deploy-config ./hunyuan_image_3_moe_2gpu_tp1.yaml \
--prompt "A cute cat" \
--num-inference-steps 8 --guidance-scale 5.0 --seed 42 \
--init-timeout 1800 --stage-init-timeout 1500 \
--output ./ar_dit.png
# ── AR-only, 1 GPU (different example script; it does NOT accept --init-timeout) ──
export CUDA_VISIBLE_DEVICES=2
cd /tmp
python /path/to/vllm-omni/examples/offline_inference/x_to_text/x_to_text.py \
--model /path/to/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound \
--deploy-config ./hunyuan_image3_ar_tp1.yaml \
--trust-remote-code \
--prompt "A cute cat" --max-tokens 256 \
--output ./ar_only.txt
--deploy-config files and the wrapper scripts (run_dit.sh / run_ar.sh) with fully expanded commands,
timeout rationale and per-mode log expectations are in
quantization_analysis/latest/runpack/README.md.
Parameter notes — --num-inference-steps 8 (distilled checkpoint: do not use 50);
--guidance-scale is a real input (cfg_distilled=true; vLLM-Omni feeds 1000 × guidance_scale as the
guidance embedding; the images above used 5.0, upstream's Distil e2e reference uses 2.5);
--prompt reaches only the AR stage.
Fidelity
Weight side — tuning moves the weights away from the base, on purpose
Dequantizing this build and -MXFP4-ar (same E2M1/MXFP4 nibble + E8M0 group-32 and E4M3 + E8M0 group-32
codewords) and comparing against the local BF16 base:
| Module | -MXFP4-ar (model_free / RTN) |
this build (iters=50 tuned) |
|---|---|---|
self_attn.qkv_proj L0 |
0.0267 | 0.0572 |
mlp.shared_mlp.gate_and_up_proj L0 |
0.0267 | 0.0476 |
self_attn.o_proj L31 |
0.0267 | 0.0640 |
mlp.shared_mlp.down_proj L31 |
0.0267 | 0.0465 |
mlp.experts.3.gate_and_up_proj L0 |
0.1123 | 0.1285 |
mlp.experts.40.down_proj L0 |
0.1125 | 0.1260 |
mlp.experts.11.gate_and_up_proj L15 |
0.1113 | 0.1239 |
mlp.experts.63.down_proj L31 |
0.1124 | 0.1267 |
| mean MXFP8 dense | 0.0267 | 0.0538 (2.01×) |
| mean MXFP4 experts | 0.1121 | 0.1263 (+12.7 %) |
mlp.gate.wg router (L0, L31) |
0.0000 | 0.0000 (max|Δ| = 0) |
Read this correctly:
- The
-MXFP4-arcolumn reproduces the format floor exactly (0.0267 / ≈0.1122 = ideal per-group E8M0 RTN), which is what validates the measurement harness. - The tuned weights deviate more from the base, not less. That is expected:
iters=50optimizes layer output error on 128 real diffusion forwards per block, so it deliberately moves weights (and rounds differently) instead of minimizing weight-space distance. Weight rel-L2 vs the base is no longer a valid proxy for quality on a tuned checkpoint — do not compare a tuned build against an RTN build with it. - Router is bit-exact with the base in both builds ⇒ the router was never tuned, only re-serialized as FP32.
Image side — ⚠️ cross-session numbers, read the calibration row first
Pixel-level agreement in this workspace is bounded by session, not by quantization. Re-shooting the same build with the same command in a different session gives:
| Situation | PSNR vs the old capture | SSIM@1/8 |
|---|---|---|
| same build, same session re-run | 41.35 – 44.51 dB (BF16: 74.2 dB, `max | Δ |
| same build, different session | 12.2 – 12.5 dB | 0.223 – 0.233 |
⇒ Any pair measured in different sessions sits at ≈12 dB / ≈0.23 by drift alone. The DiT-only table below mixes 09-17, 09-18 and 09-22 captures, so treat the two quantized rows as "≈ drift level", not as a ranking.
DiT-only, TP2 on cards 0,1 for all cells (reference = local BF16, 09-17 22:01):
| Candidate | captured | PSNR | SSIM@1 | SSIM@1/8 | mean|Δ| | std |
|---|---|---|---|---|---|---|
MXFP8-ar (-MXFP8, auto-round/INC) |
09-18 00:38 | 31.79 | 0.9811 | 0.9842 | 3.56 | 53.74 |
MXFP8-ct (-MXFP8-ct, compressed-tensors) |
09-18 | 31.59 | 0.9788 | 0.9823 | 3.62 | 53.62 |
MXFP4-ar (-MXFP4-ar, model_free) |
09-22 | 13.00 | 0.6720 | 0.2704 | 44.20 | 49.09 |
| this build (MXFP4 + tuning) | 09-22 07:15 | 12.88 | 0.6640 | 0.2568 | 45.31 | 45.70 |
Figure A — DiT-only: BF16 / MXFP8-ar / this build.
Direct pairings of this build against other captures (DiT-only): vs MXFP8-ar 12.69 dB / SSIM@1/8 0.2561; vs MXFP4-ar 22.06 dB / SSIM@1/8 0.7904 — i.e. the tuned MXFP4 build is much closer to the untuned MXFP4 build than either is to BF16 or to MXFP8, which is the signature you expect when the dominant difference is the experts' 4-bit codebook rather than the 50 tuning iterations.
AR+DiT (reference = BF16 on 4 GPUs, TP2+TP2 — cross-topology, and the AR stage re-draws its CoT every run, so this mode is the least reproducible of the three):
| Candidate | topology | PSNR | SSIM@1 | SSIM@1/8 | std |
|---|---|---|---|---|---|
| MXFP8-ar | 2 GPUs TP1+TP1 | 26.85 | 0.9227 | 0.8715 | 49.88 |
| MXFP4-ar | 2 GPUs TP1+TP1 | 12.89 | 0.6934 | 0.1708 | 46.44 |
| this build | 2 GPUs TP1+TP1 | 13.54 | 0.6423 | 0.1766 | 43.02 |
Figure B — AR+DiT: BF16 / MXFP8-ar / this build.
Individual cells, uncropped:
All four cells in the bottom row show a coherent, un-degraded cute-cat image; the MXFP4 cells differ from
BF16/MXFP8 in composition (SSIM falls as the scale gets coarser ⇒ structure, not texture), which is
the same qualitative finding as the -MXFP4-ar build. agent_smoke_dit_only.png is this agent's own
independent DiT-only reproduction of the tuned checkpoint (std 46.06), kept for cross-checking.
Not measured
No task-level benchmark (GenEval / DPG-Bench / CVTG-2K / DrawBench / WISE) has been run for this
checkpoint. "Composition diverges from BF16" is measured; "quality is better or worse than -MXFP4-ar
because of tuning" is not established — and cannot be from one prompt.
Reproduce every number in this card
PY=/home/kaokaolv/.venv-omini-latest/bin/python
QA=/home/kaokaolv/quantization_analysis
TUNE=/home/kaokaolv/HunyuanImage-3-Instruct-Distil-MXFP4-Mixed-Tuning-AutoRound
AR=/home/kaokaolv/HunyuanImage-3.0-Instruct-Distil-MXFP4-ar
# 0) does the checkpoint resolve to 4-bit experts on the installed vLLM? (exit 0 = yes)
$PY $QA/scripts/check_autoround_experts_key.py $TUNE
# 1) coverage census ("Coverage vs. the sibling -MXFP4-ar build")
$PY $QA/scripts/census_family_compare.py $AR $TUNE
# 2) weight-space error ("Weight side")
$PY $QA/scripts/dequant_error_vs_base.py $AR $TUNE /home/kaokaolv/models/HunyuanImage-3.0-Instruct-Distil
# 3) image metrics (any pair; --scales 1,8 = fine + coarse SSIM)
$PY $QA/bench/eval_scripts/compare_images.py \
--ref $QA/latest/runpack/out/bf16_dit_tp2.png \
--cands $QA/latest/runpack/out/pr_ar_dit.png $TUNE/images/user_dit_only.png --scales 1,8
# 4) the two montages (Figure A / Figure B)
$PY $QA/envkit/make_card_grid.py --out $QA/latest/runpack/out/grid_tuning_ditonly.png \
--ref "① BF16 base|DiT-only TP2|captured 09-17 22:01"=$QA/latest/runpack/out/bf16_dit_tp2.png \
--cell "② MXFP8-ar (auto-round/INC)|DiT-only TP2|captured 09-18 00:38"=$QA/latest/runpack/out/pr_ar_dit.png \
--cell "③ MXFP4-mixed +tuning (0.16.0)|DiT-only TP2|captured 09-22 07:15"=$TUNE/images/user_dit_only.png \
--scales 1,8 --width 460
Known limitations
- Memory-only on Hopper — experts run W4A16 through Marlin; the declared
act_bits: 4is inert. - Cross-session pixel drift (~12 dB) is larger than the tuning effect being measured here, so the
image tables are illustrative, not quantitative. Same-session re-shoots are the only way to rank
iters=50againstiters=0. - The full-pipeline images are not comparable at the pixel level at all — the AR stage regenerates its CoT/ratio tokens per run and those tokens are the DiT conditioning.
- Composition diverges from BF16/MXFP8 (SSIM@1/8 ≈ 0.26 in DiT-only) — the MXFP4 experts, not the tuning, are the dominant cause.
- The ViT stays BF16 (
vision_model.encoder.*.mlp.fc2input dim 4304 is not divisible by the 32-element MX group) ⇒ not a whole-model MXFP4 build. - Router tensors are FP32 here vs BF16 in the 0.15.0 builds — harmless for vLLM, but it breaks naive byte-hash comparisons across builds (a real trap: an md5-based "different base model" conclusion drawn from this was wrong; the values are equal).
block_name_to_quantize = [model.layers.0 … model.layers.31]scopes quantization to the 32 shared backbone layers; AR and DiT load the same tensors through different stacks.images/holds verification snapshots only; it is not read when loading weights.
Contents
config.json quant_method=auto-round, extra_config (4164 entries) — stock export,
nothing edited; loads as-is
quantization_config.json flat secondary discovery carrier
calibration_recipe.json AutoRound's own dump of the tuning run (the source of the
command above)
model-000NN-of-00011.safetensors 52.61 GB total, 9393 tensors
model.safetensors.index.json
images/
user_ar_dit.png AR+DiT, 2 GPUs (hero)
user_dit_only.png DiT-only TP2, this build
agent_smoke_dit_only.png independent DiT-only re-run by the validating agent
ar_only_output.txt AR-only output text (`It's a typical Photography`, 27 B)
bf16_dit_only_reference.png BF16 base, DiT-only TP2
mxfp8_ar_dit_only.png -MXFP8 (auto_round) DiT-only TP2
bf16_ar_dit_4gpu.png BF16 base, AR+DiT on 4 GPUs
mxfp8_ar_ar_dit.png -MXFP8 (auto_round), AR+DiT on 2 GPUs
grid_tuning_ditonly.png Figure A (BF16 / MXFP8-ar / this build, DiT-only)
grid_tuning_ardit.png Figure B (same three, AR+DiT)
*.py / tokenizer* / utils/ inherited from the base model
- Downloads last month
- 14







