Qwen3.8-27B Intel Arc Tuned Weights (Native Systolic DPAS GGUF)

Hardware-Accelerated Weights for Intel Arc GPUs (Battlemage Xe2 / Alchemist Xe / oneAPI SYCL)
Engineered by Greatjedi & Gemini (DeepMind Antigravity Team)


Why These Weights Exist: Solving the Intel Arc Performance Cliff

If you own an Intel Arc GPU (such as the Intel Arc Pro B60 24GB, Arc A770 16GB, or Battlemage graphics cards) and have attempted to run popular 27B quantization variants (such as IQ2_S, IQ3_XXS, or Unsloth's "Dynamic" UD-Q4_K_S), you likely noticed your generation speed crawling at an agonizing 8.1 to 12.2 tokens/second, despite the GPU having more than enough raw compute and 512 GB/s memory bandwidth.

The Silicon Root Cause: Non-Linear Lattice vs. Systolic DPAS

  1. Intel Arc Matrix Engines (XMX) utilize DPAS (Dot Product Accumulate Systolic) hardware. DPAS requires dense, linearly indexed integers (INT2, INT4, INT8) to feed its systolic pipeline at thousands of operations per clock cycle.
  2. I-Quants (IQ2_S, IQ3_XXS, IQ4_XS, IQ3_S) use non-linear 8-dimensional vector lattice codebooks ($E_8$). They require runtime table lookups per weight element.
  3. In llama.cpp's SYCL backend, there are zero kernel implementations for I-Quants. Whenever an IQ tensor is evaluated, SYCL drops out of hardware systolic GEMM into slow CPU/host dequantization loops.
  4. The Hidden Catch in Unsloth "UD" Models: When we reverse-engineered Unsloth's popular Qwen3.8-27B-UD-Q4_K_S.gguf, we discovered it secretly contains 201 I-Quant tensors (IQ4_XS, IQ3_S, IQ3_XXS), causing it to suffer from the exact same dequantization stall on Intel Arc!

The Solution: Selective Software Transquantization

Using a custom selective software parser in llama.cpp, we dynamically intercepted and transquantized every unsupported non-linear I-Quant into native Q4_K super-blocks (with independent 6-bit scales and mins per 32 weights), while preserving rotary embeddings (RoPE), layernorms, linear SSM states, high-precision attention projections, and โ€” in the MTP-included checkpoints โ€” the embedded NextN prediction head bit-for-bit.


Benchmark Results on Intel Arc Pro B60 (24GB VRAM)

All benchmarks measured on device level_zero:0 (UR_LOADER_USE_LEVEL_ZERO_V2=0) with -ngl 99 -fa on:

Model Variant File Size Quant Composition Pure Autoregressive Decode Prompt Prefill Speculative Decode (MTP / DFlash2)
ISTA-DASLab IQ3_XXS (Stock) 9.39 GiB 355 IQ tensors (81% of model) 8.10 tok/s 110.05 tok/s N/A (Stalled)
Qwen3.8-27B-Ridge (Stock) 11.72 GiB 192 IQ tensors 12.22 tok/s 133.45 tok/s ~13.2 tok/s
Unsloth UD-Q4_K_S (Stock) 14.29 GiB 201 IQ tensors 12.97 tok/s 162.75 tok/s ~18.9 tok/s (Verify bottlenecked)
Ridge Intel Arc Tuned Q4_K (Ours) 15.29 GiB 0 IQ (336 Q4_K, 51 Q5_K, 23 Q6_K) 22.23 tok/s 245.54 tok/s 41.28 tok/s (Embedded MTP, 93.4% accept)
DAS Lab Intel Arc Tuned Q4_K (Ours) 13.72 GiB 0 IQ (374 Q4_K, 28 Q2_K, 96 BF16) 22.36 tok/s 243.61 tok/s 36.03 tok/s (DFlash2, 93.5% accept)
DAS Lab IQ3_S Intel Arc Tuned MTP (Ours, NEW) 14.31 GiB 0 IQ (350 IQ tensors โ†’ Q4_K) 21.42 tok/s โ€” โ€  41.88 tok/s (Embedded MTP, 92.9% accept)

โ€  Prefill for this variant was not separately benchmarked; it shares the same architecture-lane and transquantized tensor layout as the other 0-IQ variants (~240+ tok/s expected).

  • Pure Decode: Leaped from 8.10 tok/s $\to$ 22.36 tok/s (+176% / ~2.8ร— faster).
  • Speculative Decode: Smashes past 41 tokens/second using the embedded NextN prediction head โ€” now available on both the Ridge and the DAS Lab IQ3_S tuned bases, with zero external drafter VRAM.

New: Embedded MTP Now Shipping in the DAS Lab Line (Measured Sweep, 2026-09-17)

The latest tuned checkpoint, Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf, is the DAS Lab GSQ-RCO IQ3_S-MTP checkpoint transquantized to pure Q4_K โ€” the MTP head (blk.64.nextn.*, 41 MiB) rides along untouched. Earlier DAS Lab checkpoints (IQ3_XXS) shipped with their NextN heads stripped, which is why the previous DAS Lab release needed an external DFlash2 drafter. This one is self-contained.

Speculative topology sweep on the new checkpoint (real prompts: prime-factorization reasoning 1024 tok, codegen, JSON extraction):

Configuration Speculative Topology Reasoning Code Gen JSON Extract Draft Accept %
Config 0 Pure Autoregressive (OFF) 21.42 tok/s 20.29 tok/s 16.54 tok/s โ€”
Config 1 โ˜… Embedded MTP (ngram-mod,draft-mtp, p-min=0.5, n-max=3) 41.88 tok/s 38.23 tok/s 26.38 tok/s 92.9%
Config 2 External DFlash2 (DFlash2-Q4_K_M) 39.68 tok/s 39.19 tok/s 27.11 tok/s 94.7%
Config 3 External DSpark (DSpark-Q8_0) โŒ SYCL kernel crash (common.hpp:145) โ€” โ€” โ€”

Takeaways, all measured (not estimated):

  • Embedded MTP wins: 41.88 tok/s crushes the external DFlash2 setup (39.68 tok/s) with zero drafter VRAM โ€” and beats the Ridge embedded-MTP release (41.28) by a whisker while being ~1 GiB smaller.
  • DSpark is unusable on this stack: it hard-crashes the SYCL backend at first kernel sync. Drop the Q8_0 DSpark draft model entirely.
  • Small draft windows win: n-max 2โ€“3 with a strict acceptance floor (p-min โ‰ฅ 0.5) beat both wider windows and looser floors, at shallow and deep context.

Deep-Context Truth: Attention Cost Doesn't Care About Your Spec Method

We swept decode speed at ~150k-token depth on the same hardware (the production lane regularly serves ~148k contexts, so this is real traffic, not a hypothetical):

Config @ ~150k depth Decode Acceptance
DFlash2 n-max=3, p-min=0.7 (deployed production config) 7.91 tok/s 84.7%
Native MTP n-max=2, p-min=0.5 + SYCL graph 7.21 tok/s 78.4%
DFlash2 n-max=2, p-min=0.5 6.88 tok/s 67.0%

Every topologies loses ~59% of its shallow speed by 150k tokens. That is attention-cost-of-depth โ€” the KV cache grows linearly, and no draft method or graph mode fixed it in our sweeps. If you serve extremely deep contexts, budget for ~7โ€“8 tok/s, not the shallow number.


Long-Context Needle Retrieval Restoration

Many users observed that extreme 2-bit I-Quants suffer from catastrophic Needle-In-A-Haystack (NIAH) failures past 32k context:

  • Transformer Feed-Forward Networks (FFNs) function mechanistically as associative key-value factual memory banks (Geva et al., 2021).
  • Coarse 8D lattice codebooks truncate high-magnitude outlier weights. Across deep contexts (50kโ€“200k+ tokens), quantization noise accumulates across all 64 layers, destroying the signal-to-noise ratio and creating "blind spots" where the model fails to retrieve facts.
  • Q4_K Super-Blocks Restore Retrieval: By mapping weights into linear super-blocks with independent 6-bit scales and minimum offsets per 32 weights, dynamic range is fully preserved, restoring 100% green needle retrieval across long context windows.

The Intel Arc Pro B60 Sweet Spot: 200k Production Context on a Single 24GB Card

If you opted for the Intel Arc Pro B60 (24GB VRAM) instead of spending thousands more on the B70 or enterprise workstation cards, you might have worried about being forced to truncate your context window to 32k or 64k, or suffer massive speed penalties from CPU RAM offloading.

Here is the big win: You can run an enormous 200,000-token context window completely resident on a single Intel Arc Pro B60 24GB GPU with zero CPU offloading, BF16 multimodal vision, and maximum prefill throughput (-ub 1024, 448 tok/s prefill) โ€” while keeping decode generation blazing fast (41+ tokens/second with embedded MTP)!

The KV Cache Arithmetic: Why Long Context Fits in 24GB

Standard transformers require tens of gigabytes of VRAM for 200k+ contexts because every layer allocates quadratic key-value caches. Qwen 3.8 27B changes this math entirely through Hybrid Linear Attention (full_attention_interval = 4):

  • 48 of the 64 layers are linear recurrent layers with a fixed, constant $O(1)$ state regardless of context depth.
  • Only 16 layers allocate a dynamic token-by-token KV cache.
  • Using quantized q4_0 KV cache (--cache-type-k q4_0 --cache-type-v q4_0), each token across all 16 attention layers costs only 18,432 bytes (~18.0 KiB / token): $$\text{KV Footprint} = 16 \text{ layers} \times 2 \times 4 \text{ heads} \times 256 \text{ dim} \times 0.5625 \text{ bytes/element} = 18,432\text{ bytes/token}$$

Measured Live Hardware VRAM Accounting on Intel Arc Pro B60 (24,480 MiB) โ€” Ridge Tuned (largest variant):

Component 200,000 Tokens (q4_0 KV) Notes
Model Weights (Transquant Q4_K) 15.29 GiB (16,046 MiB) 64 layers + norms (-ngl 99)
Vision Projector (mmproj BF16) 0.87 GiB (889 MiB) Unsloth BF16 vision head (-ngld 99)
Embedded MTP Head & Recurrent State 0.07 GiB (71 MiB) 4 NextN tensors + 48 recurrent states
Target + MTP KV Cache (q4_0) 3.65 GiB (3,739 MiB) 16 trunk layers + 1 MTP draft layer
Compute Buffers (-ub 1024) & Workspaces 2.37 GiB (2,425 MiB) Full 448 tok/s XMX systolic batch prefill
Total Active VRAM Usage 22.54 GiB (23,079 MiB) Verified live via Vulkan / Level Zero
Remaining Free VRAM Headroom ~1.40 GB (1,401 MiB) Safe, rock-solid headroom for production!

The new IQ3_S-MTP tuned base is ~1 GiB lighter (14.31 GiB), buying ~1 GiB of extra headroom on the same math โ€” you can bump KV precision to q4_1 or push context further.

Production Recommendation: -c 200000 is the tested, verified sweet spot for production. It preserves ~1.4 GB of healthy headroom for high-resolution vision tokens and desktop tasks while running -ub 1024 for maximum prefill speed (448 t/s).


Available Models in this Repository

1. Qwen3.8-27B-Ridge-Intel-Arc-Tuned-Q4_K.gguf (15.29 GiB) โ€” Max Throughput, Long Context

  • Embedded MTP Included: Contains 4 embedded NextN prediction tensors (blk.64.nextn.*, 41 MiB).
  • Zero External Drafter VRAM Overhead: Runs speculative decoding without needing an external -md file, saving 1.1โ€“1.3 GiB of VRAM.
  • Speed: Delivers 41.28 tok/s at 93.4% acceptance.
  • Context: Fully verified with -c 200000 (leaves ~1.4 GB headroom with BF16 vision and full 448 t/s prefill batching on a 24GB B60).

2. Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf (14.31 GiB) โ€” DAS Lab + Embedded MTP (Q4_K) โญ

  • Based on ISTA-DASLab's GSQ-RCO IQ3_S-MTP checkpoint transquantized to pure Q4_K โ€” 0 IQ tensors remain, full XMX DPAS acceleration.
  • Embedded MTP: 41.88 tok/s at 92.9% acceptance โ€” top measured result in the repo, zero external drafter VRAM.

3. Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-Q4_K.gguf (13.72 GiB) โ€” Ultra-Compact Base

  • Based on ISTA-DASLab's optimized GSQ-RCO weights (374 Q4_K, 28 Q2_K, 0 IQ).
  • Pure decode at 22.36 tok/s, pairs with external DFlash2 drafter for 36.03 tok/s average.

4. Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.7BPW.gguf (11.88 GiB) โ€” Profile A (Mid-Compression)

  • Quantized directly from virgin BF16 base weights using ISTA-DASLab's RCO sensitivity mapping (135 Q4_K, 223 Q3_K, 44 Q2_K, 8 Q6_K, 0 IQ).
  • VRAM Footprint: Leaves ~9.5 GB free VRAM on 24GB GPUs; leaves ~4.2 GB free on 16GB cards.
  • Embedded MTP Included: 89.0% draft acceptance.
  • โš ๏ธ See Architecture Advisory below regarding Intel Arc Xe2 decode speed on sub-4 BPW weights.

5. Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.2BPW.gguf (10.45 GiB) โ€” Profile B (Ultra-Compact for 12โ€“16GB Cards)

  • Quantized directly from virgin BF16 base weights using ISTA-DASLab's high-compression RCO mapping (52 Q4_K, 181 Q3_K, 169 Q2_K, 8 Q6_K, 0 IQ).
  • VRAM Footprint: Fits inside 12GB and 16GB GPUs (Intel Arc Pro B50, Arc A770 16GB, RTX 4070 Ti) with room for 32kโ€“128k context without CPU offloading.
  • Embedded MTP Included: 90.5% draft acceptance.
  • โš ๏ธ See Architecture Advisory below regarding Intel Arc Xe2 decode speed on sub-4 BPW weights.

โš ๏ธ CRITICAL ARCHITECTURAL ADVISORY: Sub-4 BPW Models (3.7BPW & 3.2BPW) on Intel Arc

Performance & Compatibility Notice for Intel Arc Users: The sub-4 BPW models (Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.7BPW.gguf and Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.2BPW.gguf) are not optimal for Intel Arc Pro architectures and DO NOT hit the stipulated 44+ tok/s speculative decode speeds achieved by our native Q4_K checkpoints.

The Silicon Root Cause:

  • Native INT4/INT8 Alignment: Intel Arc Xe2 XMX matrix engines use DPAS (Dot Product Accumulate Systolic) hardware that operates natively on packed 4-bit (int4) and 8-bit (int8) matrix dot products (dpas.8x8).
  • In Q4_K, weights stream directly into DPAS execution units with zero unpacking friction, enabling 38โ€“40 tok/s sustained decode and 44โ€“48 tok/s peak speculative bursts.
  • In sub-4 BPW formats (Q3_K and Q2_K), weights are packed into irregular, non-power-of-two bit widths. In GGML SYCL compute kernels, Execution Units (EUs) must perform heavy bitwise shifting, masking, and scale lookups per super-block before loading data into DPAS.
  • The Compute vs. Bandwidth Bottleneck: On cards with ample memory bandwidth (like the Arc Pro B60 with 456 GB/s), this ALU unpack overhead bottlenecks the pipeline, capping decode speeds at ~24โ€“28 tok/s despite transferring fewer bytes from VRAM.

When to use Sub-4 BPW: Use 3.7BPW or 3.2BPW strictly if you are VRAM-constrained (e.g. 12GBโ€“16GB cards like Arc Pro B50 or Arc A770 16GB) where fitting 32kโ€“128k context into VRAM without host RAM spillover takes priority over peak decode speed. For 24GB cards (Arc Pro B60, RTX 3090/4090), always use the native Q4_K models for maximum throughput.


Quickstart for Intel Arc Users

Critical Environment Variables for Intel Arc (Battlemage Xe2 / oneAPI 2026.x)

# Sourcing oneAPI
source /opt/intel/oneapi/setvars.sh --force

# CRITICAL FOR BATTLEMAGE: Force Level Zero V1 UR adapter (prevents error 44)
export UR_LOADER_USE_LEVEL_ZERO_V2=0
export ONEAPI_DEVICE_SELECTOR=level_zero:0
export ZE_AFFINITY_MASK=0
export ZES_ENABLE_SYSMAN=1

Running with 200k Context, Embedded MTP & Vision (41+ tok/s)

./llama-server \
  -m Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf \
  --mmproj mmproj-Qwen3.8-27B-unsloth-BF16.gguf \
  --spec-type ngram-mod,draft-mtp \
  --spec-ngram-mod-n-match 24 \
  --spec-ngram-mod-n-min 1 \
  --spec-ngram-mod-n-max 3 \
  --spec-draft-n-max 3 \
  --spec-draft-p-min 0.5 \
  -c 200000 \
  --cache-type-k q4_0 --cache-type-v q4_0 \
  -ngl 99 -ngld 99 -fa on \
  -ub 1024 -b 1024 -t 4 -tb 6 --parallel 1 \
  --jinja --chat-template-file chat_templates/qwen-sharp.jinja --reasoning-preserve \
  --host 127.0.0.1 --port 8183

Optional: SYCL Graph mode (GGML_SYCL_ENABLE_GRAPH=1) gives +19% with embedded MTP (14.68 โ†’ 17.51 tok/s measured shallow). โš ๏ธ Do not combine it with external draft models (DFlash2/DSpark) โ€” it crashes with Graph nodes cannot depend on events from outside the graph. Pure decode gains nothing from it. Safe only with embedded MTP (draft-mtp) or no speculation.

Chat Template Included: chat_templates/qwen-sharp.jinja (by froggeric) is included directly in this repository. Pass it via --chat-template-file to enable dynamic <|think_on|> / <|think_off|> reasoning toggles, robust agent tool parsing, and avoid empty-think context degradation in llama.cpp.


Credits, Attribution & Upstream Acknowledgments

This repository provides hardware-adapted conversions specifically optimized for Intel Arc GPUs. We give full credit, attribution, and gratitude to the original creators whose breakthrough models made this work possible:

1. Base Model Architecture & Foundation Weights

2. Ridge Quantization & Multi-Token Prediction Head

  • empero-ai: Creators of the Qwen3.8-27B-Ridge quantization and embedded multi-token prediction configuration.
  • Original Repository: empero-ai/Qwen3.8-27B-Ridge-GGUF
  • License: Apache 2.0
  • Notice of Modification (Apache 2.0 ยง4b): Transquantized 192 unsupported non-linear FFN lattice tensors (IQ2_S / IQ3_S) into hardware-accelerated linear Q4_K super-blocks for Intel Arc systolic DPAS execution.

3. GSQ & Randomized Coordinate Optimization (RCO)

  • ISTA-DASLab: Creators of the GSQ-RCO quantization method and checkpoints (both the IQ3_XXS and IQ3_S-MTP lines).
  • Original Repository: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
  • Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Vector Lattices
  • License: Apache 2.0
  • Notice of Modification (Apache 2.0 ยง4b): Transquantized unsupported non-linear lattice tensors into linear Q4_K super-blocks โ€” 355 tensors on the IQ3_XXS base (IQ1_S, IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_XS), and 350 tensors on the IQ3_S-MTP base (IQ2_S, IQ3_S, IQ4_XS). The embedded NextN prediction head (blk.64.nextn.*) is preserved bit-for-bit on the MTP line.

4. Chat Template & Reasoning Protocol

  • froggeric: Creator of the optimized Jinja2 chat templates for Qwen models (qwen-sharp / qwen3.8-froggeric).
  • Original Repository: froggeric/Qwen-Fixed-Chat-Templates
  • Features: C++ engine compatibility in llama.cpp, dynamic reasoning toggles (<|think_on|>, <|think_off|>, <|think_xhigh|>), mitigation of empty-think KV cache bloat, and resilient tool-call serialization.

5. Inference Engine & Toolchain

  • Georgi Gerganov & the llama.cpp community: For llama.cpp and the open-source GGML SYCL backend.
  • Intel oneAPI Team: For the DPC++ / SYCL compiler and Level Zero runtime drivers.

Methodology Notes (how these numbers were measured)

  • Hardware: Intel Arc Pro B60 24GB (Battlemage Xe2, 24,480 MiB), Ubuntu 24.04, oneAPI 2026.x, llama.cpp build-sycl (SYCL backend).
  • llama-bench -p 128 -n 32 -r 1 -ngl 99 -fa on for the pure AR/prefill table; server-side /completion timings with --metrics for the speculative sweeps (1,024-token reasoning / 400-token codegen / 69-token JSON extraction prompts).
  • Depth tests: synthetic ~150k-token prompt with ignore_eos:true, predicted_n sanity-checked against n_predict (a silent false-negative failure mode was caught and gated in our harness).
  • The transquantizer runs inside llama.cpp's llama-quantize (branch sycl-iquants-mmq): --allow-requantize --tensor-type "^iq=q4_K" <in> <out> COPY 12.
Downloads last month
4,567
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(5)
this model

Space using Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF 1

Paper for Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF