Instructions to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S # Run inference directly in the terminal: llama cli -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S # Run inference directly in the terminal: llama cli -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S # Run inference directly in the terminal: ./llama-cli -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Use Docker
docker model run hf.co/Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
- LM Studio
- Jan
- vLLM
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
- Ollama
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with Ollama:
ollama run hf.co/Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
- Unsloth Desktop
- Pi
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with Docker Model Runner:
docker model run hf.co/Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
- Lemonade
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Run and chat with the model
lemonade run user.Qwen3.8-27B-Intel-Arc-Tuned-GGUF-IQ3_S
List all available models
lemonade list
- Hermes Agent
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Greatjedi/Qwen3.8-27B-Intel-Arc-Tuned-GGUF:IQ3_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Qwen3.8-27B Intel Arc Tuned Weights (Native Systolic DPAS GGUF)
- Why These Weights Exist: Solving the Intel Arc Performance Cliff
- Benchmark Results on Intel Arc Pro B60 (24GB VRAM)
- New: Embedded MTP Now Shipping in the DAS Lab Line (Measured Sweep, 2026-09-17)
- Deep-Context Truth: Attention Cost Doesn't Care About Your Spec Method
- Long-Context Needle Retrieval Restoration
- The Intel Arc Pro B60 Sweet Spot: 200k Production Context on a Single 24GB Card
- Available Models in this Repository
- 1.
Qwen3.8-27B-Ridge-Intel-Arc-Tuned-Q4_K.gguf(15.29 GiB) โ Max Throughput, Long Context - 2.
Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf(14.31 GiB) โ DAS Lab + Embedded MTP (Q4_K) โญ - 3.
Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-Q4_K.gguf(13.72 GiB) โ Ultra-Compact Base - 4.
Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.7BPW.gguf(11.88 GiB) โ Profile A (Mid-Compression) - 5.
Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.2BPW.gguf(10.45 GiB) โ Profile B (Ultra-Compact for 12โ16GB Cards)
- 1.
- โ ๏ธ CRITICAL ARCHITECTURAL ADVISORY: Sub-4 BPW Models (
3.7BPW&3.2BPW) on Intel Arc - Quickstart for Intel Arc Users
- Credits, Attribution & Upstream Acknowledgments
- Methodology Notes (how these numbers were measured)
Qwen3.8-27B Intel Arc Tuned Weights (Native Systolic DPAS GGUF)
Hardware-Accelerated Weights for Intel Arc GPUs (Battlemage Xe2 / Alchemist Xe / oneAPI SYCL)
Engineered by Greatjedi & Gemini (DeepMind Antigravity Team)
Why These Weights Exist: Solving the Intel Arc Performance Cliff
If you own an Intel Arc GPU (such as the Intel Arc Pro B60 24GB, Arc A770 16GB, or Battlemage graphics cards) and have attempted to run popular 27B quantization variants (such as IQ2_S, IQ3_XXS, or Unsloth's "Dynamic" UD-Q4_K_S), you likely noticed your generation speed crawling at an agonizing 8.1 to 12.2 tokens/second, despite the GPU having more than enough raw compute and 512 GB/s memory bandwidth.
The Silicon Root Cause: Non-Linear Lattice vs. Systolic DPAS
- Intel Arc Matrix Engines (XMX) utilize DPAS (Dot Product Accumulate Systolic) hardware. DPAS requires dense, linearly indexed integers (
INT2,INT4,INT8) to feed its systolic pipeline at thousands of operations per clock cycle. - I-Quants (
IQ2_S,IQ3_XXS,IQ4_XS,IQ3_S) use non-linear 8-dimensional vector lattice codebooks ($E_8$). They require runtime table lookups per weight element. - In
llama.cpp's SYCL backend, there are zero kernel implementations for I-Quants. Whenever an IQ tensor is evaluated, SYCL drops out of hardware systolic GEMM into slow CPU/host dequantization loops. - The Hidden Catch in Unsloth "UD" Models: When we reverse-engineered Unsloth's popular
Qwen3.8-27B-UD-Q4_K_S.gguf, we discovered it secretly contains 201 I-Quant tensors (IQ4_XS,IQ3_S,IQ3_XXS), causing it to suffer from the exact same dequantization stall on Intel Arc!
The Solution: Selective Software Transquantization
Using a custom selective software parser in llama.cpp, we dynamically intercepted and transquantized every unsupported non-linear I-Quant into native Q4_K super-blocks (with independent 6-bit scales and mins per 32 weights), while preserving rotary embeddings (RoPE), layernorms, linear SSM states, high-precision attention projections, and โ in the MTP-included checkpoints โ the embedded NextN prediction head bit-for-bit.
Benchmark Results on Intel Arc Pro B60 (24GB VRAM)
All benchmarks measured on device level_zero:0 (UR_LOADER_USE_LEVEL_ZERO_V2=0) with -ngl 99 -fa on:
| Model Variant | File Size | Quant Composition | Pure Autoregressive Decode | Prompt Prefill | Speculative Decode (MTP / DFlash2) |
|---|---|---|---|---|---|
| ISTA-DASLab IQ3_XXS (Stock) | 9.39 GiB | 355 IQ tensors (81% of model) | 8.10 tok/s | 110.05 tok/s | N/A (Stalled) |
| Qwen3.8-27B-Ridge (Stock) | 11.72 GiB | 192 IQ tensors | 12.22 tok/s | 133.45 tok/s | ~13.2 tok/s |
| Unsloth UD-Q4_K_S (Stock) | 14.29 GiB | 201 IQ tensors | 12.97 tok/s | 162.75 tok/s | ~18.9 tok/s (Verify bottlenecked) |
| Ridge Intel Arc Tuned Q4_K (Ours) | 15.29 GiB | 0 IQ (336 Q4_K, 51 Q5_K, 23 Q6_K) |
22.23 tok/s | 245.54 tok/s | 41.28 tok/s (Embedded MTP, 93.4% accept) |
| DAS Lab Intel Arc Tuned Q4_K (Ours) | 13.72 GiB | 0 IQ (374 Q4_K, 28 Q2_K, 96 BF16) |
22.36 tok/s | 243.61 tok/s | 36.03 tok/s (DFlash2, 93.5% accept) |
| DAS Lab IQ3_S Intel Arc Tuned MTP (Ours, NEW) | 14.31 GiB | 0 IQ (350 IQ tensors โ Q4_K) |
21.42 tok/s | โ โ | 41.88 tok/s (Embedded MTP, 92.9% accept) |
โ Prefill for this variant was not separately benchmarked; it shares the same architecture-lane and transquantized tensor layout as the other 0-IQ variants (~240+ tok/s expected).
- Pure Decode: Leaped from 8.10 tok/s $\to$ 22.36 tok/s (+176% / ~2.8ร faster).
- Speculative Decode: Smashes past 41 tokens/second using the embedded NextN prediction head โ now available on both the Ridge and the DAS Lab IQ3_S tuned bases, with zero external drafter VRAM.
New: Embedded MTP Now Shipping in the DAS Lab Line (Measured Sweep, 2026-09-17)
The latest tuned checkpoint, Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf, is the DAS Lab GSQ-RCO IQ3_S-MTP checkpoint transquantized to pure Q4_K โ the MTP head (blk.64.nextn.*, 41 MiB) rides along untouched. Earlier DAS Lab checkpoints (IQ3_XXS) shipped with their NextN heads stripped, which is why the previous DAS Lab release needed an external DFlash2 drafter. This one is self-contained.
Speculative topology sweep on the new checkpoint (real prompts: prime-factorization reasoning 1024 tok, codegen, JSON extraction):
| Configuration | Speculative Topology | Reasoning | Code Gen | JSON Extract | Draft Accept % |
|---|---|---|---|---|---|
| Config 0 | Pure Autoregressive (OFF) | 21.42 tok/s | 20.29 tok/s | 16.54 tok/s | โ |
| Config 1 โ | Embedded MTP (ngram-mod,draft-mtp, p-min=0.5, n-max=3) |
41.88 tok/s | 38.23 tok/s | 26.38 tok/s | 92.9% |
| Config 2 | External DFlash2 (DFlash2-Q4_K_M) |
39.68 tok/s | 39.19 tok/s | 27.11 tok/s | 94.7% |
| Config 3 | External DSpark (DSpark-Q8_0) |
โ SYCL kernel crash (common.hpp:145) |
โ | โ | โ |
Takeaways, all measured (not estimated):
- Embedded MTP wins: 41.88 tok/s crushes the external DFlash2 setup (39.68 tok/s) with zero drafter VRAM โ and beats the Ridge embedded-MTP release (41.28) by a whisker while being ~1 GiB smaller.
- DSpark is unusable on this stack: it hard-crashes the SYCL backend at first kernel sync. Drop the
Q8_0DSpark draft model entirely. - Small draft windows win: n-max 2โ3 with a strict acceptance floor (p-min โฅ 0.5) beat both wider windows and looser floors, at shallow and deep context.
Deep-Context Truth: Attention Cost Doesn't Care About Your Spec Method
We swept decode speed at ~150k-token depth on the same hardware (the production lane regularly serves ~148k contexts, so this is real traffic, not a hypothetical):
| Config @ ~150k depth | Decode | Acceptance |
|---|---|---|
| DFlash2 n-max=3, p-min=0.7 (deployed production config) | 7.91 tok/s | 84.7% |
| Native MTP n-max=2, p-min=0.5 + SYCL graph | 7.21 tok/s | 78.4% |
| DFlash2 n-max=2, p-min=0.5 | 6.88 tok/s | 67.0% |
Every topologies loses ~59% of its shallow speed by 150k tokens. That is attention-cost-of-depth โ the KV cache grows linearly, and no draft method or graph mode fixed it in our sweeps. If you serve extremely deep contexts, budget for ~7โ8 tok/s, not the shallow number.
Long-Context Needle Retrieval Restoration
Many users observed that extreme 2-bit I-Quants suffer from catastrophic Needle-In-A-Haystack (NIAH) failures past 32k context:
- Transformer Feed-Forward Networks (FFNs) function mechanistically as associative key-value factual memory banks (Geva et al., 2021).
- Coarse 8D lattice codebooks truncate high-magnitude outlier weights. Across deep contexts (50kโ200k+ tokens), quantization noise accumulates across all 64 layers, destroying the signal-to-noise ratio and creating "blind spots" where the model fails to retrieve facts.
Q4_KSuper-Blocks Restore Retrieval: By mapping weights into linear super-blocks with independent 6-bit scales and minimum offsets per 32 weights, dynamic range is fully preserved, restoring 100% green needle retrieval across long context windows.
The Intel Arc Pro B60 Sweet Spot: 200k Production Context on a Single 24GB Card
If you opted for the Intel Arc Pro B60 (24GB VRAM) instead of spending thousands more on the B70 or enterprise workstation cards, you might have worried about being forced to truncate your context window to 32k or 64k, or suffer massive speed penalties from CPU RAM offloading.
Here is the big win: You can run an enormous 200,000-token context window completely resident on a single Intel Arc Pro B60 24GB GPU with zero CPU offloading, BF16 multimodal vision, and maximum prefill throughput (-ub 1024, 448 tok/s prefill) โ while keeping decode generation blazing fast (41+ tokens/second with embedded MTP)!
The KV Cache Arithmetic: Why Long Context Fits in 24GB
Standard transformers require tens of gigabytes of VRAM for 200k+ contexts because every layer allocates quadratic key-value caches. Qwen 3.8 27B changes this math entirely through Hybrid Linear Attention (full_attention_interval = 4):
- 48 of the 64 layers are linear recurrent layers with a fixed, constant $O(1)$ state regardless of context depth.
- Only 16 layers allocate a dynamic token-by-token KV cache.
- Using quantized
q4_0KV cache (--cache-type-k q4_0 --cache-type-v q4_0), each token across all 16 attention layers costs only 18,432 bytes (~18.0 KiB / token): $$\text{KV Footprint} = 16 \text{ layers} \times 2 \times 4 \text{ heads} \times 256 \text{ dim} \times 0.5625 \text{ bytes/element} = 18,432\text{ bytes/token}$$
Measured Live Hardware VRAM Accounting on Intel Arc Pro B60 (24,480 MiB) โ Ridge Tuned (largest variant):
| Component | 200,000 Tokens (q4_0 KV) | Notes |
|---|---|---|
Model Weights (Transquant Q4_K) |
15.29 GiB (16,046 MiB) | 64 layers + norms (-ngl 99) |
Vision Projector (mmproj BF16) |
0.87 GiB (889 MiB) | Unsloth BF16 vision head (-ngld 99) |
| Embedded MTP Head & Recurrent State | 0.07 GiB (71 MiB) | 4 NextN tensors + 48 recurrent states |
Target + MTP KV Cache (q4_0) |
3.65 GiB (3,739 MiB) | 16 trunk layers + 1 MTP draft layer |
Compute Buffers (-ub 1024) & Workspaces |
2.37 GiB (2,425 MiB) | Full 448 tok/s XMX systolic batch prefill |
| Total Active VRAM Usage | 22.54 GiB (23,079 MiB) | Verified live via Vulkan / Level Zero |
| Remaining Free VRAM Headroom | ~1.40 GB (1,401 MiB) | Safe, rock-solid headroom for production! |
The new IQ3_S-MTP tuned base is ~1 GiB lighter (14.31 GiB), buying ~1 GiB of extra headroom on the same math โ you can bump KV precision to q4_1 or push context further.
Production Recommendation:
-c 200000is the tested, verified sweet spot for production. It preserves ~1.4 GB of healthy headroom for high-resolution vision tokens and desktop tasks while running-ub 1024for maximum prefill speed (448 t/s).
Available Models in this Repository
1. Qwen3.8-27B-Ridge-Intel-Arc-Tuned-Q4_K.gguf (15.29 GiB) โ Max Throughput, Long Context
- Embedded MTP Included: Contains 4 embedded NextN prediction tensors (
blk.64.nextn.*, 41 MiB). - Zero External Drafter VRAM Overhead: Runs speculative decoding without needing an external
-mdfile, saving 1.1โ1.3 GiB of VRAM. - Speed: Delivers 41.28 tok/s at 93.4% acceptance.
- Context: Fully verified with
-c 200000(leaves ~1.4 GB headroom with BF16 vision and full 448 t/s prefill batching on a 24GB B60).
2. Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf (14.31 GiB) โ DAS Lab + Embedded MTP (Q4_K) โญ
- Based on ISTA-DASLab's GSQ-RCO IQ3_S-MTP checkpoint transquantized to pure
Q4_Kโ 0 IQ tensors remain, full XMX DPAS acceleration. - Embedded MTP: 41.88 tok/s at 92.9% acceptance โ top measured result in the repo, zero external drafter VRAM.
3. Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-Q4_K.gguf (13.72 GiB) โ Ultra-Compact Base
- Based on ISTA-DASLab's optimized GSQ-RCO weights (374
Q4_K, 28Q2_K, 0 IQ). - Pure decode at 22.36 tok/s, pairs with external DFlash2 drafter for 36.03 tok/s average.
4. Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.7BPW.gguf (11.88 GiB) โ Profile A (Mid-Compression)
- Quantized directly from virgin BF16 base weights using ISTA-DASLab's RCO sensitivity mapping (135
Q4_K, 223Q3_K, 44Q2_K, 8Q6_K, 0 IQ). - VRAM Footprint: Leaves ~9.5 GB free VRAM on 24GB GPUs; leaves ~4.2 GB free on 16GB cards.
- Embedded MTP Included: 89.0% draft acceptance.
- โ ๏ธ See Architecture Advisory below regarding Intel Arc Xe2 decode speed on sub-4 BPW weights.
5. Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.2BPW.gguf (10.45 GiB) โ Profile B (Ultra-Compact for 12โ16GB Cards)
- Quantized directly from virgin BF16 base weights using ISTA-DASLab's high-compression RCO mapping (52
Q4_K, 181Q3_K, 169Q2_K, 8Q6_K, 0 IQ). - VRAM Footprint: Fits inside 12GB and 16GB GPUs (Intel Arc Pro B50, Arc A770 16GB, RTX 4070 Ti) with room for 32kโ128k context without CPU offloading.
- Embedded MTP Included: 90.5% draft acceptance.
- โ ๏ธ See Architecture Advisory below regarding Intel Arc Xe2 decode speed on sub-4 BPW weights.
โ ๏ธ CRITICAL ARCHITECTURAL ADVISORY: Sub-4 BPW Models (3.7BPW & 3.2BPW) on Intel Arc
Performance & Compatibility Notice for Intel Arc Users: The sub-4 BPW models (
Qwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.7BPW.ggufandQwen3.8-27B-GSQ-RCO-Intel-Arc-Tuned-MTP-3.2BPW.gguf) are not optimal for Intel Arc Pro architectures and DO NOT hit the stipulated 44+ tok/s speculative decode speeds achieved by our nativeQ4_Kcheckpoints.The Silicon Root Cause:
- Native INT4/INT8 Alignment: Intel Arc Xe2 XMX matrix engines use DPAS (Dot Product Accumulate Systolic) hardware that operates natively on packed 4-bit (
int4) and 8-bit (int8) matrix dot products (dpas.8x8).- In
Q4_K, weights stream directly into DPAS execution units with zero unpacking friction, enabling 38โ40 tok/s sustained decode and 44โ48 tok/s peak speculative bursts.- In sub-4 BPW formats (
Q3_KandQ2_K), weights are packed into irregular, non-power-of-two bit widths. In GGML SYCL compute kernels, Execution Units (EUs) must perform heavy bitwise shifting, masking, and scale lookups per super-block before loading data into DPAS.- The Compute vs. Bandwidth Bottleneck: On cards with ample memory bandwidth (like the Arc Pro B60 with 456 GB/s), this ALU unpack overhead bottlenecks the pipeline, capping decode speeds at ~24โ28 tok/s despite transferring fewer bytes from VRAM.
When to use Sub-4 BPW: Use
3.7BPWor3.2BPWstrictly if you are VRAM-constrained (e.g. 12GBโ16GB cards like Arc Pro B50 or Arc A770 16GB) where fitting 32kโ128k context into VRAM without host RAM spillover takes priority over peak decode speed. For 24GB cards (Arc Pro B60, RTX 3090/4090), always use the nativeQ4_Kmodels for maximum throughput.
Quickstart for Intel Arc Users
Critical Environment Variables for Intel Arc (Battlemage Xe2 / oneAPI 2026.x)
# Sourcing oneAPI
source /opt/intel/oneapi/setvars.sh --force
# CRITICAL FOR BATTLEMAGE: Force Level Zero V1 UR adapter (prevents error 44)
export UR_LOADER_USE_LEVEL_ZERO_V2=0
export ONEAPI_DEVICE_SELECTOR=level_zero:0
export ZE_AFFINITY_MASK=0
export ZES_ENABLE_SYSMAN=1
Running with 200k Context, Embedded MTP & Vision (41+ tok/s)
./llama-server \
-m Qwen3.8-27B-GSQ-RCO-IQ3_S-Intel-Arc-Tuned-MTP-Q4_K.gguf \
--mmproj mmproj-Qwen3.8-27B-unsloth-BF16.gguf \
--spec-type ngram-mod,draft-mtp \
--spec-ngram-mod-n-match 24 \
--spec-ngram-mod-n-min 1 \
--spec-ngram-mod-n-max 3 \
--spec-draft-n-max 3 \
--spec-draft-p-min 0.5 \
-c 200000 \
--cache-type-k q4_0 --cache-type-v q4_0 \
-ngl 99 -ngld 99 -fa on \
-ub 1024 -b 1024 -t 4 -tb 6 --parallel 1 \
--jinja --chat-template-file chat_templates/qwen-sharp.jinja --reasoning-preserve \
--host 127.0.0.1 --port 8183
Optional: SYCL Graph mode (
GGML_SYCL_ENABLE_GRAPH=1) gives +19% with embedded MTP (14.68 โ 17.51 tok/s measured shallow). โ ๏ธ Do not combine it with external draft models (DFlash2/DSpark) โ it crashes withGraph nodes cannot depend on events from outside the graph. Pure decode gains nothing from it. Safe only with embedded MTP (draft-mtp) or no speculation.
Chat Template Included:
chat_templates/qwen-sharp.jinja(byfroggeric) is included directly in this repository. Pass it via--chat-template-fileto enable dynamic<|think_on|>/<|think_off|>reasoning toggles, robust agent tool parsing, and avoid empty-think context degradation inllama.cpp.
Credits, Attribution & Upstream Acknowledgments
This repository provides hardware-adapted conversions specifically optimized for Intel Arc GPUs. We give full credit, attribution, and gratitude to the original creators whose breakthrough models made this work possible:
1. Base Model Architecture & Foundation Weights
- Qwen Team (Alibaba Cloud): Creators of the original Qwen3.8-27B foundation model.
- Repository: Qwen/Qwen2.5-32B / Qwen Series
- License: Apache 2.0
2. Ridge Quantization & Multi-Token Prediction Head
- empero-ai: Creators of the
Qwen3.8-27B-Ridgequantization and embedded multi-token prediction configuration. - Original Repository: empero-ai/Qwen3.8-27B-Ridge-GGUF
- License: Apache 2.0
- Notice of Modification (Apache 2.0 ยง4b): Transquantized 192 unsupported non-linear FFN lattice tensors (
IQ2_S/IQ3_S) into hardware-accelerated linearQ4_Ksuper-blocks for Intel Arc systolic DPAS execution.
3. GSQ & Randomized Coordinate Optimization (RCO)
- ISTA-DASLab: Creators of the GSQ-RCO quantization method and checkpoints (both the
IQ3_XXSandIQ3_S-MTPlines). - Original Repository: ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF
- Paper: QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Vector Lattices
- License: Apache 2.0
- Notice of Modification (Apache 2.0 ยง4b): Transquantized unsupported non-linear lattice tensors into linear
Q4_Ksuper-blocks โ 355 tensors on theIQ3_XXSbase (IQ1_S,IQ1_M,IQ2_XXS,IQ2_XS,IQ2_S,IQ3_XXS,IQ3_S,IQ4_XS), and 350 tensors on theIQ3_S-MTPbase (IQ2_S,IQ3_S,IQ4_XS). The embedded NextN prediction head (blk.64.nextn.*) is preserved bit-for-bit on the MTP line.
4. Chat Template & Reasoning Protocol
- froggeric: Creator of the optimized Jinja2 chat templates for Qwen models (
qwen-sharp/qwen3.8-froggeric). - Original Repository: froggeric/Qwen-Fixed-Chat-Templates
- Features: C++ engine compatibility in
llama.cpp, dynamic reasoning toggles (<|think_on|>,<|think_off|>,<|think_xhigh|>), mitigation of empty-think KV cache bloat, and resilient tool-call serialization.
5. Inference Engine & Toolchain
- Georgi Gerganov & the llama.cpp community: For
llama.cppand the open-source GGML SYCL backend. - Intel oneAPI Team: For the DPC++ / SYCL compiler and Level Zero runtime drivers.
Methodology Notes (how these numbers were measured)
- Hardware: Intel Arc Pro B60 24GB (Battlemage Xe2, 24,480 MiB), Ubuntu 24.04, oneAPI 2026.x, llama.cpp
build-sycl(SYCL backend). llama-bench -p 128 -n 32 -r 1 -ngl 99 -fa onfor the pure AR/prefill table; server-side/completiontimings with--metricsfor the speculative sweeps (1,024-token reasoning / 400-token codegen / 69-token JSON extraction prompts).- Depth tests: synthetic ~150k-token prompt with
ignore_eos:true,predicted_nsanity-checked againstn_predict(a silent false-negative failure mode was caught and gated in our harness). - The transquantizer runs inside
llama.cpp'sllama-quantize(branchsycl-iquants-mmq):--allow-requantize --tensor-type "^iq=q4_K" <in> <out> COPY 12.
- Downloads last month
- 4,567
3-bit