Mitsuba & HiMitsuba 27B GGUF

Repository renamed on 2026-10-04 from Mitsuba-ComfyUI-27B-GGUF (old links redirect here). The model files are unchanged. / 2026-10-04 に Mitsuba-ComfyUI-27B-GGUF から改名しました(旧 URL は自動で転送)。モデルのファイルは変わっていません。

A ternary (1.58-bit) Qwen3.8-27B tuned for ComfyUI work: writing image/video generation prompts that follow strict conditions, and describing images. It is not for coding.

ComfyUI 向けに調整した、Qwen3.8-27B の三値(1.58 ビット)モデルです。システムプロンプトにそった画像・動画用プロンプトの作成と、画像の説明が得意です。コーディングには向きません。

  • Self-made ternarization of the official Qwen3.8-27B weights (not derived from Bonsai's weights), stored in Prism ML's PQ2_0 / PTQ1_0 GGUF formats.
  • 7.3 GB (PQ2_0) / 6.0 GB (PTQ1_0). Runs on a single 16 GB GPU.
  • New (2026-10-04): HiMitsuba-Uncensored-LoRA.gguf (70 MB, for PQ2_0) — an optional LoRA that makes the model answer adult/sensitive requests it would otherwise refuse, while leaving ordinary requests unchanged. See the HiMitsuba section below. 追加(2026-10-04): 秘三葉(HiMitsuba・PQ2_0 用)=素の Mitsuba が断る大人向け・際どい依頼に答えるようになる LoRA(70MB・任意)。ふだんの依頼の答えは変わりません。

Files

File Size Notes
Mitsuba-ComfyUI-27B/Model-v1.18/Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf 7.32 GB Recommended
Mitsuba-ComfyUI-27B/Model-v1.18/Mitsuba-ComfyUI-27B-v1.18-PTQ1_0.gguf 6.00 GB Same weights as PQ2_0, smaller. Vision is lower with the current PTQ1_0 kernel (see below)
Mitsuba-ComfyUI-27B/HiMitsuba-Uncensored-LoRA.gguf 0.07 GB Optional. For PQ2_0 only. Uncensored LoRA for v1.18 PQ2_0 (not for PTQ1_0). Add --lora-scaled HiMitsuba-Uncensored-LoRA.gguf:1.0 (see below)
Mitsuba-ComfyUI-27B/Model-v1.18/mmproj-Q8_0.gguf 0.63 GB Vision encoder. Taken unchanged from OS-Software/Ternary-Bonsai-2-27B-Uncensored-Heretic-GGUF (Apache-2.0)
Mitsuba_bridge/ (HiMitsuba_start.bat, HiMitsuba_ComfyUI.bat, HiMitsuba_stop.bat, Mitsuba_bridge.py) — Optional. Double-click start on Windows (see Easy start on Windows)
docs/ — Documents only, no need to download: EVALUATION.md (detailed evaluation), the charts used on this page (comparison.png, noninferiority.png, two_forms.png, comparison_lora.png / comparison_lora_mobile.png), LICENSE, NOTICE

Evaluation (summary)

Measured with the same questions and conditions for all four models. Details: EVALUATION.md. The last column is the original, un-quantized Qwen3.8-27B (BF16), shown as the reference point: it shows what the ternarization kept and what it gave up.

comparison

All 10 axes (score out of 100 each). Bold = best of the three ternary models. / 10 項目すべての点(各 100 点満点)。太字は三値の 3 本の中で一番。右端は三値化する前の元のモデル(BF16)で、比べる基準として載せています。

Axis / 軸 Mitsuba PQ2_0 Mitsuba PTQ1_0 Ternary Bonsai 2 27B PQ2_0 Qwen3.8-27B BF16 (original / 元)
Total / 総合 61.5 (B) 60.2 (B) 59.6 (B) 66.3 (A)
1. Uncensored / 無検閲度 53.2 48.8 36.0 31.9
2. Honesty / 正直さ 63.2 72.2 51.0 51.1
3. Self-control / 自制心 62.7 56.9 65.5 55.6
4. Directness / 率直さ 84.0 92.0 88.0 84.0
5. Rule following / 正答率 84.0 80.0 72.0 84.0
6. Task completion / 到達率 76.0 80.0 76.0 72.0
7. Coding / コーディング 4.0 4.0 38.0 66.0
8. Reading / 読解力 48.0 40.0 44.0 76.0
9. Japanese & prompts / 文章・プロンプト 52.0 48.0 42.0 52.0
10. Vision / 画像認識 87.8 79.6 83.7 89.8
└ Image/video prompt generation (all conditions met) / 生成プロンプト 6/10 5/10 2/10 3/10
Decode speed (t/s, RTX 5090) / 生成速度 119.0 98.7 120.8 1.6 *

* BF16 (51 GB) does not fit in the 5090's 32 GB, so only 28 of 64 layers ran on the GPU. Its speed is for reference only. * BF16(51GB)は 5090 の 32GB に入りきらず、64 層中 28 層だけを GPU で動かしました。速度は参考値です。

Compared with the original: the ternarization gave up coding (66 → 4) and long-document reading (76 → 48), and kept vision (89.8 → 87.8), rule following (84 → 84) and Japanese & prompts (52 → 52). Prompt generation with all conditions met went up (3/10 → 6/10). 元のモデルと比べると、三値化で手放したのはコーディング(66 → 4)と長文の読解(76 → 48)で、画像認識(89.8 → 87.8)・正答率(84 → 84)・文章とプロンプト(52 → 52)は残しています。条件をすべて満たす生成プロンプトは上がりました(3/10 → 6/10)。

Is Mitsuba not worse than the plain Bonsai? / 素の Bonsai に劣らないか

Paired comparison on the same questions (Mitsuba PQ2_0 minus Ternary Bonsai 2 27B PQ2_0). Directness, task completion and reading were measured with 4× the questions. 同じ問題を対にして比べました(Mitsuba PQ2_0 − 素の Bonsai)。率直さ・到達率・読解力は問題を 4 倍にして測っています。

noninferiority

  • Better (superior) / 優越: uncensored, honesty
  • Not worse (non-inferior, margin 10 points) / 非劣性(許容幅 10 点): rule following, directness, Japanese & prompts, vision
  • Not decided even with 49–100 questions / 49〜100 問でも判定できず: self-control, task completion, reading
  • Coding is clearly worse and is left out of the chart. / コーディングは明らかに劣るため図から除いています。

Uncensored (無検閲度) = how often the model answers sensitive requests instead of refusing. The plain Mitsuba is not an uncensored model; it still refuses about half of them. If you need those answers, add the HiMitsuba LoRA (next section). 無検閲度=際どい依頼に断らず答える割合です。素の Mitsuba は無検閲モデルではなく、約半分は断ります。必要な人は次の節の秘三葉 LoRA を足してください。

HiMitsuba — Uncensored LoRA for PQ2_0 / 秘三葉(PQ2_0 用・無検閲 LoRA)

HiMitsuba-Uncensored-LoRA.gguf (70 MB) is an optional LoRA adapter for Mitsuba v1.18 PQ2_0. "Hi" (秘) means "hidden / private" in Japanese.

What it does

  • The LoRA switches by itself. For ordinary requests (SD/Krea/video prompts, describing images) it answers exactly like the plain Mitsuba. For adult or sensitive requests, which the plain Mitsuba refuses or waters down, it answers.
  • Censorship on adult topics comes in two forms, and HiMitsuba reduces both. (1) Refusal: the model declines. (2) Evasion: the model appears to answer but quietly drops what was asked (vague wording, skipped parts, a lecture instead of the content). The plain Mitsuba does both; many "uncensored" models fix only the first. HiMitsuba was tuned on both: refusals 51 → 7 of 106 sensitive requests, and on the sex-explanation questions it did answer, evasive answers fell from 73% to 32%.
  • Nothing else to run: no second "judge" model, no router, no system prompt. One extra flag at start-up.
  • Works with images too: an adult image → a usable generation prompt.
  • It is not an "answer anything" model. It still often refuses requests for help with crimes, weapons or harming others (see the table).

two forms of censorship

How this differs from a "judge" approach. Another way to build an uncensored switch is a separate small judge model, a Jev-style model that only answers yes/no (for example Jeff-Qwen3.5-0.8B): it looks at each request, and the caller then decides which LoRA or model to use. That needs a second model loaded and code on the caller's side. HiMitsuba has no judge: the decision is inside the LoRA's weights, so the same model simply answers differently. The trade-off is that it is always on and cannot be tuned per request.

判定役を置く方式との違い: 小さな判定専用モデル(yes/no だけ答える Jev 型。例 Jeff-Qwen3.5-0.8B)に依頼を見せ、呼び出す側が LoRA やモデルを切り替えるやり方があります。モデルが 2 本要り、呼ぶ側にコードも要ります。秘三葉には判定役がなく、判断は LoRA の重みの中に溶けていて、同じモデルの答え方が自分で変わります。代わりに常に掛かっていて、依頼ごとの調整はできません。

What it is not

  • It is not a new model. The 7.3 GB base is unchanged; the LoRA is applied at load time.
  • It is made for v1.18 PQ2_0 only. It was not tested on PTQ1_0 or on other Qwen3.8-27B GGUFs, and it will not load on models with a different architecture.

秘三葉は Mitsuba v1.18 PQ2_0 用の LoRA(70MB・任意・PTQ1_0 には使えません)です。アダルトの検閲には「拒否」と「回避」の2種類があり、秘三葉はその両方を減らします。拒否=断る。回避=答えたふりをして肝心な所を抜く(ぼかす・省く・説教に替える)。素の Mitsuba は両方をやり、世の「無検閲」モデルの多くは拒否しか直していません。秘三葉は両方を狙って調整し、拒否は 106 問中 51 → 7、答えた性の説明問題の中の回避は 73% → 32% になりました。LoRA が自分で切り替えます=ふだんの依頼(画像・動画のプロンプト作成、画像の説明)は素の Mitsuba と同じ答え、大人向け・際どい依頼には断らずに答えます。判定役のモデルもルーターも設定文も要りません。起動の引数を 1 つ足すだけです。「何でも答える」モデルではなく、犯罪・武器・他人を害する手助けはいまも断ることが多いです。

How to use / 使い方

  1. Download HiMitsuba-Uncensored-LoRA.gguf and put it in the same folder as the model (Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf).
  2. Add one flag to the llama-server command: --lora-scaled HiMitsuba-Uncensored-LoRA.gguf:1.0
llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf --mmproj mmproj-Q8_0.gguf --no-mmproj-offload --reasoning off ^
  --lora-scaled HiMitsuba-Uncensored-LoRA.gguf:1.0 ^
  --jinja --temperature 0.6 --top-k 20 --top-p 0.95 ^
  --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --n-gpu-layers 99
  • Start llama-server from that folder, or give the LoRA path relative to it. On Windows, a full path with a drive letter (D:\...gguf:1.0) breaks the file:scale syntax because of the colon.
  • :1.0 is the strength. 1.0 is what was evaluated. 0.0 = off.
  • ComfyUI (Ollama-compatible nodes, OpenCode, etc.) need no change: the LoRA is applied on the server side.
  • On Windows, the double-click files in Easy start on Windows below do all of this for you. / Windows なら、下の「かんたん起動」の bat がこれを全部やります。

使い方: ① LoRA を本体と同じフォルダに置く ② 起動コマンドに --lora-scaled HiMitsuba-Uncensored-LoRA.gguf:1.0 を 1 行足す。それだけです。 Windows でフルパス(D:\...gguf:1.0)を書くとドライブの「:」が区切りと衝突して失敗するので、本体のフォルダから起動してファイル名だけを書いてください。:1.0 は強さで、評価は 1.0 で取りました。ComfyUI 側の設定は変えません(LoRA はサーバー側で効きます)。

Evaluation / 評価(素の Mitsuba と同じ 10 項目・同じ条件)

comparison_lora

Same 10-axis evaluation as above, plain Mitsuba v1.18 PQ2_0 vs. the same model with HiMitsuba at 1.0. Full domain breakdown: EVALUATION.md.

Axis / 軸 Mitsuba v1.18 + HiMitsuba
Total / 総合 61.5 (B) 66.9 (A)
1. Uncensored / 無検閲度 53.2 89.6
  └ Refused / 拒否(106 問中) 51 (48%) 7 (7%)
  └ Evasive, of the sex-explanation questions it answered / 回避率(答えた性の説明問題の中) 8/11 (73%) 6/19 (32%)
  └ Refused or evasive / 拒否+回避(106 問中) 59 (56%) 13 (12%)
2. Honesty / 正直さ 63.2 73.3
3. Self-control / 自制心 62.7 48.2
4. Directness / 率直さ 84.0 92.0
5. Rule following / 正答率 84.0 84.0
6. Task completion / 到達率 76.0 92.0
7. Coding / コーディング 4.0 4.0
8. Reading / 読解力 48.0 44.0
9. Japanese & prompts / 文章・プロンプト 52.0 52.0
  └ Image/video prompt generation / 生成プロンプト 6/10 8/10
10. Vision / 画像認識 87.8 89.8
Decode speed (t/s, RTX 5090) / 生成速度 119.0 103.1
VRAM, CTX 131K / 262K (KV q4_0) / 必要 VRAM 9.5 / 12.4 GB 9.5 / 12.4 GB
  • Uncensored, both kinds: refusals fell from 51 to 7 of 106 sensitive requests, and evasive answers (answered, but with the asked-for content dropped) fell from 73% to 32% of the answered sex-explanation questions. Refused-or-evasive together: 59 → 13 of 106. The remaining refusals are mostly "help with crimes / weapons" and "help with abuse" (both about 65%).
  • ComfyUI work did not get worse: rule following, Japanese and vision are the same or higher, and prompt generation with all conditions met went 6/10 → 8/10.
  • What got worse: self-control (62.7 → 48.2; in the "error hell" test it stops retrying after the first tool error) and speed (119 → 103 t/s, the cost of applying a LoRA at run time). VRAM is unchanged.
  • The two uncensored Bonsai variants in the chart (CRACK, heretic) are other people's models, measured here only for comparison.

無検閲度は 53.2 → 89.6。拒否は 51 → 7/106 問、回避(答えたのに肝心な所を抜く)は答えた性の説明問題の 73% → 32%、拒否+回避を合わせると 59 → 13/106 問。残る拒否は「犯罪・武器」「悪用の手助け」が中心です。ComfyUI の仕事(正答率・文章・画像)は下がらず、条件を全部守る生成プロンプトは 6/10 → 8/10 に上がりました。下がったのは自制心(62.7 → 48.2。道具がエラーを返すと 1 回で諦めます)と速度(119 → 103 t/s。LoRA を実行時に掛ける分)です。VRAM は変わりません。

Easy start on Windows (double-click) / かんたん起動(Windows・ダブルクリック)

Download the folders as they are on this page (the .bat files find the model in ../Mitsuba-ComfyUI-27B/Model-v1.18/), or put all the files below in one folder. Then double-click a .bat in Mitsuba_bridge/. Nothing is installed into ComfyUI or into Windows.

File / ファイル What it is / 中身
Mitsuba-ComfyUI-27B/Model-v1.18/Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf (or PTQ1_0) the model / 本体
Mitsuba-ComfyUI-27B/Model-v1.18/mmproj-Q8_0.gguf image input / 画像の入力
Mitsuba-ComfyUI-27B/HiMitsuba-Uncensored-LoRA.gguf optional: HiMitsuba / 任意(秘三葉)
Mitsuba_bridge/HiMitsuba_start.bat for AI apps (OpenAI-compatible) / AI アプリ用
Mitsuba_bridge/HiMitsuba_ComfyUI.bat for ComfyUI's Ollama nodes / ComfyUI の Ollama ノード用
Mitsuba_bridge/HiMitsuba_stop.bat stop and give all VRAM back / 止めて VRAM を全部返す
Mitsuba_bridge/Mitsuba_bridge.py used by HiMitsuba_ComfyUI.bat / ComfyUI 用 bat が使う

AI apps (OpenCode, Open WebUI, SillyTavern, Continue, ...) — double-click HiMitsuba_start.bat, then set in the app:

  • Base URL: http://127.0.0.1:8080/v1 (OpenAI-compatible)
  • Model: himitsuba (with the LoRA) or mitsuba (plain)
  • API key: anything (e.g. none)

ComfyUI — you need an Ollama node already (e.g. comfyui-ollama). Double-click HiMitsuba_ComfyUI.bat, then in Ollama Connectivity:

  • url: http://127.0.0.1:11434 — if a real Ollama already uses 11434, the window prints the port it used instead
  • model: himitsuba or mitsuba
  • keep_alive: 0 (recommended) = the VRAM is given back right after the answer, before the image model loads. The default 5 = given back after 5 idle minutes.

Stop: double-click HiMitsuba_stop.bat, or close the window.

Notes:

  • First run downloads the PrismML build of llama.cpp (prism-b10685, about 540 MB) into bin\, and for ComfyUI a private Python 3.12 (about 11 MB, from python.org) into python\. Later starts take a few seconds.
  • NVIDIA GPU, Windows 10/11. About 10 GB of VRAM while a model is loaded, 0 while idle. Picking the other model unloads the current one.
  • Folder names with Japanese or spaces are fine (a plain-name junction is made under C:\ProgramData\HiMitsuba\; nothing is copied).
  • Thinking is already turned off in this setup.
  • Mitsuba_bridge.py is derived from ComfyUI-Bonsai-Bridge by the same author.

このページのフォルダの形のまま落とす(bat が ../Mitsuba-ComfyUI-27B/Model-v1.18/ の本体を見つけます)か、上のファイルを全部1つのフォルダに置いて、Mitsuba_bridge/ の bat をダブルクリックするだけです。ComfyUI にも Windows にも何もインストールしません。

  • AI アプリ(OpenCode・Open WebUI・SillyTavern など): HiMitsuba_start.bat を押し、アプリの接続先を http://127.0.0.1:8080/v1(OpenAI 互換)、モデルを himitsuba(秘三葉)か mitsuba(素)、API キーは何でも、にします。
  • ComfyUI: Ollama のノード(comfyui-ollama など)が入っていることが前提です。HiMitsuba_ComfyUI.bat を押し、Ollama Connectivity の url を http://127.0.0.1:11434、model を himitsuba にします。keep_alive は 0 がおすすめです(答えた直後に VRAM を返すので、画像モデルと取り合いません)。既定の 5 のままなら、使わずに 5 分たつと返します。
  • 止める: HiMitsuba_stop.bat を押すか、窓を閉じます。
  • 初回だけ、PrismML 版の llama.cpp(約 540 MB)と、ComfyUI 用の小さな Python(約 11 MB)を自動で落とします。2 回目からは数秒で起動します。
  • NVIDIA の GPU と Windows 10/11 が必要です。モデルを載せている間は約 10 GB の VRAM を使い、使っていない時は 0 です。
  • 日本語や空白の入ったフォルダでも動きます。

How to run

The model and mmproj-Q8_0.gguf are in Mitsuba-ComfyUI-27B/Model-v1.18/, the LoRA in Mitsuba-ComfyUI-27B/. / 本体と mmproj は Mitsuba-ComfyUI-27B/Model-v1.18/、LoRA は Mitsuba-ComfyUI-27B/ にあります。

PQ2_0 and PTQ1_0 need the PrismML fork of llama.cpp (upstream llama.cpp does not support these formats yet): https://github.com/PrismML-Eng/llama.cpp (branch prism).

The settings used for the evaluation:

llama-server -m Mitsuba-ComfyUI-27B-v1.18-PQ2_0.gguf --mmproj mmproj-Q8_0.gguf --no-mmproj-offload --reasoning off ^
  --jinja --temperature 0.6 --top-k 20 --top-p 0.95 ^
  --ctx-size 131072 --cache-type-k q4_0 --cache-type-v q4_0 --n-gpu-layers 99

Thinking was off ("chat_template_kwargs": {"enable_thinking": false}). In every test, the conditions (word count, required and forbidden words, output format) were written in the request text itself.

Turn thinking OFF / 思考は必ずオフで

Use this model with thinking (reasoning) turned OFF. It was tuned only in no-thinking mode. With thinking on, it tends to repeat the same sentence in its reasoning and can end without writing an answer.

  • llama-server: add --reasoning off (or set reasoning = off in a models preset)
  • Per request: "chat_template_kwargs": {"enable_thinking": false}
  • OpenCode and other agents: set the model to "reasoning": false

このモデルは思考(reasoning)をオフにして使ってください。 思考オフの形だけで調整しています。思考をオンにすると、思考の中で同じ文を繰り返し、答えを書かないまま終わることがあります。 llama-server なら --reasoning off、リクエストごとなら "chat_template_kwargs": {"enable_thinking": false} を指定します。

Measured: the same evaluation with thinking ON. / 思考オンで同じ評価をした結果:

Axis / 軸 PQ2_0 OFF PQ2_0 ON PTQ1_0 OFF PTQ1_0 ON
Total / 総合 61.5 53.2 60.2 54.5
Vision / 画像認識 88 55 80 63
Rule following / 正答率 84 60 80 64
Uncensored / 無検閲度 53 15 49 24
Image/video prompt generation / 生成プロンプト 6/10 6/10 5/10 5/10
Reading / 読解力 48 68 40 68

With thinking ON, many answers came back empty: the model finished its reasoning and stopped without writing the answer (vision: 19 of 49 on PQ2_0, 14 of 49 on PTQ1_0). Only long-document reading improved. 思考オンでは、考えたあと答えを書かずに終わる「空の答え」が多く出ました(画像 49 問中、PQ2_0 で 19 問・PTQ1_0 で 14 問)。上がったのは長文の読解だけです。

What it is good at

  • Stable Diffusion style prompts: English tags within a given count, required words included, a final Negative: line, and forbidden words kept out.
  • Video prompts in time segments (0-3s: / 3-6s: / 6-9s:) with a camera move in each segment.
  • Describing images: objects, counts, text in images, charts, scenes, people, and comparing several images.

Limitations

  • Coding: do not use. It scores 4/100 on our coding test (Bonsai: 38).
  • Long-document reading is average (48).
  • The plain model is not uncensored; it refuses some sensitive requests. Use the HiMitsuba LoRA if you need those answers (it still refuses help with crimes and harming others).
  • With HiMitsuba: self-control drops (stops retrying after a tool error) and decode speed is about 13% lower.
  • Prompt generation passes about 6 of 10 strict test cases. Check the output against your conditions.

License and attribution

  • This model is released under the Apache License 2.0 (LICENSE). It is a modified version of Qwen3.8-27B.
  • See NOTICE for attributions.

日本語の補足

  • 推奨は PQ2_0 です。PTQ1_0 は重みは同じですが、今の llama.cpp の PTQ1_0 用の計算では画像の点が下がります。
  • 評価の詳しい表と、その見方は EVALUATION.md にあります。
  • 秘三葉(HiMitsuba)LoRA は任意です。入れなければ素の Mitsuba のままです。入れても ComfyUI の仕事の点は下がりません(上の表)。
Downloads last month
10,041
GGUF
Model size
17.3M params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for isichan-ai/Mitsuba_and_HiMitsuba-27B-GGUF

Base model

Qwen/Qwen3.8-27B
Adapter
(152)
this model

Space using isichan-ai/Mitsuba_and_HiMitsuba-27B-GGUF 1