Back to the comparison比較ページに戻る

How it works仕組みの解説

What happens inside Qwen-Image 2.1, what the Heretic encoder changes, and how LoRAs work, so the comparison makes sense.比較の意味が分かるように、Qwen-Image 2.1 の中で何が起きているか、Heretic エンコーダが何を変えるか、LoRA がどう働くかをまとめました。

Image generation in LLM terms

LLM と比べて見る画像生成

An LLM writes text one token at a time, left to right, each token sampled from what the transformer predicts next. A diffusion model such as Qwen-Image 2.1 also uses a transformer, but it produces the whole image at once: it starts from a canvas of pure noise and repeatedly predicts how to make all of it a little cleaner, until after 25 steps an image remains.

LLM は文章を、左から右へ1トークンずつ書きます。各トークンは、トランスフォーマーが予測した「次」から選ばれます。Qwen-Image 2.1 のような拡散モデルもトランスフォーマーを使いますが、画像は全体を一度に作ります。純粋なノイズのキャンバスから始め、「全体を少しきれいにする方向」の予測を繰り返して、25ステップ後に画像が残る、という作り方です。

LLMImage diffusion (Qwen-Image 2.1)画像の拡散モデル(Qwen-Image 2.1)
Unit単位token (one of ~150k ids)トークン(約15万種の id のどれか)latent patch: 64 numbers for a 16×16-pixel area潜在パッチ:16×16画素ぶんを64個の数で表したもの
“Tokenizer”「トークナイザ」BPEVAE: compresses pixels to patches and backVAE:画素とパッチを相互に変換
Modelモデルdecoder transformerデコーダ型トランスフォーマーdiffusion transformer (DiT), 7.1B parameters拡散トランスフォーマー(DiT)、71億パラメータ
Generation生成のしかたone token per forward pass, left to right1回の計算で1トークン、左から右へall 4,056 patches every pass, refined over 25 passes毎回4,056パッチ全部を計算し、25回かけて仕上げる
Training objective学習の目標predict the next token次のトークンを当てるfrom a noisy image, predict the direction back to the clean oneノイズを混ぜた画像から、元のきれいな画像へ戻る方向を当てる
Randomnessランダム性sampling (temperature, seed)サンプリング(温度、seed)the starting noise, fixed by the seed出発点のノイズ。seed で固定
The promptプロンプトpart of the same token sequence同じトークン列の一部read by a separate LLM, the text encoder別の LLM(テキストエンコーダ)が読む

How it learns

どうやって学習するか

Training is as simple as next-token prediction. Take a real image and its caption, mix the image with random noise at a random strength, and ask the transformer which way leads back to the original. Repeat over billions of image and caption pairs. Generating is the same move run in reverse: start from all noise and follow the predicted direction a step at a time. What the model can draw is therefore decided by its training images, just as an LLM's knowledge comes from its training text.

学習は、次のトークン予測と同じくらい単純です。実際の画像と説明文を用意し、画像にランダムな強さでノイズを混ぜて、「元の画像に戻るにはどちらへ進めばいいか」をトランスフォーマーに当てさせます。これを何十億組もの画像と説明文で繰り返します。生成はその逆回しです。全部ノイズの状態から始め、予測された方向へ1歩ずつ進みます。そのため、何を描けるかは学習に使った画像で決まります。LLM の知識が学習データの文章から来るのと同じです。

How the image model reads the prompt

画像モデルがプロンプトを読むしくみ

The image transformer doesn't understand words itself. A full LLM, Qwen3-VL-8B, reads the prompt, and the image transformer attends to that LLM's hidden states as extra tokens in its own sequence, the same way an LLM attends to earlier tokens. Because the text doesn't change between steps, its keys and values are computed once and reused for all 25 steps.

画像側のトランスフォーマーは、自分では言葉を理解しません。プロンプトは本物の LLM(Qwen3-VL-8B)が読み、画像側はその LLM の隠れ状態を、自分の系列に加えたトークンとして参照(アテンション)します。LLM が前のトークンを参照するのと同じしくみです。文章はステップの間で変わらないので、その key と value は一度計算すれば25ステップすべてで使い回せます。

This is the link to Heretic: everything the image model knows about your prompt passes through that LLM's internal activations. Change how the LLM represents a prompt, and you change what the image model is told, without touching the image model at all.

ここが Heretic につながるポイントです。画像モデルがプロンプトについて知ることは、すべてその LLM の内部の活性化を通ってきます。LLM がプロンプトをどう表すかを変えれば、画像モデルに伝わる内容が変わります。画像モデル自体には一切手を加えずにです。

How Qwen-Image 2.1 makes an image

Qwen-Image 2.1 で画像ができるまで

Qwen-Image 2.1 is three models working in a row. A text encoder turns the prompt into numbers that describe its meaning. A diffusion transformer starts from random noise and, step by step, turns it into an image that fits those numbers. A VAE decoder turns the result, which lives in a compressed form, into pixels.

Qwen-Image 2.1 は、3つのモデルを順番に通して画像を作ります。テキストエンコーダがプロンプトを、意味を表す数の並びに変えます。拡散トランスフォーマーはランダムなノイズから始めて、その数に合う画像へ少しずつ変えていきます。最後に VAE デコーダが、圧縮された形の結果を画素に戻します。

PipelinePromptプロンプトyour text入力した文章Text encoderテキストエンコーダQwen3-VL-8BQwen3-VL-8BHeretic replaces thisHeretic はここを差し替えDiffusion transformer拡散トランスフォーマー7.1B · 25 steps (Turbo 8)7.1B · 25ステップ(Turbo 8)LoRAs add to these weightsLoRA はここの重みに足すVAEVAEdecoderデコーダImage画像832×1248832×1248Random noiseランダムなノイズfixed by the seedseed で固定prompt features (hidden states)プロンプトの特徴(隠れ状態)latent: 52×78×64潜在表現:52×78×64
Where the two kinds of changes in this project apply.この実験の2種類の変更が、どこに効くか。

The image model never sees your words, only the encoder's hidden states (its internal representation of the prompt). It works on a latent image 16 times smaller than the output in each direction, with 64 numbers per position: 832×1248 pixels become a 52×78 grid, 4,056 positions in total.

画像モデルが受け取るのは文章そのものではなく、エンコーダの隠れ状態(プロンプトの内部表現)だけです。また、画像そのものではなく、縦横それぞれ16分の1の「潜在表現」の上で計算します。1マスあたり64個の数を持ち、832×1248 画素なら 52×78 マス、計4,056マスです。

The text encoder and Heretic

テキストエンコーダと Heretic

Where refusals come from

「断る」性質はどこから来るか

The encoder of Qwen-Image 2.1 is Qwen3-VL-8B-Instruct, a chat model. Chat models are trained to refuse some requests. Arditi et al. (2024) showed that in many chat models this behaviour runs through a single direction in the model's internal activations: when the direction lights up, the model refuses. An image model conditioned on such an encoder can inherit some of that steering, because it reads the same activations.

Qwen-Image 2.1 のエンコーダは、チャット用モデルの Qwen3-VL-8B-Instruct です。チャット用モデルは、一部の依頼を断るように学習されています。Arditi ら(2024)は、多くのチャット用モデルでこの性質が、内部の活性化の中のたった1つの方向を通って働いていることを示しました。その方向が強く出ると、モデルは断ります。画像モデルは同じ活性化を読むので、その「寄せ」の影響を一部受け継ぐ可能性があります。

Abliteration and what Heretic adds

abliteration と Heretic

Abliteration (directional ablation) estimates that direction by comparing activations on harmful and harmless prompts, then edits the weights so the model can no longer write along it. Nothing is retrained. Heretic automates the choice of how strongly and in which layers to do this: an Optuna (TPE) search tries many settings and keeps those that minimise two things at once, the number of refusals and the KL divergence, which measures how far the model's answers to harmless prompts drift from the original.

abliteration(directional ablation、方向の除去)は、問題のある依頼と無害な依頼に対する活性化を比べてその方向を推定し、モデルがその方向へ書き込めなくなるよう重みを編集する手法です。再学習はしません。Heretic は、どの層にどれだけ強くかけるかを自動で決めるツールです。Optuna(TPE)で多数の設定を試し、「断った回数」と「KL ダイバージェンス」(無害な依頼への答えが元からどれだけずれたか)の両方が小さいものを選びます。

Refusals (of 100)断った回数(100回中)KL divergenceKL ダイバージェンス
Official Qwen3-VL-8B公式 Qwen3-VL-8B1000
Heretic encoder used here (pottokao)ここで使う Heretic エンコーダ(pottokao)50.022

Because the search is randomised, two people running Heretic get slightly different models. The two numbers above are how you judge one: few refusals, and a KL divergence close to zero.

探索にランダム性があるので、Heretic を誰が回しても完全に同じモデルにはなりません。良し悪しは上の2つの数で判断します。断る回数が少なく、KL ダイバージェンスが 0 に近いほど良い結果です。

How this project applies it

この実験での使い方

The Heretic encoder has exactly the same 750 weight tensors, with the same names and shapes, as the official one. So nothing is converted: generate.py builds a model folder out of links, with the official transformer and VAE and the Heretic encoder in place of the official one. mflux checks every tensor name and shape when it loads. The image model itself is untouched.

Heretic エンコーダは、公式のものと重みの数(750個)も名前も形もまったく同じです。そのため変換は不要で、generate.py が「公式の変換器と VAE + Heretic エンコーダ」をリンクで組み合わせたモデルフォルダを作ります。mflux は読み込み時に、すべての名前と形を確認します。画像モデル本体には一切手を加えていません。

models/base-heretic/
  transformer  -> Qwen/Qwen-Image-2.1/transformer
  vae          -> Qwen/Qwen-Image-2.1/vae
  text_encoder -> pottokao/Qwen-Image-2.1-Text-Encoder-Heretic

On the prompts in this comparison the swap barely matters: the Heretic images differ from the official ones by 5.1 on average (out of 255), photos less, the manga a little more.

この比較のプロンプトでは、差し替えの影響はほとんどありません。公式エンコーダとの差は平均 5.1(255段階中)でした。写真系はそれより小さく、4コマ漫画は少し大きめです。

What a LoRA is

LoRA とは

Fine-tuning a whole 7.1-billion-parameter model is expensive and produces a 14 GB file. LoRA (low-rank adaptation, Hu et al. 2021) keeps the original weights frozen and learns only a small correction for selected layers. For a weight matrix W, the correction is the product of two thin matrices, B·A; their inner size, the rank, is small (16 for every LoRA here).

70億個以上の数を持つモデル全体を学習し直すのは大変で、14GB のファイルになります。LoRA(low-rank adaptation、Hu ら 2021)は元の重みを固定したまま、選んだ層への小さな補正だけを学習します。重み行列 W への補正は、細長い2つの行列の積 B·A で表します。間の大きさをランクと呼び、ここで使う LoRA はすべて 16 です。

LoRAW (frozen)W(固定)4096 × 40964096 × 409616.8 M numbers1,680万個の数++s ×s ×BB4096×164096×16··AA16×409616×4096==W′ (used)W′(実際に使う)same shape as WW と同じ形131 k numbers (0.8%)13.1万個の数(0.8%)
One layer: the LoRA stores only B and A, about 0.8% of the numbers in W.1つの層の例:LoRA が持つのは B と A だけで、W の約0.8%の量。

At load time each patched layer becomes W′ = W + s·B·A, where s is the strength you set (0.9 adds 90% of the learned change, 0 turns it off). The LoRAs here patch the attention layers (q, k, v, out) and the feed-forward layers of all 32 transformer blocks: 384 small matrices, 34 to 160 MB per file, against 14 GB for the transformer.

読み込むとき、対象の各層は W′ = W + s·B·A になります。s は設定した強さで、0.9 なら学習した変化の90%を足し、0 なら何もしません。ここで使う LoRA は、32個あるブロックすべての注意機構(q・k・v・out)と全結合層を補正します。小さな行列が384個で、ファイルは 34〜160MB です。変換器本体は 14GB あります。

Things that follow from the math

この仕組みから分かること

  • Stacking adds up. Two LoRAs give W + s₁B₁A₁ + s₂B₂A₂; where they touch the same layers they interfere, so keep the total strength moderate.
  • 重ねがけは足し算。2つなら W + s₁B₁A₁ + s₂B₂A₂ になります。同じ層に効く部分は干渉するので、強さの合計は控えめにします。
  • Strength is not a percentage of visual change. In our tests, two LoRAs at about the same strength (0.8 and 0.9) moved the picture by 38 and by 1.3: what matters is how large and how targeted B·A is.
  • 強さは見た目の変化の割合ではない。実験では、ほぼ同じ強さ(0.8 と 0.9)の2つの LoRA で、変化量が 38 と 1.3 でした。効き方は B·A の大きさと、どこに効くかで決まります。
  • Baking can lose tiny updates. mflux adds B·A into W once (“baking”) for speed. W is stored in bfloat16, which rounds away changes smaller than its precision; for one LoRA in our tests only 25% survived, so that kind of LoRA is better applied live (unmerged).
  • 合成すると小さな変化は消えることがある。mflux は速さのため、B·A を W に一度だけ足し込みます(合成、bake)。W は bfloat16 で保存されていて、その精度より小さい変化は丸めで消えます。実験では、ある LoRA で25%しか残らなかったので、そういう LoRA は合成せず、直接かける方式にします。
  • A LoRA fits the model it was trained on. These were trained on base Qwen-Image 2.1; Turbo has different weights, so the same correction can act differently there.
  • LoRA は学習したモデル用。ここで使う LoRA は通常版の Qwen-Image 2.1 で学習されています。Turbo は重みが違うので、同じ補正でも効き方が変わることがあります。

Steps, seeds and Turbo

ステップ・seed・Turbo

Generation starts from pure noise, chosen by the seed. At every step the transformer predicts which way to move toward a clean image, and the sampler takes a step along a schedule of noise levels (flow matching with Euler steps). The same seed means the same starting noise, which is why images in a column share their composition and differences come from the setup.

生成は、seed で決まる純粋なノイズから始まります。各ステップで変換器が「きれいな画像へ向かう方向」を予測し、ノイズの量を段階的に減らす予定表に沿って1歩ずつ進みます(flow matching、Euler 法)。seed が同じならスタートのノイズも同じなので、同じ列の画像は構図が似て、違いは設定の差から生まれます。

Turbo is the official distilled version: trained to take 8 large steps on a fixed schedule instead of 25 to 40 small ones, about three times faster. It needs that exact schedule, which stock mflux doesn't have yet, so this project adds it (turbo_scheduler.py).

Turbo は公式の高速版(蒸留版)です。25〜40回の小さな歩みの代わりに、決まった予定表で8回の大きな歩みをするよう学習されていて、約3倍速くなります。その予定表どおりに進める必要があり、標準の mflux にはまだないため、この実験で追加しています(turbo_scheduler.py)。

Reading the comparison page

比較ページの読み方

  • Every image uses the same prompt, seed and size; only the setup changes.
  • すべての画像は同じプロンプト・seed・サイズで、変わるのは設定だけです。
  • Each section asks one question and compares its setups with a baseline, outlined in green.
  • 各セクションは1つの問いを立て、緑の枠の基準画像と比べます。
  • Change is the average pixel difference from that baseline on small grayscale copies (0 = identical, 255 = opposite). It shows how far the picture moved, not whether it got better.
  • 変化量は、基準画像との平均ピクセル差です(小さく縮めた白黒画像で計算。0 = 同じ、255 = 正反対)。絵がどれだけ動いたかを示すもので、良くなったかどうかではありません。
  • Closed models such as Nano Banana have no seed control, so compare their quality and text, not composition.
  • Nano Banana のようなクローズドなモデルは seed を指定できないので、構図ではなく画質と文字を比べます。

Sources

出典