Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
57 commits
Select commit Hold shift + click to select a range
93ecda8
feat(fp8): enable FP8 storage for Z-Image
Pfannkuchensack Jul 31, 2026
cd8f052
Chore openapi
Pfannkuchensack Jul 31, 2026
b47a92c
feat(fp8): enable FP8 storage for Anima
Pfannkuchensack Jul 31, 2026
3feb027
fix(fp8): never apply FP8 storage to already-quantized weights
Pfannkuchensack Jul 31, 2026
1bec901
feat(quantization): run ComfyUI scaled-fp8 checkpoints on the fp8 ten…
Pfannkuchensack Jul 31, 2026
580813c
feat(krea2): keep ComfyUI scaled-fp8 checkpoints quantized and run th…
Pfannkuchensack Jul 31, 2026
efdf552
test(fp8): verify LoRA sidecar patching over fp8 weights, harden meta…
Pfannkuchensack Jul 31, 2026
37b9d0d
docs(fp8): correct the plan against the shipped implementation
Pfannkuchensack Aug 1, 2026
b2a31df
feat(qwen-image): add a tiling option to the image-to-latents node
Pfannkuchensack Aug 1, 2026
8a31eca
Add test
Pfannkuchensack Aug 1, 2026
39b2375
feat(qwen-image): make VAE tiling usable on both Qwen-Image VAE nodes
Pfannkuchensack Aug 1, 2026
a843a37
Merge branch 'feat/qwen_image_i2l_tiling' into feat/fp8_scaled_compute
Pfannkuchensack Aug 1, 2026
8826e8d
Merge branch 'main' into feat/fp8_quantized_guard
Pfannkuchensack Aug 2, 2026
ec0b1f3
Merge branch 'main' into feat/fp8_anima
Pfannkuchensack Aug 2, 2026
8734f29
Merge branch 'main' into feat/fp8_zimage
Pfannkuchensack Aug 2, 2026
a3784eb
fix(fp8): keep the Qwen3-VL encoder's unquantized layers fp8-resident
Pfannkuchensack Aug 4, 2026
cfd8a05
fix(fp8): ignore uncalibrated input_scale placeholders, read both spe…
Pfannkuchensack Aug 5, 2026
49990d6
docs(fp8): document that FP8 Compute needs the model fully resident
Pfannkuchensack Aug 5, 2026
6b52aec
Merge branch 'main' into feat/fp8_zimage
Pfannkuchensack Aug 6, 2026
6c38314
Merge branch 'main' into feat/fp8_anima
Pfannkuchensack Aug 6, 2026
5d81932
Merge branch 'main' into feat/fp8_quantized_guard
Pfannkuchensack Aug 6, 2026
e93854e
feat(fp8): run raw fp8 checkpoints on the tensor cores
Pfannkuchensack Aug 7, 2026
781cc9f
Merge branch 'main' into feat/fp8_compute_raw
Pfannkuchensack Aug 8, 2026
526cb36
fix(fp8): apply the keep-fp8 filters to raw weights only, probe the d…
Pfannkuchensack Aug 8, 2026
c4dec24
Merge remote-tracking branch 'upstream/main' into feat/fp8_zimage
Pfannkuchensack Aug 14, 2026
a8695cc
Merge branch 'feat/fp8_zimage' into feat/fp8_anima
Pfannkuchensack Aug 14, 2026
1f54512
Merge branch 'feat/fp8_anima' into feat/fp8_quantized_guard
Pfannkuchensack Aug 14, 2026
9cc87c2
Merge remote-tracking branch 'upstream/main' into feat/fp8_compute_raw
Pfannkuchensack Aug 14, 2026
0eca15f
fix(fp8): handle both weight-scale spellings in every loader
Pfannkuchensack Aug 14, 2026
772e3b0
test(fp8): capture a real scaled-fp8 Z-Image key layout
Pfannkuchensack Aug 14, 2026
83dd95e
test(fp8): capture a mixed fp8 FLUX.2 layout with input scales
Pfannkuchensack Aug 15, 2026
cb826ff
Merge remote-tracking branch 'upstream/main' into fix/9416-rebase
Pfannkuchensack Aug 17, 2026
a0c8b43
Merge branch 'main' into feat/fp8_quantized_guard
Pfannkuchensack Aug 17, 2026
4ccb6b1
Merge remote-tracking branch 'upstream/main' into feat/fp8_zimage
Pfannkuchensack Aug 19, 2026
2c00114
Merge remote-tracking branch 'upstream/main' into feat/fp8_compute_raw
Pfannkuchensack Aug 19, 2026
9ef7a2d
chore(fp8): regenerate openapi schema for the fp8_compute settings
Pfannkuchensack Aug 19, 2026
5c3f8aa
test(fp8): drop the Z-Image entry from the exclusion parametrize
Pfannkuchensack Aug 19, 2026
6db534a
Merge branch 'feat/fp8_zimage' into feat/fp8_anima
Pfannkuchensack Aug 19, 2026
93d3ce5
Merge branch 'feat/fp8_anima' into feat/fp8_quantized_guard
Pfannkuchensack Aug 19, 2026
decabd5
Merge branch 'feat/fp8_quantized_guard' into feat/fp8_compute_raw
Pfannkuchensack Aug 19, 2026
060f879
docs(config): list fp8_compute_full_precision_hints in the config doc…
Pfannkuchensack Aug 19, 2026
b2854a1
feat(fp8): run scaled fp8 FLUX.1 checkpoints on the tensor cores
Pfannkuchensack Aug 19, 2026
587c0fa
feat(fp8): run scaled fp8 FLUX.2 checkpoints on the tensor cores
Pfannkuchensack Aug 19, 2026
27a138f
feat(fp8): run scaled fp8 Mistral encoders on the tensor cores
Pfannkuchensack Aug 19, 2026
9f25f1e
feat(fp8): run scaled fp8 Anima checkpoints on the tensor cores
Pfannkuchensack Aug 19, 2026
d8f1618
Merge branch 'main' into feat/fp8_compute_raw
Pfannkuchensack Aug 24, 2026
5a5ddaf
Merge remote-tracking branch 'upstream/main' into feat/fp8_compute_raw
Pfannkuchensack Aug 25, 2026
21b39ab
fix(fp8): recover the weight scales the loaders were silently dropping
Pfannkuchensack Aug 25, 2026
3a0707d
test(fp8): make the matmul-probe tests run without a GPU
Pfannkuchensack Aug 25, 2026
9985274
fix(fp8): decode MXFP8 block scales instead of reading the exponent b…
Pfannkuchensack Aug 25, 2026
0a47cc4
fix(fp8): refuse MXFP8 block scales instead of loading them as noise
Pfannkuchensack Aug 25, 2026
fc3cc3f
fix(fp8): refuse MXFP8 block scales instead of loading them as noise
Pfannkuchensack Aug 25, 2026
922fb13
Merge branch 'main' into feat/fp8_compute_raw
Pfannkuchensack Aug 26, 2026
04b6f14
Merge branch 'main' into feat/fp8_compute_raw
Pfannkuchensack Aug 28, 2026
6fa8ac8
fix(fp8): address round-3 review — Qwen3-VL split, scale-axis and pro…
Pfannkuchensack Aug 29, 2026
863e96f
Merge branch 'main' into feat/fp8_compute_raw
Pfannkuchensack Aug 31, 2026
9cd24e0
fix(fp8): make the RAM prediction split-aware, pin the two unpinned f…
Pfannkuchensack Aug 31, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
38 changes: 34 additions & 4 deletions docs/src/content/docs/configuration/fp8-storage.mdx
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
---
title: FP8 Storage
title: FP8 Storage & Compute
sidebar:
order: 3
---
Expand Down Expand Up @@ -29,9 +29,9 @@ InvokeAI's FP8 path stores weights in FP8 and casts them back to BF16/FP16 on ea

The toggle works as advertised: the UNet / transformer drops by roughly 50% on the GPU. Per-step latency is the same or marginally slower because every forward pass adds an FP8 → BF16 cast on entry and a BF16 → FP8 cast on exit. This is the **largest target group**: 3090 owners squeezing FLUX into 24 GB benefit the most.

### RTX 40-series, RTX 50-series, and Hopper — VRAM win today, compute win possible later
### RTX 40-series, RTX 50-series, and Hopper — VRAM win, plus a compute win via a separate setting

These GPUs have native FP8 tensor cores. The toggle still buys you the same ~50% VRAM reduction today, because the forward pass still runs in BF16 — the hook casts weights back up to compute precision before each layer. If InvokeAI later wires up a true FP8 matmul path (e.g. via `torchao`), the same toggle will *also* unlock compute speedups on this hardware. Until then, treat the benefit as "VRAM only, same as Ampere".
These GPUs have native FP8 tensor cores. FP8 *Storage* on its own still only buys the ~50% VRAM reduction, because the forward pass runs in BF16 — the hook casts weights back up to compute precision before each layer. To actually use the tensor cores you need the separate [FP8 Compute](#fp8-compute) setting, which applies to checkpoints that ship pre-quantized ("scaled fp8") rather than to full-precision ones.

### Older CUDA cards — still a VRAM win

Expand Down Expand Up @@ -105,10 +105,40 @@ If you see unexpected quality regressions, disable FP8 Storage on the affected m

## Combining with Low-VRAM mode

**FP8 + partial loading**: fully supported. FP8 Storage shrinks the layers; partial loading streams them between RAM and VRAM as needed. Use both on tight VRAM budgets.
**FP8 Storage + partial loading**: fully supported. FP8 Storage shrinks the layers; partial loading streams them between RAM and VRAM as needed. Use both on tight VRAM budgets.

**FP8 Compute + partial loading** is a different story — it still works, but it costs both speed and reproducibility. See [FP8 Compute](#fp8-compute) below.

(For why FP8 Storage doesn't stack on top of GGUF / NF4 / int8 checkpoints, see the callout at the top of this page.)

## FP8 Compute

Everything above describes FP8 *Storage*, which changes how weights are **stored** while the math still runs in BF16. `fp8_compute` is a separate, global setting in `invokeai.yaml` that also does the **math** in FP8, on GPUs that have hardware for it. That makes generation faster, not just smaller.

It is for models that were **already saved in FP8** by whoever published them — you'll often see these labelled "fp8" or "fp8_scaled" in the filename. Normally InvokeAI unpacks them back to BF16 while loading; with `fp8_compute` on, they stay as they are and run directly on the GPU's FP8 hardware. FP8 Storage is the toggle for the other case: a full-precision model you want to shrink yourself.

Not every part of a model stays in FP8 — the small, precision-sensitive pieces are always unpacked. Those are a tiny share of the weights, so you still get nearly the full VRAM saving.

You need an FP8-capable GPU: RTX 40-series or newer, the datacenter cards of those generations, or AMD MI300 and newer. InvokeAI checks your card the first time it needs to — by actually trying a small FP8 operation rather than going by the model name — so a card that can't do it quietly falls back to the normal path instead of failing partway through a generation.

:::danger[Reproducibility requires the model to be fully in VRAM]
With `fp8_compute` enabled, **the same seed only reproduces the same image if the model is 100% resident in VRAM.**

FP8 math only happens for the parts of the model that are actually on the GPU. Anything still sitting in system RAM takes the normal path instead, which gives slightly different numbers — and *which* parts that affects depends on how much of the model happened to fit at that moment, which changes from run to run.

Measured on a 24 GB card with a ~12 GB model at 88–95% loaded: two runs with an identical seed and identical settings differed in **98.7% of all pixels**. With the model fully loaded, repeated runs came out **identical**.

To get repeatable output, check the model's load line in the log for `VRAM: … (100.0%)`. If it is below that, free up VRAM (lower resolution, fewer models loaded at once, a smaller text encoder) or set `enable_partial_loading: false`.
:::

Keeping the whole model on the GPU is worth it for speed as well: on the same 24 GB card, having to stream the last 5–12% of the model over PCIe cost **+47% per step** (1.03 → 1.51 s/it at 1024², 8 steps).

### When the model asks for full precision

Some FP8 models come with a note from whoever made them, marking certain layers as ones that should not use FP8 math. InvokeAI follows those notes by default, so those layers run the slower way. On a model that marks a lot of layers, this can eat much of the FP8 Compute speedup.

Setting `fp8_compute_full_precision_hints: false` ignores the notes and runs everything on the FP8 hardware. It is faster, but you are overriding the model author's judgement about which layers are sensitive — so compare a few images before sticking with it.

## Troubleshooting

### "I toggled FP8 Storage but VRAM usage didn't change"
Expand Down
22 changes: 22 additions & 0 deletions docs/src/generated/settings.json
Original file line number Diff line number Diff line change
Expand Up @@ -468,6 +468,28 @@
"type": "<class 'bool'>",
"validation": {}
},
{
"category": "CACHE",
"default": false,
"description": "Keep ComfyUI 'scaled fp8' checkpoints quantized instead of dequantizing them at load, and run their matmuls on the fp8 tensor cores (requires an Ada/SM 8.9 or newer NVIDIA GPU; falls back automatically otherwise). Roughly halves the transformer's VRAM and speeds up denoising, but quantizes activations as well, so images will differ from previous versions at the same seed. Reproducibility also requires the model to be FULLY resident in VRAM: a layer whose weights are still in RAM falls back to the dequantized path, and since which layers are resident shifts from run to run, the same seed then yields visibly different images. For repeatable output, ensure the model loads at 100% (e.g. enable_partial_loading=false with enough free VRAM).",
"env_var": "INVOKEAI_FP8_COMPUTE",
"literal_values": [],
"name": "fp8_compute",
"required": false,
"type": "<class 'bool'>",
"validation": {}
},
{
"category": "CACHE",
"default": true,
"description": "Honor the per-layer 'full_precision_matrix_mult' flags that some scaled-fp8 checkpoints ship. Those layers then dequantize on every forward instead of using the fp8 tensor cores, which can cost a large part of the fp8_compute speedup - on checkpoints that mark many layers, most of it. Set to false to run every quantized layer on the fp8 tensor cores, ignoring the producer's instruction; faster, but the marked layers were flagged as numerically sensitive, so quality may suffer. Only has an effect when fp8_compute is enabled.",
"env_var": "INVOKEAI_FP8_COMPUTE_FULL_PRECISION_HINTS",
"literal_values": [],
"name": "fp8_compute_full_precision_hints",
"required": false,
"type": "<class 'bool'>",
"validation": {}
},
{
"category": "CACHE",
"default": null,
Expand Down
4 changes: 4 additions & 0 deletions invokeai/app/services/config/config_default.py
Original file line number Diff line number Diff line change
Expand Up @@ -106,6 +106,8 @@ class InvokeAIAppConfig(BaseSettings):
device_working_mem_gb: The amount of working memory to keep available on the compute device (in GB). Has no effect if running on CPU. If you are experiencing OOM errors, try increasing this value.
enable_partial_loading: Enable partial loading of models. This enables models to run with reduced VRAM requirements (at the cost of slower speed) by streaming the model from RAM to VRAM as its used. In some edge cases, partial loading can cause models to run more slowly if they were previously being fully loaded into VRAM.
keep_ram_copy_of_weights: Whether to keep a full RAM copy of a model's weights when the model is loaded in VRAM. Keeping a RAM copy increases average RAM usage, but speeds up model switching and LoRA patching (assuming there is sufficient RAM). Set this to False if RAM pressure is consistently high.
fp8_compute: Keep ComfyUI 'scaled fp8' checkpoints quantized instead of dequantizing them at load, and run their matmuls on the fp8 tensor cores (requires an Ada/SM 8.9 or newer NVIDIA GPU; falls back automatically otherwise). Roughly halves the transformer's VRAM and speeds up denoising, but quantizes activations as well, so images will differ from previous versions at the same seed. Reproducibility also requires the model to be FULLY resident in VRAM: a layer whose weights are still in RAM falls back to the dequantized path, and since which layers are resident shifts from run to run, the same seed then yields visibly different images. For repeatable output, ensure the model loads at 100% (e.g. enable_partial_loading=false with enough free VRAM).
fp8_compute_full_precision_hints: Honor the per-layer 'full_precision_matrix_mult' flags that some scaled-fp8 checkpoints ship. Those layers then dequantize on every forward instead of using the fp8 tensor cores, which can cost a large part of the fp8_compute speedup - on checkpoints that mark many layers, most of it. Set to false to run every quantized layer on the fp8 tensor cores, ignoring the producer's instruction; faster, but the marked layers were flagged as numerically sensitive, so quality may suffer. Only has an effect when fp8_compute is enabled.
ram: DEPRECATED: This setting is no longer used. It has been replaced by `max_cache_ram_gb`, but most users will not need to use this config since automatic cache size limits should work well in most cases. This config setting will be removed once the new model cache behavior is stable.
vram: DEPRECATED: This setting is no longer used. It has been replaced by `max_cache_vram_gb`, but most users will not need to use this config since automatic cache size limits should work well in most cases. This config setting will be removed once the new model cache behavior is stable.
lazy_offload: DEPRECATED: This setting is no longer used. Lazy-offloading is enabled by default. This config setting will be removed once the new model cache behavior is stable.
Expand Down Expand Up @@ -210,6 +212,8 @@ class InvokeAIAppConfig(BaseSettings):
device_working_mem_gb: float = Field(default=3, description="The amount of working memory to keep available on the compute device (in GB). Has no effect if running on CPU. If you are experiencing OOM errors, try increasing this value.")
enable_partial_loading: bool = Field(default=True, description="Enable partial loading of models. This enables models to run with reduced VRAM requirements (at the cost of slower speed) by streaming the model from RAM to VRAM as its used. In some edge cases, partial loading can cause models to run more slowly if they were previously being fully loaded into VRAM.")
keep_ram_copy_of_weights: bool = Field(default=True, description="Whether to keep a full RAM copy of a model's weights when the model is loaded in VRAM. Keeping a RAM copy increases average RAM usage, but speeds up model switching and LoRA patching (assuming there is sufficient RAM). Set this to False if RAM pressure is consistently high.")
fp8_compute: bool = Field(default=False, description="Keep ComfyUI 'scaled fp8' checkpoints quantized instead of dequantizing them at load, and run their matmuls on the fp8 tensor cores (requires an Ada/SM 8.9 or newer NVIDIA GPU; falls back automatically otherwise). Roughly halves the transformer's VRAM and speeds up denoising, but quantizes activations as well, so images will differ from previous versions at the same seed. Reproducibility also requires the model to be FULLY resident in VRAM: a layer whose weights are still in RAM falls back to the dequantized path, and since which layers are resident shifts from run to run, the same seed then yields visibly different images. For repeatable output, ensure the model loads at 100% (e.g. enable_partial_loading=false with enough free VRAM).")
fp8_compute_full_precision_hints: bool = Field(default=True, description="Honor the per-layer 'full_precision_matrix_mult' flags that some scaled-fp8 checkpoints ship. Those layers then dequantize on every forward instead of using the fp8 tensor cores, which can cost a large part of the fp8_compute speedup - on checkpoints that mark many layers, most of it. Set to false to run every quantized layer on the fp8 tensor cores, ignoring the producer's instruction; faster, but the marked layers were flagged as numerically sensitive, so quality may suffer. Only has an effect when fp8_compute is enabled.")
# Deprecated CACHE configs
ram: Optional[float] = Field(default=None, gt=0, description="DEPRECATED: This setting is no longer used. It has been replaced by `max_cache_ram_gb`, but most users will not need to use this config since automatic cache size limits should work well in most cases. This config setting will be removed once the new model cache behavior is stable.")
vram: Optional[float] = Field(default=None, ge=0, description="DEPRECATED: This setting is no longer used. It has been replaced by `max_cache_vram_gb`, but most users will not need to use this config since automatic cache size limits should work well in most cases. This config setting will be removed once the new model cache behavior is stable.")
Expand Down
28 changes: 27 additions & 1 deletion invokeai/backend/model_manager/load/load_default.py
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
import re
from logging import Logger
from pathlib import Path
from typing import Optional
from typing import Callable, Optional

import torch

Expand All @@ -27,6 +27,7 @@
AnyModel,
SubModelType,
)
from invokeai.backend.quantization.fp8_scaled import count_fp8_weights, should_keep_fp8_weights
from invokeai.backend.util.devices import TorchDevice
from invokeai.backend.util.fp8 import FP8_COMPUTE_DTYPE_ATTR, set_fp8_compute_dtype

Expand Down Expand Up @@ -472,6 +473,22 @@ def _apply_fp8_layerwise_casting(
if isinstance(model, torch.nn.Module) and getattr(model, FP8_COMPUTE_DTYPE_ATTR, None) is not None:
return model

# A checkpoint that already ships fp8 weights is running (or is about to run) on the fp8
# tensor cores. Layerwise casting would install hooks that restore the compute dtype before
# every forward, so `CustomLinear._can_use_fp8_matmul` would no longer see an fp8 weight and
# would silently fall back to the dequantized path — the VRAM toggle would make the model
# *slower* with no indication why. Storage has nothing to add here anyway: the weights are
# already 1 byte per parameter.
if isinstance(model, torch.nn.Module) and should_keep_fp8_weights(self._torch_device):
already_fp8 = count_fp8_weights(model)
if already_fp8:
self._logger.info(
f"FP8 storage skipped for {config.name}: {already_fp8} weight(s) are already fp8 and "
"are being run on the fp8 tensor cores (fp8_compute). Layerwise casting would "
"disable that matmul without saving any further VRAM."
)
return model

storage_dtype = torch.float8_e4m3fn
compute_dtype = self._torch_dtype

Expand Down Expand Up @@ -517,6 +534,7 @@ def _apply_fp8_to_nn_module(
storage_dtype: torch.dtype,
compute_dtype: torch.dtype,
extra_skip_patterns: tuple[str, ...] = (),
skip: Optional[Callable[[str, torch.nn.Module], bool]] = None,
) -> None:
"""Apply FP8 layerwise casting to a plain nn.Module.

Expand All @@ -530,6 +548,12 @@ def _apply_fp8_to_nn_module(
`_model_declared_skip_patterns`), which are model-specific and cannot be inferred from
layer types or generic name patterns.

`skip` excludes further modules by (dotted name, module). Its one caller uses it to leave
scaled-fp8 layers alone: those already hold fp8 weights plus a `weight_scale`, and the cast
hooks installed here would upcast them *without* applying that scale — a silently wrong
weight. Casting only the remainder lets a partly-quantized checkpoint (fp8 language model,
bf16 visual tower) end up fully fp8-resident.

Modules holding already-quantized weights are skipped regardless of their class. This is a
backstop behind the format check in `_should_use_fp8`, which cannot see quantization that
is not reflected in the model's format (e.g. a `diffusers`-format checkpoint whose weights
Expand All @@ -548,6 +572,8 @@ def _apply_fp8_to_nn_module(
continue
if any(re.search(pattern, module_name) for pattern in skip_patterns):
continue
if skip is not None and skip(module_name, module):
continue
params = list(module.parameters(recurse=False))
if not params:
continue
Expand Down
Loading
Loading