Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion .github/workflows/build_documentation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -17,7 +17,7 @@ permissions:

jobs:
build:
uses: huggingface/doc-builder/.github/workflows/build_main_documentation.yml@e60a538eea9817ab312196d0d233604b01697265 # main
uses: huggingface/doc-builder/.github/workflows/build_main_documentation.yml@9fc8a41d7729c424fdb2192f2f85ff63102ec9a3 # main
with:
commit_sha: ${{ github.sha }}
install_libgl1: true
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/build_pr_documentation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -44,7 +44,7 @@ jobs:
build:
needs: check-links
uses: huggingface/doc-builder/.github/workflows/build_pr_documentation.yml@e60a538eea9817ab312196d0d233604b01697265 # main
uses: huggingface/doc-builder/.github/workflows/build_pr_documentation.yml@9fc8a41d7729c424fdb2192f2f85ff63102ec9a3 # main
with:
commit_sha: ${{ github.event.pull_request.head.sha }}
pr_number: ${{ github.event.number }}
Expand Down
2 changes: 1 addition & 1 deletion .github/workflows/upload_pr_documentation.yml
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@ permissions:

jobs:
build:
uses: huggingface/doc-builder/.github/workflows/upload_pr_documentation.yml@9ad2de8582b56c017cb530c1165116d40433f1c6 # main
uses: huggingface/doc-builder/.github/workflows/upload_pr_documentation.yml@9fc8a41d7729c424fdb2192f2f85ff63102ec9a3 # main
with:
package_name: diffusers
secrets:
Expand Down
30 changes: 16 additions & 14 deletions docs/source/en/_toctree.yml
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,8 @@
title: Schedulers
- local: using-diffusers/weighted_prompts
title: Prompting
- local: using-diffusers/image_quality
title: FreeU
- local: using-diffusers/reusing_seeds
title: Reproducibility
- local: using-diffusers/callback
Expand Down Expand Up @@ -68,24 +70,14 @@
title: Inference
- isExpanded: false
sections:
- local: optimization/pruna
title: Pruna
- local: optimization/xformers
title: xFormers
- local: optimization/tome
title: Token merging
- local: optimization/deepcache
title: DeepCache
- local: optimization/cache_dit
title: CacheDiT
- local: optimization/tgate
title: TGATE
- local: optimization/xdit
title: xDiT
- local: optimization/para_attn
title: ParaAttention
- local: using-diffusers/image_quality
title: FreeU
- local: optimization/pruna
title: Pruna
- local: optimization/xdit
title: xDiT
title: Community methods
- isExpanded: false
sections:
Expand Down Expand Up @@ -210,6 +202,8 @@
title: Training methods
- local: training/nemo_automodel
title: NeMo Automodel
- local: training/jobs
title: Hugging Face Jobs
title: Train and fine-tune
- isExpanded: false
sections:
Expand Down Expand Up @@ -375,6 +369,8 @@
title: JoyImageEditPlusTransformer3DModel
- local: api/models/transformer_joyimage
title: JoyImageEditTransformer3DModel
- local: api/models/kandinsky6_transformer
title: Kandinsky 6 Transformers
- local: api/models/krea2_transformer2d
title: Krea2Transformer2DModel
- local: api/models/latte_transformer3d
Expand Down Expand Up @@ -501,6 +497,8 @@
title: AutoencoderSAME
- local: api/models/consistency_decoder_vae
title: ConsistencyDecoderVAE
- local: api/models/kandinsky6_vae
title: Kandinsky 6 VAEs
- local: api/models/ltx2_diffusion_decoder
title: LTX2VideoDiffusionDecoderModel
- local: api/models/autoencoder_oobleck
Expand Down Expand Up @@ -725,6 +723,8 @@
title: HunyuanVideo1.5
- local: api/pipelines/kandinsky5_video
title: Kandinsky 5.0 Video
- local: api/pipelines/kandinsky6
title: Kandinsky 6
- local: api/pipelines/latte
title: Latte
- local: api/pipelines/ltx2
Expand Down Expand Up @@ -816,6 +816,8 @@
title: LMSDiscreteScheduler
- local: api/schedulers/minimax_h3
title: MiniMaxH3Scheduler
- local: api/schedulers/piflow
title: PiflowScheduler
- local: api/schedulers/pndm
title: PNDMScheduler
- local: api/schedulers/repaint
Expand Down
53 changes: 53 additions & 0 deletions docs/source/en/api/models/kandinsky6_transformer.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,53 @@
<!--Copyright 2026 The Kandinsky Team and The HuggingFace Team. All rights reserved.
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with
the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on
an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the
specific language governing permissions and limitations under the License.
-->

# Kandinsky 6 Transformers

Kandinsky 6 uses a multimodal diffusion transformer that denoises video and audio latents together for
text/image-to-video-and-audio generation, and a text-free diffusion transformer for video super-resolution.

## Kandinsky6Transformer3DModel

The multimodal transformer used by [`Kandinsky6TI2VAPipeline`].

```python
import torch
from diffusers import Kandinsky6Transformer3DModel

transformer = Kandinsky6Transformer3DModel.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", subfolder="transformer", torch_dtype=torch.bfloat16
)
```

[[autodoc]] Kandinsky6Transformer3DModel
- all
- forward

## Kandinsky6SRTransformer3DModel

The text-free transformer used by [`Kandinsky6SRPipeline`] to refine one tile of the upscaled video at a time.

```python
import torch
from diffusers import Kandinsky6SRTransformer3DModel

transformer = Kandinsky6SRTransformer3DModel.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", subfolder="transformer", torch_dtype=torch.bfloat16
)
# The transformer always runs NABLA sparse attention (`nabla_threshold`, 0.8 by default) on the `flex` backend.
# Compile it, otherwise flex falls back to an eager implementation that needs far more memory at video resolutions.
transformer.compile_repeated_blocks(fullgraph=True)
```

[[autodoc]] Kandinsky6SRTransformer3DModel
- all
- forward
66 changes: 66 additions & 0 deletions docs/source/en/api/models/kandinsky6_vae.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
<!--Copyright 2026 The Kandinsky Team and The HuggingFace Team. All rights reserved.
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with
the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on
an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the
specific language governing permissions and limitations under the License.
-->

# Kandinsky 6 VAEs

Kandinsky 6 uses a causal 3D K-VAE for video super-resolution and the MMAudio mel-spectrogram VAE, paired with a
separate BigVGAN [`MMAudioVocoder`], for synchronized audio generation.

## Kandinsky6SRVAE

The causal 3D K-VAE used by [`Kandinsky6SRPipeline`]. It processes arbitrarily long videos in bounded-memory
segments while reproducing the exact output of a single, non-segmented pass.

```python
import torch
from diffusers import Kandinsky6SRVAE

vae = Kandinsky6SRVAE.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers", subfolder="vae", torch_dtype=torch.bfloat16
)
```

[[autodoc]] Kandinsky6SRVAE
- encode
- decode
- all

## MMAudioVAE

The mel-spectrogram VAE used by [`Kandinsky6TI2VAPipeline`] when `sample_audio=True`. Its `decode` output is a mel
spectrogram; pass it through [`MMAudioVocoder`] to get a waveform.

The reference implementation can be found at [hkchengrex/MMAudio](https://github.com/hkchengrex/MMAudio) (MIT
license).

```python
import torch
from diffusers import MMAudioVAE

audio_vae = MMAudioVAE.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", subfolder="audio_vae", torch_dtype=torch.bfloat16
)
```

[[autodoc]] MMAudioVAE
- encode
- decode
- all

## MMAudioVocoder

Adapted from the BigVGAN-v2 vocoder MMAudio bundles, itself from
[NVIDIA/BigVGAN](https://github.com/NVIDIA/BigVGAN) (MIT license), with the anti-aliased Snake activations of
[alias-free-torch](https://github.com/junjun3518/alias-free-torch) (Apache License 2.0).

[[autodoc]] MMAudioVocoder
- forward
2 changes: 1 addition & 1 deletion docs/source/en/api/pipelines/kandinsky.md
Original file line number Diff line number Diff line change
Expand Up @@ -712,7 +712,7 @@ make_image_grid([img.resize((512, 512)), image.resize((512, 512))], rows=1, cols

Kandinsky is unique because it requires a prior pipeline to generate the mappings, and a second pipeline to decode the latents into an image. Optimization efforts should be focused on the second pipeline because that is where the bulk of the computation is done. Here are some tips to improve Kandinsky during inference.

1. Enable [xFormers](../../optimization/xformers) if you're using PyTorch < 2.0:
1. Enable [xFormers](../../optimization/attention_backends) if you're using PyTorch < 2.0:

```diff
from diffusers import DiffusionPipeline
Expand Down
150 changes: 150 additions & 0 deletions docs/source/en/api/pipelines/kandinsky6.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,150 @@
<!--Copyright 2026 The Kandinsky Team and The HuggingFace Team. All rights reserved.
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with
the License. You may obtain a copy of the License at
http://www.apache.org/licenses/LICENSE-2.0
Unless required by applicable law or agreed to in writing, software distributed under the License is distributed on
an "AS IS" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the License for the
specific language governing permissions and limitations under the License.
-->

# Kandinsky 6

Kandinsky 6 is a family of video generation models from [Kandinsky Lab](https://huggingface.co/kandinskylab). The
main model generates video and synchronized audio from text or a reference image with a single multimodal diffusion
transformer: video and audio latents are denoised together through fused blocks that cross-attend between the two
modalities, each conditioned on its own Qwen2.5-VL text branch and a CLIP pooled embedding. A separate
super-resolution model upscales the generated video tile by tile in the latent space of a causal 3D K-VAE.

> [!TIP]
> Check out the [Kandinsky Lab](https://huggingface.co/kandinskylab) organization on the Hub for the full set of
> official checkpoints, including flow-matching and distilled variants of both the base and super-resolution models.
>
> Distilled checkpoints ship with the few-step [`PiflowScheduler`] and must be run with `guidance_scale=1.0`.
## Available models

| Model | Pipeline | Notes |
|---|---|---|
| [`kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-5s-Diffusers) | [`Kandinsky6TI2VAPipeline`] | Flow matching, `guidance_scale=5.0`, 50 steps |
| [`kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers) | [`Kandinsky6TI2VAPipeline`] | Distilled, `guidance_scale=1.0`, 10 steps |
| [`kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-Lite-5s-Diffusers) | [`Kandinsky6TI2VAPipeline`] | Flow matching, `guidance_scale=5.0`, 50 steps |
| [`kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-Lite-distill-5s-Diffusers) | [`Kandinsky6TI2VAPipeline`] | Distilled, `guidance_scale=1.0`, 10 steps |
| [`kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-Pro-pretrain-5s-Diffusers) | [`Kandinsky6TI2VAPipeline`] | Flow matching, `guidance_scale=5.0`, 50 steps |
| [`kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-Lite-pretrain-5s-Diffusers) | [`Kandinsky6TI2VAPipeline`] | Flow matching, `guidance_scale=5.0`, 50 steps |
| [`kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers) | [`Kandinsky6SRPipeline`] | Flow matching super-resolution |
| [`kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers`](https://huggingface.co/kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers) | [`Kandinsky6SRPipeline`] | Distilled super-resolution, 2 steps |

## Text/image-to-video-and-audio

```python
import torch
from diffusers import Kandinsky6TI2VAPipeline
from diffusers.utils import encode_video

pipe = Kandinsky6TI2VAPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()

output = pipe(
prompt="A cat and a dog baking a cake together in a kitchen.",
height=480,
width=864,
num_frames=121,
num_inference_steps=10,
guidance_scale=1.0,
)
encode_video(
output.frames[0],
fps=24,
output_path="output.mp4",
audio=output.audio[0][None],
audio_sample_rate=pipe.audio_sample_rate,
)
```

Pass `image=` to condition the first frame on a reference image, `sample_audio=False` to generate video only, and
`expand_prompts=True` to let the Qwen2.5-VL text encoder rewrite short prompts into detailed ones first.

## Video super-resolution

[`Kandinsky6SRPipeline`] takes the frames produced by [`Kandinsky6TI2VAPipeline`] and upscales them by `2`, `4`, or
`2.25` (a 1.125x bilinear pre-upscale followed by the 2x path). The video is split into overlapping tiles, every tile
is refined at one of the tile sizes the SR transformer was trained on, and the tiles are blended back with Hann
windows.

```python
# required: lets inductor pick flex-attention tiles that fit the SR block mask
torch._inductor.config.max_autotune = True

sr_pipe = Kandinsky6SRPipeline.from_pretrained(
"kandinskylab/Kandinsky-6.0-VSR-distilled2steps-5s-Diffusers", torch_dtype=torch.bfloat16
)
# The SR transformer always runs NABLA sparse attention on the `flex` backend. Compile it, otherwise flex falls
# back to an eager implementation that needs far more memory at video resolutions.
sr_pipe.enable_model_cpu_offload()
sr_pipe.transformer.set_attention_backend("flex")
sr_pipe.transformer.compile_repeated_blocks(fullgraph=True)

upscaled = sr_pipe(video=output.frames[0], resolution_scale=2.25, num_inference_steps=2).frames[0]
```

## Memory optimization

Refer to the [Reduce memory usage](../../optimization/memory) guide for the general set of techniques. Both
[`Kandinsky6TI2VAPipeline`] and [`Kandinsky6SRPipeline`] support [model offloading](../../optimization/memory#model-offloading)
(used above) and, for a smaller footprint at the cost of speed, [sequential CPU offloading](../../optimization/memory#cpu-offloading):

```python
pipe.enable_sequential_cpu_offload()
```

[`Kandinsky6TI2VAPipeline`]'s video VAE also supports [tiled decoding](../../optimization/memory#vae-tiling) for high
resolutions or long videos:

```python
pipe.vae.enable_tiling()
```

## Notes

- `height` and `width` must be divisible by the video VAE's spatial compression ratio times the transformer's patch
size — `16` with the default [`Kandinsky6TI2VAPipeline`] configuration (`AutoencoderKLHunyuanVideo` at a
compression ratio of `8`, `patch_size=(1, 2, 2)`). `480x864`, used in the example above, satisfies this.
- [`Kandinsky6SRPipeline`]'s input `video` must have `1 + k * vae_scale_factor_temporal` frames for some integer `k`
(a temporal compression ratio of `4` with the default K-VAE configuration, so `121` frames works but `120` doesn't)
— trim or pad a video that doesn't already satisfy this before upscaling it.
- `expand_prompts=True` reuses the already-loaded Qwen2.5-VL text encoder for an extra generation pass before
denoising, so it adds latency but no extra model weights.
- Compile the repeated transformer blocks for faster repeated inference:
```python
pipe.transformer.compile_repeated_blocks(fullgraph=True)
```

## Kandinsky6TI2VAPipeline

[[autodoc]] Kandinsky6TI2VAPipeline
- all
- __call__

## Kandinsky6SRPipeline

[[autodoc]] Kandinsky6SRPipeline
- all
- __call__

## Kandinsky6SRLatentUpscalerBank

[[autodoc]] Kandinsky6SRLatentUpscalerBank
- forward

## Kandinsky6TI2VAPipelineOutput

[[autodoc]] pipelines.kandinsky6.pipeline_output.Kandinsky6TI2VAPipelineOutput

## Kandinsky6SRPipelineOutput

[[autodoc]] pipelines.kandinsky6.pipeline_output.Kandinsky6SRPipelineOutput
Original file line number Diff line number Diff line change
Expand Up @@ -448,7 +448,7 @@ SDXL is a large model, and you may need to optimize memory to get it to run on y
+ refiner.unet = torch.compile(refiner.unet, mode="reduce-overhead", fullgraph=True)
```

3. Enable [xFormers](../../../optimization/xformers) to run SDXL if `torch<2.0`:
3. Enable [xFormers](../../../optimization/attention_backends) to run SDXL if `torch<2.0`:

```diff
+ base.enable_xformers_memory_efficient_attention()
Expand Down
Loading
Loading