Skip to content

[LoRA] add LoKr adapter support (Z-Image, Flux2/Klein) - #3

Open
christopher5106 wants to merge 9 commits into
mainfrom
lokr_support
Open

christopher5106 wants to merge 9 commits into
mainfrom
lokr_support

Conversation

@christopher5106

Copy link
Copy Markdown

Upstream: huggingface#14163 (open since 2026-07-10). Both review comments were addressed in 96bb95c on 2026-09-26, and a re-review was requested.

Problem: diffusers only loads LoRA adapters, so LoKr (LyCORIS Kronecker-product) checkpoints for Z-Image and FLUX.2 / Klein cannot be loaded.

Change:

  • load_lora_adapter detects lokr_ keys and injects a peft LoKrConfig inferred from the tensor shapes.
  • Converters cover ai-toolkit Z-Image, ai-toolkit BFL FLUX.2 (fused QKV, loaded by fusing the model's QKV so the Kronecker factors map one to one), and the LyCORIS underscore and dotted-diffusers layouts. The LyCORIS alpha convention is applied at conversion.
  • Review follow-ups: one adapter-agnostic error message, and a fused-QKV load is refused when an adapter already sits on the unfused projections.

Test: tests/lora/test_lora_lokr.py.

In scenario: yes, cherry-picked as one commit. The branch has two commits, to squash at the next re-sync.

This PR documents a fix branch of this fork and is not meant to be merged. It is closed once the change is merged upstream, and the branch is then deleted.

Adds loading of LoKr (LyCORIS Kronecker product) adapters:

- `load_lora_adapter` detects `lokr_` keys and injects a peft `LoKrConfig`,
  inferred from the tensor shapes via `_create_lokr_config` (decompose
  factor, per-module rank/alpha patterns).
- State dict conversions for the formats in the wild: ai-toolkit Z-Image
  (dotted diffusers paths under `diffusion_model.`), ai-toolkit BFL Flux2
  (fused qkv), LyCORIS underscore format, and bare dotted diffusers paths.
- BFL fused-QKV LoKr cannot be split exactly into separate Q/K/V Kronecker
  factors, so `Flux2LoraLoaderMixin.load_lora_weights` fuses the model's
  QKV projections and maps the adapter 1:1 (exact).
- Alpha follows the LyCORIS convention: scaling applies only to
  rank-decomposed factors and is baked into the weights at conversion.

Fixes huggingface#13221
sayakpaul and others added 4 commits October 6, 2026 17:34
* implement Kandinsky 6
 - add TI2VA and SR pipelines
 - additonally implement PiflowScheduler and MMAudioVAE

* update docs for Kandinsky 6

* remove redundant bigvgan code and manual checkpoint loading inside model
code

* add missing _no_split_modules

* use attention processor in SR

* fix unguarded dependencies
 - removed: einops, pydantic
 - lazy: torchvision, av, librosa

* remove dead manual weight loading code; flattent MMAudioVAE structure

* move attention processors to transformer files

* remove duplicate code in K6 latent_upscaler

* address 1st batch of comments

* patch transformer k6

* patch SR

* patch pipeline style

* replace TimestepEmbedding and fix batching

* refactor ti2va dit and pipeline

* refactor piflow scheduler

* refactor SR

* patch sr transformer autocast

* major refactor

* untrack notebooks directory

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* patch mmaudio VAE

* patch mmaudio VAE

* fix inits

* fix flex attn

* refactoring

* remove __all__

* separate MMAudio VAE and Vocoder

* fix naming

* refactor checkpointing, magcache, transformer, sr vae for K6

* refactor

* address apply_scale_shift_norm problem

* update docs

* apply style chanes

* address review feedback on transformer, SR VAE, and flex attention

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* address comments

* fix critical bugs

* Revert checkpoint-breaking module renames from 28c5326

Commit 28c5326 (today, 14:03 +0300) renamed Kandinsky6FusedTransformerDecoderBlock's
videoT/audioT submodules to video_dec_block/audio_dec_block, and replaced both
Kandinsky6TimeEmbeddings and Kandinsky6SRTimeEmbeddings's hand-rolled in_layer/out_layer
with diffusers' Timesteps/TimestepEmbedding, purely for naming clarity. No checkpoint
conversion was re-run to match, so every published K6 checkpoint (unchanged since Sep 27,
confirmed via identical blob SHA256 across recent Hub commits) still ships the old names.

Verified against kandinskylab/Kandinsky-6.0-Pro-distill-5s-Diffusers and
kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers: this was silently dropping ~65% of the main
transformer's weights (every visual_transformer_blocks.*.{video,audio}T.* parameter) and
the SR transformer's time embeddings, then crashing with "Cannot copy out of meta tensor"
on enable_model_cpu_offload(). After this revert, all checkpoint keys match exactly
(4471/4471 for the main transformer, 459/459 for the SR transformer).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Revert checkpoint-breaking latent-upscaler renames from 372f290

Same class of regression as 9ae71f0, this time in Kandinsky6SRLatentUpscalerX2Branch
and Kandinsky6SRLatentUpscalerOutputHead: today's "address comments" commit (18:48 +0300)
dropped the private_ prefix from private_mid_blocks/private_upsample/private_blocks/
private_output_proj, and converted OutputHead from nn.Sequential to a plain nn.Module with
named norm/activation/conv submodules, again without a matching checkpoint re-conversion.

Verified against kandinskylab/Kandinsky-6.0-VSR-5s-Diffusers's latent_upscaler: this left
every Kandinsky6SRLatentUpscalerX2Branch parameter (and the OutputHead-based output_proj
used standalone by the x2/x4 branches) unfilled on the meta device, crashing
enable_model_cpu_offload() the same way as the main transformer did before 9ae71f0.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* move vocoder and LU

* add coppied from resolve i2v cond mode

* rename variables

* fix lu rename bug

* Update src/diffusers/pipelines/kandinsky6/pipeline_kandinsky6_ti2va.py

Co-authored-by: YiYi Xu <yixu310@gmail.com>

* Update src/diffusers/models/transformers/transformer_kandinsky6.py

Co-authored-by: YiYi Xu <yixu310@gmail.com>

* address yiyixuxu's comments

* fix audio channels bug

* fix tests

* Fix CI failures on the Kandinsky6 PR

- tests/others/test_utils.py: assert the caller file path without assuming the checkout directory is named `diffusers`.
- .github/workflows/pr_tests.yml: raise PYTEST_TIMEOUT to 300 s for the example and pipeline CPU steps. Both run several workers with four threads each on an 8-vCPU runner, and the checkpointing example tests run several training subprocesses in sequence.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Avoid staging LFS dedup race in the org push test

The second push in TestModelPushToHub.test_push_to_hub_in_organization sent the same model bytes as the first push. The staging server can deduplicate those bytes against the first push's LFS object and then reject the commit with "LFS pointer pointed to a file that does not exist". Change the weights before the second push so its bytes differ.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Revert "Avoid staging LFS dedup race in the org push test"

This reverts commit 35e6cc4.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Revert "Fix CI failures on the Kandinsky6 PR"

This reverts commit 60cad43.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

* Apply suggestion from @yiyixuxu

* Apply suggestion from @yiyixuxu

* Apply batched suggestions from code review

Co-authored-by: YiYi Xu <yixu310@gmail.com>

* remove review-only files from the Kandinsky6 PR

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCdbvRpL9fv3h3WwSUPpfS

* Apply suggestion from @yiyixuxu

* restore .gitignore to match main

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCdbvRpL9fv3h3WwSUPpfS

* simplify the mono audio handling in _write_audio

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCdbvRpL9fv3h3WwSUPpfS

* remove the no-op set_attention_backend calls from the Kandinsky6 docs

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCdbvRpL9fv3h3WwSUPpfS

* give the Kandinsky6 transformer its own output class instead of reusing LTX-2's

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCdbvRpL9fv3h3WwSUPpfS

* fix the stale pipeline exports and add return annotations for the Kandinsky6 pipelines

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HCdbvRpL9fv3h3WwSUPpfS

* address feedback from @leffff

---------

Co-authored-by: Denis Koposov <denis.koposov@phystech.edu>
Co-authored-by: leffff <levnovitskiy@gmail.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Lev Novitskiy <57654885+leffff@users.noreply.github.com>
Co-authored-by: YiYi Xu <yixu310@gmail.com>
* Bump doc-builder pin to 9fc8a41 in build_documentation.yml

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Bump doc-builder pin to 9fc8a41 in build_pr_documentation.yml

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Bump doc-builder pin to 9fc8a41 in upload_pr_documentation.yml

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Sayak Paul <spsayakpaul@gmail.com>
… fused-QKV load over an unfused adapter

- "Invalid adapter checkpoint. We currently support LoRA and LoKr." replaces the
  message that still said "LoRA checkpoint" and described the substring check.
- Loading a fused-QKV LoKr checkpoint now refuses when an adapter is already
  injected on to_q/to_k/to_v or add_{q,k,v}_proj: fuse_qkv_projections() would
  replace those modules and orphan it. Covered by a test.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The file exercises checkpoint conversion and loading for LoKr adapters, not LoRA.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
stevhliu and others added 4 commits October 6, 2026 12:04
hf jobs

Co-authored-by: Sayak Paul <spsayakpaul@gmail.com>
…sed QKV, LoKrTesterMixin

- Only the Z-Image and Flux2 loader mixins accept LoKr checkpoints, the two that
  convert them; the other mixins are back to their LoRA-only check.
- The fused-QKV handling moves to `_maybe_fuse_qkv_projections_for_lokr` in
  peft_utils, called from `load_lora_adapter`, so the model fuses its projections
  whichever entry point loads the adapter. The quantized-model error explains why.
- `_bake_lokr_alpha` becomes `_bake_lokr_alpha_` as it edits the dict in place.
- The BFL Flux2 converter no longer accepts expanded diffusers block names under
  `diffusion_model.`: no LoKr checkpoint uses that layout; diffusers-named
  checkpoints have no prefix and take the generic converter.
- The LyCORIS converter names the keys it does not recognize.
- Tests move from tests/lora to a `LoKrTesterMixin` in tests/models/testing_utils,
  with the checkpoint formats tested on the Flux2 and Z-Image model test classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation pipelines models CI schedulers labels Oct 8, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants