Skip to content

[LTX-2.5] Keyframe-aware diffusion decoding and SDR-To-HDR seam keyframes - #14975

Closed
christopher5106 wants to merge 4 commits into
huggingface:mainfrom
scenario-labs:fix_ltx2_hdr_keyframe_decoding
Closed

christopher5106 wants to merge 4 commits into
huggingface:mainfrom
scenario-labs:fix_ltx2_hdr_keyframe_decoding

Conversation

@christopher5106

Copy link
Copy Markdown
Contributor

Stacked on #14974 (SDR-To-HDR in LTX2HDRPipeline), itself stacked on #14966. Please review those first; this PR adds the last two commits: keyframe-aware diffusion decoding, then the seam keyframes in the pipeline.

What

The Lightricks reference HDRICLoraPipeline (ltx_pipelines/hdr_ic_lora.py, LTX-2 @ 9ec55f9f) adds seam keyframes by default (CLI --keyframe-strength 0.95). At each 24- or 32-frame DFR segment seam of the clip it appends two tokens:

  • a 1-frame SDR guide: the ACEScct source frame VAE-encoded on its own, held at strength 0.95;
  • an empty generated HDR slot at the same RoPE position, carrying the learned keyframe embedding.

The slots are denoised with the video, then anchor the keyframe-aware DiffVAE decode. This PR wires that into the hdr_transform="acescct" path:

  • keyframe_strength: float | None = 0.95. None runs the plain IC-LoRA. Clips with no seam (for example 9 or 17 frames) log a warning and run the plain IC-LoRA.
  • high_quality_hdr: bool = False: frame-doubled source, 2N - 1 generated frames, seam positions doubled, every second decoded frame kept.
  • Token order follows the reference: [video | reference | guides | slots]. Initial noising at noise_scale = sigmas[0] (1.0) gives the reference's result: clean reference tokens, pure noise on the video and slots, and 0.95 * clean + 0.05 * noise on the guides. Each step blends x0 with the clean tokens, and the guides' velocity is converted with their own per-token timestep, as the reference's X0 model does.
  • After denoising, the slots are extracted, denormalized like the video, and passed to LTX2VideoDiffusionDecoderModel.decode(..., keyframe_latents=, keyframe_frame_indices=) in float32.
  • Validation: keyframe_strength must be in [0, 1]. When the clip has seams, the transformer must set use_keyframes_abs_pos_embedding and there must be exactly one reference video. high_quality_hdr is rejected on LogC3.
  • The keyframe-coordinate and slot-unpack helpers are reused from the DFR pipeline through # Copied from.

The LogC3 (LTX-2.3) path ignores keyframe_strength and is bit-identical to before. keyframe_strength=None is bit-identical to PR 2.

Checkpoint note: the published LTX-2.5-Diffusers diffusion decoder

Keyframe-aware decoding needs the decoder weights that were trained with it. The original Lightricks/LTX-2.5 vae/ltx-2.5-video-vae-bf16.safetensors has them, including the keyframe tag decoder.type_emb, and the converter now carries the tag (previous commit). The diffusion_decoder currently published in Lightricks/LTX-2.5-Diffusers does not match that file:

  • it has no decoder.type_emb;
  • none of its 161 decoder tensors equals the current original file's (relative differences from 2–6% on conv_in/conv_out up to 10–40% on w_down, context_proj and upsamples.*). The latent statistics are identical.

The converter on main also cannot convert the current original VAE: its strict load fails on decoder.type_emb, which suggests the Hub copy was converted from an earlier file.

Decoder-only check on real weights (H200, 49 frames at 480x736, identical decoder inputs and identical noise on both sides, compared with Lightricks' decoder):

published -Diffusers decoder original VAE converted with this PR's converter
plain decode 44.4 dB 62.2 dB
keyframe decode (2 planes) 35.2 dB 68.6 dB

With matching weights, conv_in is bit-identical, decoder stages 1–4 agree to about 97 dB for both the video and keyframe streams, and the keyframe times agree at every stage. Re-converting the Hub diffusion_decoder from the current original VAE (convert_ltx2_diffusion_video_vae, which now sets decoder_keyframe_type_embedding=True) is what makes keyframe decoding, and plain decoding, match the reference.

Interaction with #14694

#14694 refactors the forward methods of the LTX-2.5 diffusion decoder. The previous commit adds keyframe variants of those forwards (*_with_keyframes, joint attention, per-tile plane selection), so whichever PR lands second needs a rebase. This PR only calls the public decode(z, generator=, keyframe_latents=, keyframe_frame_indices=) and does not depend on the internals.

Deviations from the reference

  • Decode generator: the reference seeds a fresh Generator(seed) for decoding. Here the decoder continues the user's generator, as PR 2 already does.
  • VAE tiling: the guides are tiled above 512x768 only when the user enabled vae.enable_tiling(), because diffusers tiling is opt-in, unlike the reference's AUTO_TILING. The full-clip encode keeps PR 2's vae.encode behaviour.
  • Guide denoise mask precision: diffusers uses a float32 1 - 0.95. The reference stores 1 - strength in the guide latent's dtype, which is bf16 at runtime (0.05005).
  • Keyframe embedding check: it only applies when the clip has seams. The reference also raises for seamless clips whenever keyframe_strength is set.
  • Without diffusion_decoder: the pipeline warns and decodes with the VAE, so the slots only take part in denoising. The reference always decodes with the keyframe-aware DiffVAE.
  • Noising form: this keeps PR 2's single-expression noising rather than the reference's two lerps. At noise_scale=1 the two are identical on the video, reference and slot tokens, and within 1 ulp on the guides.
  • Seam grid: seams are computed with the VAE's own temporal ratio, like LTX2DFRPipeline. This equals the reference's fixed x8 for every LTX VAE.
  • Latent output: output_type="latent" returns the video latents only, without the slots.

Tests

tests/pipelines/ltx2/test_ltx2_hdr_sdr_to_hdr.py now covers:

  • token layout, keyframe marker, per-token timesteps and shared coordinates for a one-seam clip;
  • the guide as a standalone 1-frame encode of the clip;
  • slots reaching decode, with indices and values;
  • keyframe_strength=None against a slice recorded on the parent commit;
  • the empty-seam fallback;
  • the warning without a diffusion decoder;
  • high_quality_hdr frame doubling, guide frame, decode and stride;
  • the tiled guide threshold;
  • input validation;
  • LogC3 unchanged.

On random tiny shapes with a stand-in encoder and velocity model, the following match the reference's own builders to within 6e-8:

  • the initial token sequence (same seed);
  • RoPE positions and the keyframe mask;
  • per-token timesteps;
  • the full 8-step Euler trajectory;
  • the extracted slots, for 25, 49 and 97 frames (60 fps capped to 30) and for HQ.

Video tokens and slots are bitwise equal. The builders used are VideoConditionByReferenceLatent, _keyframe_conditionings_from_pixel_frames, generated_keyframe_conditionings, create_noised_state, post_process_latent, EulerDiffusionStep, clear_conditioning and decode_keyframes_from_slots.

End to end on real weights (H200, hiker.mp4 from documentation-images, seed 42, seams at frames 24 and 48, against HDRICLoraPipeline with keyframe_strength=0.95): the final latents match at PSNR 52.4 dB, and token-level checks at 70 dB or better. The decoded output (34.8 dB with the published decoder) is dominated by the decoder-weight mismatch above.

🤖 Generated with Claude Code

christopher5106 and others added 4 commits October 6, 2026 17:44
…-HDR

Port the colour pipeline of the LTX-2.5 SDR-To-HDR IC-LoRA from the Lightricks
LTX-2 reference (ltx_core.hdr, ltx_core.color, ltx_pipelines media_io):

- image_processor: sRGB EOTF, Bradford Rec.709/AP1/Rec.2020 matrices, ACEScct
  encode/decode, the srgb_gamma/srgb/acescg/acescct input transforms and the
  ACEScct -> ACEScg/Rec.709 linear output transform, as pure torch functions.
  LTX2VideoHDRProcessor gains hdr_transform="acescct" with input_colorspace /
  output_colorspace options; the LogC3 default is unchanged.
- export_utils: encode_hdr_tensor_to_hlg_mp4 (BT.2020 HLG 10-bit HEVC via
  PyAV/libx265) and save_exr_frame / export_to_exr_sequence (half-float ZIP
  EXR with chromaticities and colorSpace), behind a new optional
  is_openexr_available() guard.
- tests: exact values computed with colour-science, round trips, HLG stream
  properties and EXR read-back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run the Lightricks LTX-2.5 SDR-To-HDR IC-LoRA (ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline,
without its optional seam keyframes) when the pipeline is built with hdr_transform="acescct":

- the reference video goes through the ACEScct input transform (new `input_colorspace`, default
  "srgb_gamma"), is reflect-padded up to a multiple of the VAE spatial compression ratio and
  VAE-encoded in float32; the decoded video is cropped back to height x width;
- conditioning comes from precomputed `connector_video_embeds` only (a 2D `video_context` is
  accepted): no prompt, no text encoder call, `connector_audio_embeds` optional;
- video-only denoising: `isolate_modalities=True` with a single placeholder audio token, so the
  audio stream cannot reach the video;
- the distilled schedule used verbatim (DISTILLED_SIGMA_VALUES, no mu shift), no CFG/STG/modality
  guidance, RoPE frame rate capped at 30 fps, first target latent frame marked for the keyframe
  position embedding;
- decoding in float32 with the optional `diffusion_decoder` component (LTX-2.5) or the VAE, then
  `postprocess_hdr_video(output_colorspace=...)` (new argument, default "rec709").

`hdr_transform` is now registered in the pipeline config so it survives save/load. The LogC3
(LTX-2.3) path is unchanged: its outputs are bit-identical to the parent commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port Lightricks' keyframe-aware DiffVAE decode (LTX-2 @ 9ec55f9f,
`ltx_core/model/video_vae/{keyframes,diffusion_video_decoder}.py`,
`transformer/fallback_na/joint_eager.py`) to `LTX2VideoDiffusionDecoderModel`.

- `decoder_keyframe_type_embedding` config flag (default off) creates the
  learned keyframe tag `decoder.type_emb` `(latent_channels,)`; existing
  checkpoints load unchanged.
- `decode(..., keyframe_latents=, keyframe_frame_indices=)`: planes are tagged,
  share `conv_in` and every stage with the video, and attend jointly with it
  (each video position sees its 2 nearest planes, each plane its 2 nearest
  frames, one softmax), with per-stage plane times. Tiled decode keeps the
  planes inside each temporal tile plus the nearest on each side, with times
  rebased on the tile. Without planes the decode is bit-identical to before.
- Joint attention is a port of the reference's pure-torch backend, bitwise
  equal to it; full-decoder parity on random tiny weights is within 2.4e-6.
- Converter carries `decoder.type_emb` and sets the flag when present (the
  current original checkpoint has it, so the strict load used to fail).
- `dfr_layout.resolve_seam_positions`: seam keyframes of a single-window clip
  (24/32 segments clipped to the clip, high-quality doubling), equal to the
  reference for every 8k+1 frame count up to 2001.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port the seam keyframes of Lightricks' `HDRICLoraPipeline` (LTX-2 @ 9ec55f9f, `ltx_pipelines/hdr_ic_lora.py`)
to the `hdr_transform="acescct"` path:

- `keyframe_strength` (default 0.95, `None` = plain IC-LoRA): seams from `resolve_seam_positions`; a clip
  without seams warns and runs plain. Every seam gets a guide (the ACEScct source frame VAE-encoded alone in
  float32, tiled above 512x768 when VAE tiling is on, held at `keyframe_strength`) and a generated slot
  (mask 0, same one-pixel-frame RoPE span), appended as [video | reference | guides | slots] like the reference.
  The slots and the first latent frame carry the keyframe position embedding; seams need a transformer with
  `use_keyframes_abs_pos_embedding` and a single reference video.
- Guide velocities are converted to x0 with each token's own timestep, as the reference X0 model does.
- After denoising, the slots are cut out, denormalized like the video and passed to `diffusion_decoder.decode`
  as `keyframe_latents` / `keyframe_frame_indices`. Without a diffusion decoder the pipeline warns and the VAE
  decodes the video without them.
- `high_quality_hdr`: frame-doubled source, `2N - 1` generated frames, doubled seams, every second frame kept.
- Reuses the DFR helpers (`_prepare_keyframe_coords`, `_unpack_video_and_slots`) through "Copied from".

`keyframe_strength=None` reproduces the parent commit bit for bit; the LogC3 path ignores `keyframe_strength`
and rejects `high_quality_hdr`. On tiny shapes with a stand-in encoder and velocity model, the token sequence,
RoPE positions, keyframe marker, per-token timesteps, the 8-step Euler trajectory and the extracted slots match
the reference builders to within 6e-8 (video tokens and slots bitwise).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Hi @christopher5106, thanks for the PR! It does not appear to link an issue it fixes. If this PR addresses an existing issue, please add a closing keyword (e.g. Fixes #1234) to the PR description so the issue is linked. See the contribution guide for more details. If this PR intentionally does not fix a tracked issue, a maintainer can add the no-issue-needed label to silence this reminder.

Please note that PRs without a linked issue are likely to be automatically closed 10 days after this notice.

Once the PR links an issue (or gets the no-issue-needed label), you can ignore this message — it stays here as a comment, but it no longer applies.

@sayakpaul

Copy link
Copy Markdown
Member

@christopher5106 appreciate the PRs you have been opening up recently but can we first discuss them in issues, first? This also keeps our reviewing queue manageable without ghosting contributors like yourself.

@christopher5106

Copy link
Copy Markdown
Contributor Author

Closing for now, as asked: I should have opened an issue first. The scope and structure are in #14981, and I'll reopen in whatever shape you prefer there. Thanks for the patience.

The decoder checkpoint mismatch documented here is tracked on the Hub in https://huggingface.co/Lightricks/LTX-2.5-Diffusers/discussions/19.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation models pipelines size/L PR with diff > 200 LOC tests utils

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants