Skip to content

[LTX-2.5] Keyframe-aware decoding + SDR-To-HDR seams (fix branch, not to be merged) - #8

Open
christopher5106 wants to merge 4 commits into
mainfrom
fix_ltx2_hdr_keyframe_decoding
Open

christopher5106 wants to merge 4 commits into
mainfrom
fix_ltx2_hdr_keyframe_decoding

Conversation

@christopher5106

Copy link
Copy Markdown

Tracking PR for fix_ltx2_hdr_keyframe_decoding, opened against main (the upstream mirror). It is never merged here.

Upstream PR: huggingface#14975, stacked on huggingface#14974 (tracked in #7), which is stacked on huggingface#14966 (tracked in #5).

Adds keyframe-aware diffusion decoding: decoder.type_emb and joint keyframe-plane attention in LTX2VideoDiffusionDecoderModel, with the converter carrying the tag. It also adds the SDR-To-HDR seam keyframes in LTX2HDRPipeline: keyframe_strength, high_quality_hdr, guides and slots.

Notes:

When upstream merges it: delete this branch, close this PR, sync main, and rebase scenario dropping the cherry-picks.

christopher5106 and others added 4 commits October 6, 2026 17:44
…-HDR

Port the colour pipeline of the LTX-2.5 SDR-To-HDR IC-LoRA from the Lightricks
LTX-2 reference (ltx_core.hdr, ltx_core.color, ltx_pipelines media_io):

- image_processor: sRGB EOTF, Bradford Rec.709/AP1/Rec.2020 matrices, ACEScct
  encode/decode, the srgb_gamma/srgb/acescg/acescct input transforms and the
  ACEScct -> ACEScg/Rec.709 linear output transform, as pure torch functions.
  LTX2VideoHDRProcessor gains hdr_transform="acescct" with input_colorspace /
  output_colorspace options; the LogC3 default is unchanged.
- export_utils: encode_hdr_tensor_to_hlg_mp4 (BT.2020 HLG 10-bit HEVC via
  PyAV/libx265) and save_exr_frame / export_to_exr_sequence (half-float ZIP
  EXR with chromaticities and colorSpace), behind a new optional
  is_openexr_available() guard.
- tests: exact values computed with colour-science, round trips, HLG stream
  properties and EXR read-back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run the Lightricks LTX-2.5 SDR-To-HDR IC-LoRA (ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline,
without its optional seam keyframes) when the pipeline is built with hdr_transform="acescct":

- the reference video goes through the ACEScct input transform (new `input_colorspace`, default
  "srgb_gamma"), is reflect-padded up to a multiple of the VAE spatial compression ratio and
  VAE-encoded in float32; the decoded video is cropped back to height x width;
- conditioning comes from precomputed `connector_video_embeds` only (a 2D `video_context` is
  accepted): no prompt, no text encoder call, `connector_audio_embeds` optional;
- video-only denoising: `isolate_modalities=True` with a single placeholder audio token, so the
  audio stream cannot reach the video;
- the distilled schedule used verbatim (DISTILLED_SIGMA_VALUES, no mu shift), no CFG/STG/modality
  guidance, RoPE frame rate capped at 30 fps, first target latent frame marked for the keyframe
  position embedding;
- decoding in float32 with the optional `diffusion_decoder` component (LTX-2.5) or the VAE, then
  `postprocess_hdr_video(output_colorspace=...)` (new argument, default "rec709").

`hdr_transform` is now registered in the pipeline config so it survives save/load. The LogC3
(LTX-2.3) path is unchanged: its outputs are bit-identical to the parent commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port Lightricks' keyframe-aware DiffVAE decode (LTX-2 @ 9ec55f9f,
`ltx_core/model/video_vae/{keyframes,diffusion_video_decoder}.py`,
`transformer/fallback_na/joint_eager.py`) to `LTX2VideoDiffusionDecoderModel`.

- `decoder_keyframe_type_embedding` config flag (default off) creates the
  learned keyframe tag `decoder.type_emb` `(latent_channels,)`; existing
  checkpoints load unchanged.
- `decode(..., keyframe_latents=, keyframe_frame_indices=)`: planes are tagged,
  share `conv_in` and every stage with the video, and attend jointly with it
  (each video position sees its 2 nearest planes, each plane its 2 nearest
  frames, one softmax), with per-stage plane times. Tiled decode keeps the
  planes inside each temporal tile plus the nearest on each side, with times
  rebased on the tile. Without planes the decode is bit-identical to before.
- Joint attention is a port of the reference's pure-torch backend, bitwise
  equal to it; full-decoder parity on random tiny weights is within 2.4e-6.
- Converter carries `decoder.type_emb` and sets the flag when present (the
  current original checkpoint has it, so the strict load used to fail).
- `dfr_layout.resolve_seam_positions`: seam keyframes of a single-window clip
  (24/32 segments clipped to the clip, high-quality doubling), equal to the
  reference for every 8k+1 frame count up to 2001.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port the seam keyframes of Lightricks' `HDRICLoraPipeline` (LTX-2 @ 9ec55f9f, `ltx_pipelines/hdr_ic_lora.py`)
to the `hdr_transform="acescct"` path:

- `keyframe_strength` (default 0.95, `None` = plain IC-LoRA): seams from `resolve_seam_positions`; a clip
  without seams warns and runs plain. Every seam gets a guide (the ACEScct source frame VAE-encoded alone in
  float32, tiled above 512x768 when VAE tiling is on, held at `keyframe_strength`) and a generated slot
  (mask 0, same one-pixel-frame RoPE span), appended as [video | reference | guides | slots] like the reference.
  The slots and the first latent frame carry the keyframe position embedding; seams need a transformer with
  `use_keyframes_abs_pos_embedding` and a single reference video.
- Guide velocities are converted to x0 with each token's own timestep, as the reference X0 model does.
- After denoising, the slots are cut out, denormalized like the video and passed to `diffusion_decoder.decode`
  as `keyframe_latents` / `keyframe_frame_indices`. Without a diffusion decoder the pipeline warns and the VAE
  decodes the video without them.
- `high_quality_hdr`: frame-doubled source, `2N - 1` generated frames, doubled seams, every second frame kept.
- Reuses the DFR helpers (`_prepare_keyframe_coords`, `_unpack_video_and_slots`) through "Copied from".

`keyframe_strength=None` reproduces the parent commit bit for bit; the LogC3 path ignores `keyframe_strength`
and rejects `high_quality_hdr`. On tiny shapes with a stand-in encoder and velocity model, the token sequence,
RoPE positions, keyframe marker, per-token timesteps, the 8-step Euler trajectory and the extracted slots match
the reference builders to within 6e-8 (video tokens and slots bitwise).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation tests utils pipelines models labels Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation models pipelines tests utils

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant