Repository navigation
[LTX-2.5] Keyframe-aware diffusion decoding and SDR-To-HDR seam keyframes - #14975
christopher5106 wants to merge 4 commits into
Conversation
…-HDR Port the colour pipeline of the LTX-2.5 SDR-To-HDR IC-LoRA from the Lightricks LTX-2 reference (ltx_core.hdr, ltx_core.color, ltx_pipelines media_io): - image_processor: sRGB EOTF, Bradford Rec.709/AP1/Rec.2020 matrices, ACEScct encode/decode, the srgb_gamma/srgb/acescg/acescct input transforms and the ACEScct -> ACEScg/Rec.709 linear output transform, as pure torch functions. LTX2VideoHDRProcessor gains hdr_transform="acescct" with input_colorspace / output_colorspace options; the LogC3 default is unchanged. - export_utils: encode_hdr_tensor_to_hlg_mp4 (BT.2020 HLG 10-bit HEVC via PyAV/libx265) and save_exr_frame / export_to_exr_sequence (half-float ZIP EXR with chromaticities and colorSpace), behind a new optional is_openexr_available() guard. - tests: exact values computed with colour-science, round trips, HLG stream properties and EXR read-back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run the Lightricks LTX-2.5 SDR-To-HDR IC-LoRA (ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline, without its optional seam keyframes) when the pipeline is built with hdr_transform="acescct": - the reference video goes through the ACEScct input transform (new `input_colorspace`, default "srgb_gamma"), is reflect-padded up to a multiple of the VAE spatial compression ratio and VAE-encoded in float32; the decoded video is cropped back to height x width; - conditioning comes from precomputed `connector_video_embeds` only (a 2D `video_context` is accepted): no prompt, no text encoder call, `connector_audio_embeds` optional; - video-only denoising: `isolate_modalities=True` with a single placeholder audio token, so the audio stream cannot reach the video; - the distilled schedule used verbatim (DISTILLED_SIGMA_VALUES, no mu shift), no CFG/STG/modality guidance, RoPE frame rate capped at 30 fps, first target latent frame marked for the keyframe position embedding; - decoding in float32 with the optional `diffusion_decoder` component (LTX-2.5) or the VAE, then `postprocess_hdr_video(output_colorspace=...)` (new argument, default "rec709"). `hdr_transform` is now registered in the pipeline config so it survives save/load. The LogC3 (LTX-2.3) path is unchanged: its outputs are bit-identical to the parent commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port Lightricks' keyframe-aware DiffVAE decode (LTX-2 @ 9ec55f9f,
`ltx_core/model/video_vae/{keyframes,diffusion_video_decoder}.py`,
`transformer/fallback_na/joint_eager.py`) to `LTX2VideoDiffusionDecoderModel`.
- `decoder_keyframe_type_embedding` config flag (default off) creates the
learned keyframe tag `decoder.type_emb` `(latent_channels,)`; existing
checkpoints load unchanged.
- `decode(..., keyframe_latents=, keyframe_frame_indices=)`: planes are tagged,
share `conv_in` and every stage with the video, and attend jointly with it
(each video position sees its 2 nearest planes, each plane its 2 nearest
frames, one softmax), with per-stage plane times. Tiled decode keeps the
planes inside each temporal tile plus the nearest on each side, with times
rebased on the tile. Without planes the decode is bit-identical to before.
- Joint attention is a port of the reference's pure-torch backend, bitwise
equal to it; full-decoder parity on random tiny weights is within 2.4e-6.
- Converter carries `decoder.type_emb` and sets the flag when present (the
current original checkpoint has it, so the strict load used to fail).
- `dfr_layout.resolve_seam_positions`: seam keyframes of a single-window clip
(24/32 segments clipped to the clip, high-quality doubling), equal to the
reference for every 8k+1 frame count up to 2001.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port the seam keyframes of Lightricks' `HDRICLoraPipeline` (LTX-2 @ 9ec55f9f, `ltx_pipelines/hdr_ic_lora.py`) to the `hdr_transform="acescct"` path: - `keyframe_strength` (default 0.95, `None` = plain IC-LoRA): seams from `resolve_seam_positions`; a clip without seams warns and runs plain. Every seam gets a guide (the ACEScct source frame VAE-encoded alone in float32, tiled above 512x768 when VAE tiling is on, held at `keyframe_strength`) and a generated slot (mask 0, same one-pixel-frame RoPE span), appended as [video | reference | guides | slots] like the reference. The slots and the first latent frame carry the keyframe position embedding; seams need a transformer with `use_keyframes_abs_pos_embedding` and a single reference video. - Guide velocities are converted to x0 with each token's own timestep, as the reference X0 model does. - After denoising, the slots are cut out, denormalized like the video and passed to `diffusion_decoder.decode` as `keyframe_latents` / `keyframe_frame_indices`. Without a diffusion decoder the pipeline warns and the VAE decodes the video without them. - `high_quality_hdr`: frame-doubled source, `2N - 1` generated frames, doubled seams, every second frame kept. - Reuses the DFR helpers (`_prepare_keyframe_coords`, `_unpack_video_and_slots`) through "Copied from". `keyframe_strength=None` reproduces the parent commit bit for bit; the LogC3 path ignores `keyframe_strength` and rejects `high_quality_hdr`. On tiny shapes with a stand-in encoder and velocity model, the token sequence, RoPE positions, keyframe marker, per-token timesteps, the 8-step Euler trajectory and the extracted slots match the reference builders to within 6e-8 (video tokens and slots bitwise). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Hi @christopher5106, thanks for the PR! It does not appear to link an issue it fixes. If this PR addresses an existing issue, please add a closing keyword (e.g. Please note that PRs without a linked issue are likely to be automatically closed 10 days after this notice. Once the PR links an issue (or gets the |
|
@christopher5106 appreciate the PRs you have been opening up recently but can we first discuss them in issues, first? This also keeps our reviewing queue manageable without ghosting contributors like yourself. |
|
Closing for now, as asked: I should have opened an issue first. The scope and structure are in #14981, and I'll reopen in whatever shape you prefer there. Thanks for the patience. The decoder checkpoint mismatch documented here is tracked on the Hub in https://huggingface.co/Lightricks/LTX-2.5-Diffusers/discussions/19. |
What
The Lightricks reference
HDRICLoraPipeline(ltx_pipelines/hdr_ic_lora.py, LTX-2 @ 9ec55f9f) adds seam keyframes by default (CLI--keyframe-strength 0.95). At each 24- or 32-frame DFR segment seam of the clip it appends two tokens:The slots are denoised with the video, then anchor the keyframe-aware DiffVAE decode. This PR wires that into the
hdr_transform="acescct"path:keyframe_strength: float | None = 0.95.Noneruns the plain IC-LoRA. Clips with no seam (for example 9 or 17 frames) log a warning and run the plain IC-LoRA.high_quality_hdr: bool = False: frame-doubled source,2N - 1generated frames, seam positions doubled, every second decoded frame kept.[video | reference | guides | slots]. Initial noising atnoise_scale = sigmas[0](1.0) gives the reference's result: clean reference tokens, pure noise on the video and slots, and0.95 * clean + 0.05 * noiseon the guides. Each step blends x0 with the clean tokens, and the guides' velocity is converted with their own per-token timestep, as the reference's X0 model does.LTX2VideoDiffusionDecoderModel.decode(..., keyframe_latents=, keyframe_frame_indices=)in float32.keyframe_strengthmust be in [0, 1]. When the clip has seams, the transformer must setuse_keyframes_abs_pos_embeddingand there must be exactly one reference video.high_quality_hdris rejected on LogC3.# Copied from.The LogC3 (LTX-2.3) path ignores
keyframe_strengthand is bit-identical to before.keyframe_strength=Noneis bit-identical to PR 2.Checkpoint note: the published
LTX-2.5-Diffusersdiffusion decoderKeyframe-aware decoding needs the decoder weights that were trained with it. The original
Lightricks/LTX-2.5vae/ltx-2.5-video-vae-bf16.safetensorshas them, including the keyframe tagdecoder.type_emb, and the converter now carries the tag (previous commit). Thediffusion_decodercurrently published inLightricks/LTX-2.5-Diffusersdoes not match that file:decoder.type_emb;conv_in/conv_outup to 10–40% onw_down,context_projandupsamples.*). The latent statistics are identical.The converter on main also cannot convert the current original VAE: its strict load fails on
decoder.type_emb, which suggests the Hub copy was converted from an earlier file.Decoder-only check on real weights (H200, 49 frames at 480x736, identical decoder inputs and identical noise on both sides, compared with Lightricks' decoder):
-DiffusersdecoderWith matching weights,
conv_inis bit-identical, decoder stages 1–4 agree to about 97 dB for both the video and keyframe streams, and the keyframe times agree at every stage. Re-converting the Hubdiffusion_decoderfrom the current original VAE (convert_ltx2_diffusion_video_vae, which now setsdecoder_keyframe_type_embedding=True) is what makes keyframe decoding, and plain decoding, match the reference.Interaction with #14694
#14694 refactors the forward methods of the LTX-2.5 diffusion decoder. The previous commit adds keyframe variants of those forwards (
*_with_keyframes, joint attention, per-tile plane selection), so whichever PR lands second needs a rebase. This PR only calls the publicdecode(z, generator=, keyframe_latents=, keyframe_frame_indices=)and does not depend on the internals.Deviations from the reference
Generator(seed)for decoding. Here the decoder continues the user'sgenerator, as PR 2 already does.vae.enable_tiling(), because diffusers tiling is opt-in, unlike the reference'sAUTO_TILING. The full-clip encode keeps PR 2'svae.encodebehaviour.1 - 0.95. The reference stores1 - strengthin the guide latent's dtype, which is bf16 at runtime (0.05005).keyframe_strengthis set.diffusion_decoder: the pipeline warns and decodes with the VAE, so the slots only take part in denoising. The reference always decodes with the keyframe-aware DiffVAE.lerps. Atnoise_scale=1the two are identical on the video, reference and slot tokens, and within 1 ulp on the guides.LTX2DFRPipeline. This equals the reference's fixed x8 for every LTX VAE.output_type="latent"returns the video latents only, without the slots.Tests
tests/pipelines/ltx2/test_ltx2_hdr_sdr_to_hdr.pynow covers:decode, with indices and values;keyframe_strength=Noneagainst a slice recorded on the parent commit;high_quality_hdrframe doubling, guide frame, decode and stride;On random tiny shapes with a stand-in encoder and velocity model, the following match the reference's own builders to within 6e-8:
Video tokens and slots are bitwise equal. The builders used are
VideoConditionByReferenceLatent,_keyframe_conditionings_from_pixel_frames,generated_keyframe_conditionings,create_noised_state,post_process_latent,EulerDiffusionStep,clear_conditioninganddecode_keyframes_from_slots.End to end on real weights (H200,
hiker.mp4fromdocumentation-images, seed 42, seams at frames 24 and 48, againstHDRICLoraPipelinewithkeyframe_strength=0.95): the final latents match at PSNR 52.4 dB, and token-level checks at 70 dB or better. The decoded output (34.8 dB with the published decoder) is dominated by the decoder-weight mismatch above.🤖 Generated with Claude Code