Repository navigation
[LTX-2.5] Keyframe-aware decoding + SDR-To-HDR seams (fix branch, not to be merged) - #8
Open
christopher5106 wants to merge 4 commits into
Open
christopher5106 wants to merge 4 commits into
christopher5106 wants to merge 4 commits into
Conversation
…-HDR Port the colour pipeline of the LTX-2.5 SDR-To-HDR IC-LoRA from the Lightricks LTX-2 reference (ltx_core.hdr, ltx_core.color, ltx_pipelines media_io): - image_processor: sRGB EOTF, Bradford Rec.709/AP1/Rec.2020 matrices, ACEScct encode/decode, the srgb_gamma/srgb/acescg/acescct input transforms and the ACEScct -> ACEScg/Rec.709 linear output transform, as pure torch functions. LTX2VideoHDRProcessor gains hdr_transform="acescct" with input_colorspace / output_colorspace options; the LogC3 default is unchanged. - export_utils: encode_hdr_tensor_to_hlg_mp4 (BT.2020 HLG 10-bit HEVC via PyAV/libx265) and save_exr_frame / export_to_exr_sequence (half-float ZIP EXR with chromaticities and colorSpace), behind a new optional is_openexr_available() guard. - tests: exact values computed with colour-science, round trips, HLG stream properties and EXR read-back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run the Lightricks LTX-2.5 SDR-To-HDR IC-LoRA (ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline, without its optional seam keyframes) when the pipeline is built with hdr_transform="acescct": - the reference video goes through the ACEScct input transform (new `input_colorspace`, default "srgb_gamma"), is reflect-padded up to a multiple of the VAE spatial compression ratio and VAE-encoded in float32; the decoded video is cropped back to height x width; - conditioning comes from precomputed `connector_video_embeds` only (a 2D `video_context` is accepted): no prompt, no text encoder call, `connector_audio_embeds` optional; - video-only denoising: `isolate_modalities=True` with a single placeholder audio token, so the audio stream cannot reach the video; - the distilled schedule used verbatim (DISTILLED_SIGMA_VALUES, no mu shift), no CFG/STG/modality guidance, RoPE frame rate capped at 30 fps, first target latent frame marked for the keyframe position embedding; - decoding in float32 with the optional `diffusion_decoder` component (LTX-2.5) or the VAE, then `postprocess_hdr_video(output_colorspace=...)` (new argument, default "rec709"). `hdr_transform` is now registered in the pipeline config so it survives save/load. The LogC3 (LTX-2.3) path is unchanged: its outputs are bit-identical to the parent commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port Lightricks' keyframe-aware DiffVAE decode (LTX-2 @ 9ec55f9f,
`ltx_core/model/video_vae/{keyframes,diffusion_video_decoder}.py`,
`transformer/fallback_na/joint_eager.py`) to `LTX2VideoDiffusionDecoderModel`.
- `decoder_keyframe_type_embedding` config flag (default off) creates the
learned keyframe tag `decoder.type_emb` `(latent_channels,)`; existing
checkpoints load unchanged.
- `decode(..., keyframe_latents=, keyframe_frame_indices=)`: planes are tagged,
share `conv_in` and every stage with the video, and attend jointly with it
(each video position sees its 2 nearest planes, each plane its 2 nearest
frames, one softmax), with per-stage plane times. Tiled decode keeps the
planes inside each temporal tile plus the nearest on each side, with times
rebased on the tile. Without planes the decode is bit-identical to before.
- Joint attention is a port of the reference's pure-torch backend, bitwise
equal to it; full-decoder parity on random tiny weights is within 2.4e-6.
- Converter carries `decoder.type_emb` and sets the flag when present (the
current original checkpoint has it, so the strict load used to fail).
- `dfr_layout.resolve_seam_positions`: seam keyframes of a single-window clip
(24/32 segments clipped to the clip, high-quality doubling), equal to the
reference for every 8k+1 frame count up to 2001.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Port the seam keyframes of Lightricks' `HDRICLoraPipeline` (LTX-2 @ 9ec55f9f, `ltx_pipelines/hdr_ic_lora.py`) to the `hdr_transform="acescct"` path: - `keyframe_strength` (default 0.95, `None` = plain IC-LoRA): seams from `resolve_seam_positions`; a clip without seams warns and runs plain. Every seam gets a guide (the ACEScct source frame VAE-encoded alone in float32, tiled above 512x768 when VAE tiling is on, held at `keyframe_strength`) and a generated slot (mask 0, same one-pixel-frame RoPE span), appended as [video | reference | guides | slots] like the reference. The slots and the first latent frame carry the keyframe position embedding; seams need a transformer with `use_keyframes_abs_pos_embedding` and a single reference video. - Guide velocities are converted to x0 with each token's own timestep, as the reference X0 model does. - After denoising, the slots are cut out, denormalized like the video and passed to `diffusion_decoder.decode` as `keyframe_latents` / `keyframe_frame_indices`. Without a diffusion decoder the pipeline warns and the VAE decodes the video without them. - `high_quality_hdr`: frame-doubled source, `2N - 1` generated frames, doubled seams, every second frame kept. - Reuses the DFR helpers (`_prepare_keyframe_coords`, `_unpack_video_and_slots`) through "Copied from". `keyframe_strength=None` reproduces the parent commit bit for bit; the LogC3 path ignores `keyframe_strength` and rejects `high_quality_hdr`. On tiny shapes with a stand-in encoder and velocity model, the token sequence, RoPE positions, keyframe marker, per-token timesteps, the 8-step Euler trajectory and the extracted slots match the reference builders to within 6e-8 (video tokens and slots bitwise). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Tracking PR for
fix_ltx2_hdr_keyframe_decoding, opened againstmain(the upstream mirror). It is never merged here.Upstream PR: huggingface#14975, stacked on huggingface#14974 (tracked in #7), which is stacked on huggingface#14966 (tracked in #5).
Adds keyframe-aware diffusion decoding:
decoder.type_emband joint keyframe-plane attention inLTX2VideoDiffusionDecoderModel, with the converter carrying the tag. It also adds the SDR-To-HDR seam keyframes inLTX2HDRPipeline:keyframe_strength,high_quality_hdr, guides and slots.Notes:
Lightricks/LTX-2.5-Diffusersdiffusion_decoderdoes not match the current original VAE. Re-converted, the decoder matches Lightricks' at 62–69 dB instead of 35–44 dB.When upstream merges it: delete this branch, close this PR, sync
main, and rebasescenariodropping the cherry-picks.