Skip to content

[LTX-2.5] Support the SDR-To-HDR IC-LoRA in LTX2HDRPipeline - #14974

Closed
christopher5106 wants to merge 2 commits into
huggingface:mainfrom
scenario-labs:fix_ltx2_hdr_sdr_to_hdr
Closed

christopher5106 wants to merge 2 commits into
huggingface:mainfrom
scenario-labs:fix_ltx2_hdr_sdr_to_hdr

Conversation

@christopher5106

@christopher5106 christopher5106 commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #14966 (ACEScct colour transforms and HLG/EXR export): the first commit of this branch is #14966's, so please review that PR first and only the last commit () here. A follow-up PR, stacked on this one, adds seam keyframes and keyframe-aware diffusion decoding.

What does this PR do?

Runs Lightricks' LTX-2.5-22b-IC-LoRA-SDR-To-HDR in the existing LTX2HDRPipeline. It follows ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline without seam keyframes, and is selected with hdr_transform="acescct":

  • Input: the SDR reference is mapped to ACEScct (input_colorspace, default "srgb_gamma"), reflect-padded to a multiple of 32 and encoded with the VAE in float32. The decoded video is cropped back to height x width, so any clip size works.
  • Conditioning: only the LoRA's precomputed video_context, passed as connector_video_embeds (2D accepted). No prompt and no text encoder are needed; text_encoder, tokenizer and connectors can be loaded as None, and connector_audio_embeds is optional.
  • Video-only denoising: audio↔video cross-attention is off (isolate_modalities=True), so the placeholder audio stream cannot influence the video.
  • Sampling: one distilled stage, with DISTILLED_SIGMA_VALUES used as given (no shift) and no CFG/STG/modality guidance. RoPE is capped at 30 fps for faster sources, and the first latent frame is marked for the keyframe position embedding, as in the reference.
  • Decoding: in float32, through a new optional diffusion_decoder component when it is loaded, and through the VAE otherwise. Output goes through postprocess_hdr_video(output_colorspace=...): linear Rec.709 by default, or linear ACEScg / raw ACEScct.
  • Config: hdr_transform is now saved in the pipeline config.

The LTX-2.3 LogC3 path is unchanged and bit-identical across six configurations. The LoRA's ComfyUI-style keys are already handled by LTX2LoraLoaderMixin; a test pins the format.

Parity with the reference, real weights

Run on one H200: hiker.mp4 from documentation-images (720x480, 49 frames), seed 42, the LoRA at 1.0 and its scene embedding, 8 distilled steps. Both sides use bf16 weights. The reference is Lightricks/LTX-2 9ec55f9f, HDRICLoraPipeline with keyframe_strength=None.

Stage diffusers vs reference
ACEScct input pixels after pad max abs 1.2e-7
VAE-encoded reference latents cosine 0.99999998
RoPE positions, keyframe mask, denoise mask, token order identical
Initial noise same draw (cosine 0.9999986; reference holds it in bf16)
Final latent PSNR 50.0 dB, cosine 0.99982, flat across latent frames
Output (ACEScct codes) PSNR 38.4 dB, cosine 0.99973, mean abs 0.0076; flat per frame (0.0068–0.0081), no drift

The remaining gap comes from precision and noise handling, not logic:

  • the reference keeps its latent state and the encoded reference latents in bf16, while this pipeline keeps them in float32;
  • the reference re-seeds the decoder's noise, while here it continues the user's generator.

The difference concentrates on high-frequency texture. LoRA loading reports no missing or unexpected keys.

Not included

Tests

tests/pipelines/ltx2/test_ltx2_hdr_sdr_to_hdr.py (30) covers:

  • end to end, with and without the diffusion decoder;
  • no prompt encoding (the text components are patched to raise);
  • audio isolation (random audio leaves the video bit-equal, with a LogC3 control);
  • the float32 VAE and decoder;
  • reflect padding and crop;
  • the RoPE fps cap;
  • the distilled schedule;
  • the keyframe marker;
  • save/load;
  • the LoRA key format;
  • a pinned LogC3 output slice.

test_ltx2_hdr.py results are unchanged. Ruff, doc-builder, check_copies, check_dummies and check_forward_call_docstrings are clean.

Before submitting

  • Did you read the contributor guideline?
  • Did you write any new necessary tests?
  • Did you update the documentation? (docstrings, example)

Who can review?

@DN6 @sayakpaul @yiyixuxu

🤖 Generated with Claude Code

christopher5106 and others added 2 commits October 6, 2026 17:44
…-HDR

Port the colour pipeline of the LTX-2.5 SDR-To-HDR IC-LoRA from the Lightricks
LTX-2 reference (ltx_core.hdr, ltx_core.color, ltx_pipelines media_io):

- image_processor: sRGB EOTF, Bradford Rec.709/AP1/Rec.2020 matrices, ACEScct
  encode/decode, the srgb_gamma/srgb/acescg/acescct input transforms and the
  ACEScct -> ACEScg/Rec.709 linear output transform, as pure torch functions.
  LTX2VideoHDRProcessor gains hdr_transform="acescct" with input_colorspace /
  output_colorspace options; the LogC3 default is unchanged.
- export_utils: encode_hdr_tensor_to_hlg_mp4 (BT.2020 HLG 10-bit HEVC via
  PyAV/libx265) and save_exr_frame / export_to_exr_sequence (half-float ZIP
  EXR with chromaticities and colorSpace), behind a new optional
  is_openexr_available() guard.
- tests: exact values computed with colour-science, round trips, HLG stream
  properties and EXR read-back.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run the Lightricks LTX-2.5 SDR-To-HDR IC-LoRA (ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline,
without its optional seam keyframes) when the pipeline is built with hdr_transform="acescct":

- the reference video goes through the ACEScct input transform (new `input_colorspace`, default
  "srgb_gamma"), is reflect-padded up to a multiple of the VAE spatial compression ratio and
  VAE-encoded in float32; the decoded video is cropped back to height x width;
- conditioning comes from precomputed `connector_video_embeds` only (a 2D `video_context` is
  accepted): no prompt, no text encoder call, `connector_audio_embeds` optional;
- video-only denoising: `isolate_modalities=True` with a single placeholder audio token, so the
  audio stream cannot reach the video;
- the distilled schedule used verbatim (DISTILLED_SIGMA_VALUES, no mu shift), no CFG/STG/modality
  guidance, RoPE frame rate capped at 30 fps, first target latent frame marked for the keyframe
  position embedding;
- decoding in float32 with the optional `diffusion_decoder` component (LTX-2.5) or the VAE, then
  `postprocess_hdr_video(output_colorspace=...)` (new argument, default "rec709").

`hdr_transform` is now registered in the pipeline config so it survives save/load. The LogC3
(LTX-2.3) path is unchanged: its outputs are bit-identical to the parent commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@christopher5106

Copy link
Copy Markdown
Contributor Author

Follow-up stacked on this PR: #14975 adds keyframe-aware diffusion decoding and the SDR-To-HDR seam keyframes. It also documents that the published LTX-2.5-Diffusers diffusion decoder doesn't match the current original VAE.

@christopher5106
christopher5106 marked this pull request as ready for review October 7, 2026 00:30
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Hi @christopher5106, thanks for the PR! It does not appear to link an issue it fixes. If this PR addresses an existing issue, please add a closing keyword (e.g. Fixes #1234) to the PR description so the issue is linked. See the contribution guide for more details. If this PR intentionally does not fix a tracked issue, a maintainer can add the no-issue-needed label to silence this reminder.

Please note that PRs without a linked issue are likely to be automatically closed 10 days after this notice.

Once the PR links an issue (or gets the no-issue-needed label), you can ignore this message — it stays here as a comment, but it no longer applies.

@christopher5106

Copy link
Copy Markdown
Contributor Author

Closing for now, as asked: I should have opened an issue first. The scope and structure are in #14981, and I'll reopen in whatever shape you prefer there. Thanks for the patience.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant