Repository navigation
[LTX-2.5] Support the SDR-To-HDR IC-LoRA in LTX2HDRPipeline - #14974
christopher5106 wants to merge 2 commits into
Conversation
…-HDR Port the colour pipeline of the LTX-2.5 SDR-To-HDR IC-LoRA from the Lightricks LTX-2 reference (ltx_core.hdr, ltx_core.color, ltx_pipelines media_io): - image_processor: sRGB EOTF, Bradford Rec.709/AP1/Rec.2020 matrices, ACEScct encode/decode, the srgb_gamma/srgb/acescg/acescct input transforms and the ACEScct -> ACEScg/Rec.709 linear output transform, as pure torch functions. LTX2VideoHDRProcessor gains hdr_transform="acescct" with input_colorspace / output_colorspace options; the LogC3 default is unchanged. - export_utils: encode_hdr_tensor_to_hlg_mp4 (BT.2020 HLG 10-bit HEVC via PyAV/libx265) and save_exr_frame / export_to_exr_sequence (half-float ZIP EXR with chromaticities and colorSpace), behind a new optional is_openexr_available() guard. - tests: exact values computed with colour-science, round trips, HLG stream properties and EXR read-back. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Run the Lightricks LTX-2.5 SDR-To-HDR IC-LoRA (ltx_pipelines.hdr_ic_lora.HDRICLoraPipeline, without its optional seam keyframes) when the pipeline is built with hdr_transform="acescct": - the reference video goes through the ACEScct input transform (new `input_colorspace`, default "srgb_gamma"), is reflect-padded up to a multiple of the VAE spatial compression ratio and VAE-encoded in float32; the decoded video is cropped back to height x width; - conditioning comes from precomputed `connector_video_embeds` only (a 2D `video_context` is accepted): no prompt, no text encoder call, `connector_audio_embeds` optional; - video-only denoising: `isolate_modalities=True` with a single placeholder audio token, so the audio stream cannot reach the video; - the distilled schedule used verbatim (DISTILLED_SIGMA_VALUES, no mu shift), no CFG/STG/modality guidance, RoPE frame rate capped at 30 fps, first target latent frame marked for the keyframe position embedding; - decoding in float32 with the optional `diffusion_decoder` component (LTX-2.5) or the VAE, then `postprocess_hdr_video(output_colorspace=...)` (new argument, default "rec709"). `hdr_transform` is now registered in the pipeline config so it survives save/load. The LogC3 (LTX-2.3) path is unchanged: its outputs are bit-identical to the parent commit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Follow-up stacked on this PR: #14975 adds keyframe-aware diffusion decoding and the SDR-To-HDR seam keyframes. It also documents that the published |
|
Hi @christopher5106, thanks for the PR! It does not appear to link an issue it fixes. If this PR addresses an existing issue, please add a closing keyword (e.g. Please note that PRs without a linked issue are likely to be automatically closed 10 days after this notice. Once the PR links an issue (or gets the |
|
Closing for now, as asked: I should have opened an issue first. The scope and structure are in #14981, and I'll reopen in whatever shape you prefer there. Thanks for the patience. |
What does this PR do?
Runs Lightricks' LTX-2.5-22b-IC-LoRA-SDR-To-HDR in the existing
LTX2HDRPipeline. It followsltx_pipelines.hdr_ic_lora.HDRICLoraPipelinewithout seam keyframes, and is selected withhdr_transform="acescct":input_colorspace, default"srgb_gamma"), reflect-padded to a multiple of 32 and encoded with the VAE in float32. The decoded video is cropped back toheightxwidth, so any clip size works.video_context, passed asconnector_video_embeds(2D accepted). No prompt and no text encoder are needed;text_encoder,tokenizerandconnectorscan be loaded asNone, andconnector_audio_embedsis optional.isolate_modalities=True), so the placeholder audio stream cannot influence the video.DISTILLED_SIGMA_VALUESused as given (no shift) and no CFG/STG/modality guidance. RoPE is capped at 30 fps for faster sources, and the first latent frame is marked for the keyframe position embedding, as in the reference.diffusion_decodercomponent when it is loaded, and through the VAE otherwise. Output goes throughpostprocess_hdr_video(output_colorspace=...): linear Rec.709 by default, or linear ACEScg / raw ACEScct.hdr_transformis now saved in the pipeline config.The LTX-2.3 LogC3 path is unchanged and bit-identical across six configurations. The LoRA's ComfyUI-style keys are already handled by
LTX2LoraLoaderMixin; a test pins the format.Parity with the reference, real weights
Run on one H200:
hiker.mp4fromdocumentation-images(720x480, 49 frames), seed 42, the LoRA at 1.0 and its scene embedding, 8 distilled steps. Both sides use bf16 weights. The reference is Lightricks/LTX-29ec55f9f,HDRICLoraPipelinewithkeyframe_strength=None.The remaining gap comes from precision and noise handling, not logic:
The difference concentrates on high-frequency texture. LoRA loading reports no missing or unexpected keys.
Not included
high_quality_hdr(frame doubling), seam keyframes and keyframe-aware decoding: these come in the follow-up PR.encode_hdr_tensor_to_hlg_mp4/export_to_exr_sequence. The second docstring example shows it.Tests
tests/pipelines/ltx2/test_ltx2_hdr_sdr_to_hdr.py(30) covers:test_ltx2_hdr.pyresults are unchanged. Ruff, doc-builder,check_copies,check_dummiesandcheck_forward_call_docstringsare clean.Before submitting
Who can review?
@DN6 @sayakpaul @yiyixuxu
🤖 Generated with Claude Code