Skip to content

[LTX-2] Add conditioning_frame_rate to decouple the positional time base from the playback frame rate - #14967

Open
christopher5106 wants to merge 1 commit into
huggingface:mainfrom
scenario-labs:fix_ltx2_conditioning_frame_rate
Open

christopher5106 wants to merge 1 commit into
huggingface:mainfrom
scenario-labs:fix_ltx2_conditioning_frame_rate

Conversation

@christopher5106

@christopher5106 christopher5106 commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #14980

What does this PR do?

LTX-2 places every video token on a time axis in seconds (pixel_frame / fps in the RoPE coordinates), and the pipelines derive that fps from frame_rate, the playback rate. Some adapters need the two apart. Lightricks/LTX-2.5-22b-LoRA-Slow-Motion-Control was trained on high-speed footage labelled with its effective capture rate, and its reference driver conditions the diffusion stage on fps / speed while the video is still encoded at fps (distilled_speed_demo.py). The model card warns that "with a generic loader that ignores the speed value you get a plain LoRA and none of the intended motion control". There is no way to express that with the current pipelines: raising frame_rate also shortens the generated audio, changes the duration head's frame count, and is the rate callers encode the output at.

This PR adds an optional conditioning_frame_rate:

  • Modular (LTX2* / LTX25* blocks): a new InputParam on the six blocks that place tokens in time. These are LTX2PrepareCoordsStep, LTX2ConditionPrepareCoordsStep, LTX2ConditionPrepareLatentsStep (keyframe coords), LTX2InContextPrepareLatentsStep (keyframe coords), LTX2ReferenceEncoderStep (reference coords) and LTX2LoopDenoiser (the transformer's fps).
  • Classic (LTX2Pipeline, LTX2ImageToVideoPipeline, LTX2ConditionPipeline, LTX2InContextPipeline): a new __call__ argument applied to the same sites. It reaches the video coords, the transformer's fps and the keyframe/reference coords through prepare_latents.

Unchanged, still following frame_rate: the audio latent length (num_frames / frame_rate seconds), the duration head's seconds → frames, and the returned video. When conditioning_frame_rate is None (the default) it resolves to frame_rate, so existing calls are bit-identical.

This mirrors the split the DFR pipelines already make internally (conditioning_fps vs playback fps in pipeline_ltx2_dfr.py), exposed as an argument instead of derived. The same caveat applies: the base model is trained around 24/25/30/60 fps (see MAX_CONDITIONING_FPS in ltx2/utils.py), so a conditioning rate far from those is only meaningful with an adapter trained for it, which the docstring says.

pipe.load_lora_weights(
    "Lightricks/LTX-2.5-22b-LoRA-Slow-Motion-Control",
    weight_name="ltx-2.5-22b-lora-slow-motion-control-1.0.safetensors",
)
speed = 0.2  # 5x slow motion
video, audio = pipe(
    image=image,
    prompt=prompt,
    num_frames=121,
    frame_rate=24.0,                         # playback rate: encode the video at this
    conditioning_frame_rate=24.0 / speed,    # what the model is conditioned on
    sigmas=DISTILLED_SIGMA_VALUES,
    output_type="np",
    return_dict=False,
)

Tests

  • tests/pipelines/ltx2/test_ltx2.py::test_conditioning_frame_rate_rescales_only_the_video_time_axis and the modular counterpart in tests/modular_pipelines/ltx2/test_modular_pipeline_ltx25.py:
    • conditioning_frame_rate=frame_rate reproduces the default output exactly
    • a doubled rate halves the time axis of the video coords and leaves the spatial axes unchanged
    • the transformer receives the new fps
    • the audio token count and the output shapes are unchanged
  • Existing LTX-2 classic and modular suites pass unchanged (CPU slice tests included).
  • make fix-copies, utils/modular_auto_docstring.py (the 12 regenerated LTX-2 blockset docstrings), utils/check_forward_call_docstrings.py, ruff and doc-builder style are clean.

Before submitting

  • Did you read the contributor guideline?
  • Did you write any new necessary tests?
  • Did you update the documentation? (docstrings; InputParam descriptions)

Who can review?

@yiyixuxu @DN6 @sayakpaul (LTX-2 / modular)

🤖 Generated with Claude Code

@christopher5106

Copy link
Copy Markdown
Contributor Author

End-to-end check with the Slow-Motion-Control LoRA

Setup:

  • Model: LTX-2.5 distilled (8 sigmas).
  • LoRA: Lightricks/LTX-2.5-22b-LoRA-Slow-Motion-Control at scale 1.0.
  • Image to video from documentation-images' astronaut.jpg, prompt "the astronaut moves slowly while the camera gently pushes in".
  • 121 frames, frame_rate=24, same seed for every run.

The only thing that varies is conditioning_frame_rate. Motion is measured as the mean absolute per-pixel change between consecutive decoded frames, on 0–255 RGB:

conditioning_frame_rate mean change / frame first → last frame
LoRA, speed 1.0 24 (= frame_rate) 2.75 36.0
LoRA, speed 0.2 120 (= frame_rate / 0.2) 0.90 26.5

At speed 0.2 there is about 3× less motion per frame. At speed 1.0 the subject turns and walks across the frame; at speed 0.2 the same motion only begins within the same 5 s of playback.

The frame count, the generated audio track and its length are identical across the two runs, matching the unit tests: only the positional time axis moves. Happy to share the clips if useful.

@yiyixuxu yiyixuxu left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for the PR!

can you help me understand a bit more: can't we just pass 120 as frame_rate to the pipeline and 24 to encode_video()?

@christopher5106

Copy link
Copy Markdown
Contributor Author

Thanks for the review! Passing frame_rate=120 does give the same video time axis, but frame_rate also sizes the generated audio and the duration head, which should stay on the playback rate:

  1. Audio: the audio latents are sized from num_frames / frame_rate (duration_s in LTX2Pipeline.__call__). At frame_rate=120 with 121 frames the model generates ~1 s of audio for a clip that plays 5 s at 24 fps, so it would have to be time-stretched 5× or dropped. With conditioning_frame_rate=120, frame_rate=24 the audio is generated for the real 5 s (in the run above: 240,640 samples at 48 kHz = 5.01 s).
  2. Duration head: with num_frames omitted, the predicted seconds are converted with frame_rate, so 120 would yield 5× the frames for the same predicted duration.

That's also how Lightricks wires this LoRA: their distilled_speed_demo.py only wraps the diffusion stage with fps / speed, while the target (incl. audio sizing) and the decode keep the real fps. It's the same split the DFR pipelines already make internally with conditioning_fps; this PR exposes it as an argument.

For silent video with an explicit num_frames, your workaround is equivalent; the parameter matters once audio or auto-duration is involved. Happy to rename it or trim the scope if you'd prefer a different shape.

@christopher5106

Copy link
Copy Markdown
Contributor Author

Measured on the tiny LTX2Pipeline from the test suite (41 frames):

audio tokens transformer fps frames
frame_rate=24 43 24 41
frame_rate=120 9 120 41
frame_rate=24, conditioning_frame_rate=120 43 120 41

And with num_frames omitted (duration head, 1–2 s bounds): frame_rate=24 → 25 frames, frame_rate=120 → 121 frames, frame_rate=24, conditioning_frame_rate=120 → 25 frames.

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

Hi @christopher5106, thanks for the PR! It does not appear to link an issue it fixes. If this PR addresses an existing issue, please add a closing keyword (e.g. Fixes #1234) to the PR description so the issue is linked. See the contribution guide for more details. If this PR intentionally does not fix a tracked issue, a maintainer can add the no-issue-needed label to silence this reminder.

Please note that PRs without a linked issue are likely to be automatically closed 10 days after this notice.

Once the PR links an issue (or gets the no-issue-needed label), you can ignore this message — it stays here as a comment, but it no longer applies.

… base from the playback frame rate

LTX-2 places video tokens on a time axis in seconds (`pixel_frame / fps`
in the RoPE coordinates), derived from `frame_rate`. Adapters such as the
LTX-2.5 Slow-Motion-Control LoRA are conditioned on a different rate
(`fps / speed`) than the video plays at. Add an optional
`conditioning_frame_rate` to the LTX-2 classic pipelines and modular
blocks: it drives the video, keyframe and reference coordinates and the
transformer's `fps`, while the audio length, the duration head and the
returned video keep following `frame_rate`. `None` resolves to
`frame_rate`, leaving current outputs unchanged.
@christopher5106

Copy link
Copy Markdown
Contributor Author

Self-review against .ai/review-rules.md: the new tests monkeypatched transformer.forward to record its inputs, which testing.md rules out. They now assert on public outputs only: the modular test requests video_coords / audio_coords and checks that a doubled conditioning_frame_rate halves the video time axis and nothing else; the classic tests check that the default is unchanged, that shapes and audio length are preserved, and that the duration head still follows frame_rate.

pytest tests/pipelines/ltx2/test_ltx2.py tests/modular_pipelines/ltx2/test_modular_pipeline_ltx25.py -k conditioning_frame_rate
3 passed

ruff check / ruff format --check clean. Issue: #14980.

@yiyixuxu

yiyixuxu commented Oct 7, 2026

Copy link
Copy Markdown
Collaborator

thanks @christopher5106
would it work if we just add a motion_speed argument instead? frame_rate is a petty standard argument across our pipelines and should have one consistent meaning for the user , i.e. the playback rate. It's usually same as what the model conditioned on, but that's not something the end user has to know

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LTX-2: decouple the RoPE time base from the playback frame rate (Slow-Motion-Control LoRA)

2 participants