Repository navigation
[LTX-2] Add conditioning_frame_rate to decouple the positional time base from the playback frame rate - #14967
Conversation
ac1aa33 to
446cbaf
Compare
|
End-to-end check with the Slow-Motion-Control LoRA Setup:
The only thing that varies is
At speed 0.2 there is about 3× less motion per frame. At speed 1.0 the subject turns and walks across the frame; at speed 0.2 the same motion only begins within the same 5 s of playback. The frame count, the generated audio track and its length are identical across the two runs, matching the unit tests: only the positional time axis moves. Happy to share the clips if useful. |
yiyixuxu
left a comment
There was a problem hiding this comment.
thanks for the PR!
can you help me understand a bit more: can't we just pass 120 as frame_rate to the pipeline and 24 to encode_video()?
|
Thanks for the review! Passing
That's also how Lightricks wires this LoRA: their For silent video with an explicit |
|
Measured on the tiny
And with |
|
Hi @christopher5106, thanks for the PR! It does not appear to link an issue it fixes. If this PR addresses an existing issue, please add a closing keyword (e.g. Please note that PRs without a linked issue are likely to be automatically closed 10 days after this notice. Once the PR links an issue (or gets the |
… base from the playback frame rate LTX-2 places video tokens on a time axis in seconds (`pixel_frame / fps` in the RoPE coordinates), derived from `frame_rate`. Adapters such as the LTX-2.5 Slow-Motion-Control LoRA are conditioned on a different rate (`fps / speed`) than the video plays at. Add an optional `conditioning_frame_rate` to the LTX-2 classic pipelines and modular blocks: it drives the video, keyframe and reference coordinates and the transformer's `fps`, while the audio length, the duration head and the returned video keep following `frame_rate`. `None` resolves to `frame_rate`, leaving current outputs unchanged.
446cbaf to
05f0e34
Compare
|
Self-review against
|
|
thanks @christopher5106 |
Fixes #14980
What does this PR do?
LTX-2 places every video token on a time axis in seconds (
pixel_frame / fpsin the RoPE coordinates), and the pipelines derive thatfpsfromframe_rate, the playback rate. Some adapters need the two apart.Lightricks/LTX-2.5-22b-LoRA-Slow-Motion-Controlwas trained on high-speed footage labelled with its effective capture rate, and its reference driver conditions the diffusion stage onfps / speedwhile the video is still encoded atfps(distilled_speed_demo.py). The model card warns that "with a generic loader that ignores the speed value you get a plain LoRA and none of the intended motion control". There is no way to express that with the current pipelines: raisingframe_ratealso shortens the generated audio, changes the duration head's frame count, and is the rate callers encode the output at.This PR adds an optional
conditioning_frame_rate:LTX2*/LTX25*blocks): a newInputParamon the six blocks that place tokens in time. These areLTX2PrepareCoordsStep,LTX2ConditionPrepareCoordsStep,LTX2ConditionPrepareLatentsStep(keyframe coords),LTX2InContextPrepareLatentsStep(keyframe coords),LTX2ReferenceEncoderStep(reference coords) andLTX2LoopDenoiser(the transformer'sfps).LTX2Pipeline,LTX2ImageToVideoPipeline,LTX2ConditionPipeline,LTX2InContextPipeline): a new__call__argument applied to the same sites. It reaches the video coords, the transformer'sfpsand the keyframe/reference coords throughprepare_latents.Unchanged, still following
frame_rate: the audio latent length (num_frames / frame_rateseconds), the duration head's seconds → frames, and the returned video. Whenconditioning_frame_rateisNone(the default) it resolves toframe_rate, so existing calls are bit-identical.This mirrors the split the DFR pipelines already make internally (
conditioning_fpsvs playback fps inpipeline_ltx2_dfr.py), exposed as an argument instead of derived. The same caveat applies: the base model is trained around 24/25/30/60 fps (seeMAX_CONDITIONING_FPSinltx2/utils.py), so a conditioning rate far from those is only meaningful with an adapter trained for it, which the docstring says.Tests
tests/pipelines/ltx2/test_ltx2.py::test_conditioning_frame_rate_rescales_only_the_video_time_axisand the modular counterpart intests/modular_pipelines/ltx2/test_modular_pipeline_ltx25.py:conditioning_frame_rate=frame_ratereproduces the default output exactlyfpsmake fix-copies,utils/modular_auto_docstring.py(the 12 regenerated LTX-2 blockset docstrings),utils/check_forward_call_docstrings.py, ruff and doc-builder style are clean.Before submitting
InputParamdescriptions)Who can review?
@yiyixuxu @DN6 @sayakpaul (LTX-2 / modular)
🤖 Generated with Claude Code