Repository navigation
[TPU] TorchTPU backend integration - eager / torch.compile / tp - #14039
JingyaHuang wants to merge 74 commits into
Conversation
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
…ster tpu.md; drop redundant execution_device check - Propagate the text_encoder.device-based fix (introduced for TPU CPU-offload support) from FluxPipeline/Flux2KleinPipeline/WanPipeline into their `# Copied from` copies (flux/*, flux2_klein_inpaint, visualcloze, anyflow, chronoedit, lucy_edit, skyreels_v2/*). SDXL-family copies of StableDiffusionXLPipeline.encode_prompt are intentionally left untouched; they'll be handled in a follow-up PR that fixes device placement for every pipeline component (not just text encoders). - Register docs/source/en/optimization/tpu.md in _toctree.yml (was breaking the docs build: "not present in the table of contents"). - Remove the redundant "prefer non-CPU, non-meta component" loop from DiffusionPipeline._execution_device: PR huggingface#14383 already fixed this in DiffusionPipeline.device, which _execution_device falls back to. Verified on TPU hardware that _execution_device still resolves correctly for a split-placement pipeline after the removal. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
There was a problem hiding this comment.
Nice, thanks! The changes are minimal and we should be able to merge soon.
(Ran the updated tests on 2 A10Gs and they worked: https://huggingface.co/jobs/sayakpaul/6ac30cecfbc85ba6823a4556)
| image.save("output.png") | ||
| ``` | ||
|
|
||
| If the text encoder alone is too large for a single chip(eg. FLUX.2-dev's Mistral-3-Small is ~45GB), |
There was a problem hiding this comment.
I think we should provide a working example here. Or we could use transformers' support for TP?
Also, the same argument (a model to be placed on a TPU chip is too large) applies to the DiT. I think we could make an entirely separate section to talk about TP, etc.
Or, we could include a statement like "For details on how to shard the denoising module, refer to the "## Tensor parallelism" section."
There was a problem hiding this comment.
Yeah, it would be better just show a simple example with everything ran on tpu w/o tp in this section, and have an example with TP on the TP dedicated section.
sayakpaul
left a comment
There was a problem hiding this comment.
Thanks! I left some further comments.
Co-authored-by: Steven Liu <59462357+stevhliu@users.noreply.github.com>
…n the TPU examples
…step, thus we can use it for the tpu.md compile mode snippet
…/diffusers into add-torchtpu-support
What does this PR do?
Need the fix #14739 and perferrably merge the sharding improvement PR #14544 first.
torch.compileThis is a preparation based on TorchTPU beta before the official release.
Who can review?
Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.