Skip to content

fix: align SD3 encoder token chunks before conditioning - #2111

Merged
leejet merged 1 commit into
leejet:masterfrom
srelus:fix/sd3-encoder-token-alignment
Oct 8, 2026
Merged

leejet merged 1 commit into
leejet:masterfrom
srelus:fix/sd3-encoder-token-alignment

Conversation

@srelus

@srelus srelus commented Oct 8, 2026

Copy link
Copy Markdown

Summary

Fix out-of-bounds token and weight slicing in SD3/SD3.5 conditioning when active encoders produce different numbers of 77-token chunks.

For example, CLIP-L and CLIP-G can produce 462 tokens while T5 produces 539. The conditioning loop processes seven chunks, reading past the end of both CLIP sequences on the seventh chunk.

Pad shorter active sequences with tokenizer-generated empty chunks before slicing. This preserves the full prompt, existing weights, and encoder-specific special tokens. Equal-length sequences remain unchanged.

Related Issue / Discussion

Related to #1105.

Additional Information

  • Eight isolated C++ regression cases passed on base commit a1ded76 with AddressSanitizer and UndefinedBehaviorSanitizer (leak detection disabled).
  • Removing the fix makes the 462/462/539 regression case fail.
  • git diff --check passed.
  • The same padding algorithm was tested successfully with full SD3.5 medium prompts on Linux RTX 4090 FE (59c23bc) and Windows RTX 3080 Ti (07a85c7).
  • Full backend compilation and GPU generation on a1ded76 have not been run.
  • Only conditioner.hpp is changed; no test scripts, binaries, models, or dependency updates are included.

Checklist

@leejet
leejet merged commit e16d26a into leejet:master Oct 8, 2026
10 checks passed
@leejet

leejet commented Oct 8, 2026

Copy link
Copy Markdown
Owner

Thank you for your contribution.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants