Skip to content

Release 0.61.2: improve dubbed speech timing and reliability - #314

Merged
BartWojtowicz merged 1 commit into
mainfrom
perf/dubbing-speaker-cloning
Sep 9, 2026
Merged

BartWojtowicz merged 1 commit into
mainfrom
perf/dubbing-speaker-cloning

Conversation

@BartWojtowicz

Copy link
Copy Markdown
Owner

No description provided.

Anchor phrases to source words, preserve complete generated speech at natural
speed when it fits, and smooth phrase boundaries with 5 ms fades. Report
necessary speedups and translation or synthesis failures explicitly.

Bound translation and speech requests, validate model output, guard speech
tokens, and reduce CUDA decoder overhead. Support audio-only MP4 inputs and
select dubbed audio for default playback when retaining the original track.

Require the published Chatterbox short-text fix. Document timing migrations,
the tested TranslateGemma 12B configuration, and measured quality and
performance limits without changing the default translation model.

Validation: 1,285 tests; final translator checks; real-model dubbing with
diarization and cloned voices; pre-commit; strict docs; lock validation;
wheel/sdist builds and clean-wheel render/import/MCP smoke checks.
@BartWojtowicz BartWojtowicz self-assigned this Sep 9, 2026
@BartWojtowicz
BartWojtowicz merged commit 9de4358 into main Sep 9, 2026
22 checks passed
@BartWojtowicz
BartWojtowicz deleted the perf/dubbing-speaker-cloning branch September 9, 2026 16:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant