The emotional TTS path applied a librosa post-hoc pitch_shift (±1–3 semitones) on top of a 1.3x-fast Melo base, producing a robotic "monster" delivery. Zero out the pitch column for every emotion so pitch_shift is never invoked (the worker's _pitch_shift already no-ops on 0.0), and drop the default synthesis speed to 1.0. Emotion is now conveyed by speed alone — natural, artefact-free. The pitch column is retained so a proper pitch method can be re-enabled later. Verified: 29 tests pass; real 2-emotion synthesis on CUDA yields a clean wav with pitch=0.0 on all segments. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
11 KiB
11 KiB