Files
watch_sceen_ai/wsai/backends
EJClaw 80a83d9944 fix(tts): remove monster-voice artefact — disable librosa pitch shift, base speed 1.3→1.0
The emotional TTS path applied a librosa post-hoc pitch_shift (±1–3 semitones)
on top of a 1.3x-fast Melo base, producing a robotic "monster" delivery. Zero
out the pitch column for every emotion so pitch_shift is never invoked (the
worker's _pitch_shift already no-ops on 0.0), and drop the default synthesis
speed to 1.0. Emotion is now conveyed by speed alone — natural, artefact-free.
The pitch column is retained so a proper pitch method can be re-enabled later.

Verified: 29 tests pass; real 2-emotion synthesis on CUDA yields a clean wav
with pitch=0.0 on all segments.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 23:10:52 +09:00
..