Ports the 4 independent voice controls documented in tts_site's manual into the
MeloTTS backend, layered on top of the existing emotion-tag segments:
- glyph speed: generation-stage length_scale (existing speed path); default
bumped 1.2 -> 1.25 to match the manual (still well below the 1.5 that slurred)
- word_gap (sec, -0.2..0.5): scale intra-sentence silences after natural synth
- sentence_gap (sec, -0.5..1.5): insert/trim silence at sentence boundaries,
replacing the old fixed 120ms inter-segment gap
- pitch (semitones, -12..12): global offset added on top of per-emotion pitch
Each reply is split into sentences, synthesised per-sentence at its segment's
speed, word_gap applied, joined with sentence_gap, then pitch-shifted. Exposed
via WSAI_TTS_SPEED/WORD_GAP/SENTENCE_GAP/PITCH env vars and MeloTTS ctor args.
Verified by real synthesis: each control independently changes wav duration
(speed, sentence_gap, word_gap-with-pauses, pitch); 29 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>