feat(tts): per-emotion voice controls (speed/word-gap/sentence-gap/pitch)

Each emotion can now be tuned independently. parse_segments resolves the four
controls per segment from a base (공통) dict plus an optional per-emotion
override; a missing override key inherits base. By default there are no
overrides, so every emotion delivers with the base values (모든 감정 = 기본값).

- emotion.py: Segment now carries all 4 controls + the canonical emotion name;
  parse_segments(text, base, overrides). Adds EMOTION_LABELS/EMOTIONS for the UI.
- melo.py: MeloTTS.emotion_overrides store; synth resolves per-segment controls
  and sends them per segment.
- melo_worker.py: _render applies each segment's own word_gap/sentence_gap/pitch
  (previously reply-global).
- dashboard.py: emotion dropdown in the TTS panel; GET returns base + overrides
  + emotion list; POST {emotion,...} stores an override (or {reset:true} clears
  it); base is set when emotion is omitted/"base".

Verified: an override on one emotion slows only that emotion (happy@0.7=5.66s vs
base 3.02s; sad unchanged at 3.06s); dashboard store/reset/base all work; 34
tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
EJClaw
2026-08-26 23:10:33 +09:00
parent 39b743d976
commit 067efc7abe
5 changed files with 242 additions and 99 deletions

View File

@@ -107,6 +107,10 @@ class MeloTTS:
else os.environ.get("WSAI_TTS_SENTENCE_GAP", "-0.30"))
self.pitch = float(pitch if pitch is not None
else os.environ.get("WSAI_TTS_PITCH", "0.0"))
# Per-emotion control overrides: canonical emotion -> partial dict of
# {speed, word_gap, sentence_gap, pitch}. Empty by default, so every
# emotion inherits the base controls above (모든 감정 = 기본값).
self.emotion_overrides: dict[str, dict] = {}
self.sink = sink or _log_sink
self._proc: asyncio.subprocess.Process | None = None
self._lock = asyncio.Lock()
@@ -214,35 +218,36 @@ class MeloTTS:
per call (used by the dashboard preview so tuning does not disturb the
live bot voice until explicitly applied)."""
await self._ensure()
spd = self.speed if speed is None else float(speed)
wg = self.word_gap if word_gap is None else float(word_gap)
sg = self.sentence_gap if sentence_gap is None else float(sentence_gap)
pt = self.pitch if pitch is None else float(pitch)
base = {
"speed": self.speed if speed is None else float(speed),
"word_gap": self.word_gap if word_gap is None else float(word_gap),
"sentence_gap": self.sentence_gap if sentence_gap is None else float(sentence_gap),
"pitch": self.pitch if pitch is None else float(pitch),
}
text = normalize_for_speech(text)
# Split on [감정] tags: each tag steers pitch/speed for the text that
# Split on [감정] tags: each tag switches delivery for the text that
# follows (and is itself not spoken); non-emotion brackets stay as words.
segments = parse_segments(text, spd)
# Each segment carries its own 4 controls (base + per-emotion override).
segments = parse_segments(text, base, self.emotion_overrides)
self._n += 1
out = str(self.out_dir / f"tts-{self._n:06d}.wav")
if segments:
payload = {
"segments": [
{"text": s.text, "speed": s.speed, "pitch": s.pitch}
{"text": s.text, "speed": s.speed, "word_gap": s.word_gap,
"sentence_gap": s.sentence_gap, "pitch": s.pitch}
for s in segments
],
"out": out,
"word_gap": wg,
"sentence_gap": sg,
"pitch": pt,
}
else: # empty/whitespace reply: keep legacy single-utterance behaviour
payload = {
"text": text,
"out": out,
"speed": spd,
"word_gap": wg,
"sentence_gap": sg,
"pitch": pt,
"speed": base["speed"],
"word_gap": base["word_gap"],
"sentence_gap": base["sentence_gap"],
"pitch": base["pitch"],
}
req = json.dumps(payload)
s = time.monotonic()