feat(tts): per-emotion voice controls (speed/word-gap/sentence-gap/pitch)

Each emotion can now be tuned independently. parse_segments resolves the four
controls per segment from a base (공통) dict plus an optional per-emotion
override; a missing override key inherits base. By default there are no
overrides, so every emotion delivers with the base values (모든 감정 = 기본값).

- emotion.py: Segment now carries all 4 controls + the canonical emotion name;
  parse_segments(text, base, overrides). Adds EMOTION_LABELS/EMOTIONS for the UI.
- melo.py: MeloTTS.emotion_overrides store; synth resolves per-segment controls
  and sends them per segment.
- melo_worker.py: _render applies each segment's own word_gap/sentence_gap/pitch
  (previously reply-global).
- dashboard.py: emotion dropdown in the TTS panel; GET returns base + overrides
  + emotion list; POST {emotion,...} stores an override (or {reset:true} clears
  it); base is set when emotion is omitted/"base".

Verified: an override on one emotion slows only that emotion (happy@0.7=5.66s vs
base 3.02s; sad unchanged at 3.06s); dashboard store/reset/base all work; 34
tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
EJClaw
2026-08-26 23:10:33 +09:00
parent 39b743d976
commit 067efc7abe
5 changed files with 242 additions and 99 deletions

View File

@@ -11,28 +11,27 @@ library chatter lands on stderr instead.
Protocol (one JSON object per line, on the protocol channel):
<- {"text": "...", "out": "/abs/path.wav", "speed": 1.3,
"word_gap": 0.25, "sentence_gap": 0.75, "pitch": 0.0}
<- {"segments": [{"text": "...", "speed": 1.3, "pitch": 2.0}, ...],
"out": "/abs/path.wav",
"word_gap": 0.25, "sentence_gap": 0.75, "pitch": 0.0}
"word_gap": -0.07, "sentence_gap": -0.30, "pitch": 0.0}
<- {"segments": [{"text": "...", "speed": 1.3, "word_gap": -0.07,
"sentence_gap": -0.30, "pitch": 2.0}, ...], "out": "/abs/path.wav"}
-> {"ok": true, "out": "/abs/path.wav", "ms": 123}
-> {"ok": false, "error": "..."}
On startup, once the model is ready, it emits exactly one line:
-> {"ready": true, "ms": <load-ms>, "device": "cpu"}
Four independent voice controls (ported from tts_site, see docs manual):
speed per-segment glyph speed -> generation-stage length_scale (1/speed)
word_gap reply-global, sec (-0.2..0.5): grow/shrink intra-sentence pauses
sentence_gap reply-global, sec (-0.5..1.5): insert/trim silence at sentence
boundaries (also used between emotion segments)
pitch reply-global semitone offset, added on top of each segment's own
``pitch`` (from its emotion tag; 0 == no shift)
Four independent voice controls (ported from tts_site, see docs manual). They
are PER SEGMENT: the caller resolves each segment's controls from the base
values plus that emotion's override, so different emotions can have different
speed/gaps/pitch within one reply.
speed glyph speed -> generation-stage length_scale (1/speed)
word_gap sec (-0.2..0.5): grow/shrink intra-sentence pauses
sentence_gap sec (-0.5..1.5): insert/trim silence at sentence boundaries
(within a segment, and before the next segment)
pitch semitone offset applied to the segment's wav (0 == no shift)
Each segment's ``pitch`` is a semitone offset applied to that segment's wav so
emotion tags can raise/lower the voice without changing the words. A reply is
split into sentences, each sentence synthesised at its segment's speed, and the
pieces concatenated with ``sentence_gap`` so one reply can carry several emotions
and a controllable rhythm.
A reply is split into sentences, each sentence synthesised at its segment's
speed, and the pieces concatenated with that segment's ``sentence_gap`` so one
reply can carry several emotions each with its own rhythm and pitch.
"""
import json
@@ -188,23 +187,21 @@ def main() -> None:
tts.tts_to_file(text, speaker_id, None, speed=speed), dtype=np.float32
)
def _render(segments, out, word_gap, sentence_gap, pitch_offset):
"""Render one wav from emotion segments applying the 4 controls.
Each segment carries its own glyph ``speed`` and ``pitch`` (from emotion
tags); ``word_gap``/``sentence_gap``/``pitch_offset`` are reply-global.
Sentences within a segment, and the segments themselves, are joined with
``sentence_gap`` (positive inserts silence, negative trims it)."""
word_gap = _clamp(word_gap, -0.2, 0.5)
sentence_gap = _clamp(sentence_gap, -0.5, 1.5)
pitch_offset = _clamp(pitch_offset, -12.0, 12.0)
def _render(segments, out):
"""Render one wav from emotion segments. Each segment carries its own
four controls (speed, word_gap, sentence_gap, pitch) — resolved on the
caller side from the base values plus that emotion's override. Sentences
within a segment join with that segment's ``sentence_gap``; between
segments the incoming segment's ``sentence_gap`` sets the pause."""
pieces = []
for seg in segments:
text = seg.get("text", "")
if not text.strip():
continue
speed = _clamp(seg.get("speed", 1.0), 0.5, 2.0)
pitch = _clamp(float(seg.get("pitch", 0.0)) + pitch_offset, -12.0, 12.0)
word_gap = _clamp(seg.get("word_gap", 0.0), -0.2, 0.5)
sentence_gap = _clamp(seg.get("sentence_gap", 0.0), -0.5, 1.5)
pitch = _clamp(seg.get("pitch", 0.0), -12.0, 12.0)
sent_pieces = []
for sent in split_sentences(text):
a = _synth_one(sent, speed)
@@ -254,14 +251,17 @@ def main() -> None:
if out.startswith("/tmp") or out.startswith("/dev/shm"):
raise ValueError(f"refusing RAM-backed tmpfs path: {out}")
s = time.monotonic()
word_gap = float(req.get("word_gap", 0.0))
sentence_gap = float(req.get("sentence_gap", 0.0))
pitch_offset = float(req.get("pitch", 0.0))
if "segments" in req:
_render(req["segments"], out, word_gap, sentence_gap, pitch_offset)
else: # legacy single-utterance form
seg = {"text": req["text"], "speed": float(req.get("speed", 1.0)), "pitch": 0.0}
_render([seg], out, word_gap, sentence_gap, pitch_offset)
_render(req["segments"], out)
else: # legacy single-utterance form: one segment carrying all controls
seg = {
"text": req["text"],
"speed": float(req.get("speed", 1.0)),
"word_gap": float(req.get("word_gap", 0.0)),
"sentence_gap": float(req.get("sentence_gap", 0.0)),
"pitch": float(req.get("pitch", 0.0)),
}
_render([seg], out)
ms = int((time.monotonic() - s) * 1000)
_emit({"ok": True, "out": out, "ms": ms})
except Exception as exc: # keep the worker alive across bad requests