feat(dashboard): live TTS voice controls (speed/word-gap/sentence-gap/pitch)

Adds a "봇 목소리(TTS) 조절" panel to the voice-server dashboard so the four
controls can be tuned from the browser instead of only via env vars:

- 4 sliders (glyph speed, word gap, sentence gap, pitch) with live labels
- 미리듣기: synthesises a sample with the slider values WITHOUT changing the
  live bot voice (new per-call overrides on MeloTTS.synth)
- 봇에 적용: commits the slider values to the live TTS instance; next reply uses
  them. 기본값 button resets to the manual defaults.

Backend: GET/POST /api/tts/settings (clamped to the manual ranges) and POST
/api/tts/preview (returns audio/wav). Panel shows only when tts is a real
(non-mock) backend.

Verified end-to-end against a live dashboard instance: page renders the panel,
GET returns defaults, POST applies+clamps (pitch 99->12), preview returns a
valid wav and leaves the live settings unchanged; 29 tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
EJClaw
2026-08-26 22:48:00 +09:00
parent 0b92284ff8
commit 2ce2806102
2 changed files with 215 additions and 10 deletions

View File

@@ -198,14 +198,30 @@ class MeloTTS:
self._ready = True
log.info("melo worker ready in %s ms on %s", self.load_ms, info.get("device"))
async def synth(self, text: str) -> str:
async def synth(
self,
text: str,
*,
speed: float | None = None,
word_gap: float | None = None,
sentence_gap: float | None = None,
pitch: float | None = None,
) -> str:
"""Synthesize `text` to a wav and return its path (no sink). Reusable by
callers that want the wav directly (e.g. the Discord voice bridge)."""
callers that want the wav directly (e.g. the Discord voice bridge).
The four controls default to the instance settings but may be overridden
per call (used by the dashboard preview so tuning does not disturb the
live bot voice until explicitly applied)."""
await self._ensure()
spd = self.speed if speed is None else float(speed)
wg = self.word_gap if word_gap is None else float(word_gap)
sg = self.sentence_gap if sentence_gap is None else float(sentence_gap)
pt = self.pitch if pitch is None else float(pitch)
text = normalize_for_speech(text)
# Split on [감정] tags: each tag steers pitch/speed for the text that
# follows (and is itself not spoken); non-emotion brackets stay as words.
segments = parse_segments(text, self.speed)
segments = parse_segments(text, spd)
self._n += 1
out = str(self.out_dir / f"tts-{self._n:06d}.wav")
if segments:
@@ -215,18 +231,18 @@ class MeloTTS:
for s in segments
],
"out": out,
"word_gap": self.word_gap,
"sentence_gap": self.sentence_gap,
"pitch": self.pitch,
"word_gap": wg,
"sentence_gap": sg,
"pitch": pt,
}
else: # empty/whitespace reply: keep legacy single-utterance behaviour
payload = {
"text": text,
"out": out,
"speed": self.speed,
"word_gap": self.word_gap,
"sentence_gap": self.sentence_gap,
"pitch": self.pitch,
"speed": spd,
"word_gap": wg,
"sentence_gap": sg,
"pitch": pt,
}
req = json.dumps(payload)
s = time.monotonic()