feat(tts): real Korean TTS via persistent MeloTTS worker
Adds a MeloTTS backend that runs the model in its own melo311 interpreter as a long-lived worker (melo_worker.py), loaded once and fed synthesis requests over a stdin/stdout JSON protocol. fd1 is split from fd2 in the worker so MeloTTS's stdout progress chatter can't corrupt the protocol. Each speak() writes a wav and hands the path to a pluggable sink (the Discord voice step will swap in "play into the call"). factory wires tts=melo; pipeline.aclose now also tears down the tts worker. Verified (CPU): model load ~7.9s once, then a short reply synthesizes in ~0.86s (within the ~1s budget); wav is valid 44.1kHz PCM. GPU (cuda) is selectable via WSAI_MELO_DEVICE for lower latency, pending GPU approval. 7 smoke tests still pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -77,6 +77,10 @@ def _stt(s: Settings):
|
||||
def _tts(s: Settings):
|
||||
if s.tts in (None, "none"):
|
||||
return None
|
||||
if s.tts == "melo":
|
||||
from .backends.melo import MeloTTS
|
||||
|
||||
return MeloTTS()
|
||||
from .backends.mock import MockTTS
|
||||
|
||||
return MockTTS()
|
||||
|
||||
Reference in New Issue
Block a user