feat(stt): real Korean STT via persistent faster-whisper worker
Step 3 (귀): add WhisperSTT + whisper_worker, a warm out-of-venv worker mirroring the MeloTTS shape (whisper312 venv, small/int8 on CPU). transcribe() closes the voice round trip (MeloTTS wav -> whisper text); utterances() turns an injected audio_source into Utterances (Discord voice feed pending). Wired into factory as WSAI_STT=whisper. Also address the arbiter's TTS follow-ups: - melo worker error handling: capture stderr (drained in a bounded background task so the pipe can't fill), surface the real failure cause, and defend against an empty/invalid ready line instead of dying on JSONDecodeError. - pipeline pre-warm: load slow backends (warmup()) at startup so the first utterance is answered warm; a warmup failure is logged, not fatal. Verified: real TTS->STT round trip recovers the sentence near-perfectly; warm transcribe ~1.2s (CPU). 12 tests pass. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -18,7 +18,7 @@ from dataclasses import dataclass
|
||||
class Settings:
|
||||
source: str | None = "mock" # mock | mss | None (eyes-free)
|
||||
vision: str | None = "mock" # mock | claude | None (eyes-free)
|
||||
stt: str | None = "mock" # mock | (whisper) | None
|
||||
stt: str | None = "mock" # mock | whisper | None
|
||||
tts: str | None = "mock" # mock | melo | None
|
||||
brain: str = "mock" # mock | claude
|
||||
text: str | None = None # None | (discord)
|
||||
|
||||
Reference in New Issue
Block a user