Files
watch_sceen_ai/requirements.txt
EJClaw 63fcfb7ba2 feat(stt): real Korean STT via persistent faster-whisper worker
Step 3 (귀): add WhisperSTT + whisper_worker, a warm out-of-venv worker
mirroring the MeloTTS shape (whisper312 venv, small/int8 on CPU). transcribe()
closes the voice round trip (MeloTTS wav -> whisper text); utterances() turns an
injected audio_source into Utterances (Discord voice feed pending). Wired into
factory as WSAI_STT=whisper.

Also address the arbiter's TTS follow-ups:
- melo worker error handling: capture stderr (drained in a bounded background
  task so the pipe can't fill), surface the real failure cause, and defend
  against an empty/invalid ready line instead of dying on JSONDecodeError.
- pipeline pre-warm: load slow backends (warmup()) at startup so the first
  utterance is answered warm; a warmup failure is logged, not fatal.

Verified: real TTS->STT round trip recovers the sentence near-perfectly;
warm transcribe ~1.2s (CPU). 12 tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 18:29:32 +09:00

21 lines
843 B
Plaintext

# Core skeleton has NO required third-party deps (mock mode is pure stdlib).
# Install extras per backend you enable:
# --- screen capture (WSAI_SOURCE=mss) ---
# mss
# pillow
# --- cloud eyes + brain (WSAI_VISION=claude / WSAI_BRAIN=claude) ---
# anthropic
# --- voice backends (run in their OWN venvs; loaded as persistent workers) ---
# STT: faster-whisper (WSAI_STT=whisper) — installed in a dedicated venv, e.g.
# uv venv --python 3.12 /home/claude/jarvis-stt/whisper312
# uv pip install --python /home/claude/jarvis-stt/whisper312/bin/python faster-whisper
# (override interpreter/model via WSAI_WHISPER_PYTHON / WSAI_WHISPER_MODEL)
# TTS: MeloTTS (WSAI_TTS=melo) — in /home/claude/jarvis-tts/melo311
# (Opus decode for Discord voice audio, e.g. via the selfbot/ffmpeg path — pending)
# --- dev ---
# pytest