Compare commits

..

6 Commits

Author SHA1 Message Date
EJClaw
361dce70bb feat(tts): add female voice candidates (XTTS v2) + emotion tone samples
- samples/voice/: 6 XTTS v2 female studio voices speaking the same Korean
  line, as clearer-pronunciation alternatives to the single MeloTTS speaker
- samples/emotion/: per-[emotion] delivery clips from the current MeloTTS engine
- samples/README.md documents both sets and the tradeoffs

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 00:27:03 +09:00
EJClaw
e81ce1ac97 fix(tts): lower base speed 1.5→1.2 — 1.5x slurred Korean pronunciation
User reported the 1.5x base made Korean pronunciation mushy/slurred. Dial the
default WSAI_TTS_SPEED back to 1.2: still noticeably faster than the original
1.0, but clear. Emotion multipliers scale off base as before.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 00:18:47 +09:00
EJClaw
e67cd2e93f feat(tts): add [기본] neutral emotion + raise base speed to 1.5x
User reported the emotion samples sounded the same and the speed change wasn't
noticeable. Two changes:
- Add a "base" (기본/neutral) emotion at speed 1.0x so the plain base voice can
  be selected explicitly via a [기본] tag and auditioned against the others.
- Bump the WSAI_TTS_SPEED default 1.15 -> 1.5 for a clearly faster base voice.
  Emotion multipliers scale off base, so every emotion speeds up together.

Also extends gen_emotion_samples.py to emit one wav per emotion (incl. 기본)
plus a stitched all-in-one, so each emotion can be delivered as a separate clip.

Verified: 29 tests pass; match_emotion('기본') == 'base'; per-emotion synthesis
at base 1.5 produces distinct clip durations (base 4.9s vs happy 3.3s).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 00:08:21 +09:00
EJClaw
d410d4a6b5 feat(tts): bump base speed 1.0→1.15 for slightly faster delivery
User asked for a bit faster speech. Pitch shift is already disabled, so the
earlier "monster voice" risk from stacking pitch on a fast base is gone — a
modest 1.15x base is natural. Emotion multipliers scale off base, so every
emotion gets the bump too. Also adds tests/gen_emotion_samples.py, a dev
utility that synthesizes one clip per canonical emotion (announce name at
neutral speed, then a sample sentence steered by that emotion's tag) and
stitches them into a single wav for auditioning the emotion palette.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 23:54:31 +09:00
EJClaw
80a83d9944 fix(tts): remove monster-voice artefact — disable librosa pitch shift, base speed 1.3→1.0
The emotional TTS path applied a librosa post-hoc pitch_shift (±1–3 semitones)
on top of a 1.3x-fast Melo base, producing a robotic "monster" delivery. Zero
out the pitch column for every emotion so pitch_shift is never invoked (the
worker's _pitch_shift already no-ops on 0.0), and drop the default synthesis
speed to 1.0. Emotion is now conveyed by speed alone — natural, artefact-free.
The pitch column is retained so a proper pitch method can be re-enabled later.

Verified: 29 tests pass; real 2-emotion synthesis on CUDA yields a clean wav
with pitch=0.0 on all segments.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 23:10:52 +09:00
EJClaw
6e20f2cd79 fix: serialize worker warmup handshake to stop concurrent stdout reads
The STT/TTS worker _ensure() treated a spawned-but-not-yet-handshaked
subprocess as ready, so a voice turn arriving during warmup read the same
stdout StreamReader concurrently with the warmup handshake and crashed with
"readuntil() called while another coroutine is already waiting for incoming
data". Add a _start_lock + _ready flag so (re)start and the ready handshake
run atomically and callers wait for real readiness before reading stdout.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 22:43:43 +09:00
33 changed files with 336 additions and 106 deletions

32
samples/README.md Normal file
View File

@@ -0,0 +1,32 @@
# 음성 샘플 (voice / emotion samples)
브라우저나 로컬에서 들어보며 목소리·감정 톤을 고르기 위한 오디오 샘플 모음이다.
Discord로 올리는 대신 여기에 보관한다.
## voice/ — 다른 여자 목소리 후보 (XTTS v2)
현재 라이브 TTS(MeloTTS 한국어)는 화자가 하나뿐이라 "다른 여자 목소리"를 낼 수 없다.
대안으로 XTTS v2의 내장 여성 스튜디오 보이스로 같은 한국어 문장을 합성한 후보들이다.
문장은 모두 동일하다: "안녕하세요, 저는 새로운 목소리예요. 한국어 발음이 또렷하게
들리는지 한번 들어봐 주세요."
| 파일 | 화자 |
|------|------|
| xtts_01_ana_florence.mp3 | Ana Florence |
| xtts_02_daisy_studious.mp3 | Daisy Studious |
| xtts_03_sofia_hellen.mp3 | Sofia Hellen |
| xtts_04_alexandra_hisakawa.mp3 | Alexandra Hisakawa |
| xtts_05_nova_hogarth.mp3 | Nova Hogarth |
| xtts_06_rosemary_okafor.mp3 | Rosemary Okafor |
트레이드오프: XTTS는 MeloTTS보다 발음이 또렷하고 자연스러운 여성 음색을 고를 수 있지만,
합성이 더 무겁다(실시간 1초 예산과 충돌 가능). 라이브 채택 시 GPU 스트리밍으로 첫 소리
지연을 실측해 맞춰야 한다. 생성 스크립트: `/home/claude/jarvis-tts/gen_xtts_voices.py`.
## emotion/ — 감정별 톤 샘플 (현재 라이브 MeloTTS)
현재 엔진(MeloTTS 한국어)의 `[감정]` 태그별 델리버리를 하나씩 들어보는 샘플이다.
각 클립은 감정 이름을 기본 속도로 말한 뒤 그 감정 톤으로 예시 문장을 말한다.
감정 구분은 피치가 아니라 말 빠르기로만 표현된다(피치 변조는 잡음 때문에 비활성).
파일명이 감정을 그대로 담는다(예: `emo_02_happy.mp3` = 기쁨). 생성 스크립트:
`tests/gen_emotion_samples.py`.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

Binary file not shown.

View File

@@ -0,0 +1,83 @@
"""One-off: synthesize a short Korean sample for every canonical emotion.
Each clip announces the emotion name at neutral (base) speed, then speaks a
sample sentence steered by that emotion's ``[태그]`` — exactly the path the live
voice server uses. Writes one wav per emotion into an output directory (plus a
stitched all-in-one) so each emotion can be auditioned separately. Run with the
orchestrator venv:
.venv/bin/python -m tests.gen_emotion_samples /abs/out_dir
"""
from __future__ import annotations
import asyncio
import sys
import wave
from pathlib import Path
from wsai.backends.melo import MeloTTS
# (canonical english for filename, announced Korean name, emotion tag, sample)
SAMPLES: list[tuple[str, str, str, str]] = [
("base", "기본", "기본", "이건 기본 목소리예요, 감정 없이 이렇게 말해요."),
("happy", "기쁨", "기쁨", "오늘은 정말 기분 좋은 하루예요!"),
("excited", "신남", "신남", "우와, 이거 진짜 신난다! 빨리 하자!"),
("hopeful", "희망", "희망", "우리 분명히 잘 해낼 수 있어요!"),
("sad", "슬픔", "슬픔", "조금 속상한 일이 있었어요."),
("angry", "화남", "화남", "정말 너무하잖아요, 화가 나요."),
("fearful", "두려움", "두려움", "어떡하지, 너무 무서워요."),
("surprised", "놀람", "놀람", "어머, 이게 정말이에요?"),
("disgust", "혐오", "혐오", "으, 이건 좀 별로예요."),
("calm", "차분", "차분", "천천히 하나씩 정리해 볼게요."),
("friendly", "다정", "다정", "언제든지 편하게 말해 주세요."),
("serious", "진지", "진지", "이건 정말 중요한 이야기예요."),
("disappointed", "실망", "실망", "조금 아쉬운 결과네요."),
("tired", "피곤", "피곤", "아, 오늘 너무 피곤하네요."),
("affectionate", "사랑스럽게", "사랑스럽게", "당신은 정말 소중한 사람이에요."),
("playful", "장난스럽게", "장난스럽게", "히히, 한번 맞혀 보세요!"),
("curious", "궁금", "궁금", "그건 대체 왜 그런 걸까요?"),
("whisper", "속삭임", "속삭임", "조용히, 우리끼리만 아는 비밀이에요."),
("shout", "외침", "외침", "다 같이 힘내자, 파이팅!"),
("determined", "단호", "단호", "이번엔 반드시 해내겠어요."),
("relieved", "안도", "안도", "휴, 이제야 마음이 놓이네요."),
]
def stitch(paths: list[str], out: str, gap_s: float = 0.45) -> None:
with wave.open(paths[0], "rb") as w0:
nch, sw, fr = w0.getnchannels(), w0.getsampwidth(), w0.getframerate()
silence = b"\x00" * (int(fr * gap_s) * sw * nch)
with wave.open(out, "wb") as wo:
wo.setnchannels(nch)
wo.setsampwidth(sw)
wo.setframerate(fr)
for i, p in enumerate(paths):
with wave.open(p, "rb") as w:
wo.writeframes(w.readframes(w.getnframes()))
if i < len(paths) - 1:
wo.writeframes(silence)
async def main(out_dir: str) -> None:
d = Path(out_dir)
d.mkdir(parents=True, exist_ok=True)
tts = MeloTTS() # base speed comes from WSAI_TTS_SPEED default
await tts.warmup()
print(f"melo ready ({tts.load_ms} ms), base speed {tts.speed}")
paths: list[str] = []
for i, (canon, name, tag, sample) in enumerate(SAMPLES, 1):
text = f"{name}. [{tag}] {sample}"
src = await tts.synth(text)
dst = d / f"emo_{i:02d}_{canon}.wav"
Path(src).replace(dst)
paths.append(str(dst))
print(f" {name:8s} -> {dst}")
await tts.aclose()
stitch(paths, str(d / "emotion_samples_all.wav"))
print(f"wrote {len(paths)} per-emotion wavs + stitched all -> {d}")
if __name__ == "__main__":
out_dir = sys.argv[1] if len(sys.argv) > 1 else "/tmp/emotion_samples"
asyncio.run(main(out_dir))

View File

@@ -7,6 +7,7 @@ behaves with no source wired yet.
"""
import asyncio
import json
from typing import AsyncIterator
from wsai.backends.whisper import WhisperSTT
@@ -58,3 +59,75 @@ def test_empty_transcript_is_skipped(monkeypatch):
utts = _collect(stt)
assert [u.text for u in utts] == ["안녕"]
def test_request_during_warmup_does_not_overlap_stdout(monkeypatch):
"""Regression: a transcribe() arriving while warmup() is still awaiting the
worker's ready line must NOT read the same stdout StreamReader concurrently.
Before the fix, _ensure()'s fast path returned as soon as the subprocess was
spawned (proc set, returncode None) even though the ready handshake was still
in flight, so the request's stdout.readline() overlapped warmup's and asyncio
raised "readuntil() called while another coroutine is already waiting for
incoming data" — the exact crash seen in the Discord voice server."""
async def run():
stt = WhisperSTT()
stdout = asyncio.StreamReader()
stderr = asyncio.StreamReader()
stderr.feed_eof() # nothing on stderr; let the drain task finish cleanly
class FakeStdin:
def write(self, _b):
pass
async def drain(self):
pass
class FakeProc:
returncode = None
def __init__(self):
self.stdin = FakeStdin()
self.stdout = stdout
self.stderr = stderr
def terminate(self):
self.returncode = 0
async def wait(self):
return 0
spawns = []
async def fake_create(*_a, **_k):
spawns.append(1)
return FakeProc()
monkeypatch.setattr(asyncio, "create_subprocess_exec", fake_create)
# warmup enters _ensure and blocks awaiting the ready line on stdout.
warm = asyncio.create_task(stt.warmup())
await asyncio.sleep(0.05)
# A concurrent request lands mid-warmup. It must wait for readiness, not
# crash and not read stdout yet.
tr = asyncio.create_task(stt.transcribe("x.wav"))
await asyncio.sleep(0.05)
assert not tr.done() # blocked on the start lock, no overlapping read
# Complete the handshake -> warmup finishes and releases the request.
stdout.feed_data(
(json.dumps({"ready": True, "ms": 1, "device": "cpu"}) + "\n").encode()
)
await asyncio.wait_for(warm, timeout=1)
await asyncio.sleep(0.02)
stdout.feed_data(
(json.dumps({"ok": True, "text": "안녕", "ms": 2}) + "\n").encode()
)
assert await asyncio.wait_for(tr, timeout=1) == "안녕"
assert sum(spawns) == 1 # one worker, not one-per-concurrent-caller
await stt.aclose()
asyncio.run(asyncio.wait_for(run(), timeout=5))

View File

@@ -23,27 +23,34 @@ from dataclasses import dataclass
# Canonical emotion -> (speed multiplier relative to base, pitch shift in semitones).
# Kept deliberately modest so delivery stays natural, not cartoonish.
#
# Pitch is held at 0.0 for every emotion: librosa's post-hoc pitch_shift on Melo
# output produced a robotic, "monster"-sounding artefact (worse when stacked on a
# fast base speed). Emotion is therefore conveyed by speed only — natural and
# artefact-free. The pitch column is retained (rather than removed) so the effect
# can be re-enabled per-emotion later with a real, artefact-free pitch method.
EMOTION_PARAMS: dict[str, tuple[float, float]] = {
"happy": (1.08, 2.0), # 기쁨 / cheerful
"excited": (1.15, 3.0), # 신남 / excited
"hopeful": (1.10, 1.5), # 희망 / 힘차게
"sad": (0.90, -2.5), # 슬픔 / sad
"angry": (1.12, 1.0), # 화남 / angry
"fearful": (1.12, 2.0), # 두려움 / terrified
"surprised": (1.05, 3.0), # 놀람 / surprise
"disgust": (0.96, -1.0), # 혐오 / disgust
"calm": (0.95, -1.0), # 차분 / calm
"friendly": (1.00, 1.0), # 다정 / friendly
"serious": (0.97, -1.0), # 진지 / serious
"disappointed":(0.92, -2.0), # 실망 / disappointed
"tired": (0.90, -2.0), # 피곤 / 지침
"affectionate":(0.98, 1.0), # 사랑스럽게 / affectionate
"playful": (1.08, 2.0), # 장난스럽게 / playful
"whisper": (0.92, -1.5), # 속삭임 / whispering
"shout": (1.05, 2.5), # 외침 / shouting
"determined": (1.05, 0.5), # 단호 / determined
"relieved": (0.95, 0.5), # 안도 / relieved
"curious": (1.03, 1.5), # 궁금 / curious
"base": (1.00, 0.0), # 기본 / neutral — the plain base voice, no colour
"happy": (1.08, 0.0), # 기쁨 / cheerful
"excited": (1.15, 0.0), # 신남 / excited
"hopeful": (1.10, 0.0), # 희망 / 힘차게
"sad": (0.90, 0.0), # 슬픔 / sad
"angry": (1.12, 0.0), # 화남 / angry
"fearful": (1.12, 0.0), # 두려움 / terrified
"surprised": (1.05, 0.0), # 놀람 / surprise
"disgust": (0.96, 0.0), # 혐오 / disgust
"calm": (0.95, 0.0), # 차분 / calm
"friendly": (1.00, 0.0), # 다정 / friendly
"serious": (0.97, 0.0), # 진지 / serious
"disappointed":(0.92, 0.0), # 실망 / disappointed
"tired": (0.90, 0.0), # 피곤 / 지침
"affectionate":(0.98, 0.0), # 사랑스럽게 / affectionate
"playful": (1.08, 0.0), # 장난스럽게 / playful
"whisper": (0.92, 0.0), # 속삭임 / whispering
"shout": (1.05, 0.0), # 외침 / shouting
"determined": (1.05, 0.0), # 단호 / determined
"relieved": (0.95, 0.0), # 안도 / relieved
"curious": (1.03, 0.0), # 궁금 / curious
}
# Every spelling Claude might realistically emit, mapped to a canonical emotion.
@@ -63,6 +70,7 @@ def _norm(word: str) -> str:
return re.sub(r"\s+", "", word).lower()
_register("base", "기본", "기본목소리", "기본톤", "보통", "평범", "무감정", "default", "neutral", "normal", "plain")
_register("happy", "기쁨", "기쁘게", "기뻐", "기뻐하며", "행복", "행복하게", "행복하게도", "즐겁게", "즐거움", "밝게", "반가움", "반갑게", "반가워", "cheerful", "happy", "joyful")
_register("excited", "신남", "신나게", "신나서", "흥분", "들뜬", "들떠서", "설렘", "설레며", "excited", "thrilled")
_register("hopeful", "희망", "희망차게", "힘차게", "힘내", "힘내서", "응원", "응원하며", "격려", "hopeful", "encouraging")

View File

@@ -12,7 +12,7 @@ Env:
WSAI_MELO_DEVICE cpu | cuda | auto (default auto: GPU if torch sees one,
else CPU; the worker falls back to CPU if CUDA fails)
WSAI_TTS_OUT_DIR where wavs are written (default ~/.cache/wsai/tts)
WSAI_TTS_SPEED synthesis speed multiplier (default 1.3)
WSAI_TTS_SPEED synthesis speed multiplier (default 1.2)
"""
from __future__ import annotations
@@ -93,10 +93,15 @@ class MeloTTS:
self.out_dir = Path(out_dir or os.environ.get("WSAI_TTS_OUT_DIR")
or (Path.home() / ".cache/wsai/tts"))
self.speed = float(speed if speed is not None
else os.environ.get("WSAI_TTS_SPEED", "1.3"))
else os.environ.get("WSAI_TTS_SPEED", "1.2"))
self.sink = sink or _log_sink
self._proc: asyncio.subprocess.Process | None = None
self._lock = asyncio.Lock()
# Serialises worker (re)start + the ready handshake so a caller that
# arrives mid-warmup waits for readiness instead of reading the same
# stdout StreamReader concurrently (asyncio forbids overlapping reads).
self._start_lock = asyncio.Lock()
self._ready = False # True only after the ready handshake completes
self._n = 0
self.load_ms: int | None = None
# Keep the worker's most recent stderr lines so a crash reports its real
@@ -132,8 +137,18 @@ class MeloTTS:
await self._ensure()
async def _ensure(self) -> None:
if self._proc is not None and self._proc.returncode is None:
# Fast path: only skip when the worker is not just spawned but fully
# handshaked. Checking `_proc` alone would let a caller sail past while
# another coroutine (e.g. warmup) is still awaiting the ready line on
# this same stdout, causing overlapping StreamReader reads.
if self._proc is not None and self._proc.returncode is None and self._ready:
return
async with self._start_lock:
# Re-check under the lock: another coroutine may have finished the
# (re)start + handshake while we waited.
if self._proc is not None and self._proc.returncode is None and self._ready:
return
self._ready = False
self.out_dir.mkdir(parents=True, exist_ok=True)
env = {**os.environ, "WSAI_MELO_DEVICE": self.device}
# Run the worker module from the wsai source tree with the melo venv.
@@ -167,6 +182,7 @@ class MeloTTS:
f"melo worker failed to start: {info}.{self._stderr_hint()}"
)
self.load_ms = info.get("ms")
self._ready = True
log.info("melo worker ready in %s ms on %s", self.load_ms, info.get("device"))
async def synth(self, text: str) -> str:
@@ -223,3 +239,4 @@ class MeloTTS:
pass
self._stderr_task = None
self._proc = None
self._ready = False

View File

@@ -75,6 +75,11 @@ class WhisperSTT:
self.audio_source = audio_source
self._proc: asyncio.subprocess.Process | None = None
self._lock = asyncio.Lock()
# Serialises worker (re)start + the ready handshake so a caller that
# arrives mid-warmup waits for readiness instead of reading the same
# stdout StreamReader concurrently (asyncio forbids overlapping reads).
self._start_lock = asyncio.Lock()
self._ready = False # True only after the ready handshake completes
self.load_ms: int | None = None
self.resolved_device: str | None = None # "cuda" | "cpu", known after start
# Keep the worker's most recent stderr so a crash reports its real cause
@@ -106,8 +111,18 @@ class WhisperSTT:
await self._ensure()
async def _ensure(self) -> None:
if self._proc is not None and self._proc.returncode is None:
# Fast path: only skip when the worker is not just spawned but fully
# handshaked. Checking `_proc` alone would let a caller sail past while
# another coroutine (e.g. warmup) is still awaiting the ready line on
# this same stdout, causing overlapping StreamReader reads.
if self._proc is not None and self._proc.returncode is None and self._ready:
return
async with self._start_lock:
# Re-check under the lock: another coroutine may have finished the
# (re)start + handshake while we waited.
if self._proc is not None and self._proc.returncode is None and self._ready:
return
self._ready = False
env = {
**os.environ,
"WSAI_WHISPER_MODEL": self.model,
@@ -155,6 +170,7 @@ class WhisperSTT:
)
self.load_ms = info.get("ms")
self.resolved_device = info.get("device")
self._ready = True
log.info(
"whisper worker ready in %s ms on %s (model %s)",
self.load_ms, info.get("device"), info.get("model"),
@@ -211,3 +227,4 @@ class WhisperSTT:
pass
self._stderr_task = None
self._proc = None
self._ready = False