diff --git a/docs/TTS_속도_간격_적용_매뉴얼.md b/docs/TTS_속도_간격_적용_매뉴얼.md new file mode 100644 index 0000000..f89b0bc --- /dev/null +++ b/docs/TTS_속도_간격_적용_매뉴얼.md @@ -0,0 +1,290 @@ +# TTS 속도·간격·피치 제어 적용 매뉴얼 (watch_screen_ai 이식용) + +이 문서는 tts_site에 구현된 4개 음성 제어(글자 속도 · 단어 간격 · 문장 간격 · 피치)를 +다른 프로젝트(예: `watch_screen_ai`)에 그대로 이식하기 위한 매뉴얼이다. +MeloTTS(한국어) 기준의 레퍼런스 구현과, API/UI 계약, 적용 절차, 주의점을 담는다. + +- 원본 구현: `tts_site` (repo `git.tkrmagid.kr/tkrmagid/tts_site`), 커밋 `5ed2d5f` 기준 +- 핵심 파일: `app/engines/melo_engine.py`, `app/server.py`, `frontend/index.html`, `frontend/app.js` + +--- + +## 1. 4개 컨트롤 요약 + +| 컨트롤 | 파라미터 | 범위 | 기본값 | 동작 원리 | +|---|---|---|---|---| +| 글자 속도 | `speed` | 0.5 ~ 2.0 | 1.25 | 생성 단계 `length_scale = 1/speed` 로 각 음절 발화 길이 조절 (피치 보존) | +| 단어 간격 | `word_gap` (초) | -0.2 ~ 0.5 | 0.25 | 자연 합성 결과의 **어절 사이 무음**을 늘리거나(양수) 줄임(음수) | +| 문장 간격 | `sentence_gap` (초) | -0.5 ~ 1.5 | 0.75 | 문장 경계에 무음 삽입(양수) 또는 경계 무음 트리밍(음수) | +| 피치 | `pitch` (반음) | -12 ~ +12 | 0 | `librosa.effects.pitch_shift` (선택) | + +핵심 설계 원칙: +- **글자 속도**는 후처리 배속(atempo)이 아니라 **생성 단계 length_scale**로 처리한다. + 각 음절("안·녕·하")의 발음 자체가 빨라지고 느려지며, 배속 특유의 뭉개짐/금속성이 없다. +- **단어/문장 간격은 글자 속도와 독립**이다. 간격은 삽입/삭제한 "초" 단위 그대로 유지되고 + 발화 속도 변경에 영향받지 않는다. +- **음수 간격**은 경계/어절 사이 무음을 잘라 자연 상태(0)보다 **더 촘촘하게** 만든다. + +--- + +## 2. 왜 이렇게 하는가 (원리) + +### 2.1 글자 속도 = length_scale (후처리 배속 금지) +빠르게 만드는 방법은 크게 3가지이고 왜곡 성격이 다르다. +- 단순 리샘플링: 피치까지 올라가 "다람쥐 소리" → **사용 금지** +- 후처리 타임스트레치(atempo/rubberband): 피치는 보존되나 자음 뭉개짐·금속성 아티팩트 +- **생성 단계 length_scale**: 모델이 처음부터 빠른/느린 발화를 합성 → 음색 보존, 아티팩트 없음 → **채택** + +MeloTTS는 `tts_to_file(..., speed=s)` 내부에서 `length_scale = 1/s`로 모든 음소 길이를 스케일한다. + +### 2.2 간격은 오디오 레벨에서 (무음 삽입/트리밍) +- 문장 간격: 문장을 개별 합성 → 사이에 무음(zeros)을 넣거나(양수), 경계 무음을 잘라낸다(음수). +- 단어 간격: 문장을 **통째로 자연 합성**(억양 보존)한 뒤, 내부 어절 무음 구간만 늘리거나 줄인다. + - 어절을 개별 합성하면 억양이 밋밋해지고 자연 상태보다 더 붙일 수 없으므로 쓰지 않는다. + - 자음 폐쇄음 같은 짧은 무음(< 80ms)과 문장 앞뒤 경계 무음은 건드리지 않아 발음이 안전하다. + +--- + +## 3. 레퍼런스 구현 (그대로 복붙 가능) + +MeloTTS `TTS` 객체(`from melo.api import TTS`)와 `numpy`만 있으면 된다. +`model.tts_to_file(text, speaker_id, output_path=None, speed=s, quiet=True)` 는 +`float32` 파형(numpy)을 반환하고, `model.hps.data.sampling_rate` 가 sr 이다. + +```python +import re +import numpy as np + +# ── 문장 분리 ──────────────────────────────────────────────── +_SENT_SPLIT_RE = re.compile(r"(?<=[.!?。!?…])\s+|\n+") + +def split_sentences(text: str) -> list[str]: + parts = [p.strip() for p in _SENT_SPLIT_RE.split(text) if p and p.strip()] + return parts or ([text.strip()] if text.strip() else []) + +# ── 경계 무음 트리밍(문장 간격 음수용) ────────────────────────── +def trim_end_silence(a, sr, max_sec, thresh=0.02): + n = a.size + if n == 0 or max_sec <= 0: + return a, 0.0 + max_n = min(n, int(sr * max_sec)) + if max_n <= 0: + return a, 0.0 + tail = np.abs(a[n - max_n:]) + nz = np.where(tail >= thresh)[0] + cut = max_n if nz.size == 0 else (max_n - 1 - int(nz[-1])) + return (a, 0.0) if cut <= 0 else (a[: n - cut], cut / sr) + +def trim_start_silence(a, sr, max_sec, thresh=0.02): + n = a.size + if n == 0 or max_sec <= 0: + return a, 0.0 + max_n = min(n, int(sr * max_sec)) + if max_n <= 0: + return a, 0.0 + head = np.abs(a[:max_n]) + nz = np.where(head >= thresh)[0] + cut = max_n if nz.size == 0 else int(nz[0]) + return (a, 0.0) if cut <= 0 else (a[cut:], cut / sr) + +# ── 조각 이어붙이기(문장 간격: 양수=무음삽입, 음수=경계 트리밍) ────── +def append_unit(pieces, unit, gap, sr): + if not pieces: + pieces.append(unit); return + if gap >= 0: + if gap > 1e-4: + pieces.append(np.zeros(int(sr * gap), dtype=np.float32)) + pieces.append(unit); return + budget = -gap + prev, removed = trim_end_silence(pieces[-1], sr, budget) + pieces[-1] = prev + budget -= removed + if budget > 1e-4: + unit, _ = trim_start_silence(unit, sr, budget) + pieces.append(unit) + +# ── 단어 간격: 문장 내부 무음 스케일링(양수=늘림, 음수=줄임) ───────── +def scale_word_gaps(a, sr, delta, thresh=0.02, min_pause=0.08): + n = a.size + if abs(delta) < 1e-4 or n == 0: + return a + silent = np.abs(a) < thresh + changes = np.flatnonzero(np.diff(silent.astype(np.int8)) != 0) + 1 + bounds = [0, *changes.tolist(), n] + min_n = int(sr * min_pause) + floor_n = int(sr * 0.015) + add_n = int(delta * sr) + out, last = [], len(bounds) - 2 + for k in range(len(bounds) - 1): + s, e = bounds[k], bounds[k + 1] + seg = a[s:e] + is_internal = 0 < k < last # 앞/뒤 경계 무음은 제외 + if silent[s] and is_internal and (e - s) >= min_n: + new_n = max(floor_n, (e - s) + add_n) + if new_n >= (e - s): + seg = np.concatenate([seg, np.zeros(new_n - (e - s), dtype=np.float32)]) + else: + seg = seg[:new_n] + out.append(seg) + return np.concatenate(out) if out else a + +# ── 메인 렌더 함수 ────────────────────────────────────────────── +def render(model, text, speaker_id, *, + speed=1.25, word_gap=0.25, sentence_gap=0.75, pitch=0.0): + """model: melo.api.TTS 인스턴스. float32 파형(numpy)과 sr 을 만든다.""" + speed = float(max(0.5, min(2.0, speed))) + word_gap = float(max(-0.2, min(0.5, word_gap))) + sentence_gap = float(max(-0.5, min(1.5, sentence_gap))) + pitch = float(max(-12., min(12., pitch))) + sr = model.hps.data.sampling_rate + + pieces, first = [], True + for sent in (split_sentences(text) or [text]): + # 글자 속도 = 생성 단계 length_scale + a = np.asarray( + model.tts_to_file(sent, speaker_id, output_path=None, speed=speed, quiet=True), + dtype=np.float32, + ) + if abs(word_gap) > 1e-4: + a = scale_word_gaps(a, sr, word_gap) + if first: + pieces.append(a); first = False + else: + append_unit(pieces, a, sentence_gap, sr) + + audio = np.concatenate(pieces) if pieces else np.zeros(1, dtype=np.float32) + + if abs(pitch) > 1e-3: + import librosa + audio = librosa.effects.pitch_shift(audio, sr, n_steps=pitch) + return audio, sr +``` + +WAV(PCM16) 바이트로 만들려면: +```python +import io, soundfile as sf +buf = io.BytesIO(); sf.write(buf, audio, sr, format="WAV", subtype="PCM_16"); buf.seek(0) +wav_bytes = buf.read() +``` + +> AUTO(한/영/일/중 혼합) 언어 분기가 필요하면 원본 `melo_engine.py`의 `split_by_language` +> + `_synth_natural` 를 참고. 단일 언어면 위 `render` 로 충분하다. + +--- + +## 4. API 계약 + +요청(JSON, `POST /api/tts`): +```json +{ + "text": "읽을 내용", + "engine": "melo", + "language": "KR", + "speaker": "KR", + "speed": 1.25, + "pitch": 0.0, + "word_gap": 0.25, + "sentence_gap": 0.75 +} +``` + +FastAPI/Pydantic 필드 정의(범위·기본값 포함): +```python +speed: float = Field(1.25, ge=0.5, le=2.0) +pitch: float = Field(0.0, ge=-12.0, le=12.0) +word_gap: float = Field(0.25, ge=-0.2, le=0.5) # 음수=더 붙임 +sentence_gap: float = Field(0.75, ge=-0.5, le=1.5) # 음수=더 붙임 +``` +응답: `audio/wav` (PCM16) 바이트. + +--- + +## 5. UI 슬라이더 스펙 + +```html + + + + + + + + +``` + +표시 포맷(예): +```js +speedVal.textContent = parseFloat(speed.value).toFixed(2) + "x"; // 1.25x +wordGapVal.textContent = Math.round(parseFloat(wordGap.value)*1000) + " ms"; // -70 ms +sentGapVal.textContent = parseFloat(sentGap.value).toFixed(2) + " s"; // -0.30 s +pitchVal.textContent = (v>0? "+"+v : v) + " 반음"; +``` +전송 시 `word_gap`, `sentence_gap` 은 **초 단위 float** 로 보낸다(ms 아님). + +엔진 지원 플래그로 슬라이더 활성/비활성 처리(선택): +```js +wordGap.disabled = supports.word_gap !== true; +sentGap.disabled = supports.sentence_gap !== true; +``` +엔진 `describe().supports` 예: `{"speed": true, "pitch": true, "word_gap": true, "sentence_gap": true}` + +--- + +## 6. watch_screen_ai 적용 절차 + +1. **의존성**: `melo`(MeloTTS), `numpy`, `soundfile`, (피치 쓰면) `librosa`. GPU면 torch cu128. +2. **모델 로드**: `model = TTS(language="KR", device="cuda:0" if torch.cuda.is_available() else "cpu")` + - `speaker_id = model.hps.data.spk2id["KR"]` +3. **렌더 함수 이식**: 위 §3 코드를 그대로 넣고, TTS 호출부를 `render(model, text, speaker_id, ...)` 로 교체. +4. **파라미터 노출**: + - 기존에 "속도" 하나만 있었다면 `speed`(글자 속도)로 매핑하고, `word_gap`/`sentence_gap` 을 추가. + - API/설정에 §4 필드를, UI가 있으면 §5 슬라이더를 추가. +5. **검증**: §8 스니펫으로 각 파라미터가 독립적으로 duration을 바꾸는지 확인. +6. **배포**: 이미지/서비스 재빌드·재기동 후 실제 합성으로 확인. + +기존에 후처리 배속(atempo/rubberband/리샘플)으로 속도를 주고 있었다면 그 코드는 제거하고 +`speed`(length_scale) 경로로 교체할 것. 배속과 length_scale을 동시에 걸면 이중 왜곡이 된다. + +--- + +## 7. 주의점 / 한계 + +- **글자 속도 vs 전체 배속**: 이 방식은 "음절 발화 속도"다. 완성 음성을 통째로 빠르게(전체 배속) + 하고 싶으면 그건 별도의 atempo 슬라이더로 분리해야 한다(두 개념은 한 슬라이더로 공존 불가). +- **단어 간격의 효과 범위**: 문장 내부에 실제로 존재하는 무음(어절/구 경계 pause)만 조절한다. + 쉼표 등으로 pause가 있으면 효과가 크고, 완전 연속 발화 구간은 조절 여지가 적다. + 음수는 그 pause를 자연 상태보다 더 줄인다. +- **min_pause(기본 80ms)**: 이보다 짧은 무음은 자음 폐쇄음일 수 있어 건드리지 않는다. + 더 촘촘히 줄이고 싶으면 낮추되, 너무 낮추면 파열음이 뭉개질 수 있다. +- **thresh(기본 0.02)**: 무음 판정 임계값(파형 진폭, float32 [-1,1] 기준 ~1~2%). 배경 잡음이 + 있는 음성이면 올리고, 아주 조용하면 내린다. +- **MeloTTS 확률성**: 내부 duration predictor가 확률적이라 같은 문장도 길이가 미세하게 다르다. + 검증 시 여러 번 평균으로 비교할 것. +- **문장 간격 트리밍 한도**: 경계에 존재하는 무음 이상으로는 못 줄인다(겹침/크로스페이드 미구현). + +--- + +## 8. 검증 스니펫 + +```python +import io, wave, statistics +def dur(wav_bytes): + w = wave.open(io.BytesIO(wav_bytes)); return w.getnframes()/w.getframerate() + +# 글자 속도(간격 0): 1.5배가 1.0배보다 짧아야 +# 문장 간격: 0.75 > 0.0 > -0.5 순으로 짧아져야 +# 단어 간격: +0.3 > 0.0 > -0.2 순으로 짧아져야 (확률성 있어 3~4회 평균) +``` + +측정 예(라이브, 참고값): +- 글자 속도 0.7 / 1.0 / 1.5 → 3.82 / 2.74 / 1.96 s +- 문장 간격 0.75 / 0.0 / -0.5 → 4.64 / 3.86 / 3.36 s +- 단어 간격 +0.3 / 0.0 / -0.2 → 6.82 / 6.00 / 5.64 s (같은 문장) + +--- + +## 9. 파라미터 치트시트 + +- 또박또박 천천히: `speed 0.9`, `word_gap 0.15`, `sentence_gap 0.6` +- 빠르고 촘촘히(요약 낭독): `speed 1.5`, `word_gap -0.1`, `sentence_gap -0.2` +- 자연스러운 기본: `speed 1.1~1.25`, `word_gap 0.0`, `sentence_gap 0.3~0.5`