# TTS 속도·간격·피치 제어 적용 매뉴얼 (watch_screen_ai 이식용) 이 문서는 tts_site에 구현된 4개 음성 제어(글자 속도 · 단어 간격 · 문장 간격 · 피치)를 다른 프로젝트(예: `watch_screen_ai`)에 그대로 이식하기 위한 매뉴얼이다. MeloTTS(한국어) 기준의 레퍼런스 구현과, API/UI 계약, 적용 절차, 주의점을 담는다. - 원본 구현: `tts_site` (repo `git.tkrmagid.kr/tkrmagid/tts_site`), 커밋 `5ed2d5f` 기준 - 핵심 파일: `app/engines/melo_engine.py`, `app/server.py`, `frontend/index.html`, `frontend/app.js` --- ## 1. 4개 컨트롤 요약 | 컨트롤 | 파라미터 | 범위 | 기본값 | 동작 원리 | |---|---|---|---|---| | 글자 속도 | `speed` | 0.5 ~ 2.0 | 1.25 | 생성 단계 `length_scale = 1/speed` 로 각 음절 발화 길이 조절 (피치 보존) | | 단어 간격 | `word_gap` (초) | -0.2 ~ 0.5 | 0.25 | 자연 합성 결과의 **어절 사이 무음**을 늘리거나(양수) 줄임(음수) | | 문장 간격 | `sentence_gap` (초) | -0.5 ~ 1.5 | 0.75 | 문장 경계에 무음 삽입(양수) 또는 경계 무음 트리밍(음수) | | 피치 | `pitch` (반음) | -12 ~ +12 | 0 | `librosa.effects.pitch_shift` (선택) | 핵심 설계 원칙: - **글자 속도**는 후처리 배속(atempo)이 아니라 **생성 단계 length_scale**로 처리한다. 각 음절("안·녕·하")의 발음 자체가 빨라지고 느려지며, 배속 특유의 뭉개짐/금속성이 없다. - **단어/문장 간격은 글자 속도와 독립**이다. 간격은 삽입/삭제한 "초" 단위 그대로 유지되고 발화 속도 변경에 영향받지 않는다. - **음수 간격**은 경계/어절 사이 무음을 잘라 자연 상태(0)보다 **더 촘촘하게** 만든다. --- ## 2. 왜 이렇게 하는가 (원리) ### 2.1 글자 속도 = length_scale (후처리 배속 금지) 빠르게 만드는 방법은 크게 3가지이고 왜곡 성격이 다르다. - 단순 리샘플링: 피치까지 올라가 "다람쥐 소리" → **사용 금지** - 후처리 타임스트레치(atempo/rubberband): 피치는 보존되나 자음 뭉개짐·금속성 아티팩트 - **생성 단계 length_scale**: 모델이 처음부터 빠른/느린 발화를 합성 → 음색 보존, 아티팩트 없음 → **채택** MeloTTS는 `tts_to_file(..., speed=s)` 내부에서 `length_scale = 1/s`로 모든 음소 길이를 스케일한다. ### 2.2 간격은 오디오 레벨에서 (무음 삽입/트리밍) - 문장 간격: 문장을 개별 합성 → 사이에 무음(zeros)을 넣거나(양수), 경계 무음을 잘라낸다(음수). - 단어 간격: 문장을 **통째로 자연 합성**(억양 보존)한 뒤, 내부 어절 무음 구간만 늘리거나 줄인다. - 어절을 개별 합성하면 억양이 밋밋해지고 자연 상태보다 더 붙일 수 없으므로 쓰지 않는다. - 자음 폐쇄음 같은 짧은 무음(< 80ms)과 문장 앞뒤 경계 무음은 건드리지 않아 발음이 안전하다. --- ## 3. 레퍼런스 구현 (그대로 복붙 가능) MeloTTS `TTS` 객체(`from melo.api import TTS`)와 `numpy`만 있으면 된다. `model.tts_to_file(text, speaker_id, output_path=None, speed=s, quiet=True)` 는 `float32` 파형(numpy)을 반환하고, `model.hps.data.sampling_rate` 가 sr 이다. ```python import re import numpy as np # ── 문장 분리 ──────────────────────────────────────────────── _SENT_SPLIT_RE = re.compile(r"(?<=[.!?。!?…])\s+|\n+") def split_sentences(text: str) -> list[str]: parts = [p.strip() for p in _SENT_SPLIT_RE.split(text) if p and p.strip()] return parts or ([text.strip()] if text.strip() else []) # ── 경계 무음 트리밍(문장 간격 음수용) ────────────────────────── def trim_end_silence(a, sr, max_sec, thresh=0.02): n = a.size if n == 0 or max_sec <= 0: return a, 0.0 max_n = min(n, int(sr * max_sec)) if max_n <= 0: return a, 0.0 tail = np.abs(a[n - max_n:]) nz = np.where(tail >= thresh)[0] cut = max_n if nz.size == 0 else (max_n - 1 - int(nz[-1])) return (a, 0.0) if cut <= 0 else (a[: n - cut], cut / sr) def trim_start_silence(a, sr, max_sec, thresh=0.02): n = a.size if n == 0 or max_sec <= 0: return a, 0.0 max_n = min(n, int(sr * max_sec)) if max_n <= 0: return a, 0.0 head = np.abs(a[:max_n]) nz = np.where(head >= thresh)[0] cut = max_n if nz.size == 0 else int(nz[0]) return (a, 0.0) if cut <= 0 else (a[cut:], cut / sr) # ── 조각 이어붙이기(문장 간격: 양수=무음삽입, 음수=경계 트리밍) ────── def append_unit(pieces, unit, gap, sr): if not pieces: pieces.append(unit); return if gap >= 0: if gap > 1e-4: pieces.append(np.zeros(int(sr * gap), dtype=np.float32)) pieces.append(unit); return budget = -gap prev, removed = trim_end_silence(pieces[-1], sr, budget) pieces[-1] = prev budget -= removed if budget > 1e-4: unit, _ = trim_start_silence(unit, sr, budget) pieces.append(unit) # ── 단어 간격: 문장 내부 무음 스케일링(양수=늘림, 음수=줄임) ───────── def scale_word_gaps(a, sr, delta, thresh=0.02, min_pause=0.08): n = a.size if abs(delta) < 1e-4 or n == 0: return a silent = np.abs(a) < thresh changes = np.flatnonzero(np.diff(silent.astype(np.int8)) != 0) + 1 bounds = [0, *changes.tolist(), n] min_n = int(sr * min_pause) floor_n = int(sr * 0.015) add_n = int(delta * sr) out, last = [], len(bounds) - 2 for k in range(len(bounds) - 1): s, e = bounds[k], bounds[k + 1] seg = a[s:e] is_internal = 0 < k < last # 앞/뒤 경계 무음은 제외 if silent[s] and is_internal and (e - s) >= min_n: new_n = max(floor_n, (e - s) + add_n) if new_n >= (e - s): seg = np.concatenate([seg, np.zeros(new_n - (e - s), dtype=np.float32)]) else: seg = seg[:new_n] out.append(seg) return np.concatenate(out) if out else a # ── 메인 렌더 함수 ────────────────────────────────────────────── def render(model, text, speaker_id, *, speed=1.25, word_gap=0.25, sentence_gap=0.75, pitch=0.0): """model: melo.api.TTS 인스턴스. float32 파형(numpy)과 sr 을 만든다.""" speed = float(max(0.5, min(2.0, speed))) word_gap = float(max(-0.2, min(0.5, word_gap))) sentence_gap = float(max(-0.5, min(1.5, sentence_gap))) pitch = float(max(-12., min(12., pitch))) sr = model.hps.data.sampling_rate pieces, first = [], True for sent in (split_sentences(text) or [text]): # 글자 속도 = 생성 단계 length_scale a = np.asarray( model.tts_to_file(sent, speaker_id, output_path=None, speed=speed, quiet=True), dtype=np.float32, ) if abs(word_gap) > 1e-4: a = scale_word_gaps(a, sr, word_gap) if first: pieces.append(a); first = False else: append_unit(pieces, a, sentence_gap, sr) audio = np.concatenate(pieces) if pieces else np.zeros(1, dtype=np.float32) if abs(pitch) > 1e-3: import librosa audio = librosa.effects.pitch_shift(audio, sr, n_steps=pitch) return audio, sr ``` WAV(PCM16) 바이트로 만들려면: ```python import io, soundfile as sf buf = io.BytesIO(); sf.write(buf, audio, sr, format="WAV", subtype="PCM_16"); buf.seek(0) wav_bytes = buf.read() ``` > AUTO(한/영/일/중 혼합) 언어 분기가 필요하면 원본 `melo_engine.py`의 `split_by_language` > + `_synth_natural` 를 참고. 단일 언어면 위 `render` 로 충분하다. --- ## 4. API 계약 요청(JSON, `POST /api/tts`): ```json { "text": "읽을 내용", "engine": "melo", "language": "KR", "speaker": "KR", "speed": 1.25, "pitch": 0.0, "word_gap": 0.25, "sentence_gap": 0.75 } ``` FastAPI/Pydantic 필드 정의(범위·기본값 포함): ```python speed: float = Field(1.25, ge=0.5, le=2.0) pitch: float = Field(0.0, ge=-12.0, le=12.0) word_gap: float = Field(0.25, ge=-0.2, le=0.5) # 음수=더 붙임 sentence_gap: float = Field(0.75, ge=-0.5, le=1.5) # 음수=더 붙임 ``` 응답: `audio/wav` (PCM16) 바이트. --- ## 5. UI 슬라이더 스펙 ```html ``` 표시 포맷(예): ```js speedVal.textContent = parseFloat(speed.value).toFixed(2) + "x"; // 1.25x wordGapVal.textContent = Math.round(parseFloat(wordGap.value)*1000) + " ms"; // -70 ms sentGapVal.textContent = parseFloat(sentGap.value).toFixed(2) + " s"; // -0.30 s pitchVal.textContent = (v>0? "+"+v : v) + " 반음"; ``` 전송 시 `word_gap`, `sentence_gap` 은 **초 단위 float** 로 보낸다(ms 아님). 엔진 지원 플래그로 슬라이더 활성/비활성 처리(선택): ```js wordGap.disabled = supports.word_gap !== true; sentGap.disabled = supports.sentence_gap !== true; ``` 엔진 `describe().supports` 예: `{"speed": true, "pitch": true, "word_gap": true, "sentence_gap": true}` --- ## 6. watch_screen_ai 적용 절차 1. **의존성**: `melo`(MeloTTS), `numpy`, `soundfile`, (피치 쓰면) `librosa`. GPU면 torch cu128. 2. **모델 로드**: `model = TTS(language="KR", device="cuda:0" if torch.cuda.is_available() else "cpu")` - `speaker_id = model.hps.data.spk2id["KR"]` 3. **렌더 함수 이식**: 위 §3 코드를 그대로 넣고, TTS 호출부를 `render(model, text, speaker_id, ...)` 로 교체. 4. **파라미터 노출**: - 기존에 "속도" 하나만 있었다면 `speed`(글자 속도)로 매핑하고, `word_gap`/`sentence_gap` 을 추가. - API/설정에 §4 필드를, UI가 있으면 §5 슬라이더를 추가. 5. **검증**: §8 스니펫으로 각 파라미터가 독립적으로 duration을 바꾸는지 확인. 6. **배포**: 이미지/서비스 재빌드·재기동 후 실제 합성으로 확인. 기존에 후처리 배속(atempo/rubberband/리샘플)으로 속도를 주고 있었다면 그 코드는 제거하고 `speed`(length_scale) 경로로 교체할 것. 배속과 length_scale을 동시에 걸면 이중 왜곡이 된다. --- ## 7. 주의점 / 한계 - **글자 속도 vs 전체 배속**: 이 방식은 "음절 발화 속도"다. 완성 음성을 통째로 빠르게(전체 배속) 하고 싶으면 그건 별도의 atempo 슬라이더로 분리해야 한다(두 개념은 한 슬라이더로 공존 불가). - **단어 간격의 효과 범위**: 문장 내부에 실제로 존재하는 무음(어절/구 경계 pause)만 조절한다. 쉼표 등으로 pause가 있으면 효과가 크고, 완전 연속 발화 구간은 조절 여지가 적다. 음수는 그 pause를 자연 상태보다 더 줄인다. - **min_pause(기본 80ms)**: 이보다 짧은 무음은 자음 폐쇄음일 수 있어 건드리지 않는다. 더 촘촘히 줄이고 싶으면 낮추되, 너무 낮추면 파열음이 뭉개질 수 있다. - **thresh(기본 0.02)**: 무음 판정 임계값(파형 진폭, float32 [-1,1] 기준 ~1~2%). 배경 잡음이 있는 음성이면 올리고, 아주 조용하면 내린다. - **MeloTTS 확률성**: 내부 duration predictor가 확률적이라 같은 문장도 길이가 미세하게 다르다. 검증 시 여러 번 평균으로 비교할 것. - **문장 간격 트리밍 한도**: 경계에 존재하는 무음 이상으로는 못 줄인다(겹침/크로스페이드 미구현). --- ## 8. 검증 스니펫 ```python import io, wave, statistics def dur(wav_bytes): w = wave.open(io.BytesIO(wav_bytes)); return w.getnframes()/w.getframerate() # 글자 속도(간격 0): 1.5배가 1.0배보다 짧아야 # 문장 간격: 0.75 > 0.0 > -0.5 순으로 짧아져야 # 단어 간격: +0.3 > 0.0 > -0.2 순으로 짧아져야 (확률성 있어 3~4회 평균) ``` 측정 예(라이브, 참고값): - 글자 속도 0.7 / 1.0 / 1.5 → 3.82 / 2.74 / 1.96 s - 문장 간격 0.75 / 0.0 / -0.5 → 4.64 / 3.86 / 3.36 s - 단어 간격 +0.3 / 0.0 / -0.2 → 6.82 / 6.00 / 5.64 s (같은 문장) --- ## 9. 파라미터 치트시트 - 또박또박 천천히: `speed 0.9`, `word_gap 0.15`, `sentence_gap 0.6` - 빠르고 촘촘히(요약 낭독): `speed 1.5`, `word_gap -0.1`, `sentence_gap -0.2` - 자연스러운 기본: `speed 1.1~1.25`, `word_gap 0.0`, `sentence_gap 0.3~0.5`