docs: TTS 속도·간격·피치 제어 이식 매뉴얼(watch_screen_ai 적용용) 추가
This commit is contained in:
290
docs/TTS_속도_간격_적용_매뉴얼.md
Normal file
290
docs/TTS_속도_간격_적용_매뉴얼.md
Normal file
@@ -0,0 +1,290 @@
|
||||
# TTS 속도·간격·피치 제어 적용 매뉴얼 (watch_screen_ai 이식용)
|
||||
|
||||
이 문서는 tts_site에 구현된 4개 음성 제어(글자 속도 · 단어 간격 · 문장 간격 · 피치)를
|
||||
다른 프로젝트(예: `watch_screen_ai`)에 그대로 이식하기 위한 매뉴얼이다.
|
||||
MeloTTS(한국어) 기준의 레퍼런스 구현과, API/UI 계약, 적용 절차, 주의점을 담는다.
|
||||
|
||||
- 원본 구현: `tts_site` (repo `git.tkrmagid.kr/tkrmagid/tts_site`), 커밋 `5ed2d5f` 기준
|
||||
- 핵심 파일: `app/engines/melo_engine.py`, `app/server.py`, `frontend/index.html`, `frontend/app.js`
|
||||
|
||||
---
|
||||
|
||||
## 1. 4개 컨트롤 요약
|
||||
|
||||
| 컨트롤 | 파라미터 | 범위 | 기본값 | 동작 원리 |
|
||||
|---|---|---|---|---|
|
||||
| 글자 속도 | `speed` | 0.5 ~ 2.0 | 1.25 | 생성 단계 `length_scale = 1/speed` 로 각 음절 발화 길이 조절 (피치 보존) |
|
||||
| 단어 간격 | `word_gap` (초) | -0.2 ~ 0.5 | 0.25 | 자연 합성 결과의 **어절 사이 무음**을 늘리거나(양수) 줄임(음수) |
|
||||
| 문장 간격 | `sentence_gap` (초) | -0.5 ~ 1.5 | 0.75 | 문장 경계에 무음 삽입(양수) 또는 경계 무음 트리밍(음수) |
|
||||
| 피치 | `pitch` (반음) | -12 ~ +12 | 0 | `librosa.effects.pitch_shift` (선택) |
|
||||
|
||||
핵심 설계 원칙:
|
||||
- **글자 속도**는 후처리 배속(atempo)이 아니라 **생성 단계 length_scale**로 처리한다.
|
||||
각 음절("안·녕·하")의 발음 자체가 빨라지고 느려지며, 배속 특유의 뭉개짐/금속성이 없다.
|
||||
- **단어/문장 간격은 글자 속도와 독립**이다. 간격은 삽입/삭제한 "초" 단위 그대로 유지되고
|
||||
발화 속도 변경에 영향받지 않는다.
|
||||
- **음수 간격**은 경계/어절 사이 무음을 잘라 자연 상태(0)보다 **더 촘촘하게** 만든다.
|
||||
|
||||
---
|
||||
|
||||
## 2. 왜 이렇게 하는가 (원리)
|
||||
|
||||
### 2.1 글자 속도 = length_scale (후처리 배속 금지)
|
||||
빠르게 만드는 방법은 크게 3가지이고 왜곡 성격이 다르다.
|
||||
- 단순 리샘플링: 피치까지 올라가 "다람쥐 소리" → **사용 금지**
|
||||
- 후처리 타임스트레치(atempo/rubberband): 피치는 보존되나 자음 뭉개짐·금속성 아티팩트
|
||||
- **생성 단계 length_scale**: 모델이 처음부터 빠른/느린 발화를 합성 → 음색 보존, 아티팩트 없음 → **채택**
|
||||
|
||||
MeloTTS는 `tts_to_file(..., speed=s)` 내부에서 `length_scale = 1/s`로 모든 음소 길이를 스케일한다.
|
||||
|
||||
### 2.2 간격은 오디오 레벨에서 (무음 삽입/트리밍)
|
||||
- 문장 간격: 문장을 개별 합성 → 사이에 무음(zeros)을 넣거나(양수), 경계 무음을 잘라낸다(음수).
|
||||
- 단어 간격: 문장을 **통째로 자연 합성**(억양 보존)한 뒤, 내부 어절 무음 구간만 늘리거나 줄인다.
|
||||
- 어절을 개별 합성하면 억양이 밋밋해지고 자연 상태보다 더 붙일 수 없으므로 쓰지 않는다.
|
||||
- 자음 폐쇄음 같은 짧은 무음(< 80ms)과 문장 앞뒤 경계 무음은 건드리지 않아 발음이 안전하다.
|
||||
|
||||
---
|
||||
|
||||
## 3. 레퍼런스 구현 (그대로 복붙 가능)
|
||||
|
||||
MeloTTS `TTS` 객체(`from melo.api import TTS`)와 `numpy`만 있으면 된다.
|
||||
`model.tts_to_file(text, speaker_id, output_path=None, speed=s, quiet=True)` 는
|
||||
`float32` 파형(numpy)을 반환하고, `model.hps.data.sampling_rate` 가 sr 이다.
|
||||
|
||||
```python
|
||||
import re
|
||||
import numpy as np
|
||||
|
||||
# ── 문장 분리 ────────────────────────────────────────────────
|
||||
_SENT_SPLIT_RE = re.compile(r"(?<=[.!?。!?…])\s+|\n+")
|
||||
|
||||
def split_sentences(text: str) -> list[str]:
|
||||
parts = [p.strip() for p in _SENT_SPLIT_RE.split(text) if p and p.strip()]
|
||||
return parts or ([text.strip()] if text.strip() else [])
|
||||
|
||||
# ── 경계 무음 트리밍(문장 간격 음수용) ──────────────────────────
|
||||
def trim_end_silence(a, sr, max_sec, thresh=0.02):
|
||||
n = a.size
|
||||
if n == 0 or max_sec <= 0:
|
||||
return a, 0.0
|
||||
max_n = min(n, int(sr * max_sec))
|
||||
if max_n <= 0:
|
||||
return a, 0.0
|
||||
tail = np.abs(a[n - max_n:])
|
||||
nz = np.where(tail >= thresh)[0]
|
||||
cut = max_n if nz.size == 0 else (max_n - 1 - int(nz[-1]))
|
||||
return (a, 0.0) if cut <= 0 else (a[: n - cut], cut / sr)
|
||||
|
||||
def trim_start_silence(a, sr, max_sec, thresh=0.02):
|
||||
n = a.size
|
||||
if n == 0 or max_sec <= 0:
|
||||
return a, 0.0
|
||||
max_n = min(n, int(sr * max_sec))
|
||||
if max_n <= 0:
|
||||
return a, 0.0
|
||||
head = np.abs(a[:max_n])
|
||||
nz = np.where(head >= thresh)[0]
|
||||
cut = max_n if nz.size == 0 else int(nz[0])
|
||||
return (a, 0.0) if cut <= 0 else (a[cut:], cut / sr)
|
||||
|
||||
# ── 조각 이어붙이기(문장 간격: 양수=무음삽입, 음수=경계 트리밍) ──────
|
||||
def append_unit(pieces, unit, gap, sr):
|
||||
if not pieces:
|
||||
pieces.append(unit); return
|
||||
if gap >= 0:
|
||||
if gap > 1e-4:
|
||||
pieces.append(np.zeros(int(sr * gap), dtype=np.float32))
|
||||
pieces.append(unit); return
|
||||
budget = -gap
|
||||
prev, removed = trim_end_silence(pieces[-1], sr, budget)
|
||||
pieces[-1] = prev
|
||||
budget -= removed
|
||||
if budget > 1e-4:
|
||||
unit, _ = trim_start_silence(unit, sr, budget)
|
||||
pieces.append(unit)
|
||||
|
||||
# ── 단어 간격: 문장 내부 무음 스케일링(양수=늘림, 음수=줄임) ─────────
|
||||
def scale_word_gaps(a, sr, delta, thresh=0.02, min_pause=0.08):
|
||||
n = a.size
|
||||
if abs(delta) < 1e-4 or n == 0:
|
||||
return a
|
||||
silent = np.abs(a) < thresh
|
||||
changes = np.flatnonzero(np.diff(silent.astype(np.int8)) != 0) + 1
|
||||
bounds = [0, *changes.tolist(), n]
|
||||
min_n = int(sr * min_pause)
|
||||
floor_n = int(sr * 0.015)
|
||||
add_n = int(delta * sr)
|
||||
out, last = [], len(bounds) - 2
|
||||
for k in range(len(bounds) - 1):
|
||||
s, e = bounds[k], bounds[k + 1]
|
||||
seg = a[s:e]
|
||||
is_internal = 0 < k < last # 앞/뒤 경계 무음은 제외
|
||||
if silent[s] and is_internal and (e - s) >= min_n:
|
||||
new_n = max(floor_n, (e - s) + add_n)
|
||||
if new_n >= (e - s):
|
||||
seg = np.concatenate([seg, np.zeros(new_n - (e - s), dtype=np.float32)])
|
||||
else:
|
||||
seg = seg[:new_n]
|
||||
out.append(seg)
|
||||
return np.concatenate(out) if out else a
|
||||
|
||||
# ── 메인 렌더 함수 ──────────────────────────────────────────────
|
||||
def render(model, text, speaker_id, *,
|
||||
speed=1.25, word_gap=0.25, sentence_gap=0.75, pitch=0.0):
|
||||
"""model: melo.api.TTS 인스턴스. float32 파형(numpy)과 sr 을 만든다."""
|
||||
speed = float(max(0.5, min(2.0, speed)))
|
||||
word_gap = float(max(-0.2, min(0.5, word_gap)))
|
||||
sentence_gap = float(max(-0.5, min(1.5, sentence_gap)))
|
||||
pitch = float(max(-12., min(12., pitch)))
|
||||
sr = model.hps.data.sampling_rate
|
||||
|
||||
pieces, first = [], True
|
||||
for sent in (split_sentences(text) or [text]):
|
||||
# 글자 속도 = 생성 단계 length_scale
|
||||
a = np.asarray(
|
||||
model.tts_to_file(sent, speaker_id, output_path=None, speed=speed, quiet=True),
|
||||
dtype=np.float32,
|
||||
)
|
||||
if abs(word_gap) > 1e-4:
|
||||
a = scale_word_gaps(a, sr, word_gap)
|
||||
if first:
|
||||
pieces.append(a); first = False
|
||||
else:
|
||||
append_unit(pieces, a, sentence_gap, sr)
|
||||
|
||||
audio = np.concatenate(pieces) if pieces else np.zeros(1, dtype=np.float32)
|
||||
|
||||
if abs(pitch) > 1e-3:
|
||||
import librosa
|
||||
audio = librosa.effects.pitch_shift(audio, sr, n_steps=pitch)
|
||||
return audio, sr
|
||||
```
|
||||
|
||||
WAV(PCM16) 바이트로 만들려면:
|
||||
```python
|
||||
import io, soundfile as sf
|
||||
buf = io.BytesIO(); sf.write(buf, audio, sr, format="WAV", subtype="PCM_16"); buf.seek(0)
|
||||
wav_bytes = buf.read()
|
||||
```
|
||||
|
||||
> AUTO(한/영/일/중 혼합) 언어 분기가 필요하면 원본 `melo_engine.py`의 `split_by_language`
|
||||
> + `_synth_natural` 를 참고. 단일 언어면 위 `render` 로 충분하다.
|
||||
|
||||
---
|
||||
|
||||
## 4. API 계약
|
||||
|
||||
요청(JSON, `POST /api/tts`):
|
||||
```json
|
||||
{
|
||||
"text": "읽을 내용",
|
||||
"engine": "melo",
|
||||
"language": "KR",
|
||||
"speaker": "KR",
|
||||
"speed": 1.25,
|
||||
"pitch": 0.0,
|
||||
"word_gap": 0.25,
|
||||
"sentence_gap": 0.75
|
||||
}
|
||||
```
|
||||
|
||||
FastAPI/Pydantic 필드 정의(범위·기본값 포함):
|
||||
```python
|
||||
speed: float = Field(1.25, ge=0.5, le=2.0)
|
||||
pitch: float = Field(0.0, ge=-12.0, le=12.0)
|
||||
word_gap: float = Field(0.25, ge=-0.2, le=0.5) # 음수=더 붙임
|
||||
sentence_gap: float = Field(0.75, ge=-0.5, le=1.5) # 음수=더 붙임
|
||||
```
|
||||
응답: `audio/wav` (PCM16) 바이트.
|
||||
|
||||
---
|
||||
|
||||
## 5. UI 슬라이더 스펙
|
||||
|
||||
```html
|
||||
<!-- 글자 속도 -->
|
||||
<input type="range" id="speed" min="0.5" max="2.0" step="0.05" value="1.25" />
|
||||
<!-- 단어 간격(초) -->
|
||||
<input type="range" id="wordGap" min="-0.2" max="0.5" step="0.01" value="0.25" />
|
||||
<!-- 문장 간격(초) -->
|
||||
<input type="range" id="sentGap" min="-0.5" max="1.5" step="0.05" value="0.75" />
|
||||
<!-- 피치(반음) -->
|
||||
<input type="range" id="pitch" min="-12" max="12" step="1" value="0" />
|
||||
```
|
||||
|
||||
표시 포맷(예):
|
||||
```js
|
||||
speedVal.textContent = parseFloat(speed.value).toFixed(2) + "x"; // 1.25x
|
||||
wordGapVal.textContent = Math.round(parseFloat(wordGap.value)*1000) + " ms"; // -70 ms
|
||||
sentGapVal.textContent = parseFloat(sentGap.value).toFixed(2) + " s"; // -0.30 s
|
||||
pitchVal.textContent = (v>0? "+"+v : v) + " 반음";
|
||||
```
|
||||
전송 시 `word_gap`, `sentence_gap` 은 **초 단위 float** 로 보낸다(ms 아님).
|
||||
|
||||
엔진 지원 플래그로 슬라이더 활성/비활성 처리(선택):
|
||||
```js
|
||||
wordGap.disabled = supports.word_gap !== true;
|
||||
sentGap.disabled = supports.sentence_gap !== true;
|
||||
```
|
||||
엔진 `describe().supports` 예: `{"speed": true, "pitch": true, "word_gap": true, "sentence_gap": true}`
|
||||
|
||||
---
|
||||
|
||||
## 6. watch_screen_ai 적용 절차
|
||||
|
||||
1. **의존성**: `melo`(MeloTTS), `numpy`, `soundfile`, (피치 쓰면) `librosa`. GPU면 torch cu128.
|
||||
2. **모델 로드**: `model = TTS(language="KR", device="cuda:0" if torch.cuda.is_available() else "cpu")`
|
||||
- `speaker_id = model.hps.data.spk2id["KR"]`
|
||||
3. **렌더 함수 이식**: 위 §3 코드를 그대로 넣고, TTS 호출부를 `render(model, text, speaker_id, ...)` 로 교체.
|
||||
4. **파라미터 노출**:
|
||||
- 기존에 "속도" 하나만 있었다면 `speed`(글자 속도)로 매핑하고, `word_gap`/`sentence_gap` 을 추가.
|
||||
- API/설정에 §4 필드를, UI가 있으면 §5 슬라이더를 추가.
|
||||
5. **검증**: §8 스니펫으로 각 파라미터가 독립적으로 duration을 바꾸는지 확인.
|
||||
6. **배포**: 이미지/서비스 재빌드·재기동 후 실제 합성으로 확인.
|
||||
|
||||
기존에 후처리 배속(atempo/rubberband/리샘플)으로 속도를 주고 있었다면 그 코드는 제거하고
|
||||
`speed`(length_scale) 경로로 교체할 것. 배속과 length_scale을 동시에 걸면 이중 왜곡이 된다.
|
||||
|
||||
---
|
||||
|
||||
## 7. 주의점 / 한계
|
||||
|
||||
- **글자 속도 vs 전체 배속**: 이 방식은 "음절 발화 속도"다. 완성 음성을 통째로 빠르게(전체 배속)
|
||||
하고 싶으면 그건 별도의 atempo 슬라이더로 분리해야 한다(두 개념은 한 슬라이더로 공존 불가).
|
||||
- **단어 간격의 효과 범위**: 문장 내부에 실제로 존재하는 무음(어절/구 경계 pause)만 조절한다.
|
||||
쉼표 등으로 pause가 있으면 효과가 크고, 완전 연속 발화 구간은 조절 여지가 적다.
|
||||
음수는 그 pause를 자연 상태보다 더 줄인다.
|
||||
- **min_pause(기본 80ms)**: 이보다 짧은 무음은 자음 폐쇄음일 수 있어 건드리지 않는다.
|
||||
더 촘촘히 줄이고 싶으면 낮추되, 너무 낮추면 파열음이 뭉개질 수 있다.
|
||||
- **thresh(기본 0.02)**: 무음 판정 임계값(파형 진폭, float32 [-1,1] 기준 ~1~2%). 배경 잡음이
|
||||
있는 음성이면 올리고, 아주 조용하면 내린다.
|
||||
- **MeloTTS 확률성**: 내부 duration predictor가 확률적이라 같은 문장도 길이가 미세하게 다르다.
|
||||
검증 시 여러 번 평균으로 비교할 것.
|
||||
- **문장 간격 트리밍 한도**: 경계에 존재하는 무음 이상으로는 못 줄인다(겹침/크로스페이드 미구현).
|
||||
|
||||
---
|
||||
|
||||
## 8. 검증 스니펫
|
||||
|
||||
```python
|
||||
import io, wave, statistics
|
||||
def dur(wav_bytes):
|
||||
w = wave.open(io.BytesIO(wav_bytes)); return w.getnframes()/w.getframerate()
|
||||
|
||||
# 글자 속도(간격 0): 1.5배가 1.0배보다 짧아야
|
||||
# 문장 간격: 0.75 > 0.0 > -0.5 순으로 짧아져야
|
||||
# 단어 간격: +0.3 > 0.0 > -0.2 순으로 짧아져야 (확률성 있어 3~4회 평균)
|
||||
```
|
||||
|
||||
측정 예(라이브, 참고값):
|
||||
- 글자 속도 0.7 / 1.0 / 1.5 → 3.82 / 2.74 / 1.96 s
|
||||
- 문장 간격 0.75 / 0.0 / -0.5 → 4.64 / 3.86 / 3.36 s
|
||||
- 단어 간격 +0.3 / 0.0 / -0.2 → 6.82 / 6.00 / 5.64 s (같은 문장)
|
||||
|
||||
---
|
||||
|
||||
## 9. 파라미터 치트시트
|
||||
|
||||
- 또박또박 천천히: `speed 0.9`, `word_gap 0.15`, `sentence_gap 0.6`
|
||||
- 빠르고 촘촘히(요약 낭독): `speed 1.5`, `word_gap -0.1`, `sentence_gap -0.2`
|
||||
- 자연스러운 기본: `speed 1.1~1.25`, `word_gap 0.0`, `sentence_gap 0.3~0.5`
|
||||
Reference in New Issue
Block a user