리뷰 지적(포터블에서 transformers 제외 -> NLLB 로드 실패)을 고치려고 실제
NLLB 를 받아 돌려봤고, 그 과정에서 용어집이 사실상 동작하지 않고 있었다는
것을 발견했다. 단위 테스트는 "모델이 자리표시자를 통과시킨다"는 틀린 전제
위에 서 있었다.
1) 포터블 패키징 (리뷰 지적)
transformers 를 제외 목록에서 빼고 hiddenimports 에 넣었다. NLLB
토크나이저가 AutoTokenizer 를 쓰기 때문이다. transformers 는 torch 가
없으면 토크나이저 전용 모드로 뜨며 그게 우리 용도와 정확히 맞는다.
torch 없는 환경에서 ctranslate2/transformers/faster-whisper import 와
앱 기동을 검증하는 test_portable.py 를 추가했다.
2) 자리표시자 형식 (실측으로 발견)
`⟦0⟧` 는 NLLB 가 괄호를 날려 생존률 0/3 이었다. 용어가 자막에서 그냥
사라지고 있었다 ("Third party incoming" -> "0 들어오는"). 후보 8종을
실제 모델로 비교해 `#0#` 로 교체 (3/3, 다중 4/5).
3) 소실 대비
모델이 문장 일부를 누락하면 자리표시자도 사라진다. 그대로 복원하면
용어가 증발하므로, 하나라도 없으면 보호 없이 재번역한다.
4) 서술어는 문장 전체일 때만 (whole_only)
절/서술어를 문장 중간에서 치환하면 문법이 무너진다.
before: "탄 필요해와 구급상자"
after : "탄약과 구급상자가 필요합니다"
해당 56개 항목을 whole_only 로 지정해 단독 발화일 때만 적용한다.
("Cover me!" -> "엄호해줘" 는 그대로 유지)
5) 조사 교정
역어 받침이 달라 "자기장를" 이 남던 것을 fix_particles() 로 고친다.
을/를, 이/가, 은/는, 과/와, (으)로 — 한글 코드에서 받침을 읽어 판정.
검증: pytest 189개 통과, ruff clean
실제 NLLB-600M(torch 없이 CPU)로 번역 품질 직접 확인
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
279 lines
9.4 KiB
Python
279 lines
9.4 KiB
Python
"""말투(존댓말/반말) 모드.
|
|
|
|
사용자 규칙:
|
|
- 기본은 존댓말.
|
|
- 높임이 없거나(영어·중국어) 알 수 없으면 고른 모드를 따른다.
|
|
- **원문이 실제로 존댓말이면 반말 모드여도 존댓말을 지킨다.**
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
import pytest
|
|
|
|
from livesub.models.speech_level import (
|
|
Politeness,
|
|
SpeechLevel,
|
|
apply_speech_level,
|
|
detect_politeness,
|
|
prompt_instruction,
|
|
to_casual,
|
|
)
|
|
|
|
# --- 원문 높임 감지 ---------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"text",
|
|
["안녕하세요", "같이 가시죠", "도와주세요", "감사합니다", "지금 가겠습니다", "괜찮으세요?"],
|
|
)
|
|
def test_korean_polite_is_detected(text):
|
|
assert detect_politeness(text, "ko") is Politeness.POLITE
|
|
|
|
|
|
@pytest.mark.parametrize("text", ["같이 가자", "빨리 와", "적이 온다", "내가 할게"])
|
|
def test_korean_casual_is_detected(text):
|
|
assert detect_politeness(text, "ko") is Politeness.CASUAL
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
"text", ["こんにちは、お願いします", "行きますよ", "ありがとうございます", "待ってください"]
|
|
)
|
|
def test_japanese_polite_is_detected(text):
|
|
assert detect_politeness(text, "ja") is Politeness.POLITE
|
|
|
|
|
|
@pytest.mark.parametrize("text", ["早く来いよ", "行くぞ", "危ないだろう"])
|
|
def test_japanese_casual_is_detected(text):
|
|
assert detect_politeness(text, "ja") is Politeness.CASUAL
|
|
|
|
|
|
@pytest.mark.parametrize("lang", ["en", "zh"])
|
|
def test_languages_without_honorifics_are_always_unknown(lang):
|
|
"""영어·중국어는 문법적 높임이 없으므로 항상 모드를 따라야 한다.
|
|
|
|
"please" 를 존댓말 근거로 삼으면 오탐이 너무 많아진다.
|
|
"""
|
|
for text in ["Could you please help me, sir?", "Get down now!", "请帮我一下", "快走"]:
|
|
assert detect_politeness(text, lang) is Politeness.UNKNOWN
|
|
|
|
|
|
def test_empty_text_is_unknown():
|
|
assert detect_politeness("", "ko") is Politeness.UNKNOWN
|
|
assert detect_politeness(" ", "ja") is Politeness.UNKNOWN
|
|
|
|
|
|
# --- 반말 변환 --------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
("polite", "casual"),
|
|
[
|
|
("적이 왼쪽에서 옵니다", "적이 왼쪽에서 와"),
|
|
("지금 바로 빠지세요", "지금 바로 빠져"),
|
|
("바론을 치겠습니다", "바론을 칠게"),
|
|
("적을 눕혔습니다", "적을 눕혔어"),
|
|
("탄약이 없습니다", "탄약이 없어"),
|
|
("여기 적이 있습니다", "여기 적이 있어"),
|
|
("제가 하겠습니다", "제가 할게"),
|
|
("좋습니다", "좋아"),
|
|
("빨리 먹습니다", "빨리 먹어"),
|
|
("저건 함정이에요", "저건 함정이야"),
|
|
("제 차례예요", "제 차례야"),
|
|
("조심하세요", "조심해"),
|
|
("잘 했네요", "잘 했네"),
|
|
("같이 가죠", "같이 가지"),
|
|
],
|
|
)
|
|
def test_polite_to_casual(polite, casual):
|
|
assert to_casual(polite) == casual
|
|
|
|
|
|
def test_trailing_yo_is_dropped():
|
|
assert to_casual("같이 가요") == "같이 가"
|
|
assert to_casual("빨리 와요!") == "빨리 와!"
|
|
|
|
|
|
def test_mid_sentence_yo_is_preserved():
|
|
"""'중요', '필요' 의 '요'를 떼면 말이 망가진다."""
|
|
assert "중요" in to_casual("이게 제일 중요합니다")
|
|
assert "필요" in to_casual("탄약이 필요합니다")
|
|
|
|
|
|
def test_conversion_is_idempotent_on_already_casual_text():
|
|
casual = "적이 왼쪽에서 와"
|
|
assert to_casual(casual) == casual
|
|
|
|
|
|
def test_empty_stays_empty():
|
|
assert to_casual("") == ""
|
|
|
|
|
|
# --- 모드 적용 규칙 ---------------------------------------------------------
|
|
|
|
|
|
def test_polite_mode_never_alters_output():
|
|
"""존댓말 모드는 손대지 않는다 — 모델 출력이 이미 격식체다."""
|
|
text = "적이 왼쪽에서 옵니다"
|
|
for politeness in Politeness:
|
|
assert apply_speech_level(text, "ko", SpeechLevel.POLITE, politeness) == text
|
|
|
|
|
|
def test_casual_mode_lowers_when_source_is_unknown():
|
|
"""영어 원문 → 높임 정보 없음 → 고른 모드(반말)를 따른다."""
|
|
out = apply_speech_level(
|
|
"적이 왼쪽에서 옵니다", "ko", SpeechLevel.CASUAL, Politeness.UNKNOWN
|
|
)
|
|
assert out == "적이 왼쪽에서 와"
|
|
|
|
|
|
def test_casual_mode_keeps_polite_when_source_was_polite():
|
|
"""핵심 규칙 — 실제로 존댓말로 말했으면 반말 모드여도 존댓말."""
|
|
text = "도와주시겠습니까"
|
|
assert apply_speech_level(text, "ko", SpeechLevel.CASUAL, Politeness.POLITE) == text
|
|
|
|
|
|
def test_casual_mode_lowers_when_source_was_casual():
|
|
out = apply_speech_level("빨리 옵니다", "ko", SpeechLevel.CASUAL, Politeness.CASUAL)
|
|
assert out == "빨리 와"
|
|
|
|
|
|
@pytest.mark.parametrize("target", ["en", "ja", "zh"])
|
|
def test_non_korean_targets_are_untouched(target):
|
|
"""한국어 외에는 적용할 규칙이 없으므로 원본을 그대로 둔다."""
|
|
text = "Enemy incoming"
|
|
assert apply_speech_level(text, target, SpeechLevel.CASUAL, Politeness.UNKNOWN) == text
|
|
|
|
|
|
# --- LLM 프롬프트 지시문 ----------------------------------------------------
|
|
|
|
|
|
def test_prompt_asks_for_polite_when_source_was_polite():
|
|
hint = prompt_instruction(SpeechLevel.CASUAL, Politeness.POLITE)
|
|
assert "존댓말" in hint
|
|
|
|
|
|
def test_prompt_asks_for_casual_in_casual_mode():
|
|
assert "반말" in prompt_instruction(SpeechLevel.CASUAL, Politeness.UNKNOWN)
|
|
|
|
|
|
def test_prompt_asks_for_polite_in_polite_mode():
|
|
assert "존댓말" in prompt_instruction(SpeechLevel.POLITE, Politeness.UNKNOWN)
|
|
|
|
|
|
# --- 엔진 파이프라인 끝까지 도달하는지 ---------------------------------------
|
|
|
|
|
|
class _PoliteTranslator:
|
|
"""한국어 격식체를 내놓는 번역기 (NLLB/Seed-X 가 실제로 그렇다)."""
|
|
|
|
loaded = True
|
|
|
|
def load(self):
|
|
pass
|
|
|
|
def unload(self):
|
|
pass
|
|
|
|
def translate(self, text, source_lang, target_lang, glossary=None, **kwargs):
|
|
self.kwargs = kwargs
|
|
return "적이 왼쪽에서 옵니다"
|
|
|
|
|
|
def _run(monkeypatch, tmp_path, level: str, source_text: str, source_lang: str) -> str:
|
|
from livesub.config import AppConfig
|
|
from livesub.core.engine import TranslationEngine
|
|
from livesub.models.asr import Transcript
|
|
|
|
cfg = AppConfig()
|
|
cfg.models.preload_on_start = False
|
|
cfg.models.speech_level = level
|
|
cfg.glossary.path = str(tmp_path / "g.json")
|
|
cfg.glossary.enabled_packs = []
|
|
|
|
engine = TranslationEngine(cfg)
|
|
translator = _PoliteTranslator()
|
|
|
|
class Recognizer:
|
|
def transcribe(self, audio, language=None, fast=False, prompt=""):
|
|
return Transcript(text=source_text, language=source_lang)
|
|
|
|
lines = []
|
|
engine._on_line = lines.append
|
|
|
|
class Segment:
|
|
is_final = True
|
|
audio = None
|
|
ended_at = 0.0
|
|
|
|
engine._process(Segment(), Recognizer(), translator, cfg.models)
|
|
return lines[0].translated_text
|
|
|
|
|
|
def test_english_source_casual_mode_lowers(monkeypatch, tmp_path):
|
|
"""영어는 높임이 없으니 고른 모드(반말)를 따른다."""
|
|
assert _run(monkeypatch, tmp_path, "casual", "Enemy from the left", "en") == (
|
|
"적이 왼쪽에서 와"
|
|
)
|
|
|
|
|
|
def test_english_source_polite_mode_stays_polite(monkeypatch, tmp_path):
|
|
assert _run(monkeypatch, tmp_path, "polite", "Enemy from the left", "en") == (
|
|
"적이 왼쪽에서 옵니다"
|
|
)
|
|
|
|
|
|
def test_japanese_polite_source_stays_polite_even_in_casual_mode(monkeypatch, tmp_path):
|
|
"""핵심 규칙 — 실제로 존댓말로 말했으면 반말 모드여도 존댓말."""
|
|
assert _run(monkeypatch, tmp_path, "casual", "左から来ます", "ja") == (
|
|
"적이 왼쪽에서 옵니다"
|
|
)
|
|
|
|
|
|
def test_japanese_casual_source_is_lowered_in_casual_mode(monkeypatch, tmp_path):
|
|
assert _run(monkeypatch, tmp_path, "casual", "左から来るぞ", "ja") == (
|
|
"적이 왼쪽에서 와"
|
|
)
|
|
|
|
|
|
def test_unknown_speech_level_value_falls_back_to_polite(monkeypatch, tmp_path):
|
|
"""설정 파일이 손상돼도 안전한 존댓말로 떨어져야 한다."""
|
|
assert _run(monkeypatch, tmp_path, "무엇인가이상한값", "Enemy", "en") == (
|
|
"적이 왼쪽에서 옵니다"
|
|
)
|
|
|
|
|
|
# --- 조사 교정 --------------------------------------------------------------
|
|
|
|
|
|
@pytest.mark.parametrize(
|
|
("wrong", "right"),
|
|
[
|
|
("자기장를 봐요", "자기장을 봐요"),
|
|
("바론를 치자", "바론을 치자"),
|
|
("데드섹가 해킹했다", "데드섹이 해킹했다"),
|
|
("구급상자을 주세요", "구급상자를 주세요"),
|
|
("넥서스은 저기", "넥서스는 저기"),
|
|
("서울으로 간다", "서울로 간다"),
|
|
("칼으로 베다", "칼로 베다"),
|
|
("적으로 간다", "적으로 간다"),
|
|
],
|
|
)
|
|
def test_particles_are_corrected(wrong, right):
|
|
from livesub.models.speech_level import fix_particles
|
|
|
|
assert fix_particles(wrong) == right
|
|
|
|
|
|
def test_correct_particles_are_left_alone():
|
|
from livesub.models.speech_level import fix_particles
|
|
|
|
for text in ["적이 왼쪽에서 온다", "넥서스를 밀어", "스파이크를 설치했다", "바론을 치자"]:
|
|
assert fix_particles(text) == text
|
|
|
|
|
|
def test_particle_fix_ignores_non_hangul_and_empty():
|
|
from livesub.models.speech_level import fix_particles
|
|
|
|
assert fix_particles("") == ""
|
|
assert fix_particles("ctOS를 해킹") == "ctOS를 해킹"
|