Files
live-app-translator/tests/test_glossary.py
EJClaw bede94ee48 fix: 포터블에 transformers 포함 + 실제 모델로 드러난 용어집 결함 수정
리뷰 지적(포터블에서 transformers 제외 -> NLLB 로드 실패)을 고치려고 실제
NLLB 를 받아 돌려봤고, 그 과정에서 용어집이 사실상 동작하지 않고 있었다는
것을 발견했다. 단위 테스트는 "모델이 자리표시자를 통과시킨다"는 틀린 전제
위에 서 있었다.

1) 포터블 패키징 (리뷰 지적)
   transformers 를 제외 목록에서 빼고 hiddenimports 에 넣었다. NLLB
   토크나이저가 AutoTokenizer 를 쓰기 때문이다. transformers 는 torch 가
   없으면 토크나이저 전용 모드로 뜨며 그게 우리 용도와 정확히 맞는다.
   torch 없는 환경에서 ctranslate2/transformers/faster-whisper import 와
   앱 기동을 검증하는 test_portable.py 를 추가했다.

2) 자리표시자 형식 (실측으로 발견)
   `⟦0⟧` 는 NLLB 가 괄호를 날려 생존률 0/3 이었다. 용어가 자막에서 그냥
   사라지고 있었다 ("Third party incoming" -> "0 들어오는"). 후보 8종을
   실제 모델로 비교해 `#0#` 로 교체 (3/3, 다중 4/5).

3) 소실 대비
   모델이 문장 일부를 누락하면 자리표시자도 사라진다. 그대로 복원하면
   용어가 증발하므로, 하나라도 없으면 보호 없이 재번역한다.

4) 서술어는 문장 전체일 때만 (whole_only)
   절/서술어를 문장 중간에서 치환하면 문법이 무너진다.
     before: "탄 필요해와 구급상자"
     after : "탄약과 구급상자가 필요합니다"
   해당 56개 항목을 whole_only 로 지정해 단독 발화일 때만 적용한다.
   ("Cover me!" -> "엄호해줘" 는 그대로 유지)

5) 조사 교정
   역어 받침이 달라 "자기장를" 이 남던 것을 fix_particles() 로 고친다.
   을/를, 이/가, 은/는, 과/와, (으)로 — 한글 코드에서 받침을 읽어 판정.

검증: pytest 189개 통과, ruff clean
      실제 NLLB-600M(torch 없이 CPU)로 번역 품질 직접 확인

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-09-23 01:20:42 +09:00

94 lines
2.9 KiB
Python

from livesub.models.glossary import Glossary, GlossaryEntry
def make() -> Glossary:
return Glossary(
[
GlossaryEntry("headshot", {"ko": "헤드샷"}),
GlossaryEntry("head", {"ko": "머리"}),
GlossaryEntry("Nexus", {"ko": "넥서스", "ja": "ネクサス"}),
]
)
def test_protect_and_restore_roundtrip():
g = make()
protected, repl = g.protect("Nice headshot on the Nexus", "ko")
assert "headshot" not in protected
assert "Nexus" not in protected
assert repl == ["헤드샷", "넥서스"]
# 번역 모델이 플레이스홀더를 그대로 통과시켰다고 가정
assert Glossary.restore(protected, repl) == "Nice 헤드샷 on the 넥서스"
def test_longest_match_wins():
"""'headshot'이 'head'에 먼저 잡아먹히면 안 된다."""
g = make()
_, repl = g.protect("headshot", "ko")
assert repl == ["헤드샷"]
def test_word_boundary_for_latin_terms():
g = make()
_, repl = g.protect("overheated", "ko")
assert repl == []
def test_missing_target_language_is_left_alone():
g = make()
protected, repl = g.protect("headshot", "ja")
assert protected == "headshot"
assert repl == []
def test_prompt_hint_only_includes_present_terms():
g = make()
hint = g.prompt_hint("push to the Nexus", "ko")
assert "Nexus -> 넥서스" in hint
assert "headshot" not in hint
assert g.prompt_hint("nothing here", "ko") == ""
def test_case_insensitive_by_default():
g = make()
_, repl = g.protect("NEXUS down", "ko")
assert repl == ["넥서스"]
def test_save_and_load(tmp_path):
path = tmp_path / "glossary.json"
make().save(path)
loaded = Glossary.load(path)
assert len(loaded) == 3
assert loaded.entries[2].target_for("ja") == "ネクサス"
def test_add_replaces_existing_entry():
g = make()
g.add(GlossaryEntry("Nexus", {"ko": "본진"}))
assert len(g) == 3
_, repl = g.protect("Nexus", "ko")
assert repl == ["본진"]
def test_restore_ignores_out_of_range_placeholder():
assert Glossary.restore("a #5# b", ["x"]) == "a b"
def test_missing_placeholder_is_reported():
"""번역이 문장 일부를 날리면 자리표시자도 사라진다 — 호출부가 알아야 한다."""
assert Glossary.missing_placeholders("#0# 만 남음", ["가", "나"]) == [1]
assert Glossary.missing_placeholders("#0# #1#", ["가", "나"]) == []
assert Glossary.missing_placeholders("아무것도 없음", ["가"]) == [0]
def test_placeholder_format_is_the_measured_one():
"""NLLB 실측에서 살아남은 형식(#0#)을 유지해야 한다.
`⟦0⟧` 같은 특수 괄호는 NLLB 가 통째로 날려 용어가 자막에서 사라졌다.
형식을 바꾸려면 실제 모델로 생존률을 다시 재고 이 테스트를 고칠 것.
"""
g = Glossary([GlossaryEntry("baron", {"ko": "바론"})])
protected, _ = g.protect("take baron", "ko")
assert protected == "take #0#"