리뷰 지적(포터블에서 transformers 제외 -> NLLB 로드 실패)을 고치려고 실제
NLLB 를 받아 돌려봤고, 그 과정에서 용어집이 사실상 동작하지 않고 있었다는
것을 발견했다. 단위 테스트는 "모델이 자리표시자를 통과시킨다"는 틀린 전제
위에 서 있었다.
1) 포터블 패키징 (리뷰 지적)
transformers 를 제외 목록에서 빼고 hiddenimports 에 넣었다. NLLB
토크나이저가 AutoTokenizer 를 쓰기 때문이다. transformers 는 torch 가
없으면 토크나이저 전용 모드로 뜨며 그게 우리 용도와 정확히 맞는다.
torch 없는 환경에서 ctranslate2/transformers/faster-whisper import 와
앱 기동을 검증하는 test_portable.py 를 추가했다.
2) 자리표시자 형식 (실측으로 발견)
`⟦0⟧` 는 NLLB 가 괄호를 날려 생존률 0/3 이었다. 용어가 자막에서 그냥
사라지고 있었다 ("Third party incoming" -> "0 들어오는"). 후보 8종을
실제 모델로 비교해 `#0#` 로 교체 (3/3, 다중 4/5).
3) 소실 대비
모델이 문장 일부를 누락하면 자리표시자도 사라진다. 그대로 복원하면
용어가 증발하므로, 하나라도 없으면 보호 없이 재번역한다.
4) 서술어는 문장 전체일 때만 (whole_only)
절/서술어를 문장 중간에서 치환하면 문법이 무너진다.
before: "탄 필요해와 구급상자"
after : "탄약과 구급상자가 필요합니다"
해당 56개 항목을 whole_only 로 지정해 단독 발화일 때만 적용한다.
("Cover me!" -> "엄호해줘" 는 그대로 유지)
5) 조사 교정
역어 받침이 달라 "자기장를" 이 남던 것을 fix_particles() 로 고친다.
을/를, 이/가, 은/는, 과/와, (으)로 — 한글 코드에서 받침을 읽어 판정.
검증: pytest 189개 통과, ruff clean
실제 NLLB-600M(torch 없이 CPU)로 번역 품질 직접 확인
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
94 lines
2.9 KiB
Python
94 lines
2.9 KiB
Python
from livesub.models.glossary import Glossary, GlossaryEntry
|
|
|
|
|
|
def make() -> Glossary:
|
|
return Glossary(
|
|
[
|
|
GlossaryEntry("headshot", {"ko": "헤드샷"}),
|
|
GlossaryEntry("head", {"ko": "머리"}),
|
|
GlossaryEntry("Nexus", {"ko": "넥서스", "ja": "ネクサス"}),
|
|
]
|
|
)
|
|
|
|
|
|
def test_protect_and_restore_roundtrip():
|
|
g = make()
|
|
protected, repl = g.protect("Nice headshot on the Nexus", "ko")
|
|
assert "headshot" not in protected
|
|
assert "Nexus" not in protected
|
|
assert repl == ["헤드샷", "넥서스"]
|
|
# 번역 모델이 플레이스홀더를 그대로 통과시켰다고 가정
|
|
assert Glossary.restore(protected, repl) == "Nice 헤드샷 on the 넥서스"
|
|
|
|
|
|
def test_longest_match_wins():
|
|
"""'headshot'이 'head'에 먼저 잡아먹히면 안 된다."""
|
|
g = make()
|
|
_, repl = g.protect("headshot", "ko")
|
|
assert repl == ["헤드샷"]
|
|
|
|
|
|
def test_word_boundary_for_latin_terms():
|
|
g = make()
|
|
_, repl = g.protect("overheated", "ko")
|
|
assert repl == []
|
|
|
|
|
|
def test_missing_target_language_is_left_alone():
|
|
g = make()
|
|
protected, repl = g.protect("headshot", "ja")
|
|
assert protected == "headshot"
|
|
assert repl == []
|
|
|
|
|
|
def test_prompt_hint_only_includes_present_terms():
|
|
g = make()
|
|
hint = g.prompt_hint("push to the Nexus", "ko")
|
|
assert "Nexus -> 넥서스" in hint
|
|
assert "headshot" not in hint
|
|
assert g.prompt_hint("nothing here", "ko") == ""
|
|
|
|
|
|
def test_case_insensitive_by_default():
|
|
g = make()
|
|
_, repl = g.protect("NEXUS down", "ko")
|
|
assert repl == ["넥서스"]
|
|
|
|
|
|
def test_save_and_load(tmp_path):
|
|
path = tmp_path / "glossary.json"
|
|
make().save(path)
|
|
loaded = Glossary.load(path)
|
|
assert len(loaded) == 3
|
|
assert loaded.entries[2].target_for("ja") == "ネクサス"
|
|
|
|
|
|
def test_add_replaces_existing_entry():
|
|
g = make()
|
|
g.add(GlossaryEntry("Nexus", {"ko": "본진"}))
|
|
assert len(g) == 3
|
|
_, repl = g.protect("Nexus", "ko")
|
|
assert repl == ["본진"]
|
|
|
|
|
|
def test_restore_ignores_out_of_range_placeholder():
|
|
assert Glossary.restore("a #5# b", ["x"]) == "a b"
|
|
|
|
|
|
def test_missing_placeholder_is_reported():
|
|
"""번역이 문장 일부를 날리면 자리표시자도 사라진다 — 호출부가 알아야 한다."""
|
|
assert Glossary.missing_placeholders("#0# 만 남음", ["가", "나"]) == [1]
|
|
assert Glossary.missing_placeholders("#0# #1#", ["가", "나"]) == []
|
|
assert Glossary.missing_placeholders("아무것도 없음", ["가"]) == [0]
|
|
|
|
|
|
def test_placeholder_format_is_the_measured_one():
|
|
"""NLLB 실측에서 살아남은 형식(#0#)을 유지해야 한다.
|
|
|
|
`⟦0⟧` 같은 특수 괄호는 NLLB 가 통째로 날려 용어가 자막에서 사라졌다.
|
|
형식을 바꾸려면 실제 모델로 생존률을 다시 재고 이 테스트를 고칠 것.
|
|
"""
|
|
g = Glossary([GlossaryEntry("baron", {"ko": "바론"})])
|
|
protected, _ = g.protect("take baron", "ko")
|
|
assert protected == "take #0#"
|