Files
live-app-translator/tests/test_prompts.py
EJClaw bede94ee48 fix: 포터블에 transformers 포함 + 실제 모델로 드러난 용어집 결함 수정
리뷰 지적(포터블에서 transformers 제외 -> NLLB 로드 실패)을 고치려고 실제
NLLB 를 받아 돌려봤고, 그 과정에서 용어집이 사실상 동작하지 않고 있었다는
것을 발견했다. 단위 테스트는 "모델이 자리표시자를 통과시킨다"는 틀린 전제
위에 서 있었다.

1) 포터블 패키징 (리뷰 지적)
   transformers 를 제외 목록에서 빼고 hiddenimports 에 넣었다. NLLB
   토크나이저가 AutoTokenizer 를 쓰기 때문이다. transformers 는 torch 가
   없으면 토크나이저 전용 모드로 뜨며 그게 우리 용도와 정확히 맞는다.
   torch 없는 환경에서 ctranslate2/transformers/faster-whisper import 와
   앱 기동을 검증하는 test_portable.py 를 추가했다.

2) 자리표시자 형식 (실측으로 발견)
   `⟦0⟧` 는 NLLB 가 괄호를 날려 생존률 0/3 이었다. 용어가 자막에서 그냥
   사라지고 있었다 ("Third party incoming" -> "0 들어오는"). 후보 8종을
   실제 모델로 비교해 `#0#` 로 교체 (3/3, 다중 4/5).

3) 소실 대비
   모델이 문장 일부를 누락하면 자리표시자도 사라진다. 그대로 복원하면
   용어가 증발하므로, 하나라도 없으면 보호 없이 재번역한다.

4) 서술어는 문장 전체일 때만 (whole_only)
   절/서술어를 문장 중간에서 치환하면 문법이 무너진다.
     before: "탄 필요해와 구급상자"
     after : "탄약과 구급상자가 필요합니다"
   해당 56개 항목을 whole_only 로 지정해 단독 발화일 때만 적용한다.
   ("Cover me!" -> "엄호해줘" 는 그대로 유지)

5) 조사 교정
   역어 받침이 달라 "자기장를" 이 남던 것을 fix_particles() 로 고친다.
   을/를, 이/가, 은/는, 과/와, (으)로 — 한글 코드에서 받침을 읽어 판정.

검증: pytest 189개 통과, ruff clean
      실제 NLLB-600M(torch 없이 CPU)로 번역 품질 직접 확인

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-09-23 01:20:42 +09:00

83 lines
3.3 KiB
Python

"""번역 프롬프트 형식 회귀 테스트.
Seed-X 는 chat template 없는 번역 전용 completion 모델이라 모델 카드가 정한
형식을 한 글자도 벗어나면 안 된다. 특히 끝의 `<언어코드>` 태그는 PPO 학습에
쓰인 것이라 빠지면 품질이 무너진다.
"""
from __future__ import annotations
import pytest
from livesub.constants import LANGUAGE_CODES, LANGUAGES
from livesub.models.glossary import Glossary, GlossaryEntry
from livesub.models.tiers import TIERS, MTBackend, PromptStyle, get_tier
from livesub.models.translator import build_seedx_prompt
def test_seedx_prompt_matches_model_card_exactly():
"""모델 카드 예시: "Translate the following English sentence into Chinese:\\nMay the force be with you <zh>" """
assert build_seedx_prompt("May the force be with you", "en", "zh") == (
"Translate the following English sentence into Chinese:\n"
"May the force be with you <zh>"
)
def test_seedx_prompt_always_ends_with_target_language_tag():
for target in LANGUAGE_CODES:
prompt = build_seedx_prompt("hello", "en", target)
assert prompt.endswith(f" <{target}>"), f"{target} 태그 누락"
def test_seedx_prompt_has_no_extra_instructions():
"""지시문을 끼워 넣으면 Seed-X 의 학습 분포를 벗어난다."""
prompt = build_seedx_prompt("fall back now", "en", "ko")
lowered = prompt.lower()
for forbidden in ("casual", "output only", "glossary", "game or broadcast", "용어"):
assert forbidden not in lowered
assert prompt.count("\n") == 1 # 지시문 한 줄 + 본문 한 줄
@pytest.mark.parametrize("code", LANGUAGE_CODES)
def test_every_language_has_a_seedx_tag(code):
assert LANGUAGES[code]["seedx"], f"{code} 에 seedx 태그가 없습니다"
def test_seedx_tier_does_not_use_prompt_glossary():
"""Seed-X 는 지시문을 못 알아들으므로 프롬프트 용어집을 쓰면 안 된다."""
precision = get_tier("precision")
assert precision.mt.prompt_style is PromptStyle.SEEDX
assert precision.mt.supports_prompt_glossary is False
def test_instruct_tier_uses_prompt_glossary():
ultimate = get_tier("ultimate")
assert ultimate.mt.prompt_style is PromptStyle.INSTRUCT
assert ultimate.mt.supports_prompt_glossary is True
def test_seq2seq_tiers_never_use_prompt_glossary():
for tier in TIERS.values():
if tier.mt.backend is MTBackend.CTRANSLATE2:
assert tier.mt.prompt_style is PromptStyle.NONE
assert tier.mt.supports_prompt_glossary is False
def test_every_llm_tier_declares_a_prompt_style():
for tier in TIERS.values():
if tier.mt.backend is MTBackend.TRANSFORMERS:
assert tier.mt.prompt_style is not PromptStyle.NONE, tier.key
def test_glossary_survives_seedx_placeholder_path():
"""Seed-X 경로에서 용어집이 프롬프트가 아니라 치환으로 동작하는지."""
g = Glossary([GlossaryEntry("Nexus", {"ko": "넥서스"})])
protected, repl = g.protect("push to the Nexus", "ko")
prompt = build_seedx_prompt(protected, "en", "ko")
assert "Nexus" not in prompt # 모델이 건드릴 수 없게 가려짐
assert prompt.endswith(" <ko>")
# 모델이 플레이스홀더를 그대로 통과시켰다고 가정
assert Glossary.restore("#0#로 밀어", repl) == "넥서스로 밀어"