프로그램별 오디오를 캡처해 로컬 GPU에서 음성인식→번역하고 화면 위 자막으로 보여주는 데스크톱 앱. 한/영/일/중 4개 언어. 구성 - audio: WASAPI 프로그램별 캡처(C++ 보조 프로그램) + 장치 루프백 폴백, 적응형 VAD 발화 분할 - models: 속도~품질 5단계 티어, faster-whisper + CTranslate2/LLM 2백엔드, 용어집(플레이스홀더 보호 + 프롬프트 주입) - core: Qt 비의존 파이프라인 엔진 (캡처/분할/추론 3스레드, 큐 연결) - ui: 사이드바 5화면 + 무테두리 항상위 자막 오버레이, 자체 다크 테마 모델 선정 근거는 docs/MODELS.md, 추가학습 가능 여부와 방법은 docs/FINETUNING.md 참고. 검증: pytest 39개 통과 (GPU·오디오 장치 없이 실행), ruff clean Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
23 lines
752 B
Plaintext
23 lines
752 B
Plaintext
# 기본 실행에 필요한 최소 구성
|
|
PySide6>=6.6
|
|
numpy>=1.24
|
|
|
|
# --- GPU 추론 ---------------------------------------------------------
|
|
# torch 는 반드시 CUDA 빌드로 먼저 설치하세요 (README 참고):
|
|
# pip install torch --index-url https://download.pytorch.org/whl/cu128
|
|
faster-whisper>=1.0
|
|
ctranslate2>=4.4
|
|
transformers>=4.44
|
|
huggingface-hub>=0.24
|
|
sentencepiece>=0.2
|
|
accelerate>=0.33
|
|
|
|
# --- Windows 오디오 캡처 ---------------------------------------------
|
|
PyAudioWPatch>=0.2.12.7; sys_platform == "win32"
|
|
pycaw>=20240210; sys_platform == "win32"
|
|
comtypes>=1.4; sys_platform == "win32"
|
|
psutil>=5.9; sys_platform == "win32"
|
|
|
|
# --- 선택 ------------------------------------------------------------
|
|
webrtcvad-wheels>=2.0.14
|