Compare commits
2 Commits
ec1bc70b04
...
1a87ec6677
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
1a87ec6677 | ||
|
|
c24e5f938d |
31
.gitignore
vendored
Normal file
31
.gitignore
vendored
Normal file
@@ -0,0 +1,31 @@
|
|||||||
|
# Python
|
||||||
|
__pycache__/
|
||||||
|
*.py[cod]
|
||||||
|
*.egg-info/
|
||||||
|
build/
|
||||||
|
dist/
|
||||||
|
.venv/
|
||||||
|
venv/
|
||||||
|
.pytest_cache/
|
||||||
|
.ruff_cache/
|
||||||
|
|
||||||
|
# 네이티브 빌드 산출물
|
||||||
|
native/process_loopback/build/
|
||||||
|
src/hearo/resources/bin/*.exe
|
||||||
|
src/hearo/resources/bin/*.pdb
|
||||||
|
|
||||||
|
# 모델 캐시 / 학습 산출물 / 사용자 데이터
|
||||||
|
models/
|
||||||
|
adapters/
|
||||||
|
data/
|
||||||
|
*.jsonl
|
||||||
|
config.json
|
||||||
|
glossary.json
|
||||||
|
transcripts/
|
||||||
|
*.log
|
||||||
|
|
||||||
|
# 에디터
|
||||||
|
.vscode/
|
||||||
|
.idea/
|
||||||
|
*.swp
|
||||||
|
.DS_Store
|
||||||
21
LICENSE
Normal file
21
LICENSE
Normal file
@@ -0,0 +1,21 @@
|
|||||||
|
MIT License
|
||||||
|
|
||||||
|
Copyright (c) 2026 tkrmagid
|
||||||
|
|
||||||
|
Permission is hereby granted, free of charge, to any person obtaining a copy
|
||||||
|
of this software and associated documentation files (the "Software"), to deal
|
||||||
|
in the Software without restriction, including without limitation the rights
|
||||||
|
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
||||||
|
copies of the Software, and to permit persons to whom the Software is
|
||||||
|
furnished to do so, subject to the following conditions:
|
||||||
|
|
||||||
|
The above copyright notice and this permission notice shall be included in all
|
||||||
|
copies or substantial portions of the Software.
|
||||||
|
|
||||||
|
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
||||||
|
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
||||||
|
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
||||||
|
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
||||||
|
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
||||||
|
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
||||||
|
SOFTWARE.
|
||||||
169
README.md
169
README.md
@@ -1,2 +1,169 @@
|
|||||||
# live-app-translator
|
# Hearo — 듣고, 바로 이해하다
|
||||||
|
|
||||||
|
특정 프로그램에서 나오는 소리를 실시간으로 받아 번역해 자막으로 보여줍니다.
|
||||||
|
번역은 전부 **내 컴퓨터의 GPU에서** 돌아갑니다. 인터넷도, API 키도 필요 없습니다.
|
||||||
|
|
||||||
|
한국어 · English · 日本語 · 中文 사이를 번역합니다.
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 무엇을 하는 프로그램인가
|
||||||
|
|
||||||
|
게임을 하는데 영어 음성만 나올 때, 해외 방송을 보는데 자막이 없을 때,
|
||||||
|
그 프로그램을 골라서 `번역 시작`만 누르면 화면 아래에 한국어 자막이 뜹니다.
|
||||||
|
|
||||||
|
```
|
||||||
|
게임 소리 → 음성인식 → 번역 → 화면 위 자막
|
||||||
|
(Whisper) (Seed-X / NLLB)
|
||||||
|
```
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
## 주요 기능
|
||||||
|
|
||||||
|
- **프로그램 단위 소리 캡처** — 게임 소리만 받고 디스코드 음성은 안 받습니다
|
||||||
|
- **5단계 품질 선택** — 속도 우선(0.6초)부터 품질 우선까지, GPU 사양에 맞춰 고릅니다
|
||||||
|
- **자막 자유 설정** — 글꼴·크기·색·외곽선·투명도·위치·줄 수
|
||||||
|
- **용어집** — 게임 고유명사를 원하는 번역으로 고정. 학습 없이 즉시 적용
|
||||||
|
- **추가학습** — 게임/방송 말투로 번역 모델을 LoRA 학습시킬 수 있습니다
|
||||||
|
- **완전 로컬** — 음성이 외부로 나가지 않습니다
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 설치
|
||||||
|
|
||||||
|
### 1. 사전 준비
|
||||||
|
|
||||||
|
- Windows 10 (2004 이상) 또는 Windows 11
|
||||||
|
- NVIDIA GPU (권장 VRAM 6GB 이상) + 최신 드라이버
|
||||||
|
- Python 3.10 ~ 3.12
|
||||||
|
|
||||||
|
### 2. PyTorch (CUDA 빌드) 먼저
|
||||||
|
|
||||||
|
CPU 버전이 깔리면 GPU를 못 씁니다. 반드시 이 순서로 설치하세요.
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
py -3.12 -m venv .venv
|
||||||
|
.venv\Scripts\activate
|
||||||
|
|
||||||
|
# Blackwell(RTX 50 시리즈)은 cu128, 그 이전 세대는 cu124
|
||||||
|
pip install torch --index-url https://download.pytorch.org/whl/cu128
|
||||||
|
```
|
||||||
|
|
||||||
|
### 3. 나머지 설치
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
pip install -r requirements.txt
|
||||||
|
pip install -e .
|
||||||
|
```
|
||||||
|
|
||||||
|
### 4. 프로그램별 캡처 켜기 (선택, 권장)
|
||||||
|
|
||||||
|
이걸 빌드하지 않으면 출력 장치 전체 소리를 받습니다 (다른 앱 소리도 섞임).
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
# Visual Studio 2022 Build Tools (C++ 데스크톱) + CMake 필요
|
||||||
|
winget install Microsoft.VisualStudio.2022.BuildTools
|
||||||
|
winget install Kitware.CMake
|
||||||
|
|
||||||
|
powershell -ExecutionPolicy Bypass -File native\process_loopback\build.ps1
|
||||||
|
```
|
||||||
|
|
||||||
|
### 5. 실행
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
hearo
|
||||||
|
```
|
||||||
|
|
||||||
|
또는 `python -m hearo`
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 처음 쓸 때
|
||||||
|
|
||||||
|
1. **모델** 화면에서 티어를 고릅니다. VRAM 8GB면 `3. 균형`, 10GB 이상이면 `4. 정밀`
|
||||||
|
2. **홈** 화면에서 소리를 받아올 프로그램을 고릅니다
|
||||||
|
3. 원본 언어(자동 감지 권장)와 번역할 언어를 고릅니다
|
||||||
|
4. `번역 시작`
|
||||||
|
|
||||||
|
모델은 처음 한 번만 자동으로 내려받습니다 (2~17GB, 티어에 따라 다름).
|
||||||
|
저장 위치는 `%APPDATA%\Hearo\models` 입니다.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 모델 티어
|
||||||
|
|
||||||
|
| # | 이름 | VRAM | 지연 | 설명 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 1 | 번개 | 2GB | 0.6초 | 속도 최우선. 저사양 GPU |
|
||||||
|
| 2 | 신속 | 4GB | 0.9초 | 인식률은 최상급, 번역만 경량 |
|
||||||
|
| 3 | **균형** | 6GB | 1.2초 | 기본값. 8GB GPU에 가장 무난 |
|
||||||
|
| 4 | **정밀** | 10GB | 2.0초 | **추천.** 현 시점 최고 번역 품질 |
|
||||||
|
| 5 | 극한 | 16GB | 3.0초 | 추가학습(LoRA)에 최적 |
|
||||||
|
|
||||||
|

|
||||||
|
|
||||||
|
어떤 모델을 왜 골랐는지는 [docs/MODELS.md](docs/MODELS.md)에 정리했습니다.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 용어집
|
||||||
|
|
||||||
|
게임 고유명사가 이상하게 번역될 때 씁니다. **학습이 필요 없고 즉시 적용됩니다.**
|
||||||
|
|
||||||
|
`용어집` 화면에서 직접 입력하거나 CSV로 가져오세요.
|
||||||
|
|
||||||
|
```csv
|
||||||
|
원문 용어,한국어 역어,English 역어,日本語 역어,中文 역어,메모
|
||||||
|
Nexus,넥서스,Nexus,ネクサス,基地,LoL
|
||||||
|
baron,바론,Baron,バロン,男爵,LoL
|
||||||
|
```
|
||||||
|
|
||||||
|
게임 하나당 200~500개만 등록해도 체감 품질이 크게 달라집니다.
|
||||||
|
|
||||||
|
## 추가학습
|
||||||
|
|
||||||
|
말투나 문장 구조까지 바꾸고 싶다면 LoRA로 학습시킬 수 있습니다.
|
||||||
|
언어쌍당 문장 1,000~3,000개가 필요합니다.
|
||||||
|
|
||||||
|
```powershell
|
||||||
|
pip install -e ".[finetune]"
|
||||||
|
python scripts\finetune_mt.py --data data\game.jsonl --tier ultimate --output .\adapters\game-ko
|
||||||
|
```
|
||||||
|
|
||||||
|
데이터를 어떻게 모으고 어느 모델을 학습시켜야 하는지는
|
||||||
|
[docs/FINETUNING.md](docs/FINETUNING.md)를 참고하세요.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 자막 창 조작
|
||||||
|
|
||||||
|
| 동작 | 결과 |
|
||||||
|
|---|---|
|
||||||
|
| 드래그 | 위치 이동 |
|
||||||
|
| 우하단 모서리 드래그 | 크기 조절 |
|
||||||
|
| 마우스 휠 | 글자 크기 |
|
||||||
|
| 우클릭 | 잠금 / 클릭 통과 / 항상 위에 / 숨기기 |
|
||||||
|
|
||||||
|
전체화면 게임에서 자막이 안 보이면 게임을 **테두리 없는 창 모드**로 바꾸세요.
|
||||||
|
독점 전체화면(exclusive fullscreen)에서는 어떤 오버레이도 표시되지 않습니다.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 개발
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install -e ".[dev]"
|
||||||
|
pytest # GPU·오디오 장치 없이 실행됩니다
|
||||||
|
ruff check src tests scripts
|
||||||
|
```
|
||||||
|
|
||||||
|
구조는 [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md)를 보세요.
|
||||||
|
|
||||||
|
## 라이선스
|
||||||
|
|
||||||
|
MIT. 사용하는 모델은 각자의 라이선스를 따릅니다
|
||||||
|
(Whisper: MIT, NLLB: CC-BY-NC, Seed-X: OpenMDW, Qwen3: Apache-2.0).
|
||||||
|
**NLLB는 비상업적 이용만 허용됩니다.** 상업적으로 쓰려면 4~5티어를 사용하세요.
|
||||||
|
|||||||
91
docs/ARCHITECTURE.md
Normal file
91
docs/ARCHITECTURE.md
Normal file
@@ -0,0 +1,91 @@
|
|||||||
|
# 구조
|
||||||
|
|
||||||
|
## 데이터 흐름
|
||||||
|
|
||||||
|
```
|
||||||
|
[캡처 스레드] [분할 스레드] [추론 스레드] [GUI 스레드]
|
||||||
|
hearo_capture.exe → Segmenter (VAD) → Whisper → 번역 모델 → 자막 오버레이
|
||||||
|
또는 WASAPI 발화 단위로 절단 용어집 적용 컨트롤 창 로그
|
||||||
|
↓ ↓ ↓ ↑
|
||||||
|
오디오 큐 구간 큐(최대 32) TranslationLine Qt Signal
|
||||||
|
(16kHz mono f32) 확정 / 중간결과
|
||||||
|
```
|
||||||
|
|
||||||
|
세 워커 스레드는 모두 큐로만 연결된다. 어느 한 단계가 밀려도 나머지가 멈추지 않는다.
|
||||||
|
큐가 가득 차면 **오래된 오디오를 버린다** — 실시간 자막에서는 밀린 소리보다 지금 나는
|
||||||
|
소리가 항상 더 가치 있기 때문이다.
|
||||||
|
|
||||||
|
## 패키지
|
||||||
|
|
||||||
|
```
|
||||||
|
src/hearo/
|
||||||
|
constants.py 전역 상수, 언어 표, 경로
|
||||||
|
config.py 설정 dataclass + JSON 영속화
|
||||||
|
app.py 진입점
|
||||||
|
|
||||||
|
audio/
|
||||||
|
base.py CaptureBackend 추상 클래스, 리샘플링
|
||||||
|
process_loopback.py 프로그램별 캡처 (네이티브 보조 프로그램 구동)
|
||||||
|
wasapi_loopback.py 출력 장치 전체 캡처 (PyAudioWPatch)
|
||||||
|
file_source.py WAV 재생 (테스트·데모용)
|
||||||
|
segmenter.py 적응형 VAD로 발화 구간 절단
|
||||||
|
|
||||||
|
models/
|
||||||
|
tiers.py 5단계 티어 정의
|
||||||
|
asr.py faster-whisper 래퍼
|
||||||
|
translator.py CTranslate2 / Transformers 두 백엔드
|
||||||
|
glossary.py 용어집 (플레이스홀더 보호 + 프롬프트 주입)
|
||||||
|
manager.py 티어별 모델 쌍 로드·해제, GPU 탐지
|
||||||
|
|
||||||
|
core/
|
||||||
|
engine.py 파이프라인 오케스트레이션 (Qt 비의존)
|
||||||
|
events.py TranslationLine, EngineStatus
|
||||||
|
|
||||||
|
ui/
|
||||||
|
theme.py 디자인 토큰 + QSS
|
||||||
|
overlay.py 자막 창 + 공용 렌더링 함수
|
||||||
|
main_window.py 사이드바 + 페이지 스택
|
||||||
|
pages/ 홈 / 모델 / 자막 / 용어집 / 설정
|
||||||
|
widgets/ Card, StatusPill, LevelMeter
|
||||||
|
|
||||||
|
native/process_loopback/ WASAPI process loopback C++ 보조 프로그램
|
||||||
|
scripts/finetune_mt.py LoRA 추가학습
|
||||||
|
```
|
||||||
|
|
||||||
|
## 설계 판단
|
||||||
|
|
||||||
|
### 캡처를 별도 실행 파일로 뺀 이유
|
||||||
|
|
||||||
|
WASAPI의 `ActivateAudioInterfaceAsync` + `AUDIOCLIENT_ACTIVATION_TYPE_PROCESS_LOOPBACK`
|
||||||
|
은 COM 비동기 콜백을 요구한다. ctypes로 COM vtable을 흉내 내는 것보다 C++ 200줄이
|
||||||
|
훨씬 안전하고 디버깅이 쉽다. 파이프로 생 PCM만 주고받으므로 인터페이스도 단순하다.
|
||||||
|
|
||||||
|
보조 프로그램이 없으면 `capture_capabilities()` 가 이를 알려주고 장치 루프백으로
|
||||||
|
자동 폴백한다. 빌드 없이도 프로그램은 동작한다.
|
||||||
|
|
||||||
|
### 중간 결과는 번역하지 않는다
|
||||||
|
|
||||||
|
말하는 도중에도 인식 결과를 흐리게 보여주면 체감 지연이 크게 줄어든다. 하지만
|
||||||
|
매번 번역까지 돌리면 GPU가 몇 배로 바빠지고, 문장이 완성되기 전 번역은 어차피
|
||||||
|
틀린다. 그래서 **중간 결과는 원문만, 확정 문장만 번역**한다.
|
||||||
|
|
||||||
|
### UI 코드와 렌더링 코드 분리
|
||||||
|
|
||||||
|
자막 설정 미리보기와 실제 오버레이가 `paint_subtitle()` 하나를 공유한다.
|
||||||
|
QSS로 미리보기를 흉내 내면 `text-shadow` 미지원 같은 이유로 실제와 어긋난다.
|
||||||
|
|
||||||
|
### 엔진은 Qt를 모른다
|
||||||
|
|
||||||
|
`core/engine.py` 는 콜백만 받는다. Qt 의존은 `MainWindow` 가 콜백을 Signal로
|
||||||
|
다시 던지는 지점에만 있다. 덕분에 GPU도 오디오 장치도 없는 환경에서 파이프라인
|
||||||
|
전체를 테스트할 수 있다 (`tests/test_pipeline.py`).
|
||||||
|
|
||||||
|
## 테스트
|
||||||
|
|
||||||
|
`pytest` 전체가 GPU·오디오 장치 없이 돈다.
|
||||||
|
|
||||||
|
- `test_segmenter.py` — 합성 사인파로 VAD 절단 검증
|
||||||
|
- `test_pipeline.py` — 가짜 캡처/인식/번역으로 엔진 종단 검증
|
||||||
|
- `test_glossary.py` — 최장일치, 단어경계, 왕복 변환
|
||||||
|
- `test_config.py` — 스키마 변경 내성
|
||||||
|
- `test_ui_smoke.py` — offscreen 렌더링으로 화면 생성 및 자막 픽셀 검증
|
||||||
137
docs/FINETUNING.md
Normal file
137
docs/FINETUNING.md
Normal file
@@ -0,0 +1,137 @@
|
|||||||
|
# 게임·방송 용어 추가학습
|
||||||
|
|
||||||
|
## 결론부터
|
||||||
|
|
||||||
|
**가능합니다.** 다만 대부분의 경우 **추가학습보다 용어집이 먼저입니다.**
|
||||||
|
|
||||||
|
| 방법 | 준비 | 효과가 나타나는 시점 | 무엇에 좋은가 |
|
||||||
|
|---|---|---|---|
|
||||||
|
| **1. 용어집** (구현 완료) | 단어 목록만 | **즉시** | 고유명사, 스킬명, 아이템명, 캐릭터명 |
|
||||||
|
| **2. LoRA 추가학습** | 문장 쌍 1,000~3,000개 | 30분~3시간 학습 | 말투, 문장 구조, 도메인 어조 |
|
||||||
|
| **3. 풀 파인튜닝** | 문장 쌍 10,000개 이상 | 수 시간~하루 | 번역 스타일 전면 교체 |
|
||||||
|
|
||||||
|
용어집과 추가학습은 **경쟁 관계가 아니라 보완 관계**입니다. 둘 다 켜는 게 가장 좋습니다.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 1. 용어집 — 먼저 이것부터
|
||||||
|
|
||||||
|
`용어집` 화면에서 "원문 → 각 언어 역어"를 등록하면 끝입니다. 학습 없이 바로 적용됩니다.
|
||||||
|
|
||||||
|
동작 방식은 번역 백엔드에 따라 다릅니다.
|
||||||
|
|
||||||
|
- **NLLB 계열(1~3티어)** — 등록 단어를 `⟦0⟧` 같은 토큰으로 바꿔치기해 모델이 아예
|
||||||
|
건드리지 못하게 한 뒤, 번역이 끝나면 지정한 역어로 되돌립니다. 100% 보장됩니다.
|
||||||
|
- **LLM 계열(4~5티어)** — 그 문장에 실제로 나온 용어만 골라 프롬프트에
|
||||||
|
"이 용어는 이렇게 옮겨라"로 넣어줍니다. 조사·어미까지 문맥에 맞게 붙습니다.
|
||||||
|
|
||||||
|
CSV로 한 번에 가져올 수 있습니다.
|
||||||
|
|
||||||
|
```csv
|
||||||
|
원문 용어,한국어 역어,English 역어,日本語 역어,中文 역어,메모
|
||||||
|
Nexus,넥서스,Nexus,ネクサス,基地,LoL
|
||||||
|
baron,바론,Baron,バロン,男爵,LoL
|
||||||
|
ult,궁,ultimate,アルティメット,大招,궁극기
|
||||||
|
gg,잘 싸웠다,gg,gg,打得好,
|
||||||
|
```
|
||||||
|
|
||||||
|
**게임 하나당 200~500개만 등록해도 체감 품질이 크게 올라갑니다.** 추가학습으로
|
||||||
|
같은 효과를 내려면 훨씬 많은 데이터가 필요합니다. 용어집을 먼저 채우세요.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 2. LoRA 추가학습 — 용어집으로 안 되는 것
|
||||||
|
|
||||||
|
용어집은 "이 단어를 저 단어로"만 고칩니다. 아래는 못 잡습니다.
|
||||||
|
|
||||||
|
- 말투 — "You're getting rolled" → "탈탈 털리고 있네" (직역 아닌 게임 말투)
|
||||||
|
- 생략된 주어·목적어 복원 — 게임 음성에는 생략이 많습니다
|
||||||
|
- 방송 특유의 감탄·리액션 어조
|
||||||
|
|
||||||
|
이건 문장 쌍으로 학습시켜야 합니다.
|
||||||
|
|
||||||
|
### 필요한 데이터 양
|
||||||
|
|
||||||
|
공개 연구 기준으로,
|
||||||
|
|
||||||
|
- 언어쌍당 **약 650~1,000 문장**만으로도 LoRA 도메인 적응이 동작합니다.
|
||||||
|
- 언어쌍당 **2,000 문장 / 15~20 epoch** 수준에서 chrF++ 평균 **+6.1점** 개선이 보고됩니다.
|
||||||
|
|
||||||
|
즉 **언어쌍당 1,000~3,000 문장이 현실적인 목표**입니다. 4개 언어 전부가 아니라,
|
||||||
|
실제로 많이 쓰는 방향(예: 영어→한국어) 하나만 먼저 하는 게 효율적입니다.
|
||||||
|
|
||||||
|
### 데이터 만드는 법
|
||||||
|
|
||||||
|
1. **게임 공식 현지화 자산** — 한국어판이 있는 게임의 자막/UI 텍스트. 품질이 가장 좋습니다.
|
||||||
|
2. **자막 파일 쌍** — 같은 영상의 영어 자막 + 한국어 자막(.srt)을 시간축으로 정렬.
|
||||||
|
3. **이 프로그램의 기록** — `설정 > 번역 기록을 파일로 남기기`를 켜두면
|
||||||
|
`%APPDATA%/Hearo/transcripts/` 에 "원문 / 번역" 쌍이 쌓입니다.
|
||||||
|
**틀린 번역만 손으로 고쳐서** 학습 데이터로 쓰는 게 가장 현실적인 경로입니다.
|
||||||
|
4. 위키·커뮤니티 용어 사전 — 용어집으로 쓰는 게 더 낫습니다. 문장이 아니므로.
|
||||||
|
|
||||||
|
형식은 JSONL 한 줄에 한 쌍입니다.
|
||||||
|
|
||||||
|
```jsonl
|
||||||
|
{"source": "Enemy missing from mid", "target": "미드 실종", "source_lang": "en", "target_lang": "ko"}
|
||||||
|
{"source": "I'll take baron", "target": "바론 내가 먹을게", "source_lang": "en", "target_lang": "ko"}
|
||||||
|
```
|
||||||
|
|
||||||
|
### 학습 실행
|
||||||
|
|
||||||
|
```bash
|
||||||
|
pip install -e ".[finetune]"
|
||||||
|
|
||||||
|
python scripts/finetune_mt.py \
|
||||||
|
--data data/game_terms.jsonl \
|
||||||
|
--tier ultimate \
|
||||||
|
--output ./adapters/game-ko \
|
||||||
|
--epochs 3
|
||||||
|
```
|
||||||
|
|
||||||
|
끝나면 `모델` 화면의 **추가학습 어댑터** 칸에 `./adapters/game-ko` 를 넣으면 적용됩니다.
|
||||||
|
|
||||||
|
### 티어별 난이도
|
||||||
|
|
||||||
|
| 티어 | 번역 모델 | 방식 | 학습에 필요한 VRAM | 난이도 |
|
||||||
|
|---|---|---|---|---|
|
||||||
|
| 1~2 | NLLB-600M | 풀 파인튜닝 가능 | ~8GB | ★ 쉬움 |
|
||||||
|
| 3 | NLLB-1.3B | LoRA | ~10GB | ★★ |
|
||||||
|
| 4 | Seed-X-PPO-7B | LoRA (권장 안 함) | ~16GB | ★★★ 까다로움 |
|
||||||
|
| 5 | **Qwen3-8B** | **LoRA** | **~16GB** | **★ 쉬움 (권장)** |
|
||||||
|
|
||||||
|
**4티어 Seed-X는 추가학습 대상으로 권하지 않습니다.** 이미 PPO(강화학습)까지 마친
|
||||||
|
모델이라 그 위에 SFT를 얹으면 기존 번역 품질이 무너지기 쉽습니다 (catastrophic forgetting).
|
||||||
|
Int4 양자화 가중치라 학습 자체도 번거롭습니다.
|
||||||
|
|
||||||
|
**추가학습을 할 거라면 5티어(Qwen3-8B)**, **VRAM이 부족하면 1~2티어(NLLB-600M 풀 파인튜닝)**
|
||||||
|
로 가는 게 맞습니다.
|
||||||
|
|
||||||
|
### 권장 하이퍼파라미터
|
||||||
|
|
||||||
|
연구에서 널리 쓰이는 설정입니다. `scripts/finetune_mt.py` 의 기본값이기도 합니다.
|
||||||
|
|
||||||
|
```
|
||||||
|
LoRA rank r = 16, alpha = 32, dropout = 0.05
|
||||||
|
대상 모듈: attention 의 q_proj, v_proj
|
||||||
|
learning rate = 2e-4 (LoRA) / 5e-5 (풀 파인튜닝)
|
||||||
|
epochs = 3 (데이터 2,000개 이상) / 10~20 (수백 개)
|
||||||
|
```
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## 3. 어느 것부터 할지
|
||||||
|
|
||||||
|
```
|
||||||
|
1주차 용어집 300개 등록 → 이것만으로 충분한지 확인
|
||||||
|
2주차 번역 기록 켜고 실사용, 틀린 문장 수집
|
||||||
|
3주차 고친 문장 1,000개 모이면 LoRA 학습
|
||||||
|
```
|
||||||
|
|
||||||
|
**용어집만으로 만족스러우면 추가학습은 안 해도 됩니다.** 실제로 고유명사 오역이
|
||||||
|
체감 불만의 대부분입니다.
|
||||||
|
|
||||||
|
## 참고
|
||||||
|
|
||||||
|
- [Fine-Tuning NLLB-200 with LoRA on a 650-Sentence Corpus](https://medium.com/@meinnps/fine-tuning-nllb-200-with-lora-on-a-650-sentence-turkmen-english-corpus-082f68bdec71)
|
||||||
|
- [SemiAdapt / SemiLoRA: Efficient Domain Adaptation for Low-Resource MT (arXiv)](https://arxiv.org/pdf/2510.18725)
|
||||||
|
- [How to fine-tune a NLLB-200 model](https://cointegrated.medium.com/how-to-fine-tune-a-nllb-200-model-for-translating-a-new-language-a37fc706b865)
|
||||||
81
docs/MODELS.md
Normal file
81
docs/MODELS.md
Normal file
@@ -0,0 +1,81 @@
|
|||||||
|
# 모델 선정 근거
|
||||||
|
|
||||||
|
## 왜 두 단계인가
|
||||||
|
|
||||||
|
"소리 → 번역"을 한 모델로 끝내는 방법도 있다. Whisper에는 `translate` 태스크가 있고
|
||||||
|
Voxtral·Qwen-Omni 같은 음성 LLM도 있다. 하지만 이 프로그램에는 맞지 않는다.
|
||||||
|
|
||||||
|
- Whisper의 `translate`는 **X → 영어만** 된다. 한국어로 받아볼 수 없다.
|
||||||
|
- 음성 LLM 한 방에 처리하면 번역만 따로 교체하거나 추가학습시킬 수 없다.
|
||||||
|
- 용어집을 꽂아 넣을 지점이 사라진다.
|
||||||
|
|
||||||
|
그래서 **음성인식(ASR) → 번역(MT)** 2단 구조로 간다. 두 단계를 따로 고를 수 있어서
|
||||||
|
"인식은 최고급, 번역은 경량" 같은 조합이 가능하고, 이게 실제로 가장 가성비가 좋다.
|
||||||
|
|
||||||
|
## 음성인식
|
||||||
|
|
||||||
|
| 후보 | 판단 |
|
||||||
|
|---|---|
|
||||||
|
| **Whisper large-v3-turbo** | 디코더를 32층 → 4층으로 줄여 large-v3 대비 약 6배 빠르면서 정확도 손실은 1~2%. 809M 파라미터, int8이면 VRAM 1.6GB. **한/영/일/중 모두 지원.** |
|
||||||
|
| Whisper large-v3 | 가장 정확하지만 느리다. 최상위 티어에만. |
|
||||||
|
| distil-whisper | 빠르지만 **영어 전용**이라 탈락. |
|
||||||
|
| NVIDIA Parakeet | 실시간성은 최고지만 **영어 전용**이라 탈락. |
|
||||||
|
|
||||||
|
실행은 `faster-whisper`(CTranslate2)로 한다. 순정 `openai-whisper` 대비 4배 빠르고
|
||||||
|
VRAM을 절반만 쓴다.
|
||||||
|
|
||||||
|
→ **2~4티어의 기본은 large-v3-turbo.** 4개 언어를 모두 지원하면서 실시간을 만족하는
|
||||||
|
유일한 지점이다.
|
||||||
|
|
||||||
|
## 번역
|
||||||
|
|
||||||
|
| 후보 | VRAM | 판단 |
|
||||||
|
|---|---|---|
|
||||||
|
| **Seed-X-PPO-7B** (ByteDance) | ~5.6GB (Int4) | 28개 언어 **번역 전용**으로 학습된 7B. 자체 평가에서 Gemma3-27B, Llama4-Scout, Qwen3-235B를 앞서고 사람 평가에서 GPT-4o·Claude-3.5·Gemini-2.5-Pro와 대등. 한/영/일/중이 모두 주력 언어. |
|
||||||
|
| **Qwen3-8B** | ~11GB (fp16) | 범용 LLM이라 Seed-X보다 기본 번역은 약간 아래. 대신 **추가학습(LoRA) 생태계가 가장 두껍고**, 프롬프트 지시를 잘 따른다. |
|
||||||
|
| **NLLB-200-distilled 600M / 1.3B** | 0.8~2.9GB | 200개 언어 seq2seq. 구어체는 LLM보다 딱딱하지만 압도적으로 빠르고 가볍다. 게다가 **풀 파인튜닝이 소비자 GPU에서 된다.** |
|
||||||
|
| Gemma 3 12B | ~8GB | 140개 언어로 넓게 강하지만, 한/영/일/중만 필요한 우리에게는 Seed-X가 더 낫다. |
|
||||||
|
| opus-mt (Helsinki) | ~0.3GB | 언어쌍마다 모델이 따로라 4×3=12개를 관리해야 한다. 티어 구조와 안 맞아 탈락. |
|
||||||
|
|
||||||
|
## 다섯 티어
|
||||||
|
|
||||||
|
| # | 이름 | 음성인식 | 번역 | VRAM | 지연 |
|
||||||
|
|---|---|---|---|---|---|
|
||||||
|
| 1 | 번개 | whisper-small (int8) | NLLB-600M (int8) | ~2GB | 0.6초 |
|
||||||
|
| 2 | 신속 | large-v3-turbo (int8) | NLLB-600M (fp16) | ~4GB | 0.9초 |
|
||||||
|
| 3 | **균형** (기본) | large-v3-turbo (fp16) | NLLB-1.3B (fp16) | ~6GB | 1.2초 |
|
||||||
|
| 4 | **정밀** (추천) | large-v3-turbo (fp16) | Seed-X-PPO-7B (Int4) | ~10GB | 2.0초 |
|
||||||
|
| 5 | 극한 | large-v3 (fp16) | Qwen3-8B (fp16) + LoRA | ~16GB | 3.0초 |
|
||||||
|
|
||||||
|
## 결론: 무엇을 고를 것인가
|
||||||
|
|
||||||
|
**추가학습을 안 한다면 → 4티어 "정밀"이 최선이다.**
|
||||||
|
Seed-X는 번역만 하도록 만들어진 모델이라 7B치고 품질이 비정상적으로 좋다.
|
||||||
|
게임 대사처럼 짧고 구어체인 문장에서 NLLB 계열과 체감 차이가 크다.
|
||||||
|
|
||||||
|
**추가학습을 전제로 하면 → 5티어 "극한"(Qwen3-8B + LoRA)으로 간다.**
|
||||||
|
이유는 세 가지다.
|
||||||
|
|
||||||
|
1. Seed-X는 이미 PPO(강화학습)로 조율이 끝난 모델이다. 그 위에 다시 SFT를 얹으면
|
||||||
|
기존 품질이 깨지기 쉽다. 반면 Qwen3-8B는 instruct 베이스라 추가 학습을 전제로 만들어졌다.
|
||||||
|
2. peft / unsloth / TRL 등 LoRA 도구가 Qwen 계열에 가장 잘 맞춰져 있다.
|
||||||
|
3. 용어집을 프롬프트로 직접 지시할 수 있어 "학습 + 프롬프트" 이중으로 용어를 잡을 수 있다.
|
||||||
|
|
||||||
|
**VRAM이 8GB 이하라면 → 3티어 "균형".** 그리고 이 경우의 추가학습은
|
||||||
|
NLLB-1.3B를 **풀 파인튜닝**하는 쪽이 오히려 유리하다 (docs/FINETUNING.md 참고).
|
||||||
|
|
||||||
|
## 주의 — 벤치마크를 곧이곧대로 믿지 말 것
|
||||||
|
|
||||||
|
위 비교는 공개 벤치마크와 모델 카드에 근거한 것이고, **실제 게임/방송 음성에서의
|
||||||
|
체감 품질은 다를 수 있다.** 특히 Seed-X의 우위는 자체 발표 수치에 크게 기대고 있다.
|
||||||
|
그래서 프로그램에 5개 티어를 모두 넣었다. 실제 쓰는 콘텐츠로 3·4·5티어를 직접
|
||||||
|
번갈아 써보고 정하는 것이 가장 정확하다.
|
||||||
|
|
||||||
|
## 참고
|
||||||
|
|
||||||
|
- [Best open source STT model in 2026 (benchmarks)](https://northflank.com/blog/best-open-source-speech-to-text-stt-model-in-2026-benchmarks)
|
||||||
|
- [Whisper Large-v3 vs Turbo: Speed, WER & Cost](https://vexascribe.com/whisper-large-v3-vs-turbo)
|
||||||
|
- [Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters (arXiv)](https://arxiv.org/html/2507.13618v1)
|
||||||
|
- [ByteDance-Seed/Seed-X-PPO-7B (Hugging Face)](https://huggingface.co/ByteDance-Seed/Seed-X-PPO-7B)
|
||||||
|
- [Local translation benchmark 2026 (cctrans)](https://github.com/kargnas/cctrans/blob/main/docs/local-translation-benchmark-2026.md)
|
||||||
|
- [CTranslate2 지원 모델](https://opennmt.net/CTranslate2/guides/transformers.html)
|
||||||
BIN
docs/images/glossary.png
Normal file
BIN
docs/images/glossary.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 61 KiB |
BIN
docs/images/home.png
Normal file
BIN
docs/images/home.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 53 KiB |
BIN
docs/images/models.png
Normal file
BIN
docs/images/models.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 134 KiB |
BIN
docs/images/overlay.png
Normal file
BIN
docs/images/overlay.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 26 KiB |
BIN
docs/images/subtitle.png
Normal file
BIN
docs/images/subtitle.png
Normal file
Binary file not shown.
|
After Width: | Height: | Size: 67 KiB |
22
native/process_loopback/CMakeLists.txt
Normal file
22
native/process_loopback/CMakeLists.txt
Normal file
@@ -0,0 +1,22 @@
|
|||||||
|
cmake_minimum_required(VERSION 3.21)
|
||||||
|
project(hearo_capture LANGUAGES CXX)
|
||||||
|
|
||||||
|
set(CMAKE_CXX_STANDARD 17)
|
||||||
|
set(CMAKE_CXX_STANDARD_REQUIRED ON)
|
||||||
|
|
||||||
|
add_executable(hearo_capture main.cpp)
|
||||||
|
|
||||||
|
if(MSVC)
|
||||||
|
target_compile_options(hearo_capture PRIVATE /W3 /permissive- /EHsc)
|
||||||
|
target_compile_definitions(hearo_capture PRIVATE _WIN32_WINNT=0x0A00 NOMINMAX)
|
||||||
|
endif()
|
||||||
|
|
||||||
|
target_link_libraries(hearo_capture PRIVATE ole32 mmdevapi)
|
||||||
|
|
||||||
|
# 빌드 결과를 파이썬 패키지가 찾는 위치로 복사한다.
|
||||||
|
add_custom_command(TARGET hearo_capture POST_BUILD
|
||||||
|
COMMAND ${CMAKE_COMMAND} -E make_directory
|
||||||
|
"${CMAKE_SOURCE_DIR}/../../src/hearo/resources/bin"
|
||||||
|
COMMAND ${CMAKE_COMMAND} -E copy_if_different
|
||||||
|
"$<TARGET_FILE:hearo_capture>"
|
||||||
|
"${CMAKE_SOURCE_DIR}/../../src/hearo/resources/bin/")
|
||||||
21
native/process_loopback/build.ps1
Normal file
21
native/process_loopback/build.ps1
Normal file
@@ -0,0 +1,21 @@
|
|||||||
|
# 프로그램별 오디오 캡처 보조 프로그램 빌드 (Windows 전용)
|
||||||
|
#
|
||||||
|
# 필요: Visual Studio 2022 Build Tools (C++ 데스크톱 워크로드) + CMake
|
||||||
|
# winget install Microsoft.VisualStudio.2022.BuildTools
|
||||||
|
# winget install Kitware.CMake
|
||||||
|
#
|
||||||
|
# 사용: powershell -ExecutionPolicy Bypass -File native\process_loopback\build.ps1
|
||||||
|
|
||||||
|
$ErrorActionPreference = "Stop"
|
||||||
|
$here = Split-Path -Parent $MyInvocation.MyCommand.Path
|
||||||
|
$build = Join-Path $here "build"
|
||||||
|
|
||||||
|
cmake -S $here -B $build -A x64
|
||||||
|
cmake --build $build --config Release
|
||||||
|
|
||||||
|
$out = Join-Path $here "..\..\src\hearo\resources\bin\hearo_capture.exe"
|
||||||
|
if (Test-Path $out) {
|
||||||
|
Write-Host "빌드 완료: $out" -ForegroundColor Green
|
||||||
|
} else {
|
||||||
|
Write-Error "빌드는 끝났지만 결과물을 찾지 못했습니다."
|
||||||
|
}
|
||||||
229
native/process_loopback/main.cpp
Normal file
229
native/process_loopback/main.cpp
Normal file
@@ -0,0 +1,229 @@
|
|||||||
|
// hearo_capture.exe — 특정 프로세스(와 그 자식 프로세스)의 렌더 오디오만
|
||||||
|
// 뽑아내 stdout 으로 흘려보내는 보조 프로그램.
|
||||||
|
//
|
||||||
|
// 출력 포맷: 16000Hz / mono / 32-bit float, 헤더 없는 생 PCM.
|
||||||
|
// 사용법: hearo_capture.exe --pid 1234 [--exclude]
|
||||||
|
// --exclude 를 주면 해당 프로세스 트리를 "제외한" 나머지 소리를 받는다.
|
||||||
|
//
|
||||||
|
// Windows 10 2004(build 19041) 이상에서 동작하는 WASAPI process loopback API를
|
||||||
|
// 사용한다. Microsoft 의 ApplicationLoopback 샘플이 기반이다.
|
||||||
|
|
||||||
|
#define WIN32_LEAN_AND_MEAN
|
||||||
|
#include <windows.h>
|
||||||
|
|
||||||
|
#include <audioclient.h>
|
||||||
|
#include <audioclientactivationparams.h>
|
||||||
|
#include <mmdeviceapi.h>
|
||||||
|
#include <wrl/implements.h>
|
||||||
|
|
||||||
|
#include <atomic>
|
||||||
|
#include <cstdio>
|
||||||
|
#include <cstdlib>
|
||||||
|
#include <cstring>
|
||||||
|
#include <fcntl.h>
|
||||||
|
#include <io.h>
|
||||||
|
#include <vector>
|
||||||
|
|
||||||
|
#pragma comment(lib, "ole32.lib")
|
||||||
|
#pragma comment(lib, "mmdevapi.lib")
|
||||||
|
|
||||||
|
using Microsoft::WRL::ComPtr;
|
||||||
|
using Microsoft::WRL::RuntimeClass;
|
||||||
|
using Microsoft::WRL::RuntimeClassFlags;
|
||||||
|
using Microsoft::WRL::ClassicCom;
|
||||||
|
using Microsoft::WRL::FtmBase;
|
||||||
|
|
||||||
|
namespace {
|
||||||
|
|
||||||
|
constexpr int kSampleRate = 16000;
|
||||||
|
constexpr int kChannels = 1;
|
||||||
|
|
||||||
|
std::atomic<bool> g_stop{false};
|
||||||
|
|
||||||
|
BOOL WINAPI ConsoleHandler(DWORD type) {
|
||||||
|
if (type == CTRL_C_EVENT || type == CTRL_BREAK_EVENT || type == CTRL_CLOSE_EVENT) {
|
||||||
|
g_stop = true;
|
||||||
|
return TRUE;
|
||||||
|
}
|
||||||
|
return FALSE;
|
||||||
|
}
|
||||||
|
|
||||||
|
// ActivateAudioInterfaceAsync 완료 콜백. 결과를 이벤트로 넘겨준다.
|
||||||
|
class ActivationHandler
|
||||||
|
: public RuntimeClass<RuntimeClassFlags<ClassicCom | FtmBase>,
|
||||||
|
IActivateAudioInterfaceCompletionHandler> {
|
||||||
|
public:
|
||||||
|
explicit ActivationHandler(HANDLE done) : done_(done) {}
|
||||||
|
|
||||||
|
STDMETHODIMP ActivateCompleted(IActivateAudioInterfaceAsyncOperation* op) override {
|
||||||
|
HRESULT activate_hr = S_OK;
|
||||||
|
ComPtr<IUnknown> unknown;
|
||||||
|
HRESULT hr = op->GetActivateResult(&activate_hr, &unknown);
|
||||||
|
if (SUCCEEDED(hr) && SUCCEEDED(activate_hr)) {
|
||||||
|
unknown.As(&client_);
|
||||||
|
}
|
||||||
|
result_ = FAILED(hr) ? hr : activate_hr;
|
||||||
|
SetEvent(done_);
|
||||||
|
return S_OK;
|
||||||
|
}
|
||||||
|
|
||||||
|
HRESULT result() const { return result_; }
|
||||||
|
ComPtr<IAudioClient> client() const { return client_; }
|
||||||
|
|
||||||
|
private:
|
||||||
|
HANDLE done_;
|
||||||
|
HRESULT result_ = E_FAIL;
|
||||||
|
ComPtr<IAudioClient> client_;
|
||||||
|
};
|
||||||
|
|
||||||
|
void Fail(const char* what, HRESULT hr) {
|
||||||
|
std::fprintf(stderr, "[hearo_capture] %s (hr=0x%08lX)\n", what, static_cast<unsigned long>(hr));
|
||||||
|
}
|
||||||
|
|
||||||
|
int Run(DWORD pid, bool exclude) {
|
||||||
|
HANDLE activated = CreateEventW(nullptr, FALSE, FALSE, nullptr);
|
||||||
|
HANDLE sample_ready = CreateEventW(nullptr, FALSE, FALSE, nullptr);
|
||||||
|
if (!activated || !sample_ready) {
|
||||||
|
Fail("이벤트 생성 실패", HRESULT_FROM_WIN32(GetLastError()));
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
|
||||||
|
AUDIOCLIENT_ACTIVATION_PARAMS params{};
|
||||||
|
params.ActivationType = AUDIOCLIENT_ACTIVATION_TYPE_PROCESS_LOOPBACK;
|
||||||
|
params.ProcessLoopbackParams.TargetProcessId = pid;
|
||||||
|
params.ProcessLoopbackParams.ProcessLoopbackMode =
|
||||||
|
exclude ? PROCESS_LOOPBACK_MODE_EXCLUDE_TARGET_PROCESS_TREE
|
||||||
|
: PROCESS_LOOPBACK_MODE_INCLUDE_TARGET_PROCESS_TREE;
|
||||||
|
|
||||||
|
PROPVARIANT activate_params{};
|
||||||
|
activate_params.vt = VT_BLOB;
|
||||||
|
activate_params.blob.cbSize = sizeof(params);
|
||||||
|
activate_params.blob.pBlobData = reinterpret_cast<BYTE*>(¶ms);
|
||||||
|
|
||||||
|
auto handler = Microsoft::WRL::Make<ActivationHandler>(activated);
|
||||||
|
ComPtr<IActivateAudioInterfaceAsyncOperation> async_op;
|
||||||
|
HRESULT hr = ActivateAudioInterfaceAsync(VIRTUAL_AUDIO_DEVICE_PROCESS_LOOPBACK,
|
||||||
|
__uuidof(IAudioClient), &activate_params,
|
||||||
|
handler.Get(), &async_op);
|
||||||
|
if (FAILED(hr)) {
|
||||||
|
Fail("ActivateAudioInterfaceAsync 실패", hr);
|
||||||
|
return 2;
|
||||||
|
}
|
||||||
|
if (WaitForSingleObject(activated, 5000) != WAIT_OBJECT_0 || FAILED(handler->result())) {
|
||||||
|
Fail("오디오 인터페이스 활성화 실패 (Windows 10 2004 이상 필요)", handler->result());
|
||||||
|
return 3;
|
||||||
|
}
|
||||||
|
|
||||||
|
ComPtr<IAudioClient> client = handler->client();
|
||||||
|
if (!client) {
|
||||||
|
Fail("IAudioClient 를 받지 못했습니다", E_POINTER);
|
||||||
|
return 3;
|
||||||
|
}
|
||||||
|
|
||||||
|
// process loopback 경로에는 GetMixFormat 이 없으므로 포맷을 직접 지정한다.
|
||||||
|
WAVEFORMATEX format{};
|
||||||
|
format.wFormatTag = WAVE_FORMAT_IEEE_FLOAT;
|
||||||
|
format.nChannels = kChannels;
|
||||||
|
format.nSamplesPerSec = kSampleRate;
|
||||||
|
format.wBitsPerSample = 32;
|
||||||
|
format.nBlockAlign = format.nChannels * format.wBitsPerSample / 8;
|
||||||
|
format.nAvgBytesPerSec = format.nSamplesPerSec * format.nBlockAlign;
|
||||||
|
format.cbSize = 0;
|
||||||
|
|
||||||
|
constexpr REFERENCE_TIME kBufferDuration = 2 * 10'000'000LL; // 200ms
|
||||||
|
hr = client->Initialize(AUDCLNT_SHAREMODE_SHARED,
|
||||||
|
AUDCLNT_STREAMFLAGS_LOOPBACK | AUDCLNT_STREAMFLAGS_EVENTCALLBACK,
|
||||||
|
kBufferDuration, 0, &format, nullptr);
|
||||||
|
if (FAILED(hr)) {
|
||||||
|
Fail("IAudioClient::Initialize 실패", hr);
|
||||||
|
return 4;
|
||||||
|
}
|
||||||
|
|
||||||
|
ComPtr<IAudioCaptureClient> capture;
|
||||||
|
hr = client->GetService(IID_PPV_ARGS(&capture));
|
||||||
|
if (FAILED(hr)) {
|
||||||
|
Fail("IAudioCaptureClient 획득 실패", hr);
|
||||||
|
return 4;
|
||||||
|
}
|
||||||
|
hr = client->SetEventHandle(sample_ready);
|
||||||
|
if (FAILED(hr)) {
|
||||||
|
Fail("SetEventHandle 실패", hr);
|
||||||
|
return 4;
|
||||||
|
}
|
||||||
|
hr = client->Start();
|
||||||
|
if (FAILED(hr)) {
|
||||||
|
Fail("스트림 시작 실패", hr);
|
||||||
|
return 4;
|
||||||
|
}
|
||||||
|
|
||||||
|
std::fprintf(stderr, "[hearo_capture] pid=%lu 캡처 시작 (%dHz mono f32)\n",
|
||||||
|
static_cast<unsigned long>(pid), kSampleRate);
|
||||||
|
std::fflush(stderr);
|
||||||
|
|
||||||
|
while (!g_stop) {
|
||||||
|
// 무음 구간에서는 이벤트가 오지 않으므로 타임아웃으로 빠져나와 종료 여부를 확인한다.
|
||||||
|
DWORD wait = WaitForSingleObject(sample_ready, 200);
|
||||||
|
if (wait == WAIT_TIMEOUT) continue;
|
||||||
|
if (wait != WAIT_OBJECT_0) break;
|
||||||
|
|
||||||
|
UINT32 packet = 0;
|
||||||
|
while (SUCCEEDED(capture->GetNextPacketSize(&packet)) && packet > 0) {
|
||||||
|
BYTE* data = nullptr;
|
||||||
|
UINT32 frames = 0;
|
||||||
|
DWORD flags = 0;
|
||||||
|
hr = capture->GetBuffer(&data, &frames, &flags, nullptr, nullptr);
|
||||||
|
if (FAILED(hr)) break;
|
||||||
|
|
||||||
|
const size_t bytes = static_cast<size_t>(frames) * format.nBlockAlign;
|
||||||
|
if (flags & AUDCLNT_BUFFERFLAGS_SILENT) {
|
||||||
|
static thread_local std::vector<BYTE> zeros;
|
||||||
|
zeros.assign(bytes, 0);
|
||||||
|
std::fwrite(zeros.data(), 1, bytes, stdout);
|
||||||
|
} else if (data && bytes) {
|
||||||
|
std::fwrite(data, 1, bytes, stdout);
|
||||||
|
}
|
||||||
|
std::fflush(stdout);
|
||||||
|
capture->ReleaseBuffer(frames);
|
||||||
|
|
||||||
|
if (std::ferror(stdout)) { // 부모 프로세스가 파이프를 닫음
|
||||||
|
g_stop = true;
|
||||||
|
break;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
}
|
||||||
|
|
||||||
|
client->Stop();
|
||||||
|
CloseHandle(activated);
|
||||||
|
CloseHandle(sample_ready);
|
||||||
|
return 0;
|
||||||
|
}
|
||||||
|
|
||||||
|
} // namespace
|
||||||
|
|
||||||
|
int main(int argc, char** argv) {
|
||||||
|
DWORD pid = 0;
|
||||||
|
bool exclude = false;
|
||||||
|
for (int i = 1; i < argc; ++i) {
|
||||||
|
if (std::strcmp(argv[i], "--pid") == 0 && i + 1 < argc) {
|
||||||
|
pid = static_cast<DWORD>(std::strtoul(argv[++i], nullptr, 10));
|
||||||
|
} else if (std::strcmp(argv[i], "--exclude") == 0) {
|
||||||
|
exclude = true;
|
||||||
|
}
|
||||||
|
}
|
||||||
|
if (pid == 0) {
|
||||||
|
std::fprintf(stderr, "사용법: hearo_capture.exe --pid <PID> [--exclude]\n");
|
||||||
|
return 64;
|
||||||
|
}
|
||||||
|
|
||||||
|
_setmode(_fileno(stdout), _O_BINARY);
|
||||||
|
SetConsoleCtrlHandler(ConsoleHandler, TRUE);
|
||||||
|
|
||||||
|
HRESULT hr = CoInitializeEx(nullptr, COINIT_MULTITHREADED);
|
||||||
|
if (FAILED(hr)) {
|
||||||
|
Fail("CoInitializeEx 실패", hr);
|
||||||
|
return 1;
|
||||||
|
}
|
||||||
|
int rc = Run(pid, exclude);
|
||||||
|
CoUninitialize();
|
||||||
|
return rc;
|
||||||
|
}
|
||||||
66
pyproject.toml
Normal file
66
pyproject.toml
Normal file
@@ -0,0 +1,66 @@
|
|||||||
|
[build-system]
|
||||||
|
requires = ["setuptools>=68", "wheel"]
|
||||||
|
build-backend = "setuptools.build_meta"
|
||||||
|
|
||||||
|
[project]
|
||||||
|
name = "hearo"
|
||||||
|
version = "0.1.0"
|
||||||
|
description = "프로그램 소리를 실시간으로 듣고 번역해 자막으로 보여주는 도구"
|
||||||
|
readme = "README.md"
|
||||||
|
requires-python = ">=3.10"
|
||||||
|
license = { text = "MIT" }
|
||||||
|
authors = [{ name = "tkrmagid" }]
|
||||||
|
|
||||||
|
dependencies = [
|
||||||
|
"PySide6>=6.6",
|
||||||
|
"numpy>=1.24",
|
||||||
|
]
|
||||||
|
|
||||||
|
[project.optional-dependencies]
|
||||||
|
# GPU 추론 (실사용에 필요). torch 는 CUDA 빌드를 따로 설치해야 한다 — README 참고.
|
||||||
|
gpu = [
|
||||||
|
"faster-whisper>=1.0",
|
||||||
|
"ctranslate2>=4.4",
|
||||||
|
"transformers>=4.44",
|
||||||
|
"huggingface-hub>=0.24",
|
||||||
|
"sentencepiece>=0.2",
|
||||||
|
]
|
||||||
|
# Windows 오디오 캡처
|
||||||
|
windows = [
|
||||||
|
"PyAudioWPatch>=0.2.12.7",
|
||||||
|
"pycaw>=20240210",
|
||||||
|
"comtypes>=1.4",
|
||||||
|
"psutil>=5.9",
|
||||||
|
]
|
||||||
|
# 더 정확한 발화 구간 판정 (없어도 동작)
|
||||||
|
vad = ["webrtcvad-wheels>=2.0.14"]
|
||||||
|
# 추가학습 (LoRA)
|
||||||
|
finetune = [
|
||||||
|
"peft>=0.12",
|
||||||
|
"datasets>=2.20",
|
||||||
|
"accelerate>=0.33",
|
||||||
|
"bitsandbytes>=0.43; platform_system != 'Darwin'",
|
||||||
|
]
|
||||||
|
dev = ["pytest>=8.0", "ruff>=0.6"]
|
||||||
|
|
||||||
|
[project.scripts]
|
||||||
|
hearo = "hearo.app:main"
|
||||||
|
|
||||||
|
[tool.setuptools.packages.find]
|
||||||
|
where = ["src"]
|
||||||
|
|
||||||
|
[tool.setuptools.package-data]
|
||||||
|
hearo = ["resources/bin/*"]
|
||||||
|
|
||||||
|
[tool.pytest.ini_options]
|
||||||
|
testpaths = ["tests"]
|
||||||
|
pythonpath = ["src"]
|
||||||
|
|
||||||
|
[tool.ruff]
|
||||||
|
line-length = 100
|
||||||
|
target-version = "py310"
|
||||||
|
src = ["src", "tests"]
|
||||||
|
|
||||||
|
[tool.ruff.lint]
|
||||||
|
select = ["E", "F", "W", "I", "UP", "B", "SIM"]
|
||||||
|
ignore = ["E501", "B008"]
|
||||||
22
requirements.txt
Normal file
22
requirements.txt
Normal file
@@ -0,0 +1,22 @@
|
|||||||
|
# 기본 실행에 필요한 최소 구성
|
||||||
|
PySide6>=6.6
|
||||||
|
numpy>=1.24
|
||||||
|
|
||||||
|
# --- GPU 추론 ---------------------------------------------------------
|
||||||
|
# torch 는 반드시 CUDA 빌드로 먼저 설치하세요 (README 참고):
|
||||||
|
# pip install torch --index-url https://download.pytorch.org/whl/cu128
|
||||||
|
faster-whisper>=1.0
|
||||||
|
ctranslate2>=4.4
|
||||||
|
transformers>=4.44
|
||||||
|
huggingface-hub>=0.24
|
||||||
|
sentencepiece>=0.2
|
||||||
|
accelerate>=0.33
|
||||||
|
|
||||||
|
# --- Windows 오디오 캡처 ---------------------------------------------
|
||||||
|
PyAudioWPatch>=0.2.12.7; sys_platform == "win32"
|
||||||
|
pycaw>=20240210; sys_platform == "win32"
|
||||||
|
comtypes>=1.4; sys_platform == "win32"
|
||||||
|
psutil>=5.9; sys_platform == "win32"
|
||||||
|
|
||||||
|
# --- 선택 ------------------------------------------------------------
|
||||||
|
webrtcvad-wheels>=2.0.14
|
||||||
248
scripts/finetune_mt.py
Normal file
248
scripts/finetune_mt.py
Normal file
@@ -0,0 +1,248 @@
|
|||||||
|
#!/usr/bin/env python3
|
||||||
|
"""번역 모델 추가학습 (게임·방송 용어).
|
||||||
|
|
||||||
|
사용법:
|
||||||
|
python scripts/finetune_mt.py --data data/game.jsonl --tier ultimate \
|
||||||
|
--output ./adapters/game-ko --epochs 3
|
||||||
|
|
||||||
|
데이터 형식 (JSONL, 한 줄에 한 쌍):
|
||||||
|
{"source": "...", "target": "...", "source_lang": "en", "target_lang": "ko"}
|
||||||
|
|
||||||
|
자세한 배경은 docs/FINETUNING.md 참고.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import argparse
|
||||||
|
import json
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "src"))
|
||||||
|
|
||||||
|
from hearo.constants import LANGUAGES, models_dir # noqa: E402
|
||||||
|
from hearo.models.tiers import MTBackend, get_tier, ordered_tiers # noqa: E402
|
||||||
|
|
||||||
|
#: NLLB(seq2seq)와 Qwen(causal LM)에서 공통으로 존재하는 attention 투영 이름
|
||||||
|
LORA_TARGETS = ["q_proj", "v_proj"]
|
||||||
|
|
||||||
|
|
||||||
|
def load_pairs(path: Path) -> list[dict]:
|
||||||
|
pairs = []
|
||||||
|
with path.open(encoding="utf-8") as fh:
|
||||||
|
for lineno, raw in enumerate(fh, 1):
|
||||||
|
raw = raw.strip()
|
||||||
|
if not raw:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
item = json.loads(raw)
|
||||||
|
except json.JSONDecodeError as exc:
|
||||||
|
raise SystemExit(f"{path}:{lineno} JSON 오류: {exc}") from exc
|
||||||
|
for key in ("source", "target"):
|
||||||
|
if not item.get(key):
|
||||||
|
raise SystemExit(f"{path}:{lineno} '{key}' 가 비어 있습니다")
|
||||||
|
item.setdefault("source_lang", "en")
|
||||||
|
item.setdefault("target_lang", "ko")
|
||||||
|
for key in ("source_lang", "target_lang"):
|
||||||
|
if item[key] not in LANGUAGES:
|
||||||
|
raise SystemExit(f"{path}:{lineno} 지원하지 않는 언어: {item[key]}")
|
||||||
|
pairs.append(item)
|
||||||
|
if not pairs:
|
||||||
|
raise SystemExit(f"{path} 에 학습 데이터가 없습니다")
|
||||||
|
return pairs
|
||||||
|
|
||||||
|
|
||||||
|
def build_prompt(item: dict) -> tuple[str, str]:
|
||||||
|
"""LLM 학습용 (프롬프트, 정답) 쌍. translator.py 의 추론 프롬프트와 맞춰야 한다."""
|
||||||
|
src = LANGUAGES[item["source_lang"]]["english"]
|
||||||
|
tgt = LANGUAGES[item["target_lang"]]["english"]
|
||||||
|
prompt = (
|
||||||
|
f"Translate the following {src} text into {tgt}. "
|
||||||
|
"It is a live spoken line from a game or broadcast, so keep the tone "
|
||||||
|
"casual and natural. Output only the translation.\n\n"
|
||||||
|
f"{src}: {item['source']}\n{tgt}:"
|
||||||
|
)
|
||||||
|
return prompt, " " + item["target"]
|
||||||
|
|
||||||
|
|
||||||
|
def train_llm(pairs, repo, output: Path, epochs: int, lr: float, rank: int, batch: int):
|
||||||
|
import torch
|
||||||
|
from datasets import Dataset
|
||||||
|
from peft import LoraConfig, get_peft_model
|
||||||
|
from transformers import (
|
||||||
|
AutoModelForCausalLM,
|
||||||
|
AutoTokenizer,
|
||||||
|
DataCollatorForLanguageModeling,
|
||||||
|
Trainer,
|
||||||
|
TrainingArguments,
|
||||||
|
)
|
||||||
|
|
||||||
|
tokenizer = AutoTokenizer.from_pretrained(repo, cache_dir=str(models_dir()))
|
||||||
|
if tokenizer.pad_token is None:
|
||||||
|
tokenizer.pad_token = tokenizer.eos_token
|
||||||
|
|
||||||
|
def encode(item):
|
||||||
|
prompt, answer = build_prompt(item)
|
||||||
|
prompt_ids = tokenizer(prompt, add_special_tokens=False)["input_ids"]
|
||||||
|
answer_ids = tokenizer(answer, add_special_tokens=False)["input_ids"]
|
||||||
|
answer_ids.append(tokenizer.eos_token_id)
|
||||||
|
ids = (prompt_ids + answer_ids)[:512]
|
||||||
|
# 프롬프트 부분은 손실에서 제외해야 '번역 결과'만 학습된다.
|
||||||
|
labels = ([-100] * len(prompt_ids) + answer_ids)[:512]
|
||||||
|
return {"input_ids": ids, "labels": labels, "attention_mask": [1] * len(ids)}
|
||||||
|
|
||||||
|
dataset = Dataset.from_list([encode(p) for p in pairs])
|
||||||
|
|
||||||
|
model = AutoModelForCausalLM.from_pretrained(
|
||||||
|
repo,
|
||||||
|
cache_dir=str(models_dir()),
|
||||||
|
torch_dtype=torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float16,
|
||||||
|
device_map="auto",
|
||||||
|
)
|
||||||
|
model.enable_input_require_grads()
|
||||||
|
model = get_peft_model(
|
||||||
|
model,
|
||||||
|
LoraConfig(
|
||||||
|
r=rank,
|
||||||
|
lora_alpha=rank * 2,
|
||||||
|
lora_dropout=0.05,
|
||||||
|
bias="none",
|
||||||
|
task_type="CAUSAL_LM",
|
||||||
|
target_modules=LORA_TARGETS,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
model.print_trainable_parameters()
|
||||||
|
|
||||||
|
Trainer(
|
||||||
|
model=model,
|
||||||
|
args=TrainingArguments(
|
||||||
|
output_dir=str(output / "checkpoints"),
|
||||||
|
num_train_epochs=epochs,
|
||||||
|
per_device_train_batch_size=batch,
|
||||||
|
gradient_accumulation_steps=max(1, 8 // batch),
|
||||||
|
learning_rate=lr,
|
||||||
|
warmup_ratio=0.05,
|
||||||
|
logging_steps=20,
|
||||||
|
save_strategy="no",
|
||||||
|
gradient_checkpointing=True,
|
||||||
|
report_to=[],
|
||||||
|
),
|
||||||
|
train_dataset=dataset,
|
||||||
|
data_collator=DataCollatorForLanguageModeling(tokenizer, mlm=False),
|
||||||
|
).train()
|
||||||
|
|
||||||
|
model.save_pretrained(output)
|
||||||
|
tokenizer.save_pretrained(output)
|
||||||
|
|
||||||
|
|
||||||
|
def train_seq2seq(pairs, repo, output: Path, epochs: int, lr: float, rank: int, batch: int):
|
||||||
|
"""NLLB 계열. CTranslate2 변환 모델이 아니라 원본 HF 모델로 학습해야 한다."""
|
||||||
|
from datasets import Dataset
|
||||||
|
from peft import LoraConfig, get_peft_model
|
||||||
|
from transformers import (
|
||||||
|
AutoModelForSeq2SeqLM,
|
||||||
|
AutoTokenizer,
|
||||||
|
DataCollatorForSeq2Seq,
|
||||||
|
Seq2SeqTrainer,
|
||||||
|
Seq2SeqTrainingArguments,
|
||||||
|
)
|
||||||
|
|
||||||
|
tokenizer = AutoTokenizer.from_pretrained(repo, cache_dir=str(models_dir()))
|
||||||
|
|
||||||
|
def encode(item):
|
||||||
|
tokenizer.src_lang = LANGUAGES[item["source_lang"]]["nllb"]
|
||||||
|
tokenizer.tgt_lang = LANGUAGES[item["target_lang"]]["nllb"]
|
||||||
|
batch_enc = tokenizer(
|
||||||
|
item["source"], text_target=item["target"], truncation=True, max_length=256
|
||||||
|
)
|
||||||
|
return dict(batch_enc)
|
||||||
|
|
||||||
|
dataset = Dataset.from_list([encode(p) for p in pairs])
|
||||||
|
model = AutoModelForSeq2SeqLM.from_pretrained(repo, cache_dir=str(models_dir()))
|
||||||
|
model = get_peft_model(
|
||||||
|
model,
|
||||||
|
LoraConfig(
|
||||||
|
r=rank,
|
||||||
|
lora_alpha=rank * 2,
|
||||||
|
lora_dropout=0.05,
|
||||||
|
bias="none",
|
||||||
|
task_type="SEQ_2_SEQ_LM",
|
||||||
|
target_modules=LORA_TARGETS,
|
||||||
|
),
|
||||||
|
)
|
||||||
|
model.print_trainable_parameters()
|
||||||
|
|
||||||
|
Seq2SeqTrainer(
|
||||||
|
model=model,
|
||||||
|
args=Seq2SeqTrainingArguments(
|
||||||
|
output_dir=str(output / "checkpoints"),
|
||||||
|
num_train_epochs=epochs,
|
||||||
|
per_device_train_batch_size=batch,
|
||||||
|
learning_rate=lr,
|
||||||
|
warmup_ratio=0.05,
|
||||||
|
logging_steps=20,
|
||||||
|
save_strategy="no",
|
||||||
|
report_to=[],
|
||||||
|
),
|
||||||
|
train_dataset=dataset,
|
||||||
|
data_collator=DataCollatorForSeq2Seq(tokenizer, model=model),
|
||||||
|
).train()
|
||||||
|
|
||||||
|
model.save_pretrained(output)
|
||||||
|
tokenizer.save_pretrained(output)
|
||||||
|
|
||||||
|
|
||||||
|
def main() -> int:
|
||||||
|
parser = argparse.ArgumentParser(description="번역 모델 LoRA 추가학습")
|
||||||
|
parser.add_argument("--data", required=True, type=Path, help="JSONL 학습 데이터")
|
||||||
|
parser.add_argument(
|
||||||
|
"--tier",
|
||||||
|
default="ultimate",
|
||||||
|
choices=[t.key for t in ordered_tiers()],
|
||||||
|
help="어느 티어의 번역 모델을 학습할지 (기본: ultimate)",
|
||||||
|
)
|
||||||
|
parser.add_argument("--output", required=True, type=Path, help="어댑터 저장 폴더")
|
||||||
|
parser.add_argument("--epochs", type=int, default=3)
|
||||||
|
parser.add_argument("--lr", type=float, default=2e-4)
|
||||||
|
parser.add_argument("--rank", type=int, default=16)
|
||||||
|
parser.add_argument("--batch", type=int, default=2)
|
||||||
|
parser.add_argument(
|
||||||
|
"--base-model",
|
||||||
|
default="",
|
||||||
|
help="티어 기본값 대신 쓸 HF 모델 (NLLB는 CT2 변환본이 아닌 원본을 지정해야 함)",
|
||||||
|
)
|
||||||
|
args = parser.parse_args()
|
||||||
|
|
||||||
|
tier = get_tier(args.tier)
|
||||||
|
pairs = load_pairs(args.data)
|
||||||
|
args.output.mkdir(parents=True, exist_ok=True)
|
||||||
|
|
||||||
|
if tier.mt.finetune_ease >= 3:
|
||||||
|
print(
|
||||||
|
f"경고: '{tier.name}' 티어의 번역 모델은 추가학습에 적합하지 않습니다.\n"
|
||||||
|
" docs/FINETUNING.md 참고 — 'ultimate' 티어를 권장합니다.\n",
|
||||||
|
file=sys.stderr,
|
||||||
|
)
|
||||||
|
|
||||||
|
if tier.mt.backend is MTBackend.TRANSFORMERS:
|
||||||
|
repo = args.base_model or tier.mt.repo
|
||||||
|
print(f"LLM LoRA 학습: {repo} · {len(pairs)}쌍 · {args.epochs} epoch")
|
||||||
|
train_llm(pairs, repo, args.output, args.epochs, args.lr, args.rank, args.batch)
|
||||||
|
else:
|
||||||
|
# CT2 변환본에는 학습에 필요한 가중치가 없으므로 원본 리포로 바꾼다.
|
||||||
|
repo = args.base_model or _hf_original(tier.mt.repo)
|
||||||
|
print(f"seq2seq LoRA 학습: {repo} · {len(pairs)}쌍 · {args.epochs} epoch")
|
||||||
|
train_seq2seq(pairs, repo, args.output, args.epochs, args.lr, args.rank, args.batch)
|
||||||
|
|
||||||
|
print(f"\n완료. '모델' 화면의 추가학습 어댑터 칸에 입력하세요:\n {args.output.resolve()}")
|
||||||
|
return 0
|
||||||
|
|
||||||
|
|
||||||
|
def _hf_original(ct2_repo: str) -> str:
|
||||||
|
"""CTranslate2 변환 리포 이름에서 원본 facebook/ 리포를 추정한다."""
|
||||||
|
name = ct2_repo.split("/")[-1].replace("-ctranslate2", "")
|
||||||
|
return f"facebook/{name}"
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
6
src/hearo/__init__.py
Normal file
6
src/hearo/__init__.py
Normal file
@@ -0,0 +1,6 @@
|
|||||||
|
"""Hearo — 프로그램 소리를 실시간으로 듣고 번역해 자막으로 보여주는 도구."""
|
||||||
|
|
||||||
|
from .constants import APP_NAME, APP_SLOGAN, APP_VERSION
|
||||||
|
|
||||||
|
__all__ = ["APP_NAME", "APP_SLOGAN", "APP_VERSION", "__version__"]
|
||||||
|
__version__ = APP_VERSION
|
||||||
4
src/hearo/__main__.py
Normal file
4
src/hearo/__main__.py
Normal file
@@ -0,0 +1,4 @@
|
|||||||
|
from .app import main
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
53
src/hearo/app.py
Normal file
53
src/hearo/app.py
Normal file
@@ -0,0 +1,53 @@
|
|||||||
|
"""애플리케이션 진입점."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import contextlib
|
||||||
|
import logging
|
||||||
|
import sys
|
||||||
|
|
||||||
|
from .config import AppConfig
|
||||||
|
from .constants import APP_NAME, ORG_NAME, user_data_dir
|
||||||
|
|
||||||
|
|
||||||
|
def setup_logging(verbose: bool = False) -> None:
|
||||||
|
log_file = user_data_dir() / "hearo.log"
|
||||||
|
handlers: list[logging.Handler] = [logging.StreamHandler(sys.stderr)]
|
||||||
|
with contextlib.suppress(OSError):
|
||||||
|
handlers.append(logging.FileHandler(log_file, encoding="utf-8"))
|
||||||
|
logging.basicConfig(
|
||||||
|
level=logging.DEBUG if verbose else logging.INFO,
|
||||||
|
format="%(asctime)s %(levelname)-7s %(name)s: %(message)s",
|
||||||
|
handlers=handlers,
|
||||||
|
force=True,
|
||||||
|
)
|
||||||
|
# 모델 라이브러리들이 INFO로 너무 시끄럽다.
|
||||||
|
for noisy in ("urllib3", "filelock", "huggingface_hub", "transformers"):
|
||||||
|
logging.getLogger(noisy).setLevel(logging.WARNING)
|
||||||
|
|
||||||
|
|
||||||
|
def main(argv: list[str] | None = None) -> int:
|
||||||
|
argv = list(sys.argv if argv is None else argv)
|
||||||
|
verbose = "--verbose" in argv or "-v" in argv
|
||||||
|
setup_logging(verbose)
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt
|
||||||
|
from PySide6.QtWidgets import QApplication
|
||||||
|
|
||||||
|
from .ui import MainWindow
|
||||||
|
|
||||||
|
QApplication.setAttribute(Qt.ApplicationAttribute.AA_UseHighDpiPixmaps, True)
|
||||||
|
app = QApplication(argv)
|
||||||
|
app.setApplicationName(APP_NAME)
|
||||||
|
app.setOrganizationName(ORG_NAME)
|
||||||
|
app.setQuitOnLastWindowClosed(True)
|
||||||
|
|
||||||
|
config = AppConfig.load()
|
||||||
|
window = MainWindow(config)
|
||||||
|
if not config.start_minimized:
|
||||||
|
window.show()
|
||||||
|
return app.exec()
|
||||||
|
|
||||||
|
|
||||||
|
if __name__ == "__main__":
|
||||||
|
raise SystemExit(main())
|
||||||
72
src/hearo/audio/__init__.py
Normal file
72
src/hearo/audio/__init__.py
Normal file
@@ -0,0 +1,72 @@
|
|||||||
|
"""오디오 캡처 레이어."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
|
||||||
|
from ..config import AudioConfig
|
||||||
|
from .base import AudioSource, CaptureBackend, CaptureError
|
||||||
|
from .file_source import FileCapture
|
||||||
|
from .process_loopback import ProcessLoopbackCapture
|
||||||
|
from .segmenter import Segment, Segmenter, SegmenterConfig
|
||||||
|
from .wasapi_loopback import WasapiLoopbackCapture
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
__all__ = [
|
||||||
|
"AudioSource",
|
||||||
|
"CaptureBackend",
|
||||||
|
"CaptureError",
|
||||||
|
"FileCapture",
|
||||||
|
"ProcessLoopbackCapture",
|
||||||
|
"Segment",
|
||||||
|
"Segmenter",
|
||||||
|
"SegmenterConfig",
|
||||||
|
"WasapiLoopbackCapture",
|
||||||
|
"capture_capabilities",
|
||||||
|
"create_capture",
|
||||||
|
"list_sources",
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
def capture_capabilities() -> dict[str, bool]:
|
||||||
|
"""현재 환경에서 어떤 캡처가 가능한지."""
|
||||||
|
return {
|
||||||
|
"process": ProcessLoopbackCapture.available(),
|
||||||
|
"device": WasapiLoopbackCapture.available(),
|
||||||
|
"file": True,
|
||||||
|
}
|
||||||
|
|
||||||
|
|
||||||
|
def list_sources() -> list[AudioSource]:
|
||||||
|
"""프로그램 + 출력 장치를 한 목록으로."""
|
||||||
|
sources: list[AudioSource] = []
|
||||||
|
if ProcessLoopbackCapture.available():
|
||||||
|
sources.extend(ProcessLoopbackCapture.list_sources())
|
||||||
|
sources.extend(WasapiLoopbackCapture.list_sources())
|
||||||
|
return sources
|
||||||
|
|
||||||
|
|
||||||
|
def create_capture(config: AudioConfig) -> CaptureBackend:
|
||||||
|
"""설정에 맞는 캡처 백엔드를 만든다. 불가능하면 장치 루프백으로 폴백."""
|
||||||
|
backend = config.backend
|
||||||
|
if backend == "file":
|
||||||
|
return FileCapture(config.target_process_name)
|
||||||
|
|
||||||
|
want_process = backend in ("process", "auto") and config.target_pid > 0
|
||||||
|
if want_process:
|
||||||
|
if ProcessLoopbackCapture.available():
|
||||||
|
return ProcessLoopbackCapture(config.target_pid)
|
||||||
|
if backend == "process":
|
||||||
|
raise CaptureError(
|
||||||
|
"프로그램별 캡처를 쓰려면 hearo_capture.exe 가 필요합니다. "
|
||||||
|
"native/process_loopback/build.ps1 로 빌드하거나 '출력 장치 전체'를 선택하세요."
|
||||||
|
)
|
||||||
|
log.warning("프로그램별 캡처를 쓸 수 없어 장치 루프백으로 전환합니다.")
|
||||||
|
|
||||||
|
if not WasapiLoopbackCapture.available():
|
||||||
|
raise CaptureError(
|
||||||
|
"이 환경에서는 오디오 캡처를 할 수 없습니다. "
|
||||||
|
"Windows에서 PyAudioWPatch 설치 여부를 확인하세요."
|
||||||
|
)
|
||||||
|
return WasapiLoopbackCapture(config.device_index)
|
||||||
146
src/hearo/audio/base.py
Normal file
146
src/hearo/audio/base.py
Normal file
@@ -0,0 +1,146 @@
|
|||||||
|
"""오디오 캡처 백엔드 공통 인터페이스.
|
||||||
|
|
||||||
|
모든 백엔드는 16kHz / mono / float32(-1.0~1.0) 프레임을 큐로 흘려보낸다.
|
||||||
|
리샘플링과 다운믹스는 각 백엔드가 책임진다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import abc
|
||||||
|
import queue
|
||||||
|
import threading
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..constants import SAMPLE_RATE
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class AudioSource:
|
||||||
|
"""UI에 노출되는 캡처 대상 한 줄."""
|
||||||
|
|
||||||
|
kind: str # "process" | "device"
|
||||||
|
identifier: str # pid 문자열 또는 장치 인덱스
|
||||||
|
label: str
|
||||||
|
detail: str = ""
|
||||||
|
icon_path: str = ""
|
||||||
|
|
||||||
|
@property
|
||||||
|
def pid(self) -> int:
|
||||||
|
return int(self.identifier) if self.kind == "process" else 0
|
||||||
|
|
||||||
|
@property
|
||||||
|
def device_index(self) -> int:
|
||||||
|
return int(self.identifier) if self.kind == "device" else -1
|
||||||
|
|
||||||
|
|
||||||
|
class CaptureError(RuntimeError):
|
||||||
|
"""캡처를 시작할 수 없을 때."""
|
||||||
|
|
||||||
|
|
||||||
|
class CaptureBackend(abc.ABC):
|
||||||
|
"""오디오 캡처 백엔드 베이스."""
|
||||||
|
|
||||||
|
name = "base"
|
||||||
|
#: 특정 프로그램 소리만 분리해서 받을 수 있는가
|
||||||
|
per_process = False
|
||||||
|
|
||||||
|
def __init__(self, max_queue_frames: int = 400) -> None:
|
||||||
|
self._queue: queue.Queue[np.ndarray] = queue.Queue(maxsize=max_queue_frames)
|
||||||
|
self._stop = threading.Event()
|
||||||
|
self._thread: threading.Thread | None = None
|
||||||
|
self._error: Exception | None = None
|
||||||
|
|
||||||
|
# --- 하위 클래스가 구현 -------------------------------------------
|
||||||
|
@abc.abstractmethod
|
||||||
|
def _run(self) -> None:
|
||||||
|
"""블로킹 캡처 루프. self._stop 이 set 될 때까지 _emit() 호출."""
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
@abc.abstractmethod
|
||||||
|
def available() -> bool:
|
||||||
|
"""현재 환경에서 이 백엔드를 쓸 수 있는가."""
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
@abc.abstractmethod
|
||||||
|
def list_sources() -> list[AudioSource]:
|
||||||
|
"""선택 가능한 캡처 대상 목록."""
|
||||||
|
|
||||||
|
# --- 공통 동작 -----------------------------------------------------
|
||||||
|
def start(self) -> None:
|
||||||
|
if self._thread and self._thread.is_alive():
|
||||||
|
return
|
||||||
|
self._stop.clear()
|
||||||
|
self._error = None
|
||||||
|
self._thread = threading.Thread(
|
||||||
|
target=self._thread_main, name=f"capture-{self.name}", daemon=True
|
||||||
|
)
|
||||||
|
self._thread.start()
|
||||||
|
|
||||||
|
def _thread_main(self) -> None:
|
||||||
|
try:
|
||||||
|
self._run()
|
||||||
|
except Exception as exc: # noqa: BLE001 - 워커 스레드 경계
|
||||||
|
self._error = exc
|
||||||
|
|
||||||
|
def stop(self, timeout: float = 2.0) -> None:
|
||||||
|
self._stop.set()
|
||||||
|
if self._thread:
|
||||||
|
self._thread.join(timeout=timeout)
|
||||||
|
self._thread = None
|
||||||
|
|
||||||
|
@property
|
||||||
|
def running(self) -> bool:
|
||||||
|
return bool(self._thread and self._thread.is_alive())
|
||||||
|
|
||||||
|
@property
|
||||||
|
def error(self) -> Exception | None:
|
||||||
|
return self._error
|
||||||
|
|
||||||
|
def _emit(self, frame: np.ndarray) -> None:
|
||||||
|
"""캡처 프레임 투입. 큐가 가득 차면 가장 오래된 프레임을 버린다.
|
||||||
|
|
||||||
|
실시간 자막에서는 밀린 오디오보다 최신 오디오가 항상 더 가치 있다.
|
||||||
|
"""
|
||||||
|
try:
|
||||||
|
self._queue.put_nowait(frame)
|
||||||
|
except queue.Full:
|
||||||
|
try:
|
||||||
|
self._queue.get_nowait()
|
||||||
|
self._queue.put_nowait(frame)
|
||||||
|
except (queue.Empty, queue.Full):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def read(self, timeout: float = 0.5) -> np.ndarray | None:
|
||||||
|
try:
|
||||||
|
return self._queue.get(timeout=timeout)
|
||||||
|
except queue.Empty:
|
||||||
|
return None
|
||||||
|
|
||||||
|
def drain(self) -> None:
|
||||||
|
while not self._queue.empty():
|
||||||
|
try:
|
||||||
|
self._queue.get_nowait()
|
||||||
|
except queue.Empty:
|
||||||
|
break
|
||||||
|
|
||||||
|
|
||||||
|
def to_mono_16k(data: np.ndarray, channels: int, src_rate: int) -> np.ndarray:
|
||||||
|
"""인터리브된 float32 PCM을 16kHz 모노로 변환."""
|
||||||
|
if data.size == 0:
|
||||||
|
return data.astype(np.float32, copy=False)
|
||||||
|
audio = data.astype(np.float32, copy=False)
|
||||||
|
if channels > 1:
|
||||||
|
usable = (audio.size // channels) * channels
|
||||||
|
audio = audio[:usable].reshape(-1, channels).mean(axis=1)
|
||||||
|
if src_rate != SAMPLE_RATE and audio.size:
|
||||||
|
# 선형 보간 리샘플. 음성인식 입력으로는 충분한 품질이며 의존성이 없다.
|
||||||
|
duration = audio.size / src_rate
|
||||||
|
target_len = max(1, int(duration * SAMPLE_RATE))
|
||||||
|
audio = np.interp(
|
||||||
|
np.linspace(0.0, audio.size - 1, target_len, dtype=np.float64),
|
||||||
|
np.arange(audio.size, dtype=np.float64),
|
||||||
|
audio,
|
||||||
|
).astype(np.float32)
|
||||||
|
return audio
|
||||||
66
src/hearo/audio/file_source.py
Normal file
66
src/hearo/audio/file_source.py
Normal file
@@ -0,0 +1,66 @@
|
|||||||
|
"""파일/합성 오디오 캡처 백엔드.
|
||||||
|
|
||||||
|
실제 오디오 장치 없이 파이프라인 전체를 돌려볼 수 있게 해준다.
|
||||||
|
개발·테스트·데모용이며 어느 OS에서나 동작한다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import time
|
||||||
|
import wave
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..constants import FRAME_SAMPLES, SAMPLE_RATE
|
||||||
|
from .base import AudioSource, CaptureBackend, CaptureError, to_mono_16k
|
||||||
|
|
||||||
|
|
||||||
|
class FileCapture(CaptureBackend):
|
||||||
|
"""WAV 파일을 실시간 속도로 재생하듯 흘려보낸다."""
|
||||||
|
|
||||||
|
name = "file"
|
||||||
|
per_process = False
|
||||||
|
|
||||||
|
def __init__(self, path: str | Path, loop: bool = False, realtime: bool = True) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.path = Path(path)
|
||||||
|
self.loop = loop
|
||||||
|
self.realtime = realtime
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def available() -> bool:
|
||||||
|
return True
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def list_sources() -> list[AudioSource]:
|
||||||
|
return []
|
||||||
|
|
||||||
|
def _read_wav(self) -> np.ndarray:
|
||||||
|
if not self.path.is_file():
|
||||||
|
raise CaptureError(f"오디오 파일을 찾을 수 없습니다: {self.path}")
|
||||||
|
with wave.open(str(self.path), "rb") as wf:
|
||||||
|
channels = wf.getnchannels()
|
||||||
|
rate = wf.getframerate()
|
||||||
|
width = wf.getsampwidth()
|
||||||
|
raw = wf.readframes(wf.getnframes())
|
||||||
|
if width == 2:
|
||||||
|
data = np.frombuffer(raw, dtype=np.int16).astype(np.float32) / 32768.0
|
||||||
|
elif width == 4:
|
||||||
|
data = np.frombuffer(raw, dtype=np.float32)
|
||||||
|
else:
|
||||||
|
raise CaptureError(f"지원하지 않는 샘플 폭입니다: {width * 8}bit")
|
||||||
|
return to_mono_16k(data, channels, rate)
|
||||||
|
|
||||||
|
def _run(self) -> None:
|
||||||
|
audio = self._read_wav()
|
||||||
|
chunk = FRAME_SAMPLES * 5 # 100ms
|
||||||
|
while not self._stop.is_set():
|
||||||
|
for start in range(0, audio.size, chunk):
|
||||||
|
if self._stop.is_set():
|
||||||
|
return
|
||||||
|
self._emit(audio[start : start + chunk].copy())
|
||||||
|
if self.realtime:
|
||||||
|
time.sleep(chunk / SAMPLE_RATE)
|
||||||
|
if not self.loop:
|
||||||
|
return
|
||||||
168
src/hearo/audio/process_loopback.py
Normal file
168
src/hearo/audio/process_loopback.py
Normal file
@@ -0,0 +1,168 @@
|
|||||||
|
"""프로그램별 오디오 캡처 (Windows WASAPI process loopback).
|
||||||
|
|
||||||
|
`native/process_loopback` 에서 빌드한 `hearo_capture.exe` 를 자식 프로세스로 띄우고
|
||||||
|
stdout 으로 흘러나오는 16kHz mono float32 PCM 을 그대로 읽는다.
|
||||||
|
|
||||||
|
보조 프로그램이 없으면 `available()` 이 False 를 돌려주고, 상위 레이어는
|
||||||
|
장치 루프백으로 자동 폴백한다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import subprocess
|
||||||
|
import sys
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..constants import FRAME_SAMPLES
|
||||||
|
from .base import AudioSource, CaptureBackend, CaptureError
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
HELPER_NAME = "hearo_capture.exe"
|
||||||
|
READ_BYTES = FRAME_SAMPLES * 4 * 5 # 100ms 분량
|
||||||
|
|
||||||
|
|
||||||
|
def helper_path() -> Path | None:
|
||||||
|
"""보조 프로그램 위치. 패키지 내부 → 저장소 빌드 폴더 순으로 찾는다."""
|
||||||
|
candidates = [
|
||||||
|
Path(__file__).resolve().parent.parent / "resources" / "bin" / HELPER_NAME,
|
||||||
|
Path(__file__).resolve().parents[3]
|
||||||
|
/ "native"
|
||||||
|
/ "process_loopback"
|
||||||
|
/ "build"
|
||||||
|
/ "Release"
|
||||||
|
/ HELPER_NAME,
|
||||||
|
]
|
||||||
|
for path in candidates:
|
||||||
|
if path.is_file():
|
||||||
|
return path
|
||||||
|
return None
|
||||||
|
|
||||||
|
|
||||||
|
class ProcessLoopbackCapture(CaptureBackend):
|
||||||
|
name = "process"
|
||||||
|
per_process = True
|
||||||
|
|
||||||
|
def __init__(self, pid: int, exclude: bool = False) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.pid = pid
|
||||||
|
self.exclude = exclude
|
||||||
|
self._proc: subprocess.Popen[bytes] | None = None
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def available() -> bool:
|
||||||
|
return sys.platform == "win32" and helper_path() is not None
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def list_sources() -> list[AudioSource]:
|
||||||
|
"""현재 소리를 내고 있는(또는 낼 수 있는) 프로세스 목록."""
|
||||||
|
if sys.platform != "win32":
|
||||||
|
return []
|
||||||
|
try:
|
||||||
|
from pycaw.pycaw import AudioUtilities
|
||||||
|
except ImportError:
|
||||||
|
log.debug("pycaw 미설치 — 프로세스 목록을 만들 수 없습니다.")
|
||||||
|
return []
|
||||||
|
|
||||||
|
seen: dict[int, AudioSource] = {}
|
||||||
|
for session in AudioUtilities.GetAllSessions():
|
||||||
|
proc = session.Process
|
||||||
|
if proc is None or proc.pid in seen:
|
||||||
|
continue
|
||||||
|
try:
|
||||||
|
name = proc.name()
|
||||||
|
except Exception: # noqa: BLE001 - 종료된 프로세스
|
||||||
|
continue
|
||||||
|
title = _window_title(proc.pid) or name
|
||||||
|
seen[proc.pid] = AudioSource(
|
||||||
|
kind="process",
|
||||||
|
identifier=str(proc.pid),
|
||||||
|
label=title,
|
||||||
|
detail=f"{name} (PID {proc.pid})",
|
||||||
|
)
|
||||||
|
return sorted(seen.values(), key=lambda s: s.label.lower())
|
||||||
|
|
||||||
|
def _run(self) -> None: # pragma: no cover - Windows 전용
|
||||||
|
exe = helper_path()
|
||||||
|
if exe is None:
|
||||||
|
raise CaptureError(
|
||||||
|
f"{HELPER_NAME} 를 찾지 못했습니다. "
|
||||||
|
"native/process_loopback/build.ps1 로 먼저 빌드하세요."
|
||||||
|
)
|
||||||
|
cmd = [str(exe), "--pid", str(self.pid)]
|
||||||
|
if self.exclude:
|
||||||
|
cmd.append("--exclude")
|
||||||
|
|
||||||
|
self._proc = subprocess.Popen(
|
||||||
|
cmd,
|
||||||
|
stdout=subprocess.PIPE,
|
||||||
|
stderr=subprocess.PIPE,
|
||||||
|
bufsize=0,
|
||||||
|
creationflags=getattr(subprocess, "CREATE_NO_WINDOW", 0),
|
||||||
|
)
|
||||||
|
assert self._proc.stdout is not None
|
||||||
|
try:
|
||||||
|
while not self._stop.is_set():
|
||||||
|
raw = self._proc.stdout.read(READ_BYTES)
|
||||||
|
if not raw:
|
||||||
|
code = self._proc.poll()
|
||||||
|
err = b""
|
||||||
|
if self._proc.stderr is not None:
|
||||||
|
err = self._proc.stderr.read() or b""
|
||||||
|
raise CaptureError(
|
||||||
|
f"캡처 보조 프로그램이 종료되었습니다(코드 {code}): "
|
||||||
|
f"{err.decode('utf-8', 'replace').strip()}"
|
||||||
|
)
|
||||||
|
# 4바이트 정렬이 깨지면 float32 해석이 밀린다.
|
||||||
|
usable = len(raw) - (len(raw) % 4)
|
||||||
|
self._emit(np.frombuffer(raw[:usable], dtype=np.float32).copy())
|
||||||
|
finally:
|
||||||
|
self._terminate()
|
||||||
|
|
||||||
|
def _terminate(self) -> None:
|
||||||
|
proc, self._proc = self._proc, None
|
||||||
|
if proc is None:
|
||||||
|
return
|
||||||
|
if proc.poll() is None:
|
||||||
|
proc.terminate()
|
||||||
|
try:
|
||||||
|
proc.wait(timeout=2)
|
||||||
|
except subprocess.TimeoutExpired:
|
||||||
|
proc.kill()
|
||||||
|
for pipe in (proc.stdout, proc.stderr):
|
||||||
|
if pipe is not None:
|
||||||
|
pipe.close()
|
||||||
|
|
||||||
|
|
||||||
|
def _window_title(pid: int) -> str:
|
||||||
|
"""해당 PID의 최상위 창 제목. 실패하면 빈 문자열."""
|
||||||
|
if sys.platform != "win32":
|
||||||
|
return ""
|
||||||
|
try:
|
||||||
|
import ctypes
|
||||||
|
from ctypes import wintypes
|
||||||
|
|
||||||
|
user32 = ctypes.windll.user32
|
||||||
|
result: list[str] = []
|
||||||
|
|
||||||
|
@ctypes.WINFUNCTYPE(wintypes.BOOL, wintypes.HWND, wintypes.LPARAM)
|
||||||
|
def callback(hwnd, _lparam):
|
||||||
|
owner = wintypes.DWORD()
|
||||||
|
user32.GetWindowThreadProcessId(hwnd, ctypes.byref(owner))
|
||||||
|
if owner.value != pid or not user32.IsWindowVisible(hwnd):
|
||||||
|
return True
|
||||||
|
length = user32.GetWindowTextLengthW(hwnd)
|
||||||
|
if length <= 0:
|
||||||
|
return True
|
||||||
|
buf = ctypes.create_unicode_buffer(length + 1)
|
||||||
|
user32.GetWindowTextW(hwnd, buf, length + 1)
|
||||||
|
result.append(buf.value)
|
||||||
|
return False
|
||||||
|
|
||||||
|
user32.EnumWindows(callback, 0)
|
||||||
|
return result[0] if result else ""
|
||||||
|
except Exception: # noqa: BLE001 - 창 제목은 부가 정보일 뿐
|
||||||
|
return ""
|
||||||
188
src/hearo/audio/segmenter.py
Normal file
188
src/hearo/audio/segmenter.py
Normal file
@@ -0,0 +1,188 @@
|
|||||||
|
"""발화 구간 분할기.
|
||||||
|
|
||||||
|
캡처된 연속 오디오를 "한 문장" 단위로 잘라 음성인식에 넘긴다.
|
||||||
|
외부 의존성 없이 동작하도록 적응형 노이즈 플로어 기반 에너지 VAD를 쓴다.
|
||||||
|
webrtcvad 가 설치돼 있으면 더 정확한 판정을 위해 함께 사용한다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import time
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..constants import FRAME_MS, FRAME_SAMPLES, SAMPLE_RATE
|
||||||
|
|
||||||
|
try: # pragma: no cover - 선택 의존성
|
||||||
|
import webrtcvad
|
||||||
|
except ImportError: # pragma: no cover
|
||||||
|
webrtcvad = None # type: ignore[assignment]
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class Segment:
|
||||||
|
"""음성인식에 넘길 오디오 한 덩어리."""
|
||||||
|
|
||||||
|
audio: np.ndarray
|
||||||
|
started_at: float
|
||||||
|
ended_at: float
|
||||||
|
is_final: bool = True
|
||||||
|
|
||||||
|
@property
|
||||||
|
def duration_s(self) -> float:
|
||||||
|
return self.audio.size / SAMPLE_RATE
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SegmenterConfig:
|
||||||
|
silence_ms: int = 600
|
||||||
|
min_segment_ms: int = 400
|
||||||
|
max_segment_ms: int = 12_000
|
||||||
|
partial_interval_ms: int = 900
|
||||||
|
#: 노이즈 플로어 대비 몇 배 이상이면 발화로 볼지
|
||||||
|
speech_ratio: float = 3.0
|
||||||
|
#: 절대 무음 판정 하한 (RMS)
|
||||||
|
absolute_floor: float = 2e-4
|
||||||
|
vad_aggressiveness: int = 2
|
||||||
|
#: 발화 시작 앞쪽에 붙일 여유분 (앞 음절 잘림 방지)
|
||||||
|
lead_in_ms: int = 240
|
||||||
|
|
||||||
|
|
||||||
|
class Segmenter:
|
||||||
|
"""프레임을 받아 발화 구간을 뱉는 상태 기계.
|
||||||
|
|
||||||
|
`push()` 는 확정 구간(is_final=True) 또는 중간 결과(is_final=False)를
|
||||||
|
돌려주거나, 아직 내보낼 게 없으면 None 을 돌려준다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
def __init__(self, config: SegmenterConfig | None = None) -> None:
|
||||||
|
self.config = config or SegmenterConfig()
|
||||||
|
self._vad = None
|
||||||
|
if webrtcvad is not None:
|
||||||
|
self._vad = webrtcvad.Vad(max(0, min(3, self.config.vad_aggressiveness)))
|
||||||
|
|
||||||
|
self._buffer: list[np.ndarray] = []
|
||||||
|
self._lead_in: list[np.ndarray] = []
|
||||||
|
self._pending = np.zeros(0, dtype=np.float32)
|
||||||
|
self._noise_floor = 1e-3
|
||||||
|
self._silence_ms = 0
|
||||||
|
self._speech_ms = 0
|
||||||
|
self._in_speech = False
|
||||||
|
self._started_at = 0.0
|
||||||
|
self._last_partial_ms = 0
|
||||||
|
self._lead_in_frames = max(1, self.config.lead_in_ms // FRAME_MS)
|
||||||
|
|
||||||
|
# --- 입력 ----------------------------------------------------------
|
||||||
|
def push(self, chunk: np.ndarray) -> list[Segment]:
|
||||||
|
"""임의 길이의 오디오를 넣고 완성된 구간 목록을 받는다."""
|
||||||
|
out: list[Segment] = []
|
||||||
|
if chunk.size:
|
||||||
|
self._pending = (
|
||||||
|
chunk.astype(np.float32, copy=False)
|
||||||
|
if self._pending.size == 0
|
||||||
|
else np.concatenate((self._pending, chunk))
|
||||||
|
)
|
||||||
|
while self._pending.size >= FRAME_SAMPLES:
|
||||||
|
frame = self._pending[:FRAME_SAMPLES]
|
||||||
|
self._pending = self._pending[FRAME_SAMPLES:]
|
||||||
|
segment = self._push_frame(frame)
|
||||||
|
if segment is not None:
|
||||||
|
out.append(segment)
|
||||||
|
return out
|
||||||
|
|
||||||
|
def flush(self) -> Segment | None:
|
||||||
|
"""남은 버퍼를 강제로 확정 구간으로 만든다 (캡처 종료 시)."""
|
||||||
|
if not self._in_speech:
|
||||||
|
return None
|
||||||
|
return self._finish()
|
||||||
|
|
||||||
|
def reset(self) -> None:
|
||||||
|
"""발화 상태만 초기화한다.
|
||||||
|
|
||||||
|
`_pending` 은 건드리지 않는다. 아직 프레임으로 쪼개지 못한 '입력' 이라
|
||||||
|
여기서 버리면 한 구간을 확정한 직후의 오디오가 통째로 사라진다.
|
||||||
|
"""
|
||||||
|
self._buffer.clear()
|
||||||
|
self._lead_in.clear()
|
||||||
|
self._silence_ms = 0
|
||||||
|
self._speech_ms = 0
|
||||||
|
self._in_speech = False
|
||||||
|
self._last_partial_ms = 0
|
||||||
|
|
||||||
|
# --- 내부 ----------------------------------------------------------
|
||||||
|
def _push_frame(self, frame: np.ndarray) -> Segment | None:
|
||||||
|
speech = self._is_speech(frame)
|
||||||
|
cfg = self.config
|
||||||
|
|
||||||
|
if not self._in_speech:
|
||||||
|
self._lead_in.append(frame)
|
||||||
|
if len(self._lead_in) > self._lead_in_frames:
|
||||||
|
self._lead_in.pop(0)
|
||||||
|
if speech:
|
||||||
|
self._in_speech = True
|
||||||
|
self._started_at = time.monotonic()
|
||||||
|
self._buffer = list(self._lead_in)
|
||||||
|
self._lead_in = []
|
||||||
|
self._silence_ms = 0
|
||||||
|
self._speech_ms = FRAME_MS
|
||||||
|
self._last_partial_ms = 0
|
||||||
|
return None
|
||||||
|
|
||||||
|
self._buffer.append(frame)
|
||||||
|
self._speech_ms += FRAME_MS
|
||||||
|
self._silence_ms = 0 if speech else self._silence_ms + FRAME_MS
|
||||||
|
|
||||||
|
buffered_ms = len(self._buffer) * FRAME_MS
|
||||||
|
if self._silence_ms >= cfg.silence_ms:
|
||||||
|
if buffered_ms - self._silence_ms < cfg.min_segment_ms:
|
||||||
|
self.reset() # 잡음 한 번 튄 것 — 버린다
|
||||||
|
return None
|
||||||
|
return self._finish()
|
||||||
|
|
||||||
|
if buffered_ms >= cfg.max_segment_ms:
|
||||||
|
return self._finish()
|
||||||
|
|
||||||
|
if cfg.partial_interval_ms > 0:
|
||||||
|
since = buffered_ms - self._last_partial_ms
|
||||||
|
if since >= cfg.partial_interval_ms and buffered_ms >= cfg.min_segment_ms:
|
||||||
|
self._last_partial_ms = buffered_ms
|
||||||
|
return Segment(
|
||||||
|
audio=np.concatenate(self._buffer),
|
||||||
|
started_at=self._started_at,
|
||||||
|
ended_at=time.monotonic(),
|
||||||
|
is_final=False,
|
||||||
|
)
|
||||||
|
return None
|
||||||
|
|
||||||
|
def _finish(self) -> Segment:
|
||||||
|
audio = np.concatenate(self._buffer) if self._buffer else np.zeros(0, np.float32)
|
||||||
|
segment = Segment(
|
||||||
|
audio=audio,
|
||||||
|
started_at=self._started_at,
|
||||||
|
ended_at=time.monotonic(),
|
||||||
|
is_final=True,
|
||||||
|
)
|
||||||
|
self.reset()
|
||||||
|
return segment
|
||||||
|
|
||||||
|
def _is_speech(self, frame: np.ndarray) -> bool:
|
||||||
|
rms = float(np.sqrt(np.mean(np.square(frame), dtype=np.float64)))
|
||||||
|
cfg = self.config
|
||||||
|
|
||||||
|
if rms < cfg.absolute_floor:
|
||||||
|
self._noise_floor = min(self._noise_floor, max(rms, 1e-6))
|
||||||
|
return False
|
||||||
|
|
||||||
|
energetic = rms > max(self._noise_floor * cfg.speech_ratio, cfg.absolute_floor)
|
||||||
|
if not energetic:
|
||||||
|
# 무음 구간에서만 노이즈 플로어를 천천히 따라가게 한다.
|
||||||
|
self._noise_floor = 0.95 * self._noise_floor + 0.05 * rms
|
||||||
|
|
||||||
|
if self._vad is None:
|
||||||
|
return energetic
|
||||||
|
try:
|
||||||
|
pcm16 = (np.clip(frame, -1.0, 1.0) * 32767.0).astype(np.int16).tobytes()
|
||||||
|
return energetic and self._vad.is_speech(pcm16, SAMPLE_RATE)
|
||||||
|
except Exception: # noqa: BLE001 - webrtcvad 는 프레임 길이에 민감
|
||||||
|
return energetic
|
||||||
95
src/hearo/audio/wasapi_loopback.py
Normal file
95
src/hearo/audio/wasapi_loopback.py
Normal file
@@ -0,0 +1,95 @@
|
|||||||
|
"""WASAPI 장치 루프백 캡처 (Windows).
|
||||||
|
|
||||||
|
출력 장치 전체의 소리를 받는다. 프로그램별 분리는 안 되지만 의존성이
|
||||||
|
`PyAudioWPatch` 하나뿐이라 어디서나 바로 동작하는 기본 백엔드다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import contextlib
|
||||||
|
import logging
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..constants import FRAME_SAMPLES
|
||||||
|
from .base import AudioSource, CaptureBackend, CaptureError, to_mono_16k
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
try: # pragma: no cover - 플랫폼 의존
|
||||||
|
import pyaudiowpatch as pyaudio
|
||||||
|
except ImportError: # pragma: no cover
|
||||||
|
pyaudio = None # type: ignore[assignment]
|
||||||
|
|
||||||
|
|
||||||
|
class WasapiLoopbackCapture(CaptureBackend):
|
||||||
|
name = "device"
|
||||||
|
per_process = False
|
||||||
|
|
||||||
|
def __init__(self, device_index: int = -1) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.device_index = device_index
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def available() -> bool:
|
||||||
|
return pyaudio is not None
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def list_sources() -> list[AudioSource]:
|
||||||
|
if pyaudio is None:
|
||||||
|
return []
|
||||||
|
sources: list[AudioSource] = []
|
||||||
|
pa = pyaudio.PyAudio()
|
||||||
|
try:
|
||||||
|
default_index = -1
|
||||||
|
with contextlib.suppress(OSError, LookupError):
|
||||||
|
default_index = pa.get_default_wasapi_loopback()["index"]
|
||||||
|
for info in pa.get_loopback_device_info_generator():
|
||||||
|
idx = int(info["index"])
|
||||||
|
sources.append(
|
||||||
|
AudioSource(
|
||||||
|
kind="device",
|
||||||
|
identifier=str(idx),
|
||||||
|
label=str(info["name"]).replace(" [Loopback]", ""),
|
||||||
|
detail="기본 출력 장치" if idx == default_index else "출력 장치",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
finally:
|
||||||
|
pa.terminate()
|
||||||
|
return sources
|
||||||
|
|
||||||
|
def _resolve_device(self, pa) -> dict:
|
||||||
|
if self.device_index >= 0:
|
||||||
|
return pa.get_device_info_by_index(self.device_index)
|
||||||
|
return pa.get_default_wasapi_loopback()
|
||||||
|
|
||||||
|
def _run(self) -> None: # pragma: no cover - 실제 오디오 장치 필요
|
||||||
|
if pyaudio is None:
|
||||||
|
raise CaptureError("PyAudioWPatch가 설치되지 않았습니다.")
|
||||||
|
pa = pyaudio.PyAudio()
|
||||||
|
stream = None
|
||||||
|
try:
|
||||||
|
info = self._resolve_device(pa)
|
||||||
|
if info is None:
|
||||||
|
raise CaptureError("루프백 장치를 찾지 못했습니다.")
|
||||||
|
rate = int(info["defaultSampleRate"])
|
||||||
|
channels = int(info["maxInputChannels"]) or 2
|
||||||
|
chunk = max(FRAME_SAMPLES, int(rate * 0.02))
|
||||||
|
stream = pa.open(
|
||||||
|
format=pyaudio.paFloat32,
|
||||||
|
channels=channels,
|
||||||
|
rate=rate,
|
||||||
|
input=True,
|
||||||
|
input_device_index=int(info["index"]),
|
||||||
|
frames_per_buffer=chunk,
|
||||||
|
)
|
||||||
|
log.info("장치 루프백 캡처 시작: %s (%dHz, %dch)", info["name"], rate, channels)
|
||||||
|
while not self._stop.is_set():
|
||||||
|
raw = stream.read(chunk, exception_on_overflow=False)
|
||||||
|
data = np.frombuffer(raw, dtype=np.float32)
|
||||||
|
self._emit(to_mono_16k(data, channels, rate))
|
||||||
|
finally:
|
||||||
|
if stream is not None:
|
||||||
|
stream.stop_stream()
|
||||||
|
stream.close()
|
||||||
|
pa.terminate()
|
||||||
140
src/hearo/config.py
Normal file
140
src/hearo/config.py
Normal file
@@ -0,0 +1,140 @@
|
|||||||
|
"""설정 모델과 JSON 영속화."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import json
|
||||||
|
import logging
|
||||||
|
from dataclasses import asdict, dataclass, field, fields, is_dataclass
|
||||||
|
from pathlib import Path
|
||||||
|
from typing import Any
|
||||||
|
|
||||||
|
from .constants import DEFAULT_TIER, user_data_dir
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
CONFIG_PATH = user_data_dir() / "config.json"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class SubtitleStyle:
|
||||||
|
"""자막 텍스트 표현 설정."""
|
||||||
|
|
||||||
|
font_family: str = "Pretendard"
|
||||||
|
font_size: int = 34
|
||||||
|
bold: bool = True
|
||||||
|
text_color: str = "#FFFFFF"
|
||||||
|
source_text_color: str = "#9AA4B2"
|
||||||
|
outline_color: str = "#000000"
|
||||||
|
outline_width: int = 3
|
||||||
|
background_color: str = "#000000"
|
||||||
|
background_opacity: int = 55 # 0-100
|
||||||
|
line_spacing: int = 6
|
||||||
|
max_lines: int = 2
|
||||||
|
show_source: bool = False # 원문 동시 표시
|
||||||
|
align: str = "center" # left | center | right
|
||||||
|
fade_out_ms: int = 4000 # 0이면 계속 유지
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class OverlayConfig:
|
||||||
|
"""자막 창(오버레이) 위치/동작."""
|
||||||
|
|
||||||
|
x: int = -1 # -1 = 화면 하단 중앙 자동 배치
|
||||||
|
y: int = -1
|
||||||
|
width: int = 1100
|
||||||
|
height: int = 200
|
||||||
|
always_on_top: bool = True
|
||||||
|
click_through: bool = False
|
||||||
|
locked: bool = False
|
||||||
|
visible: bool = True
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class AudioConfig:
|
||||||
|
backend: str = "auto" # auto | process | device | file
|
||||||
|
target_pid: int = 0
|
||||||
|
target_process_name: str = ""
|
||||||
|
device_index: int = -1 # -1 = 기본 출력 장치 루프백
|
||||||
|
input_gain: float = 1.0
|
||||||
|
vad_aggressiveness: int = 2 # 0-3
|
||||||
|
silence_ms: int = 600 # 이 시간만큼 조용하면 한 문장 종료
|
||||||
|
min_segment_ms: int = 400
|
||||||
|
max_segment_ms: int = 12_000
|
||||||
|
partial_interval_ms: int = 900 # 중간 결과 갱신 주기, 0이면 끔
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class ModelConfig:
|
||||||
|
tier: str = DEFAULT_TIER
|
||||||
|
device: str = "cuda" # cuda | cpu
|
||||||
|
device_index: int = 0
|
||||||
|
source_lang: str = "auto" # auto 또는 언어코드
|
||||||
|
target_lang: str = "ko"
|
||||||
|
preload_on_start: bool = True
|
||||||
|
lora_adapter_path: str = "" # 추가학습 어댑터 (비어있으면 미사용)
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class GlossaryConfig:
|
||||||
|
enabled: bool = True
|
||||||
|
path: str = "" # 비어있으면 user_data_dir()/glossary.json
|
||||||
|
case_sensitive: bool = False
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class AppConfig:
|
||||||
|
audio: AudioConfig = field(default_factory=AudioConfig)
|
||||||
|
models: ModelConfig = field(default_factory=ModelConfig)
|
||||||
|
subtitle: SubtitleStyle = field(default_factory=SubtitleStyle)
|
||||||
|
overlay: OverlayConfig = field(default_factory=OverlayConfig)
|
||||||
|
glossary: GlossaryConfig = field(default_factory=GlossaryConfig)
|
||||||
|
theme: str = "dark" # dark | light
|
||||||
|
start_minimized: bool = False
|
||||||
|
log_transcripts: bool = True
|
||||||
|
|
||||||
|
# --- 영속화 ---------------------------------------------------------
|
||||||
|
def to_dict(self) -> dict[str, Any]:
|
||||||
|
return asdict(self)
|
||||||
|
|
||||||
|
def save(self, path: Path | None = None) -> Path:
|
||||||
|
target = path or CONFIG_PATH
|
||||||
|
target.parent.mkdir(parents=True, exist_ok=True)
|
||||||
|
tmp = target.with_suffix(".json.tmp")
|
||||||
|
tmp.write_text(
|
||||||
|
json.dumps(self.to_dict(), ensure_ascii=False, indent=2), encoding="utf-8"
|
||||||
|
)
|
||||||
|
tmp.replace(target)
|
||||||
|
return target
|
||||||
|
|
||||||
|
@classmethod
|
||||||
|
def load(cls, path: Path | None = None) -> AppConfig:
|
||||||
|
target = path or CONFIG_PATH
|
||||||
|
if not target.exists():
|
||||||
|
return cls()
|
||||||
|
try:
|
||||||
|
raw = json.loads(target.read_text(encoding="utf-8"))
|
||||||
|
except (OSError, json.JSONDecodeError) as exc:
|
||||||
|
log.warning("설정 파일을 읽지 못해 기본값을 사용합니다: %s", exc)
|
||||||
|
return cls()
|
||||||
|
return _from_dict(cls, raw)
|
||||||
|
|
||||||
|
|
||||||
|
def _from_dict(cls: type, raw: Any) -> Any:
|
||||||
|
"""알 수 없는 키는 버리고 누락된 키는 기본값으로 채우는 관대한 역직렬화.
|
||||||
|
|
||||||
|
기본값 인스턴스를 먼저 만든 뒤 알고 있는 필드만 덮어쓴다. 덕분에
|
||||||
|
설정 스키마가 바뀌어도 사용자의 기존 config.json이 깨지지 않는다.
|
||||||
|
"""
|
||||||
|
instance = cls()
|
||||||
|
if not isinstance(raw, dict):
|
||||||
|
return instance
|
||||||
|
known = {f.name for f in fields(cls)}
|
||||||
|
for key, value in raw.items():
|
||||||
|
if key not in known:
|
||||||
|
continue
|
||||||
|
current = getattr(instance, key)
|
||||||
|
if is_dataclass(current) and not isinstance(current, type):
|
||||||
|
setattr(instance, key, _from_dict(type(current), value))
|
||||||
|
else:
|
||||||
|
setattr(instance, key, value)
|
||||||
|
return instance
|
||||||
52
src/hearo/constants.py
Normal file
52
src/hearo/constants.py
Normal file
@@ -0,0 +1,52 @@
|
|||||||
|
"""전역 상수. 제품명 변경은 이 파일만 고치면 됩니다."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
APP_NAME = "Hearo"
|
||||||
|
APP_NAME_KO = "히어로"
|
||||||
|
APP_SLOGAN = "듣고, 바로 이해하다"
|
||||||
|
APP_VERSION = "0.1.0"
|
||||||
|
ORG_NAME = "tkrmagid"
|
||||||
|
|
||||||
|
# 지원 언어 (1차: 4개)
|
||||||
|
LANGUAGES: dict[str, dict[str, str]] = {
|
||||||
|
"ko": {"label": "한국어", "english": "Korean", "nllb": "kor_Hang", "flag": "KO"},
|
||||||
|
"en": {"label": "English", "english": "English", "nllb": "eng_Latn", "flag": "EN"},
|
||||||
|
"ja": {"label": "日本語", "english": "Japanese", "nllb": "jpn_Jpan", "flag": "JA"},
|
||||||
|
"zh": {"label": "中文", "english": "Chinese", "nllb": "zho_Hans", "flag": "ZH"},
|
||||||
|
}
|
||||||
|
LANGUAGE_CODES = tuple(LANGUAGES.keys())
|
||||||
|
|
||||||
|
# 기본 모델 품질 티어. 실제 티어 정의는 models/tiers.py 에 있지만,
|
||||||
|
# config 가 models 패키지를 import 하면 순환 의존이 생기므로 키만 여기에 둔다.
|
||||||
|
DEFAULT_TIER = "balance"
|
||||||
|
|
||||||
|
# 오디오 파이프라인 규격
|
||||||
|
SAMPLE_RATE = 16_000
|
||||||
|
CHANNELS = 1
|
||||||
|
FRAME_MS = 20
|
||||||
|
FRAME_SAMPLES = SAMPLE_RATE * FRAME_MS // 1000
|
||||||
|
|
||||||
|
|
||||||
|
def user_data_dir() -> Path:
|
||||||
|
"""설정/로그/모델 캐시를 두는 사용자 디렉터리."""
|
||||||
|
import os
|
||||||
|
import sys
|
||||||
|
|
||||||
|
if sys.platform == "win32":
|
||||||
|
base = Path(os.environ.get("APPDATA", Path.home() / "AppData" / "Roaming"))
|
||||||
|
elif sys.platform == "darwin":
|
||||||
|
base = Path.home() / "Library" / "Application Support"
|
||||||
|
else:
|
||||||
|
base = Path(os.environ.get("XDG_CONFIG_HOME", Path.home() / ".config"))
|
||||||
|
path = base / APP_NAME
|
||||||
|
path.mkdir(parents=True, exist_ok=True)
|
||||||
|
return path
|
||||||
|
|
||||||
|
|
||||||
|
def models_dir() -> Path:
|
||||||
|
path = user_data_dir() / "models"
|
||||||
|
path.mkdir(parents=True, exist_ok=True)
|
||||||
|
return path
|
||||||
6
src/hearo/core/__init__.py
Normal file
6
src/hearo/core/__init__.py
Normal file
@@ -0,0 +1,6 @@
|
|||||||
|
"""파이프라인 코어."""
|
||||||
|
|
||||||
|
from .engine import TranslationEngine
|
||||||
|
from .events import EngineState, EngineStatus, TranslationLine
|
||||||
|
|
||||||
|
__all__ = ["EngineState", "EngineStatus", "TranslationEngine", "TranslationLine"]
|
||||||
259
src/hearo/core/engine.py
Normal file
259
src/hearo/core/engine.py
Normal file
@@ -0,0 +1,259 @@
|
|||||||
|
"""실시간 번역 파이프라인 오케스트레이션.
|
||||||
|
|
||||||
|
캡처 스레드 → [오디오 큐] → 분할 스레드 → [구간 큐] → 인식/번역 워커 → 콜백
|
||||||
|
|
||||||
|
UI 프레임워크에 의존하지 않는다. Qt 쪽은 콜백을 시그널로 다시 던져준다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
import queue
|
||||||
|
import threading
|
||||||
|
import time
|
||||||
|
from collections.abc import Callable
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from ..audio import (
|
||||||
|
CaptureBackend,
|
||||||
|
CaptureError,
|
||||||
|
Segment,
|
||||||
|
Segmenter,
|
||||||
|
SegmenterConfig,
|
||||||
|
create_capture,
|
||||||
|
)
|
||||||
|
from ..config import AppConfig
|
||||||
|
from ..constants import LANGUAGE_CODES, user_data_dir
|
||||||
|
from ..models import Glossary, ModelManager
|
||||||
|
from .events import EngineState, EngineStatus, TranslationLine
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
StatusCallback = Callable[[EngineStatus], None]
|
||||||
|
LineCallback = Callable[[TranslationLine], None]
|
||||||
|
|
||||||
|
|
||||||
|
class TranslationEngine:
|
||||||
|
def __init__(
|
||||||
|
self,
|
||||||
|
config: AppConfig,
|
||||||
|
on_line: LineCallback | None = None,
|
||||||
|
on_status: StatusCallback | None = None,
|
||||||
|
) -> None:
|
||||||
|
self.config = config
|
||||||
|
self._on_line = on_line
|
||||||
|
self._on_status = on_status
|
||||||
|
|
||||||
|
self.models = ModelManager(config.models)
|
||||||
|
self.glossary = self._load_glossary()
|
||||||
|
|
||||||
|
self._capture: CaptureBackend | None = None
|
||||||
|
self._segments: queue.Queue[Segment] = queue.Queue(maxsize=32)
|
||||||
|
self._threads: list[threading.Thread] = []
|
||||||
|
self._stop = threading.Event()
|
||||||
|
self._status = EngineStatus()
|
||||||
|
self._level = 0.0
|
||||||
|
#: 인식 정확도를 높이기 위해 직전 확정 문장을 프롬프트로 넘긴다.
|
||||||
|
self._last_final_text = ""
|
||||||
|
|
||||||
|
# --- 수명주기 -------------------------------------------------------
|
||||||
|
@property
|
||||||
|
def status(self) -> EngineStatus:
|
||||||
|
return self._status
|
||||||
|
|
||||||
|
@property
|
||||||
|
def running(self) -> bool:
|
||||||
|
return self._status.state in (EngineState.RUNNING, EngineState.LOADING)
|
||||||
|
|
||||||
|
def start(self) -> None:
|
||||||
|
if self.running:
|
||||||
|
return
|
||||||
|
self._stop.clear()
|
||||||
|
self._emit_status(EngineState.LOADING, "모델을 준비하고 있습니다…", 0.0)
|
||||||
|
|
||||||
|
thread = threading.Thread(target=self._bootstrap, name="engine-boot", daemon=True)
|
||||||
|
thread.start()
|
||||||
|
self._threads = [thread]
|
||||||
|
|
||||||
|
def stop(self) -> None:
|
||||||
|
if self._status.state is EngineState.IDLE:
|
||||||
|
return
|
||||||
|
self._emit_status(EngineState.STOPPING, "정지 중…")
|
||||||
|
self._stop.set()
|
||||||
|
if self._capture is not None:
|
||||||
|
self._capture.stop()
|
||||||
|
self._capture = None
|
||||||
|
for t in self._threads:
|
||||||
|
if t is not threading.current_thread():
|
||||||
|
t.join(timeout=3)
|
||||||
|
self._threads = []
|
||||||
|
self._drain_segments()
|
||||||
|
self._emit_status(EngineState.IDLE, "대기 중")
|
||||||
|
|
||||||
|
def shutdown(self) -> None:
|
||||||
|
self.stop()
|
||||||
|
self.models.unload()
|
||||||
|
|
||||||
|
def reload_glossary(self) -> None:
|
||||||
|
self.glossary = self._load_glossary()
|
||||||
|
|
||||||
|
# --- 내부 스레드 ----------------------------------------------------
|
||||||
|
def _bootstrap(self) -> None:
|
||||||
|
try:
|
||||||
|
if self.config.models.preload_on_start:
|
||||||
|
self.models.preload(
|
||||||
|
lambda msg, pct: self._emit_status(EngineState.LOADING, msg, pct)
|
||||||
|
)
|
||||||
|
self._capture = create_capture(self.config.audio)
|
||||||
|
self._capture.start()
|
||||||
|
|
||||||
|
workers = [
|
||||||
|
threading.Thread(target=self._segment_loop, name="engine-segment", daemon=True),
|
||||||
|
threading.Thread(target=self._infer_loop, name="engine-infer", daemon=True),
|
||||||
|
]
|
||||||
|
for w in workers:
|
||||||
|
w.start()
|
||||||
|
self._threads.extend(workers)
|
||||||
|
|
||||||
|
kind = "프로그램별 캡처" if self._capture.per_process else "출력 장치 전체"
|
||||||
|
self._emit_status(EngineState.RUNNING, f"듣는 중 · {kind}", 1.0)
|
||||||
|
except (CaptureError, OSError, RuntimeError, ImportError) as exc:
|
||||||
|
log.exception("엔진 시작 실패")
|
||||||
|
self._emit_status(EngineState.ERROR, str(exc))
|
||||||
|
|
||||||
|
def _segment_loop(self) -> None:
|
||||||
|
cfg = self.config.audio
|
||||||
|
segmenter = Segmenter(
|
||||||
|
SegmenterConfig(
|
||||||
|
silence_ms=cfg.silence_ms,
|
||||||
|
min_segment_ms=cfg.min_segment_ms,
|
||||||
|
max_segment_ms=cfg.max_segment_ms,
|
||||||
|
partial_interval_ms=cfg.partial_interval_ms,
|
||||||
|
vad_aggressiveness=cfg.vad_aggressiveness,
|
||||||
|
)
|
||||||
|
)
|
||||||
|
while not self._stop.is_set():
|
||||||
|
capture = self._capture
|
||||||
|
if capture is None:
|
||||||
|
break
|
||||||
|
chunk = capture.read(timeout=0.3)
|
||||||
|
if chunk is None:
|
||||||
|
if capture.error is not None:
|
||||||
|
self._emit_status(EngineState.ERROR, str(capture.error))
|
||||||
|
return
|
||||||
|
continue
|
||||||
|
|
||||||
|
if cfg.input_gain != 1.0:
|
||||||
|
chunk = np.clip(chunk * cfg.input_gain, -1.0, 1.0)
|
||||||
|
self._level = float(np.sqrt(np.mean(np.square(chunk, dtype=np.float64))))
|
||||||
|
|
||||||
|
for segment in segmenter.push(chunk):
|
||||||
|
self._enqueue(segment)
|
||||||
|
|
||||||
|
leftover = segmenter.flush()
|
||||||
|
if leftover is not None:
|
||||||
|
self._enqueue(leftover)
|
||||||
|
|
||||||
|
def _enqueue(self, segment: Segment) -> None:
|
||||||
|
try:
|
||||||
|
self._segments.put_nowait(segment)
|
||||||
|
except queue.Full:
|
||||||
|
# 밀렸다면 중간 결과부터 버린다. 확정 문장은 지켜야 한다.
|
||||||
|
if not segment.is_final:
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
self._segments.get_nowait()
|
||||||
|
self._segments.put_nowait(segment)
|
||||||
|
except (queue.Empty, queue.Full):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def _infer_loop(self) -> None:
|
||||||
|
recognizer = self.models.recognizer()
|
||||||
|
translator = self.models.translator()
|
||||||
|
mcfg = self.config.models
|
||||||
|
|
||||||
|
while not self._stop.is_set():
|
||||||
|
try:
|
||||||
|
segment = self._segments.get(timeout=0.3)
|
||||||
|
except queue.Empty:
|
||||||
|
continue
|
||||||
|
|
||||||
|
# 처리 중 더 새로운 확정 구간이 들어왔다면 오래된 중간 결과는 버린다.
|
||||||
|
if not segment.is_final and not self._segments.empty():
|
||||||
|
continue
|
||||||
|
|
||||||
|
try:
|
||||||
|
self._process(segment, recognizer, translator, mcfg)
|
||||||
|
except Exception as exc: # noqa: BLE001 - 한 문장 실패로 엔진이 죽으면 안 된다
|
||||||
|
log.warning("구간 처리 실패: %s", exc, exc_info=log.isEnabledFor(logging.DEBUG))
|
||||||
|
|
||||||
|
def _process(self, segment: Segment, recognizer, translator, mcfg) -> None:
|
||||||
|
source_lang = None if mcfg.source_lang == "auto" else mcfg.source_lang
|
||||||
|
|
||||||
|
t0 = time.perf_counter()
|
||||||
|
transcript = recognizer.transcribe(
|
||||||
|
segment.audio,
|
||||||
|
language=source_lang,
|
||||||
|
fast=not segment.is_final,
|
||||||
|
prompt=self._last_final_text if segment.is_final else "",
|
||||||
|
)
|
||||||
|
asr_ms = int((time.perf_counter() - t0) * 1000)
|
||||||
|
if not transcript.text:
|
||||||
|
return
|
||||||
|
|
||||||
|
detected = transcript.language if transcript.language in LANGUAGE_CODES else mcfg.target_lang
|
||||||
|
target = mcfg.target_lang
|
||||||
|
|
||||||
|
translated = ""
|
||||||
|
mt_ms = 0
|
||||||
|
if segment.is_final:
|
||||||
|
# 중간 결과는 원문만 흘려보낸다. 번역은 확정 문장에만 돌려 GPU를 아낀다.
|
||||||
|
t1 = time.perf_counter()
|
||||||
|
translated = translator.translate(
|
||||||
|
transcript.text,
|
||||||
|
detected,
|
||||||
|
target,
|
||||||
|
self.glossary if self.config.glossary.enabled else None,
|
||||||
|
)
|
||||||
|
mt_ms = int((time.perf_counter() - t1) * 1000)
|
||||||
|
self._last_final_text = transcript.text
|
||||||
|
|
||||||
|
line = TranslationLine(
|
||||||
|
source_text=transcript.text,
|
||||||
|
translated_text=translated,
|
||||||
|
source_lang=detected,
|
||||||
|
target_lang=target,
|
||||||
|
is_final=segment.is_final,
|
||||||
|
latency_s=time.monotonic() - segment.ended_at,
|
||||||
|
asr_ms=asr_ms,
|
||||||
|
mt_ms=mt_ms,
|
||||||
|
)
|
||||||
|
if self._on_line:
|
||||||
|
self._on_line(line)
|
||||||
|
self._status.level = self._level
|
||||||
|
self._status.queue_depth = self._segments.qsize()
|
||||||
|
|
||||||
|
# --- 유틸 -----------------------------------------------------------
|
||||||
|
def _load_glossary(self) -> Glossary:
|
||||||
|
cfg = self.config.glossary
|
||||||
|
path = cfg.path or str(user_data_dir() / "glossary.json")
|
||||||
|
return Glossary.load(path, cfg.case_sensitive)
|
||||||
|
|
||||||
|
def _drain_segments(self) -> None:
|
||||||
|
while not self._segments.empty():
|
||||||
|
try:
|
||||||
|
self._segments.get_nowait()
|
||||||
|
except queue.Empty:
|
||||||
|
break
|
||||||
|
|
||||||
|
def _emit_status(self, state: EngineState, message: str, progress: float | None = None) -> None:
|
||||||
|
self._status = EngineStatus(
|
||||||
|
state=state,
|
||||||
|
message=message,
|
||||||
|
progress=self._status.progress if progress is None else progress,
|
||||||
|
level=self._level,
|
||||||
|
queue_depth=self._segments.qsize(),
|
||||||
|
)
|
||||||
|
if self._on_status:
|
||||||
|
self._on_status(self._status)
|
||||||
45
src/hearo/core/events.py
Normal file
45
src/hearo/core/events.py
Normal file
@@ -0,0 +1,45 @@
|
|||||||
|
"""엔진이 UI로 올려보내는 이벤트 타입."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import time
|
||||||
|
from dataclasses import dataclass, field
|
||||||
|
from enum import Enum
|
||||||
|
|
||||||
|
|
||||||
|
class EngineState(str, Enum):
|
||||||
|
IDLE = "idle"
|
||||||
|
LOADING = "loading"
|
||||||
|
RUNNING = "running"
|
||||||
|
STOPPING = "stopping"
|
||||||
|
ERROR = "error"
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class TranslationLine:
|
||||||
|
"""화면에 뿌릴 자막 한 줄."""
|
||||||
|
|
||||||
|
source_text: str
|
||||||
|
translated_text: str
|
||||||
|
source_lang: str
|
||||||
|
target_lang: str
|
||||||
|
is_final: bool = True
|
||||||
|
created_at: float = field(default_factory=time.time)
|
||||||
|
#: 발화가 끝난 시점부터 번역이 나오기까지 걸린 시간(초)
|
||||||
|
latency_s: float = 0.0
|
||||||
|
asr_ms: int = 0
|
||||||
|
mt_ms: int = 0
|
||||||
|
|
||||||
|
@property
|
||||||
|
def display(self) -> str:
|
||||||
|
return self.translated_text or self.source_text
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass
|
||||||
|
class EngineStatus:
|
||||||
|
state: EngineState = EngineState.IDLE
|
||||||
|
message: str = ""
|
||||||
|
progress: float = 0.0
|
||||||
|
#: 최근 입력 오디오 레벨 (0.0 ~ 1.0), VU 미터용
|
||||||
|
level: float = 0.0
|
||||||
|
queue_depth: int = 0
|
||||||
6
src/hearo/ui/__init__.py
Normal file
6
src/hearo/ui/__init__.py
Normal file
@@ -0,0 +1,6 @@
|
|||||||
|
"""Qt UI 레이어."""
|
||||||
|
|
||||||
|
from .main_window import MainWindow
|
||||||
|
from .overlay import SubtitleOverlay, SubtitlePreview
|
||||||
|
|
||||||
|
__all__ = ["MainWindow", "SubtitleOverlay", "SubtitlePreview"]
|
||||||
258
src/hearo/ui/main_window.py
Normal file
258
src/hearo/ui/main_window.py
Normal file
@@ -0,0 +1,258 @@
|
|||||||
|
"""메인 컨트롤 창 — 사이드바 + 페이지 스택.
|
||||||
|
|
||||||
|
디자인 의도: 자막은 게임 위 오버레이가 전담하고, 이 창은 '설정과 상태'만
|
||||||
|
담당한다. 그래서 두 창이 서로 방해하지 않는다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import logging
|
||||||
|
from datetime import datetime
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt, QTimer, Signal
|
||||||
|
from PySide6.QtGui import QCloseEvent
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QButtonGroup,
|
||||||
|
QFrame,
|
||||||
|
QHBoxLayout,
|
||||||
|
QLabel,
|
||||||
|
QMainWindow,
|
||||||
|
QMessageBox,
|
||||||
|
QPushButton,
|
||||||
|
QScrollArea,
|
||||||
|
QStackedWidget,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ..config import AppConfig
|
||||||
|
from ..constants import APP_NAME, APP_SLOGAN, APP_VERSION, user_data_dir
|
||||||
|
from ..core import EngineState, EngineStatus, TranslationEngine, TranslationLine
|
||||||
|
from .overlay import SubtitleOverlay
|
||||||
|
from .pages import GlossaryPage, HomePage, ModelsPage, SettingsPage, SubtitlePage
|
||||||
|
from .theme import SPACING, palette, stylesheet
|
||||||
|
|
||||||
|
log = logging.getLogger(__name__)
|
||||||
|
|
||||||
|
NAV = [
|
||||||
|
("홈", "실시간 번역"),
|
||||||
|
("모델", "품질 티어"),
|
||||||
|
("자막", "표시 설정"),
|
||||||
|
("용어집", "고유명사 고정"),
|
||||||
|
("설정", "세부 조정"),
|
||||||
|
]
|
||||||
|
|
||||||
|
|
||||||
|
class MainWindow(QMainWindow):
|
||||||
|
#: 엔진 콜백은 워커 스레드에서 오므로 시그널로 GUI 스레드에 넘긴다.
|
||||||
|
line_received = Signal(object)
|
||||||
|
status_received = Signal(object)
|
||||||
|
|
||||||
|
def __init__(self, config: AppConfig) -> None:
|
||||||
|
super().__init__()
|
||||||
|
self.config = config
|
||||||
|
self.setWindowTitle(f"{APP_NAME} · {APP_SLOGAN}")
|
||||||
|
self.resize(1060, 820)
|
||||||
|
|
||||||
|
self.engine = TranslationEngine(
|
||||||
|
config,
|
||||||
|
on_line=self.line_received.emit,
|
||||||
|
on_status=self.status_received.emit,
|
||||||
|
)
|
||||||
|
self.overlay = SubtitleOverlay(config.subtitle, config.overlay)
|
||||||
|
self._transcript_file = None
|
||||||
|
|
||||||
|
self._build_ui()
|
||||||
|
self._connect()
|
||||||
|
|
||||||
|
self._save_timer = QTimer(self)
|
||||||
|
self._save_timer.setSingleShot(True)
|
||||||
|
self._save_timer.setInterval(800) # 입력 중 계속 쓰지 않도록 묶어서 저장
|
||||||
|
self._save_timer.timeout.connect(lambda: self.config.save())
|
||||||
|
|
||||||
|
# --- UI -------------------------------------------------------------
|
||||||
|
def _build_ui(self) -> None:
|
||||||
|
self.setStyleSheet(stylesheet(self.config.theme))
|
||||||
|
|
||||||
|
central = QWidget()
|
||||||
|
layout = QHBoxLayout(central)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(0)
|
||||||
|
layout.addWidget(self._build_sidebar())
|
||||||
|
|
||||||
|
self.stack = QStackedWidget()
|
||||||
|
self.home_page = HomePage(self.config)
|
||||||
|
self.models_page = ModelsPage(self.config)
|
||||||
|
self.subtitle_page = SubtitlePage(self.config)
|
||||||
|
self.glossary_page = GlossaryPage(self.config)
|
||||||
|
self.settings_page = SettingsPage(self.config)
|
||||||
|
for page in (
|
||||||
|
self.home_page,
|
||||||
|
self.models_page,
|
||||||
|
self.subtitle_page,
|
||||||
|
self.glossary_page,
|
||||||
|
self.settings_page,
|
||||||
|
):
|
||||||
|
self.stack.addWidget(_scrollable(page))
|
||||||
|
layout.addWidget(self.stack, 1)
|
||||||
|
self.setCentralWidget(central)
|
||||||
|
|
||||||
|
def _build_sidebar(self) -> QFrame:
|
||||||
|
p = palette(self.config.theme)
|
||||||
|
bar = QFrame()
|
||||||
|
bar.setObjectName("Sidebar")
|
||||||
|
bar.setFixedWidth(212)
|
||||||
|
|
||||||
|
layout = QVBoxLayout(bar)
|
||||||
|
layout.setContentsMargins(14, 22, 14, 18)
|
||||||
|
layout.setSpacing(6)
|
||||||
|
|
||||||
|
brand = QLabel(APP_NAME)
|
||||||
|
brand.setStyleSheet(
|
||||||
|
f"font-size: 24px; font-weight: 800; color: {p.accent}; padding-left: 6px;"
|
||||||
|
)
|
||||||
|
layout.addWidget(brand)
|
||||||
|
tag = QLabel(APP_SLOGAN)
|
||||||
|
tag.setObjectName("Hint")
|
||||||
|
tag.setStyleSheet(f"color: {p.text_dim}; padding-left: 7px;")
|
||||||
|
layout.addWidget(tag)
|
||||||
|
layout.addSpacing(SPACING + 8)
|
||||||
|
|
||||||
|
self._nav_group = QButtonGroup(self)
|
||||||
|
self._nav_group.setExclusive(True)
|
||||||
|
for index, (title, hint) in enumerate(NAV):
|
||||||
|
button = QPushButton(title)
|
||||||
|
button.setObjectName("NavButton")
|
||||||
|
button.setCheckable(True)
|
||||||
|
button.setToolTip(hint)
|
||||||
|
button.setCursor(Qt.CursorShape.PointingHandCursor)
|
||||||
|
self._nav_group.addButton(button, index)
|
||||||
|
layout.addWidget(button)
|
||||||
|
self._nav_group.button(0).setChecked(True)
|
||||||
|
self._nav_group.idClicked.connect(self._goto_page)
|
||||||
|
|
||||||
|
layout.addStretch(1)
|
||||||
|
version = QLabel(f"v{APP_VERSION}")
|
||||||
|
version.setObjectName("Hint")
|
||||||
|
version.setAlignment(Qt.AlignmentFlag.AlignCenter)
|
||||||
|
layout.addWidget(version)
|
||||||
|
return bar
|
||||||
|
|
||||||
|
def _connect(self) -> None:
|
||||||
|
self.line_received.connect(self._on_line)
|
||||||
|
self.status_received.connect(self._on_status)
|
||||||
|
|
||||||
|
self.home_page.start_requested.connect(self.start_engine)
|
||||||
|
self.home_page.stop_requested.connect(self.stop_engine)
|
||||||
|
self.home_page.config_changed.connect(self._schedule_save)
|
||||||
|
|
||||||
|
self.models_page.tier_changed.connect(self._on_tier_changed)
|
||||||
|
self.models_page.config_changed.connect(self._schedule_save)
|
||||||
|
|
||||||
|
self.subtitle_page.style_changed.connect(self._on_style_changed)
|
||||||
|
self.subtitle_page.overlay_toggled.connect(self._on_overlay_toggled)
|
||||||
|
|
||||||
|
self.glossary_page.glossary_changed.connect(self._on_glossary_changed)
|
||||||
|
|
||||||
|
self.settings_page.config_changed.connect(self._schedule_save)
|
||||||
|
self.settings_page.restart_needed.connect(self._on_restart_needed)
|
||||||
|
|
||||||
|
self.overlay.geometry_changed.connect(lambda *_: self._schedule_save())
|
||||||
|
self.overlay.closed.connect(
|
||||||
|
lambda: self.subtitle_page.overlay_check.setChecked(False)
|
||||||
|
)
|
||||||
|
|
||||||
|
# --- 엔진 -----------------------------------------------------------
|
||||||
|
def start_engine(self) -> None:
|
||||||
|
if self.config.overlay.visible:
|
||||||
|
self.overlay.show()
|
||||||
|
self.config.save()
|
||||||
|
self.engine.start()
|
||||||
|
|
||||||
|
def stop_engine(self) -> None:
|
||||||
|
self.engine.stop()
|
||||||
|
self.overlay.clear()
|
||||||
|
self._close_transcript()
|
||||||
|
|
||||||
|
def _on_tier_changed(self, tier_key: str) -> None:
|
||||||
|
was_running = self.engine.running
|
||||||
|
if was_running:
|
||||||
|
self.engine.stop()
|
||||||
|
self.engine.models.switch_tier(tier_key)
|
||||||
|
if was_running:
|
||||||
|
self.engine.start()
|
||||||
|
|
||||||
|
def _on_restart_needed(self) -> None:
|
||||||
|
if self.engine.running:
|
||||||
|
self.engine.stop()
|
||||||
|
self.engine.start()
|
||||||
|
|
||||||
|
# --- 이벤트 ---------------------------------------------------------
|
||||||
|
def _on_line(self, line: TranslationLine) -> None:
|
||||||
|
self.overlay.set_line(line.translated_text, line.source_text, line.is_final)
|
||||||
|
self.home_page.on_line(line)
|
||||||
|
if line.is_final and self.config.log_transcripts:
|
||||||
|
self._write_transcript(line)
|
||||||
|
|
||||||
|
def _on_status(self, status: EngineStatus) -> None:
|
||||||
|
self.home_page.on_status(status)
|
||||||
|
if status.state is EngineState.ERROR:
|
||||||
|
QMessageBox.warning(self, "오류", status.message)
|
||||||
|
|
||||||
|
def _on_style_changed(self) -> None:
|
||||||
|
self.overlay.update()
|
||||||
|
self._schedule_save()
|
||||||
|
|
||||||
|
def _on_overlay_toggled(self, visible: bool) -> None:
|
||||||
|
self.config.overlay.visible = visible
|
||||||
|
self.overlay.setVisible(visible)
|
||||||
|
self._schedule_save()
|
||||||
|
|
||||||
|
def _on_glossary_changed(self) -> None:
|
||||||
|
self.engine.reload_glossary()
|
||||||
|
self._schedule_save()
|
||||||
|
|
||||||
|
def _goto_page(self, index: int) -> None:
|
||||||
|
self.stack.setCurrentIndex(index)
|
||||||
|
|
||||||
|
def _schedule_save(self) -> None:
|
||||||
|
self._save_timer.start()
|
||||||
|
|
||||||
|
# --- 기록 -----------------------------------------------------------
|
||||||
|
def _write_transcript(self, line: TranslationLine) -> None:
|
||||||
|
if self._transcript_file is None:
|
||||||
|
folder = user_data_dir() / "transcripts"
|
||||||
|
folder.mkdir(parents=True, exist_ok=True)
|
||||||
|
name = datetime.now().strftime("%Y%m%d-%H%M%S.txt")
|
||||||
|
# 세션 내내 열어두는 핸들이다. _close_transcript() 에서 닫는다.
|
||||||
|
self._transcript_file = open(folder / name, "a", encoding="utf-8") # noqa: SIM115
|
||||||
|
stamp = datetime.now().strftime("%H:%M:%S")
|
||||||
|
self._transcript_file.write(
|
||||||
|
f"[{stamp}] {line.source_text}\n {line.translated_text}\n"
|
||||||
|
)
|
||||||
|
self._transcript_file.flush()
|
||||||
|
|
||||||
|
def _close_transcript(self) -> None:
|
||||||
|
if self._transcript_file is not None:
|
||||||
|
self._transcript_file.close()
|
||||||
|
self._transcript_file = None
|
||||||
|
|
||||||
|
# --- 종료 -----------------------------------------------------------
|
||||||
|
def closeEvent(self, event: QCloseEvent) -> None: # noqa: N802
|
||||||
|
self.config.save()
|
||||||
|
self.engine.shutdown()
|
||||||
|
self._close_transcript()
|
||||||
|
self.overlay.close()
|
||||||
|
super().closeEvent(event)
|
||||||
|
|
||||||
|
|
||||||
|
def _scrollable(widget: QWidget) -> QScrollArea:
|
||||||
|
area = QScrollArea()
|
||||||
|
area.setWidgetResizable(True)
|
||||||
|
area.setHorizontalScrollBarPolicy(Qt.ScrollBarPolicy.ScrollBarAlwaysOff)
|
||||||
|
holder = QWidget()
|
||||||
|
layout = QVBoxLayout(holder)
|
||||||
|
layout.setContentsMargins(26, 24, 26, 24)
|
||||||
|
layout.addWidget(widget)
|
||||||
|
area.setWidget(holder)
|
||||||
|
return area
|
||||||
346
src/hearo/ui/overlay.py
Normal file
346
src/hearo/ui/overlay.py
Normal file
@@ -0,0 +1,346 @@
|
|||||||
|
"""자막 오버레이 창.
|
||||||
|
|
||||||
|
게임/영상 위에 얹는 무테두리 반투명 창. 기본 동작은 유튜브 자막과 같은
|
||||||
|
'외곽선 있는 큰 흰 글씨'로, 어떤 배경 위에서도 읽히는 것을 최우선으로 한다.
|
||||||
|
|
||||||
|
조작:
|
||||||
|
* 드래그 — 위치 이동
|
||||||
|
* 우하단 모서리 — 크기 조절
|
||||||
|
* 휠 — 글자 크기 조절
|
||||||
|
* 우클릭 — 잠금 / 클릭 통과 / 숨기기 메뉴
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from PySide6.QtCore import QPoint, QRect, Qt, QTimer, Signal
|
||||||
|
from PySide6.QtGui import (
|
||||||
|
QAction,
|
||||||
|
QColor,
|
||||||
|
QFont,
|
||||||
|
QFontMetrics,
|
||||||
|
QPainter,
|
||||||
|
QPainterPath,
|
||||||
|
QPen,
|
||||||
|
)
|
||||||
|
from PySide6.QtWidgets import QApplication, QMenu, QWidget
|
||||||
|
|
||||||
|
from ..config import OverlayConfig, SubtitleStyle
|
||||||
|
|
||||||
|
GRIP = 18 # 우하단 리사이즈 핸들 크기
|
||||||
|
|
||||||
|
|
||||||
|
def paint_subtitle(
|
||||||
|
painter: QPainter,
|
||||||
|
rect: QRect,
|
||||||
|
st: SubtitleStyle,
|
||||||
|
translated: str,
|
||||||
|
source: str = "",
|
||||||
|
partial: str = "",
|
||||||
|
) -> None:
|
||||||
|
"""자막 한 화면을 그린다. 오버레이와 설정 미리보기가 이 함수를 공유한다."""
|
||||||
|
painter.setRenderHint(QPainter.RenderHint.Antialiasing, True)
|
||||||
|
painter.setRenderHint(QPainter.RenderHint.TextAntialiasing, True)
|
||||||
|
|
||||||
|
if st.background_opacity > 0:
|
||||||
|
bg = QColor(st.background_color)
|
||||||
|
bg.setAlpha(int(255 * st.background_opacity / 100))
|
||||||
|
painter.setBrush(bg)
|
||||||
|
painter.setPen(Qt.PenStyle.NoPen)
|
||||||
|
painter.drawRoundedRect(rect, 12, 12)
|
||||||
|
|
||||||
|
rows: list[tuple[str, QColor, int]] = []
|
||||||
|
if st.show_source and source and translated:
|
||||||
|
rows.append((source, QColor(st.source_text_color), int(st.font_size * 0.62)))
|
||||||
|
main = translated or partial
|
||||||
|
if main:
|
||||||
|
color = QColor(st.text_color)
|
||||||
|
if not translated: # 아직 확정 전 — 흐리게
|
||||||
|
color.setAlpha(170)
|
||||||
|
rows.append((main, color, st.font_size))
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
|
||||||
|
margin = 18
|
||||||
|
available = rect.width() - margin * 2
|
||||||
|
blocks: list[tuple[list[str], QColor, QFont]] = []
|
||||||
|
total_h = 0
|
||||||
|
for text, color, size in rows:
|
||||||
|
font = QFont(st.font_family, size)
|
||||||
|
font.setBold(st.bold)
|
||||||
|
lines = _wrap(text, QFontMetrics(font), available, st.max_lines)
|
||||||
|
blocks.append((lines, color, font))
|
||||||
|
total_h += len(lines) * (QFontMetrics(font).height() + st.line_spacing)
|
||||||
|
|
||||||
|
y = rect.y() + max(margin, (rect.height() - total_h) // 2)
|
||||||
|
for lines, color, font in blocks:
|
||||||
|
painter.setFont(font)
|
||||||
|
fm = QFontMetrics(font)
|
||||||
|
for line in lines:
|
||||||
|
width = fm.horizontalAdvance(line)
|
||||||
|
if st.align == "left":
|
||||||
|
x = rect.x() + margin
|
||||||
|
elif st.align == "right":
|
||||||
|
x = rect.right() - margin - width
|
||||||
|
else:
|
||||||
|
x = rect.x() + (rect.width() - width) // 2
|
||||||
|
_draw_outlined_text(painter, x, y + fm.ascent(), line, color, font, st)
|
||||||
|
y += fm.height() + st.line_spacing
|
||||||
|
|
||||||
|
|
||||||
|
def _draw_outlined_text(painter, x, y, text, color, font, st: SubtitleStyle) -> None:
|
||||||
|
"""외곽선 있는 글자.
|
||||||
|
|
||||||
|
스트로크는 글자 윤곽선 '가운데'에 그려지므로 절반이 글자 안쪽을 파먹는다.
|
||||||
|
한글처럼 획이 얇은 글꼴은 그대로 두면 글자가 통째로 외곽선 색이 된다.
|
||||||
|
그래서 스트로크를 먼저 깔고 그 위에 글자를 채우는 2패스로 그린다.
|
||||||
|
"""
|
||||||
|
if st.outline_width <= 0:
|
||||||
|
painter.setPen(color)
|
||||||
|
painter.setBrush(Qt.BrushStyle.NoBrush)
|
||||||
|
painter.drawText(x, y, text)
|
||||||
|
return
|
||||||
|
|
||||||
|
path = QPainterPath()
|
||||||
|
path.addText(x, y, font, text)
|
||||||
|
|
||||||
|
pen = QPen(QColor(st.outline_color))
|
||||||
|
pen.setWidth(st.outline_width * 2) # 절반이 안쪽으로 먹히므로 2배
|
||||||
|
pen.setJoinStyle(Qt.PenJoinStyle.RoundJoin)
|
||||||
|
pen.setCapStyle(Qt.PenCapStyle.RoundCap)
|
||||||
|
painter.setPen(pen)
|
||||||
|
painter.setBrush(Qt.BrushStyle.NoBrush)
|
||||||
|
painter.drawPath(path)
|
||||||
|
|
||||||
|
painter.setPen(Qt.PenStyle.NoPen)
|
||||||
|
painter.setBrush(color)
|
||||||
|
painter.drawPath(path)
|
||||||
|
|
||||||
|
|
||||||
|
class SubtitlePreview(QWidget):
|
||||||
|
"""설정 화면에서 쓰는 미리보기. 실제 자막과 같은 코드로 그린다."""
|
||||||
|
|
||||||
|
def __init__(self, style: SubtitleStyle, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.style_cfg = style
|
||||||
|
self.translated = ""
|
||||||
|
self.source = ""
|
||||||
|
self.setMinimumHeight(140)
|
||||||
|
|
||||||
|
def set_sample(self, translated: str, source: str) -> None:
|
||||||
|
self.translated = translated
|
||||||
|
self.source = source
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def paintEvent(self, event) -> None: # noqa: N802
|
||||||
|
painter = QPainter(self)
|
||||||
|
paint_subtitle(
|
||||||
|
painter, self.rect(), self.style_cfg, self.translated, self.source
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class SubtitleOverlay(QWidget):
|
||||||
|
"""항상 위에 떠 있는 자막 창."""
|
||||||
|
|
||||||
|
geometry_changed = Signal(int, int, int, int)
|
||||||
|
closed = Signal()
|
||||||
|
|
||||||
|
def __init__(self, style: SubtitleStyle, overlay: OverlayConfig) -> None:
|
||||||
|
super().__init__(None)
|
||||||
|
self.style_cfg = style
|
||||||
|
self.overlay_cfg = overlay
|
||||||
|
|
||||||
|
self._final_text = ""
|
||||||
|
self._partial_text = ""
|
||||||
|
self._source_text = ""
|
||||||
|
self._drag_from: QPoint | None = None
|
||||||
|
self._resize_from: QRect | None = None
|
||||||
|
|
||||||
|
self._fade = QTimer(self)
|
||||||
|
self._fade.setSingleShot(True)
|
||||||
|
self._fade.timeout.connect(self._clear_text)
|
||||||
|
|
||||||
|
self.setWindowFlags(
|
||||||
|
Qt.WindowType.FramelessWindowHint
|
||||||
|
| Qt.WindowType.Tool
|
||||||
|
| Qt.WindowType.WindowStaysOnTopHint
|
||||||
|
)
|
||||||
|
self.setAttribute(Qt.WidgetAttribute.WA_TranslucentBackground, True)
|
||||||
|
self.setMinimumSize(320, 90)
|
||||||
|
self.apply_config()
|
||||||
|
|
||||||
|
# --- 외부 API -------------------------------------------------------
|
||||||
|
def apply_config(self) -> None:
|
||||||
|
cfg = self.overlay_cfg
|
||||||
|
flags = self.windowFlags()
|
||||||
|
flags = (
|
||||||
|
flags | Qt.WindowType.WindowStaysOnTopHint
|
||||||
|
if cfg.always_on_top
|
||||||
|
else flags & ~Qt.WindowType.WindowStaysOnTopHint
|
||||||
|
)
|
||||||
|
self.setWindowFlags(flags)
|
||||||
|
self.setAttribute(
|
||||||
|
Qt.WidgetAttribute.WA_TransparentForMouseEvents, cfg.click_through
|
||||||
|
)
|
||||||
|
|
||||||
|
if cfg.x < 0 or cfg.y < 0:
|
||||||
|
screen = QApplication.primaryScreen()
|
||||||
|
area = screen.availableGeometry() if screen else QRect(0, 0, 1920, 1080)
|
||||||
|
cfg.width = min(cfg.width, area.width() - 80)
|
||||||
|
cfg.x = area.x() + (area.width() - cfg.width) // 2
|
||||||
|
cfg.y = area.y() + area.height() - cfg.height - 90
|
||||||
|
self.setGeometry(cfg.x, cfg.y, cfg.width, cfg.height)
|
||||||
|
if cfg.visible:
|
||||||
|
self.show()
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def set_line(self, translated: str, source: str = "", is_final: bool = True) -> None:
|
||||||
|
if is_final:
|
||||||
|
self._final_text = translated or source
|
||||||
|
self._partial_text = ""
|
||||||
|
self._source_text = source
|
||||||
|
else:
|
||||||
|
self._partial_text = source
|
||||||
|
self._restart_fade()
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def clear(self) -> None:
|
||||||
|
self._clear_text()
|
||||||
|
|
||||||
|
# --- 그리기 ---------------------------------------------------------
|
||||||
|
def paintEvent(self, event) -> None: # noqa: N802 - Qt 규약
|
||||||
|
painter = QPainter(self)
|
||||||
|
paint_subtitle(
|
||||||
|
painter,
|
||||||
|
self.rect(),
|
||||||
|
self.style_cfg,
|
||||||
|
self._final_text,
|
||||||
|
self._source_text,
|
||||||
|
self._partial_text,
|
||||||
|
)
|
||||||
|
if not (self._final_text or self._partial_text) and not self.overlay_cfg.locked:
|
||||||
|
self._draw_placeholder(painter)
|
||||||
|
self._draw_grip(painter)
|
||||||
|
|
||||||
|
def _draw_placeholder(self, painter) -> None:
|
||||||
|
painter.setPen(QColor(255, 255, 255, 90))
|
||||||
|
font = QFont(self.style_cfg.font_family, 13)
|
||||||
|
painter.setFont(font)
|
||||||
|
painter.drawText(
|
||||||
|
self.rect(),
|
||||||
|
Qt.AlignmentFlag.AlignCenter,
|
||||||
|
"자막이 여기에 표시됩니다\n드래그로 이동 · 휠로 크기 · 우클릭으로 메뉴",
|
||||||
|
)
|
||||||
|
|
||||||
|
def _draw_grip(self, painter) -> None:
|
||||||
|
if self.overlay_cfg.locked or self.overlay_cfg.click_through:
|
||||||
|
return
|
||||||
|
painter.setPen(QPen(QColor(255, 255, 255, 110), 2))
|
||||||
|
w, h = self.width(), self.height()
|
||||||
|
for offset in (5, 10, 15):
|
||||||
|
painter.drawLine(w - offset, h - 4, w - 4, h - offset)
|
||||||
|
|
||||||
|
# --- 마우스 ---------------------------------------------------------
|
||||||
|
def mousePressEvent(self, event) -> None: # noqa: N802
|
||||||
|
if self.overlay_cfg.locked or event.button() != Qt.MouseButton.LeftButton:
|
||||||
|
return
|
||||||
|
pos = event.position().toPoint()
|
||||||
|
if pos.x() > self.width() - GRIP and pos.y() > self.height() - GRIP:
|
||||||
|
self._resize_from = QRect(event.globalPosition().toPoint(), self.size())
|
||||||
|
else:
|
||||||
|
self._drag_from = event.globalPosition().toPoint() - self.pos()
|
||||||
|
|
||||||
|
def mouseMoveEvent(self, event) -> None: # noqa: N802
|
||||||
|
if self._resize_from is not None:
|
||||||
|
delta = event.globalPosition().toPoint() - self._resize_from.topLeft()
|
||||||
|
self.resize(
|
||||||
|
max(self.minimumWidth(), self._resize_from.width() + delta.x()),
|
||||||
|
max(self.minimumHeight(), self._resize_from.height() + delta.y()),
|
||||||
|
)
|
||||||
|
elif self._drag_from is not None:
|
||||||
|
self.move(event.globalPosition().toPoint() - self._drag_from)
|
||||||
|
|
||||||
|
def mouseReleaseEvent(self, event) -> None: # noqa: N802
|
||||||
|
if self._drag_from is not None or self._resize_from is not None:
|
||||||
|
self._drag_from = None
|
||||||
|
self._resize_from = None
|
||||||
|
g = self.geometry()
|
||||||
|
self.overlay_cfg.x, self.overlay_cfg.y = g.x(), g.y()
|
||||||
|
self.overlay_cfg.width, self.overlay_cfg.height = g.width(), g.height()
|
||||||
|
self.geometry_changed.emit(g.x(), g.y(), g.width(), g.height())
|
||||||
|
|
||||||
|
def wheelEvent(self, event) -> None: # noqa: N802
|
||||||
|
if self.overlay_cfg.locked:
|
||||||
|
return
|
||||||
|
step = 2 if event.angleDelta().y() > 0 else -2
|
||||||
|
self.style_cfg.font_size = max(12, min(120, self.style_cfg.font_size + step))
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def contextMenuEvent(self, event) -> None: # noqa: N802
|
||||||
|
menu = QMenu(self)
|
||||||
|
cfg = self.overlay_cfg
|
||||||
|
|
||||||
|
lock = QAction("위치 잠금", self, checkable=True, checked=cfg.locked)
|
||||||
|
lock.toggled.connect(self._set_locked)
|
||||||
|
menu.addAction(lock)
|
||||||
|
|
||||||
|
through = QAction("클릭 통과", self, checkable=True, checked=cfg.click_through)
|
||||||
|
through.toggled.connect(self._set_click_through)
|
||||||
|
menu.addAction(through)
|
||||||
|
|
||||||
|
top = QAction("항상 위에", self, checkable=True, checked=cfg.always_on_top)
|
||||||
|
top.toggled.connect(self._set_always_on_top)
|
||||||
|
menu.addAction(top)
|
||||||
|
|
||||||
|
menu.addSeparator()
|
||||||
|
menu.addAction("자막 지우기", self.clear)
|
||||||
|
menu.addAction("자막 창 숨기기", self.hide_overlay)
|
||||||
|
menu.exec(event.globalPos())
|
||||||
|
|
||||||
|
# --- 메뉴 동작 ------------------------------------------------------
|
||||||
|
def _set_locked(self, value: bool) -> None:
|
||||||
|
self.overlay_cfg.locked = value
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def _set_click_through(self, value: bool) -> None:
|
||||||
|
self.overlay_cfg.click_through = value
|
||||||
|
self.setAttribute(Qt.WidgetAttribute.WA_TransparentForMouseEvents, value)
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def _set_always_on_top(self, value: bool) -> None:
|
||||||
|
self.overlay_cfg.always_on_top = value
|
||||||
|
self.apply_config()
|
||||||
|
|
||||||
|
def hide_overlay(self) -> None:
|
||||||
|
self.overlay_cfg.visible = False
|
||||||
|
self.hide()
|
||||||
|
self.closed.emit()
|
||||||
|
|
||||||
|
def _restart_fade(self) -> None:
|
||||||
|
self._fade.stop()
|
||||||
|
if self.style_cfg.fade_out_ms > 0:
|
||||||
|
self._fade.start(self.style_cfg.fade_out_ms)
|
||||||
|
|
||||||
|
def _clear_text(self) -> None:
|
||||||
|
self._final_text = ""
|
||||||
|
self._partial_text = ""
|
||||||
|
self._source_text = ""
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
|
||||||
|
def _wrap(text: str, fm: QFontMetrics, width: int, max_lines: int) -> list[str]:
|
||||||
|
"""단어 단위 줄바꿈. 마지막 줄이 넘치면 앞부분을 버리고 최신 내용을 남긴다."""
|
||||||
|
if width <= 0 or not text:
|
||||||
|
return [text] if text else []
|
||||||
|
lines: list[str] = []
|
||||||
|
current = ""
|
||||||
|
for word in text.split():
|
||||||
|
candidate = f"{current} {word}".strip()
|
||||||
|
if fm.horizontalAdvance(candidate) <= width or not current:
|
||||||
|
current = candidate
|
||||||
|
else:
|
||||||
|
lines.append(current)
|
||||||
|
current = word
|
||||||
|
if current:
|
||||||
|
lines.append(current)
|
||||||
|
# 실시간 자막에서는 앞이 아니라 뒤(최신)를 보여주는 게 맞다.
|
||||||
|
return lines[-max_lines:] if max_lines > 0 else lines
|
||||||
7
src/hearo/ui/pages/__init__.py
Normal file
7
src/hearo/ui/pages/__init__.py
Normal file
@@ -0,0 +1,7 @@
|
|||||||
|
from .glossary_page import GlossaryPage
|
||||||
|
from .home import HomePage
|
||||||
|
from .models_page import ModelsPage
|
||||||
|
from .settings_page import SettingsPage
|
||||||
|
from .subtitle_page import SubtitlePage
|
||||||
|
|
||||||
|
__all__ = ["GlossaryPage", "HomePage", "ModelsPage", "SettingsPage", "SubtitlePage"]
|
||||||
256
src/hearo/ui/pages/glossary_page.py
Normal file
256
src/hearo/ui/pages/glossary_page.py
Normal file
@@ -0,0 +1,256 @@
|
|||||||
|
"""용어집 — 게임·방송 고유명사를 원하는 번역으로 고정한다."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import csv
|
||||||
|
from pathlib import Path
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt, Signal
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QAbstractItemView,
|
||||||
|
QCheckBox,
|
||||||
|
QFileDialog,
|
||||||
|
QHBoxLayout,
|
||||||
|
QHeaderView,
|
||||||
|
QLabel,
|
||||||
|
QLineEdit,
|
||||||
|
QMessageBox,
|
||||||
|
QPushButton,
|
||||||
|
QTableWidget,
|
||||||
|
QTableWidgetItem,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ...config import AppConfig
|
||||||
|
from ...constants import LANGUAGE_CODES, LANGUAGES, user_data_dir
|
||||||
|
from ...models import Glossary, GlossaryEntry
|
||||||
|
from ..theme import SPACING
|
||||||
|
from ..widgets import Card, PageHeader
|
||||||
|
|
||||||
|
COLUMNS = ["원문 용어", *[f"{LANGUAGES[c]['label']} 역어" for c in LANGUAGE_CODES], "메모"]
|
||||||
|
|
||||||
|
|
||||||
|
class GlossaryPage(QWidget):
|
||||||
|
glossary_changed = Signal()
|
||||||
|
|
||||||
|
def __init__(self, config: AppConfig, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.config = config
|
||||||
|
self.path = Path(config.glossary.path or user_data_dir() / "glossary.json")
|
||||||
|
self.glossary = Glossary.load(self.path, config.glossary.case_sensitive)
|
||||||
|
|
||||||
|
root = QVBoxLayout(self)
|
||||||
|
root.setContentsMargins(0, 0, 0, 0)
|
||||||
|
root.setSpacing(SPACING + 4)
|
||||||
|
root.addWidget(
|
||||||
|
PageHeader(
|
||||||
|
"용어집",
|
||||||
|
"여기 등록한 단어는 번역 모델이 마음대로 바꾸지 못하고 지정한 역어로 고정됩니다. "
|
||||||
|
"모델을 추가로 학습시키지 않아도 바로 적용됩니다.",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
root.addWidget(self._build_options_card())
|
||||||
|
root.addWidget(self._build_table_card(), 1)
|
||||||
|
self._reload_table()
|
||||||
|
|
||||||
|
# --- 구성 -----------------------------------------------------------
|
||||||
|
def _build_options_card(self) -> Card:
|
||||||
|
card = Card()
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(10)
|
||||||
|
|
||||||
|
self.enabled_check = QCheckBox("용어집 사용")
|
||||||
|
self.enabled_check.setChecked(self.config.glossary.enabled)
|
||||||
|
self.enabled_check.toggled.connect(self._on_enabled)
|
||||||
|
layout.addWidget(self.enabled_check)
|
||||||
|
|
||||||
|
self.case_check = QCheckBox("대소문자 구분")
|
||||||
|
self.case_check.setChecked(self.config.glossary.case_sensitive)
|
||||||
|
self.case_check.toggled.connect(self._on_case)
|
||||||
|
layout.addWidget(self.case_check)
|
||||||
|
layout.addStretch(1)
|
||||||
|
|
||||||
|
for label, slot in (
|
||||||
|
("CSV 가져오기", self._import_csv),
|
||||||
|
("CSV 내보내기", self._export_csv),
|
||||||
|
):
|
||||||
|
button = QPushButton(label)
|
||||||
|
button.clicked.connect(slot)
|
||||||
|
layout.addWidget(button)
|
||||||
|
|
||||||
|
card.add(row)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_table_card(self) -> Card:
|
||||||
|
card = Card(
|
||||||
|
"등록된 용어",
|
||||||
|
"역어를 비워두면 그 언어에서는 규칙을 적용하지 않습니다. "
|
||||||
|
"CSV 형식: 원문,한국어,English,日本語,中文,메모",
|
||||||
|
)
|
||||||
|
|
||||||
|
search_row = QWidget()
|
||||||
|
search_layout = QHBoxLayout(search_row)
|
||||||
|
search_layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
search_layout.setSpacing(10)
|
||||||
|
self.search = QLineEdit()
|
||||||
|
self.search.setPlaceholderText("용어 검색…")
|
||||||
|
self.search.textChanged.connect(self._filter)
|
||||||
|
search_layout.addWidget(self.search, 1)
|
||||||
|
for label, slot in (("행 추가", self._add_row), ("선택 삭제", self._delete_rows)):
|
||||||
|
button = QPushButton(label)
|
||||||
|
button.clicked.connect(slot)
|
||||||
|
search_layout.addWidget(button)
|
||||||
|
card.add(search_row)
|
||||||
|
|
||||||
|
self.table = QTableWidget(0, len(COLUMNS))
|
||||||
|
self.table.setHorizontalHeaderLabels(COLUMNS)
|
||||||
|
self.table.verticalHeader().setVisible(False)
|
||||||
|
self.table.setSelectionBehavior(QAbstractItemView.SelectionBehavior.SelectRows)
|
||||||
|
self.table.horizontalHeader().setSectionResizeMode(
|
||||||
|
QHeaderView.ResizeMode.Stretch
|
||||||
|
)
|
||||||
|
self.table.itemChanged.connect(self._on_item_changed)
|
||||||
|
card.add(self.table)
|
||||||
|
|
||||||
|
self.count_label = card.add(
|
||||||
|
_hint(f"{len(self.glossary)}개 등록됨")
|
||||||
|
)
|
||||||
|
return card
|
||||||
|
|
||||||
|
# --- 데이터 ---------------------------------------------------------
|
||||||
|
def _reload_table(self) -> None:
|
||||||
|
self.table.blockSignals(True)
|
||||||
|
self.table.setRowCount(0)
|
||||||
|
for entry in self.glossary.entries:
|
||||||
|
self._append_row(entry)
|
||||||
|
self.table.blockSignals(False)
|
||||||
|
self.count_label.setText(f"{len(self.glossary)}개 등록됨")
|
||||||
|
|
||||||
|
def _append_row(self, entry: GlossaryEntry) -> None:
|
||||||
|
row = self.table.rowCount()
|
||||||
|
self.table.insertRow(row)
|
||||||
|
values = [entry.source, *[entry.target_for(c) for c in LANGUAGE_CODES], entry.note]
|
||||||
|
for col, value in enumerate(values):
|
||||||
|
self.table.setItem(row, col, QTableWidgetItem(value))
|
||||||
|
|
||||||
|
def _collect(self) -> list[GlossaryEntry]:
|
||||||
|
entries: list[GlossaryEntry] = []
|
||||||
|
for row in range(self.table.rowCount()):
|
||||||
|
source_item = self.table.item(row, 0)
|
||||||
|
source = source_item.text().strip() if source_item else ""
|
||||||
|
if not source:
|
||||||
|
continue
|
||||||
|
targets: dict[str, str] = {}
|
||||||
|
for i, code in enumerate(LANGUAGE_CODES, start=1):
|
||||||
|
item = self.table.item(row, i)
|
||||||
|
text = item.text().strip() if item else ""
|
||||||
|
if text:
|
||||||
|
targets[code] = text
|
||||||
|
note_item = self.table.item(row, len(COLUMNS) - 1)
|
||||||
|
entries.append(
|
||||||
|
GlossaryEntry(
|
||||||
|
source=source,
|
||||||
|
targets=targets,
|
||||||
|
note=note_item.text().strip() if note_item else "",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
return entries
|
||||||
|
|
||||||
|
def _save(self) -> None:
|
||||||
|
self.glossary = Glossary(self._collect(), self.config.glossary.case_sensitive)
|
||||||
|
self.glossary.save(self.path)
|
||||||
|
self.count_label.setText(f"{len(self.glossary)}개 등록됨")
|
||||||
|
self.glossary_changed.emit()
|
||||||
|
|
||||||
|
# --- 동작 -----------------------------------------------------------
|
||||||
|
def _on_item_changed(self, *_) -> None:
|
||||||
|
self._save()
|
||||||
|
|
||||||
|
def _on_enabled(self, value: bool) -> None:
|
||||||
|
self.config.glossary.enabled = value
|
||||||
|
self.glossary_changed.emit()
|
||||||
|
|
||||||
|
def _on_case(self, value: bool) -> None:
|
||||||
|
self.config.glossary.case_sensitive = value
|
||||||
|
self._save()
|
||||||
|
|
||||||
|
def _add_row(self) -> None:
|
||||||
|
self.table.blockSignals(True)
|
||||||
|
self._append_row(GlossaryEntry(source=""))
|
||||||
|
self.table.blockSignals(False)
|
||||||
|
self.table.editItem(self.table.item(self.table.rowCount() - 1, 0))
|
||||||
|
|
||||||
|
def _delete_rows(self) -> None:
|
||||||
|
rows = sorted({i.row() for i in self.table.selectedIndexes()}, reverse=True)
|
||||||
|
if not rows:
|
||||||
|
return
|
||||||
|
self.table.blockSignals(True)
|
||||||
|
for row in rows:
|
||||||
|
self.table.removeRow(row)
|
||||||
|
self.table.blockSignals(False)
|
||||||
|
self._save()
|
||||||
|
|
||||||
|
def _filter(self, text: str) -> None:
|
||||||
|
needle = text.strip().casefold()
|
||||||
|
for row in range(self.table.rowCount()):
|
||||||
|
haystack = " ".join(
|
||||||
|
(self.table.item(row, c).text() if self.table.item(row, c) else "")
|
||||||
|
for c in range(self.table.columnCount())
|
||||||
|
).casefold()
|
||||||
|
self.table.setRowHidden(row, bool(needle) and needle not in haystack)
|
||||||
|
|
||||||
|
def _import_csv(self) -> None:
|
||||||
|
path, _ = QFileDialog.getOpenFileName(self, "CSV 가져오기", "", "CSV (*.csv)")
|
||||||
|
if not path:
|
||||||
|
return
|
||||||
|
try:
|
||||||
|
with open(path, encoding="utf-8-sig", newline="") as fh:
|
||||||
|
rows = [r for r in csv.reader(fh) if r and r[0].strip()]
|
||||||
|
except OSError as exc:
|
||||||
|
QMessageBox.warning(self, "가져오기 실패", str(exc))
|
||||||
|
return
|
||||||
|
|
||||||
|
# 첫 줄이 헤더처럼 보이면 건너뛴다.
|
||||||
|
if rows and rows[0][0].strip() in ("원문", "원문 용어", "source"):
|
||||||
|
rows = rows[1:]
|
||||||
|
|
||||||
|
added = 0
|
||||||
|
for row in rows:
|
||||||
|
cells = [c.strip() for c in row] + [""] * (len(COLUMNS) - len(row))
|
||||||
|
targets = {
|
||||||
|
code: cells[i]
|
||||||
|
for i, code in enumerate(LANGUAGE_CODES, start=1)
|
||||||
|
if cells[i]
|
||||||
|
}
|
||||||
|
self.glossary.add(
|
||||||
|
GlossaryEntry(source=cells[0], targets=targets, note=cells[len(COLUMNS) - 1])
|
||||||
|
)
|
||||||
|
added += 1
|
||||||
|
self.glossary.save(self.path)
|
||||||
|
self._reload_table()
|
||||||
|
self.glossary_changed.emit()
|
||||||
|
QMessageBox.information(self, "가져오기 완료", f"{added}개 용어를 반영했습니다.")
|
||||||
|
|
||||||
|
def _export_csv(self) -> None:
|
||||||
|
path, _ = QFileDialog.getSaveFileName(
|
||||||
|
self, "CSV 내보내기", "glossary.csv", "CSV (*.csv)"
|
||||||
|
)
|
||||||
|
if not path:
|
||||||
|
return
|
||||||
|
with open(path, "w", encoding="utf-8-sig", newline="") as fh:
|
||||||
|
writer = csv.writer(fh)
|
||||||
|
writer.writerow(COLUMNS)
|
||||||
|
for entry in self.glossary.entries:
|
||||||
|
writer.writerow(
|
||||||
|
[entry.source, *[entry.target_for(c) for c in LANGUAGE_CODES], entry.note]
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def _hint(text: str) -> QLabel:
|
||||||
|
label = QLabel(text)
|
||||||
|
label.setObjectName("Hint")
|
||||||
|
label.setAlignment(Qt.AlignmentFlag.AlignRight)
|
||||||
|
return label
|
||||||
226
src/hearo/ui/pages/home.py
Normal file
226
src/hearo/ui/pages/home.py
Normal file
@@ -0,0 +1,226 @@
|
|||||||
|
"""홈 — 대상 선택, 언어, 시작/정지, 실시간 로그."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt, Signal
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QComboBox,
|
||||||
|
QHBoxLayout,
|
||||||
|
QLabel,
|
||||||
|
QListWidget,
|
||||||
|
QListWidgetItem,
|
||||||
|
QPushButton,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ...audio import capture_capabilities, list_sources
|
||||||
|
from ...config import AppConfig
|
||||||
|
from ...constants import LANGUAGES
|
||||||
|
from ...core.events import EngineState, EngineStatus, TranslationLine
|
||||||
|
from ..theme import SPACING, palette
|
||||||
|
from ..widgets import Card, LevelMeter, PageHeader, StatusPill
|
||||||
|
|
||||||
|
MAX_LOG_ROWS = 300
|
||||||
|
|
||||||
|
|
||||||
|
class HomePage(QWidget):
|
||||||
|
start_requested = Signal()
|
||||||
|
stop_requested = Signal()
|
||||||
|
config_changed = Signal()
|
||||||
|
|
||||||
|
def __init__(self, config: AppConfig, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.config = config
|
||||||
|
self._sources: list = []
|
||||||
|
|
||||||
|
root = QVBoxLayout(self)
|
||||||
|
root.setContentsMargins(0, 0, 0, 0)
|
||||||
|
root.setSpacing(SPACING + 4)
|
||||||
|
root.addWidget(
|
||||||
|
PageHeader("실시간 번역", "소리를 받아올 프로그램과 번역할 언어를 고르세요.")
|
||||||
|
)
|
||||||
|
|
||||||
|
root.addWidget(self._build_source_card())
|
||||||
|
root.addWidget(self._build_control_card())
|
||||||
|
root.addWidget(self._build_log_card(), 1)
|
||||||
|
|
||||||
|
self.refresh_sources()
|
||||||
|
|
||||||
|
# --- 구성 -----------------------------------------------------------
|
||||||
|
def _build_source_card(self) -> Card:
|
||||||
|
card = Card("소리 받아올 곳", "")
|
||||||
|
self.capability_label = QLabel()
|
||||||
|
self.capability_label.setObjectName("Hint")
|
||||||
|
self.capability_label.setWordWrap(True)
|
||||||
|
card.add(self.capability_label)
|
||||||
|
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(10)
|
||||||
|
|
||||||
|
self.source_combo = QComboBox()
|
||||||
|
self.source_combo.setMinimumWidth(360)
|
||||||
|
self.source_combo.currentIndexChanged.connect(self._on_source_changed)
|
||||||
|
layout.addWidget(self.source_combo, 1)
|
||||||
|
|
||||||
|
refresh = QPushButton("새로고침")
|
||||||
|
refresh.clicked.connect(self.refresh_sources)
|
||||||
|
layout.addWidget(refresh)
|
||||||
|
card.add(row)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_control_card(self) -> Card:
|
||||||
|
card = Card()
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(12)
|
||||||
|
|
||||||
|
self.source_lang = QComboBox()
|
||||||
|
self.source_lang.addItem("자동 감지", "auto")
|
||||||
|
for code, meta in LANGUAGES.items():
|
||||||
|
self.source_lang.addItem(meta["label"], code)
|
||||||
|
self.source_lang.setCurrentIndex(
|
||||||
|
max(0, self.source_lang.findData(self.config.models.source_lang))
|
||||||
|
)
|
||||||
|
self.source_lang.currentIndexChanged.connect(self._on_lang_changed)
|
||||||
|
|
||||||
|
self.target_lang = QComboBox()
|
||||||
|
for code, meta in LANGUAGES.items():
|
||||||
|
self.target_lang.addItem(meta["label"], code)
|
||||||
|
self.target_lang.setCurrentIndex(
|
||||||
|
max(0, self.target_lang.findData(self.config.models.target_lang))
|
||||||
|
)
|
||||||
|
self.target_lang.currentIndexChanged.connect(self._on_lang_changed)
|
||||||
|
|
||||||
|
layout.addWidget(QLabel("원본"))
|
||||||
|
layout.addWidget(self.source_lang)
|
||||||
|
arrow = QLabel("→")
|
||||||
|
arrow.setStyleSheet(f"color: {palette('dark').accent}; font-size: 18px;")
|
||||||
|
layout.addWidget(arrow)
|
||||||
|
layout.addWidget(QLabel("번역"))
|
||||||
|
layout.addWidget(self.target_lang)
|
||||||
|
layout.addStretch(1)
|
||||||
|
|
||||||
|
self.status_pill = StatusPill()
|
||||||
|
layout.addWidget(self.status_pill)
|
||||||
|
|
||||||
|
self.toggle_button = QPushButton("번역 시작")
|
||||||
|
self.toggle_button.setObjectName("Primary")
|
||||||
|
self.toggle_button.setMinimumWidth(140)
|
||||||
|
self.toggle_button.clicked.connect(self._on_toggle)
|
||||||
|
layout.addWidget(self.toggle_button)
|
||||||
|
|
||||||
|
card.add(row)
|
||||||
|
self.level_meter = LevelMeter()
|
||||||
|
card.add(self.level_meter)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_log_card(self) -> Card:
|
||||||
|
card = Card("최근 자막", "확정된 문장이 위로 쌓입니다. 지연 시간도 함께 보여줍니다.")
|
||||||
|
self.log_list = QListWidget()
|
||||||
|
self.log_list.setWordWrap(True)
|
||||||
|
self.log_list.setAlternatingRowColors(False)
|
||||||
|
card.add(self.log_list)
|
||||||
|
return card
|
||||||
|
|
||||||
|
# --- 동작 -----------------------------------------------------------
|
||||||
|
def refresh_sources(self) -> None:
|
||||||
|
caps = capture_capabilities()
|
||||||
|
self._sources = list_sources()
|
||||||
|
|
||||||
|
self.source_combo.blockSignals(True)
|
||||||
|
self.source_combo.clear()
|
||||||
|
for src in self._sources:
|
||||||
|
prefix = "🎮" if src.kind == "process" else "🔊"
|
||||||
|
self.source_combo.addItem(f"{prefix} {src.label} — {src.detail}", src)
|
||||||
|
if not self._sources:
|
||||||
|
self.source_combo.addItem("사용 가능한 오디오 대상이 없습니다", None)
|
||||||
|
self._select_saved_source()
|
||||||
|
self.source_combo.blockSignals(False)
|
||||||
|
self._on_source_changed()
|
||||||
|
|
||||||
|
if caps["process"]:
|
||||||
|
msg = "프로그램별 캡처 사용 가능 — 선택한 프로그램의 소리만 정확히 받아옵니다."
|
||||||
|
elif caps["device"]:
|
||||||
|
msg = (
|
||||||
|
"프로그램별 캡처 보조 프로그램(hearo_capture.exe)이 없어 "
|
||||||
|
"출력 장치 전체 소리를 받습니다. native/process_loopback/build.ps1 로 빌드하면 "
|
||||||
|
"프로그램 단위로 분리됩니다."
|
||||||
|
)
|
||||||
|
else:
|
||||||
|
msg = "이 환경에서는 오디오 캡처를 할 수 없습니다. Windows에서 실행하세요."
|
||||||
|
self.capability_label.setText(msg)
|
||||||
|
|
||||||
|
def _select_saved_source(self) -> None:
|
||||||
|
cfg = self.config.audio
|
||||||
|
for i, src in enumerate(self._sources):
|
||||||
|
if cfg.backend == "process" and src.kind == "process" and src.pid == cfg.target_pid:
|
||||||
|
self.source_combo.setCurrentIndex(i)
|
||||||
|
return
|
||||||
|
if cfg.backend == "device" and src.kind == "device" and src.device_index == cfg.device_index:
|
||||||
|
self.source_combo.setCurrentIndex(i)
|
||||||
|
return
|
||||||
|
|
||||||
|
def _on_source_changed(self, *_) -> None:
|
||||||
|
src = self.source_combo.currentData()
|
||||||
|
cfg = self.config.audio
|
||||||
|
if src is None:
|
||||||
|
return
|
||||||
|
if src.kind == "process":
|
||||||
|
cfg.backend = "process"
|
||||||
|
cfg.target_pid = src.pid
|
||||||
|
cfg.target_process_name = src.label
|
||||||
|
else:
|
||||||
|
cfg.backend = "device"
|
||||||
|
cfg.device_index = src.device_index
|
||||||
|
cfg.target_pid = 0
|
||||||
|
self.config_changed.emit()
|
||||||
|
|
||||||
|
def _on_lang_changed(self, *_) -> None:
|
||||||
|
self.config.models.source_lang = self.source_lang.currentData()
|
||||||
|
self.config.models.target_lang = self.target_lang.currentData()
|
||||||
|
self.config_changed.emit()
|
||||||
|
|
||||||
|
def _on_toggle(self) -> None:
|
||||||
|
if self.toggle_button.property("running"):
|
||||||
|
self.stop_requested.emit()
|
||||||
|
else:
|
||||||
|
self.start_requested.emit()
|
||||||
|
|
||||||
|
# --- 엔진 연동 ------------------------------------------------------
|
||||||
|
def on_status(self, status: EngineStatus) -> None:
|
||||||
|
running = status.state in (EngineState.RUNNING, EngineState.LOADING)
|
||||||
|
self.toggle_button.setProperty("running", running)
|
||||||
|
self.toggle_button.setText("번역 정지" if running else "번역 시작")
|
||||||
|
self.toggle_button.setObjectName("Danger" if running else "Primary")
|
||||||
|
self.toggle_button.style().polish(self.toggle_button)
|
||||||
|
|
||||||
|
kind = {
|
||||||
|
EngineState.RUNNING: "ok",
|
||||||
|
EngineState.LOADING: "busy",
|
||||||
|
EngineState.STOPPING: "busy",
|
||||||
|
EngineState.ERROR: "error",
|
||||||
|
}.get(status.state, "idle")
|
||||||
|
self.status_pill.set_status(status.message or status.state.value, kind)
|
||||||
|
self.level_meter.set_level(status.level)
|
||||||
|
|
||||||
|
for widget in (self.source_combo, self.source_lang, self.target_lang):
|
||||||
|
widget.setEnabled(not running)
|
||||||
|
|
||||||
|
def on_line(self, line: TranslationLine) -> None:
|
||||||
|
if not line.is_final:
|
||||||
|
return
|
||||||
|
item = QListWidgetItem(
|
||||||
|
f"{line.translated_text}\n"
|
||||||
|
f" {line.source_text} · {line.latency_s:.1f}초 "
|
||||||
|
f"(인식 {line.asr_ms}ms · 번역 {line.mt_ms}ms)"
|
||||||
|
)
|
||||||
|
item.setTextAlignment(Qt.AlignmentFlag.AlignLeft | Qt.AlignmentFlag.AlignVCenter)
|
||||||
|
item.setToolTip(line.source_text)
|
||||||
|
self.log_list.insertItem(0, item)
|
||||||
|
while self.log_list.count() > MAX_LOG_ROWS:
|
||||||
|
self.log_list.takeItem(self.log_list.count() - 1)
|
||||||
|
self.log_list.scrollToTop()
|
||||||
221
src/hearo/ui/pages/models_page.py
Normal file
221
src/hearo/ui/pages/models_page.py
Normal file
@@ -0,0 +1,221 @@
|
|||||||
|
"""모델 — 5단계 품질 티어 선택과 GPU 상태."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt, Signal
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QButtonGroup,
|
||||||
|
QComboBox,
|
||||||
|
QFrame,
|
||||||
|
QHBoxLayout,
|
||||||
|
QLabel,
|
||||||
|
QLineEdit,
|
||||||
|
QPushButton,
|
||||||
|
QRadioButton,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ...config import AppConfig
|
||||||
|
from ...models import Tier, detect_gpu, ordered_tiers
|
||||||
|
from ..theme import SPACING, palette
|
||||||
|
from ..widgets import Card, PageHeader
|
||||||
|
|
||||||
|
|
||||||
|
class TierCard(QFrame):
|
||||||
|
"""티어 하나를 라디오 버튼 카드로."""
|
||||||
|
|
||||||
|
def __init__(self, tier: Tier, gpu_vram_gb: float, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.tier = tier
|
||||||
|
self.setObjectName("Card")
|
||||||
|
p = palette("dark")
|
||||||
|
|
||||||
|
layout = QHBoxLayout(self)
|
||||||
|
layout.setContentsMargins(16, 14, 16, 14)
|
||||||
|
layout.setSpacing(14)
|
||||||
|
|
||||||
|
self.radio = QRadioButton()
|
||||||
|
layout.addWidget(self.radio, 0, Qt.AlignmentFlag.AlignTop)
|
||||||
|
|
||||||
|
text = QVBoxLayout()
|
||||||
|
text.setSpacing(3)
|
||||||
|
|
||||||
|
head = QHBoxLayout()
|
||||||
|
head.setSpacing(8)
|
||||||
|
name = QLabel(f"{tier.order}. {tier.name}")
|
||||||
|
name.setStyleSheet("font-size: 16px; font-weight: 700;")
|
||||||
|
head.addWidget(name)
|
||||||
|
if tier.recommended:
|
||||||
|
head.addWidget(_badge("추천", p.accent))
|
||||||
|
if tier.best_after_finetune:
|
||||||
|
head.addWidget(_badge("추가학습 최적", p.success))
|
||||||
|
if gpu_vram_gb and gpu_vram_gb < tier.min_vram_gb:
|
||||||
|
head.addWidget(_badge(f"VRAM {tier.min_vram_gb:g}GB 필요", p.warning))
|
||||||
|
head.addStretch(1)
|
||||||
|
text.addLayout(head)
|
||||||
|
|
||||||
|
tagline = QLabel(tier.tagline)
|
||||||
|
tagline.setObjectName("Subtitle")
|
||||||
|
text.addWidget(tagline)
|
||||||
|
|
||||||
|
spec = QLabel(
|
||||||
|
f"예상 지연 {tier.approx_latency_s:g}초 · 권장 VRAM {tier.min_vram_gb:g}GB "
|
||||||
|
f"· 내려받기 약 {tier.download_gb:g}GB"
|
||||||
|
)
|
||||||
|
spec.setObjectName("Hint")
|
||||||
|
text.addWidget(spec)
|
||||||
|
|
||||||
|
if tier.notes:
|
||||||
|
note = QLabel(tier.notes)
|
||||||
|
note.setObjectName("Hint")
|
||||||
|
note.setWordWrap(True)
|
||||||
|
text.addWidget(note)
|
||||||
|
|
||||||
|
layout.addLayout(text, 1)
|
||||||
|
|
||||||
|
def set_selected(self, selected: bool) -> None:
|
||||||
|
self.setObjectName("CardSelected" if selected else "Card")
|
||||||
|
self.style().unpolish(self)
|
||||||
|
self.style().polish(self)
|
||||||
|
|
||||||
|
|
||||||
|
class ModelsPage(QWidget):
|
||||||
|
tier_changed = Signal(str)
|
||||||
|
config_changed = Signal()
|
||||||
|
|
||||||
|
def __init__(self, config: AppConfig, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.config = config
|
||||||
|
self._cards: list[TierCard] = []
|
||||||
|
|
||||||
|
root = QVBoxLayout(self)
|
||||||
|
root.setContentsMargins(0, 0, 0, 0)
|
||||||
|
root.setSpacing(SPACING + 4)
|
||||||
|
root.addWidget(
|
||||||
|
PageHeader(
|
||||||
|
"모델",
|
||||||
|
"속도 우선(1)부터 품질 우선(5)까지 다섯 단계입니다. "
|
||||||
|
"모델은 처음 사용할 때 자동으로 내려받습니다.",
|
||||||
|
)
|
||||||
|
)
|
||||||
|
|
||||||
|
root.addWidget(self._build_gpu_card())
|
||||||
|
|
||||||
|
self._group = QButtonGroup(self)
|
||||||
|
self._group.setExclusive(True)
|
||||||
|
for tier in ordered_tiers():
|
||||||
|
card = TierCard(tier, self._gpu_vram_gb)
|
||||||
|
self._group.addButton(card.radio, tier.order)
|
||||||
|
card.radio.toggled.connect(
|
||||||
|
lambda checked, t=tier: self._on_tier_selected(t) if checked else None
|
||||||
|
)
|
||||||
|
self._cards.append(card)
|
||||||
|
root.addWidget(card)
|
||||||
|
|
||||||
|
root.addWidget(self._build_lora_card())
|
||||||
|
root.addStretch(1)
|
||||||
|
self._apply_selection(config.models.tier)
|
||||||
|
|
||||||
|
# --- 구성 -----------------------------------------------------------
|
||||||
|
def _build_gpu_card(self) -> Card:
|
||||||
|
gpu = detect_gpu()
|
||||||
|
self._gpu_vram_gb = gpu.total_vram_mb / 1024 if gpu.available else 0.0
|
||||||
|
p = palette("dark")
|
||||||
|
|
||||||
|
card = Card("실행 환경")
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(12)
|
||||||
|
|
||||||
|
if gpu.available:
|
||||||
|
text = (
|
||||||
|
f"GPU: {gpu.name} · VRAM {gpu.total_vram_mb / 1024:.1f}GB "
|
||||||
|
f"(여유 {gpu.free_vram_mb / 1024:.1f}GB)"
|
||||||
|
)
|
||||||
|
color = p.success
|
||||||
|
else:
|
||||||
|
text = f"GPU를 쓸 수 없습니다 — {gpu.reason} (CPU로도 동작하지만 많이 느립니다)"
|
||||||
|
color = p.warning
|
||||||
|
info = QLabel(text)
|
||||||
|
info.setStyleSheet(f"color: {color}; font-weight: 600;")
|
||||||
|
info.setWordWrap(True)
|
||||||
|
layout.addWidget(info, 1)
|
||||||
|
|
||||||
|
self.device_combo = QComboBox()
|
||||||
|
self.device_combo.addItem("GPU (CUDA)", "cuda")
|
||||||
|
self.device_combo.addItem("CPU", "cpu")
|
||||||
|
self.device_combo.setCurrentIndex(
|
||||||
|
max(0, self.device_combo.findData(self.config.models.device))
|
||||||
|
)
|
||||||
|
self.device_combo.setEnabled(gpu.available)
|
||||||
|
self.device_combo.currentIndexChanged.connect(self._on_device_changed)
|
||||||
|
layout.addWidget(self.device_combo)
|
||||||
|
|
||||||
|
card.add(row)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_lora_card(self) -> Card:
|
||||||
|
card = Card(
|
||||||
|
"추가학습 어댑터 (선택)",
|
||||||
|
"게임·방송 용어로 따로 학습시킨 LoRA 어댑터 폴더를 지정하면 번역 모델에 얹습니다. "
|
||||||
|
"scripts/finetune_mt.py 로 만들 수 있습니다.",
|
||||||
|
)
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(10)
|
||||||
|
|
||||||
|
self.lora_edit = QLineEdit(self.config.models.lora_adapter_path)
|
||||||
|
self.lora_edit.setPlaceholderText("비워두면 사용하지 않습니다")
|
||||||
|
self.lora_edit.editingFinished.connect(self._on_lora_changed)
|
||||||
|
layout.addWidget(self.lora_edit, 1)
|
||||||
|
|
||||||
|
browse = QPushButton("폴더 선택")
|
||||||
|
browse.clicked.connect(self._browse_lora)
|
||||||
|
layout.addWidget(browse)
|
||||||
|
|
||||||
|
card.add(row)
|
||||||
|
return card
|
||||||
|
|
||||||
|
# --- 동작 -----------------------------------------------------------
|
||||||
|
def _apply_selection(self, tier_key: str) -> None:
|
||||||
|
for card in self._cards:
|
||||||
|
selected = card.tier.key == tier_key
|
||||||
|
card.radio.setChecked(selected)
|
||||||
|
card.set_selected(selected)
|
||||||
|
|
||||||
|
def _on_tier_selected(self, tier: Tier) -> None:
|
||||||
|
if self.config.models.tier == tier.key:
|
||||||
|
return
|
||||||
|
self.config.models.tier = tier.key
|
||||||
|
for card in self._cards:
|
||||||
|
card.set_selected(card.tier.key == tier.key)
|
||||||
|
self.tier_changed.emit(tier.key)
|
||||||
|
self.config_changed.emit()
|
||||||
|
|
||||||
|
def _on_device_changed(self) -> None:
|
||||||
|
self.config.models.device = self.device_combo.currentData()
|
||||||
|
self.config_changed.emit()
|
||||||
|
|
||||||
|
def _on_lora_changed(self) -> None:
|
||||||
|
self.config.models.lora_adapter_path = self.lora_edit.text().strip()
|
||||||
|
self.config_changed.emit()
|
||||||
|
|
||||||
|
def _browse_lora(self) -> None:
|
||||||
|
from PySide6.QtWidgets import QFileDialog
|
||||||
|
|
||||||
|
path = QFileDialog.getExistingDirectory(self, "LoRA 어댑터 폴더 선택")
|
||||||
|
if path:
|
||||||
|
self.lora_edit.setText(path)
|
||||||
|
self._on_lora_changed()
|
||||||
|
|
||||||
|
|
||||||
|
def _badge(text: str, color: str) -> QLabel:
|
||||||
|
label = QLabel(text)
|
||||||
|
label.setStyleSheet(
|
||||||
|
f"color: {color}; border: 1px solid {color}; border-radius: 9px;"
|
||||||
|
f"padding: 1px 8px; font-size: 11px; font-weight: 700;"
|
||||||
|
)
|
||||||
|
return label
|
||||||
176
src/hearo/ui/pages/settings_page.py
Normal file
176
src/hearo/ui/pages/settings_page.py
Normal file
@@ -0,0 +1,176 @@
|
|||||||
|
"""설정 — 인식 민감도, 응답 속도, 기타 동작."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt, Signal
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QCheckBox,
|
||||||
|
QComboBox,
|
||||||
|
QHBoxLayout,
|
||||||
|
QLabel,
|
||||||
|
QPushButton,
|
||||||
|
QSlider,
|
||||||
|
QSpinBox,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ...config import AppConfig
|
||||||
|
from ...constants import APP_NAME, APP_SLOGAN, APP_VERSION, user_data_dir
|
||||||
|
from ..theme import SPACING
|
||||||
|
from ..widgets import Card, PageHeader
|
||||||
|
|
||||||
|
|
||||||
|
class SettingsPage(QWidget):
|
||||||
|
config_changed = Signal()
|
||||||
|
#: 오디오 파이프라인 설정이 바뀌어 엔진 재시작이 필요할 때
|
||||||
|
restart_needed = Signal()
|
||||||
|
|
||||||
|
def __init__(self, config: AppConfig, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.config = config
|
||||||
|
|
||||||
|
root = QVBoxLayout(self)
|
||||||
|
root.setContentsMargins(0, 0, 0, 0)
|
||||||
|
root.setSpacing(SPACING + 4)
|
||||||
|
root.addWidget(PageHeader("설정", "인식 민감도와 응답 속도를 조절합니다."))
|
||||||
|
|
||||||
|
root.addWidget(self._build_audio_card())
|
||||||
|
root.addWidget(self._build_behavior_card())
|
||||||
|
root.addWidget(self._build_about_card())
|
||||||
|
root.addStretch(1)
|
||||||
|
|
||||||
|
# --- 구성 -----------------------------------------------------------
|
||||||
|
def _build_audio_card(self) -> Card:
|
||||||
|
cfg = self.config.audio
|
||||||
|
card = Card(
|
||||||
|
"음성 인식",
|
||||||
|
"소리가 작아 인식이 안 되면 입력 증폭을 올리고, 주변 소음까지 잡히면 "
|
||||||
|
"민감도를 낮추세요. 값을 바꾸면 다음 시작부터 적용됩니다.",
|
||||||
|
)
|
||||||
|
form = card.add_form()
|
||||||
|
|
||||||
|
self.gain_slider, gain_row = _slider_row(
|
||||||
|
int(cfg.input_gain * 100), 25, 400, "%", self._on_audio_changed
|
||||||
|
)
|
||||||
|
form.addRow("입력 증폭", gain_row)
|
||||||
|
|
||||||
|
self.vad_combo = QComboBox()
|
||||||
|
for label, value in (
|
||||||
|
("낮음 — 작은 소리도 잡음", 0),
|
||||||
|
("보통", 1),
|
||||||
|
("높음 (기본)", 2),
|
||||||
|
("매우 높음 — 또렷한 말만", 3),
|
||||||
|
):
|
||||||
|
self.vad_combo.addItem(label, value)
|
||||||
|
self.vad_combo.setCurrentIndex(max(0, self.vad_combo.findData(cfg.vad_aggressiveness)))
|
||||||
|
self.vad_combo.currentIndexChanged.connect(self._on_audio_changed)
|
||||||
|
form.addRow("잡음 억제", self.vad_combo)
|
||||||
|
|
||||||
|
self.silence_spin = _spin(cfg.silence_ms, 200, 2000, 50, " ms", self._on_audio_changed)
|
||||||
|
self.silence_spin.setToolTip(
|
||||||
|
"이 시간만큼 조용하면 한 문장이 끝난 것으로 봅니다. "
|
||||||
|
"짧게 하면 자막이 빨리 뜨지만 문장이 자주 끊깁니다."
|
||||||
|
)
|
||||||
|
form.addRow("문장 끊는 침묵", self.silence_spin)
|
||||||
|
|
||||||
|
self.max_spin = _spin(cfg.max_segment_ms, 3000, 30_000, 1000, " ms", self._on_audio_changed)
|
||||||
|
self.max_spin.setToolTip("말이 계속 이어져도 이 길이에서 강제로 끊습니다.")
|
||||||
|
form.addRow("문장 최대 길이", self.max_spin)
|
||||||
|
|
||||||
|
self.partial_spin = _spin(
|
||||||
|
cfg.partial_interval_ms, 0, 3000, 100, " ms", self._on_audio_changed
|
||||||
|
)
|
||||||
|
self.partial_spin.setToolTip(
|
||||||
|
"말하는 도중에도 인식 중인 내용을 흐리게 미리 보여줍니다. "
|
||||||
|
"0이면 끕니다. GPU 부하가 조금 늘어납니다."
|
||||||
|
)
|
||||||
|
form.addRow("중간 결과 갱신", self.partial_spin)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_behavior_card(self) -> Card:
|
||||||
|
card = Card("동작")
|
||||||
|
form = card.add_form()
|
||||||
|
|
||||||
|
self.theme_combo = QComboBox()
|
||||||
|
self.theme_combo.addItem("어두운 테마", "dark")
|
||||||
|
self.theme_combo.addItem("밝은 테마", "light")
|
||||||
|
self.theme_combo.setCurrentIndex(max(0, self.theme_combo.findData(self.config.theme)))
|
||||||
|
self.theme_combo.currentIndexChanged.connect(self._on_behavior_changed)
|
||||||
|
form.addRow("테마", self.theme_combo)
|
||||||
|
|
||||||
|
self.preload_check = QCheckBox("시작할 때 모델을 미리 올려두기")
|
||||||
|
self.preload_check.setChecked(self.config.models.preload_on_start)
|
||||||
|
self.preload_check.setToolTip(
|
||||||
|
"끄면 프로그램이 빨리 뜨지만 첫 문장 번역이 오래 걸립니다."
|
||||||
|
)
|
||||||
|
self.preload_check.toggled.connect(self._on_behavior_changed)
|
||||||
|
form.addRow("", self.preload_check)
|
||||||
|
|
||||||
|
self.log_check = QCheckBox("번역 기록을 파일로 남기기")
|
||||||
|
self.log_check.setChecked(self.config.log_transcripts)
|
||||||
|
self.log_check.toggled.connect(self._on_behavior_changed)
|
||||||
|
form.addRow("", self.log_check)
|
||||||
|
|
||||||
|
open_folder = QPushButton("설정 폴더 열기")
|
||||||
|
open_folder.clicked.connect(self._open_folder)
|
||||||
|
form.addRow("", open_folder)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_about_card(self) -> Card:
|
||||||
|
card = Card(f"{APP_NAME} v{APP_VERSION}", APP_SLOGAN)
|
||||||
|
path = QLabel(f"설정 위치: {user_data_dir()}")
|
||||||
|
path.setObjectName("Hint")
|
||||||
|
path.setTextInteractionFlags(Qt.TextInteractionFlag.TextSelectableByMouse)
|
||||||
|
path.setWordWrap(True)
|
||||||
|
card.add(path)
|
||||||
|
return card
|
||||||
|
|
||||||
|
# --- 동작 -----------------------------------------------------------
|
||||||
|
def _on_audio_changed(self, *_) -> None:
|
||||||
|
cfg = self.config.audio
|
||||||
|
cfg.input_gain = self.gain_slider.value() / 100
|
||||||
|
cfg.vad_aggressiveness = self.vad_combo.currentData()
|
||||||
|
cfg.silence_ms = self.silence_spin.value()
|
||||||
|
cfg.max_segment_ms = self.max_spin.value()
|
||||||
|
cfg.partial_interval_ms = self.partial_spin.value()
|
||||||
|
self.config_changed.emit()
|
||||||
|
self.restart_needed.emit()
|
||||||
|
|
||||||
|
def _on_behavior_changed(self, *_) -> None:
|
||||||
|
self.config.theme = self.theme_combo.currentData()
|
||||||
|
self.config.models.preload_on_start = self.preload_check.isChecked()
|
||||||
|
self.config.log_transcripts = self.log_check.isChecked()
|
||||||
|
self.config_changed.emit()
|
||||||
|
|
||||||
|
def _open_folder(self) -> None:
|
||||||
|
from PySide6.QtCore import QUrl
|
||||||
|
from PySide6.QtGui import QDesktopServices
|
||||||
|
|
||||||
|
QDesktopServices.openUrl(QUrl.fromLocalFile(str(user_data_dir())))
|
||||||
|
|
||||||
|
|
||||||
|
def _spin(value, low, high, step, suffix, slot) -> QSpinBox:
|
||||||
|
spin = QSpinBox()
|
||||||
|
spin.setRange(low, high)
|
||||||
|
spin.setSingleStep(step)
|
||||||
|
spin.setSuffix(suffix)
|
||||||
|
spin.setValue(value)
|
||||||
|
spin.valueChanged.connect(slot)
|
||||||
|
return spin
|
||||||
|
|
||||||
|
|
||||||
|
def _slider_row(value, low, high, suffix, slot) -> tuple[QSlider, QWidget]:
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
slider = QSlider(Qt.Orientation.Horizontal)
|
||||||
|
slider.setRange(low, high)
|
||||||
|
slider.setValue(value)
|
||||||
|
label = QLabel(f"{value}{suffix}")
|
||||||
|
label.setMinimumWidth(48)
|
||||||
|
slider.valueChanged.connect(lambda v: label.setText(f"{v}{suffix}"))
|
||||||
|
slider.valueChanged.connect(slot)
|
||||||
|
layout.addWidget(slider, 1)
|
||||||
|
layout.addWidget(label)
|
||||||
|
return slider, row
|
||||||
216
src/hearo/ui/pages/subtitle_page.py
Normal file
216
src/hearo/ui/pages/subtitle_page.py
Normal file
@@ -0,0 +1,216 @@
|
|||||||
|
"""자막 — 글꼴·색·크기 등 표시 설정과 실시간 미리보기."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt, Signal
|
||||||
|
from PySide6.QtGui import QColor, QFont, QFontDatabase
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QCheckBox,
|
||||||
|
QColorDialog,
|
||||||
|
QComboBox,
|
||||||
|
QFontComboBox,
|
||||||
|
QHBoxLayout,
|
||||||
|
QLabel,
|
||||||
|
QPushButton,
|
||||||
|
QSlider,
|
||||||
|
QSpinBox,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ...config import AppConfig
|
||||||
|
from ..overlay import SubtitlePreview
|
||||||
|
from ..theme import SPACING
|
||||||
|
from ..widgets import Card, PageHeader
|
||||||
|
|
||||||
|
PREVIEW_TEXT = "적이 왼쪽에서 들어온다, 지금 바로 빠져!"
|
||||||
|
PREVIEW_SOURCE = "Enemy coming from the left, fall back now!"
|
||||||
|
|
||||||
|
|
||||||
|
class ColorButton(QPushButton):
|
||||||
|
color_picked = Signal(str)
|
||||||
|
|
||||||
|
def __init__(self, color: str, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.setFixedSize(52, 30)
|
||||||
|
self.setCursor(Qt.CursorShape.PointingHandCursor)
|
||||||
|
self.set_color(color)
|
||||||
|
self.clicked.connect(self._pick)
|
||||||
|
|
||||||
|
@property
|
||||||
|
def color(self) -> str:
|
||||||
|
return self._color
|
||||||
|
|
||||||
|
def set_color(self, color: str) -> None:
|
||||||
|
self._color = color
|
||||||
|
self.setStyleSheet(
|
||||||
|
f"background: {color}; border: 1px solid #3A4152; border-radius: 6px;"
|
||||||
|
)
|
||||||
|
|
||||||
|
def _pick(self) -> None:
|
||||||
|
chosen = QColorDialog.getColor(QColor(self._color), self, "색 선택")
|
||||||
|
if chosen.isValid():
|
||||||
|
self.set_color(chosen.name())
|
||||||
|
self.color_picked.emit(chosen.name())
|
||||||
|
|
||||||
|
|
||||||
|
class SubtitlePage(QWidget):
|
||||||
|
style_changed = Signal()
|
||||||
|
overlay_toggled = Signal(bool)
|
||||||
|
|
||||||
|
def __init__(self, config: AppConfig, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.config = config
|
||||||
|
|
||||||
|
root = QVBoxLayout(self)
|
||||||
|
root.setContentsMargins(0, 0, 0, 0)
|
||||||
|
root.setSpacing(SPACING + 4)
|
||||||
|
root.addWidget(
|
||||||
|
PageHeader("자막", "자막 창에 표시될 글씨의 모양을 정합니다. 아래에서 바로 확인하세요.")
|
||||||
|
)
|
||||||
|
|
||||||
|
root.addWidget(self._build_preview_card())
|
||||||
|
root.addWidget(self._build_text_card())
|
||||||
|
root.addWidget(self._build_layout_card())
|
||||||
|
root.addStretch(1)
|
||||||
|
|
||||||
|
self._sync_preview()
|
||||||
|
|
||||||
|
# --- 구성 -----------------------------------------------------------
|
||||||
|
def _build_preview_card(self) -> Card:
|
||||||
|
card = Card("미리보기", "실제 자막 창과 같은 방식으로 그립니다.")
|
||||||
|
self.preview = SubtitlePreview(self.config.subtitle)
|
||||||
|
self.preview.set_sample(PREVIEW_TEXT, PREVIEW_SOURCE)
|
||||||
|
card.add(self.preview)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_text_card(self) -> Card:
|
||||||
|
st = self.config.subtitle
|
||||||
|
card = Card("글씨")
|
||||||
|
form = card.add_form()
|
||||||
|
|
||||||
|
self.font_combo = QFontComboBox()
|
||||||
|
available = set(QFontDatabase.families())
|
||||||
|
self.font_combo.setCurrentFont(
|
||||||
|
QFont(st.font_family if st.font_family in available else "Segoe UI")
|
||||||
|
)
|
||||||
|
self.font_combo.currentFontChanged.connect(self._on_font_changed)
|
||||||
|
form.addRow("글꼴", self.font_combo)
|
||||||
|
|
||||||
|
self.size_spin = QSpinBox()
|
||||||
|
self.size_spin.setRange(12, 120)
|
||||||
|
self.size_spin.setSuffix(" px")
|
||||||
|
self.size_spin.setValue(st.font_size)
|
||||||
|
self.size_spin.valueChanged.connect(self._on_change)
|
||||||
|
form.addRow("크기", self.size_spin)
|
||||||
|
|
||||||
|
self.bold_check = QCheckBox("굵게")
|
||||||
|
self.bold_check.setChecked(st.bold)
|
||||||
|
self.bold_check.toggled.connect(self._on_change)
|
||||||
|
form.addRow("", self.bold_check)
|
||||||
|
|
||||||
|
self.text_color = ColorButton(st.text_color)
|
||||||
|
self.text_color.color_picked.connect(self._on_change)
|
||||||
|
form.addRow("번역문 색", self.text_color)
|
||||||
|
|
||||||
|
self.source_color = ColorButton(st.source_text_color)
|
||||||
|
self.source_color.color_picked.connect(self._on_change)
|
||||||
|
form.addRow("원문 색", self.source_color)
|
||||||
|
|
||||||
|
self.outline_color = ColorButton(st.outline_color)
|
||||||
|
self.outline_color.color_picked.connect(self._on_change)
|
||||||
|
form.addRow("외곽선 색", self.outline_color)
|
||||||
|
|
||||||
|
self.outline_spin = QSpinBox()
|
||||||
|
self.outline_spin.setRange(0, 8)
|
||||||
|
self.outline_spin.setSuffix(" px")
|
||||||
|
self.outline_spin.setValue(st.outline_width)
|
||||||
|
self.outline_spin.setToolTip("배경이 밝든 어둡든 글씨가 읽히게 해줍니다. 0이면 끕니다.")
|
||||||
|
self.outline_spin.valueChanged.connect(self._on_change)
|
||||||
|
form.addRow("외곽선 두께", self.outline_spin)
|
||||||
|
return card
|
||||||
|
|
||||||
|
def _build_layout_card(self) -> Card:
|
||||||
|
st = self.config.subtitle
|
||||||
|
card = Card("배치와 배경")
|
||||||
|
form = card.add_form()
|
||||||
|
|
||||||
|
self.bg_color = ColorButton(st.background_color)
|
||||||
|
self.bg_color.color_picked.connect(self._on_change)
|
||||||
|
form.addRow("배경 색", self.bg_color)
|
||||||
|
|
||||||
|
opacity_row = QWidget()
|
||||||
|
layout = QHBoxLayout(opacity_row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
self.opacity_slider = QSlider(Qt.Orientation.Horizontal)
|
||||||
|
self.opacity_slider.setRange(0, 100)
|
||||||
|
self.opacity_slider.setValue(st.background_opacity)
|
||||||
|
self.opacity_slider.valueChanged.connect(self._on_change)
|
||||||
|
self.opacity_value = QLabel(f"{st.background_opacity}%")
|
||||||
|
self.opacity_value.setMinimumWidth(42)
|
||||||
|
layout.addWidget(self.opacity_slider, 1)
|
||||||
|
layout.addWidget(self.opacity_value)
|
||||||
|
form.addRow("배경 투명도", opacity_row)
|
||||||
|
|
||||||
|
self.align_combo = QComboBox()
|
||||||
|
for label, value in (("가운데", "center"), ("왼쪽", "left"), ("오른쪽", "right")):
|
||||||
|
self.align_combo.addItem(label, value)
|
||||||
|
self.align_combo.setCurrentIndex(max(0, self.align_combo.findData(st.align)))
|
||||||
|
self.align_combo.currentIndexChanged.connect(self._on_change)
|
||||||
|
form.addRow("정렬", self.align_combo)
|
||||||
|
|
||||||
|
self.lines_spin = QSpinBox()
|
||||||
|
self.lines_spin.setRange(1, 5)
|
||||||
|
self.lines_spin.setSuffix(" 줄")
|
||||||
|
self.lines_spin.setValue(st.max_lines)
|
||||||
|
self.lines_spin.valueChanged.connect(self._on_change)
|
||||||
|
form.addRow("최대 줄 수", self.lines_spin)
|
||||||
|
|
||||||
|
self.fade_spin = QSpinBox()
|
||||||
|
self.fade_spin.setRange(0, 30_000)
|
||||||
|
self.fade_spin.setSingleStep(500)
|
||||||
|
self.fade_spin.setSuffix(" ms")
|
||||||
|
self.fade_spin.setValue(st.fade_out_ms)
|
||||||
|
self.fade_spin.setToolTip("이 시간이 지나면 자막을 지웁니다. 0이면 계속 남겨둡니다.")
|
||||||
|
self.fade_spin.valueChanged.connect(self._on_change)
|
||||||
|
form.addRow("자동 지우기", self.fade_spin)
|
||||||
|
|
||||||
|
self.show_source_check = QCheckBox("원문도 같이 표시")
|
||||||
|
self.show_source_check.setChecked(st.show_source)
|
||||||
|
self.show_source_check.toggled.connect(self._on_change)
|
||||||
|
form.addRow("", self.show_source_check)
|
||||||
|
|
||||||
|
self.overlay_check = QCheckBox("자막 창 보이기")
|
||||||
|
self.overlay_check.setChecked(self.config.overlay.visible)
|
||||||
|
self.overlay_check.toggled.connect(self.overlay_toggled.emit)
|
||||||
|
form.addRow("", self.overlay_check)
|
||||||
|
return card
|
||||||
|
|
||||||
|
# --- 동작 -----------------------------------------------------------
|
||||||
|
def _on_font_changed(self, font: QFont) -> None:
|
||||||
|
self.config.subtitle.font_family = font.family()
|
||||||
|
self._on_change()
|
||||||
|
|
||||||
|
def _on_change(self, *_) -> None:
|
||||||
|
st = self.config.subtitle
|
||||||
|
st.font_size = self.size_spin.value()
|
||||||
|
st.bold = self.bold_check.isChecked()
|
||||||
|
st.text_color = self.text_color.color
|
||||||
|
st.source_text_color = self.source_color.color
|
||||||
|
st.outline_color = self.outline_color.color
|
||||||
|
st.outline_width = self.outline_spin.value()
|
||||||
|
st.background_color = self.bg_color.color
|
||||||
|
st.background_opacity = self.opacity_slider.value()
|
||||||
|
st.align = self.align_combo.currentData()
|
||||||
|
st.max_lines = self.lines_spin.value()
|
||||||
|
st.fade_out_ms = self.fade_spin.value()
|
||||||
|
st.show_source = self.show_source_check.isChecked()
|
||||||
|
|
||||||
|
self.opacity_value.setText(f"{st.background_opacity}%")
|
||||||
|
self._sync_preview()
|
||||||
|
self.style_changed.emit()
|
||||||
|
|
||||||
|
def _sync_preview(self) -> None:
|
||||||
|
# 미리보기는 설정 객체를 그대로 참조하므로 다시 그리기만 하면 된다.
|
||||||
|
self.preview.setMinimumHeight(max(140, self.config.subtitle.font_size * 4))
|
||||||
|
self.preview.update()
|
||||||
209
src/hearo/ui/theme.py
Normal file
209
src/hearo/ui/theme.py
Normal file
@@ -0,0 +1,209 @@
|
|||||||
|
"""디자인 토큰과 QSS 스타일시트.
|
||||||
|
|
||||||
|
의도적으로 외부 테마 패키지를 쓰지 않는다. 자막 오버레이가 반투명·무테두리라
|
||||||
|
플랫폼 기본 테마와 섞이면 오히려 지저분해지고, 색/간격을 한곳에서 관리하는 편이
|
||||||
|
설정 화면의 '미리보기'와 실제 자막을 일치시키기에도 좋다.
|
||||||
|
"""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from dataclasses import dataclass
|
||||||
|
|
||||||
|
|
||||||
|
@dataclass(frozen=True)
|
||||||
|
class Palette:
|
||||||
|
bg: str
|
||||||
|
surface: str
|
||||||
|
surface_alt: str
|
||||||
|
border: str
|
||||||
|
text: str
|
||||||
|
text_dim: str
|
||||||
|
accent: str
|
||||||
|
accent_hover: str
|
||||||
|
accent_text: str
|
||||||
|
success: str
|
||||||
|
warning: str
|
||||||
|
danger: str
|
||||||
|
|
||||||
|
|
||||||
|
DARK = Palette(
|
||||||
|
bg="#0F1117",
|
||||||
|
surface="#171A22",
|
||||||
|
surface_alt="#1E222C",
|
||||||
|
border="#2A2F3C",
|
||||||
|
text="#E7EAF0",
|
||||||
|
text_dim="#8B93A5",
|
||||||
|
accent="#5B8CFF",
|
||||||
|
accent_hover="#7AA2FF",
|
||||||
|
accent_text="#0B0E14",
|
||||||
|
success="#3DD68C",
|
||||||
|
warning="#F5B84E",
|
||||||
|
danger="#F2555A",
|
||||||
|
)
|
||||||
|
|
||||||
|
LIGHT = Palette(
|
||||||
|
bg="#F4F6FA",
|
||||||
|
surface="#FFFFFF",
|
||||||
|
surface_alt="#EDF0F6",
|
||||||
|
border="#D9DEE9",
|
||||||
|
text="#131722",
|
||||||
|
text_dim="#5C6578",
|
||||||
|
accent="#3B6EF0",
|
||||||
|
accent_hover="#2A5AD6",
|
||||||
|
accent_text="#FFFFFF",
|
||||||
|
success="#17A46A",
|
||||||
|
warning="#C9821B",
|
||||||
|
danger="#D63A3F",
|
||||||
|
)
|
||||||
|
|
||||||
|
RADIUS = 10
|
||||||
|
SPACING = 12
|
||||||
|
|
||||||
|
|
||||||
|
def palette(theme: str) -> Palette:
|
||||||
|
return LIGHT if theme == "light" else DARK
|
||||||
|
|
||||||
|
|
||||||
|
def stylesheet(theme: str = "dark") -> str:
|
||||||
|
p = palette(theme)
|
||||||
|
return f"""
|
||||||
|
QWidget {{
|
||||||
|
background: {p.bg};
|
||||||
|
color: {p.text};
|
||||||
|
font-family: "Pretendard", "Segoe UI Variable", "Segoe UI", "Malgun Gothic", sans-serif;
|
||||||
|
font-size: 14px;
|
||||||
|
}}
|
||||||
|
/* 라벨류는 부모(카드/사이드바) 배경이 비쳐야 한다. 지정하지 않으면
|
||||||
|
QWidget 규칙의 창 배경색을 상속해 카드 위에 검은 박스가 생긴다. */
|
||||||
|
QLabel, QCheckBox, QRadioButton, QScrollArea > QWidget > QWidget {{
|
||||||
|
background: transparent;
|
||||||
|
}}
|
||||||
|
QLabel#Title {{ font-size: 22px; font-weight: 700; }}
|
||||||
|
QLabel#Subtitle {{ font-size: 13px; color: {p.text_dim}; }}
|
||||||
|
QLabel#SectionLabel {{ font-size: 12px; font-weight: 700; color: {p.text_dim};
|
||||||
|
letter-spacing: 1px; }}
|
||||||
|
QLabel#Hint {{ font-size: 12px; color: {p.text_dim}; }}
|
||||||
|
|
||||||
|
/* --- 사이드바 --- */
|
||||||
|
QFrame#Sidebar {{
|
||||||
|
background: {p.surface};
|
||||||
|
border-right: 1px solid {p.border};
|
||||||
|
}}
|
||||||
|
QPushButton#NavButton {{
|
||||||
|
background: transparent;
|
||||||
|
border: none;
|
||||||
|
border-radius: {RADIUS}px;
|
||||||
|
padding: 11px 14px;
|
||||||
|
text-align: left;
|
||||||
|
color: {p.text_dim};
|
||||||
|
font-size: 14px;
|
||||||
|
}}
|
||||||
|
QPushButton#NavButton:hover {{ background: {p.surface_alt}; color: {p.text}; }}
|
||||||
|
QPushButton#NavButton:checked {{
|
||||||
|
background: {p.accent};
|
||||||
|
color: {p.accent_text};
|
||||||
|
font-weight: 700;
|
||||||
|
}}
|
||||||
|
|
||||||
|
/* --- 카드 --- */
|
||||||
|
QFrame#Card {{
|
||||||
|
background: {p.surface};
|
||||||
|
border: 1px solid {p.border};
|
||||||
|
border-radius: {RADIUS}px;
|
||||||
|
}}
|
||||||
|
QFrame#CardSelected {{
|
||||||
|
background: {p.surface};
|
||||||
|
border: 2px solid {p.accent};
|
||||||
|
border-radius: {RADIUS}px;
|
||||||
|
}}
|
||||||
|
|
||||||
|
/* --- 입력 위젯 --- */
|
||||||
|
QComboBox, QLineEdit, QSpinBox, QPlainTextEdit, QTextEdit, QListWidget, QTableWidget {{
|
||||||
|
background: {p.surface_alt};
|
||||||
|
border: 1px solid {p.border};
|
||||||
|
border-radius: 8px;
|
||||||
|
padding: 7px 10px;
|
||||||
|
selection-background-color: {p.accent};
|
||||||
|
selection-color: {p.accent_text};
|
||||||
|
}}
|
||||||
|
QComboBox:focus, QLineEdit:focus, QSpinBox:focus, QPlainTextEdit:focus {{
|
||||||
|
border: 1px solid {p.accent};
|
||||||
|
}}
|
||||||
|
QComboBox::drop-down {{ border: none; width: 22px; }}
|
||||||
|
QComboBox QAbstractItemView {{
|
||||||
|
background: {p.surface_alt};
|
||||||
|
border: 1px solid {p.border};
|
||||||
|
selection-background-color: {p.accent};
|
||||||
|
outline: none;
|
||||||
|
}}
|
||||||
|
QHeaderView::section {{
|
||||||
|
background: {p.surface};
|
||||||
|
color: {p.text_dim};
|
||||||
|
border: none;
|
||||||
|
border-bottom: 1px solid {p.border};
|
||||||
|
padding: 7px;
|
||||||
|
font-weight: 600;
|
||||||
|
}}
|
||||||
|
QTableWidget {{ gridline-color: {p.border}; }}
|
||||||
|
|
||||||
|
/* --- 버튼 --- */
|
||||||
|
QPushButton {{
|
||||||
|
background: {p.surface_alt};
|
||||||
|
border: 1px solid {p.border};
|
||||||
|
border-radius: 8px;
|
||||||
|
padding: 8px 16px;
|
||||||
|
color: {p.text};
|
||||||
|
}}
|
||||||
|
QPushButton:hover {{ border-color: {p.accent}; }}
|
||||||
|
QPushButton:disabled {{ color: {p.text_dim}; border-color: {p.border}; }}
|
||||||
|
QPushButton#Primary {{
|
||||||
|
background: {p.accent};
|
||||||
|
color: {p.accent_text};
|
||||||
|
border: none;
|
||||||
|
font-weight: 700;
|
||||||
|
padding: 11px 22px;
|
||||||
|
}}
|
||||||
|
QPushButton#Primary:hover {{ background: {p.accent_hover}; }}
|
||||||
|
QPushButton#Primary:disabled {{ background: {p.border}; color: {p.text_dim}; }}
|
||||||
|
QPushButton#Danger {{
|
||||||
|
background: {p.danger};
|
||||||
|
color: #FFFFFF;
|
||||||
|
border: none;
|
||||||
|
font-weight: 700;
|
||||||
|
padding: 11px 22px;
|
||||||
|
}}
|
||||||
|
|
||||||
|
/* --- 기타 --- */
|
||||||
|
QSlider::groove:horizontal {{
|
||||||
|
height: 4px; background: {p.border}; border-radius: 2px;
|
||||||
|
}}
|
||||||
|
QSlider::handle:horizontal {{
|
||||||
|
width: 16px; height: 16px; margin: -6px 0;
|
||||||
|
background: {p.accent}; border-radius: 8px;
|
||||||
|
}}
|
||||||
|
QSlider::sub-page:horizontal {{ background: {p.accent}; border-radius: 2px; }}
|
||||||
|
QProgressBar {{
|
||||||
|
background: {p.surface_alt}; border: none; border-radius: 4px;
|
||||||
|
height: 6px; text-align: center; color: transparent;
|
||||||
|
}}
|
||||||
|
QProgressBar::chunk {{ background: {p.accent}; border-radius: 4px; }}
|
||||||
|
QCheckBox::indicator, QRadioButton::indicator {{
|
||||||
|
width: 17px; height: 17px;
|
||||||
|
border: 1px solid {p.border}; border-radius: 4px;
|
||||||
|
background: {p.surface_alt};
|
||||||
|
}}
|
||||||
|
QCheckBox::indicator:checked, QRadioButton::indicator:checked {{
|
||||||
|
background: {p.accent}; border-color: {p.accent};
|
||||||
|
}}
|
||||||
|
QScrollBar:vertical {{ background: transparent; width: 10px; margin: 0; }}
|
||||||
|
QScrollBar::handle:vertical {{
|
||||||
|
background: {p.border}; border-radius: 5px; min-height: 30px;
|
||||||
|
}}
|
||||||
|
QScrollBar::handle:vertical:hover {{ background: {p.text_dim}; }}
|
||||||
|
QScrollBar::add-line, QScrollBar::sub-line {{ height: 0; }}
|
||||||
|
QScrollArea {{ border: none; }}
|
||||||
|
QToolTip {{
|
||||||
|
background: {p.surface_alt}; color: {p.text};
|
||||||
|
border: 1px solid {p.border}; padding: 6px; border-radius: 6px;
|
||||||
|
}}
|
||||||
|
"""
|
||||||
3
src/hearo/ui/widgets/__init__.py
Normal file
3
src/hearo/ui/widgets/__init__.py
Normal file
@@ -0,0 +1,3 @@
|
|||||||
|
from .common import Card, LevelMeter, PageHeader, StatusPill, labelled_row
|
||||||
|
|
||||||
|
__all__ = ["Card", "LevelMeter", "PageHeader", "StatusPill", "labelled_row"]
|
||||||
136
src/hearo/ui/widgets/common.py
Normal file
136
src/hearo/ui/widgets/common.py
Normal file
@@ -0,0 +1,136 @@
|
|||||||
|
"""여러 화면에서 재사용하는 작은 위젯들."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
from PySide6.QtCore import Qt
|
||||||
|
from PySide6.QtGui import QColor, QPainter
|
||||||
|
from PySide6.QtWidgets import (
|
||||||
|
QFormLayout,
|
||||||
|
QFrame,
|
||||||
|
QHBoxLayout,
|
||||||
|
QLabel,
|
||||||
|
QSizePolicy,
|
||||||
|
QVBoxLayout,
|
||||||
|
QWidget,
|
||||||
|
)
|
||||||
|
|
||||||
|
from ..theme import SPACING, palette
|
||||||
|
|
||||||
|
|
||||||
|
class Card(QFrame):
|
||||||
|
"""제목 + 본문 레이아웃을 가진 카드 컨테이너."""
|
||||||
|
|
||||||
|
def __init__(self, title: str = "", subtitle: str = "", parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self.setObjectName("Card")
|
||||||
|
outer = QVBoxLayout(self)
|
||||||
|
outer.setContentsMargins(SPACING + 4, SPACING + 4, SPACING + 4, SPACING + 4)
|
||||||
|
outer.setSpacing(SPACING)
|
||||||
|
|
||||||
|
if title:
|
||||||
|
label = QLabel(title)
|
||||||
|
label.setStyleSheet("font-size: 15px; font-weight: 700;")
|
||||||
|
outer.addWidget(label)
|
||||||
|
if subtitle:
|
||||||
|
hint = QLabel(subtitle)
|
||||||
|
hint.setObjectName("Hint")
|
||||||
|
hint.setWordWrap(True)
|
||||||
|
outer.addWidget(hint)
|
||||||
|
|
||||||
|
self.body = QVBoxLayout()
|
||||||
|
self.body.setSpacing(SPACING)
|
||||||
|
outer.addLayout(self.body)
|
||||||
|
|
||||||
|
def add(self, widget: QWidget) -> QWidget:
|
||||||
|
self.body.addWidget(widget)
|
||||||
|
return widget
|
||||||
|
|
||||||
|
def add_form(self) -> QFormLayout:
|
||||||
|
form = QFormLayout()
|
||||||
|
form.setSpacing(10)
|
||||||
|
form.setLabelAlignment(Qt.AlignmentFlag.AlignLeft | Qt.AlignmentFlag.AlignVCenter)
|
||||||
|
form.setFieldGrowthPolicy(QFormLayout.FieldGrowthPolicy.ExpandingFieldsGrow)
|
||||||
|
self.body.addLayout(form)
|
||||||
|
return form
|
||||||
|
|
||||||
|
|
||||||
|
class PageHeader(QWidget):
|
||||||
|
def __init__(self, title: str, subtitle: str = "", parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
layout = QVBoxLayout(self)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(4)
|
||||||
|
head = QLabel(title)
|
||||||
|
head.setObjectName("Title")
|
||||||
|
layout.addWidget(head)
|
||||||
|
if subtitle:
|
||||||
|
sub = QLabel(subtitle)
|
||||||
|
sub.setObjectName("Subtitle")
|
||||||
|
sub.setWordWrap(True)
|
||||||
|
layout.addWidget(sub)
|
||||||
|
|
||||||
|
|
||||||
|
class StatusPill(QLabel):
|
||||||
|
"""상태를 색 점 + 문구로 보여주는 배지."""
|
||||||
|
|
||||||
|
def __init__(self, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self._color = "#8B93A5"
|
||||||
|
self.setContentsMargins(10, 6, 12, 6)
|
||||||
|
self.set_status("대기 중", "idle")
|
||||||
|
|
||||||
|
def set_status(self, text: str, kind: str = "idle") -> None:
|
||||||
|
p = palette("dark")
|
||||||
|
self._color = {
|
||||||
|
"idle": p.text_dim,
|
||||||
|
"busy": p.warning,
|
||||||
|
"ok": p.success,
|
||||||
|
"error": p.danger,
|
||||||
|
}.get(kind, p.text_dim)
|
||||||
|
self.setText(f"● {text}")
|
||||||
|
self.setStyleSheet(
|
||||||
|
f"color: {self._color}; background: {p.surface_alt};"
|
||||||
|
f"border-radius: 13px; padding: 6px 14px; font-weight: 600;"
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
class LevelMeter(QWidget):
|
||||||
|
"""입력 오디오 레벨 바. 소리가 잡히는지 한눈에 확인하는 용도."""
|
||||||
|
|
||||||
|
def __init__(self, parent: QWidget | None = None):
|
||||||
|
super().__init__(parent)
|
||||||
|
self._level = 0.0
|
||||||
|
self._peak = 0.0
|
||||||
|
self.setFixedHeight(8)
|
||||||
|
self.setSizePolicy(QSizePolicy.Policy.Expanding, QSizePolicy.Policy.Fixed)
|
||||||
|
|
||||||
|
def set_level(self, value: float) -> None:
|
||||||
|
# RMS를 그대로 그리면 거의 안 움직여서 제곱근으로 펴준다.
|
||||||
|
self._level = min(1.0, max(0.0, value) ** 0.5 * 2.2)
|
||||||
|
self._peak = max(self._level, self._peak * 0.92)
|
||||||
|
self.update()
|
||||||
|
|
||||||
|
def paintEvent(self, event) -> None: # noqa: N802
|
||||||
|
p = palette("dark")
|
||||||
|
painter = QPainter(self)
|
||||||
|
painter.setRenderHint(QPainter.RenderHint.Antialiasing)
|
||||||
|
painter.setPen(Qt.PenStyle.NoPen)
|
||||||
|
painter.setBrush(QColor(p.surface_alt))
|
||||||
|
painter.drawRoundedRect(self.rect(), 4, 4)
|
||||||
|
if self._level > 0.01:
|
||||||
|
color = QColor(p.danger if self._level > 0.92 else p.success)
|
||||||
|
rect = self.rect().adjusted(0, 0, -int(self.width() * (1 - self._level)), 0)
|
||||||
|
painter.setBrush(color)
|
||||||
|
painter.drawRoundedRect(rect, 4, 4)
|
||||||
|
|
||||||
|
|
||||||
|
def labelled_row(label: str, widget: QWidget) -> QWidget:
|
||||||
|
row = QWidget()
|
||||||
|
layout = QHBoxLayout(row)
|
||||||
|
layout.setContentsMargins(0, 0, 0, 0)
|
||||||
|
layout.setSpacing(10)
|
||||||
|
text = QLabel(label)
|
||||||
|
text.setMinimumWidth(110)
|
||||||
|
layout.addWidget(text)
|
||||||
|
layout.addWidget(widget, 1)
|
||||||
|
return row
|
||||||
58
tests/test_config.py
Normal file
58
tests/test_config.py
Normal file
@@ -0,0 +1,58 @@
|
|||||||
|
import json
|
||||||
|
|
||||||
|
from hearo.config import AppConfig
|
||||||
|
from hearo.models.tiers import DEFAULT_TIER
|
||||||
|
|
||||||
|
|
||||||
|
def test_defaults_are_sane():
|
||||||
|
cfg = AppConfig()
|
||||||
|
assert cfg.models.tier == DEFAULT_TIER
|
||||||
|
assert cfg.models.target_lang == "ko"
|
||||||
|
assert cfg.subtitle.font_size > 0
|
||||||
|
assert cfg.audio.silence_ms > 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_save_load_roundtrip(tmp_path):
|
||||||
|
path = tmp_path / "config.json"
|
||||||
|
cfg = AppConfig()
|
||||||
|
cfg.subtitle.font_size = 52
|
||||||
|
cfg.subtitle.text_color = "#FF00AA"
|
||||||
|
cfg.models.tier = "precision"
|
||||||
|
cfg.audio.target_pid = 4242
|
||||||
|
cfg.save(path)
|
||||||
|
|
||||||
|
loaded = AppConfig.load(path)
|
||||||
|
assert loaded.subtitle.font_size == 52
|
||||||
|
assert loaded.subtitle.text_color == "#FF00AA"
|
||||||
|
assert loaded.models.tier == "precision"
|
||||||
|
assert loaded.audio.target_pid == 4242
|
||||||
|
|
||||||
|
|
||||||
|
def test_missing_file_gives_defaults(tmp_path):
|
||||||
|
assert AppConfig.load(tmp_path / "nope.json").models.tier == DEFAULT_TIER
|
||||||
|
|
||||||
|
|
||||||
|
def test_corrupt_file_gives_defaults(tmp_path):
|
||||||
|
path = tmp_path / "config.json"
|
||||||
|
path.write_text("{ this is not json", encoding="utf-8")
|
||||||
|
assert AppConfig.load(path).models.tier == DEFAULT_TIER
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_keys_are_ignored_and_missing_keys_defaulted(tmp_path):
|
||||||
|
"""스키마가 바뀌어도 사용자의 기존 설정이 깨지면 안 된다."""
|
||||||
|
path = tmp_path / "config.json"
|
||||||
|
path.write_text(
|
||||||
|
json.dumps(
|
||||||
|
{
|
||||||
|
"theme": "light",
|
||||||
|
"removed_option": True,
|
||||||
|
"subtitle": {"font_size": 40, "gone_field": 1},
|
||||||
|
}
|
||||||
|
),
|
||||||
|
encoding="utf-8",
|
||||||
|
)
|
||||||
|
cfg = AppConfig.load(path)
|
||||||
|
assert cfg.theme == "light"
|
||||||
|
assert cfg.subtitle.font_size == 40
|
||||||
|
assert cfg.subtitle.bold is True # 기본값 유지
|
||||||
|
assert not hasattr(cfg, "removed_option")
|
||||||
75
tests/test_glossary.py
Normal file
75
tests/test_glossary.py
Normal file
@@ -0,0 +1,75 @@
|
|||||||
|
from hearo.models.glossary import Glossary, GlossaryEntry
|
||||||
|
|
||||||
|
|
||||||
|
def make() -> Glossary:
|
||||||
|
return Glossary(
|
||||||
|
[
|
||||||
|
GlossaryEntry("headshot", {"ko": "헤드샷"}),
|
||||||
|
GlossaryEntry("head", {"ko": "머리"}),
|
||||||
|
GlossaryEntry("Nexus", {"ko": "넥서스", "ja": "ネクサス"}),
|
||||||
|
]
|
||||||
|
)
|
||||||
|
|
||||||
|
|
||||||
|
def test_protect_and_restore_roundtrip():
|
||||||
|
g = make()
|
||||||
|
protected, repl = g.protect("Nice headshot on the Nexus", "ko")
|
||||||
|
assert "headshot" not in protected
|
||||||
|
assert "Nexus" not in protected
|
||||||
|
assert repl == ["헤드샷", "넥서스"]
|
||||||
|
# 번역 모델이 플레이스홀더를 그대로 통과시켰다고 가정
|
||||||
|
assert Glossary.restore(protected, repl) == "Nice 헤드샷 on the 넥서스"
|
||||||
|
|
||||||
|
|
||||||
|
def test_longest_match_wins():
|
||||||
|
"""'headshot'이 'head'에 먼저 잡아먹히면 안 된다."""
|
||||||
|
g = make()
|
||||||
|
_, repl = g.protect("headshot", "ko")
|
||||||
|
assert repl == ["헤드샷"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_word_boundary_for_latin_terms():
|
||||||
|
g = make()
|
||||||
|
_, repl = g.protect("overheated", "ko")
|
||||||
|
assert repl == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_missing_target_language_is_left_alone():
|
||||||
|
g = make()
|
||||||
|
protected, repl = g.protect("headshot", "ja")
|
||||||
|
assert protected == "headshot"
|
||||||
|
assert repl == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_prompt_hint_only_includes_present_terms():
|
||||||
|
g = make()
|
||||||
|
hint = g.prompt_hint("push to the Nexus", "ko")
|
||||||
|
assert "Nexus -> 넥서스" in hint
|
||||||
|
assert "headshot" not in hint
|
||||||
|
assert g.prompt_hint("nothing here", "ko") == ""
|
||||||
|
|
||||||
|
|
||||||
|
def test_case_insensitive_by_default():
|
||||||
|
g = make()
|
||||||
|
_, repl = g.protect("NEXUS down", "ko")
|
||||||
|
assert repl == ["넥서스"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_save_and_load(tmp_path):
|
||||||
|
path = tmp_path / "glossary.json"
|
||||||
|
make().save(path)
|
||||||
|
loaded = Glossary.load(path)
|
||||||
|
assert len(loaded) == 3
|
||||||
|
assert loaded.entries[2].target_for("ja") == "ネクサス"
|
||||||
|
|
||||||
|
|
||||||
|
def test_add_replaces_existing_entry():
|
||||||
|
g = make()
|
||||||
|
g.add(GlossaryEntry("Nexus", {"ko": "본진"}))
|
||||||
|
assert len(g) == 3
|
||||||
|
_, repl = g.protect("Nexus", "ko")
|
||||||
|
assert repl == ["본진"]
|
||||||
|
|
||||||
|
|
||||||
|
def test_restore_ignores_out_of_range_placeholder():
|
||||||
|
assert Glossary.restore("a ⟦5⟧ b", ["x"]) == "a b"
|
||||||
129
tests/test_pipeline.py
Normal file
129
tests/test_pipeline.py
Normal file
@@ -0,0 +1,129 @@
|
|||||||
|
"""엔진 파이프라인을 가짜 모델로 끝까지 돌려본다 (GPU/오디오 장치 불필요)."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import time
|
||||||
|
|
||||||
|
import numpy as np
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
from hearo.audio.base import CaptureBackend, to_mono_16k
|
||||||
|
from hearo.config import AppConfig
|
||||||
|
from hearo.constants import SAMPLE_RATE
|
||||||
|
from hearo.core.engine import TranslationEngine
|
||||||
|
from hearo.models.asr import Transcript
|
||||||
|
|
||||||
|
|
||||||
|
class FakeCapture(CaptureBackend):
|
||||||
|
"""말-침묵-말-침묵 패턴을 즉시 흘려보낸다."""
|
||||||
|
|
||||||
|
name = "fake"
|
||||||
|
per_process = True
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def available() -> bool:
|
||||||
|
return True
|
||||||
|
|
||||||
|
@staticmethod
|
||||||
|
def list_sources():
|
||||||
|
return []
|
||||||
|
|
||||||
|
def _run(self) -> None:
|
||||||
|
def tone(ms, amp=0.35):
|
||||||
|
n = SAMPLE_RATE * ms // 1000
|
||||||
|
t = np.arange(n, dtype=np.float32) / SAMPLE_RATE
|
||||||
|
return (np.sin(2 * np.pi * 240 * t) * amp).astype(np.float32)
|
||||||
|
|
||||||
|
pattern = np.concatenate(
|
||||||
|
[np.zeros(SAMPLE_RATE // 5, np.float32), tone(900),
|
||||||
|
np.zeros(SAMPLE_RATE, np.float32)]
|
||||||
|
)
|
||||||
|
while not self._stop.is_set():
|
||||||
|
for i in range(0, pattern.size, 1600):
|
||||||
|
if self._stop.is_set():
|
||||||
|
return
|
||||||
|
self._emit(pattern[i : i + 1600].copy())
|
||||||
|
time.sleep(0.005)
|
||||||
|
|
||||||
|
|
||||||
|
class FakeRecognizer:
|
||||||
|
loaded = True
|
||||||
|
|
||||||
|
def load(self):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def unload(self):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def transcribe(self, audio, language=None, fast=False, prompt=""):
|
||||||
|
return Transcript(text="hello world", language="en", confidence=0.99)
|
||||||
|
|
||||||
|
|
||||||
|
class FakeTranslator:
|
||||||
|
loaded = True
|
||||||
|
|
||||||
|
def load(self):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def unload(self):
|
||||||
|
pass
|
||||||
|
|
||||||
|
def translate(self, text, source_lang, target_lang, glossary=None):
|
||||||
|
return f"[{target_lang}] {text}"
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def engine(tmp_path, monkeypatch):
|
||||||
|
config = AppConfig()
|
||||||
|
config.models.preload_on_start = False
|
||||||
|
config.glossary.path = str(tmp_path / "glossary.json")
|
||||||
|
config.audio.partial_interval_ms = 0
|
||||||
|
config.audio.silence_ms = 300
|
||||||
|
config.audio.min_segment_ms = 200
|
||||||
|
|
||||||
|
lines = []
|
||||||
|
eng = TranslationEngine(config, on_line=lines.append)
|
||||||
|
monkeypatch.setattr(eng.models, "recognizer", lambda: FakeRecognizer())
|
||||||
|
monkeypatch.setattr(eng.models, "translator", lambda: FakeTranslator())
|
||||||
|
monkeypatch.setattr("hearo.core.engine.create_capture", lambda cfg: FakeCapture())
|
||||||
|
eng.lines = lines
|
||||||
|
yield eng
|
||||||
|
eng.shutdown()
|
||||||
|
|
||||||
|
|
||||||
|
def test_engine_produces_translated_lines(engine):
|
||||||
|
engine.start()
|
||||||
|
deadline = time.time() + 12
|
||||||
|
while time.time() < deadline and not engine.lines:
|
||||||
|
time.sleep(0.05)
|
||||||
|
engine.stop()
|
||||||
|
|
||||||
|
assert engine.lines, "12초 안에 자막이 한 줄도 나오지 않았습니다"
|
||||||
|
line = engine.lines[0]
|
||||||
|
assert line.source_text == "hello world"
|
||||||
|
assert line.translated_text == "[ko] hello world"
|
||||||
|
assert line.source_lang == "en"
|
||||||
|
assert line.is_final
|
||||||
|
assert line.latency_s >= 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_stop_is_idempotent(engine):
|
||||||
|
engine.start()
|
||||||
|
time.sleep(0.3)
|
||||||
|
engine.stop()
|
||||||
|
engine.stop()
|
||||||
|
assert not engine.running
|
||||||
|
|
||||||
|
|
||||||
|
def test_to_mono_16k_downmixes_and_resamples():
|
||||||
|
stereo_48k = np.tile(np.array([1.0, -1.0], dtype=np.float32), 48_000)
|
||||||
|
out = to_mono_16k(stereo_48k, channels=2, src_rate=48_000)
|
||||||
|
assert out.dtype == np.float32
|
||||||
|
assert abs(out.size - 16_000) <= 2 # 1초 분량
|
||||||
|
assert np.allclose(out, 0.0, atol=1e-6) # L+R 이 상쇄
|
||||||
|
|
||||||
|
|
||||||
|
def test_to_mono_16k_passthrough_when_already_correct():
|
||||||
|
mono = np.linspace(-1, 1, 16_000, dtype=np.float32)
|
||||||
|
out = to_mono_16k(mono, channels=1, src_rate=16_000)
|
||||||
|
assert np.array_equal(out, mono)
|
||||||
78
tests/test_segmenter.py
Normal file
78
tests/test_segmenter.py
Normal file
@@ -0,0 +1,78 @@
|
|||||||
|
import numpy as np
|
||||||
|
|
||||||
|
from hearo.audio.segmenter import Segmenter, SegmenterConfig
|
||||||
|
from hearo.constants import SAMPLE_RATE
|
||||||
|
|
||||||
|
|
||||||
|
def tone(ms: int, amplitude: float = 0.3) -> np.ndarray:
|
||||||
|
n = SAMPLE_RATE * ms // 1000
|
||||||
|
t = np.arange(n, dtype=np.float32) / SAMPLE_RATE
|
||||||
|
return (np.sin(2 * np.pi * 220 * t) * amplitude).astype(np.float32)
|
||||||
|
|
||||||
|
|
||||||
|
def silence(ms: int) -> np.ndarray:
|
||||||
|
return np.zeros(SAMPLE_RATE * ms // 1000, dtype=np.float32)
|
||||||
|
|
||||||
|
|
||||||
|
def config(**kw) -> SegmenterConfig:
|
||||||
|
base = dict(silence_ms=300, min_segment_ms=200, max_segment_ms=5000,
|
||||||
|
partial_interval_ms=0)
|
||||||
|
base.update(kw)
|
||||||
|
return SegmenterConfig(**base)
|
||||||
|
|
||||||
|
|
||||||
|
def test_speech_then_silence_emits_one_final_segment():
|
||||||
|
seg = Segmenter(config())
|
||||||
|
out = seg.push(silence(200))
|
||||||
|
assert out == []
|
||||||
|
out += seg.push(tone(800))
|
||||||
|
out += seg.push(silence(600))
|
||||||
|
finals = [s for s in out if s.is_final]
|
||||||
|
assert len(finals) == 1
|
||||||
|
assert 0.7 <= finals[0].duration_s <= 1.8
|
||||||
|
|
||||||
|
|
||||||
|
def test_short_blip_is_discarded():
|
||||||
|
seg = Segmenter(config(min_segment_ms=500))
|
||||||
|
out = seg.push(tone(60)) + seg.push(silence(600))
|
||||||
|
assert out == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_max_length_forces_a_cut():
|
||||||
|
seg = Segmenter(config(max_segment_ms=1000))
|
||||||
|
out = seg.push(tone(3000))
|
||||||
|
assert len([s for s in out if s.is_final]) >= 2
|
||||||
|
|
||||||
|
|
||||||
|
def test_partial_results_are_emitted_while_speaking():
|
||||||
|
seg = Segmenter(config(partial_interval_ms=300))
|
||||||
|
out = seg.push(tone(1500))
|
||||||
|
partials = [s for s in out if not s.is_final]
|
||||||
|
assert len(partials) >= 2
|
||||||
|
# 중간 결과는 누적된다
|
||||||
|
assert partials[-1].duration_s > partials[0].duration_s
|
||||||
|
|
||||||
|
|
||||||
|
def test_flush_returns_pending_speech():
|
||||||
|
seg = Segmenter(config())
|
||||||
|
assert seg.push(tone(700)) == [] # 아직 침묵이 안 왔으니 확정 없음
|
||||||
|
leftover = seg.flush()
|
||||||
|
assert leftover is not None and leftover.is_final
|
||||||
|
assert seg.flush() is None
|
||||||
|
|
||||||
|
|
||||||
|
def test_continuous_silence_never_emits():
|
||||||
|
seg = Segmenter(config())
|
||||||
|
for _ in range(10):
|
||||||
|
assert seg.push(silence(200)) == []
|
||||||
|
|
||||||
|
|
||||||
|
def test_lead_in_is_prepended():
|
||||||
|
"""발화 직전 오디오가 붙어 첫 음절이 잘리지 않아야 한다."""
|
||||||
|
seg = Segmenter(config(lead_in_ms=200))
|
||||||
|
seg.push(silence(400))
|
||||||
|
out = seg.push(tone(600)) + seg.push(silence(600))
|
||||||
|
finals = [s for s in out if s.is_final]
|
||||||
|
assert len(finals) == 1
|
||||||
|
# 0.6초 발화 + 최대 0.2초 리드인
|
||||||
|
assert finals[0].duration_s > 0.6
|
||||||
36
tests/test_tiers.py
Normal file
36
tests/test_tiers.py
Normal file
@@ -0,0 +1,36 @@
|
|||||||
|
from hearo.models.tiers import DEFAULT_TIER, TIERS, MTBackend, get_tier, ordered_tiers
|
||||||
|
|
||||||
|
|
||||||
|
def test_exactly_five_tiers_ordered_by_quality():
|
||||||
|
tiers = ordered_tiers()
|
||||||
|
assert len(tiers) == 5
|
||||||
|
assert [t.order for t in tiers] == [1, 2, 3, 4, 5]
|
||||||
|
|
||||||
|
|
||||||
|
def test_latency_and_vram_increase_monotonically():
|
||||||
|
tiers = ordered_tiers()
|
||||||
|
latencies = [t.approx_latency_s for t in tiers]
|
||||||
|
vram = [t.min_vram_gb for t in tiers]
|
||||||
|
assert latencies == sorted(latencies)
|
||||||
|
assert vram == sorted(vram)
|
||||||
|
|
||||||
|
|
||||||
|
def test_exactly_one_recommended_and_one_finetune_pick():
|
||||||
|
assert sum(t.recommended for t in TIERS.values()) == 1
|
||||||
|
assert sum(t.best_after_finetune for t in TIERS.values()) == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_llm_backends_support_prompt_glossary():
|
||||||
|
for tier in TIERS.values():
|
||||||
|
if tier.mt.backend is MTBackend.TRANSFORMERS:
|
||||||
|
assert tier.mt.supports_prompt_glossary
|
||||||
|
else:
|
||||||
|
assert not tier.mt.supports_prompt_glossary
|
||||||
|
|
||||||
|
|
||||||
|
def test_unknown_key_falls_back_to_default():
|
||||||
|
assert get_tier("does-not-exist").key == DEFAULT_TIER
|
||||||
|
|
||||||
|
|
||||||
|
def test_default_tier_exists():
|
||||||
|
assert DEFAULT_TIER in TIERS
|
||||||
102
tests/test_ui_smoke.py
Normal file
102
tests/test_ui_smoke.py
Normal file
@@ -0,0 +1,102 @@
|
|||||||
|
"""UI가 실제로 만들어지고 자막이 그려지는지 확인한다 (offscreen 렌더링)."""
|
||||||
|
|
||||||
|
from __future__ import annotations
|
||||||
|
|
||||||
|
import os
|
||||||
|
|
||||||
|
import pytest
|
||||||
|
|
||||||
|
pytest.importorskip("PySide6")
|
||||||
|
os.environ.setdefault("QT_QPA_PLATFORM", "offscreen")
|
||||||
|
|
||||||
|
from PySide6.QtGui import QImage, QPainter # noqa: E402
|
||||||
|
from PySide6.QtWidgets import QApplication # noqa: E402
|
||||||
|
|
||||||
|
from hearo.config import AppConfig # noqa: E402
|
||||||
|
from hearo.core.events import EngineState, EngineStatus, TranslationLine # noqa: E402
|
||||||
|
from hearo.ui.main_window import MainWindow # noqa: E402
|
||||||
|
from hearo.ui.overlay import paint_subtitle # noqa: E402
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture(scope="session")
|
||||||
|
def qapp():
|
||||||
|
app = QApplication.instance() or QApplication([])
|
||||||
|
yield app
|
||||||
|
|
||||||
|
|
||||||
|
@pytest.fixture
|
||||||
|
def window(qapp, tmp_path, monkeypatch):
|
||||||
|
monkeypatch.setattr("hearo.constants.user_data_dir", lambda: tmp_path)
|
||||||
|
config = AppConfig()
|
||||||
|
config.models.preload_on_start = False
|
||||||
|
config.glossary.path = str(tmp_path / "glossary.json")
|
||||||
|
win = MainWindow(config)
|
||||||
|
yield win
|
||||||
|
win.engine.shutdown()
|
||||||
|
win.overlay.close()
|
||||||
|
win.close()
|
||||||
|
|
||||||
|
|
||||||
|
def test_main_window_builds_all_pages(window):
|
||||||
|
assert window.stack.count() == 5
|
||||||
|
|
||||||
|
|
||||||
|
def test_navigation_switches_pages(window):
|
||||||
|
for index in range(5):
|
||||||
|
window._goto_page(index)
|
||||||
|
assert window.stack.currentIndex() == index
|
||||||
|
|
||||||
|
|
||||||
|
def test_status_updates_toggle_button(window):
|
||||||
|
window._on_status(EngineStatus(EngineState.RUNNING, "듣는 중"))
|
||||||
|
assert window.home_page.toggle_button.text() == "번역 정지"
|
||||||
|
window._on_status(EngineStatus(EngineState.IDLE, "대기 중"))
|
||||||
|
assert window.home_page.toggle_button.text() == "번역 시작"
|
||||||
|
|
||||||
|
|
||||||
|
def test_final_line_reaches_overlay_and_log(window):
|
||||||
|
window.config.log_transcripts = False
|
||||||
|
window._on_line(
|
||||||
|
TranslationLine("hello", "안녕하세요", "en", "ko", is_final=True, latency_s=1.2)
|
||||||
|
)
|
||||||
|
assert window.overlay._final_text == "안녕하세요"
|
||||||
|
assert window.home_page.log_list.count() == 1
|
||||||
|
|
||||||
|
|
||||||
|
def test_partial_line_is_not_logged(window):
|
||||||
|
window.config.log_transcripts = False
|
||||||
|
window._on_line(TranslationLine("hel", "", "en", "ko", is_final=False))
|
||||||
|
assert window.home_page.log_list.count() == 0
|
||||||
|
|
||||||
|
|
||||||
|
def test_tier_selection_updates_config(window):
|
||||||
|
window.models_page._on_tier_selected(window.models_page._cards[3].tier)
|
||||||
|
assert window.config.models.tier == "precision"
|
||||||
|
|
||||||
|
|
||||||
|
def test_subtitle_style_edit_propagates(window):
|
||||||
|
window.subtitle_page.size_spin.setValue(60)
|
||||||
|
assert window.config.subtitle.font_size == 60
|
||||||
|
|
||||||
|
|
||||||
|
def test_paint_subtitle_actually_draws_pixels(qapp):
|
||||||
|
"""자막 글씨가 투명 이미지 위에 실제로 그려지는지 (빈 화면이면 실패).
|
||||||
|
|
||||||
|
배경 상자가 픽셀을 채워버리면 글씨 유무를 알 수 없으므로 배경은 끄고 잰다.
|
||||||
|
"""
|
||||||
|
st = AppConfig().subtitle
|
||||||
|
st.background_opacity = 0
|
||||||
|
image = QImage(900, 200, QImage.Format.Format_ARGB32)
|
||||||
|
image.fill(0)
|
||||||
|
|
||||||
|
painter = QPainter(image)
|
||||||
|
paint_subtitle(painter, image.rect(), st, "안녕하세요 반갑습니다", "Hello there")
|
||||||
|
painter.end()
|
||||||
|
|
||||||
|
non_transparent = sum(
|
||||||
|
1
|
||||||
|
for y in range(0, image.height(), 4)
|
||||||
|
for x in range(0, image.width(), 4)
|
||||||
|
if image.pixelColor(x, y).alpha() > 0
|
||||||
|
)
|
||||||
|
assert non_transparent > 100
|
||||||
Reference in New Issue
Block a user