- Remove the 'instruct' style field: Qwen3-TTS Base generate_voice_clone
silently ignores it, so the UI option did nothing.
- Serialize generations with a lock so concurrent requests queue instead
of doubling RAM/VRAM on one shared model.
- Validate empty text/ref_text ourselves so the API returns the Korean
400 messages instead of FastAPI's generic 422.
- UI: explain that mic recording needs https/localhost instead of a raw
TypeError on plain-http LAN access.
- README: document mic/https limit, queueing, and measured CPU speed.
Verified on CPU (Python 3.12, torch 2.14 CPU, 0.6B Base): wav and mp3
references, language ko/auto, two concurrent requests -> all 200 and
Whisper transcripts match the requested text.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Extract the voice-cloning feature from jamiepine/voicebox (Qwen3-TTS Base
engine) into a small standalone app: FastAPI backend (engine.py + app.py)
wrapping create_voice_clone_prompt/generate_voice_clone, and a single-page
UI (upload or mic-record a reference clip + transcript -> synthesize any
text in that voice). Supports 10 languages incl. Korean; model loads lazily
and downloads on first use.