• v1.0.0 ccddbd6448

    v1.0.0
    Some checks failed
    Release / semantic-release (push) Successful in 42s
    Release / build-windows (push) Has been cancelled
    Release / build-macos (arm64, macos-latest) (push) Has been cancelled
    Release / build-macos (x64, macos-15-intel) (push) Has been cancelled
    Release / build-linux (push) Has been cancelled
    Release / release-main (push) Has been cancelled
    Release / release-develop (push) Has been cancelled
    tests / Unit tests (Linux, Python 3.11) (push) Has been cancelled
    Stable

    gitea-actions released this 2026-06-16 19:56:13 +09:00 | 30 commits to main since this release

    1.0.0 (2026-06-17)

    Features

    • 1080p60 NVENC selfbot broadcast (8 Mbps default) (ad0caa8)
    • bot: delay login until MeloTTS voice is warm (932aace)
    • bot: mirror heard voice turns to a text channel (62bb0ab)
    • bot: userbot (selfbot) mode — voice conversation on one user account (stage 1) (921a757)
    • brain: add Gemini CLI OAuth path for STREAM_BROWSER=false real-time search (b88def6)
    • brain: add OUTPUT_LANGUAGE reply-language lock (006a322)
    • brain: make OUTPUT_LANGUAGE lock robust on small models (8a2a109)
    • brain: route real-time search by live broadcast state (5d45d1d)
    • brain: wire STREAM_BROWSER real-time modes into the reply engine (browser + Gemini) (702fe80)
    • bridge: gate Whisper behind Silero VAD; harden broadcast auto-start (6d72e10)
    • browser-control server on host (real input) + remote-bot routing + ignore env backups (aebf183)
    • controlBrowser 'search' action + one-sentence voice replies (c21c5b2)
    • controlBrowser tool — human-operable on-screen Chrome (3d1e56f)
    • couple broadcast to voice + voice-controlled broadcast toggle (ca86390)
    • cross-platform compose (Ubuntu CDI + Windows Docker Desktop GPU) (3bdc7d0)
    • humanise selfbot voice-join and go-live pacing (b6cf05f)
    • per-call LLM timing, speaker ID, cancel captures on leave (44ebfea)
    • robust controlBrowser — real xdotool input, tabs/back/forward/popups (109dbc7)
    • selfbot: broadcast desktop audio + smart subtitles in the browse scenario (208fbbc)
    • selfbot: dedicated DISCORD_STREAM_TOKEN for the broadcast account (863337c)
    • settings web UI (models / STT / TTS speed / language / LLM instructions) (84e435f)
    • show per-stage timing (듣기/LLM/TTS) in the transcript channel (de5384d)
    • show speaker nickname instead of raw user ID in voice logs (54c3ce7)
    • split recording vs STT-processing time in turn logs (09afc21)
    • split-deployment roles (browser-host on LAN + remote bot) (1935c1a)
    • stream-test: drive the whole browse scenario with real input (2cdd159)
    • stream-test: persistent YouTube ad auto-skipper for the broadcast (e154404)
    • stream: single-account Go-Live — broadcast on the conversation's session (ddebdd7)
    • stream: STREAM_BROWSER flag + make toolbar-hide/subtitles broadcast-wide (ef6f6ff)
    • stream: true-mode browser-action core + Gemini scaffold + mode design (c420d5d)
    • stt-log: log the WHOLE turn pipeline to the transcript channel (e8234b7)
    • tts: add MeloTTS Korean voice via warm worker with offline-baked cache (b17961e)
    • weather: romanise non-Latin place names before geocoding (ccbaed9)

    🐛 Bug Fixes

    • always offer controlBrowser in screen-share so on-screen search works (f3d34ba)
    • bot: auto-start broadcast on voice join in userbot mode (568a1ae)
    • bot: never broadcast on the conversation's account (it deafened STT) (2e21185)
    • bot: userbot voice via @discordjs/voice 0.19 + DAVE (selfbot voice handshake fix) (8354591)
    • brain: drop --approval-mode yolo from Gemini CLI search (f54e2a4)
    • brain: recover colon-JSON and single tool_call object forms (c8a04a1)
    • brain: route auxiliary small-model calls to an available model (f3a1d92)
    • bridge: gate STT on real speech so noise doesn't trigger replies (39c7a22)
    • bridge: keep decimals, versions and URLs whole in TTS sentence split (d6c029d)
    • bridge: stop dropping real speech — disable avg_logprob confidence floor (f2b43cb)
    • cap selfbot stream -maxrate at lib's 10 Mbps ceiling; add stream-test tooling (1e30a49)
    • concise Korean weather reply (current conditions, one sentence) (3d620dc)
    • deterministic browser navigate for 'go back to ' (구글로 돌아가) (49061d3)
    • deterministic on-screen site search + lock replies to Korean (11a72cb)
    • deterministic weather → one clean Korean sentence (no 'Celsius') (c522e1b)
    • docker+brain: make container Chrome search work, harden, improve toggle (03375a7)
    • docker: in-container Gemini CLI OAuth, broadcast audio, ports + hardening (35e754d)
    • don't unlock active in startup catch when a newer attempt owns it (7a148f8)
    • drop webSearch when a site is named in screen-share, forcing controlBrowser (642fa42)
    • google anti-bot flag + persistent/safe settings apply + TTS engine wiring (247edda)
    • harden Korean-only output lock (front+end, explicit script ban) (d970bf2)
    • let melo-worker honour MELO_DEVICE from env (was hardcoded cpu) (b18217f)
    • make humanised selfbot startup abort- and concurrency-safe (2c7f0a9)
    • minimal Chrome flags (drop --test-type/AutomationControlled) + policy infobar suppress (8868381)
    • native tool-calling for qwen2.5 (tools actually fire) + kill Chrome infobars (4bc7f83)
    • persona uses settings output_language, matching the reply directive (7870a76)
    • search like a person — open homepage, type in the site's search box (8dd6386)
    • selfbot: gate broadcast 'live' on real Go-Live WebRTC connect (5961fae)
    • selfbot: smooth VNC capture via keepalive + stop ffmpeg leak on stream end (4176a68)
    • settings output_language overrides the compose env default (b3088dd)
    • single-pass NVENC encode for selfbot stream (no double encode) (40fd7db)
    • stream-test: hide Chrome toolbar in fullscreen so the address bar stays off the broadcast (c6a0ca4)
    • stream-test: refuse final box when element stays off-screen (8709f40)
    • stream-test: restore audio after ads, enforce subtitle rule broadcast-wide, commit the 60fps MV path (f93b241)
    • stream-test: restore audio after ads, enforce subtitle rule broadcast-wide, commit the 60fps MV path (0241628)
    • stream: enable captions only for real Korean tracks, skip auto-generated (3e33376)
    • stream: tie broadcast-helper to the stream lifecycle, enforce STREAM_BROWSER, fix fullscreen window (8aa2e4c)
    • weather passes named city from utterance; clean navigate reply (bdb012f)

    Performance Improvements

    • brain: env to disable pre-loop planner; cut voice latency (3a4776c)
    • brain: keep ollama models resident to kill voice cold-start latency (7792be2)
    • brain: pin chat model per-request, unload embeddings; default qwen2.5:3b (b91c05a)
    • bridge: lock STT to Korean + add per-stage turn timing (f12e6b2)
    • bridge: stream TTS per sentence to cut voice reply latency (5c29542)
    • conversational fast-path (skip enrichment) + shorter silence wait (37759f2)
    • memory: keep embed model warm across turns (keep_alive 0 -> 5m) (989a4f3)
    • pre-warm ollama at the engine's num_ctx (8192) so first turn is hot (5c11c5f)
    • pre-warm Whisper + chat model + TTS at bridge startup (d4e5e7f)
    • run MeloTTS on the GPU (cu128 torch) + warm CUDA at startup (927d59f)
    • stt: default Whisper model small -> medium for better Korean accuracy (4e446c1)
    • unify Ollama num_ctx so a voice turn keeps one resident model (2c38e75)

    📝 Documentation

    • docker: clarify userbot mode in compose/run-bot, bot token optional (f89246a)
    • env: mark DISCORD_BOT_TOKEN optional — blank runs userbot mode (40877b6)
    • readme: make Discord token setup userbot-first (67d0ae7)

    ♻️ Code Refactoring

    • stream-test: real-wheel into view, no synthetic-click fallback (bbc2fa3)
    Downloads