Sonnet 5 reaches first token ~0.4s sooner than Sonnet 4.5 but answers the
same voice prompt more verbosely (measured 46-54 vs ~28 output tokens),
which erased the win in total turn time. Add a cached, persona-independent
brevity system block that pulls output back to ~30 tokens, so the faster
first token becomes a faster, lower-variance whole reply.
Measured (16-round interleaved A/B, production-shaped call):
sonnet-4-5 TTFT 1.26s total 2.02s (tail 3.26s) out 29
sonnet-5+brev TTFT 0.85s total 1.70s (tail 2.28s) out 30
- ClaudeBrain: default model claude-sonnet-5 + BREVITY block (cached with
the persona prefix so a dashboard persona edit can't drop it).
- Defaults aligned: config.Settings.anthropic_model and the voice-server
WSAI_BRAIN_MODEL default -> claude-sonnet-5.
- Dashboard: add claude-sonnet-5 to LLM_OPTIONS + JS label; fix stale
restart hint.
- tests/latency_ab.py: reproducible model-latency A/B harness (reads
CLAUDE_CREDENTIALS_PATH; makes live API calls, so not a pytest test).
Streaming TTS was intentionally not added: this bot answers in one sentence,
where sentence-level streaming has no overlap to exploit, and it would
require rearchitecting both the Python endpoint and the node playback.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Expose the barge-in sustained-speech threshold as a dashboard bot setting
"유저 음성 인식 시간" (ms). Server clamps 0..5000, persists via state_store,
and the bot reads it each report round-trip; WSAI_BARGE_IN_MS/700 stays the
fallback default.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Instead of immediately discarding a fragment like "그러면", hold it briefly
(WSAI_FRAGMENT_MERGE_MS, default 5s). If the next utterance arrives in time,
merge it in front and answer the whole thing ("그러면" + "뭐 먹지"); if nothing
follows within the window, it's silently dropped. Also redeployed the wsai-bot
container so the sustained (700ms) barge-in bot.mjs is now live.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Add a per-turn ✕ button on each 대화 카드 and a "🗑 전체 삭제" button in the turn
filter bar. Monitor gains delete_turn(id)/clear_turns() which broadcast
turn_deleted/turns_cleared; the page removes them live (all connected clients).
Routes: POST /api/turns/delete, POST /api/turns/clear.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
STT (whisper_worker): add VAD tuning (speech_pad_ms so soft first/last words
aren't clipped), condition_on_previous_text=False + temperature fallback +
no_speech/logprob/compression thresholds to reject noisy/quiet decodes, drop
per-segment non-speech, and a hallucination guard that blanks Whisper's classic
Korean silence/noise boilerplate ("감사합니다" 등) when no_speech_prob is high.
Barge-in (bot.mjs): stop the bot's TTS only when the speaking (green ring) stays
on for >= WSAI_BARGE_IN_MS (default 700ms), not on the first blip — cancelled if
speaking stops in time. AfterSilence 800->1000ms so trailing soft words finish.
Fragment gate (dashboard): a lone connective filler ("그러면") carries no
answerable intent -> discard as [대기] instead of replying; hidden by the noise
filter like [잡음].
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Voice replies were too long and slow (~1.3s LLM). Cap output at 150 tokens
(WSAI_BRAIN_MAX_TOKENS), rewrite the persona to demand one short sentence with
no boilerplate self-intro / "무엇을 도와드릴까요" padding, and mark the static
system prompt with cache_control so repeat turns skip re-processing it.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- 유저 등록/관리 search now has 4 types: 유저(non-bot members), 봇(bot members),
통화방(current voice-channel members, bot flag enriched), 역할(roles — expandable
to the members under each role; register the role OR a member under it).
- Persist settings to a JSON state store (WSAI_STATE_FILE, default
~/.config/wsai/state.json) so they survive a service/container restart:
per-guild listen lists + bargeIn (BotControl), TTS base/overrides and the
chosen STT/LLM models (Dashboard). Applied on voice-server startup before warm.
- TTS 감정 선택에 최상단 "전체변경 (모든 감정)" 추가: 적용하면 base를 그 값으로 세팅하고
모든 per-emotion override를 지워 전 감정이 한 번에 그 값으로 말합니다.
- Move the component status bar (눈/시각/귀/LLM/입) above the bot bar.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- voice_turn now records three separate timed steps (STT, LLM, TTS) instead of
one combined "STT+두뇌+TTS" step, so each stage's latency shows in the turn.
- Rename 두뇌 -> LLM everywhere user-facing: component chip, sub header, demo
banner, model label, brain-failure log, startup print, pipeline step name.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Bump the base TTS speed default (WSAI_TTS_SPEED, slider default/label, dashboard
fallbacks) so the bot speaks a bit faster by default.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The bot reports identity as {id, username, tag}; the CONNECT log concatenated
that dict onto a string, raising TypeError and 500ing /api/bot/report (breaking
the bot's command/settings round-trip). Extract a string field (tag/username/id)
before formatting.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Default STT model changed small -> medium (WhisperSTT default + service env;
medium pre-cached). UI labels/hints updated.
- Model 적용 buttons: if the picked model == current, flash "변경사항 없음" for 3s
then show the current model again; if it actually changes, STT shows a live
loading panel that polls /api/models until the worker is ready and logs
completion to the event log (set_stt_model/set_llm_model now return `changed`).
- Event/error log gains categories (READY, CONNECT, MODEL, TTS, VOICE, FILTER,
SETTING, TURN, BRAIN, VISION, PIPELINE, …): Monitor.log takes an optional `cat`,
call sites tagged, and each line shows a coloured category chip. A bot
(re)connection now logs a CONNECT event.
- Log search extended: filter by 레벨(type), 종류(category, auto-populated), and
시간(time range) in addition to free text.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Add a drag handle on the top edge of the bottom log dock so its height can be
adjusted freely (80px .. 82vh), persisted in localStorage. Dock max-height
raised to 90vh; syncDockPad keeps the page bottom padding in step so the last
turn never hides behind the dock. Handle is hidden while the dock is collapsed.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a "🧠 모델 (STT · LLM)" collapsible panel with two dropdowns:
- STT(귀): tiny/base/small/medium/large-v3. Switching swaps WhisperSTT.model,
tears the worker down and re-warms it in the BACKGROUND (first switch to a
not-yet-downloaded size fetches it, so the HTTP call returns immediately and
the next utterance waits for the reload).
- LLM(두뇌): Haiku 4.5 / Sonnet 4.5. Applied on the next reply (no reload).
Backend: GET /api/models, POST /api/models/stt, POST /api/models/llm; Dashboard
gains models_settings/set_stt_model/set_llm_model. In-memory only (a service
restart reverts to the env defaults small / claude-haiku-4-5).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Dashboard changes requested by the user:
- Merge 화이트리스트/블랙리스트 into one "👤 유저 등록/관리" button (the modal
already holds both lists).
- Remove the 음성 인식(STT) 테스트 section and its client JS.
- Add collapsible "⚙️ 봇 관련 설정" with a barge-in toggle: "유저 목소리 들을 때
봇이 말하던 것 중지" (default on). Stored via /api/bot/settings; the bot reads
it on its report round-trip and calls voicePlayer.stop() when an allowed user
starts speaking (dave/bot.mjs).
- Add collapsible "🧹 로그 관련 설정" with a "잡음 로그 표시 안 함" toggle (default
on, localStorage-backed): hides turns whose 들음 is (빈 결과) or 답변 is [잡음].
- Fix the fixed event/error log dock covering the bottom-most turn: syncDockPad()
pads main by the dock height (on load, log render, toggle, resize).
BotControl gains a settings store; /api/bot/report now also returns settings.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The per-emotion refactor renamed the settings response keys but the POST
handler's log line still referenced res["settings"], throwing KeyError and
500ing every apply. Log base + overrides instead.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Each emotion can now be tuned independently. parse_segments resolves the four
controls per segment from a base (공통) dict plus an optional per-emotion
override; a missing override key inherits base. By default there are no
overrides, so every emotion delivers with the base values (모든 감정 = 기본값).
- emotion.py: Segment now carries all 4 controls + the canonical emotion name;
parse_segments(text, base, overrides). Adds EMOTION_LABELS/EMOTIONS for the UI.
- melo.py: MeloTTS.emotion_overrides store; synth resolves per-segment controls
and sends them per segment.
- melo_worker.py: _render applies each segment's own word_gap/sentence_gap/pitch
(previously reply-global).
- dashboard.py: emotion dropdown in the TTS panel; GET returns base + overrides
+ emotion list; POST {emotion,...} stores an override (or {reset:true} clears
it); base is set when emotion is omitted/"base".
Verified: an override on one emotion slows only that emotion (happy@0.7=5.66s vs
base 3.02s; sad unchanged at 3.06s); dashboard store/reset/base all work; 34
tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Change default voice controls to the user's tuned values: glyph speed 1.35,
word_gap -0.07s (-70ms), sentence_gap -0.30s, pitch 0 (env defaults, slider
initial values, and the 기본값 reset button all updated).
- Make the "봇 목소리(TTS) 조절" dashboard panel collapsible via its header;
starts collapsed (▸), expands on click (▾).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a "봇 목소리(TTS) 조절" panel to the voice-server dashboard so the four
controls can be tuned from the browser instead of only via env vars:
- 4 sliders (glyph speed, word gap, sentence gap, pitch) with live labels
- 미리듣기: synthesises a sample with the slider values WITHOUT changing the
live bot voice (new per-call overrides on MeloTTS.synth)
- 봇에 적용: commits the slider values to the live TTS instance; next reply uses
them. 기본값 button resets to the manual defaults.
Backend: GET/POST /api/tts/settings (clamped to the manual ranges) and POST
/api/tts/preview (returns audio/wav). Panel shows only when tts is a real
(non-mock) backend.
Verified end-to-end against a live dashboard instance: page renders the panel,
GET returns defaults, POST applies+clamps (pitch 99->12), preview returns a
valid wav and leaves the live settings unchanged; 29 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Ports the 4 independent voice controls documented in tts_site's manual into the
MeloTTS backend, layered on top of the existing emotion-tag segments:
- glyph speed: generation-stage length_scale (existing speed path); default
bumped 1.2 -> 1.25 to match the manual (still well below the 1.5 that slurred)
- word_gap (sec, -0.2..0.5): scale intra-sentence silences after natural synth
- sentence_gap (sec, -0.5..1.5): insert/trim silence at sentence boundaries,
replacing the old fixed 120ms inter-segment gap
- pitch (semitones, -12..12): global offset added on top of per-emotion pitch
Each reply is split into sentences, synthesised per-sentence at its segment's
speed, word_gap applied, joined with sentence_gap, then pitch-shifted. Exposed
via WSAI_TTS_SPEED/WORD_GAP/SENTENCE_GAP/PITCH env vars and MeloTTS ctor args.
Verified by real synthesis: each control independently changes wav duration
(speed, sentence_gap, word_gap-with-pauses, pitch); 29 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Voice TTS (MeloTTS) is Korean-only, so any non-Korean reply gets
mangled. Remove the persona's language-switch escape hatch so the AI
always answers in Korean even when the user requests another language.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- samples/voice/: 6 XTTS v2 female studio voices speaking the same Korean
line, as clearer-pronunciation alternatives to the single MeloTTS speaker
- samples/emotion/: per-[emotion] delivery clips from the current MeloTTS engine
- samples/README.md documents both sets and the tradeoffs
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
User reported the 1.5x base made Korean pronunciation mushy/slurred. Dial the
default WSAI_TTS_SPEED back to 1.2: still noticeably faster than the original
1.0, but clear. Emotion multipliers scale off base as before.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
User reported the emotion samples sounded the same and the speed change wasn't
noticeable. Two changes:
- Add a "base" (기본/neutral) emotion at speed 1.0x so the plain base voice can
be selected explicitly via a [기본] tag and auditioned against the others.
- Bump the WSAI_TTS_SPEED default 1.15 -> 1.5 for a clearly faster base voice.
Emotion multipliers scale off base, so every emotion speeds up together.
Also extends gen_emotion_samples.py to emit one wav per emotion (incl. 기본)
plus a stitched all-in-one, so each emotion can be delivered as a separate clip.
Verified: 29 tests pass; match_emotion('기본') == 'base'; per-emotion synthesis
at base 1.5 produces distinct clip durations (base 4.9s vs happy 3.3s).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
User asked for a bit faster speech. Pitch shift is already disabled, so the
earlier "monster voice" risk from stacking pitch on a fast base is gone — a
modest 1.15x base is natural. Emotion multipliers scale off base, so every
emotion gets the bump too. Also adds tests/gen_emotion_samples.py, a dev
utility that synthesizes one clip per canonical emotion (announce name at
neutral speed, then a sample sentence steered by that emotion's tag) and
stitches them into a single wav for auditioning the emotion palette.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The emotional TTS path applied a librosa post-hoc pitch_shift (±1–3 semitones)
on top of a 1.3x-fast Melo base, producing a robotic "monster" delivery. Zero
out the pitch column for every emotion so pitch_shift is never invoked (the
worker's _pitch_shift already no-ops on 0.0), and drop the default synthesis
speed to 1.0. Emotion is now conveyed by speed alone — natural, artefact-free.
The pitch column is retained so a proper pitch method can be re-enabled later.
Verified: 29 tests pass; real 2-emotion synthesis on CUDA yields a clean wav
with pitch=0.0 on all segments.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The STT/TTS worker _ensure() treated a spawned-but-not-yet-handshaked
subprocess as ready, so a voice turn arriving during warmup read the same
stdout StreamReader concurrently with the warmup handshake and crashed with
"readuntil() called while another coroutine is already waiting for incoming
data". Add a _start_lock + _ready flag so (re)start and the ready handshake
run atomically and callers wait for real readiness before reading stdout.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
To show every server user (not just cached speakers) in the whitelist/blacklist
popup, the bot needs the privileged Server Members Intent. Declaring it while the
Developer Portal toggle is off breaks login, so it's gated behind
WSAI_MEMBERS_INTENT=1: when set, the bot adds GuildMembers intent, fetches each
guild's full member list on ready, and reports it (cap 200→2000). Default off =
unchanged behavior, safe to deploy.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Persists per-day token usage (usage_store, ~/.config/wsai/usage.json, survives
restarts) and surfaces it in status as claude_usage.{today,week}. The navbar now
shows two cards — 오늘 토큰 / 주간 토큰 (input+output, with per-card tooltips
breaking down input/output tokens and requests) — replacing the session-only
token count.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The voice-channel picker already switched channels / left on 없음. Now the
server picker does too: choosing 없음 (or switching to a server with no channel
selected) sends a leave so the bot exits its current voice channel. Factored the
select POST into sendSelect(guildId, channelId).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
ClaudeBrain now returns per-reply token usage (Reply.usage from the API
response), the dashboard accumulates it (monitor.add_claude_usage), and the
header shows a "클로드 토큰" stat (input+output total, with a tooltip breaking
down input/output tokens and request count).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Redefines the bot as "디스코드를 이용해 실시간으로 대화하는 AI 인공지능" and drops
all screen-share wording. Organizes the one-line prompt into sections (역할·언어·
답변방식·대화태도·안전/사실성·감정표현·정체성). Adds: always-Korean-unless-asked,
길이 정량화(한두 문장·10초), 되묻기/침묵 무시, 불확실·최신 정보는 "확인 필요",
URL·숫자·코드 풀어 읽기. Keeps the [감정] tag section (required by the emotion TTS).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds the requested multi-dimension log search. Turns now carry speaker, guild
and channel (the bot sends X-User/Guild/Channel-Name on the voice-turn POST), and
a filter bar above the conversation feed narrows by 시간(최근 N분)·유저·서버·
채널·내용. The event/error panel keeps its text+level search.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Records what's now built beyond the original plan: real GPU STT/brain/TTS via
the voice-server, bracketed-emotion TTS (pitch/speed), and the dashboard's
prompt editing, bot control bar, whitelist/blacklist, 3-row turns, and log dock,
plus the dashboard<->bot control plane.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Per-guild listen filter stored in the control plane (BotControl) with
GET/POST /api/bot/lists; the filter also rides along in the bot report
response so the bot always has the latest config.
- 화이트리스트/블랙리스트 popup: search the guild's members OR roles (type
selector), add/remove to white/black, save. Whitelist = listen to only those
(empty = everyone); blacklist = exclude. Reuses the shared 뒤로가기 modal.
- Bot reports guild roles + known members for the search UI, and filters
incoming audio via a pure, unit-tested isAllowed() (dave/filter.mjs):
blacklist always excludes; a non-empty whitelist restricts; else everyone.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
joinChannel reused a same-guild connection and re-subscribed a new player and
receiver each time, stacking duplicate speaking listeners (→ duplicate voice
turns) and error handlers. Now it no-ops if already in the target channel and
otherwise leaves the current connection first, so channel switches are clean.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a dashboard<->bot control plane (bot pushes state + polls commands, keeping
the bot's single outbound-HTTP direction):
- New bot_control.BotControl + endpoints: GET /api/bot/state, /api/bot/commands;
POST /api/bot/report, /api/bot/select.
- Dashboard header bar: bot identity/connection, server dropdown (top "없음"),
voice-channel dropdown (top "없음"), and live participant list.
- Turns record who spoke (Turn.speaker, via X-User-Name on the voice-turn POST).
- dave/bot.mjs: reports identity/guilds/voice-channels/members, polls join/leave
commands and joins dynamically, and sends the speaker's display name.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Bottom-docked collapsible terminal log panel (open/close), retains the event
log from voice-server start (event tail 200→2000).
- Log search box + level filter (전체/오류/경고/정보); per-line 삭제/수정 and
전체 삭제, backed by new monitor event ids and /api/logs/{clear,delete,edit}.
- Turns now show 들음 / 생각 / 답변 three rows; 생각 surfaces the emotion-tone
plan the bot chose (and the [잡음] decision), via a new Turn.thought field.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a persisted, runtime-editable system prompt. The brain reads the persona
on every turn (prompt_store.get_persona), so a dashboard edit applies to the
next reply with no restart; blank clears the override back to the built-in
PERSONA. New endpoints GET/POST /api/prompt, and a reusable modal popup
(뒤로가기 + 수정/저장) that later white/blacklist features will share.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The dashboard voice-turn path ran _speech_text() before MeloTTS.synth, which
rewrote a leading "[힘차게] 안녕!" into "힘차게, 안녕!" — reading the first
emotion aloud and destroying the tag before the TTS emotion parser could use it.
Make _speech_text() a pass-through so every emotion tag (including the first)
reaches synth intact and shapes pitch/speed instead of being spoken. Adds a
regression test covering the leading-tag case.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Emotion tags now steer delivery rather than being read aloud. parse_segments()
splits a reply on [감정] tags: a recognised emotion word switches the pitch and
speed of the text that follows (and is dropped), while a non-emotion bracket
(e.g. [1번]) keeps its inner words as spoken content. Emotions can change
mid-reply, so a single turn is synthesised as several pitch-shifted segments and
concatenated in the melo worker (librosa pitch_shift, warmed at startup).
The emotion vocabulary is grounded in Azure Neural TTS speaking styles plus
Ekman's basic emotions, with Korean synonyms. The brain persona is updated to
emit inline tags from that set.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Brain persona now prefixes every reply with one bracketed emotion tag
(e.g. [반가움], [궁금]) and keeps replies to one or two short sentences.
- Empty/unrecognised audio (silence/noise) is reported as reply "[잡음]" with
no TTS playback instead of an empty reply.
- TTS speaks the bracketed emotion too: "[힘차게] 안녕!" is synthesised as
"힘차게, 안녕!" via a leading-tag -> spoken-word transform.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Log voice connection errors instead of letting them surface silently.
- Collapse bursty repeated receive-stream errors (DAVE E2EE group-transition
decrypt failures) into one line + a suppressed-count summary, so a member
joining/leaving no longer floods the log.
Deploy: voice-server now runs as the wsai-voice.service user unit (STT+Claude
Haiku brain+TTS on GPU); the bot container reaches it via
host.docker.internal:8787.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Claude replies with markdown/backticks by default; MeloTTS's Korean text
normaliser has no entry for '`' and dies with KeyError: '`', so any reply
mentioning a command/code block crashed the whole voice turn (500 on
/api/voice-turn). Fix at the shared synth() choke point with
normalize_for_speech(), which flattens code fences/inline code/links/markdown
and guarantees no backtick reaches the worker — covering both the dashboard
voice turn and the Discord speak() bridge. Also add a PERSONA line asking the
model to avoid markdown (belt-and-suspenders; the code strip is the real fix).
errors_total never moved for turn-level failures: it was only bumped by
log("error") events, and the dashboard voice path calls turn.finish(error=...)
without logging. Emit one error-level log event from Turn.finish() when a turn
ends in error, so both the server counter and the browser SSE mirror stay
consistent, guarded to count at most once. Drop the now-redundant pipeline
log("error") to avoid double counting and remove the dead _publish stub.
Verified: raw backtick -> worker KeyError '`' reproduced; after fix real
MeloTTS synth of a backtick+fenced reply succeeds; /api/voice-turn returns 200
with a wav body on a backtick reply and errors_total stays 0, and an induced
synth failure returns 500 with errors_total incrementing to exactly 1. Full
suite 18 passed.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
A single Claude 529 Overloaded dropped the voice turn straight to the apology
fallback. The anthropic SDK retries >=500/429 but only twice by default, which
a busy window can outlast. Raise max_retries (WSAI_BRAIN_MAX_RETRIES, default 4)
so transient overloads recover silently, and give overloads their own spoken
fallback ("서버가 붐벼서...") distinct from generic failures. Kept modest so a
sustained outage still fails fast instead of leaving the bot silent.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The voice loop echoed the recognised text. Wire the real brain: the Discord
voice-turn now runs STT -> ClaudeBrain.respond (with rolling conversation
history) -> TTS, so the bot actually thinks and answers. --voice-server builds
the brain by default (WSAI_BRAIN=claude, WSAI_BRAIN_MODEL overridable) and
gracefully falls back to echo if anthropic/Claude auth is unavailable. A brain
error speaks a short apology instead of killing the loop.
Verified end-to-end: an utterance wav returns X-Heard plus a distinct Claude
X-Reply and a synthesised reply wav on device=cuda. 12 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Raise the voice-Ready ceiling 20s->40s: the DAVE/MLS handshake cycles
signalling<->connecting and can take ~25s, so 20s spuriously failed the join.
- Handle AudioReceiveStream 'error' (e.g. a DAVE decrypt/UDP GenericFailure on
one packet): log and free the speaker slot instead of letting the unhandled
'error' event crash the whole bot process.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The bot (dave/bot.mjs) previously only joined the channel and counted audio
frames — it never fed STT or spoke back. Wire the real loop:
- Node bot: buffer each speaker's Opus->PCM utterance until AfterSilence,
wrap as WAV, POST to the Python voice-turn endpoint, then play the returned
reply wav into the channel via an AudioPlayer (ffmpeg->Opus). Skips its own
audio, dedupes overlapping subscriptions, and ignores sub-0.35s noise.
- Python: new `python -m wsai --voice-server` serves /api/voice-turn — decode
the uploaded utterance, GPU faster-whisper STT, produce a reply (echo of what
was heard for now), GPU MeloTTS synth, return the reply wav (recognised/reply
text ride along as X-Heard/X-Reply headers). Both engines pre-warmed; turns
show in the dashboard feed. MeloTTS.synth() extracted for direct wav reuse.
Echo mode verifies listening+speaking+GPU recognition entirely in Discord; the
Claude brain is the next slice. Verified the endpoint round-trip: utterance wav
-> correct Korean X-Heard/X-Reply + a WAVE reply on device=cuda. 12 tests pass,
node --check clean.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The status page was view-only, so there was no way to actually verify Korean
recognition end-to-end. Add a live test: record from the mic (localhost/https)
or upload an audio file (works over LAN http, where browsers block getUserMedia),
POST it to a new /api/stt endpoint that ffmpeg-normalises the blob to 16 kHz
mono and runs the real GPU faster-whisper, then shows the recognised text +
latency + device. Results also land in the live turn feed.
The dashboard now optionally holds a WhisperSTT and drives it from a private
asyncio loop thread. New `python -m wsai --stt-test` serves the page with STT
enabled and pre-warms the GPU worker so the first recognition is instant.
WhisperSTT.resolved_device is exposed for the UI.
Verified: wav and browser-style webm/opus uploads both return the correct
Korean text on device=cuda in ~240-280ms. 12 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The first CUDA inference pays a large lazy cost (kernel autotune/cudnn) —
~10s for a cold TTS synth — which would blow the voice loop's ~1s budget on
the very first reply. Each worker now runs one dummy inference (TTS: a short
phrase; STT: 1s of silence) after model load and before emitting "ready", so
"ready" means "hot". Warmup failures are logged and never block startup.
Verified: first real call after startup is now TTS ~238ms / STT ~189ms
(was ~11s cold for TTS). 12 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>