Add a drag handle on the top edge of the bottom log dock so its height can be
adjusted freely (80px .. 82vh), persisted in localStorage. Dock max-height
raised to 90vh; syncDockPad keeps the page bottom padding in step so the last
turn never hides behind the dock. Handle is hidden while the dock is collapsed.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a "🧠 모델 (STT · LLM)" collapsible panel with two dropdowns:
- STT(귀): tiny/base/small/medium/large-v3. Switching swaps WhisperSTT.model,
tears the worker down and re-warms it in the BACKGROUND (first switch to a
not-yet-downloaded size fetches it, so the HTTP call returns immediately and
the next utterance waits for the reload).
- LLM(두뇌): Haiku 4.5 / Sonnet 4.5. Applied on the next reply (no reload).
Backend: GET /api/models, POST /api/models/stt, POST /api/models/llm; Dashboard
gains models_settings/set_stt_model/set_llm_model. In-memory only (a service
restart reverts to the env defaults small / claude-haiku-4-5).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Dashboard changes requested by the user:
- Merge 화이트리스트/블랙리스트 into one "👤 유저 등록/관리" button (the modal
already holds both lists).
- Remove the 음성 인식(STT) 테스트 section and its client JS.
- Add collapsible "⚙️ 봇 관련 설정" with a barge-in toggle: "유저 목소리 들을 때
봇이 말하던 것 중지" (default on). Stored via /api/bot/settings; the bot reads
it on its report round-trip and calls voicePlayer.stop() when an allowed user
starts speaking (dave/bot.mjs).
- Add collapsible "🧹 로그 관련 설정" with a "잡음 로그 표시 안 함" toggle (default
on, localStorage-backed): hides turns whose 들음 is (빈 결과) or 답변 is [잡음].
- Fix the fixed event/error log dock covering the bottom-most turn: syncDockPad()
pads main by the dock height (on load, log render, toggle, resize).
BotControl gains a settings store; /api/bot/report now also returns settings.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The per-emotion refactor renamed the settings response keys but the POST
handler's log line still referenced res["settings"], throwing KeyError and
500ing every apply. Log base + overrides instead.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Each emotion can now be tuned independently. parse_segments resolves the four
controls per segment from a base (공통) dict plus an optional per-emotion
override; a missing override key inherits base. By default there are no
overrides, so every emotion delivers with the base values (모든 감정 = 기본값).
- emotion.py: Segment now carries all 4 controls + the canonical emotion name;
parse_segments(text, base, overrides). Adds EMOTION_LABELS/EMOTIONS for the UI.
- melo.py: MeloTTS.emotion_overrides store; synth resolves per-segment controls
and sends them per segment.
- melo_worker.py: _render applies each segment's own word_gap/sentence_gap/pitch
(previously reply-global).
- dashboard.py: emotion dropdown in the TTS panel; GET returns base + overrides
+ emotion list; POST {emotion,...} stores an override (or {reset:true} clears
it); base is set when emotion is omitted/"base".
Verified: an override on one emotion slows only that emotion (happy@0.7=5.66s vs
base 3.02s; sad unchanged at 3.06s); dashboard store/reset/base all work; 34
tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Change default voice controls to the user's tuned values: glyph speed 1.35,
word_gap -0.07s (-70ms), sentence_gap -0.30s, pitch 0 (env defaults, slider
initial values, and the 기본값 reset button all updated).
- Make the "봇 목소리(TTS) 조절" dashboard panel collapsible via its header;
starts collapsed (▸), expands on click (▾).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a "봇 목소리(TTS) 조절" panel to the voice-server dashboard so the four
controls can be tuned from the browser instead of only via env vars:
- 4 sliders (glyph speed, word gap, sentence gap, pitch) with live labels
- 미리듣기: synthesises a sample with the slider values WITHOUT changing the
live bot voice (new per-call overrides on MeloTTS.synth)
- 봇에 적용: commits the slider values to the live TTS instance; next reply uses
them. 기본값 button resets to the manual defaults.
Backend: GET/POST /api/tts/settings (clamped to the manual ranges) and POST
/api/tts/preview (returns audio/wav). Panel shows only when tts is a real
(non-mock) backend.
Verified end-to-end against a live dashboard instance: page renders the panel,
GET returns defaults, POST applies+clamps (pitch 99->12), preview returns a
valid wav and leaves the live settings unchanged; 29 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Persists per-day token usage (usage_store, ~/.config/wsai/usage.json, survives
restarts) and surfaces it in status as claude_usage.{today,week}. The navbar now
shows two cards — 오늘 토큰 / 주간 토큰 (input+output, with per-card tooltips
breaking down input/output tokens and requests) — replacing the session-only
token count.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The voice-channel picker already switched channels / left on 없음. Now the
server picker does too: choosing 없음 (or switching to a server with no channel
selected) sends a leave so the bot exits its current voice channel. Factored the
select POST into sendSelect(guildId, channelId).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
ClaudeBrain now returns per-reply token usage (Reply.usage from the API
response), the dashboard accumulates it (monitor.add_claude_usage), and the
header shows a "클로드 토큰" stat (input+output total, with a tooltip breaking
down input/output tokens and request count).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds the requested multi-dimension log search. Turns now carry speaker, guild
and channel (the bot sends X-User/Guild/Channel-Name on the voice-turn POST), and
a filter bar above the conversation feed narrows by 시간(최근 N분)·유저·서버·
채널·내용. The event/error panel keeps its text+level search.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Per-guild listen filter stored in the control plane (BotControl) with
GET/POST /api/bot/lists; the filter also rides along in the bot report
response so the bot always has the latest config.
- 화이트리스트/블랙리스트 popup: search the guild's members OR roles (type
selector), add/remove to white/black, save. Whitelist = listen to only those
(empty = everyone); blacklist = exclude. Reuses the shared 뒤로가기 modal.
- Bot reports guild roles + known members for the search UI, and filters
incoming audio via a pure, unit-tested isAllowed() (dave/filter.mjs):
blacklist always excludes; a non-empty whitelist restricts; else everyone.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a dashboard<->bot control plane (bot pushes state + polls commands, keeping
the bot's single outbound-HTTP direction):
- New bot_control.BotControl + endpoints: GET /api/bot/state, /api/bot/commands;
POST /api/bot/report, /api/bot/select.
- Dashboard header bar: bot identity/connection, server dropdown (top "없음"),
voice-channel dropdown (top "없음"), and live participant list.
- Turns record who spoke (Turn.speaker, via X-User-Name on the voice-turn POST).
- dave/bot.mjs: reports identity/guilds/voice-channels/members, polls join/leave
commands and joins dynamically, and sends the speaker's display name.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Bottom-docked collapsible terminal log panel (open/close), retains the event
log from voice-server start (event tail 200→2000).
- Log search box + level filter (전체/오류/경고/정보); per-line 삭제/수정 and
전체 삭제, backed by new monitor event ids and /api/logs/{clear,delete,edit}.
- Turns now show 들음 / 생각 / 답변 three rows; 생각 surfaces the emotion-tone
plan the bot chose (and the [잡음] decision), via a new Turn.thought field.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Adds a persisted, runtime-editable system prompt. The brain reads the persona
on every turn (prompt_store.get_persona), so a dashboard edit applies to the
next reply with no restart; blank clears the override back to the built-in
PERSONA. New endpoints GET/POST /api/prompt, and a reusable modal popup
(뒤로가기 + 수정/저장) that later white/blacklist features will share.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The dashboard voice-turn path ran _speech_text() before MeloTTS.synth, which
rewrote a leading "[힘차게] 안녕!" into "힘차게, 안녕!" — reading the first
emotion aloud and destroying the tag before the TTS emotion parser could use it.
Make _speech_text() a pass-through so every emotion tag (including the first)
reaches synth intact and shapes pitch/speed instead of being spoken. Adds a
regression test covering the leading-tag case.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- Brain persona now prefixes every reply with one bracketed emotion tag
(e.g. [반가움], [궁금]) and keeps replies to one or two short sentences.
- Empty/unrecognised audio (silence/noise) is reported as reply "[잡음]" with
no TTS playback instead of an empty reply.
- TTS speaks the bracketed emotion too: "[힘차게] 안녕!" is synthesised as
"힘차게, 안녕!" via a leading-tag -> spoken-word transform.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
A single Claude 529 Overloaded dropped the voice turn straight to the apology
fallback. The anthropic SDK retries >=500/429 but only twice by default, which
a busy window can outlast. Raise max_retries (WSAI_BRAIN_MAX_RETRIES, default 4)
so transient overloads recover silently, and give overloads their own spoken
fallback ("서버가 붐벼서...") distinct from generic failures. Kept modest so a
sustained outage still fails fast instead of leaving the bot silent.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The voice loop echoed the recognised text. Wire the real brain: the Discord
voice-turn now runs STT -> ClaudeBrain.respond (with rolling conversation
history) -> TTS, so the bot actually thinks and answers. --voice-server builds
the brain by default (WSAI_BRAIN=claude, WSAI_BRAIN_MODEL overridable) and
gracefully falls back to echo if anthropic/Claude auth is unavailable. A brain
error speaks a short apology instead of killing the loop.
Verified end-to-end: an utterance wav returns X-Heard plus a distinct Claude
X-Reply and a synthesised reply wav on device=cuda. 12 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The bot (dave/bot.mjs) previously only joined the channel and counted audio
frames — it never fed STT or spoke back. Wire the real loop:
- Node bot: buffer each speaker's Opus->PCM utterance until AfterSilence,
wrap as WAV, POST to the Python voice-turn endpoint, then play the returned
reply wav into the channel via an AudioPlayer (ffmpeg->Opus). Skips its own
audio, dedupes overlapping subscriptions, and ignores sub-0.35s noise.
- Python: new `python -m wsai --voice-server` serves /api/voice-turn — decode
the uploaded utterance, GPU faster-whisper STT, produce a reply (echo of what
was heard for now), GPU MeloTTS synth, return the reply wav (recognised/reply
text ride along as X-Heard/X-Reply headers). Both engines pre-warmed; turns
show in the dashboard feed. MeloTTS.synth() extracted for direct wav reuse.
Echo mode verifies listening+speaking+GPU recognition entirely in Discord; the
Claude brain is the next slice. Verified the endpoint round-trip: utterance wav
-> correct Korean X-Heard/X-Reply + a WAVE reply on device=cuda. 12 tests pass,
node --check clean.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The status page was view-only, so there was no way to actually verify Korean
recognition end-to-end. Add a live test: record from the mic (localhost/https)
or upload an audio file (works over LAN http, where browsers block getUserMedia),
POST it to a new /api/stt endpoint that ffmpeg-normalises the blob to 16 kHz
mono and runs the real GPU faster-whisper, then shows the recognised text +
latency + device. Results also land in the live turn feed.
The dashboard now optionally holds a WhisperSTT and drives it from a private
asyncio loop thread. New `python -m wsai --stt-test` serves the page with STT
enabled and pre-warms the GPU worker so the first recognition is instant.
WhisperSTT.resolved_device is exposed for the UI.
Verified: wav and browser-style webm/opus uploads both return the correct
Korean text on device=cuda in ~240-280ms. 12 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
--dashboard defaulted to looping mock STT forever, so the status page
piled up thousands of fake "conversations" (all mock, 0ms, same reply)
that looked like real traffic. Now it plays 3 sample utterances then
idles; a loud 데모 모드 banner states the turns are mock samples, not
real STT/Brain/TTS. Continuous demo moved behind --dashboard-loop-demo.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Add a stdlib-only observability site so you can open a browser and watch,
step by step: whether it is listening, what it heard, what the brain thought
and answered, how long each stage took, and whether anything errored.
- wsai/monitor.py: thread-safe telemetry hub (per-turn timed steps, status
header, error log) with a pub/sub for live push.
- wsai/dashboard.py: stdlib http.server serving a self-contained page plus an
SSE (/events) live stream; /api/state snapshot fallback.
- Pipeline emits step-by-step turn telemetry (화면 맥락 → 두뇌 → 응답) and
listening/running status; optional monitor, so existing paths are untouched.
- `python -m wsai --dashboard` starts the site (0.0.0.0:8787, WSAI_DASHBOARD_PORT)
and loops the mock voice demo so there is always live activity to watch.
- Tests cover turn recording, per-step timing, error marking, and live push.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>