Commit Graph

26 Commits

Author SHA1 Message Date
EJClaw
7c899d3b19 feat(tts): raise default glyph speed 1.35 -> 1.4
Bump the base TTS speed default (WSAI_TTS_SPEED, slider default/label, dashboard
fallbacks) so the bot speaks a bit faster by default.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-27 00:00:06 +09:00
EJClaw
ed5f328889 fix(dashboard): CONNECT log TypeError when bot identity is a dict
The bot reports identity as {id, username, tag}; the CONNECT log concatenated
that dict onto a string, raising TypeError and 500ing /api/bot/report (breaking
the bot's command/settings round-trip). Extract a string field (tag/username/id)
before formatting.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:54:54 +09:00
EJClaw
1ac45214ce feat(dashboard): default STT=medium, apply feedback+loading, log categories+filters
- Default STT model changed small -> medium (WhisperSTT default + service env;
  medium pre-cached). UI labels/hints updated.
- Model 적용 buttons: if the picked model == current, flash "변경사항 없음" for 3s
  then show the current model again; if it actually changes, STT shows a live
  loading panel that polls /api/models until the worker is ready and logs
  completion to the event log (set_stt_model/set_llm_model now return `changed`).
- Event/error log gains categories (READY, CONNECT, MODEL, TTS, VOICE, FILTER,
  SETTING, TURN, BRAIN, VISION, PIPELINE, …): Monitor.log takes an optional `cat`,
  call sites tagged, and each line shows a coloured category chip. A bot
  (re)connection now logs a CONNECT event.
- Log search extended: filter by 레벨(type), 종류(category, auto-populated), and
  시간(time range) in addition to free text.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:51:58 +09:00
EJClaw
d164630bb8 feat(dashboard): drag-resizable event/error log dock height (persisted)
Add a drag handle on the top edge of the bottom log dock so its height can be
adjusted freely (80px .. 82vh), persisted in localStorage. Dock max-height
raised to 90vh; syncDockPad keeps the page bottom padding in step so the last
turn never hides behind the dock. Handle is hidden while the dock is collapsed.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:40:56 +09:00
EJClaw
6d8ba3ad3a feat(dashboard): live STT/LLM model switching (small<->medium, Haiku<->Sonnet)
Adds a "🧠 모델 (STT · LLM)" collapsible panel with two dropdowns:
- STT(귀): tiny/base/small/medium/large-v3. Switching swaps WhisperSTT.model,
  tears the worker down and re-warms it in the BACKGROUND (first switch to a
  not-yet-downloaded size fetches it, so the HTTP call returns immediately and
  the next utterance waits for the reload).
- LLM(두뇌): Haiku 4.5 / Sonnet 4.5. Applied on the next reply (no reload).

Backend: GET /api/models, POST /api/models/stt, POST /api/models/llm; Dashboard
gains models_settings/set_stt_model/set_llm_model. In-memory only (a service
restart reverts to the env defaults small / claude-haiku-4-5).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:32:55 +09:00
EJClaw
b0272d2171 feat(dashboard): merge lists button, drop STT test, add bot/log settings, barge-in
Dashboard changes requested by the user:
- Merge 화이트리스트/블랙리스트 into one "👤 유저 등록/관리" button (the modal
  already holds both lists).
- Remove the 음성 인식(STT) 테스트 section and its client JS.
- Add collapsible "⚙️ 봇 관련 설정" with a barge-in toggle: "유저 목소리 들을 때
  봇이 말하던 것 중지" (default on). Stored via /api/bot/settings; the bot reads
  it on its report round-trip and calls voicePlayer.stop() when an allowed user
  starts speaking (dave/bot.mjs).
- Add collapsible "🧹 로그 관련 설정" with a "잡음 로그 표시 안 함" toggle (default
  on, localStorage-backed): hides turns whose 들음 is (빈 결과) or 답변 is [잡음].
- Fix the fixed event/error log dock covering the bottom-most turn: syncDockPad()
  pads main by the dock height (on load, log render, toggle, resize).

BotControl gains a settings store; /api/bot/report now also returns settings.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:19:21 +09:00
EJClaw
b782ca70bd fix(dashboard): TTS settings POST logged res["settings"] (renamed to base/overrides)
The per-emotion refactor renamed the settings response keys but the POST
handler's log line still referenced res["settings"], throwing KeyError and
500ing every apply. Log base + overrides instead.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:11:26 +09:00
EJClaw
067efc7abe feat(tts): per-emotion voice controls (speed/word-gap/sentence-gap/pitch)
Each emotion can now be tuned independently. parse_segments resolves the four
controls per segment from a base (공통) dict plus an optional per-emotion
override; a missing override key inherits base. By default there are no
overrides, so every emotion delivers with the base values (모든 감정 = 기본값).

- emotion.py: Segment now carries all 4 controls + the canonical emotion name;
  parse_segments(text, base, overrides). Adds EMOTION_LABELS/EMOTIONS for the UI.
- melo.py: MeloTTS.emotion_overrides store; synth resolves per-segment controls
  and sends them per segment.
- melo_worker.py: _render applies each segment's own word_gap/sentence_gap/pitch
  (previously reply-global).
- dashboard.py: emotion dropdown in the TTS panel; GET returns base + overrides
  + emotion list; POST {emotion,...} stores an override (or {reset:true} clears
  it); base is set when emotion is omitted/"base".

Verified: an override on one emotion slows only that emotion (happy@0.7=5.66s vs
base 3.02s; sad unchanged at 3.06s); dashboard store/reset/base all work; 34
tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 23:10:33 +09:00
EJClaw
39b743d976 feat(tts): new default tone (1.35/-0.07/-0.30) + collapsible dashboard panel
- Change default voice controls to the user's tuned values: glyph speed 1.35,
  word_gap -0.07s (-70ms), sentence_gap -0.30s, pitch 0 (env defaults, slider
  initial values, and the 기본값 reset button all updated).
- Make the "봇 목소리(TTS) 조절" dashboard panel collapsible via its header;
  starts collapsed (▸), expands on click (▾).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 22:57:52 +09:00
EJClaw
2ce2806102 feat(dashboard): live TTS voice controls (speed/word-gap/sentence-gap/pitch)
Adds a "봇 목소리(TTS) 조절" panel to the voice-server dashboard so the four
controls can be tuned from the browser instead of only via env vars:

- 4 sliders (glyph speed, word gap, sentence gap, pitch) with live labels
- 미리듣기: synthesises a sample with the slider values WITHOUT changing the
  live bot voice (new per-call overrides on MeloTTS.synth)
- 봇에 적용: commits the slider values to the live TTS instance; next reply uses
  them. 기본값 button resets to the manual defaults.

Backend: GET/POST /api/tts/settings (clamped to the manual ranges) and POST
/api/tts/preview (returns audio/wav). Panel shows only when tts is a real
(non-mock) backend.

Verified end-to-end against a live dashboard instance: page renders the panel,
GET returns defaults, POST applies+clamps (pitch 99->12), preview returns a
valid wav and leaves the live settings unchanged; 29 tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-26 22:48:00 +09:00
EJClaw
d1d4bd1514 feat(dashboard): show Claude usage per day/week instead of per session
Persists per-day token usage (usage_store, ~/.config/wsai/usage.json, survives
restarts) and surfaces it in status as claude_usage.{today,week}. The navbar now
shows two cards — 오늘 토큰 / 주간 토큰 (input+output, with per-card tooltips
breaking down input/output tokens and requests) — replacing the session-only
token count.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 22:20:30 +09:00
EJClaw
7ea07c13ad feat(dashboard): leave voice when server/channel set to 없음
The voice-channel picker already switched channels / left on 없음. Now the
server picker does too: choosing 없음 (or switching to a server with no channel
selected) sends a leave so the bot exits its current voice channel. Factored the
select POST into sendSelect(guildId, channelId).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 22:15:42 +09:00
EJClaw
f1f059a018 feat(dashboard): show Claude usage (tokens + requests) since server start
ClaudeBrain now returns per-reply token usage (Reply.usage from the API
response), the dashboard accumulates it (monitor.add_claude_usage), and the
header shows a "클로드 토큰" stat (input+output total, with a tooltip breaking
down input/output tokens and request count).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 16:40:25 +09:00
EJClaw
be4bc87edf feat(dashboard): conversation log filter by time/user/server/channel/content
Adds the requested multi-dimension log search. Turns now carry speaker, guild
and channel (the bot sends X-User/Guild/Channel-Name on the voice-turn POST), and
a filter bar above the conversation feed narrows by 시간(최근 N분)·유저·서버·
채널·내용. The event/error panel keeps its text+level search.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 12:15:11 +09:00
EJClaw
441ab4831f feat(bot+dashboard): whitelist/blacklist listen filter (users + roles)
- Per-guild listen filter stored in the control plane (BotControl) with
  GET/POST /api/bot/lists; the filter also rides along in the bot report
  response so the bot always has the latest config.
- 화이트리스트/블랙리스트 popup: search the guild's members OR roles (type
  selector), add/remove to white/black, save. Whitelist = listen to only those
  (empty = everyone); blacklist = exclude. Reuses the shared 뒤로가기 modal.
- Bot reports guild roles + known members for the search UI, and filters
  incoming audio via a pure, unit-tested isAllowed() (dave/filter.mjs):
  blacklist always excludes; a non-empty whitelist restricts; else everyone.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 12:06:03 +09:00
EJClaw
d3cf4e01b5 feat(bot+dashboard): bot info, server/voice-channel picker, participants, speaker
Adds a dashboard<->bot control plane (bot pushes state + polls commands, keeping
the bot's single outbound-HTTP direction):
- New bot_control.BotControl + endpoints: GET /api/bot/state, /api/bot/commands;
  POST /api/bot/report, /api/bot/select.
- Dashboard header bar: bot identity/connection, server dropdown (top "없음"),
  voice-channel dropdown (top "없음"), and live participant list.
- Turns record who spoke (Turn.speaker, via X-User-Name on the voice-turn POST).
- dave/bot.mjs: reports identity/guilds/voice-channels/members, polls join/leave
  commands and joins dynamically, and sends the speaker's display name.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 11:43:48 +09:00
EJClaw
6a2865d899 feat(dashboard): VSCode-style log dock, 3-row turns, log search/edit/delete
- Bottom-docked collapsible terminal log panel (open/close), retains the event
  log from voice-server start (event tail 200→2000).
- Log search box + level filter (전체/오류/경고/정보); per-line 삭제/수정 and
  전체 삭제, backed by new monitor event ids and /api/logs/{clear,delete,edit}.
- Turns now show 들음 / 생각 / 답변 three rows; 생각 surfaces the emotion-tone
  plan the bot chose (and the [잡음] decision), via a new Turn.thought field.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 11:35:11 +09:00
EJClaw
df07224feb feat(dashboard): live-editable bot prompt via popup (view/edit/save)
Adds a persisted, runtime-editable system prompt. The brain reads the persona
on every turn (prompt_store.get_persona), so a dashboard edit applies to the
next reply with no restart; blank clears the override back to the built-in
PERSONA. New endpoints GET/POST /api/prompt, and a reusable modal popup
(뒤로가기 + 수정/저장) that later white/blacklist features will share.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 11:20:57 +09:00
EJClaw
87b77f9997 fix(voice): stop dashboard from mangling the leading [감정] tag
The dashboard voice-turn path ran _speech_text() before MeloTTS.synth, which
rewrote a leading "[힘차게] 안녕!" into "힘차게, 안녕!" — reading the first
emotion aloud and destroying the tag before the TTS emotion parser could use it.
Make _speech_text() a pass-through so every emotion tag (including the first)
reaches synth intact and shapes pitch/speed instead of being spoken. Adds a
regression test covering the leading-tag case.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 10:18:54 +09:00
EJClaw
f585ed7b76 feat(voice): bracketed emotion tags, [잡음] for noise, spoken emotion, concise replies
- Brain persona now prefixes every reply with one bracketed emotion tag
  (e.g. [반가움], [궁금]) and keeps replies to one or two short sentences.
- Empty/unrecognised audio (silence/noise) is reported as reply "[잡음]" with
  no TTS playback instead of an empty reply.
- TTS speaks the bracketed emotion too: "[힘차게] 안녕!" is synthesised as
  "힘차게, 안녕!" via a leading-tag -> spoken-word transform.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 01:44:25 +09:00
EJClaw
ee6f6b7f55 fix(brain): ride out transient Claude 529 overloads
A single Claude 529 Overloaded dropped the voice turn straight to the apology
fallback. The anthropic SDK retries >=500/429 but only twice by default, which
a busy window can outlast. Raise max_retries (WSAI_BRAIN_MAX_RETRIES, default 4)
so transient overloads recover silently, and give overloads their own spoken
fallback ("서버가 붐벼서...") distinct from generic failures. Kept modest so a
sustained outage still fails fast instead of leaving the bot silent.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 23:09:36 +09:00
EJClaw
575ac2949a feat(voice): real Claude brain in the Discord loop (think, not echo)
The voice loop echoed the recognised text. Wire the real brain: the Discord
voice-turn now runs STT -> ClaudeBrain.respond (with rolling conversation
history) -> TTS, so the bot actually thinks and answers. --voice-server builds
the brain by default (WSAI_BRAIN=claude, WSAI_BRAIN_MODEL overridable) and
gracefully falls back to echo if anthropic/Claude auth is unavailable. A brain
error speaks a short apology instead of killing the loop.

Verified end-to-end: an utterance wav returns X-Heard plus a distinct Claude
X-Reply and a synthesised reply wav on device=cuda. 12 tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 23:03:55 +09:00
EJClaw
0f94245d5f feat(voice): real Discord listen+speak loop (STT->echo->TTS)
The bot (dave/bot.mjs) previously only joined the channel and counted audio
frames — it never fed STT or spoke back. Wire the real loop:

- Node bot: buffer each speaker's Opus->PCM utterance until AfterSilence,
  wrap as WAV, POST to the Python voice-turn endpoint, then play the returned
  reply wav into the channel via an AudioPlayer (ffmpeg->Opus). Skips its own
  audio, dedupes overlapping subscriptions, and ignores sub-0.35s noise.
- Python: new `python -m wsai --voice-server` serves /api/voice-turn — decode
  the uploaded utterance, GPU faster-whisper STT, produce a reply (echo of what
  was heard for now), GPU MeloTTS synth, return the reply wav (recognised/reply
  text ride along as X-Heard/X-Reply headers). Both engines pre-warmed; turns
  show in the dashboard feed. MeloTTS.synth() extracted for direct wav reuse.

Echo mode verifies listening+speaking+GPU recognition entirely in Discord; the
Claude brain is the next slice. Verified the endpoint round-trip: utterance wav
-> correct Korean X-Heard/X-Reply + a WAVE reply on device=cuda. 12 tests pass,
node --check clean.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 21:59:01 +09:00
EJClaw
b9c929a73f feat(dashboard): browser STT recognition test on the GPU
The status page was view-only, so there was no way to actually verify Korean
recognition end-to-end. Add a live test: record from the mic (localhost/https)
or upload an audio file (works over LAN http, where browsers block getUserMedia),
POST it to a new /api/stt endpoint that ffmpeg-normalises the blob to 16 kHz
mono and runs the real GPU faster-whisper, then shows the recognised text +
latency + device. Results also land in the live turn feed.

The dashboard now optionally holds a WhisperSTT and drives it from a private
asyncio loop thread. New `python -m wsai --stt-test` serves the page with STT
enabled and pre-warms the GPU worker so the first recognition is instant.
WhisperSTT.resolved_device is exposed for the UI.

Verified: wav and browser-style webm/opus uploads both return the correct
Korean text on device=cuda in ~240-280ms. 12 tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 21:49:03 +09:00
EJClaw
9bac6d170a fix(dashboard): stop infinite mock loop; add demo-mode warning banner
--dashboard defaulted to looping mock STT forever, so the status page
piled up thousands of fake "conversations" (all mock, 0ms, same reply)
that looked like real traffic. Now it plays 3 sample utterances then
idles; a loud 데모 모드 banner states the turns are mock samples, not
real STT/Brain/TTS. Continuous demo moved behind --dashboard-loop-demo.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-16 23:06:30 +09:00
EJClaw
6b0755e1ff feat: live status dashboard for the voice loop
Add a stdlib-only observability site so you can open a browser and watch,
step by step: whether it is listening, what it heard, what the brain thought
and answered, how long each stage took, and whether anything errored.

- wsai/monitor.py: thread-safe telemetry hub (per-turn timed steps, status
  header, error log) with a pub/sub for live push.
- wsai/dashboard.py: stdlib http.server serving a self-contained page plus an
  SSE (/events) live stream; /api/state snapshot fallback.
- Pipeline emits step-by-step turn telemetry (화면 맥락 → 두뇌 → 응답) and
  listening/running status; optional monitor, so existing paths are untouched.
- `python -m wsai --dashboard` starts the site (0.0.0.0:8787, WSAI_DASHBOARD_PORT)
  and loops the mock voice demo so there is always live activity to watch.
- Tests cover turn recording, per-step timing, error marking, and live push.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-16 20:24:13 +09:00