Sonnet 5 reaches first token ~0.4s sooner than Sonnet 4.5 but answers the
same voice prompt more verbosely (measured 46-54 vs ~28 output tokens),
which erased the win in total turn time. Add a cached, persona-independent
brevity system block that pulls output back to ~30 tokens, so the faster
first token becomes a faster, lower-variance whole reply.
Measured (16-round interleaved A/B, production-shaped call):
sonnet-4-5 TTFT 1.26s total 2.02s (tail 3.26s) out 29
sonnet-5+brev TTFT 0.85s total 1.70s (tail 2.28s) out 30
- ClaudeBrain: default model claude-sonnet-5 + BREVITY block (cached with
the persona prefix so a dashboard persona edit can't drop it).
- Defaults aligned: config.Settings.anthropic_model and the voice-server
WSAI_BRAIN_MODEL default -> claude-sonnet-5.
- Dashboard: add claude-sonnet-5 to LLM_OPTIONS + JS label; fix stale
restart hint.
- tests/latency_ab.py: reproducible model-latency A/B harness (reads
CLAUDE_CREDENTIALS_PATH; makes live API calls, so not a pytest test).
Streaming TTS was intentionally not added: this bot answers in one sentence,
where sentence-level streaming has no overlap to exploit, and it would
require rearchitecting both the Python endpoint and the node playback.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
- 유저 등록/관리 search now has 4 types: 유저(non-bot members), 봇(bot members),
통화방(current voice-channel members, bot flag enriched), 역할(roles — expandable
to the members under each role; register the role OR a member under it).
- Persist settings to a JSON state store (WSAI_STATE_FILE, default
~/.config/wsai/state.json) so they survive a service/container restart:
per-guild listen lists + bargeIn (BotControl), TTS base/overrides and the
chosen STT/LLM models (Dashboard). Applied on voice-server startup before warm.
- TTS 감정 선택에 최상단 "전체변경 (모든 감정)" 추가: 적용하면 base를 그 값으로 세팅하고
모든 per-emotion override를 지워 전 감정이 한 번에 그 값으로 말합니다.
- Move the component status bar (눈/시각/귀/LLM/입) above the bot bar.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Each emotion can now be tuned independently. parse_segments resolves the four
controls per segment from a base (공통) dict plus an optional per-emotion
override; a missing override key inherits base. By default there are no
overrides, so every emotion delivers with the base values (모든 감정 = 기본값).
- emotion.py: Segment now carries all 4 controls + the canonical emotion name;
parse_segments(text, base, overrides). Adds EMOTION_LABELS/EMOTIONS for the UI.
- melo.py: MeloTTS.emotion_overrides store; synth resolves per-segment controls
and sends them per segment.
- melo_worker.py: _render applies each segment's own word_gap/sentence_gap/pitch
(previously reply-global).
- dashboard.py: emotion dropdown in the TTS panel; GET returns base + overrides
+ emotion list; POST {emotion,...} stores an override (or {reset:true} clears
it); base is set when emotion is omitted/"base".
Verified: an override on one emotion slows only that emotion (happy@0.7=5.66s vs
base 3.02s; sad unchanged at 3.06s); dashboard store/reset/base all work; 34
tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
User reported the emotion samples sounded the same and the speed change wasn't
noticeable. Two changes:
- Add a "base" (기본/neutral) emotion at speed 1.0x so the plain base voice can
be selected explicitly via a [기본] tag and auditioned against the others.
- Bump the WSAI_TTS_SPEED default 1.15 -> 1.5 for a clearly faster base voice.
Emotion multipliers scale off base, so every emotion speeds up together.
Also extends gen_emotion_samples.py to emit one wav per emotion (incl. 기본)
plus a stitched all-in-one, so each emotion can be delivered as a separate clip.
Verified: 29 tests pass; match_emotion('기본') == 'base'; per-emotion synthesis
at base 1.5 produces distinct clip durations (base 4.9s vs happy 3.3s).
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
User asked for a bit faster speech. Pitch shift is already disabled, so the
earlier "monster voice" risk from stacking pitch on a fast base is gone — a
modest 1.15x base is natural. Emotion multipliers scale off base, so every
emotion gets the bump too. Also adds tests/gen_emotion_samples.py, a dev
utility that synthesizes one clip per canonical emotion (announce name at
neutral speed, then a sample sentence steered by that emotion's tag) and
stitches them into a single wav for auditioning the emotion palette.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The STT/TTS worker _ensure() treated a spawned-but-not-yet-handshaked
subprocess as ready, so a voice turn arriving during warmup read the same
stdout StreamReader concurrently with the warmup handshake and crashed with
"readuntil() called while another coroutine is already waiting for incoming
data". Add a _start_lock + _ready flag so (re)start and the ready handshake
run atomically and callers wait for real readiness before reading stdout.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
The dashboard voice-turn path ran _speech_text() before MeloTTS.synth, which
rewrote a leading "[힘차게] 안녕!" into "힘차게, 안녕!" — reading the first
emotion aloud and destroying the tag before the TTS emotion parser could use it.
Make _speech_text() a pass-through so every emotion tag (including the first)
reaches synth intact and shapes pitch/speed instead of being spoken. Adds a
regression test covering the leading-tag case.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Emotion tags now steer delivery rather than being read aloud. parse_segments()
splits a reply on [감정] tags: a recognised emotion word switches the pitch and
speed of the text that follows (and is dropped), while a non-emotion bracket
(e.g. [1번]) keeps its inner words as spoken content. Emotions can change
mid-reply, so a single turn is synthesised as several pitch-shifted segments and
concatenated in the melo worker (librosa pitch_shift, warmed at startup).
The emotion vocabulary is grounded in Azure Neural TTS speaking styles plus
Ekman's basic emotions, with Korean synonyms. The brain persona is updated to
emit inline tags from that set.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Claude replies with markdown/backticks by default; MeloTTS's Korean text
normaliser has no entry for '`' and dies with KeyError: '`', so any reply
mentioning a command/code block crashed the whole voice turn (500 on
/api/voice-turn). Fix at the shared synth() choke point with
normalize_for_speech(), which flattens code fences/inline code/links/markdown
and guarantees no backtick reaches the worker — covering both the dashboard
voice turn and the Discord speak() bridge. Also add a PERSONA line asking the
model to avoid markdown (belt-and-suspenders; the code strip is the real fix).
errors_total never moved for turn-level failures: it was only bumped by
log("error") events, and the dashboard voice path calls turn.finish(error=...)
without logging. Emit one error-level log event from Turn.finish() when a turn
ends in error, so both the server counter and the browser SSE mirror stay
consistent, guarded to count at most once. Drop the now-redundant pipeline
log("error") to avoid double counting and remove the dead _publish stub.
Verified: raw backtick -> worker KeyError '`' reproduced; after fix real
MeloTTS synth of a backtick+fenced reply succeeds; /api/voice-turn returns 200
with a wav body on a backtick reply and errors_total stays 0, and an induced
synth failure returns 500 with errors_total incrementing to exactly 1. Full
suite 18 passed.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Step 3 (귀): add WhisperSTT + whisper_worker, a warm out-of-venv worker
mirroring the MeloTTS shape (whisper312 venv, small/int8 on CPU). transcribe()
closes the voice round trip (MeloTTS wav -> whisper text); utterances() turns an
injected audio_source into Utterances (Discord voice feed pending). Wired into
factory as WSAI_STT=whisper.
Also address the arbiter's TTS follow-ups:
- melo worker error handling: capture stderr (drained in a bounded background
task so the pipe can't fill), surface the real failure cause, and defend
against an empty/invalid ready line instead of dying on JSONDecodeError.
- pipeline pre-warm: load slow backends (warmup()) at startup so the first
utterance is answered warm; a warmup failure is logged, not fatal.
Verified: real TTS->STT round trip recovers the sentence near-perfectly;
warm transcribe ~1.2s (CPU). 12 tests pass.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Add a stdlib-only observability site so you can open a browser and watch,
step by step: whether it is listening, what it heard, what the brain thought
and answered, how long each stage took, and whether anything errored.
- wsai/monitor.py: thread-safe telemetry hub (per-turn timed steps, status
header, error log) with a pub/sub for live push.
- wsai/dashboard.py: stdlib http.server serving a self-contained page plus an
SSE (/events) live stream; /api/state snapshot fallback.
- Pipeline emits step-by-step turn telemetry (화면 맥락 → 두뇌 → 응답) and
listening/running status; optional monitor, so existing paths are untouched.
- `python -m wsai --dashboard` starts the site (0.0.0.0:8787, WSAI_DASHBOARD_PORT)
and loops the mock voice demo so there is always live activity to watch.
- Tests cover turn recording, per-step timing, error marking, and live push.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Pipeline.run() used asyncio.gather, so if one loop raised, the failing
coroutine propagated while the sibling loops kept running detached; aclose()
in the finally then closed a source/stt out from under a still-live loop.
Switch to asyncio.TaskGroup so a failing loop cancels+awaits the siblings
before teardown. Add a regression test asserting an error in the conversation
loop cancels the perception loop and still closes every source.
Make source/vision optional so the conversation loop runs with no screen
capture. Add Settings.voice() preset and `python -m wsai --voice` demo.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Modular async pipeline: FrameSource->Vision->context and STT/text->Brain->TTS.
All stages are Protocols; mock backends run end-to-end with no deps/keys.
Real backends included: mss screen capture, Claude vision+brain (guarded imports).