perf(brain): switch voice brain to Sonnet 5 with an always-on brevity rule

Sonnet 5 reaches first token ~0.4s sooner than Sonnet 4.5 but answers the
same voice prompt more verbosely (measured 46-54 vs ~28 output tokens),
which erased the win in total turn time. Add a cached, persona-independent
brevity system block that pulls output back to ~30 tokens, so the faster
first token becomes a faster, lower-variance whole reply.

Measured (16-round interleaved A/B, production-shaped call):
  sonnet-4-5      TTFT 1.26s  total 2.02s (tail 3.26s)  out 29
  sonnet-5+brev   TTFT 0.85s  total 1.70s (tail 2.28s)  out 30

- ClaudeBrain: default model claude-sonnet-5 + BREVITY block (cached with
  the persona prefix so a dashboard persona edit can't drop it).
- Defaults aligned: config.Settings.anthropic_model and the voice-server
  WSAI_BRAIN_MODEL default -> claude-sonnet-5.
- Dashboard: add claude-sonnet-5 to LLM_OPTIONS + JS label; fix stale
  restart hint.
- tests/latency_ab.py: reproducible model-latency A/B harness (reads
  CLAUDE_CREDENTIALS_PATH; makes live API calls, so not a pytest test).

Streaming TTS was intentionally not added: this bot answers in one sentence,
where sentence-level streaming has no overlap to exploit, and it would
require rearchitecting both the Python endpoint and the node playback.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
EJClaw
2026-08-28 20:57:01 +09:00
parent 0e4e7c6bb2
commit 04664ce61a
5 changed files with 150 additions and 8 deletions

View File

@@ -142,7 +142,7 @@ def _run_voice_server(host: str, port: int) -> None:
if os.environ.get("WSAI_BRAIN", "claude").lower() not in ("none", "echo"):
try:
from .backends.claude import ClaudeBrain
model = os.environ.get("WSAI_BRAIN_MODEL", "claude-sonnet-4-5")
model = os.environ.get("WSAI_BRAIN_MODEL", "claude-sonnet-5")
brain = ClaudeBrain(model=model)
brain_name = "claude"
except Exception as exc: # noqa: BLE001

View File

@@ -127,6 +127,15 @@ class ClaudeVision:
class ClaudeBrain:
# Injected as an always-on, cached system block on top of the (editable)
# persona. Sonnet 5 answers correctly but more verbosely than 4.5 for the
# same voice prompt (measured 46-54 vs 28 output tokens), which erased its
# ~0.4s time-to-first-token advantage in total turn time. This hard brevity
# rule pulls Sonnet 5 back to ~30 tokens, so the faster first token actually
# translates into a faster (and lower-variance) whole reply. Kept separate
# from PERSONA so a dashboard persona edit can never drop it.
BREVITY = "지금부터 답은 무조건 한 문장, 12단어 이내로만. 부연·재확인·군더더기 금지."
PERSONA = (
"너는 디스코드를 이용해 사용자와 실시간으로 대화하는 AI 인공지능이야.\n\n"
"1. 역할\n"
@@ -163,7 +172,7 @@ class ClaudeBrain:
"- 너는 디스코드에서 함께 대화하는 실시간 AI 인공지능이다."
)
def __init__(self, *, model: str = "claude-sonnet-4-5", api_key: str | None = None) -> None:
def __init__(self, *, model: str = "claude-sonnet-5", api_key: str | None = None) -> None:
self.model = model
self._auth = _Auth(api_key)
@@ -176,8 +185,9 @@ class ClaudeBrain:
msgs.append({"role": "user", "content": screen_note + user_text})
client = self._auth.client()
# Read the persona live each turn so a dashboard edit applies immediately
# (falls back to the built-in PERSONA when no override is saved).
system = self._auth.system(get_persona(self.PERSONA))
# (falls back to the built-in PERSONA when no override is saved). The
# brevity rule is appended as its own block so it survives persona edits.
system = self._auth.system(get_persona(self.PERSONA), self.BREVITY)
if system:
# Cache the (static) system prompt so repeat turns skip re-processing
# it — lower time-to-first-token. No-op below the model's cache

View File

@@ -24,7 +24,7 @@ class Settings:
text: str | None = None # None | (discord)
capture_interval: float = 1.5
anthropic_model: str = "claude-sonnet-4-5"
anthropic_model: str = "claude-sonnet-5"
@classmethod
def from_env(cls) -> "Settings":

View File

@@ -649,7 +649,7 @@ class Dashboard:
# -- live model switching (STT size / LLM model) --------------------- #
STT_OPTIONS = ["tiny", "base", "small", "medium", "large-v3"]
LLM_OPTIONS = ["claude-haiku-4-5", "claude-sonnet-4-5"]
LLM_OPTIONS = ["claude-haiku-4-5", "claude-sonnet-4-5", "claude-sonnet-5"]
def models_settings(self) -> dict:
stt = self.stt
@@ -1193,7 +1193,7 @@ PAGE = r"""<!DOCTYPE html>
<div id="mLoad" class="mload" style="display:none">
<span class="spin"></span><span id="mLoadText">로딩 중…</span>
</div>
<p class="cfghint" style="margin:8px 0 0">STT는 전환 시 모델을 다시 로드합니다(처음 medium/large-v3는 다운로드로 수 분 걸릴 수 있어요). LLM은 다음 답변부터 즉시 적용됩니다. 서비스 재시작 시 기본값(medium · Haiku)으로 돌아갑니다.</p>
<p class="cfghint" style="margin:8px 0 0">STT는 전환 시 모델을 다시 로드합니다(처음 medium/large-v3는 다운로드로 수 분 걸릴 수 있어요). LLM은 다음 답변부터 즉시 적용됩니다. 마지막으로 고른 모델은 재시작 후에도 유지됩니다(기본값 Sonnet 5).</p>
</div>
</section>
<div class="demobar" id="demobar" style="display:none"></div>
@@ -1519,7 +1519,8 @@ function wireCollapse(toggleId, bodyId, caretId){
const STT_LABEL = {tiny:'tiny (가장 빠름)', base:'base', small:'small (빠름)',
medium:'medium (기본·정확)', 'large-v3':'large-v3 (최고 정확도)'};
const LLM_LABEL = {'claude-haiku-4-5':'Haiku 4.5 (가장 빠름)',
'claude-sonnet-4-5':'Sonnet 4.5 (고품질·조금 느림)'};
'claude-sonnet-4-5':'Sonnet 4.5 (고품질·조금 느림)',
'claude-sonnet-5':'Sonnet 5 (기본 · 빠르고 고품질)'};
function fillSel(sel, options, current, labels){
sel.innerHTML='';
options.forEach(o=>{ const el=document.createElement('option');