perf(brain): switch voice brain to Sonnet 5 with an always-on brevity rule
Sonnet 5 reaches first token ~0.4s sooner than Sonnet 4.5 but answers the same voice prompt more verbosely (measured 46-54 vs ~28 output tokens), which erased the win in total turn time. Add a cached, persona-independent brevity system block that pulls output back to ~30 tokens, so the faster first token becomes a faster, lower-variance whole reply. Measured (16-round interleaved A/B, production-shaped call): sonnet-4-5 TTFT 1.26s total 2.02s (tail 3.26s) out 29 sonnet-5+brev TTFT 0.85s total 1.70s (tail 2.28s) out 30 - ClaudeBrain: default model claude-sonnet-5 + BREVITY block (cached with the persona prefix so a dashboard persona edit can't drop it). - Defaults aligned: config.Settings.anthropic_model and the voice-server WSAI_BRAIN_MODEL default -> claude-sonnet-5. - Dashboard: add claude-sonnet-5 to LLM_OPTIONS + JS label; fix stale restart hint. - tests/latency_ab.py: reproducible model-latency A/B harness (reads CLAUDE_CREDENTIALS_PATH; makes live API calls, so not a pytest test). Streaming TTS was intentionally not added: this bot answers in one sentence, where sentence-level streaming has no overlap to exploit, and it would require rearchitecting both the Python endpoint and the node playback. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
@@ -127,6 +127,15 @@ class ClaudeVision:
|
||||
|
||||
|
||||
class ClaudeBrain:
|
||||
# Injected as an always-on, cached system block on top of the (editable)
|
||||
# persona. Sonnet 5 answers correctly but more verbosely than 4.5 for the
|
||||
# same voice prompt (measured 46-54 vs 28 output tokens), which erased its
|
||||
# ~0.4s time-to-first-token advantage in total turn time. This hard brevity
|
||||
# rule pulls Sonnet 5 back to ~30 tokens, so the faster first token actually
|
||||
# translates into a faster (and lower-variance) whole reply. Kept separate
|
||||
# from PERSONA so a dashboard persona edit can never drop it.
|
||||
BREVITY = "지금부터 답은 무조건 한 문장, 12단어 이내로만. 부연·재확인·군더더기 금지."
|
||||
|
||||
PERSONA = (
|
||||
"너는 디스코드를 이용해 사용자와 실시간으로 대화하는 AI 인공지능이야.\n\n"
|
||||
"1. 역할\n"
|
||||
@@ -163,7 +172,7 @@ class ClaudeBrain:
|
||||
"- 너는 디스코드에서 함께 대화하는 실시간 AI 인공지능이다."
|
||||
)
|
||||
|
||||
def __init__(self, *, model: str = "claude-sonnet-4-5", api_key: str | None = None) -> None:
|
||||
def __init__(self, *, model: str = "claude-sonnet-5", api_key: str | None = None) -> None:
|
||||
self.model = model
|
||||
self._auth = _Auth(api_key)
|
||||
|
||||
@@ -176,8 +185,9 @@ class ClaudeBrain:
|
||||
msgs.append({"role": "user", "content": screen_note + user_text})
|
||||
client = self._auth.client()
|
||||
# Read the persona live each turn so a dashboard edit applies immediately
|
||||
# (falls back to the built-in PERSONA when no override is saved).
|
||||
system = self._auth.system(get_persona(self.PERSONA))
|
||||
# (falls back to the built-in PERSONA when no override is saved). The
|
||||
# brevity rule is appended as its own block so it survives persona edits.
|
||||
system = self._auth.system(get_persona(self.PERSONA), self.BREVITY)
|
||||
if system:
|
||||
# Cache the (static) system prompt so repeat turns skip re-processing
|
||||
# it — lower time-to-first-token. No-op below the model's cache
|
||||
|
||||
Reference in New Issue
Block a user