Sonnet 5 reaches first token ~0.4s sooner than Sonnet 4.5 but answers the
same voice prompt more verbosely (measured 46-54 vs ~28 output tokens),
which erased the win in total turn time. Add a cached, persona-independent
brevity system block that pulls output back to ~30 tokens, so the faster
first token becomes a faster, lower-variance whole reply.
Measured (16-round interleaved A/B, production-shaped call):
sonnet-4-5 TTFT 1.26s total 2.02s (tail 3.26s) out 29
sonnet-5+brev TTFT 0.85s total 1.70s (tail 2.28s) out 30
- ClaudeBrain: default model claude-sonnet-5 + BREVITY block (cached with
the persona prefix so a dashboard persona edit can't drop it).
- Defaults aligned: config.Settings.anthropic_model and the voice-server
WSAI_BRAIN_MODEL default -> claude-sonnet-5.
- Dashboard: add claude-sonnet-5 to LLM_OPTIONS + JS label; fix stale
restart hint.
- tests/latency_ab.py: reproducible model-latency A/B harness (reads
CLAUDE_CREDENTIALS_PATH; makes live API calls, so not a pytest test).
Streaming TTS was intentionally not added: this bot answers in one sentence,
where sentence-level streaming has no overlap to exploit, and it would
require rearchitecting both the Python endpoint and the node playback.
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>