Voice replies were too long and slow (~1.3s LLM). Cap output at 150 tokens (WSAI_BRAIN_MAX_TOKENS), rewrite the persona to demand one short sentence with no boilerplate self-intro / "무엇을 도와드릴까요" padding, and mark the static system prompt with cache_control so repeat turns skip re-processing it. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
9.7 KiB
9.7 KiB