perf: cap chat output tokens via ollama_num_predict to bound reply latency

Spoken (TTS) replies are 1-2 sentences, so an unbounded num_predict only
exposes the worst case where the chat model rambles or loops. Add an
ollama_num_predict config (default 512, 0 disables) wired into the reply
loop's chat call on both the native- and text-tool paths. The 512-token
headroom stays well above this app's short tool-call JSON, so capping never
truncates a tool call. This keeps the user's quality model instead of
downgrading it. Configurable in the container via OLLAMA_NUM_PREDICT.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
javis-bot
2026-06-23 15:33:45 +09:00
parent c189ce2e65
commit 5ee47827f3
7 changed files with 146 additions and 4 deletions

View File

@@ -287,6 +287,8 @@ Turn 4: LLM → {content: "Here's a comprehensive comparison of the iPhone 15 mo
- `llm_tools_timeout_sec` (enrichment extraction)
- `llm_embed_timeout_sec` (vector search)
- `llm_chat_timeout_sec` (messages loop turn)
- Output bound:
- `ollama_num_predict` (default `512`, `0`/negative disables) caps the chat model's generated tokens per turn via the Ollama `num_predict` option on the messages-loop call. Spoken (TTS) answers are 1-2 sentences, so this never clips a normal answer; it bounds the worst-case latency of a model that occasionally rambles or loops. The default headroom sits well above this app's short tool-call JSON, so it does not truncate tool calls. Applied uniformly to the reply loop's chat call (both native-tools and text-tools paths); the small classification passes (intent judge, digests) keep their own caps. Note: this is a worst-case guard, not the primary latency lever, which is model size and GPU residency.
- Memory enrichment:
- `memory_enrichment_max_results` limits recalled snippets.
- `memory_digest_enabled` (default `null` = auto-on for SMALL models ≤7B, off for LARGE) distils the combined diary + graph dump into a short relevance-filtered note via a cheap LLM pass before injecting into the system prompt. See **Memory Digest for Small Models** below.