Commit Graph

1536 Commits

Author SHA1 Message Date
Codex
f5d1ec302d Guard against phantom "I'll be notified" dead-end finals at delivery
An agent could end a turn with "I'll wait for the build to complete — I'll be
notified" without registering any watcher, so nothing ever resumed and the user
was left staring at a dead-end (observed in the tts_site room). The prompt rule
alone did not stop the model. At final-delivery time, when the final reads like
such a phantom-notification dead-end AND no CI watcher is active for the chat,
append an explicit notice telling the user no notification is coming and to send
a follow-up — turning a silent hang into an actionable prompt. Pure detector +
guard with unit tests; wired into the owner/reviewer final-delivery path.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 21:02:23 +09:00
Codex
03e6f8bcc6 Self-heal tribunal rooms with no work_dir so the owner never goes silent
A tribunal room whose work_dir was never provisioned (null) made
resolveOwnerTaskForHumanMessage return a null task, so the owner never ran and
the room went completely silent — the user's messages got no reply at all (seen
in the tts_site room, where days of requests were dropped). ensurePairedProject
now provisions the canonical workspace on demand when work_dir is missing,
guarded to tribunal rooms so a single-mode room never gets a spurious paired
workspace. Idempotent via ensurePairedWorkspaceProvisioned. Adds tests for both
the tribunal self-heal and the single-mode no-op.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 20:47:00 +09:00
Codex
6fad6dcfab Pin carry-forward flag off in default-behavior test; ignore test scratch
The "does not carry forward by default" test in paired-execution-context.test.ts
read the real config, so enabling PAIRED_CARRY_FORWARD_LATEST_OWNER_FINAL in
this deployment's .env flipped its result and it failed. Pin the flag to false
in that suite (mirroring the flag-on carry-forward.test.ts which pins true) so
the unit test is deterministic regardless of the ambient .env. Full suite is
green again with the flag enabled: 1595 pass / 3 skip / 0 fail.

Also gitignore the .ejclaw-*images-*/ and .ejclaw-attachment-*/ scratch dirs
that outbound-attachments tests create in the repo root; they only leak when a
test run is interrupted mid-flight and would otherwise clutter git status.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 18:48:40 +09:00
Codex
712664ca00 Fix owner context loss after finalize; green the test suite
Owner continuity:
- Seed a freshly created owner task with the previous task's latest owner
  final on a cold start after the previous task already closed, so a user
  reply to a finalized TASK_DONE no longer produces a "no context" answer.
  Activated for this deployment via PAIRED_CARRY_FORWARD_LATEST_OWNER_FINAL
  (.env); carried text is injected as clearly-marked background only.
- Skip intermediate STEP_DONE outputs when picking the carry-forward anchor.

Single-mode routing:
- enforceRoomModeOnLease strips a stale reviewer/arbiter lease from a room
  switched back to single, preventing single-mode messages from stalling in
  the paired path on a stuck execution lease.

Session auth / credentials:
- Pre-sync Claude credentials into each session dir before the agent spawns.
- Honor CLAUDE_CREDENTIALS_PATH in setup/login.ts (per-service isolation).
- Add a relogin-required gate so a permanently logged-out claude-code room
  asks the user to re-login instead of spawning a doomed agent.

Other:
- Arbiter verdicts written in the user's language (verdict keyword stays EN).
- status-dashboard chatName field; runtime-inventory credential path resolver.

Tests (make suite fully green: 1595 pass / 3 skip):
- service-routing: default owner is now the claude service and reviewer is
  codex-review; update the 7 failover/default expectations accordingly.
- migrate-room-registrations: owner inferred as claude-code (configured
  OWNER_AGENT_TYPE) for a dual legacy room; reviewer becomes codex.
- register: mock paired-workspace provisioning + reload signal (registration
  now provisions a workspace and hot-reloads); assert RELOADED status.
- paired-execution-context: force a claude-code reviewer to exercise the
  Claude read-only branch regardless of the deployment default.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 18:40:14 +09:00
Codex
80df025672 Chunk long final answers in Discord editMessage instead of failing
A final answer over 2000 chars edited straight into the tracked progress
message was rejected by Discord (content[BASE_TYPE_MAX_LENGTH]), leaving a
failed edit + stale progress stub before the send fallback. editMessage
now edits the tracked message with the first 2000-char chunk and sends the
remainder as follow-up messages, so long finals land cleanly in order.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 17:55:57 +09:00
Codex
89d68927c6 Detect bun/node runtime version by basename, not exact command file
detectDirectRuntimeVersion switched on command.file === 'bun'|'node', so
a resolved absolute path (e.g. /home/claude/.bun/bin/bun) fell through to
"host:<path>" and lost the version. Match on path.basename so the version
probe fires and runtimeVersion stays "host:bun@<ver>" regardless of how
the package-manager command was resolved.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 17:48:23 +09:00
Codex
7b27167af6 Auto-reset Claude session on context-window overflow (prompt too long)
A resumed Claude session that grew past the model's context limit failed
every turn with "Prompt is too long" — even auto-compaction could not
shrink it — so the same oversized session was reused indefinitely and the
channel looped forever on that error (no Claude reset pattern matched,
unlike Codex). Add "prompt is too long" / context-limit patterns to
SESSION_RESET_PATTERNS so the poisoned session is cleared and the next
turn starts fresh.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 14:23:41 +09:00
Codex
e81357e1df Stop phantom-notification hangs and drop noisy "0초" progress label
Agents ended turns with "I'll wait … I'll be notified" expecting a
background-completion callback that this environment never delivers, so
the turn hung and left the planning sentence stuck in the channel. Add a
claude-platform rule that no such notification exists — finish within the
turn (foreground/poll) or ask the user for a follow-up.

Also suppress the elapsed-time suffix on progress messages until at least
one 5s bucket has elapsed, so freshly-created progress no longer shows a
meaningless "0초". Update the affected message-runtime expectations.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-25 11:47:53 +09:00
Codex
08c99675c6 Add per-service Claude credential resolver and re-login helper scripts
claude-credentials-path.ts is imported by token-refresh/claude-usage/
agent-runner-environment/runtime-inventory; committed tree already
depends on it. claude_relogin.sh is referenced by token-refresh.ts as
the manual recovery step; codex_relogin.sh mirrors the Codex re-auth
runbook. Drop stray empty .codex file.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-24 20:50:18 +09:00
Codex
c016b9c2fa port: apply 7 upstream security/robustness patches
Ports from isolated upstream-port branch (base b3c5a4b), verified in
isolation via baseline-vs-port failure-set diff and re-verified live
(195 pass / 0 fail on affected tests):
- redact Discord bot tokens in outbound (router.ts SECRET_PATTERNS)
- block SSRF to private hosts in MoA base URL (moa.ts)
- refuse public dashboard bind without auth token (web-dashboard-server.ts)
- merge upstream .gitignore rules for python/build/secret noise
- real CPU utilization from /proc/stat instead of load avg (unified-dashboard.ts)
- width-safe placeholder for missing usage window on mobile (unified-dashboard.ts)
- bump direct deps to patch known vulnerabilities (discord.js/yaml/cron-parser)

Risky upstream commits (d5a94af phantom reset-time, patch 6 Codex usage)
intentionally skipped to avoid touching the credential-isolation tree.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-24 19:35:59 +09:00
Codex
b3c5a4b41b Show the error reason after the generic failure message
When an agent turn fails, users only saw "요청을 완료하지 못했습니다. 다시
시도해 주세요." with no clue why. Capture the failure reason (explicit
error field, or provider-error text like "API Error: 529 Overloaded"
classified by detectClaudeProviderFailureMessage) in the turn controller
and append it to the failure message ("...\n\n오류 내용: <reason>",
secret-redacted + length-capped). Update the paired-room loop filter to
startsWith so a failure-with-reason is still recognized and not
re-injected into prompt history. Unit-tested.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-24 17:10:18 +09:00
Codex
b9d4144b7e Default Discord client allowedMentions to users-only (block mass pings)
The raw sendMessage/editMessage paths did not route through
sanitizeForOutbound, so an agent-authored @everyone/@here in a progress
or edited message could still ping the whole channel. Set a client-level
allowedMentions default (parse: ['users']) so no normal send/edit can
mass-ping; explicit <@id> user mentions still work, and the disk-usage
alert broadcasts via a separate REST path with its own allowedMentions.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-24 11:01:10 +09:00
Codex
23c9e5101f Neutralize accidental @everyone/@here in outbound agent text
Agent-authored replies pass through sanitizeForOutbound before Discord
send, but it did not touch mass-mention tokens — so a reply that merely
*mentioned* "@everyone" while explaining a feature pinged the whole
channel. Add neutralizeMassMentions (zero-width space after @) to the
central sanitizer so normal reply/progress/edit text can never mass-ping;
intentional broadcasts (e.g. the disk-usage alert) use a dedicated path.
Leaves <@id> mentions and emails intact. Covered by unit tests.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-24 10:58:53 +09:00
Codex
8349a020ce Add humanize-korean skill (bundled from im-not-ai, MIT)
Self-contained single-call Korean AI-tell humanizer skill for EJClaw's
runner-skill sync. Derived from epoko77-ai/im-not-ai @0ac1e84 (MIT):
adapted the codex single-call path to return the rewrite inline in chat,
bundled only the 3 rulebooks it reads, and made reference paths runtime
-neutral (Claude + Codex). Includes upstream LICENSE + NOTICE for MIT
attribution. No external/paid API — runs on the session model only.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 22:19:45 +09:00
Codex
ea8e146ab6 docs(global): align deregistration wording to SIGHUP hot-reload (no full restart)
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 19:54:52 +09:00
Codex
f74536a8b2 feat(register): auto hot-reload service after channel register/deregister
Making a channel registration live no longer depends on an operator/agent
remembering to restart. setup/register.ts and scripts/deregister-room.ts now
call signalEjclawReload() after the DB write, which sends SIGHUP to the running
service's main PID (→ runtimeState.reloadRoomBindings). Best-effort and
injectable-for-tests; no-ops when the service isn't running.

Also document the canonical rule in the committed, agent-loaded platform prompt
(prompts/claude-platform.md): prefer these paths, never fake a restart, verify
via the "Room bindings reloaded" log; a plain in-turn systemctl restart is
forbidden (it kills the agent).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 19:50:24 +09:00
Codex
9020d2c8c5 feat(runtime): SIGHUP hot-reload of room bindings (no full restart)
Registering/deregistering a channel via a DB write from a non-main room did not
take effect until a full restart, and restarting from inside an agent turn is
unreliable (the agent is a child of ejclaw.service, so `systemctl restart` kills
it mid-command and it can't be confirmed) — the source of repeated "auto-restart
missing / claimed restart but it didn't happen" failures.

Add a SIGHUP handler that calls runtimeState.reloadRoomBindings(), which re-reads
ONLY room bindings from the DB (cursors/sessions untouched, no re-processing).
`kill -HUP <MainPID>` then makes a registration live instantly with no restart
and no agent death, so it is verifiable in the same turn.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-23 19:42:03 +09:00
Codex
e752524252 docs(readme): note dashboard refresh, token write-back, deregister tooling
Record the recently deployed operational changes: minute-boundary + event-driven
status dashboard refresh with edit-retry-before-repost, Claude session→canonical
token write-back (with corrected session credential path scan) to prevent
invalid_grant family revocation, and the deregister-room purge tool.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 11:03:20 +09:00
Codex
04cfe97a14 fix(auth): scan real session credential paths for fan-out/stale detection
listSessionCredentialPaths scanned <sessions>/<folder>/.claude/.credentials.json,
but real session creds live at
<sessions>/<folder>/services/<serviceId>/.claude/.credentials.json (and under
tasks/<taskId>/...). The mismatch meant writeCredentials' fan-out and
loadFreshestCredentials' stale-copy scan silently missed every real session
file — so old refresh-token copies (family-revocation landmines) were never
overwritten and the session→canonical write-back could not converge dormant
sessions.

Rewrite it via the pure, tested collectSessionCredentialPaths that walks the
actual services/<id>/.claude and services/<id>/tasks/<id>/.claude layout.
Verified on live data: now matches all 24 real session credential files (was 0).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 10:31:42 +09:00
Codex
d8abcf0621 fix(auth): write Claude session-refreshed token back to canonical creds
Root cause of recurring "Refresh token expired / invalid_grant" logouts: the
main refresh loop and each agent session's Claude CLI share one OAuth token
family but refresh independently. Anthropic rotates refresh tokens and revokes
the whole family if an already-rotated token is reused, so when a session
refreshed mid-turn the canonical copy went stale and its next refresh was
rejected — forcing a manual re-login.

Add syncClaudeSessionAuthBack (mirrors the existing Codex syncCodexSessionAuthBack):
after each Claude turn, if the session's CLAUDE_CONFIG_DIR credentials are
strictly newer than canonical (and same subscription), adopt them into the
canonical file. writeCredentials then fans the current token out to every
session dir, so no stale copy lingers to trigger family revocation. Decision
logic extracted to the pure, tested shouldAdoptSessionOAuth.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-22 10:24:57 +09:00
Codex
3f73197e6c fix(scripts): purge reviewer/arbiter session leftovers on deregister
Deregistration only removed the base group folder, leaving the tribunal
reviewer/arbiter runtime behind: DB session rows keyed as "<folder>:reviewer"
/":arbiter" and on-disk dirs data/sessions/<folder>-reviewer/-arbiter (plus
ipc/workspaces variants). Include those role-suffixed variants in both the
sessions DELETE and the disk cleanup so a deregistered room leaves nothing.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-21 14:10:38 +09:00
Codex
d57dd68fc0 feat(dashboard): align base refresh to wall-clock minute boundary
Instead of a fixed 60s interval from an arbitrary start offset, the base status
refresh now fires on each wall-clock minute boundary (:00), keeping the
minute-precision timestamp shown in the message accurate. Event-driven updates
(new message / agent activity) still refresh in between. Adds
msUntilNextMinuteBoundary with tests.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-21 13:01:02 +09:00
Codex
65ef6e8830 fix(dashboard): bind editMessage to channel to stop repost loop
The retry refactor captured `const editMessage = channel.editMessage` and
called it detached, losing `this`. Every status edit then threw
"this.client is undefined", so each cycle failed all retries and reposted a
fresh (notifying) status message — ~50 reposts in 30 minutes. Bind the method
to the channel so the edit runs in place. Adds a regression test showing a
detached method fails while a bound one succeeds.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-21 12:58:04 +09:00
Codex
8600213bcc feat(dashboard): 60s base refresh with event-driven immediate updates
Raise the status dashboard base refresh from 10s to 60s, and refresh
immediately (debounced 1.5s) on events between ticks:
- a real chat message arrives in a registered room (index.ts onMessage)
- an agent starts/finishes a run, i.e. activity moves between rooms
  (GroupQueue.setOnActivityChange fired on activeCount changes)

requestImmediateStatusUpdate() drives updateStatus out-of-band via a coalescing
trigger. The existing re-entrancy guard means the base periodic refresh does not
run while an edit-retry is in flight; a refresh requested during that window is
remembered and runs once afterward. Adds createCoalescingTrigger with tests.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-21 12:53:39 +09:00
Codex
9108174380 feat(dashboard): repost status immediately on first render after restart
The 15s x2 edit-retry-before-repost policy is meant for transient blips during
steady-state operation, not the moment right after a (re)start. On the first
render after startup, use 0 retries: still edit the stored message in place if
possible, but if that edit fails, repost a fresh status message immediately
instead of waiting through the retry cycle. Subsequent ticks use the 2-retry
policy.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-21 12:34:34 +09:00
Codex
cf6d1f89fc feat(dashboard): retry status edit before reposting
A transient Discord error (e.g. HTTP 503) on the periodic status-message edit
previously caused an immediate repost of a fresh status message. Now the edit
is retried up to 2 more times at 15s spacing, and a fresh message is only sent
if every attempt fails. Retry logic is extracted into the testable
editStatusMessageWithRetry helper (injectable sleep). Added a re-entrancy guard
so overlapping interval ticks are skipped while a slow retry runs, preventing
double-posts.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-21 12:23:14 +09:00
Codex
0064654d8f fix(scripts): recover group folder for already-unregistered rooms
When a room was unregistered first (room_settings row already gone), the
folder was unknown, so sessions rows and the on-disk group/workspace/session/
ipc folders (and any git worktree) were silently skipped — leaving orphans.
Back-trace the folder from the group_folder recorded on the chat's leftover
paired_tasks/work_items/scheduled_tasks/service_handoffs/paired_projects rows
so folder-scoped cleanup still runs. Targets now carry a folder list.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 23:44:26 +09:00
Codex
e21308db1f fix(scripts): prune scheduled-task run logs + auto-backup on deregister
task_run_logs.task_id references scheduled_tasks.id, not paired_tasks.id, so
the previous deletion (keyed by paired task ids) left orphaned run-log rows
behind — with foreign_keys OFF nothing cleaned them up. Delete task_run_logs
by the chat's scheduled_tasks ids before removing the scheduled_tasks rows.

Also auto-back up the DB to /home/claude/ejclaw-db-backup-<ts>.db before the
irreversible purge (skippable with --no-backup).

Verified end-to-end in a sandboxed DB copy: a seeded room's scheduled task +
3 run logs, paired data, work items, sessions, and router cursor are all
removed, disk folders deleted, while chats/messages and unrelated rooms stay
intact.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 23:38:35 +09:00
Codex
77426fc15c feat(scripts): add deregister-room purge script
Reusable script to fully unregister chat room(s) and purge all managed data
while keeping ONLY the chats channel row and messages history. Removes room
registration, paired tasks/turns/attempts/outputs/reservations/leases/
projects/handoffs, work items, scheduled tasks, sessions, router cursor, and
the on-disk group/workspace/session/ipc folders (git worktrees removed
cleanly). Supports --dry-run, refuses main rooms without --force, and skips
disk deletion for group folders still shared by another room.

Backs the "채팅 등록 해제" standing workflow (unregister + purge + restart).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 23:30:35 +09:00
Codex
8a71484934 style(rooms): apply prettier formatting to room-registration test
Fold in the pre-commit prettier reflow that the previous commit did not
restage. No behavior change.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 15:07:40 +09:00
Codex
3930c21625 fix(rooms): keep rooms routable when mode_source is corrupt
A room_settings row with a valid room_mode but an unrecognized mode_source
(e.g. the stray 'room' value that disabled the cgv-macro channel) was
silently dropped by getStoredRoomSettingsRowFromDatabase. That removed the
room from every binding lookup, so the router ignored the channel, it
vanished from the status list, and re-registration wedged on the
UNIQUE(chat_jid) constraint because assignRoom's "existing" probe uses the
same loader.

Coerce an invalid mode_source to 'explicit' (preserving the stored
room_mode) and log a warning so the corruption stays visible and self-heals
on the next assignRoom, instead of taking the channel silently offline.
room_mode is already protected by a column CHECK; mode_source was not.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-08-18 15:06:46 +09:00
Codex
44c54a759c fix(paired): auto-provision paired workspace so tribunal rooms get a reviewer on registration
Rooms registered as tribunal had no work_dir and no paired_projects row, so
ensurePairedProject() returned null, no paired task was ever created, and only
the owner ran — the reviewer/arbiter never fired (web-vstock and 4 other rooms).

Add ensurePairedWorkspaceProvisioned(): defaults canonical work_dir to
groups/<folder>, guarantees it is a standalone git repo with an initial commit
(detected via a LOCAL .git so we never walk up into the EJClaw checkout and
create a stray worktree), and upserts the paired_projects row. Wire it into
setup/register.ts so every newly registered room is immediately usable by the
full owner→reviewer→arbiter flow. Adds a regression test.
2026-07-30 01:45:28 +09:00
Codex
6036430c60 style(turns): apply prettier line-wrap from pre-commit hook 2026-07-26 18:07:33 +09:00
Codex
142929ba39 fix(turns): clear progress ticker on abnormal turn exit to prevent zombie progress messages
turnController.finish() is the only path that clears the 5s progress ticker,
but it lives inside the caller's try block in message-runtime-turns.ts. When
runAgent throws / is aborted / is killed (e.g. a user "중단"), control jumps to
finally and finish() is skipped, leaving the ticker editing the Discord message
forever — an orphaned "stuck at 0s" zombie progress message with no backing
task/attempt.

Add an idempotent MessageTurnController.dispose() that tears down the progress
ticker + idle timer, and always call it from the caller's finally so timers are
cleared on every exit path. Adds a regression test.
2026-07-26 18:07:17 +09:00
Codex
58811b2700 style(usage): apply prettier line-wrap from pre-commit hook 2026-07-24 23:15:58 +09:00
Codex
6ab38ca461 fix(usage): render codex usage when rate-limit response has null secondary window
Some Codex plans (e.g. Plus) return only a weekly window in `primary` with
`secondary: null`. applyCodexUsageToAccount dereferenced `secondary.usedPercent`
unconditionally, throwing a TypeError that refreshActiveCodexUsage swallowed at
debug level, so usage never applied and the dashboard row stayed blank (-1).

Guard the null secondary, and when a lone window is weekly (windowDurationMins
> 1440) route it into the 7d slot leaving 5h unknown. Accounts that report both
windows are unchanged. Adds a regression test for the secondary:null case.
2026-07-24 23:15:12 +09:00
Codex
a9095d95a8 style(paired): apply prettier line-wrap from pre-commit hook
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-22 19:30:53 +09:00
Codex
4d3ab20378 fix(paired): stop silent halt on reviewer PROCEED + reviewer-unavailable
Two causes of the paired room "keeps stopping" symptom:
- Reviewer approvals worded as "PROCEED" were parsed as 'continue'
  (a change request), causing an owner TASK_DONE <-> reviewer PROCEED
  ping-pong until the deadlock cap. parseReviewerVerdict() now treats a
  leading PROCEED as approval so the turn finalizes after one round.
- When the Codex reviewer was unavailable, the owner's answer was held
  for review and the user saw nothing. Now the held owner answer is
  emitted with a "review skipped" notice on reviewer_codex_unavailable.

Verified: tsc --noEmit clean; 27 related vitest tests pass.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-22 19:30:22 +09:00
Codex
5d60df8122 style(usage): prettier line-wrap for dashboard 429 row test
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-20 20:56:44 +09:00
Codex
2612e8a6ca style(usage): prettier line-wrap for claude-usage 429 backoff
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-20 20:56:16 +09:00
Codex
2be6c8db8d fix(usage): honor 429 Retry-After to stop self-sustaining rate-limit loop
The Claude usage poller retried /api/oauth/usage every 60s, but the endpoint
returns 429 with a longer Retry-After window (~93s). Retrying mid-cooldown
re-tripped the limit so the 429s never cleared and usage data never populated.

Record a cooldownUntil from the 429 Retry-After header (5min fallback when
absent) and skip the API until it closes; cleared on success. Dashboard now
shows a "429" indicator instead of a stale value when rate-limited.

Adds regression tests proving the cooldown outlasts the 60s throttle and
releases once the window passes.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-20 20:55:35 +09:00
Codex
f9b1e74838 fix(rooms): default new rooms to tribunal when room_mode is omitted
assign_room (MCP tool + host-side assignRoomInDatabase) and setup
register previously fell back to 'single' when no room_mode was given,
so newly added rooms silently lost the reviewer. Default to 'tribunal'
across the MCP zod schema, the IPC arg fallback, and the DB helper, and
update the ipc-auth expectation accordingly.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-19 00:00:20 +09:00
Codex
822ac34c0e fix(merge): integrate gitea/main fork — renumber migrations, dedup, fix paired_tasks insert
Resolves the gitea/main <-> deployed-line merge:
- DB migrations: renumber gitea's colliding 019/020 to 021/022
  (reviewer_failure_count -> v21, turn_progress_text_compat -> v22) so all
  four migrations have distinct versions; update ordered list + bootstrap test.
- paired_tasks INSERT: add the missing VALUES placeholder so both new columns
  (reviewer_failure_count + arbiter_intervention_count) bind (26 cols/values).
- index.ts: drop duplicate startUsagePrimer import from the auto-merge.
- discord output: keep the deployed pipeline (attachment rejection notice) and
  call sanitizeForOutbound at the channel boundary so prose escaping done in
  prepareDiscordOutbound is not double-applied; keep gitea's reviewer
  silent-failure cap + router markdown-escape helpers.
- usage-primer/codex-warmup: keep the dawn-hold removal over gitea's primer.

Full test suite: only pre-existing env/bun-path failures remain; no merge regressions.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-18 05:49:10 +09:00
Codex
e14ac3dfca Merge remote-tracking branch 'gitea/main'
# Conflicts:
#	prompts/owner-common-paired-room.md
#	src/channels/discord.ts
#	src/codex-warmup.ts
#	src/db/bootstrap.test.ts
#	src/db/migrations/index.ts
#	src/paired-execution-context-reviewer.ts
#	src/usage-primer.test.ts
#	src/usage-primer.ts
2026-06-18 05:41:07 +09:00
Codex
1d8358d1e1 style(paired): prettier line-wrap for routing-indicator test
Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-18 05:31:11 +09:00
Codex
be9f2379c0 fix(paired): remove dawn Codex primer-alignment hold; add owner routing indicator
Reviewer/arbiter/owner Codex turns are no longer deferred during the
[03:00-08:00 KST] dawn window — they dispatch immediately at all hours.
Deletes codex-primer-alignment.ts and every hold call site (gating owner
hold + alignment notice, reviewer/arbiter queue hold, warm-up skip, primer
anchor-lock). The 08:00 primer still fires; it just no longer holds other
consumers. Tradeoff: the 13:00 KST 5h-reset anchoring is no longer enforced.

Also adds a user-visible next-step indicator after owner turns that do not
end the task (review_ready -> reviewer requested, arbiter_requested ->
arbiter called) so a same-looking status line is no longer ambiguous; no
extra line on completion to avoid duplicating the owner's final message.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-18 05:24:28 +09:00
Codex
a621e85432 fix bun path for workspace installs 2026-06-17 15:12:12 +09:00
Codex
80ddb1aa0d fix(scheduler): forward only final agent output to chat
Scheduled tasks only suppressed phase 'progress', so intermediate
preambles (e.g. "I'll run the watchdog checks.") leaked to the chat
each run while the actual <internal>-wrapped result was correctly
stripped. The interactive path treats intermediate/tool-activity as
silent (toVisiblePhase); align the scheduler to forward only the final
message, with error outputs still falling through to rotation/error
handling.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-13 14:37:42 +09:00
Codex
ee4c5559b1 feat(dashboard): add GPU/VRAM usage to status channel server section
Show NVIDIA GPU utilization and VRAM used/total in the 서버 status block,
matching the existing CPU/Memory/Disk bar format. Gracefully omitted when
nvidia-smi is unavailable.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-12 19:14:36 +09:00
Codex
73fd71b39a revert: drop unsafe outbound attachment relocation, keep only the visible failure notice
The staging logic in 9c46cf6 copied any agent-declared file from outside the
room's allowed directories into a safe folder and attached it, which bypassed
the attachment directory allowlist (cross-room isolation / sensitive-file
protection). Removing it.

Kept: appendRejectionNotice / describeRejectedAttachments so rejected
attachments are surfaced in the visible message instead of being silently
dropped. This changes no security behavior — it only adds text when an
attachment was already going to be rejected.

Verified: outbound-attachments + final-delivery + discord tests 70/70, tsc clean.

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
2026-06-12 00:52:42 +09:00