26 KiB
Global Memory
This file is for mutable memory shared across Claude groups.
Use it for durable facts, preferences, and shared context that may change over time.
Do not store platform-wide operating rules here. Those now live in prompts/claude-platform.md.
Host hardware (GPU)
This host has an NVIDIA GPU available for compute:
- GPU: NVIDIA GeForce RTX 5050 (GB207, Blackwell), 8GB VRAM, compute capability 12.0 (sm_120)
- Driver 595.71.05, CUDA runtime 13.2;
libcuda.sopresent so GPU compute is usable - Ready-to-use PyTorch (CUDA 12.8 / Blackwell-capable) venv at
/home/claude/gpu-venv— run with/home/claude/gpu-venv/bin/python.torch.cuda.is_available()is True. - For CUDA C compilation,
nvccis not installed; install the CUDA Toolkit only if a task needs it.
Behavioral rule: when a task could meaningfully benefit from the GPU (ML inference/training, heavy parallel compute, large media processing), first ask the user whether they want to use the GPU before doing so. Do not silently run GPU workloads. Once confirmed, use the /home/claude/gpu-venv Python (or install what the task needs) and verify with a real GPU run.
- 2026-06-18 exception (javis_bot project): the user pre-approved GPU use for the
javis_botproject — do NOT ask before using the GPU on javis_bot tasks. Other projects still follow the ask-first rule above.
javis_bot deployment target
- 2026-06-18: the real production deployment runs on the user's Windows 11 gaming PC (CPU Ryzen 9 9950X3D, GPU RTX 5070Ti — Blackwell, same sm_120 arch as this host's RTX 5050) with Docker. Use
docker-compose.yml+docker-compose.gpu-windows.yml(Docker Desktop + WSL2 GPU). This Linux host is only the LAN browser host (JARVIS_ROLE=browser); test changes here as a Blackwell-GPU proxy but remember the bot/brain role actually runs on the Windows box. The1516396486265405548value the user pasted on 2026-06-18 was just a Discord message ID in the chat channel, not a guild/voice-channel ID.
Host RAM / disk safety (tmpfs rule)
/tmp (7.2G) and /dev/shm (13G) are RAM-backed tmpfs. Anything large written there consumes RAM, not disk.
- 2026-06-11 incident: MeloTTS Korean-TTS venv (~2.4GB) was installed into
/tmp/melo311(tmpfs), eating RAM and risking OOM. Fixed by moving it to disk at/home/claude/jarvis-tts/melo311and updating 43 script shebangs. - Recurrence guards (in place, defence-in-depth):
~/.bashrcexportsTMPDIR=$HOME/.tmp,UV_CACHE_DIR=$HOME/.cache/uv,PIP_CACHE_DIR=$HOME/.cache/pip(all ext4 disk), above the interactive-shell guard so non-interactive agent shells inherit it.- The
ejclaw.serviceuser unit (~/.config/systemd/user/ejclaw.service) sets the same threeEnvironment=vars, so every child process (agent runners, install/build steps that never source.bashrc) inherits disk temp at the systemd root. Applied viadaemon-reload; takes full effect on the nextejclaw.servicerestart. Do NOT restart it from inside a live agent turn (the agent is its descendant). - Scheduled watchdog (cron
*/30) alerts the group only if/tmpor/dev/shm> 3GB, available RAM < 4GB, or disk free < 10GB, and auto-deletes$HOME/.tmpfiles older than 3 days. This catches an explicit/tmp/...write that bypassesTMPDIR.
- Rule: NEVER install venvs, models, datasets, or any multi-GB artefact into
/tmpor/dev/shm. Use disk paths under$HOME(e.g./home/claude/jarvis-tts). If a tool ignoresTMPDIR, pass an explicit disk path. EJClaw runs directly viabununder the user systemd unit (no Docker/compose in use here). - As of 2026-06-11 the root filesystem was expanded:
/(ext4) is now 353G total with ~214G free (37% used); earlier it was 156G/84%. RAM is 30GB total, usually ~25GB free. Adding RAM is not the real fix for tmpfs blowups; keeping big artefacts on disk is. Disk is no longer tight, but the watchdog still warns if free drops below 10GB. - The two TTS scratch scripts (
/home/claude/jarvis-tts/gen_melo.py,gen_xtts.py) now route output to disk and hard-refuse any/tmpor/dev/shmwrite path via adisk_path()guard. - 2026-07-26 OOM incident +
MemoryHighchange: theejclaw.servicebase unit setMemoryHigh=3G(soft cgroup cap;MemoryMaxunset). An agent-launched GPU/diffusers job in theanyway2channel (gen_banner.py, ~2.9GB RSS on its own) ran INSIDE the bot's cgroup, pushed total past 3G, andsystemd-oomdoom-killedejclaw.service~6× in a row (systemd auto-restarted each time → restart loop; the banner also never finished). Fix: raised the soft cap toMemoryHigh=16Gvia a documented drop-in~/.config/systemd/user/ejclaw.service.d/10-memory-high.conf(applied withdaemon-reload, no restart needed for a cgroup-property change). Rationale: host has 30G, KAclaw is capped at 3G and other services ~2G, so 16G leaves ~11G+cache headroom while still capping runaways;MemoryMaxstays infinity so a burst is throttled, not hard-killed. The base unit still reads 3G — the drop-in overrides it (systemctl --user cat ejclawshows both). Cleaner long-term alternative (not done): keep the bot cap tight and launch heavy GPU jobs OUTSIDE the bot cgroup viasystemd-run --user --scope.
Outbound attachment allowed folders
Discord/outbound attachments are only delivered if the file lives inside an allowlisted directory; files elsewhere are rejected as outside-allowed-dirs and the bot now appends a visible "전송 실패" notice (it no longer drops silently).
Standing rule (2026-06-12 user decision): the single canonical outbound attachment folder is /home/claude/EJClaw/data/attachments (the EJClaw DATA_DIR attachments dir). It is already allow-listed by default — no .env registration is needed. To send ANY file, first write or copy it into that folder (a subfolder is fine, e.g. data/attachments/generated/), then reference that path in the MEDIA: directive. Do NOT widen the allowlist to other folders (the earlier idea of allow-listing /home/claude/jarvis-tts was reverted at the user's request, and auto-copying arbitrary out-of-allowlist files in the delivery layer was also reverted because it bypasses cross-room / sensitive-file protection).
- Concretely: when attaching generated media (TTS audio, images, etc.),
cpit into/home/claude/EJClaw/data/attachments/(or a subfolder) and emitMEDIA:/home/claude/EJClaw/data/attachments/.... This needs no restart, no env change, no code change. - The TTS scripts under
/home/claude/jarvis-ttsstill write their working output there; copy the resulting.wavintodata/attachmentsbefore attaching (jarvis-tts itself is NOT allow-listed). - System defaults still allowed for internal use: the codex
generated_imagesdir and the system temp dir (used for auto-generated screenshots). Onlydata/attachmentsis the place WE deliberately stage files to send. - Safety net: if a file is ever referenced from outside an allowed folder, delivery now appends a visible "전송 실패 (사유)" notice instead of silently dropping it.
- Auto-cleanup (2026-06-12 user decision): prefer
/home/claude/EJClaw/data/attachments/generated/for files we send. A user crontab entry sweeps that subfolder every 15 min and deletes files older than 60 min (find .../data/attachments/generated -type f -mmin +60 -delete). The sweep is scoped togenerated/only, so system-managed attachments underdata/attachments(e.g.paired-turn-outputs/, dashboard copies) are never touched. If a sent file must persist longer than ~1h, put it directly underdata/attachments/(notgenerated/).
Stored credentials
Shared credentials live at /home/claude/.config/ejclaw/secrets.json (chmod 600, owner-only). Read with the Read tool when a session in any channel asks to "저장해둔 계정토큰으로 로그인" or otherwise needs a registered token.
Schema: credentials.<host>.{type, host, token, note, added_at}.
Currently stored:
git.tkrmagid.kr— Gitea personal access token. Use viaAuthorization: token <value>header, or embed in HTTPS URL ashttps://<user>:<token>@git.tkrmagid.kr/.... Forgit clone/push, prefer the URL form orgit -c http.extraHeader="Authorization: token <value>" clone .... Do not paste the raw token into chat replies.sudo— local sudo password for theclaudeuser on this host. Use viaecho "$PW" | sudo -S <cmd>(read the password fromcredentials.sudo.passwordwith the Read tool, then pipe). Do not paste the raw password into chat replies.discord.com/bot/testbot— Discord test bot (테스트봇, client/app ID1538122882528321536), stored 2026-08-15 for use across all chats. Read the token fromcredentials["discord.com/bot/testbot"].tokenin secrets.json with the Read tool (token is NOT written here). Use asAuthorization: Bot <token>againsthttps://discord.com/api/v10, or as the bot token for a discord.js/gateway login. Do not paste the raw token into chat replies. Token was chat-exposed on save; if used for anything sensitive/long-term, rotate it in the Discord Developer Portal first. Re-ask the user if Discord returns 401., editsecrets.jsonand append a new entry undercredentials; update this list with the host and intended use.
User communication preferences
- 2026-05-27: 모든 채팅에서 사용자에게 답변할 때 기본적으로 존댓말을 사용한다. 특별히 반말을 요청받지 않는 한 반말로 답변하지 않는다.
Room mode policy
All rooms default to tribunal (paired) mode. Owner runs the work, reviewer/arbiter (claude-code) verifies. New rooms registered via bun setup/index.ts --step register are also tribunal by default.
If the user in a channel says any of these — "클로드 사용하지 말자", "paired 모드 끄자", "리뷰어 끄자", "이 방은 single 로 바꿔줘" — switch that channel back to single mode. Use:
bun -e "import { initDatabase, setExplicitRoomMode } from './src/db.js'; initDatabase(); setExplicitRoomMode('<chatJid>', 'single');"
The reverse phrase ("paired 켜자", "리뷰어 다시 켜자") flips it back to 'tribunal'. Acknowledge the change in chat and confirm the new mode.
Room deregistration (채팅 등록 해제 + 데이터 정리)
When the user says any of — "채팅 아이디 등록해제 해줘 ", "이 채널 등록 해제해줘", "채널 등록 해제" (with a channel id / JID) — the standing behavior (user decision 2026-08-18) is: fully unregister the room AND purge all managed data, keeping ONLY the chat channel record and its message history, then hot-reload (SIGHUP, NOT a full restart — see below) so the channel goes unregistered in the live process.
Use the canonical script (keeps chats + messages, removes everything else — room_settings/role/skill overrides, paired tasks/turns/attempts/outputs/reservations/leases/projects/handoffs, work_items, scheduled_tasks, task_run_logs, sessions, router cursor, and the on-disk groups/<folder>, data/workspaces/<folder>, data/sessions/<folder>, data/ipc/<folder> including any git worktree):
cd /home/claude/EJClaw
bun scripts/deregister-room.ts --dry-run <channelId...> # preview what will be removed
bun scripts/deregister-room.ts <channelId...> # execute the purge
Accepts bare Discord channel IDs (auto-prefixed to dc:) or full JIDs; multiple ids allowed. Safety built in: refuses a is_main room without --force; a group folder still used by another room is NOT deleted on disk (only that chat's DB rows); --dry-run writes nothing.
Applying to the live process is now automatic: after the purge the script calls signalEjclawReload(), which sends SIGHUP to the running service's main PID → runtimeState.reloadRoomBindings(), so the room drops from live bindings with NO full restart. Do NOT run a plain systemctl restart from inside a turn (it kills the agent mid-command and can't be confirmed). Verify by the Room bindings reloaded log or getRegisteredGroup(jid) → undefined. Always take a quick DB backup first for irreversible purges: cp store/messages.db /home/claude/ejclaw-db-backup-$(date +%Y%m%d-%H%M%S).db.
Channel registration + making it live (no false restarts) — MANDATORY
The agent runs INSIDE ejclaw.service, so a plain systemctl --user restart ejclaw.service from within a turn kills the agent mid-command: the restart is unreliable and you CANNOT confirm it in the same turn. Never claim a restart/reload happened without evidence — the repeated "재시작 됐다고 했는데 안 됨" complaints came from claiming success without verifying.
When you register (or deregister) a channel via a DB write (bun script / setup --step register) from a non-main room, the running process does NOT see it until its in-memory room bindings are reloaded. Preferred way — hot-reload, NO full restart, agent survives so you can verify in the same turn:
kill -HUP "$(systemctl --user show -p MainPID --value ejclaw.service)"
src/index.ts handles SIGHUP → runtimeState.reloadRoomBindings() (re-reads ONLY room bindings from the DB; message cursors/sessions untouched, so no re-processing). Sending HUP to the MainPID only (not the whole cgroup) leaves agent subprocesses alive. After it, VERIFY before reporting: look for the Room bindings reloaded log with the new groupCount, or check getRegisteredGroup(jid) resolves. Standing user rule: registering a channel must be followed immediately by this live-reload (this is the "등록하면 자동 재시작/반영" the user asked for) — do it once, verify, then report.
Only when a change truly needs a full process restart (e.g. code/env changes), use the detached form so it survives the agent's death, and verify in a SEPARATE scheduled task (new PID / start time) — do it exactly once:
systemd-run --user --collect --unit="ejclaw-restart-$$" systemctl --user restart ejclaw.service
washing_machine_app (EasyAppliance Android) deployment
- 2026-08-06 user standing rule: for the washing_machine_app project, ALWAYS deploy after a change and deliver via in-app update — do not just commit. Every meaningful change ends with a published Gitea release so the in-app updater offers it.
- Build env: default JDK on this host is Java 25, which the Gradle Kotlin compiler can't parse ("IllegalArgumentException: 25.0.3"). Build with
export JAVA_HOME=/usr/lib/jvm/java-21-openjdk-amd64. Android SDK at/home/claude/android-sdk(build-tools 35.0.0 has aapt/apksigner). - Release process: bump
versionCode/versionNamedefaults inapp/build.gradle.kts(also passable via-PverCode -PverName),./gradlew assembleRelease(auto-signed viakeystore.properties→ keystore at/home/claude/.config/ejclaw/washing_machine_app.keystore, release cert SHA-25656246bd2…; every release MUST use this same key or in-place update breaks). Then create a Gitea release + upload the APK asset namedEasyAppliance-v<ver>.apk. - Release host: Gitea
git.tkrmagid.kr, repotkrmagid/washing_machine_app. Use the storedgit.tkrmagid.krtoken (secrets.json). API:POST /api/v1/repos/tkrmagid/washing_machine_app/releasesthenPOST /releases/<id>/assets?name=...with the APK. The app'sUpdateRepositoryreads releases via the public API and offers any release whose tag (minusv) is newer than installed. Latest published: v0.3.9 (verCode 12) as of 2026-08-06. Token now stored encrypted (AndroidKeyStore AES/GCM via TokenCrypto); allowBackup=false, usesCleartextTraffic=false. - No git remote is configured in the owner worktree; the Gitea release (APK) is the delivery mechanism, separate from any source push.
- Real-device tuning: the user's SmartThings Personal Access Token is stored in secrets.json under
api.smartthings.com(Bearer,https://api.smartthings.com/v1). Use it to read actual device capabilities/supported values so preset/status mappings stay accurate. Their real devices: 세탁기b9b0b550-…(water temp enum none/cold/20/30/40/60/90; spincustom.washerSpinLevel; rinsecustom.washerRinseCycles; NO manual water level — auto), 벽걸이/스탠드 에어컨 (mode cool/dry/wind/aIComfort; fan auto/1/2/3/4/max; setpoint 16–30). SmartThings commands are type-strict: appliance enum args (washer temp "40", fan "1", rinse "3") are STRINGS; only setpoint-style commands are NUMBERS —SmartThingsDeviceRepository.encodeArghandles this via a NUMERIC_COMMANDS allowlist.
Git backups
- 2026-05-27 10:53 KST: EJClaw 정상 동작 상태를 git commit
1509108(backup current stable ejclaw state)로 백업했다. "이번 요청만 리뷰어 사용" 기능 작업은 이 백업 커밋 이후에 시작한 변경이다.
Per-service Claude credential isolation (user changes, verified 2026-06-23)
The user reworked credential handling so each service reads its OWN Claude credentials file instead of all sharing ~/.claude/.credentials.json. Goal: run each service on a different Claude account.
- Resolver:
src/claude-credentials-path.ts—getClaudeCredentialsPath(accountIndex, {allowHomeFallback})andhasExplicitClaudeCredentialsPath(). Order: account 0 →CLAUDE_CREDENTIALS_PATH; account 1+ →CLAUDE_ACCOUNTS_DIR/<n>/.credentials.json; else home fallback (~/.claudeor~/.claude-accounts/<n>). The same module exists in both EJClaw and KAclaw. - Wired into:
setup/login.ts(writes toCLAUDE_CREDENTIALS_PATH; PKCE state file is now per-path-hashedejclaw-claude-login-<hash>.jsonso different-account logins don't collide),src/token-refresh.ts,src/claude-usage.ts,src/agent-runner-environment.ts(pre-syncs creds into each session.claude/.credentials.json),src/runtime-inventory.ts. - Credential paths in use:
- claude CLI:
~/.claude/.credentials.json - EJClaw:
/home/claude/EJClaw/data/claude/.credentials.json(set viaCLAUDE_CREDENTIALS_PATHinEJClaw/.env) - KAclaw:
/home/claude/KAclaw/data/claude/.credentials.json(set viaCLAUDE_CREDENTIALS_PATHinKAclaw/.env)
- claude CLI:
- Relevant env flags (EJClaw and KAclaw both):
CLAUDE_USAGE_USE_HOME_CREDENTIALS=false,CLAUDE_TOKEN_REFRESH_USE_CREDENTIALS=true,CLAUDE_TOKEN_REFRESH_USE_HOME_CREDENTIALS=false. The token-refresh loop only runs whenCLAUDE_TOKEN_REFRESH_USE_CREDENTIALS=trueOR*_USE_HOME_CREDENTIALS=true. - To re-auth ONE service: run its login flow with that service's
CLAUDE_CREDENTIALS_PATHexported, then send the newcode#state. Do NOT copy one service's.credentials.jsonto another path — that collapses them onto one token + one shared refresh token, and when both refresh independently the refresh token rotates and the other side breaks withinvalid_grant. - Account-identity reality (verified via
https://api.anthropic.com/api/oauth/profile, 2026-06-23 ~19:35 KST): despite the intent, all three paths still resolved to the SAME accounttkrmagid@gmail.com(Max). CLI and KAclaw even held byte-identical access+refresh tokens (from an earlier~/.claude→KAclaw copy); EJClaw held a different token of the same account. Before assuming the services are on distinct accounts, verify the actual account by calling the profile API for each path's accessToken (compareaccount.email), not just by checking that the file paths differ. - KAclaw is NOT a git repo (no
.git), so its source changes can't be diffed/version-controlled. 2026-06-23 dashboard usage-row fix (utilization of exactly 1 rendering as 100%,> 1→>= 1indashboard-usage-rows.ts) lives only in the KAclaw working tree. EJClaw's copy of that file may carry the same> 1pattern — worth checking if the EJClaw dashboard ever shows a window at exactly 1%. - UPDATE 2026-06-23 ~21:26 KST: KAclaw was re-logged onto a DIFFERENT account at the user's request. Now EJClaw =
tkrmagid@gmail.com(Max) and KAclaw =hjj100411@gmail.com(NOT Max — has_claude_max=false, so Pro/Free, lower limits). Done via KAclaw's ownsetup/index.ts --step loginwithCLAUDE_CREDENTIALS_PATH=/home/claude/KAclaw/data/claude/.credentials.jsonexported, then updated KAclaw.envCLAUDE_CODE_OAUTH_TOKEN(S) to the new access token and restartedkaclaw.service. claude CLI (~/.claude) is stilltkrmagid@gmail.com. So EJClaw and KAclaw are now genuinely different accounts; verify with the profile-API email check if in doubt.
Codex (reviewer) credential isolation (2026-07-22)
EJClaw's paired-room reviewer runs as Codex (gpt-5.5) with ChatGPT-OAuth auth. It shared ~/.codex/auth.json with other consumers (manual codex CLI, KAclaw same-home), so refreshes rotated the token and the stale copy died with "access token could not be refreshed because your refresh token was already used." Every reviewer turn then failed → task ended reviewer_codex_unavailable and escalated straight to the arbiter, so chats showed owner + arbiter but the reviewer never posted. Diagnose via paired_turn_attempts.last_error in store/messages.db.
Fix (mirrors the Claude per-service isolation above): EJClaw now uses a dedicated Codex account dir ~/.codex-accounts/0/ instead of shared ~/.codex. src/codex-token-rotation.ts initCodexTokenRotation() loads numbered dirs under ~/.codex-accounts/ as the rotation pool and ignores the ~/.codex fallback when the dir has entries; EJClaw leases + writes-back refreshes to ~/.codex-accounts/0/auth.json (syncCodexSessionAuthBack), isolated from the CLI and KAclaw.
- Activation needs an ejclaw restart —
initCodexTokenRotation()runs once at process start (src/index.ts), so a running ejclaw keeps the old cached account untilsystemctl --user restart ejclaw. Do NOT restart from inside a live agent turn (the agent is a descendant). - Re-auth runbook (the procedure CHANGED): refresh the reviewer's Codex login into the isolated dir, NOT default
~/.codex:CODEX_HOME=/home/claude/.codex-accounts/0 /home/claude/EJClaw/node_modules/.bin/codex login --device-auththen open the printedhttps://auth.openai.com/codex/deviceURL and enter the one-time code. A plaincodex loginwrites to~/.codexand will NOT update the account EJClaw uses. Account: ChatGPTaccount_id 7e40ecd8-461b-49bd-bbe4-cdc6c7dc68cf. - Verify:
cd /home/claude/EJClaw && bun -e 'import {initCodexTokenRotation,getActiveCodexAuthPath} from "./src/codex-token-rotation.ts"; initCodexTokenRotation(); console.log(getActiveCodexAuthPath())'→ should print/home/claude/.codex-accounts/0/auth.json.
.5 Docker host + remote docker context (2026-07-21)
A second LAN machine runs Docker workloads; any Claude channel on .9 can drive it.
- Identity:
192.168.10.5— a Proxmox LXC container hostnameddocker, Debian 13, Docker 29.3.0, root-only shell, no GPU. This host (.9,192.168.10.9, "claude") is a QEMU/KVM VM with the RTX 5050 passed through — so .9 has the GPU, .5 does not. - SSH: passwordless key auth is installed (
claude@.9pubkey inroot@.5:~/.ssh/authorized_keys) →ssh root@192.168.10.5needs no password. Fallback root password insecrets.jsonunderdocker5.ssh. - Docker context
docker5(ssh://root@192.168.10.5) exists on .9 and is the active default context, so a plaindocker .../docker compose ...on .9 runs on .5. Usedocker --context default ...(ordocker context use default) for .9's own local docker (where the GPU Ollama service for B would run); switch back withdocker context use docker5. - .5 runs ~19 production containers — do NOT disturb: projects
bot(mamil, memil, music1/2/3, random),site(make_video_site, minecraft_launcher_site, signature),virtual-stock-site-bot(api/web/worker),mc_domain_proxy(mc-filter api/frontend/nginx/proxy),act_runner,portainer,spotify-tokener. Compose files underroot@.5:/root/other/*/docker-compose.yml. - GPU-from-.5: since .5 has no GPU device, GPU work for .5 containers must go over the network to .9 (e.g. Ollama on .9
--gpus allbound to0.0.0.0:11434, .5 callshttp://192.168.10.9:11434). In-container CUDA (e.g. faster-whisperdevice=cuda) will NOT work on .5.
CLAUDE.md
Behavioral guidelines to reduce common LLM coding mistakes. Merge with project-specific instructions as needed.
Tradeoff: These guidelines bias toward caution over speed. For trivial tasks, use judgment.
1. Think Before Coding
Don't assume. Don't hide confusion. Surface tradeoffs.
Before implementing:
- State your assumptions explicitly. If uncertain, ask.
- If multiple interpretations exist, present them - don't pick silently.
- If a simpler approach exists, say so. Push back when warranted.
- If something is unclear, stop. Name what's confusing. Ask.
2. Simplicity First
Minimum code that solves the problem. Nothing speculative.
- No features beyond what was asked.
- No abstractions for single-use code.
- No "flexibility" or "configurability" that wasn't requested.
- No error handling for impossible scenarios.
- If you write 200 lines and it could be 50, rewrite it.
Ask yourself: "Would a senior engineer say this is overcomplicated?" If yes, simplify.
3. Surgical Changes
Touch only what you must. Clean up only your own mess.
When editing existing code:
- Don't "improve" adjacent code, comments, or formatting.
- Don't refactor things that aren't broken.
- Match existing style, even if you'd do it differently.
- If you notice unrelated dead code, mention it - don't delete it.
When your changes create orphans:
- Remove imports/variables/functions that YOUR changes made unused.
- Don't remove pre-existing dead code unless asked.
The test: Every changed line should trace directly to the user's request.
4. Goal-Driven Execution
Define success criteria. Loop until verified.
Transform tasks into verifiable goals:
- "Add validation" → "Write tests for invalid inputs, then make them pass"
- "Fix the bug" → "Write a test that reproduces it, then make it pass"
- "Refactor X" → "Ensure tests pass before and after"
For multi-step tasks, state a brief plan:
1. [Step] → verify: [check]
2. [Step] → verify: [check]
3. [Step] → verify: [check]
Strong success criteria let you loop independently. Weak criteria ("make it work") require constant clarification.
These guidelines are working if: fewer unnecessary changes in diffs, fewer rewrites due to overcomplication, and clarifying questions come before implementation rather than after mistakes.