M2 안정성: 부하·장애 주입 테스트 통과 + 재시작 후 시간 초과 몰림 수정
- 2,000 연결 / 500 방 / 초당 248 행동: p95 20ms, p99 37ms, CPU 0.2코어, 메모리 141MB, 오류 0 - 부하 중 SIGKILL→재시작: 2,000 연결 자동 복구, 수 유실 0 - 복구된 방 마감을 30초 + 0~30초 무작위로 분산(몰림으로 p99 336ms → 40ms) - 보고서: docs/reports/load-20261004.md Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -96,7 +96,9 @@ describe('Room timers', () => {
|
|||||||
const restored = Room.restore('room1', '482913', seq, JSON.parse(stateJson) as RoomSnapshot, deps);
|
const restored = Room.restore('room1', '482913', seq, JSON.parse(stateJson) as RoomSnapshot, deps);
|
||||||
restored.clearTimer();
|
restored.clearTimer();
|
||||||
expect((restored.game!.state as any).moves).toHaveLength(1);
|
expect((restored.game!.state as any).moves).toHaveLength(1);
|
||||||
expect(restored.effectiveDeadline()).toBe(clock.now() + 30_000);
|
const d = restored.effectiveDeadline()!;
|
||||||
|
expect(d).toBeGreaterThanOrEqual(clock.now() + 30_000);
|
||||||
|
expect(d).toBeLessThan(clock.now() + 60_000);
|
||||||
expect(restored.roomView().seats.every((s) => s && !s.connected)).toBe(true);
|
expect(restored.roomView().seats.every((s) => s && !s.connected)).toBe(true);
|
||||||
});
|
});
|
||||||
});
|
});
|
||||||
|
|||||||
@@ -207,7 +207,9 @@ export class Room {
|
|||||||
room.game.state = def.migrate(room.game.state, room.game.stateVersion);
|
room.game.state = def.migrate(room.game.state, room.game.stateVersion);
|
||||||
room.game.stateVersion = def.stateVersion;
|
room.game.stateVersion = def.stateVersion;
|
||||||
}
|
}
|
||||||
room.game.deadlineFloor = now + 30_000;
|
// 30s to reconnect, plus up to 30s jitter so thousands of restored rooms don't all time out
|
||||||
|
// in the same instant (thundering herd after a restart, docs/04 §3).
|
||||||
|
room.game.deadlineFloor = now + 30_000 + Math.floor(Math.random() * 30_000);
|
||||||
room.scheduleTimer();
|
room.scheduleTimer();
|
||||||
}
|
}
|
||||||
return room;
|
return room;
|
||||||
|
|||||||
@@ -16,7 +16,7 @@
|
|||||||
## 3. 서버 시작 시 복구
|
## 3. 서버 시작 시 복구
|
||||||
1. `rooms`에서 `status IN ('lobby','playing')`인 방을 읽는다.
|
1. `rooms`에서 `status IN ('lobby','playing')`인 방을 읽는다.
|
||||||
2. 각 방의 `room_state`를 불러 메모리 Room을 만든다. 모든 플레이어는 "연결 끊김" 상태로 시작.
|
2. 각 방의 `room_state`를 불러 메모리 Room을 만든다. 모든 플레이어는 "연결 끊김" 상태로 시작.
|
||||||
3. 저장된 `deadline`이 이미 지났다면 복구 직후 일괄 처리하지 않고, **복구 시각 + 30초**로 마감을 미뤄 준다(재접속할 시간).
|
3. 저장된 `deadline`이 이미 지났다면 복구 직후 일괄 처리하지 않고, **복구 시각 + 30초 + 0~30초 무작위**로 마감을 미뤄 준다(재접속할 시간, 많은 방이 한순간에 시간 초과되는 몰림 방지).
|
||||||
4. 복구 실패(스키마 불일치, 게임 코드 변경으로 상태가 안 맞음)한 방은 `status='broken'`으로 표시하고, 들어온 사람에게 "서버 업데이트로 이 게임을 이어갈 수 없어요. 새로 시작해 주세요"를 보여 준다.
|
4. 복구 실패(스키마 불일치, 게임 코드 변경으로 상태가 안 맞음)한 방은 `status='broken'`으로 표시하고, 들어온 사람에게 "서버 업데이트로 이 게임을 이어갈 수 없어요. 새로 시작해 주세요"를 보여 준다.
|
||||||
5. 게임 규칙 코드가 바뀌어 저장 상태 형식이 달라지면, 그 게임 모듈의 `stateVersion`을 올리고 `migrate(oldState)`를 제공하거나 위 4번으로 처리한다.
|
5. 게임 규칙 코드가 바뀌어 저장 상태 형식이 달라지면, 그 게임 모듈의 `stateVersion`을 올리고 `migrate(oldState)`를 제공하거나 위 4번으로 처리한다.
|
||||||
|
|
||||||
|
|||||||
36
docs/reports/load-20261004.md
Normal file
36
docs/reports/load-20261004.md
Normal file
@@ -0,0 +1,36 @@
|
|||||||
|
# 부하·장애 주입 테스트 보고서 (2026-10-04)
|
||||||
|
|
||||||
|
- 대상 커밋: f72be63 + 복구 마감 분산 수정(이 보고서와 같은 커밋)
|
||||||
|
- 환경: 봇 호스트(8코어, 30GB), 서버·부하 생성기 각각 별도 systemd 스코프(서버 MemoryMax 2G), 같은 머신(localhost)
|
||||||
|
- 도구: `scripts/loadtest.ts` — 방마다 오목 선수 2 + 관전자 2 연결, 방마다 2초에 한 수, 지연 = 둔 사람 전송 → 상대 수신
|
||||||
|
|
||||||
|
## 1. 기준 부하 (docs/10-testing.md §4)
|
||||||
|
| 항목 | 기준 | 결과 |
|
||||||
|
|---|---|---|
|
||||||
|
| 동시 연결 | 2,000 | 2,000 (끝까지 2,000 유지) |
|
||||||
|
| 진행 중 방 | 500 | 500 |
|
||||||
|
| 초당 행동 | 250 | 248.2 |
|
||||||
|
| 지연 p95 | < 50ms | 20.2ms |
|
||||||
|
| 지연 p99 | < 100ms | 37.3ms |
|
||||||
|
| 서버 CPU | < 70% (4코어 기준 2.8코어) | 0.2코어 |
|
||||||
|
| 서버 메모리 | < 1GB | 141MB |
|
||||||
|
| 오류·끊김 | 0 | 거절 0, 끊김 0 |
|
||||||
|
|
||||||
|
## 2. 장애 주입: 부하 중 서버 SIGKILL → 재시작
|
||||||
|
| 항목 | 결과 |
|
||||||
|
|---|---|
|
||||||
|
| 재연결 | 2,000/2,000 자동 재연결 |
|
||||||
|
| 저장된 수 유실 | 0 (모든 방의 수 개수가 줄지 않음) |
|
||||||
|
| 복구된 방 | 500 (1회차), 1,000 (2회차: 버려진 방 500개 포함), 복구 실패 0 |
|
||||||
|
| 지연(전체 60초, 재시작 구간 포함) | 1회차 p95 18.8 / p99 37.9ms, 2회차 p95 21.2 / p99 39.6ms |
|
||||||
|
|
||||||
|
## 3. 발견하고 고친 문제
|
||||||
|
- 첫 장애 주입에서 p99가 336ms로 튀었다. 원인: 재시작 후 복구된 방 1,000여 개의 마감이 모두 "복구 시각 + 30초"로 같아, 30초 뒤 한꺼번에 시간 초과 처리됨(몰림).
|
||||||
|
- 수정: 복구한 방의 마감을 30초 + 0~30초 무작위로 분산(`Room.restore`, docs/04 §3). 재측정 p99 39.6ms.
|
||||||
|
|
||||||
|
## 4. 재현 방법
|
||||||
|
```bash
|
||||||
|
PORT=4600 DB_PATH=$HOME/.tmp/bg-load/app.db PUBLIC_ORIGIN=http://localhost:4600 LOADTEST_NO_HTTP_LIMITS=1 bun apps/server/src/index.ts &
|
||||||
|
bun scripts/loadtest.ts --rooms 500 --spectators 2 --seconds 60 --pid <서버 PID>
|
||||||
|
# 장애 주입: --chaos --restart-cmd "<서버를 다시 띄우는 명령>"
|
||||||
|
```
|
||||||
@@ -160,7 +160,13 @@ const cpu = () => {
|
|||||||
const st = require('node:fs').readFileSync(`/proc/${pid}/stat`, 'utf8').split(') ')[1].split(' ');
|
const st = require('node:fs').readFileSync(`/proc/${pid}/stat`, 'utf8').split(') ')[1].split(' ');
|
||||||
return (Number(st[11]) + Number(st[12])) / 100; // utime+stime seconds (USER_HZ=100)
|
return (Number(st[11]) + Number(st[12])) / 100; // utime+stime seconds (USER_HZ=100)
|
||||||
};
|
};
|
||||||
const rss = () => (pid ? Number(/VmRSS:\s+(\d+)/.exec(require('node:fs').readFileSync(`/proc/${pid}/status`, 'utf8'))![1]) / 1024 : NaN);
|
const rss = () => {
|
||||||
|
try {
|
||||||
|
return pid ? Number(/VmRSS:\s+(\d+)/.exec(require('node:fs').readFileSync(`/proc/${pid}/status`, 'utf8'))![1]) / 1024 : NaN;
|
||||||
|
} catch {
|
||||||
|
return NaN; // server was restarted (chaos mode)
|
||||||
|
}
|
||||||
|
};
|
||||||
const cpu0 = cpu();
|
const cpu0 = cpu();
|
||||||
const start = performance.now();
|
const start = performance.now();
|
||||||
let chaosDone = false;
|
let chaosDone = false;
|
||||||
|
|||||||
Reference in New Issue
Block a user