AI Implementation feature(918): Admin Area System: Database Backup/Restore and Danger Zone 1.01 (#63)

This commit was merged in pull request #63.
This commit is contained in:
2026-07-23 13:28:29 +00:00
parent e611d201d9
commit e2d6bb3d69
8 changed files with 382 additions and 163 deletions
-155
View File
@@ -1,155 +0,0 @@
# Implementation Plan: Job 916 — Admin System Restore Commit (stagingId stripped from request body)
## Root cause (read-only investigation)
The destructive restore never executes because the Zod schema used to validate
`POST /api/v1/admin/system/restore/commit` **drops the `stagingId` field** from
the request body before the controller runs.
`backend/src/modules/admin/system/dto/admin-system.dto.ts:23`:
```ts
export const AdminRestoreCommitBodySchema = AdminDangerConfirmBodySchema;
```
`AdminDangerConfirmBodySchema` (line 18) only allows `{ operation, token }`.
Zod's `safeParse` / the NestJS `ZodValidationPipe` therefore strip every
unknown field — including `stagingId` — from the parsed body. The controller
then receives `body.stagingId === undefined`, falls back to the empty string
(`body.stagingId ?? ''`), and `RestoreService.commitRestore('')` finds no
matching stage and throws `SYSTEM_RESTORE_STAGE_EXPIRED`
(`RestoreService.commitRestore` at `restore.service.ts:223-231`).
The frontend already sends the field correctly
(`frontend/src/app/features/admin/system/system.service.ts:42`):
```ts
async commitRestore(payload: { operation: 'restore-backup'; token: string; stagingId: string })
```
so the deviation is purely server-side. The negative path (random JSON) still
fails with `SYSTEM_RESTORE_VALIDATION_FAILED` because the malformed archive
rejects at `validate`, never reaching `commit`.
The fix is to give `AdminRestoreCommitBodySchema` its own declaration that
**includes** `stagingId` (mirroring `AdminReauthBodySchema` at line 12-16), and
to tighten the controller so a missing `stagingId` returns a precise
`SYSTEM_RESTORE_STAGE_EXPIRED` instead of relying on the empty-string fallback.
## 1. Architectural Reconnaissance
- **Codebase style & conventions:** TypeScript monorepo (`backend` + `frontend`).
Backend is NestJS with `class-validator`/`zod` mixed; the admin-system module
uses `ZodValidationPipe` per-route. Strict async/await, error handling via
`ApiError` + `ERROR_CODES` (`backend/src/common/errors/`).
- **Data Layer:** SQLite via `better-sqlite3` + TypeORM. Restore streams are
in-memory (`RestoreService.stages: Map<stagingId, StagedRestore>`) plus a
sibling directory tree under `<DATA_DIR>/.system-staging/restore-<id>/`.
- **Test Framework & Structure:** Jest with two projects in
`tests/jest.config.js`. Tests live in `tests/backend/*.spec.ts` and
`tests/frontend/*.spec.ts`. Single command from the root:
`npm test` (already wired in `package.json`).
- **Required Tools & Dependencies:** No new tools. Existing: Node 20, NestJS,
better-sqlite3, zod, Jest, ts-jest. `setup.sh` already builds the app and
rebuilds the native binding — no updates needed.
## 2. Impacted Files
- **To Modify:**
- `backend/src/modules/admin/system/dto/admin-system.dto.ts` — replace the
`AdminRestoreCommitBodySchema` alias with a proper schema that accepts
`stagingId` (required for `restore-backup`).
- `backend/src/modules/admin/system/admin-system.controller.ts` — validate
`stagingId` is present before consuming the token, and stop feeding an
empty string into `commitRestore`.
- **To Create:**
- `tests/backend/admin-system-restore-commit.spec.ts` — focused e2e test
that the `restore/commit` endpoint forwards `stagingId` to
`RestoreService.commitRestore` and atomically swaps the live DB.
## 3. Proposed Changes
1. **DTO fix**`backend/src/modules/admin/system/dto/admin-system.dto.ts`
- Remove the `AdminRestoreCommitBodySchema = AdminDangerConfirmBodySchema`
alias (line 23).
- Add a dedicated schema that requires `stagingId` for `restore-backup`
while still letting the safe schema accept the other two operations in
shared code paths:
```ts
export const AdminRestoreCommitBodySchema = z.object({
operation: z.enum(ADMIN_OPERATION_KINDS),
token: z.string().min(1).max(512),
stagingId: z.string().min(1).max(128).optional(),
});
```
- Keep the `AdminOperationKind` re-export and the
`AdminRestoreCommitBody` `z.infer` type aligned so the controller
signature remains `body: { operation; token; stagingId? }`.
2. **Controller hardening** — `backend/src/modules/admin/system/admin-system.controller.ts:95-121`
- After the `body.operation === 'restore-backup'` check, throw
`SYSTEM_RESTORE_STAGE_EXPIRED` with HTTP 400 if `body.stagingId` is
falsy. This produces a precise error instead of the empty-string
fallback that currently slips through to `RestoreService`.
- Pass `body.stagingId` directly to `commitRestore` (no `?? ''`).
- Existing behavior for `reset-scores` / `wipe-challenges` (which use
`AdminDangerConfirmBodySchema`) is unchanged.
3. **No frontend changes required.** `system.service.ts` already POSTs
`{ operation, token, stagingId }`; the modal already calls it after
`confirmations`. Once the server schema preserves `stagingId` the happy
path works end-to-end.
4. **No DB migration needed.** `admin_operation_token` already exists;
`stagingId` is bound through the in-memory `RestoreService.stages` map.
## 4. Test Strategy
- **Target Unit Test File:** `tests/backend/admin-system-restore-commit.spec.ts` (new).
Companion lightweight spec additions to
`tests/backend/admin-system-restore-validation.spec.ts` are optional; the
new file is sufficient and keeps the diff small.
- **Mocking Strategy:**
- Build a minimal `Test.createTestingModule` with just
`AdminSystemController`, `RestoreService`, `ConfirmationTokenService`,
`AuthService`, `BackupService`, `DangerZoneService`, the
`AdminOperationTokenEntity` repository, the `RestoreArchiveSchema` DTO,
and a `Role`-aware stub for `AdminGuard` (re-use the approach from
`tests/backend/admin-system-authorization.spec.ts`).
- Stub `ConfirmationTokenService.consume` to record the
`{ userId, operation, stagingId }` argument and resolve `{ id: 'tok' }`.
- Stub `RestoreService.commitRestore` to record its argument and resolve
`{ restoresPerformed: true, revokeUserId: userId }`.
- Stub `AuthService.reauthenticateAdmin` to resolve an admin user, and
`AuthService.revokeAllRefreshSessions` to a no-op.
- Stub `RestoreService.stageArchive` to return a fixed
`{ stagingId: 'stage-1', expiresAt, summary }` so the validate→confirm
→commit pipeline can be exercised when needed.
- **Cases covered (minimal, focused):**
1. `commitRestore` forwards `stagingId` to `RestoreService.commitRestore`
and returns `{ ok: true, operation: 'restore-backup' }` when the body
contains `{ operation, token, stagingId }`.
2. `commitRestore` returns
`SYSTEM_RESTORE_STAGE_EXPIRED` (HTTP 400) when `stagingId` is omitted
**before** consuming the token (i.e. the fix returns the error early
rather than reaching `RestoreService`).
3. `commitRestore` returns `SYSTEM_TOKEN_MISMATCH` when `operation` is
`reset-scores` (regression guard for the existing controller branch).
- **Run command:** `npm test` from the repo root (already wired in
`package.json:test`).
## 5. Verification
- Manual smoke test against the dev stack (`npm run dev`):
1. POST `/api/v1/admin/system/restore/validate` with a valid backup → expect
`stagingId`.
2. POST `/api/v1/admin/system/confirmations` with `{ operation: 'restore-backup', password, stagingId }` → expect `token`.
3. POST `/api/v1/admin/system/restore/commit` with `{ operation: 'restore-backup', token, stagingId }`
→ expect `{ ok: true, operation: 'restore-backup' }`, the live DB swapped,
and the admin session revoked (redirect to `/login`).
- The error path must still surface
`SYSTEM_RESTORE_VALIDATION_FAILED` for malformed JSON bodies.
+32
View File
@@ -0,0 +1,32 @@
# Implementation Plan: Job 918 — Admin Area System Database Backup/Restore and Danger Zone 1.01
## 1. Architectural Reconnaissance
- **Codebase style & conventions:** Node.js monorepo using strict TypeScript, NestJS controllers/services on the backend, Angular standalone components and signals on the frontend, async/await, dependency injection, Zod request/archive validation, and stable `ApiError` codes. The requested UI and API flow already exists; the defect is isolated to the backend restore rebuild. `RestoreService.cloneAndOverwriteDatabase()` clones the live SQLite file and executes `DELETE` statements while the live schema's `trg_user_last_admin_delete` trigger remains active. `PRAGMA foreign_keys = OFF` does not disable triggers, so deleting the sole live admin raises `LAST_ADMIN` before archived users can be inserted.
- **Data Layer:** SQLite through TypeORM's `better-sqlite3` driver. Runtime migrations create a database-level invariant with `trg_user_last_admin_update` and `trg_user_last_admin_delete`. Backups dynamically include application tables but exclude `admin_operation_token`, `migrations`, and SQLite internals. Restore uses a cloned database, direct parameterized `better-sqlite3` statements, and a database/uploads filesystem swap with rollback handles.
- **Test Framework & Structure:** Jest 29 with `ts-jest`; backend tests live under `/repo/tests/backend` and are selected by `/repo/tests/jest.config.js`. Root `npm test` runs backend and frontend projects, while `npm run test:backend` runs the focused backend suite. Tests must remain CLI-only and use isolated temporary SQLite/upload/staging paths rather than shared mutable resources.
- **Required Tools & Dependencies:** No new package, system tool, global CLI, or persistent `/data` asset is required. Existing Node.js/npm, TypeScript, Jest, Nest testing utilities, TypeORM, and `better-sqlite3` are sufficient. `setup.sh` already installs dependencies and rebuilds the native SQLite binding, so it should not be changed.
## 2. Impacted Files
- **To Modify:**
- `backend/src/modules/admin/system/restore.service.ts` — make the offline clone rebuild temporarily suspend the last-admin triggers, restore all archived rows, validate invariants, and reinstate the triggers before the file is eligible for swapping.
- `tests/backend/admin-system-restore-commit.spec.ts` — add a focused real-file restore regression covering a backup that contains the same sole admin as the live database, restored table content, uploads, and trigger preservation; retain the existing controller/DTO contract coverage.
- **To Create:** None.
## 3. Proposed Changes
1. **Database / Schema Migration:**
- Do not add or alter a migration: the `LAST_ADMIN` triggers are correct for ordinary runtime user mutations.
- In the private staged database clone only, read the exact SQL definitions of `trg_user_last_admin_update` and `trg_user_last_admin_delete` from `sqlite_master`, fail closed if the expected trigger definitions cannot be captured, and drop those triggers before clearing application tables.
- Keep foreign keys disabled during the bulk replacement, clear all known application tables, and insert archived rows in the existing dependency-safe order with parameterized values.
- Before closing the staged clone, verify the archive produced at least one admin user so a restore cannot bypass the deployment-wide invariant. Recreate both captured triggers exactly, re-enable foreign keys, run `PRAGMA foreign_key_check`, and treat any violations or trigger-restoration failure as a restore failure. This confines the temporary invariant suspension to an offline candidate file and ensures the candidate has full protections before swap.
2. **Backend Logic & APIs:**
- Preserve the existing `POST /api/v1/admin/system/restore/validate`, confirmation-token, and `POST /restore/commit` contracts; `stagingId` already survives Zod validation and reaches `RestoreService.commitRestore()`.
- Refactor the staged rebuild cleanup so database closure and pragma/trigger restoration are deterministic and original failures are not masked by cleanup. Any rebuild, integrity, swap, or verification error must continue through the existing rollback path and return `SYSTEM_RESTORE_ROLLED_BACK` without changing live data/uploads.
- Keep the existing filesystem transaction: build and validate the candidate first, copy staged uploads beside the live upload root, swap database and uploads together, verify the swapped database, then commit rollback artifacts.
- Keep post-success session revocation in `AdminSystemController.commitRestore()`. After the restored database is live, revoke refresh sessions for the authenticated admin ID; the frontend's already-wired `forceServerInvalidation()` then clears the access session and redirects all tabs to login.
3. **Frontend UI Integration:**
- No frontend changes are planned. The file picker, staging summary, reauthentication modal, confirmation phrase, restore commit request, rollback alert, data-cache invalidation, and forced-login behavior are already wired in `frontend/src/app/features/admin/system/system.component.ts`, `system.service.ts`, and the auth/system data-change services.
## 4. Test Strategy
- **Target Unit Test File:** `tests/backend/admin-system-restore-commit.spec.ts`.
- **Mocking Strategy:** Keep existing controller dependencies mocked for request forwarding/token/session-revocation assertions. Add one minimal service-level integration regression using a unique temporary directory, a real file-backed `better-sqlite3` database, real `RestoreService`, and real `FilesystemTransactionService`; mock only `ConfigService` and the live `DataSource.query` verification boundary as needed. Arrange a live database with the migrated last-admin triggers, one sole admin, pre-restore challenge/blog/setting rows, and live uploads; stage a valid archive containing that same admin plus different table values and replacement challenge/icon files; commit; then open the swapped file independently and assert archived counts/settings, exact upload replacement, and both last-admin triggers still exist and reject deleting the sole restored admin. Add one focused invalid-archive invariant case with no admin and assert `SYSTEM_RESTORE_ROLLED_BACK` plus unchanged live database/uploads. Clean all temporary files in `afterEach`/`afterAll`; no UI, network, visual checks, shared `/data`, or large mock environment.
- **Execution:** Follow red-green-refactor: first run the focused backend spec to confirm the `LAST_ADMIN` failure, implement the smallest restore-service correction, rerun the focused spec, then run `npm run test:backend`, root `npm test`, `npm --workspace backend run build`, and the repository lint/typecheck commands if present (the current package scripts expose build but no dedicated lint/typecheck command).