AI Implementation feature(918): Admin Area System: Database Backup/Restore and Danger Zone 1.01 (#63)
This commit was merged in pull request #63.
This commit is contained in:
@@ -1,155 +0,0 @@
|
||||
# Implementation Plan: Job 916 — Admin System Restore Commit (stagingId stripped from request body)
|
||||
|
||||
## Root cause (read-only investigation)
|
||||
|
||||
The destructive restore never executes because the Zod schema used to validate
|
||||
`POST /api/v1/admin/system/restore/commit` **drops the `stagingId` field** from
|
||||
the request body before the controller runs.
|
||||
|
||||
`backend/src/modules/admin/system/dto/admin-system.dto.ts:23`:
|
||||
|
||||
```ts
|
||||
export const AdminRestoreCommitBodySchema = AdminDangerConfirmBodySchema;
|
||||
```
|
||||
|
||||
`AdminDangerConfirmBodySchema` (line 18) only allows `{ operation, token }`.
|
||||
Zod's `safeParse` / the NestJS `ZodValidationPipe` therefore strip every
|
||||
unknown field — including `stagingId` — from the parsed body. The controller
|
||||
then receives `body.stagingId === undefined`, falls back to the empty string
|
||||
(`body.stagingId ?? ''`), and `RestoreService.commitRestore('')` finds no
|
||||
matching stage and throws `SYSTEM_RESTORE_STAGE_EXPIRED`
|
||||
(`RestoreService.commitRestore` at `restore.service.ts:223-231`).
|
||||
|
||||
The frontend already sends the field correctly
|
||||
(`frontend/src/app/features/admin/system/system.service.ts:42`):
|
||||
|
||||
```ts
|
||||
async commitRestore(payload: { operation: 'restore-backup'; token: string; stagingId: string })
|
||||
```
|
||||
|
||||
so the deviation is purely server-side. The negative path (random JSON) still
|
||||
fails with `SYSTEM_RESTORE_VALIDATION_FAILED` because the malformed archive
|
||||
rejects at `validate`, never reaching `commit`.
|
||||
|
||||
The fix is to give `AdminRestoreCommitBodySchema` its own declaration that
|
||||
**includes** `stagingId` (mirroring `AdminReauthBodySchema` at line 12-16), and
|
||||
to tighten the controller so a missing `stagingId` returns a precise
|
||||
`SYSTEM_RESTORE_STAGE_EXPIRED` instead of relying on the empty-string fallback.
|
||||
|
||||
## 1. Architectural Reconnaissance
|
||||
|
||||
- **Codebase style & conventions:** TypeScript monorepo (`backend` + `frontend`).
|
||||
Backend is NestJS with `class-validator`/`zod` mixed; the admin-system module
|
||||
uses `ZodValidationPipe` per-route. Strict async/await, error handling via
|
||||
`ApiError` + `ERROR_CODES` (`backend/src/common/errors/`).
|
||||
- **Data Layer:** SQLite via `better-sqlite3` + TypeORM. Restore streams are
|
||||
in-memory (`RestoreService.stages: Map<stagingId, StagedRestore>`) plus a
|
||||
sibling directory tree under `<DATA_DIR>/.system-staging/restore-<id>/`.
|
||||
- **Test Framework & Structure:** Jest with two projects in
|
||||
`tests/jest.config.js`. Tests live in `tests/backend/*.spec.ts` and
|
||||
`tests/frontend/*.spec.ts`. Single command from the root:
|
||||
`npm test` (already wired in `package.json`).
|
||||
- **Required Tools & Dependencies:** No new tools. Existing: Node 20, NestJS,
|
||||
better-sqlite3, zod, Jest, ts-jest. `setup.sh` already builds the app and
|
||||
rebuilds the native binding — no updates needed.
|
||||
|
||||
## 2. Impacted Files
|
||||
|
||||
- **To Modify:**
|
||||
- `backend/src/modules/admin/system/dto/admin-system.dto.ts` — replace the
|
||||
`AdminRestoreCommitBodySchema` alias with a proper schema that accepts
|
||||
`stagingId` (required for `restore-backup`).
|
||||
- `backend/src/modules/admin/system/admin-system.controller.ts` — validate
|
||||
`stagingId` is present before consuming the token, and stop feeding an
|
||||
empty string into `commitRestore`.
|
||||
- **To Create:**
|
||||
- `tests/backend/admin-system-restore-commit.spec.ts` — focused e2e test
|
||||
that the `restore/commit` endpoint forwards `stagingId` to
|
||||
`RestoreService.commitRestore` and atomically swaps the live DB.
|
||||
|
||||
## 3. Proposed Changes
|
||||
|
||||
1. **DTO fix** — `backend/src/modules/admin/system/dto/admin-system.dto.ts`
|
||||
- Remove the `AdminRestoreCommitBodySchema = AdminDangerConfirmBodySchema`
|
||||
alias (line 23).
|
||||
- Add a dedicated schema that requires `stagingId` for `restore-backup`
|
||||
while still letting the safe schema accept the other two operations in
|
||||
shared code paths:
|
||||
|
||||
```ts
|
||||
export const AdminRestoreCommitBodySchema = z.object({
|
||||
operation: z.enum(ADMIN_OPERATION_KINDS),
|
||||
token: z.string().min(1).max(512),
|
||||
stagingId: z.string().min(1).max(128).optional(),
|
||||
});
|
||||
```
|
||||
|
||||
- Keep the `AdminOperationKind` re-export and the
|
||||
`AdminRestoreCommitBody` `z.infer` type aligned so the controller
|
||||
signature remains `body: { operation; token; stagingId? }`.
|
||||
|
||||
2. **Controller hardening** — `backend/src/modules/admin/system/admin-system.controller.ts:95-121`
|
||||
- After the `body.operation === 'restore-backup'` check, throw
|
||||
`SYSTEM_RESTORE_STAGE_EXPIRED` with HTTP 400 if `body.stagingId` is
|
||||
falsy. This produces a precise error instead of the empty-string
|
||||
fallback that currently slips through to `RestoreService`.
|
||||
- Pass `body.stagingId` directly to `commitRestore` (no `?? ''`).
|
||||
- Existing behavior for `reset-scores` / `wipe-challenges` (which use
|
||||
`AdminDangerConfirmBodySchema`) is unchanged.
|
||||
|
||||
3. **No frontend changes required.** `system.service.ts` already POSTs
|
||||
`{ operation, token, stagingId }`; the modal already calls it after
|
||||
`confirmations`. Once the server schema preserves `stagingId` the happy
|
||||
path works end-to-end.
|
||||
|
||||
4. **No DB migration needed.** `admin_operation_token` already exists;
|
||||
`stagingId` is bound through the in-memory `RestoreService.stages` map.
|
||||
|
||||
## 4. Test Strategy
|
||||
|
||||
- **Target Unit Test File:** `tests/backend/admin-system-restore-commit.spec.ts` (new).
|
||||
Companion lightweight spec additions to
|
||||
`tests/backend/admin-system-restore-validation.spec.ts` are optional; the
|
||||
new file is sufficient and keeps the diff small.
|
||||
- **Mocking Strategy:**
|
||||
- Build a minimal `Test.createTestingModule` with just
|
||||
`AdminSystemController`, `RestoreService`, `ConfirmationTokenService`,
|
||||
`AuthService`, `BackupService`, `DangerZoneService`, the
|
||||
`AdminOperationTokenEntity` repository, the `RestoreArchiveSchema` DTO,
|
||||
and a `Role`-aware stub for `AdminGuard` (re-use the approach from
|
||||
`tests/backend/admin-system-authorization.spec.ts`).
|
||||
- Stub `ConfirmationTokenService.consume` to record the
|
||||
`{ userId, operation, stagingId }` argument and resolve `{ id: 'tok' }`.
|
||||
- Stub `RestoreService.commitRestore` to record its argument and resolve
|
||||
`{ restoresPerformed: true, revokeUserId: userId }`.
|
||||
- Stub `AuthService.reauthenticateAdmin` to resolve an admin user, and
|
||||
`AuthService.revokeAllRefreshSessions` to a no-op.
|
||||
- Stub `RestoreService.stageArchive` to return a fixed
|
||||
`{ stagingId: 'stage-1', expiresAt, summary }` so the validate→confirm
|
||||
→commit pipeline can be exercised when needed.
|
||||
|
||||
- **Cases covered (minimal, focused):**
|
||||
1. `commitRestore` forwards `stagingId` to `RestoreService.commitRestore`
|
||||
and returns `{ ok: true, operation: 'restore-backup' }` when the body
|
||||
contains `{ operation, token, stagingId }`.
|
||||
2. `commitRestore` returns
|
||||
`SYSTEM_RESTORE_STAGE_EXPIRED` (HTTP 400) when `stagingId` is omitted
|
||||
**before** consuming the token (i.e. the fix returns the error early
|
||||
rather than reaching `RestoreService`).
|
||||
3. `commitRestore` returns `SYSTEM_TOKEN_MISMATCH` when `operation` is
|
||||
`reset-scores` (regression guard for the existing controller branch).
|
||||
|
||||
- **Run command:** `npm test` from the repo root (already wired in
|
||||
`package.json:test`).
|
||||
|
||||
## 5. Verification
|
||||
|
||||
- Manual smoke test against the dev stack (`npm run dev`):
|
||||
1. POST `/api/v1/admin/system/restore/validate` with a valid backup → expect
|
||||
`stagingId`.
|
||||
2. POST `/api/v1/admin/system/confirmations` with `{ operation: 'restore-backup', password, stagingId }` → expect `token`.
|
||||
3. POST `/api/v1/admin/system/restore/commit` with `{ operation: 'restore-backup', token, stagingId }`
|
||||
→ expect `{ ok: true, operation: 'restore-backup' }`, the live DB swapped,
|
||||
and the admin session revoked (redirect to `/login`).
|
||||
- The error path must still surface
|
||||
`SYSTEM_RESTORE_VALIDATION_FAILED` for malformed JSON bodies.
|
||||
@@ -0,0 +1,32 @@
|
||||
# Implementation Plan: Job 918 — Admin Area System Database Backup/Restore and Danger Zone 1.01
|
||||
|
||||
## 1. Architectural Reconnaissance
|
||||
- **Codebase style & conventions:** Node.js monorepo using strict TypeScript, NestJS controllers/services on the backend, Angular standalone components and signals on the frontend, async/await, dependency injection, Zod request/archive validation, and stable `ApiError` codes. The requested UI and API flow already exists; the defect is isolated to the backend restore rebuild. `RestoreService.cloneAndOverwriteDatabase()` clones the live SQLite file and executes `DELETE` statements while the live schema's `trg_user_last_admin_delete` trigger remains active. `PRAGMA foreign_keys = OFF` does not disable triggers, so deleting the sole live admin raises `LAST_ADMIN` before archived users can be inserted.
|
||||
- **Data Layer:** SQLite through TypeORM's `better-sqlite3` driver. Runtime migrations create a database-level invariant with `trg_user_last_admin_update` and `trg_user_last_admin_delete`. Backups dynamically include application tables but exclude `admin_operation_token`, `migrations`, and SQLite internals. Restore uses a cloned database, direct parameterized `better-sqlite3` statements, and a database/uploads filesystem swap with rollback handles.
|
||||
- **Test Framework & Structure:** Jest 29 with `ts-jest`; backend tests live under `/repo/tests/backend` and are selected by `/repo/tests/jest.config.js`. Root `npm test` runs backend and frontend projects, while `npm run test:backend` runs the focused backend suite. Tests must remain CLI-only and use isolated temporary SQLite/upload/staging paths rather than shared mutable resources.
|
||||
- **Required Tools & Dependencies:** No new package, system tool, global CLI, or persistent `/data` asset is required. Existing Node.js/npm, TypeScript, Jest, Nest testing utilities, TypeORM, and `better-sqlite3` are sufficient. `setup.sh` already installs dependencies and rebuilds the native SQLite binding, so it should not be changed.
|
||||
|
||||
## 2. Impacted Files
|
||||
- **To Modify:**
|
||||
- `backend/src/modules/admin/system/restore.service.ts` — make the offline clone rebuild temporarily suspend the last-admin triggers, restore all archived rows, validate invariants, and reinstate the triggers before the file is eligible for swapping.
|
||||
- `tests/backend/admin-system-restore-commit.spec.ts` — add a focused real-file restore regression covering a backup that contains the same sole admin as the live database, restored table content, uploads, and trigger preservation; retain the existing controller/DTO contract coverage.
|
||||
- **To Create:** None.
|
||||
|
||||
## 3. Proposed Changes
|
||||
1. **Database / Schema Migration:**
|
||||
- Do not add or alter a migration: the `LAST_ADMIN` triggers are correct for ordinary runtime user mutations.
|
||||
- In the private staged database clone only, read the exact SQL definitions of `trg_user_last_admin_update` and `trg_user_last_admin_delete` from `sqlite_master`, fail closed if the expected trigger definitions cannot be captured, and drop those triggers before clearing application tables.
|
||||
- Keep foreign keys disabled during the bulk replacement, clear all known application tables, and insert archived rows in the existing dependency-safe order with parameterized values.
|
||||
- Before closing the staged clone, verify the archive produced at least one admin user so a restore cannot bypass the deployment-wide invariant. Recreate both captured triggers exactly, re-enable foreign keys, run `PRAGMA foreign_key_check`, and treat any violations or trigger-restoration failure as a restore failure. This confines the temporary invariant suspension to an offline candidate file and ensures the candidate has full protections before swap.
|
||||
2. **Backend Logic & APIs:**
|
||||
- Preserve the existing `POST /api/v1/admin/system/restore/validate`, confirmation-token, and `POST /restore/commit` contracts; `stagingId` already survives Zod validation and reaches `RestoreService.commitRestore()`.
|
||||
- Refactor the staged rebuild cleanup so database closure and pragma/trigger restoration are deterministic and original failures are not masked by cleanup. Any rebuild, integrity, swap, or verification error must continue through the existing rollback path and return `SYSTEM_RESTORE_ROLLED_BACK` without changing live data/uploads.
|
||||
- Keep the existing filesystem transaction: build and validate the candidate first, copy staged uploads beside the live upload root, swap database and uploads together, verify the swapped database, then commit rollback artifacts.
|
||||
- Keep post-success session revocation in `AdminSystemController.commitRestore()`. After the restored database is live, revoke refresh sessions for the authenticated admin ID; the frontend's already-wired `forceServerInvalidation()` then clears the access session and redirects all tabs to login.
|
||||
3. **Frontend UI Integration:**
|
||||
- No frontend changes are planned. The file picker, staging summary, reauthentication modal, confirmation phrase, restore commit request, rollback alert, data-cache invalidation, and forced-login behavior are already wired in `frontend/src/app/features/admin/system/system.component.ts`, `system.service.ts`, and the auth/system data-change services.
|
||||
|
||||
## 4. Test Strategy
|
||||
- **Target Unit Test File:** `tests/backend/admin-system-restore-commit.spec.ts`.
|
||||
- **Mocking Strategy:** Keep existing controller dependencies mocked for request forwarding/token/session-revocation assertions. Add one minimal service-level integration regression using a unique temporary directory, a real file-backed `better-sqlite3` database, real `RestoreService`, and real `FilesystemTransactionService`; mock only `ConfigService` and the live `DataSource.query` verification boundary as needed. Arrange a live database with the migrated last-admin triggers, one sole admin, pre-restore challenge/blog/setting rows, and live uploads; stage a valid archive containing that same admin plus different table values and replacement challenge/icon files; commit; then open the swapped file independently and assert archived counts/settings, exact upload replacement, and both last-admin triggers still exist and reject deleting the sole restored admin. Add one focused invalid-archive invariant case with no admin and assert `SYSTEM_RESTORE_ROLLED_BACK` plus unchanged live database/uploads. Clean all temporary files in `afterEach`/`afterAll`; no UI, network, visual checks, shared `/data`, or large mock environment.
|
||||
- **Execution:** Follow red-green-refactor: first run the focused backend spec to confirm the `LAST_ADMIN` failure, implement the smallest restore-service correction, rerun the focused spec, then run `npm run test:backend`, root `npm test`, `npm --workspace backend run build`, and the repository lint/typecheck commands if present (the current package scripts expose build but no dedicated lint/typecheck command).
|
||||
Reference in New Issue
Block a user