Example report for a fictional application. Not a real customer or audit. AcmeBook does not exist. Every finding, file path, person, date and measurement in the evidence was invented to show what a FINISHER report looks like. The scores, band, ceiling and fix order were computed by the real scorer from that invented evidence.
AcmeBook (fictional example app)
FINISHER Score report · generated 2026-10-09 05:22 UTC · catalog 1.3.0 · FINISHER 2.1.0
69.0 / 100
PAID BETA WITH FIXES
Uncapped score78.0
Coverage100.0%167 of 167 checks
Open P0 / P1 / P20 / 9 / 5
Capped at 69: 9 P1 issue(s) open. Any open launch blocker holds the score at 69 however good the rest is. The uncapped score shows the progress behind the cap.
Domain scores
| Domain | Score | Weight | Checks | Open P0 | Open P1 | |
|---|---|---|---|---|---|---|
| D01 Identity, Auth & Authorization | 88.2 | 12 | 17 | 0 | 0 | |
| D02 Secrets & Key Management | 93.8 | 8 | 8 | 0 | 0 | |
| D03 Data Layer, Migrations & Backups | 79.8 | 9 | 13 | 0 | 1 | |
| D04 API & Backend Correctness | 82.1 | 7 | 11 | 0 | 0 | |
| D05 Frontend, UX & Accessibility | 65.2 | 6 | 8 | 0 | 2 | |
| D06 Application Security | 81.5 | 9 | 9 | 0 | 0 | |
| D07 Supply Chain & Dependency Integrity | 74.0 | 6 | 9 | 0 | 0 | |
| D08 CI/CD & Release Engineering | 83.0 | 6 | 9 | 0 | 0 | |
| D09 Environments, Config & Infrastructure | 84.7 | 5 | 6 | 0 | 0 | |
| D10 Observability & Error Tracking | 80.6 | 7 | 8 | 0 | 0 | |
| D11 Reliability, Backup & Incident Response | 50.0 | 7 | 7 | 0 | 2 | |
| D12 Performance, Caching & Scale | 64.3 | 5 | 7 | 0 | 0 | |
| D13 Cost Control & Abuse Prevention | 73.9 | 5 | 7 | 0 | 1 | |
| D14 AI & Agent Safety and Economics | 70.0 | 6 | 10 | 0 | 1 | |
| D15 Privacy, Legal & Compliance | 73.1 | 6 | 9 | 0 | 0 | |
| D16 Testing & Verification | 86.5 | 6 | 8 | 0 | 0 | |
| D18 Documentation & Handoff | 76.5 | 4 | 6 | 0 | 1 | |
| D19 Product Truth & Outcome Measurement | 75.0 | 3 | 4 | 0 | 0 | |
| D20 Functional Correctness | 84.4 | 8 | 6 | 0 | 0 | |
| D22 Email & Messaging | 69.6 | 4 | 5 | 0 | 1 |
P0: blocks any launch (0 open)
None open.
P1: blocks a paid or public launch (9 open)
| Check | Now | Needed | What was found |
|---|---|---|---|
| DATA-10 Data retention and deletion rules are defined per data class | 1 claimed | 3 | A privacy policy promises deletion 'when no longer needed'. There is no per-class retention table: bookings, invitee contact details, LLM prompt logs and email events are all kept indefinitely, and no job deletes anything. |
| FE-05 Automated accessibility scan passes on core pages | 2 implemented | 3 | axe run on 8 core pages: 2 violations remain. The public booking page time-slot buttons have no accessible name (each reads 'button'), and the secondary text colour on the booking page fails contrast at 3.1:1. |
| FE-06 Keyboard-only and screen-reader passes completed on critical journeys | 1 claimed | 3 | Keyboard pass on the booking journey stopped at the slot picker: the calendar grid is not reachable by Tab and has no arrow-key support. No screen-reader pass has been done. |
| REL-01 RTO and RPO are defined, written down, and consistent with the backup configuration | 1 claimed | 3 | README says 'Neon handles backups'. There is no written RTO or RPO; the only measured restore time (41 minutes, May 2026) predates the database tripling in size. |
| REL-02 An incident runbook exists covering the realistic failure set | 1 claimed | 3 | No incident runbook. The 2026-08-14 incident notes show the fix was found by reading code during the outage. |
| COST-03 Rate limiting exists on all public and expensive endpoints | 2 implemented | 3 | Rate limits exist on auth, assistant and API-key routes. The public booking endpoint POST /api/public/[org]/book has no rate limit: 500 requests in a minute from one IP were all accepted and created 500 pending bookings and 500 confirmation emails in staging. |
| AI-07 Per-user and per-tier quotas are token-budgeted, not request-counted | 1 claimed | 3 | The only limit is 30 requests per user per hour. There is no token budget per workspace or plan, so one workspace with many seats and long prompts can spend the whole provider cap; in September the top workspace used 38% of LLM spend. |
| DOC-03 Deploy and rollback procedures are written down and executable by someone else | 2 implemented | 3 | docs/ops/deploy.md covers deploys. Rollback is a single line ('use vercel rollback') with nothing on the worker or a migration that must be reversed, and nobody but its author has followed it. |
| EMAIL-03 Any user-triggerable send is rate limited per recipient and per actor | 2 implemented | 3 | Magic-link and invite emails are rate limited per recipient and per actor. Booking confirmations are not: each booking sends one, so the unthrottled booking endpoint (COST-03) can send unlimited confirmations to any address. |
P2: blocks scale (5 open)
| Check | Now | Needed | What was found |
|---|---|---|---|
| APPSEC-10 A written threat model exists for the highest-value assets | 1 claimed | 2 | Threat-model notes exist as a bullet list in a pull request description from 2025; no maintained threat model for the booking, billing and LLM paths. |
| SUP-08 An SBOM is generated per release and retained | 0 absent | 2 | Looked for an SBOM step in the release workflow and the release assets: none is generated. |
| SUP-09 Published artifacts carry build provenance | 0 absent | 2 | Releases are Vercel deployments and a worker container image; neither carries build provenance or attestations. |
| REL-06 There is a way to tell customers what is happening | 1 claimed | 2 | Customers are told about incidents by individual email from support. No status page and no template. |
| PERF-05 A load test has been run at a realistic launch multiple and the breaking point is known | 0 absent | 2 | No load test has been run. The busiest customer's Monday-morning peak is the only data point. |
Highest-value next steps
| Check | Tier | Points available |
|---|---|---|
| REL-01 RTO and RPO are defined, written down, and consistent with the backup configuration | P1 | 12 |
| REL-02 An incident runbook exists covering the realistic failure set | P1 | 12 |
| AI-07 Per-user and per-tier quotas are token-budgeted, not request-counted | P1 | 9 |
| DATA-10 Data retention and deletion rules are defined per data class | P1 | 9 |
| FE-06 Keyboard-only and screen-reader passes completed on critical journeys | P1 | 9 |
| COST-03 Rate limiting exists on all public and expensive endpoints | P1 | 8 |
| DOC-03 Deploy and rollback procedures are written down and executable by someone else | P1 | 6 |
| EMAIL-03 Any user-triggerable send is rate limited per recipient and per actor | P1 | 6 |
| FE-05 Automated accessibility scan passes on core pages | P1 | 6 |
| PERF-05 A load test has been run at a realistic launch multiple and the breaking point is known | P2 | 12 |
Every assessed check
Scale: 0 absent, 1 claimed, 2 implemented, 3 verified by an artifact a stranger could re-run, 4 enforced by automation.
AUTH-01 Authentication uses a battle-tested provider or framework, not hand-rolled code · P0 · score 4
Evidence: Auth.js v5 with the Drizzle adapter and database sessions; Google OAuth and email magic links only, no passwords. Session established in src/auth.ts; no custom crypto or token code anywhere in src/.
Artifact: src/auth.ts; `pnpm test tests/auth/session.spec.ts` (9 cases passing)
Enforcement: Semgrep rule acme/no-custom-jwt in the CI sast job fails any import of jsonwebtoken or jose outside src/auth.ts
AUTH-02 Every non-public route and API endpoint enforces authentication server-side · P0 · score 4
Evidence: All 61 route handlers and server actions go through withSession() or are on the 4-entry public allowlist (booking page, slot lookup, Stripe webhook, health). Unauthenticated requests to the 57 protected routes return 401.
Artifact: docs/security/route-inventory.md; tests/auth/unauthenticated.spec.ts (57 routes, generated from the Next.js route manifest)
Enforcement: CI job test:authz regenerates the route list from the build manifest and fails when a route is neither wrapped nor allowlisted
AUTH-03 Sessions expire, logout truly invalidates, and session IDs rotate on privilege change · P0 · score 4
Evidence: Database sessions expire after 14 days idle and 30 days absolute; sign-out deletes the session row; replaying the old cookie after sign-out returns 401. Session token rotates on sign-in and on role change.
Artifact: tests/auth/session.spec.ts::logout-replay, ::rotate-on-role-change
Enforcement: Both tests run in the CI test:authz job, a required check
AUTH-05 Password reset and email verification tokens are single-use, short-lived, and unguessable · P0 · score 3
Evidence: Magic-link tokens are 32 random bytes, hashed at rest, single-use and expire in 15 minutes. Second use of a link and use after expiry both rejected.
Artifact: tests/auth/magic-link.spec.ts::single-use, ::expires-after-15m
AUTHZ-01 Every resource access checks ownership, not just authentication · P0 · score 4
Evidence: User A / User B tampering run executed by a person against staging on 2026-09-29: 214 ID substitutions across URL, body, query and headers on bookings, event types, availability, members and invoices; all returned 404 and were logged.
Artifact: docs/evidence/2026-09-29-authz-tamper-run.txt; tests/authz/cross-org.spec.ts (214 cases)
Enforcement: CI job test:authz runs tests/authz/cross-org.spec.ts on every pull request; required status check on main
AUTHZ-02 Tenant isolation is enforced at the database layer, not only in application code · P0 · score 4
Evidence: Postgres row-level security on every org-scoped table (14 tables); the app connects as acme_app, which is not the table owner and cannot bypass RLS. Each transaction sets app.current_org via set_config.
Artifact: db/policies/rls.sql; tests/db/rls.spec.ts::tenant A sees zero rows of tenant B (runs against Postgres 16 in CI)
Enforcement: CI job test:db fails if any table with an org_id column lacks an enabled RLS policy (pg_policies check in tests/db/rls-coverage.spec.ts)
AUTHZ-03 Admin capability is server-verified and cannot be obtained by client-side manipulation · P0 · score 4
Evidence: Role is read from the memberships table server-side on every admin action; the client never sends a role. Forging role=owner in the session cookie payload and request body has no effect.
Artifact: tests/authz/admin-forgery.spec.ts (member forges owner on all 11 admin endpoints, all 403)
Enforcement: CI job test:authz runs the forgery suite as a required check
AUTH-04 Session cookies use Secure, HttpOnly, SameSite, and the __Host- prefix where possible · P1 · score 3
Evidence: Login response sets __Host-acme.session with Secure; HttpOnly; SameSite=Lax; Path=/.
Artifact: docs/evidence/2026-09-29-set-cookie.txt (raw header from staging login)
AUTH-06 Login, reset, signup, and OTP endpoints are rate limited per account and per IP · P1 · score 3
Evidence: Magic-link requests limited to 5 per email per hour and 20 per IP per hour (Upstash ratelimit); the 6th request returns 429 with Retry-After.
Artifact: scripts/abuse/magic-link-flood.sh output in docs/evidence/2026-09-29-magic-link-throttle.txt
AUTH-07 Password policy follows current guidance: length over composition, breach-list blocking · P1 · score 3
Evidence: There are no passwords: sign-in is Google OAuth or a magic link. No password column exists in the schema.
Artifact: db/schema/users.ts; `grep -ri password src/ db/` returns only the OAuth provider config comment
AUTH-08 MFA or passkeys are mandatory for admin and billing-privileged accounts · P1 · score 4
Evidence: Owner and billing roles must enrol a passkey or TOTP before reaching /settings/billing or any admin route; enforced in middleware.
Artifact: tests/auth/mfa-required.spec.ts::owner-without-mfa-redirected
Enforcement: middleware.ts requireMfaForPrivileged() plus the required CI test
AUTH-10 Tokens are validated with a pinned algorithm and are revocable · P1 · score 3
Evidence: No bearer JWTs are issued; sessions are opaque database rows deletable from the admin console. API keys for the public API are hashed and revocable per key.
Artifact: tests/auth/api-keys.spec.ts::revoked-key-rejected
AUTH-11 Account deletion exists and actually works end to end · P1 · score 3
Evidence: Account deletion removes the user, memberships, sessions and calendar tokens and anonymises their bookings; verified row counts after deleting a test account. Backups age out after 14 days, stated in the privacy policy.
Artifact: tests/e2e/account-deletion.spec.ts; docs/evidence/2026-09-30-account-deletion-rows.txt
AUTHZ-04 Authorization model is written down and matches the code · P1 · score 3
Evidence: Permission matrix for owner, admin, member and billing roles written in docs/ARCHITECTURE.md and cross-checked against src/lib/permissions.ts on 2026-09-30; no mismatches.
Artifact: docs/ARCHITECTURE.md#permission-matrix; src/lib/permissions.ts
AUTH-09 OAuth/OIDC integrations use PKCE, exact redirect URI matching, and state/nonce binding · P2 · score 3
Evidence: Google OAuth via Auth.js uses PKCE (S256) and state; redirect URI registered exactly as https://app.acmebook.example/api/auth/callback/google.
Artifact: docs/evidence/2026-09-29-oauth-authorize-request.txt (code_challenge_method=S256)
AUTH-12 Privileged actions are recorded in an append-only audit log · P2 · score 3
Evidence: Admin and billing actions write to audit_log, an append-only table (UPDATE and DELETE revoked from acme_app).
Artifact: db/policies/audit_log.sql; tests/db/audit-log-append-only.spec.ts
AUTH-13 MFA or passkeys are available to users · P2 · score 3
Evidence: Passkeys and TOTP available to every user from /settings/security.
Artifact: tests/e2e/mfa-enrol.spec.ts
SEC-01 No secret is reachable from browser or mobile client code · P0 · score 4
Evidence: Built client bundle grepped for sk_, rk_, whsec_, sk-ant-, re_ and generic key/secret/token patterns: no matches. Network panel shows the browser calls only app.acmebook.example and js.stripe.com.
Artifact: scripts/ci/scan-client-bundle.sh against .next/static; docs/evidence/2026-09-29-bundle-scan.txt
Enforcement: CI job bundle-scan runs scripts/ci/scan-client-bundle.sh after build and fails on any match
SEC-02 Secret scan of the full git history passes, and anything ever exposed has been rotated · P0 · score 4
Evidence: gitleaks over the full history (1,912 commits): no findings. One Stripe test key committed in 2025 was rotated on 2025-11-03; rotation log names the key, date and person.
Artifact: `gitleaks git -v .` output in docs/evidence/2026-09-29-gitleaks.txt; docs/security/rotation-log.md
Enforcement: gitleaks pre-commit hook plus the CI secrets job on every pull request
SEC-03 Secrets live in a secret manager or platform env store, scoped per environment · P1 · score 4
Evidence: Secrets live in Vercel environment variables scoped to Production, Preview and Development, plus Doppler for the worker; .env is gitignored and .env.example has keys only.
Artifact: .env.example; docs/ops/environments.md#secrets
Enforcement: CI job env-example-check fails if .env.example contains a value
SEC-04 CI/CD uses short-lived federated credentials (OIDC), not long-lived cloud keys · P1 · score 4
Evidence: Deploys and migrations use GitHub OIDC to assume a role scoped to repo:<org>/acmebook:ref:refs/heads/main; no long-lived cloud keys in repository secrets.
Artifact: .github/workflows/deploy.yml; infra/iam/github-oidc.tf
Enforcement: Repository ruleset blocks adding new Actions secrets matching *_ACCESS_KEY; the deploy workflow has no static credentials
SEC-05 Service accounts and database credentials follow least privilege · P1 · score 3
Evidence: Three database roles: acme_app (DML on app tables, no DDL, cannot bypass RLS), acme_migrate (DDL, used only by the migrate job), acme_readonly (analytics). Grant listing reviewed 2026-09-30.
Artifact: db/roles.sql; docs/evidence/2026-09-30-grants.txt
SEC-06 Secrets never appear in logs, error messages, traces, screenshots, or AI prompts · P1 · score 4
Evidence: pino logger redacts authorization, cookie, stripe-signature, token and *_key paths; Sentry beforeSend scrubs request bodies. A forced error carrying an API key logged [REDACTED].
Artifact: tests/logging/redaction.spec.ts; docs/evidence/2026-09-29-redacted-log-line.txt
Enforcement: Redaction unit test runs in CI; Sentry data scrubbing rules enabled at the project level
SEC-08 Webhook endpoints verify provider signatures and reject replays · P1 · score 4
Evidence: Stripe and Resend webhooks verify signatures with a 5-minute tolerance; processed event IDs are stored, so a replayed valid event causes no second side effect.
Artifact: tests/webhooks/stripe.spec.ts::forged-signature-400, ::replay-is-noop
Enforcement: Both tests run in the required CI test job
SEC-07 Key rotation is documented, rehearsed, and possible without downtime · P2 · score 2
Evidence: Rotation steps are written for Stripe, Resend and the LLM provider in docs/ops/key-rotation.md. The Stripe test-key rotation in 2025 is the only rehearsal; there is no dated rehearsal for production keys.
DATA-01 Automated backups exist for every production datastore · P0 · score 4
Evidence: Neon production branch has point-in-time restore with 14-day history plus nightly logical dumps to a separate S3 bucket (30-day retention, object lock).
Artifact: infra/backups/pg-dump.yml; docs/evidence/2026-09-29-backup-config.txt
Enforcement: Scheduled backup job alerts #ops-alerts in Slack when a nightly dump is missing or under 80% of the previous size
DATA-02 A restore has actually been performed and the result verified · P0 · score 3
Evidence: One restore performed on 2026-05-12: the nightly dump was restored into a scratch database in 41 minutes and row counts on 14 tables matched production within the dump window. No restore has been done since, the database has roughly tripled, and the restore is not on a schedule.
Artifact: docs/evidence/2026-05-12-restore-log.md (elapsed time and verification queries)
Notes: Verified once; nothing keeps it true. See REL-01 for the gap this leaves.
DATA-03 Production, staging, and development use separate datastores · P0 · score 4
Evidence: Production, staging and development use separate Neon projects; connection hostnames differ per environment.
Artifact: docs/ops/environments.md#databases (hostnames redacted to the project id)
Enforcement: ENV-02 startup check refuses to boot production with a non-production DATABASE_URL host (src/env.ts rule)
DATA-04 No human or agent has standing write access to the production database · P0 · score 3
Evidence: No person has standing write access to production; the acme_migrate role is only reachable from the deploy workflow. Break-glass access is a time-boxed Neon role granted by the CTO and logged.
Artifact: docs/ops/break-glass.md; docs/evidence/2026-09-30-neon-access-list.txt
DATA-05 Every schema change is a versioned migration file in the repo · P1 · score 4
Evidence: All schema changes are Drizzle migrations in db/migrations (87 files). Empty database migrated from zero and the app booted.
Artifact: `pnpm db:migrate && pnpm start` from an empty Postgres 16, output in docs/evidence/2026-09-29-migrate-from-zero.txt
Enforcement: CI job test:db migrates an empty database from zero on every pull request
DATA-06 Migrations are linted for locking and destructive operations · P1 · score 4
Evidence: squawk lints every new migration for locking and destructive operations.
Artifact: `pnpm squawk db/migrations/*.sql` output in docs/evidence/2026-09-29-squawk.txt
Enforcement: CI job migration-lint runs squawk and blocks the pull request on errors
DATA-07 Rollback strategy for schema changes is defined and it is roll-forward-safe · P1 · score 3
Evidence: Expand-and-contract policy in docs/ops/migrations.md; CI runs the previous release's test suite against the new schema.
Artifact: tests/db/previous-release-compat.spec.ts
DATA-08 Constraints and indexes reflect real access patterns · P1 · score 3
Evidence: Foreign keys, unique (org_id, slug) constraints and composite indexes derived from pg_stat_statements; the top 10 queries all use an index.
Artifact: docs/evidence/2026-09-28-pg-stat-statements-top10.txt; db/schema/
DATA-10 Data retention and deletion rules are defined per data class · P1 · score 1
Evidence: A privacy policy promises deletion 'when no longer needed'. There is no per-class retention table: bookings, invitee contact details, LLM prompt logs and email events are all kept indefinitely, and no job deletes anything.
Notes: Open. Needs a data-class retention table and a scheduled purge job.
DATA-12 Idempotency keys prevent duplicate side effects from retries and double-clicks · P1 · score 3
Evidence: Booking creation requires an Idempotency-Key; two concurrent identical requests create one booking and one confirmation email. Stripe calls pass idempotency keys.
Artifact: tests/api/bookings-idempotency.spec.ts::concurrent-double-submit
DATA-13 Concurrent writes cannot silently lose updates · P1 · score 3
Evidence: Slot booking uses an exclusion constraint on (host_id, tstzrange) so two concurrent bookings of the same slot cannot both succeed; event-type edits use a version column and return 409 on a stale write.
Artifact: tests/db/double-booking.spec.ts; tests/api/event-type-conflict.spec.ts
DATA-14 Backfills and heavy migrations are batched, resumable, and tested at production scale · P1 · score 3
Evidence: The one large backfill (time-zone normalisation, 2026-06) ran in 5,000-row batches with a resume cursor; dry run against a production-size copy took 18 minutes and resumed after a forced kill.
Artifact: scripts/backfills/2026-06-tz-normalise.ts; docs/evidence/2026-06-14-backfill-dry-run.txt
DATA-11 Sensitive fields are encrypted at rest beyond disk-level encryption where warranted · P2 · score 2
Evidence: Google Calendar refresh tokens are encrypted with AES-256-GCM using a key held in Vercel env (src/lib/crypto.ts). Invitee phone numbers are plaintext.
API-01 All privileged operations execute server-side; the client is never trusted · P0 · score 4
Evidence: Prices, seat counts, plan and role are derived server-side; tampering with price and userId fields via curl is ignored or rejected.
Artifact: tests/api/tamper.spec.ts (curl-equivalent requests with forged price, plan, role, userId)
Enforcement: Runs in the required CI test job
API-02 Every input is validated against a schema at the trust boundary, at runtime · P1 · score 4
Evidence: Every route handler and server action parses input with a zod schema (.strict()); malformed and extra fields return 400.
Artifact: src/lib/schemas/; tests/api/validation.spec.ts
Enforcement: ESLint rule acme/require-zod-parse in the CI lint job flags handlers that read the body without a schema
API-03 Errors return safe messages; stack traces and internals never reach clients · P1 · score 3
Evidence: Forced 500 returns {error, requestId} with no stack trace; the requestId matches the Sentry event.
Artifact: tests/api/error-shape.spec.ts::500-has-no-stack
API-04 Every external call has a timeout, bounded retries with backoff and jitter, and a defined fallback · P1 · score 3
Evidence: All outbound calls go through src/lib/http.ts with a 10-second timeout and 3 retries with jitter; a black-hole endpoint times out cleanly and the booking page still renders.
Artifact: tests/resilience/blackhole.spec.ts
API-05 Each external provider is wrapped in a single service module · P1 · score 3
Evidence: Stripe, Resend, Google Calendar and the LLM provider each have one wrapper module in src/providers/; grep for SDK imports outside them is empty.
Artifact: `grep -rn "from 'stripe'\|from 'resend'\|googleapis\|@anthropic-ai/sdk" src --exclude-dir=providers` (empty)
API-06 Work that can exceed the request timeout runs in a background job, not in the request · P1 · score 3
Evidence: Calendar sync, reminder emails, invoice sync and LLM summaries run as Inngest functions; nothing over 5 s at p95 runs in a request.
Artifact: docs/ops/background-work.md (table of operations, p95 and where they run)
API-07 Background jobs, crons, and webhooks are observable and alert on failure · P1 · score 4
Evidence: Inngest failures after final retry page #ops-alerts; a deliberately failed reminder job produced an alert on 2026-09-29.
Artifact: docs/evidence/2026-09-29-failed-job-alert.png
Enforcement: Inngest failure alert routed to Slack #ops-alerts and PagerDuty
API-10 Scheduled jobs cannot overlap destructively · P1 · score 3
Evidence: Reminder and invoice-sync crons use Inngest concurrency keys of 1; a second concurrent start is skipped and logged.
Artifact: tests/jobs/no-overlap.spec.ts
API-11 Every batch run reconciles: expected, processed, failed, and skipped are recorded · P1 · score 3
Evidence: Each cron run writes a job_runs row with expected, processed, failed and skipped counts; the last 30 runs add up.
Artifact: docs/evidence/2026-09-30-job-runs.csv
API-08 An API contract exists (OpenAPI or typed RPC) and responses conform to it · P2 · score 2
Evidence: Public API has an OpenAPI 3.1 spec generated from the zod schemas (openapi.json). No conformance run against live responses yet.
API-09 Pagination, sorting, and filtering are bounded · P2 · score 3
Evidence: List endpoints clamp limit to 100 and use cursor pagination; limit=1000000 returns 100 items.
Artifact: tests/api/pagination.spec.ts::clamps-limit
FE-01 Every async surface has loading, error, empty, and success states · P1 · score 3
Evidence: Loading, error, empty and success states exist and are tested on the five core surfaces: booking page, availability editor, bookings list, team settings, billing.
Artifact: tests/e2e/states.spec.ts (20 cases)
FE-02 Error boundaries prevent one component failure from blanking the app · P1 · score 3
Evidence: Route-level error.tsx boundaries plus a component boundary around the slot picker; a forced throw in the picker leaves the page usable and reports to Sentry.
Artifact: tests/e2e/error-boundary.spec.ts
FE-03 Core flows verified on real mobile devices, Safari, and a throttled network · P1 · score 3
Evidence: Booking and onboarding flows verified on iPhone 15 Safari, Pixel 8 Chrome, desktop Safari and Firefox, and on a Slow 3G profile.
Artifact: docs/qa/2026-09-device-matrix.md
FE-04 Forms survive hostile and messy input · P1 · score 3
Evidence: Edge-case UX checklist executed on the booking form: emoji names, 500-character notes, pasted whitespace, double submit, back button after submit.
Artifact: docs/qa/2026-09-edge-case-checklist.md
FE-05 Automated accessibility scan passes on core pages · P1 · score 2
Evidence: axe run on 8 core pages: 2 violations remain. The public booking page time-slot buttons have no accessible name (each reads 'button'), and the secondary text colour on the booking page fails contrast at 3.1:1.
Notes: Open. Both violations are on the page every invitee sees.
FE-06 Keyboard-only and screen-reader passes completed on critical journeys · P1 · score 1
Evidence: Keyboard pass on the booking journey stopped at the slot picker: the calendar grid is not reachable by Tab and has no arrow-key support. No screen-reader pass has been done.
Notes: Open. An invitee who cannot use a mouse cannot book.
FE-07 Core Web Vitals measured on real users and within thresholds · P2 · score 3
Evidence: Vercel Speed Insights field data, last 28 days at p75: LCP 2.1 s, INP 140 ms, CLS 0.04 on the booking page.
Artifact: docs/evidence/2026-09-30-speed-insights.png
Attestation: Confirmed by Jordan Example (CTO, fictional), 2026-09-30: read the field data in the Vercel dashboard
FE-08 Five target users complete the core flow unaided · P2 · score 3
Evidence: Five target users (scheduling admins at small agencies) completed workspace setup and first booking-page publish unaided; median 6 minutes; two stumbled on the buffer-time setting, since relabelled.
Artifact: docs/research/2026-08-usability-sessions.md
Attestation: Confirmed by Jordan Example (CTO, fictional), 2026-08-21: ran and recorded the five sessions
APPSEC-01 Injection is prevented structurally, not by escaping strings · P0 · score 4
Evidence: All queries go through Drizzle's parameterised builder; the two sql`` template uses take bound parameters. No child_process usage. Semgrep injection rules clean.
Artifact: `semgrep --config p/typescript --config p/nextjs` output in docs/evidence/2026-09-29-semgrep.txt
Enforcement: CI sast job runs Semgrep on every pull request and blocks new high findings
APPSEC-03 Security headers are set, including a strict Content-Security-Policy · P1 · score 4
Evidence: CSP with nonces, HSTS preload, X-Content-Type-Options, Referrer-Policy and frame-ancestors set; ran report-only for 3 weeks with no legitimate violations before enforcing.
Artifact: `curl -I https://app.acmebook.example` output in docs/evidence/2026-09-29-headers.txt
Enforcement: CI job headers-check asserts the header set against the preview deployment
APPSEC-04 CORS is restricted to known origins and credentials are not exposed to wildcards · P1 · score 3
Evidence: CORS allows only the app origin on authenticated routes; the public embed API allows * without credentials. Request from an unlisted origin with credentials is blocked.
Artifact: tests/api/cors.spec.ts
APPSEC-05 CSRF protection is present on state-changing requests · P1 · score 3
Evidence: Server actions use Next.js origin checks and SameSite=Lax cookies; a cross-origin POST from a test page is rejected.
Artifact: tests/security/csrf.spec.ts
APPSEC-06 SAST runs on pull requests and blocks new high/critical findings · P1 · score 4
Evidence: Semgrep and CodeQL run on pull requests; baseline committed with zero accepted highs.
Artifact: .github/workflows/sast.yml; .semgrepignore
Enforcement: CodeQL and Semgrep are required status checks on main
APPSEC-07 SSRF is prevented anywhere the server fetches a user-supplied URL · P1 · score 3
Evidence: The only server-side fetch of a user-supplied URL is the webhook-subscription test ping; it resolves and blocks private, loopback and metadata ranges. 169.254.169.254 and 10.0.0.1 both rejected.
Artifact: tests/security/ssrf.spec.ts
APPSEC-08 A DAST baseline scan has been run against staging and findings triaged · P1 · score 3
Evidence: OWASP ZAP baseline against staging on 2026-09-22; 3 low findings triaged, 1 fixed.
Artifact: docs/security/2026-09-22-zap-baseline.html; docs/security/zap-triage.md
APPSEC-10 A written threat model exists for the highest-value assets · P2 · score 1
Evidence: Threat-model notes exist as a bullet list in a pull request description from 2025; no maintained threat model for the booking, billing and LLM paths.
APPSEC-11 A vulnerability disclosure path exists · P2 · score 3
Evidence: security.txt published with a monitored security@ alias.
Artifact: https://app.acmebook.example/.well-known/security.txt
SUP-01 Install scripts do not run for arbitrary transitive dependencies · P0 · score 4
Evidence: pnpm 10 does not run dependency install scripts by default; onlyBuiltDependencies allowlists esbuild and sharp. CI install log shows scripts skipped for everything else.
Artifact: package.json#pnpm.onlyBuiltDependencies; docs/evidence/2026-09-29-ci-install.log
Enforcement: CI install step uses pnpm 10 defaults; a pull request changing the allowlist requires CODEOWNERS review
SUP-02 Lockfile is committed and CI installs are frozen and reproducible · P1 · score 4
Evidence: pnpm-lock.yaml committed; CI runs pnpm install --frozen-lockfile.
Artifact: .github/workflows/ci.yml (install step)
Enforcement: CI fails when the lockfile is out of date
SUP-03 A dependency cooldown / minimum release age is configured · P1 · score 3
Evidence: minimumReleaseAge set to 4320 minutes (3 days) in pnpm-workspace.yaml.
Artifact: pnpm-workspace.yaml
SUP-04 Dependency vulnerability scanning runs in CI and on a schedule · P1 · score 4
Evidence: osv-scanner runs on pull requests and weekly; current result: 0 critical, 0 high, 2 moderate in dev-only dependencies, triaged.
Artifact: `osv-scanner --lockfile pnpm-lock.yaml` output in docs/evidence/2026-09-29-osv.txt; docs/security/vuln-triage.md
Enforcement: CI job deps-scan on pull requests plus a scheduled weekly workflow; Dependabot security updates on
SUP-05 Every dependency actually exists and was not hallucinated · P1 · score 3
Evidence: Pull request template requires a note for each new dependency; the last 20 additions all checked against the registry and predate their pull requests.
Artifact: .github/pull_request_template.md; docs/evidence/2026-09-30-new-deps-review.md
SUP-06 GitHub Actions are pinned to full commit SHAs · P2 · score 3
Evidence: All third-party actions pinned to full SHAs with version comments.
Artifact: .github/workflows/*.yml
SUP-07 CI workflow token permissions are least-privilege and untrusted input is never interpolated into shell · P2 · score 3
Evidence: Workflows set permissions: contents: read by default; no ${{ github.event.* }} interpolated into run steps (zizmor clean).
Artifact: `zizmor .github/workflows` output in docs/evidence/2026-09-29-zizmor.txt
SUP-08 An SBOM is generated per release and retained · P2 · score 0
Evidence: Looked for an SBOM step in the release workflow and the release assets: none is generated.
SUP-09 Published artifacts carry build provenance · P2 · score 0
Evidence: Releases are Vercel deployments and a worker container image; neither carries build provenance or attestations.
CI-01 Code is in version control with a clean, complete history · P0 · score 4
Evidence: Private GitHub repository acmebook; clean status; 1,912 commits with conventional messages.
Artifact: `git status` clean on main; history in the private acmebook repository
Enforcement: Branch protection on main; deploys only from main via the deploy workflow
CI-05 Deployments can be rolled back quickly, and rollback has been tested · P0 · score 3
Evidence: Rollback rehearsed on 2026-09-17: vercel rollback to the previous production deployment took 47 seconds; worker image rolled back by tag in 3 minutes.
Artifact: docs/evidence/2026-09-17-rollback-rehearsal.md (commands and timings)
Notes: Rehearsed once; not scheduled.
CI-02 CI runs lint, typecheck, tests, and build on every pull request · P1 · score 4
Evidence: CI runs lint, typecheck, unit, db, authz, e2e and build on every pull request; last 50 runs green except 2 fixed failures.
Artifact: .github/workflows/ci.yml
Enforcement: All seven jobs are required status checks on main
CI-03 The main branch is protected: no direct pushes, no force pushes, review required · P1 · score 4
Evidence: main ruleset: pull request required, 1 approval, no force push, no deletion, required checks.
Artifact: `gh api repos/{owner}/acmebook/rulesets` output in docs/evidence/2026-09-29-ruleset.json
Enforcement: GitHub ruleset on main enforces branch protection and required status checks
CI-04 CODEOWNERS protects auth, payments, migrations, and CI workflow paths · P1 · score 3
Evidence: CODEOWNERS covers src/auth*, src/providers/stripe*, db/migrations, db/policies and .github/workflows; ruleset requires code-owner review.
Artifact: .github/CODEOWNERS
CI-06 Changes are previewable before production · P1 · score 3
Evidence: Every pull request gets a Vercel preview with a seeded Neon branch.
Artifact: Pull request #912: preview at https://acmebook-pr-912.preview.acmebook.example
CI-07 Deploys are small, frequent, and attributable to a commit · P1 · score 3
Evidence: Release SHA shown at /api/health and attached to every Sentry event; median 6 production deploys per week.
Artifact: `curl https://app.acmebook.example/api/health` returns the commit SHA
CI-08 Risky changes ship behind feature flags with a kill switch · P2 · score 3
Evidence: Flags in PostHog; the LLM assistant and the new availability engine ship behind flags with a kill switch toggled in production on 2026-09-10.
Artifact: src/lib/flags.ts; docs/evidence/2026-09-10-kill-switch.png
CI-09 Post-deploy verification watches error rate, latency, and key funnels · P2 · score 2
Evidence: Engineers watch Sentry and the booking funnel after deploys; the checklist exists in docs/ops/deploy.md but is not recorded per deploy.
ENV-01 Development, staging, and production are genuinely separate deployments · P0 · score 4
Evidence: Separate Vercel projects, Neon projects, Stripe accounts (test vs live), Resend domains and Inngest environments for dev, staging and production; each uses its own credentials.
Artifact: docs/ops/environments.md#matrix
Enforcement: src/env.ts startup rule refuses to boot when the Stripe key mode does not match the environment
ENV-02 Required configuration is validated at startup and the app refuses to boot if it is wrong · P1 · score 4
Evidence: src/env.ts validates 23 variables with zod at boot; removing STRIPE_WEBHOOK_SECRET stops the app with a named error.
Artifact: tests/env/boot-fails.spec.ts
Enforcement: Runs in the required CI unit job
ENV-03 Infrastructure changes are reproducible, not clicked · P1 · score 3
Evidence: Vercel, Neon, DNS and IAM defined in Terraform; Inngest and Resend settings listed in a manual inventory.
Artifact: infra/*.tf; docs/ops/manual-inventory.md
ENV-04 TLS everywhere, HTTP redirects to HTTPS, and certificate renewal is automated · P1 · score 3
Evidence: SSL Labs A+; certificates managed and renewed by Vercel; HTTP redirects to HTTPS.
Artifact: https://www.ssllabs.com/ssltest/analyze.html?d=app.acmebook.example (fictional host); docs/evidence/2026-09-29-ssllabs.png
ENV-05 Serverless and platform timeouts are known and no core action exceeds them · P1 · score 3
Evidence: Function timeout 60 s; slowest core action is availability computation at p95 1.8 s.
Artifact: docs/ops/background-work.md#latency-vs-timeout
ENV-06 DNS, domains, and registrar access are documented and secured · P2 · score 3
Evidence: Domains at the registrar with registry lock and MFA; owner and backup owner documented.
Artifact: docs/ops/ownership-register.md#domains
OBS-01 Error tracking is installed on both frontend and backend and receives production errors · P0 · score 4
Evidence: Sentry on the Next.js client, server, edge and the worker; a deliberately triggered production error on 2026-09-29 appeared with a source-mapped stack trace and release tag.
Artifact: docs/evidence/2026-09-29-sentry-test-error.png
Enforcement: Sentry alert rule on new issues in production routes to #ops-alerts; source maps uploaded by the CI deploy job
OBS-02 Logs are structured and carry a correlation ID across the request path · P1 · score 3
Evidence: pino JSON logs carry requestId from middleware through Inngest job payloads; one booking traced from request to reminder email by a single ID.
Artifact: docs/evidence/2026-09-29-request-trace.log
OBS-03 Logs redact secrets and personal data and have a defined retention period · P1 · score 3
Evidence: Logs redact secrets and email addresses; Better Stack retention set to 30 days.
Artifact: tests/logging/redaction.spec.ts; docs/ops/logging.md#retention
OBS-04 Alerts exist for the failures that matter, and they reach a human · P1 · score 4
Evidence: Alerts for 5xx rate, booking failure rate, Inngest failures, Stripe webhook failures and DB CPU; a test alert on 2026-09-29 paged the on-call engineer.
Artifact: docs/ops/alerts.md; docs/evidence/2026-09-29-pagerduty-test.png
Enforcement: PagerDuty service with alert rules from Sentry, Better Stack and Inngest
OBS-05 Uptime and synthetic monitoring check the real user journey, not just a 200 · P1 · score 3
Evidence: Checkly runs a synthetic booking (load page, pick slot, book, cancel) every 5 minutes from 3 regions; it caught the 2026-08-14 Google Calendar token outage 11 minutes before the first support ticket.
Artifact: checkly/booking-journey.spec.ts; docs/incidents/2026-08-14.md
OBS-08 Data-quality metrics are monitored alongside system health · P1 · score 3
Evidence: Daily data-quality checks: bookings without a calendar event, orphaned invitees, invoices out of sync with Stripe. An alert fired on 2026-09-03 for 41 bookings missing calendar events.
Artifact: src/jobs/data-quality.ts; docs/incidents/2026-09-03.md
OBS-06 Distributed tracing covers the slow and complex paths · P2 · score 2
Evidence: Sentry performance tracing on server routes; the availability engine shows spans per calendar fetch. Worker traces are not linked to request traces.
OBS-07 Product analytics track the core funnel · P2 · score 3
Evidence: PostHog funnel: booking page view to confirmed booking, last 30 days.
Artifact: docs/evidence/2026-09-30-posthog-funnel.png
REL-01 RTO and RPO are defined, written down, and consistent with the backup configuration · P1 · score 1
Evidence: README says 'Neon handles backups'. There is no written RTO or RPO; the only measured restore time (41 minutes, May 2026) predates the database tripling in size.
Notes: Open. Write RTO/RPO and re-measure with a scheduled restore drill.
REL-02 An incident runbook exists covering the realistic failure set · P1 · score 1
Evidence: No incident runbook. The 2026-08-14 incident notes show the fix was found by reading code during the outage.
Notes: Open. Start from assets/templates/RUNBOOK.md: calendar-provider outage, Stripe webhook backlog, LLM provider outage, database restore.
REL-03 Someone is on call, or there is an explicit documented decision that nobody is · P1 · score 3
Evidence: Weekly PagerDuty rotation across three engineers; test page reached the on-call phone in 40 seconds.
Artifact: docs/ops/on-call.md; docs/evidence/2026-09-29-pagerduty-test.png
Attestation: Confirmed by Riley Example (head of operations, fictional), 2026-09-29: triggered the test page and confirmed the acknowledgement
REL-04 A breach and data-exposure response plan exists with notification timelines · P1 · score 3
Evidence: Breach response plan with named roles (incident lead, comms, counsel) and the 72-hour GDPR notification clock.
Artifact: docs/security/breach-response.md
REL-05 Incidents get blameless post-mortems and produce concrete follow-up actions · P2 · score 3
Evidence: Blameless post-mortems for the two incidents this year, each with closed follow-ups.
Artifact: docs/incidents/2026-08-14.md; docs/incidents/2026-09-03.md
REL-06 There is a way to tell customers what is happening · P2 · score 1
Evidence: Customers are told about incidents by individual email from support. No status page and no template.
REL-07 Third-party outage behavior is defined and degraded mode is tested · P2 · score 3
Evidence: Dependency table defines degraded behaviour: Google Calendar down shows cached availability with a banner; LLM down hides the assistant; Resend down queues mail. Each tested with a stubbed outage.
Artifact: docs/ops/dependencies.md; tests/resilience/provider-outage.spec.ts
PERF-01 Database connection pooling is configured and the connection math works · P1 · score 3
Evidence: Neon pooled connection string (PgBouncer, transaction mode); pool math documented: 60 concurrent functions x 1 connection, pooler max 10,000, Postgres max_connections 450.
Artifact: docs/ops/connection-math.md
PERF-02 N+1 queries and unindexed slow queries have been found and fixed · P1 · score 3
Evidence: N+1 on the bookings list fixed (41 queries to 3); heaviest endpoints measured before and after.
Artifact: docs/perf/2026-07-n-plus-one.md
PERF-03 Static assets are served from a CDN with correct cache headers · P1 · score 3
Evidence: Hashed assets served from Vercel's CDN with immutable cache headers; HTML no-store for authenticated pages.
Artifact: docs/evidence/2026-09-29-cache-headers.txt
PERF-04 Private data is never cached where another user can receive it · P1 · score 4
Evidence: Authenticated responses send Cache-Control: private, no-store; a two-user test confirms no cross-user cache hits on the dashboard.
Artifact: tests/e2e/cache-isolation.spec.ts
Enforcement: CI job headers-check asserts no-store on authenticated routes
PERF-05 A load test has been run at a realistic launch multiple and the breaking point is known · P2 · score 0
Evidence: No load test has been run. The busiest customer's Monday-morning peak is the only data point.
PERF-06 Expensive repeated work is cached with a defined invalidation strategy · P2 · score 2
Evidence: Availability results cached for 60 seconds per host in Upstash Redis, invalidated on booking and calendar webhook. Hit rate not measured.
PERF-07 A scaling plan exists for the next order of magnitude · P2 · score 2
Evidence: A scaling note names the availability engine and Google Calendar API quotas as the next bottlenecks; no plan for the third (Postgres write volume from reminders).
COST-01 No paid inference endpoint is publicly reachable without authentication and a rate limit · P0 · score 4
Evidence: The assistant endpoint /api/assistant requires a session and is limited to 30 requests per user per hour; unauthenticated calls return 401 and the 31st call returns 429.
Artifact: tests/api/assistant-auth-and-throttle.spec.ts
Enforcement: Upstash ratelimit middleware on the route; test runs in the required CI job
COST-02 Hard spend caps and billing alerts exist at every paid provider · P0 · score 3
Evidence: Provider-level monthly spend limit of $1,500 on the LLM account with alerts at 50% and 80%; Vercel spend management on; Stripe and Resend are usage-billed with billing alerts.
Artifact: docs/ops/spend-caps.md (screenshots per provider, 2026-09-30)
Notes: This is an account-wide cap. One tenant can still exhaust it for everyone; see AI-07.
COST-03 Rate limiting exists on all public and expensive endpoints · P1 · score 2
Evidence: Rate limits exist on auth, assistant and API-key routes. The public booking endpoint POST /api/public/[org]/book has no rate limit: 500 requests in a minute from one IP were all accepted and created 500 pending bookings and 500 confirmation emails in staging.
Notes: Open. The one public write path is the one without a limit.
COST-04 Cost per user and per action is measured, not estimated · P1 · score 3
Evidence: Cost per active workspace measured from provider invoices and usage logs for September: $0.38 infrastructure, $0.11 email, $0.92 LLM for workspaces using the assistant.
Artifact: docs/finance/2026-09-unit-costs.md
COST-05 Abuse paths have been tested: signup spam, scraping, resource exhaustion · P1 · score 3
Evidence: Signup spam (blocked by email verification and Turnstile), scraping of booking pages (cached and rate limited by Vercel firewall), and assistant exhaustion tested. Booking-endpoint flooding was not stopped (see COST-03).
Artifact: docs/security/2026-09-abuse-tests.md
COST-06 Billing events reconcile with usage and entitlements · P2 · score 3
Evidence: Every subscription lifecycle event (trial, upgrade, seat change, past due, cancel, reactivate) run in Stripe test mode with the resulting entitlement recorded.
Artifact: tests/billing/lifecycle.spec.ts; docs/qa/billing-matrix.md
COST-07 Infrastructure spend is forecast for the next 90 days at expected growth · P2 · score 2
Evidence: Forecast exists for infrastructure and email at 3x workspaces; LLM spend is not in it.
AI-01 Model provider calls happen server-side only, through a single gateway module · P0 · score 4
Evidence: All model calls go through src/providers/llm.ts on the server; grep for the provider SDK elsewhere is empty; the browser never contacts the provider.
Artifact: `grep -rn "@anthropic-ai/sdk" src --exclude=llm.ts` (empty); docs/evidence/2026-09-29-network-trace.har
Enforcement: ESLint no-restricted-imports rule in the CI lint job blocks the SDK outside src/providers/llm.ts
AI-05 Model output is treated as untrusted input everywhere it is used · P1 · score 3
Evidence: Assistant output is parsed into a zod schema of proposed slots; free text is rendered as plain text, never HTML; proposed times are re-checked against real availability before display.
Artifact: src/features/assistant/parse.ts; tests/assistant/output-validation.spec.ts
AI-06 Token usage and cost are logged per request with full attribution · P1 · score 3
Evidence: Every model call logs workspace, user, feature, model, input and output tokens and cost; aggregated by feature in a daily view.
Artifact: src/providers/llm.ts#logUsage; docs/evidence/2026-09-30-llm-usage-by-feature.csv
AI-07 Per-user and per-tier quotas are token-budgeted, not request-counted · P1 · score 1
Evidence: The only limit is 30 requests per user per hour. There is no token budget per workspace or plan, so one workspace with many seats and long prompts can spend the whole provider cap; in September the top workspace used 38% of LLM spend.
Notes: Open. Add per-workspace monthly token budgets by plan and a clear message at the limit.
AI-08 Provider failure degrades gracefully · P1 · score 3
Evidence: Provider timeout or 5xx hides the assistant panel and shows 'Suggestions are unavailable right now'; booking is unaffected.
Artifact: tests/resilience/provider-outage.spec.ts::llm-down
AI-09 An eval set exists for the AI feature and runs before releases · P1 · score 3
Evidence: 48-case eval set of scheduling requests with expected slot proposals; pass threshold 90%; last run 45/48 on 2026-09-28.
Artifact: evals/assistant/cases.jsonl; `pnpm eval:assistant` output in docs/evidence/2026-09-28-eval.txt
AI-10 Prompt injection has been tested against your actual attack surface · P1 · score 3
Evidence: 30 prompt-injection attempts through meeting notes and invitee names; 0 caused out-of-schema output or data from another workspace.
Artifact: tests/assistant/prompt-injection.spec.ts
AI-11 Cost is engineered: caching, routing, and batching where they apply · P2 · score 2
Evidence: Prompt caching on the system prompt; small model used for intent classification. No before/after cost measurement.
AI-12 AI behavior is traced and reviewed against real production traffic · P2 · score 2
Evidence: Langfuse traces for 10% of assistant calls; reviewed ad hoc. No failure taxonomy.
AI-13 AI features are disclosed to users where required and outputs are labelled · P2 · score 3
Evidence: The assistant panel is labelled 'AI suggestions' and proposals are marked as suggestions until confirmed.
Artifact: docs/evidence/2026-09-29-assistant-label.png
LEG-01 A privacy policy and terms of service are published and accurate · P0 · score 3
Evidence: Privacy policy and terms live at /privacy and /terms; each claim mapped to system behaviour, except the retention claim, which has nothing enforcing it (DATA-10).
Artifact: docs/legal/policy-claims-map.md; https://acmebook.example/privacy (fictional)
Attestation: Confirmed by Casey Example (outside counsel, fictional), 2026-09-24: reviewed the claim map against the published policy
LEG-02 A data inventory exists: what you collect, where it lives, who it goes to · P0 · score 3
Evidence: Data inventory lists every field collected, where it lives, and the 9 processors it goes to. The retention column is blank for every row.
Artifact: docs/legal/data-inventory.md
Attestation: Confirmed by Jordan Example (CTO, fictional), 2026-09-30: walked the inventory against the schema and the vendor list
LEG-03 Data subject requests are operable: access, export, correction, and deletion of data outside the account · P1 · score 3
Evidence: Access, export, correction and deletion executed against a test account; export took 2 minutes, deletion 4 minutes. Invitee data held by other workspaces is handled by request to support.
Artifact: docs/legal/2026-09-dsr-test.md
Attestation: Confirmed by Riley Example (head of operations, fictional), 2026-09-26: executed all four request types
LEG-04 GDPR basics are in place: lawful basis, processor agreements, transfer mechanism, cookie consent · P1 · score 3
Evidence: Lawful basis table, DPAs signed with all 9 processors, SCCs for US transfers, consent banner verified: no analytics before consent.
Artifact: docs/legal/gdpr.md; docs/evidence/2026-09-24-consent-network-trace.har
Attestation: Confirmed by Casey Example (outside counsel, fictional), 2026-09-24
LEG-05 US state privacy obligations are handled, including universal opt-out signals · P1 · score 3
Evidence: Applicability analysis for California and other state laws; Global Privacy Control detected and honoured (analytics off).
Artifact: docs/legal/us-state-privacy.md; tests/e2e/gpc.spec.ts
Attestation: Confirmed by Casey Example (outside counsel, fictional), 2026-09-24
LEG-06 Payment card scope is minimized and the applicable PCI obligations are known · P1 · score 3
Evidence: Stripe Checkout and Billing Portal only (SAQ A); no card data touches AcmeBook servers or logs.
Artifact: docs/legal/pci-scope.md
Attestation: Confirmed by Jordan Example (CTO, fictional), 2026-09-30: completed SAQ A
LEG-07 Subscription, refund, and cancellation terms are clear and cancellation is easy · P1 · score 3
Evidence: Pricing page states renewal and refund terms; cancellation is two clicks in the Stripe Billing Portal.
Artifact: docs/evidence/2026-09-29-cancel-flow.png
Attestation: Confirmed by Casey Example (outside counsel, fictional), 2026-09-24
LEG-09 Accessibility obligations have been assessed for your markets · P2 · score 2
Evidence: Assessed: EU customers bring the European Accessibility Act into scope for the booking flow. Current status is non-conformant on the slot picker (FE-05, FE-06).
Notes: Implemented as an assessment; conformance itself is open under FE-05 and FE-06.
LEG-11 Data residency is known and consistent with what customers have been told · P2 · score 3
Evidence: All stores and processors in us-east; privacy policy says data is stored in the United States.
Artifact: docs/legal/residency-map.md
Attestation: Confirmed by Casey Example (outside counsel, fictional), 2026-09-24
TEST-01 Automated tests cover authorization on every endpoint · P0 · score 4
Evidence: Authorization tests cover all 57 protected routes for unauthenticated, wrong-org and wrong-role callers.
Artifact: tests/authz/ (171 cases); coverage count printed by `pnpm test:authz --coverage-routes`
Enforcement: CI job test:authz is a required check and fails if a route has no authz test
TEST-02 The critical user journeys have end-to-end tests · P1 · score 4
Evidence: Playwright covers signup, connect calendar, publish booking page, book as invitee, reschedule, cancel, upgrade plan.
Artifact: tests/e2e/ (7 journeys); last run output in docs/evidence/2026-09-29-e2e.txt
Enforcement: CI job e2e runs against the preview deployment on every pull request
TEST-03 The full payment lifecycle is tested in the provider's test mode · P1 · score 3
Evidence: Payment test matrix in Stripe test mode: success, 3DS, decline, past due, dunning recovery, refund, cancel at period end.
Artifact: docs/qa/billing-matrix.md; tests/billing/lifecycle.spec.ts
TEST-04 Tests actually assert; coverage is measured on the diff, not chased globally · P1 · score 3
Evidence: Diff coverage at 84% on the last 20 pull requests; Stryker mutation score 71% on the availability engine.
Artifact: docs/evidence/2026-09-30-diff-coverage.txt; reports/mutation/availability.html
TEST-05 The API has been fuzzed against its schema · P1 · score 3
Evidence: Schemathesis run against the public API spec; 2 crashes found and fixed.
Artifact: `schemathesis run openapi.json` output in docs/evidence/2026-09-25-schemathesis.txt
TEST-06 Tests run against a real database, not mocks, for data-layer behavior · P1 · score 4
Evidence: Data-layer tests run against Postgres 16 in a container, in parallel with one schema per worker.
Artifact: tests/db/setup.ts
Enforcement: CI job test:db uses the postgres:16 service container
TEST-07 Flaky tests are quarantined and fixed, not retried into silence · P2 · score 3
Evidence: Flake list tracked with owners; 2 quarantined tests, trend down from 7 in July.
Artifact: docs/qa/flakes.md
TEST-08 A human has read the security-critical code, not just the tests · P2 · score 3
Evidence: Security-critical files (auth, permissions, RLS policies, webhooks) read line by line by a second engineer on 2026-09-18.
Artifact: docs/security/2026-09-18-security-code-review.md
DOC-01 README lets a new developer run the project locally from zero · P1 · score 4
Evidence: README setup run from zero in a clean devcontainer: running in 9 minutes.
Artifact: README.md#local-setup; .devcontainer/devcontainer.json
Enforcement: CI job devcontainer-smoke builds the devcontainer and runs the documented steps weekly on a schedule
DOC-02 Architecture notes exist: system diagram, data model, vendor map, and data flows · P1 · score 3
Evidence: ARCHITECTURE.md has the system diagram, data model, vendor map and data flows.
Artifact: docs/ARCHITECTURE.md
DOC-03 Deploy and rollback procedures are written down and executable by someone else · P1 · score 2
Evidence: docs/ops/deploy.md covers deploys. Rollback is a single line ('use vercel rollback') with nothing on the worker or a migration that must be reversed, and nobody but its author has followed it.
Notes: Open with REL-02: the runbooks are the main gap in this domain.
DOC-06 Bus factor: account ownership and recovery paths are recorded · P1 · score 3
Evidence: Ownership register for every account (registrar, Vercel, Neon, Stripe, Google Cloud, LLM provider, Resend) with two owners each; recovery codes in the company vault reachable by the CTO and head of operations.
Artifact: docs/ops/ownership-register.md
DOC-04 Significant decisions are recorded with their reversal path · P2 · score 3
Evidence: Twelve ADRs, each with a reversal path.
Artifact: docs/adr/
DOC-05 Known limitations, deliberate shortcuts, and technical debt are listed honestly · P2 · score 3
Evidence: Known limitations list maintained.
Artifact: docs/KNOWN-LIMITATIONS.md
PROD-01 The problem, the user, and the success metric are written down in one paragraph · P1 · score 3
Evidence: Problem, user and metric in one paragraph: small service teams lose bookings to back-and-forth email; success metric is weekly confirmed bookings per active workspace (currently 23).
Artifact: docs/PRODUCT.md
PROD-02 A baseline was captured before changes so improvement is provable · P1 · score 3
Evidence: Baseline captured 2026-04-01 before the availability-engine rewrite.
Artifact: docs/product/2026-04-01-baseline.md
PROD-03 Real user outcomes are measured after launch, not just system health · P2 · score 3
Evidence: Monthly outcome report against target: weekly bookings per workspace 23 vs target 20; booking-page conversion 31%.
Artifact: docs/product/2026-09-outcomes.md
Attestation: Confirmed by Jordan Example (CTO, fictional), 2026-10-01: read the report against PostHog
PROD-04 There is a way for users to report problems and it is monitored · P2 · score 3
Evidence: In-app 'Report a problem' posts to the support inbox, triaged daily by the head of operations.
Artifact: docs/ops/support.md
FUNC-01 The core computation has a golden fixture set: known-good, known-bad, and boundary cases · P1 · score 4
Evidence: Golden fixtures for the availability engine: 212 cases, including 5 from production defects (double-booked buffer, overnight shift, minimum notice across midnight, recurring exception, DST gap).
Artifact: tests/availability/golden/*.json; `pnpm test tests/availability` (212 passing)
Enforcement: Runs in the required CI unit job; fixture changes need CODEOWNERS review
FUNC-02 Edge and boundary cases of the domain are enumerated, not discovered in production · P1 · score 3
Evidence: Boundary list enumerated (zero-length windows, back-to-back bookings, buffers at window edges, max bookings per day); each has a test, plus a fast-check property suite with a fixed seed.
Artifact: docs/domain/availability-boundaries.md; tests/availability/properties.spec.ts
FUNC-03 Temporal correctness: time zones, DST, expiry, and date-range boundaries are tested · P1 · score 4
Evidence: DST forward and back in America/New_York and Europe/Berlin, a host and invitee in different zones across midnight, and every date-range endpoint are tested.
Artifact: tests/availability/timezones.spec.ts
Enforcement: Runs in the required CI unit job with TZ set to three different zones
FUNC-05 A legitimately empty result is distinguishable from a broken one · P1 · score 3
Evidence: When a calendar fetch fails, the booking page shows 'Availability is temporarily unavailable' instead of 'No times available', and an alert fires.
Artifact: tests/resilience/calendar-failure-not-empty.spec.ts
FUNC-06 Invariants that must always hold are asserted at runtime, not just assumed · P1 · score 3
Evidence: Invariants asserted at runtime: no overlapping confirmed bookings per host, every confirmed booking has a calendar event, seats in use never exceed seats paid. Nightly reconciliation job: zero violations in September.
Artifact: src/jobs/data-quality.ts; docs/evidence/2026-09-30-reconciliation.txt
FUNC-10 A second developer can safely CHANGE the core logic, not merely run it · P2 · score 3
Evidence: A new engineer changed the minimum-notice rule on 2026-08-05 using only the docs; one golden case caught an off-by-one, fixed before merge.
Artifact: docs/domain/2026-08-05-second-developer-exercise.md
EMAIL-01 Sending domain is authenticated with SPF, DKIM, and an enforcing DMARC policy · P1 · score 3
Evidence: SPF, DKIM (resend._domainkey) and DMARC p=reject on mail.acmebook.example; aggregate reports show 100% alignment.
Artifact: `dig TXT _dmarc.mail.acmebook.example` output in docs/evidence/2026-09-29-dns.txt
EMAIL-02 User input never reaches mail headers, and links in mail are not open redirects · P1 · score 3
Evidence: Invitee names are stripped of CR/LF before reaching templates; mail links go through a signed redirect that refuses external hosts.
Artifact: tests/email/header-injection.spec.ts; tests/email/open-redirect.spec.ts
EMAIL-03 Any user-triggerable send is rate limited per recipient and per actor · P1 · score 2
Evidence: Magic-link and invite emails are rate limited per recipient and per actor. Booking confirmations are not: each booking sends one, so the unthrottled booking endpoint (COST-03) can send unlimited confirmations to any address.
EMAIL-04 Bounces, complaints, and suppressions are processed and monitored · P1 · score 3
Evidence: Resend bounce and complaint webhooks update a suppression table; bounce rate 0.4% and complaint rate 0.01% on the dashboard.
Artifact: src/providers/email/webhooks.ts; docs/evidence/2026-09-30-resend-dashboard.png
EMAIL-05 Marketing mail is separated from transactional, carries an honest unsubscribe, and honours it immediately · P1 · score 3
Evidence: Product newsletter sent from a separate subdomain and stream with one-click unsubscribe; unsubscribing suppresses marketing only.
Artifact: tests/email/unsubscribe.spec.ts