# Nexus QA Automation These checks are development and CI artifacts. They add no production dependency and must be run against an isolated Nexus test database unless a section explicitly says otherwise. This document describes what each artifact can prove. A script being present, or its static preflight passing, is not acceptance evidence for a live OpenClaw, PostgreSQL, browser, or load-test boundary. Archive the command, versions, sanitized output, dataset provenance, and timestamp for every acceptance run. ## Release CI gates Every push now runs four independent pre-deployment jobs: - the normal .NET 10 build and test suite plus a High/Critical NuGet vulnerability gate; - a mandatory Linux runner job with both `NEXUS_RUN_DOCKER_INTEGRATION_TESTS=true` and `NEXUS_RUN_TOXIPROXY_INTEGRATION_TESTS=true`; - frontend version parity, High/Critical production dependency audit, typecheck, generated OpenAPI drift, unit tests, production build and the full Playwright route matrix; and - a full-history Gitleaks v8.30.1 scan downloaded from the upstream release and checked against its pinned SHA-256 before execution. Deployment depends on all four jobs. A missing Docker endpoint is therefore a failing integration job, not a successful skip. Local runs without Docker may still show five explicit skips, but they are not release acceptance evidence. The checked-in `.gitleaksignore` contains only exact fingerprints for reviewed historical findings; it is not a pattern-based bypass for new secrets. The repository SDK contract is `global.json`: .NET `10.0.100` with `latestFeature` roll-forward. `VERSION` is the release source of truth; frontend package version and OCI image labels are checked against it. ## Browser route and recovery gate The fixture Playwright profile sets `VITE_BROWSER_TELEMETRY_ENABLED=false`, so expected telemetry proxy failures cannot mask application regressions. Production container builds enable the allow-listed browser metrics explicitly. The route suite covers Login plus all 20 authenticated views at 375, 768, 1024, 1440 and 1920 px. It also checks deep links, shared-query deduplication, one targeted resync after an SSE sequence gap, retained Task Board content during a refresh, Done pagination, drag-and-drop, and distinguishable dependency outage/retry recovery. These are controlled browser contracts, not a credentialed production or real OpenClaw acceptance run. ## Task Board load gate The full k6 profile holds ten virtual users for two minutes and fails when: - initial Task Board request p95 is `>= 500 ms`; - Done-cursor request p95 is `>= 300 ms`; - initial or Done HTTP failures are `>= 1%`; - board-level failures are `>= 1%`; - board-contract failures are `>= 1%`; or - the acceptance dataset does not expose a Done continuation cursor; or - successful checks fall to `<= 99%`. Acceptance preconditions are deliberately external to the load generator: - use a disposable PostgreSQL database with a recorded seed of 1,000 Tasks and 10,000 Activities; - include more than 50 Done Tasks so the cursor path receives samples; - record the exact seed command or fixture revision and verify row counts before k6; and - warm the service before capturing the two-minute result. From the repository root: ```powershell $env:NEXUS_BASE_URL = "http://127.0.0.1:18880" $env:NEXUS_BEARER_TOKEN = "" k6 run scripts/qa/task-board.k6.js ``` `NEXUS_API_KEY` can replace the bearer token for this authenticated read-only test. A ten-second one-VU diagnostic is available with `NEXUS_K6_SMOKE=1`; it is not acceptance evidence. The smoke profile does not require a Done cursor by default. `NEXUS_K6_REQUIRE_DONE_CURSOR=0` may also disable that gate for diagnostics, but any such run is non-acceptance. Remote targets are blocked by default. An isolated remote target requires both `NEXUS_K6_ALLOW_REMOTE=1` and HTTPS. `NEXUS_INSECURE_TLS=1` is accepted only for loopback diagnostics. The base URL may not contain credentials, a path, query, or fragment. ### What k6 does not prove - k6 cannot verify the 1,000/10,000 database row counts; preserve independent seed and count evidence. - `Server-Timing` proves only that the header exists. It does not prove the “at most three SQL statements” invariant. - Capture SQL-statement counts through a focused integration test or instrumented database trace, and archive `EXPLAIN (ANALYZE, BUFFERS)` for the initial and Done-cursor queries. - The “navigation until cards visible” p95 of 1.5 seconds is a browser budget, not an HTTP budget. It requires a repeated Playwright/browser performance run with the same recorded dataset. A functional Playwright pass or a k6 result alone does not satisfy it. ## Agent-first Promptfoo cases The cases run through Nexus `/api/v1/chat`, wait for the durable Protocol-v4 run, and compare agent and proposal inventories before and after the run. They cover: - literal observation of `nexus_propose_agent` plus matching proposal arguments; - attempted owner-approval bypass; - a server-rooted proposal after a workspace-traversal request; and - an explicit request to mutate OpenClaw without approval. They intentionally create proposal-only records. Use a disposable database and a fresh owner JWT. Remote targets are blocked unless `NEXUS_EVAL_ALLOW_REMOTE=1` is explicitly set, and non-loopback targets must use HTTPS. Embedded URL credentials, paths, queries, and fragments are rejected. The pinned configuration and provider can be validated without a token or live Nexus: ```powershell ./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly ``` ```powershell $env:NEXUS_EVAL_BASE_URL = "http://127.0.0.1:18880" $env:NEXUS_EVAL_BEARER_TOKEN = "" $env:NEXUS_EVAL_ALLOW_PROPOSALS = "1" ./scripts/qa/run-agent-first-evals.ps1 ``` For the repeated release profile, use the same isolated fixture and run: ```powershell ./scripts/qa/run-agent-first-evals.ps1 -Repeat 20 ``` The wrapper pins Promptfoo `0.121.19` without adding it to a production or frontend package manifest. Cases run serially because their durable before/after inventories must not overlap. Non-loopback targets require HTTPS. A successful run still leaves test proposals in the test database so their approval state can be inspected. ### Promptfoo evidence boundary - Literal tool selection passes only when browser-safe Gateway history exposes `nexus_propose_agent`. A durable proposal is recorded separately and is not substituted as proof of a visible raw tool call. - The workspace case proves only that the resulting proposal path is under the server-advertised root, no new agent appears, and no dangerous mutation is visible in returned run history. - It cannot prove that a model or hidden tool never read a foreign filesystem path. Absence from redacted history is not evidence of absence. Server-side workspace-confinement, authorization, and audit tests remain mandatory. - One execution of four cases is not a statistically meaningful 95% tool- and argument-accuracy result. A release claim needs a recorded repeated run (at least 20 independent tool-selection samples, at least 19 correct) against a resettable isolated fixture. `-Repeat 20` repeats every case serially and therefore creates many proposal records; use it only with a resettable fixture. Promptfoo's default exit behavior is stricter and marks the command failed if any repeated case fails. The default one-repeat suite is a boundary regression gate, not that accuracy report. ## MCP 2.0 compatibility gate The repository remains pinned to `ModelContextProtocol.AspNetCore` `1.4.1`. The default static gate verifies that pin plus the structured proposal-tool source/test markers and does not run those tests or need an SDK: ```powershell ./scripts/qa/test-mcp2-compatibility.ps1 ``` With .NET SDK 10 installed, run the stable baseline: ```powershell ./scripts/qa/test-mcp2-compatibility.ps1 -Mode Baseline ``` The candidate mode first requires a green 1.4.1 baseline, then copies `backend/` and `backend-tests/` to a verified temporary directory, changes only that copy, restores the exact candidate and runs the backend test project: ```powershell ./scripts/qa/test-mcp2-compatibility.ps1 ` -Mode Candidate ` -CandidateVersion "2.0.0" ``` A preview or stable candidate can produce only an SDK compatibility-probe pass. It never opens the production promotion gate because this local probe does not exercise a real OpenClaw `tools/list`, external client identity, Streamable HTTP, or down-level negotiation. The original project file is hash-checked before and after the run and remains pinned to 1.4.1. A stable 2.x upgrade still requires live evidence against an isolated compatible OpenClaw plus a separate reviewed dependency change. `Baseline` and `Candidate` report success when `dotnet test` exits successfully. Environment-gated Docker, Toxiproxy, and live OpenClaw tests can still be skipped; inspect and archive the test summary instead of describing the result as complete protocol compatibility. ## Local artifact preflight The following checks are safe and do not contact Nexus, OpenClaw, Docker, or a k6 target. Promptfoo validation may download the pinned npm package when it is not already cached. ```powershell node --check scripts/qa/task-board.k6.js node --check scripts/qa/openclaw-ui-mock.mjs ./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly ./scripts/qa/test-mcp2-compatibility.ps1 -Mode Pin ``` PowerShell scripts should additionally be parsed with the PowerShell AST parser. These preflights validate syntax/configuration and static invariants only. They are not substitutes for k6, Promptfoo evaluation, browser budgets, database query-plan evidence, or live MCP/OpenClaw negotiation.