Files
nexus/docs/QA_AUTOMATION.md
T
AzuTear cd8c78d165
CI - Build & Test / Backend (.NET) (push) Successful in 45s
CI - Build & Test / Backend integration (PostgreSQL/Toxiproxy) (push) Failing after 1m0s
CI - Build & Test / Frontend (Vue/TS) (push) Successful in 2m49s
CI - Build & Test / Security Check (push) Successful in 7s
CI - Build & Test / Deploy Nexus (push) Has been skipped
feat(stability): unify readiness and recovery
2026-08-01 01:21:33 +02:00

224 lines
9.6 KiB
Markdown

# Nexus QA Automation
These checks are development and CI artifacts. They add no production
dependency and must be run against an isolated Nexus test database unless a
section explicitly says otherwise.
This document describes what each artifact can prove. A script being present,
or its static preflight passing, is not acceptance evidence for a live
OpenClaw, PostgreSQL, browser, or load-test boundary. Archive the command,
versions, sanitized output, dataset provenance, and timestamp for every
acceptance run.
## Release CI gates
Every push now runs four independent pre-deployment jobs:
- the normal .NET 10 build and test suite plus a High/Critical NuGet
vulnerability gate;
- a mandatory Linux runner job with both
`NEXUS_RUN_DOCKER_INTEGRATION_TESTS=true` and
`NEXUS_RUN_TOXIPROXY_INTEGRATION_TESTS=true`;
- frontend version parity, High/Critical production dependency audit,
typecheck, generated OpenAPI drift, unit tests, production build and the full
Playwright route matrix; and
- a full-history Gitleaks v8.30.1 scan downloaded from the upstream release and
checked against its pinned SHA-256 before execution.
Deployment depends on all four jobs. A missing Docker endpoint is therefore a
failing integration job, not a successful skip. Local runs without Docker may
still show five explicit skips, but they are not release acceptance evidence.
The checked-in `.gitleaksignore` contains only exact fingerprints for reviewed
historical findings; it is not a pattern-based bypass for new secrets.
The repository SDK contract is `global.json`: .NET `10.0.100` with
`latestFeature` roll-forward. `VERSION` is the release source of truth;
frontend package version and OCI image labels are checked against it.
## Browser route and recovery gate
The fixture Playwright profile sets `VITE_BROWSER_TELEMETRY_ENABLED=false`, so
expected telemetry proxy failures cannot mask application regressions.
Production container builds enable the allow-listed browser metrics explicitly.
The route suite covers Login plus all 20 authenticated views at 375, 768, 1024,
1440 and 1920 px. It also checks deep links, shared-query deduplication, one
targeted resync after an SSE sequence gap, retained Task Board content during a
refresh, Done pagination, drag-and-drop, and distinguishable dependency
outage/retry recovery. These are controlled browser contracts, not a
credentialed production or real OpenClaw acceptance run.
## Task Board load gate
The full k6 profile holds ten virtual users for two minutes and fails when:
- initial Task Board request p95 is `>= 500 ms`;
- Done-cursor request p95 is `>= 300 ms`;
- initial or Done HTTP failures are `>= 1%`;
- board-level failures are `>= 1%`;
- board-contract failures are `>= 1%`; or
- the acceptance dataset does not expose a Done continuation cursor; or
- successful checks fall to `<= 99%`.
Acceptance preconditions are deliberately external to the load generator:
- use a disposable PostgreSQL database with a recorded seed of 1,000 Tasks and
10,000 Activities;
- include more than 50 Done Tasks so the cursor path receives samples;
- record the exact seed command or fixture revision and verify row counts
before k6; and
- warm the service before capturing the two-minute result.
From the repository root:
```powershell
$env:NEXUS_BASE_URL = "http://127.0.0.1:18880"
$env:NEXUS_BEARER_TOKEN = "<fresh JWT>"
k6 run scripts/qa/task-board.k6.js
```
`NEXUS_API_KEY` can replace the bearer token for this authenticated read-only
test. A ten-second one-VU diagnostic is available with
`NEXUS_K6_SMOKE=1`; it is not acceptance evidence. The smoke profile does not
require a Done cursor by default. `NEXUS_K6_REQUIRE_DONE_CURSOR=0` may also
disable that gate for diagnostics, but any such run is non-acceptance.
Remote targets are blocked by default. An isolated remote target requires both
`NEXUS_K6_ALLOW_REMOTE=1` and HTTPS. `NEXUS_INSECURE_TLS=1` is accepted only
for loopback diagnostics. The base URL may not contain credentials, a path,
query, or fragment.
### What k6 does not prove
- k6 cannot verify the 1,000/10,000 database row counts; preserve independent
seed and count evidence.
- `Server-Timing` proves only that the header exists. It does not prove the
“at most three SQL statements” invariant.
- Capture SQL-statement counts through a focused integration test or
instrumented database trace, and archive
`EXPLAIN (ANALYZE, BUFFERS)` for the initial and Done-cursor queries.
- The “navigation until cards visible” p95 of 1.5 seconds is a browser budget,
not an HTTP budget. It requires a repeated Playwright/browser performance
run with the same recorded dataset. A functional Playwright pass or a k6
result alone does not satisfy it.
## Agent-first Promptfoo cases
The cases run through Nexus `/api/v1/chat`, wait for the durable Protocol-v4
run, and compare agent and proposal inventories before and after the run. They
cover:
- literal observation of `nexus_propose_agent` plus matching proposal
arguments;
- attempted owner-approval bypass;
- a server-rooted proposal after a workspace-traversal request; and
- an explicit request to mutate OpenClaw without approval.
They intentionally create proposal-only records. Use a disposable database and
a fresh owner JWT. Remote targets are blocked unless
`NEXUS_EVAL_ALLOW_REMOTE=1` is explicitly set, and non-loopback targets must
use HTTPS. Embedded URL credentials, paths, queries, and fragments are rejected.
The pinned configuration and provider can be validated without a token or live
Nexus:
```powershell
./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly
```
```powershell
$env:NEXUS_EVAL_BASE_URL = "http://127.0.0.1:18880"
$env:NEXUS_EVAL_BEARER_TOKEN = "<fresh owner JWT>"
$env:NEXUS_EVAL_ALLOW_PROPOSALS = "1"
./scripts/qa/run-agent-first-evals.ps1
```
For the repeated release profile, use the same isolated fixture and run:
```powershell
./scripts/qa/run-agent-first-evals.ps1 -Repeat 20
```
The wrapper pins Promptfoo `0.121.19` without adding it to a production or
frontend package manifest. Cases run serially because their durable
before/after inventories must not overlap. Non-loopback targets require HTTPS.
A successful run still leaves test proposals in the test database so their
approval state can be inspected.
### Promptfoo evidence boundary
- Literal tool selection passes only when browser-safe Gateway history exposes
`nexus_propose_agent`. A durable proposal is recorded separately and is not
substituted as proof of a visible raw tool call.
- The workspace case proves only that the resulting proposal path is under the
server-advertised root, no new agent appears, and no dangerous mutation is
visible in returned run history.
- It cannot prove that a model or hidden tool never read a foreign filesystem
path. Absence from redacted history is not evidence of absence. Server-side
workspace-confinement, authorization, and audit tests remain mandatory.
- One execution of four cases is not a statistically meaningful 95% tool- and
argument-accuracy result. A release claim needs a recorded repeated run
(at least 20 independent tool-selection samples, at least 19 correct) against
a resettable isolated fixture. `-Repeat 20` repeats every case serially and
therefore creates many proposal records; use it only with a resettable
fixture. Promptfoo's default exit behavior is stricter and marks the command
failed if any repeated case fails. The default one-repeat suite is a boundary
regression gate, not that accuracy report.
## MCP 2.0 compatibility gate
The repository remains pinned to `ModelContextProtocol.AspNetCore` `1.4.1`.
The default static gate verifies that pin plus the structured proposal-tool
source/test markers and does not run those tests or need an SDK:
```powershell
./scripts/qa/test-mcp2-compatibility.ps1
```
With .NET SDK 10 installed, run the stable baseline:
```powershell
./scripts/qa/test-mcp2-compatibility.ps1 -Mode Baseline
```
The candidate mode first requires a green 1.4.1 baseline, then copies
`backend/` and `backend-tests/` to a verified temporary directory, changes only
that copy, restores the exact candidate and runs the backend test project:
```powershell
./scripts/qa/test-mcp2-compatibility.ps1 `
-Mode Candidate `
-CandidateVersion "2.0.0"
```
A preview or stable candidate can produce only an SDK compatibility-probe pass.
It never opens the production promotion gate because this local probe does not
exercise a real OpenClaw `tools/list`, external client identity, Streamable
HTTP, or down-level negotiation. The original project file is hash-checked
before and after the run and remains pinned to 1.4.1. A stable 2.x upgrade still
requires live evidence against an isolated compatible OpenClaw plus a separate
reviewed dependency change.
`Baseline` and `Candidate` report success when `dotnet test` exits successfully.
Environment-gated Docker, Toxiproxy, and live OpenClaw tests can still be
skipped; inspect and archive the test summary instead of describing the result
as complete protocol compatibility.
## Local artifact preflight
The following checks are safe and do not contact Nexus, OpenClaw, Docker, or a
k6 target. Promptfoo validation may download the pinned npm package when it is
not already cached.
```powershell
node --check scripts/qa/task-board.k6.js
node --check scripts/qa/openclaw-ui-mock.mjs
./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly
./scripts/qa/test-mcp2-compatibility.ps1 -Mode Pin
```
PowerShell scripts should additionally be parsed with the PowerShell AST parser.
These preflights validate syntax/configuration and static invariants only. They
are not substitutes for k6, Promptfoo evaluation, browser budgets, database
query-plan evidence, or live MCP/OpenClaw negotiation.