Files
nexus/docs/QA_AUTOMATION.md
T
AzuTear cd8c78d165
CI - Build & Test / Backend (.NET) (push) Successful in 45s
CI - Build & Test / Backend integration (PostgreSQL/Toxiproxy) (push) Failing after 1m0s
CI - Build & Test / Frontend (Vue/TS) (push) Successful in 2m49s
CI - Build & Test / Security Check (push) Successful in 7s
CI - Build & Test / Deploy Nexus (push) Has been skipped
feat(stability): unify readiness and recovery
2026-08-01 01:21:33 +02:00

9.6 KiB

Nexus QA Automation

These checks are development and CI artifacts. They add no production dependency and must be run against an isolated Nexus test database unless a section explicitly says otherwise.

This document describes what each artifact can prove. A script being present, or its static preflight passing, is not acceptance evidence for a live OpenClaw, PostgreSQL, browser, or load-test boundary. Archive the command, versions, sanitized output, dataset provenance, and timestamp for every acceptance run.

Release CI gates

Every push now runs four independent pre-deployment jobs:

  • the normal .NET 10 build and test suite plus a High/Critical NuGet vulnerability gate;
  • a mandatory Linux runner job with both NEXUS_RUN_DOCKER_INTEGRATION_TESTS=true and NEXUS_RUN_TOXIPROXY_INTEGRATION_TESTS=true;
  • frontend version parity, High/Critical production dependency audit, typecheck, generated OpenAPI drift, unit tests, production build and the full Playwright route matrix; and
  • a full-history Gitleaks v8.30.1 scan downloaded from the upstream release and checked against its pinned SHA-256 before execution.

Deployment depends on all four jobs. A missing Docker endpoint is therefore a failing integration job, not a successful skip. Local runs without Docker may still show five explicit skips, but they are not release acceptance evidence. The checked-in .gitleaksignore contains only exact fingerprints for reviewed historical findings; it is not a pattern-based bypass for new secrets.

The repository SDK contract is global.json: .NET 10.0.100 with latestFeature roll-forward. VERSION is the release source of truth; frontend package version and OCI image labels are checked against it.

Browser route and recovery gate

The fixture Playwright profile sets VITE_BROWSER_TELEMETRY_ENABLED=false, so expected telemetry proxy failures cannot mask application regressions. Production container builds enable the allow-listed browser metrics explicitly.

The route suite covers Login plus all 20 authenticated views at 375, 768, 1024, 1440 and 1920 px. It also checks deep links, shared-query deduplication, one targeted resync after an SSE sequence gap, retained Task Board content during a refresh, Done pagination, drag-and-drop, and distinguishable dependency outage/retry recovery. These are controlled browser contracts, not a credentialed production or real OpenClaw acceptance run.

Task Board load gate

The full k6 profile holds ten virtual users for two minutes and fails when:

  • initial Task Board request p95 is >= 500 ms;
  • Done-cursor request p95 is >= 300 ms;
  • initial or Done HTTP failures are >= 1%;
  • board-level failures are >= 1%;
  • board-contract failures are >= 1%; or
  • the acceptance dataset does not expose a Done continuation cursor; or
  • successful checks fall to <= 99%.

Acceptance preconditions are deliberately external to the load generator:

  • use a disposable PostgreSQL database with a recorded seed of 1,000 Tasks and 10,000 Activities;
  • include more than 50 Done Tasks so the cursor path receives samples;
  • record the exact seed command or fixture revision and verify row counts before k6; and
  • warm the service before capturing the two-minute result.

From the repository root:

$env:NEXUS_BASE_URL = "http://127.0.0.1:18880"
$env:NEXUS_BEARER_TOKEN = "<fresh JWT>"
k6 run scripts/qa/task-board.k6.js

NEXUS_API_KEY can replace the bearer token for this authenticated read-only test. A ten-second one-VU diagnostic is available with NEXUS_K6_SMOKE=1; it is not acceptance evidence. The smoke profile does not require a Done cursor by default. NEXUS_K6_REQUIRE_DONE_CURSOR=0 may also disable that gate for diagnostics, but any such run is non-acceptance.

Remote targets are blocked by default. An isolated remote target requires both NEXUS_K6_ALLOW_REMOTE=1 and HTTPS. NEXUS_INSECURE_TLS=1 is accepted only for loopback diagnostics. The base URL may not contain credentials, a path, query, or fragment.

What k6 does not prove

  • k6 cannot verify the 1,000/10,000 database row counts; preserve independent seed and count evidence.
  • Server-Timing proves only that the header exists. It does not prove the “at most three SQL statements” invariant.
  • Capture SQL-statement counts through a focused integration test or instrumented database trace, and archive EXPLAIN (ANALYZE, BUFFERS) for the initial and Done-cursor queries.
  • The “navigation until cards visible” p95 of 1.5 seconds is a browser budget, not an HTTP budget. It requires a repeated Playwright/browser performance run with the same recorded dataset. A functional Playwright pass or a k6 result alone does not satisfy it.

Agent-first Promptfoo cases

The cases run through Nexus /api/v1/chat, wait for the durable Protocol-v4 run, and compare agent and proposal inventories before and after the run. They cover:

  • literal observation of nexus_propose_agent plus matching proposal arguments;
  • attempted owner-approval bypass;
  • a server-rooted proposal after a workspace-traversal request; and
  • an explicit request to mutate OpenClaw without approval.

They intentionally create proposal-only records. Use a disposable database and a fresh owner JWT. Remote targets are blocked unless NEXUS_EVAL_ALLOW_REMOTE=1 is explicitly set, and non-loopback targets must use HTTPS. Embedded URL credentials, paths, queries, and fragments are rejected.

The pinned configuration and provider can be validated without a token or live Nexus:

./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly
$env:NEXUS_EVAL_BASE_URL = "http://127.0.0.1:18880"
$env:NEXUS_EVAL_BEARER_TOKEN = "<fresh owner JWT>"
$env:NEXUS_EVAL_ALLOW_PROPOSALS = "1"
./scripts/qa/run-agent-first-evals.ps1

For the repeated release profile, use the same isolated fixture and run:

./scripts/qa/run-agent-first-evals.ps1 -Repeat 20

The wrapper pins Promptfoo 0.121.19 without adding it to a production or frontend package manifest. Cases run serially because their durable before/after inventories must not overlap. Non-loopback targets require HTTPS. A successful run still leaves test proposals in the test database so their approval state can be inspected.

Promptfoo evidence boundary

  • Literal tool selection passes only when browser-safe Gateway history exposes nexus_propose_agent. A durable proposal is recorded separately and is not substituted as proof of a visible raw tool call.
  • The workspace case proves only that the resulting proposal path is under the server-advertised root, no new agent appears, and no dangerous mutation is visible in returned run history.
  • It cannot prove that a model or hidden tool never read a foreign filesystem path. Absence from redacted history is not evidence of absence. Server-side workspace-confinement, authorization, and audit tests remain mandatory.
  • One execution of four cases is not a statistically meaningful 95% tool- and argument-accuracy result. A release claim needs a recorded repeated run (at least 20 independent tool-selection samples, at least 19 correct) against a resettable isolated fixture. -Repeat 20 repeats every case serially and therefore creates many proposal records; use it only with a resettable fixture. Promptfoo's default exit behavior is stricter and marks the command failed if any repeated case fails. The default one-repeat suite is a boundary regression gate, not that accuracy report.

MCP 2.0 compatibility gate

The repository remains pinned to ModelContextProtocol.AspNetCore 1.4.1. The default static gate verifies that pin plus the structured proposal-tool source/test markers and does not run those tests or need an SDK:

./scripts/qa/test-mcp2-compatibility.ps1

With .NET SDK 10 installed, run the stable baseline:

./scripts/qa/test-mcp2-compatibility.ps1 -Mode Baseline

The candidate mode first requires a green 1.4.1 baseline, then copies backend/ and backend-tests/ to a verified temporary directory, changes only that copy, restores the exact candidate and runs the backend test project:

./scripts/qa/test-mcp2-compatibility.ps1 `
  -Mode Candidate `
  -CandidateVersion "2.0.0"

A preview or stable candidate can produce only an SDK compatibility-probe pass. It never opens the production promotion gate because this local probe does not exercise a real OpenClaw tools/list, external client identity, Streamable HTTP, or down-level negotiation. The original project file is hash-checked before and after the run and remains pinned to 1.4.1. A stable 2.x upgrade still requires live evidence against an isolated compatible OpenClaw plus a separate reviewed dependency change.

Baseline and Candidate report success when dotnet test exits successfully. Environment-gated Docker, Toxiproxy, and live OpenClaw tests can still be skipped; inspect and archive the test summary instead of describing the result as complete protocol compatibility.

Local artifact preflight

The following checks are safe and do not contact Nexus, OpenClaw, Docker, or a k6 target. Promptfoo validation may download the pinned npm package when it is not already cached.

node --check scripts/qa/task-board.k6.js
node --check scripts/qa/openclaw-ui-mock.mjs
./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly
./scripts/qa/test-mcp2-compatibility.ps1 -Mode Pin

PowerShell scripts should additionally be parsed with the PowerShell AST parser. These preflights validate syntax/configuration and static invariants only. They are not substitutes for k6, Promptfoo evaluation, browser budgets, database query-plan evidence, or live MCP/OpenClaw negotiation.