cd8c78d165
CI - Build & Test / Backend (.NET) (push) Successful in 45s
CI - Build & Test / Backend integration (PostgreSQL/Toxiproxy) (push) Failing after 1m0s
CI - Build & Test / Frontend (Vue/TS) (push) Successful in 2m49s
CI - Build & Test / Security Check (push) Successful in 7s
CI - Build & Test / Deploy Nexus (push) Has been skipped
224 lines
9.6 KiB
Markdown
224 lines
9.6 KiB
Markdown
# Nexus QA Automation
|
|
|
|
These checks are development and CI artifacts. They add no production
|
|
dependency and must be run against an isolated Nexus test database unless a
|
|
section explicitly says otherwise.
|
|
|
|
This document describes what each artifact can prove. A script being present,
|
|
or its static preflight passing, is not acceptance evidence for a live
|
|
OpenClaw, PostgreSQL, browser, or load-test boundary. Archive the command,
|
|
versions, sanitized output, dataset provenance, and timestamp for every
|
|
acceptance run.
|
|
|
|
## Release CI gates
|
|
|
|
Every push now runs four independent pre-deployment jobs:
|
|
|
|
- the normal .NET 10 build and test suite plus a High/Critical NuGet
|
|
vulnerability gate;
|
|
- a mandatory Linux runner job with both
|
|
`NEXUS_RUN_DOCKER_INTEGRATION_TESTS=true` and
|
|
`NEXUS_RUN_TOXIPROXY_INTEGRATION_TESTS=true`;
|
|
- frontend version parity, High/Critical production dependency audit,
|
|
typecheck, generated OpenAPI drift, unit tests, production build and the full
|
|
Playwright route matrix; and
|
|
- a full-history Gitleaks v8.30.1 scan downloaded from the upstream release and
|
|
checked against its pinned SHA-256 before execution.
|
|
|
|
Deployment depends on all four jobs. A missing Docker endpoint is therefore a
|
|
failing integration job, not a successful skip. Local runs without Docker may
|
|
still show five explicit skips, but they are not release acceptance evidence.
|
|
The checked-in `.gitleaksignore` contains only exact fingerprints for reviewed
|
|
historical findings; it is not a pattern-based bypass for new secrets.
|
|
|
|
The repository SDK contract is `global.json`: .NET `10.0.100` with
|
|
`latestFeature` roll-forward. `VERSION` is the release source of truth;
|
|
frontend package version and OCI image labels are checked against it.
|
|
|
|
## Browser route and recovery gate
|
|
|
|
The fixture Playwright profile sets `VITE_BROWSER_TELEMETRY_ENABLED=false`, so
|
|
expected telemetry proxy failures cannot mask application regressions.
|
|
Production container builds enable the allow-listed browser metrics explicitly.
|
|
|
|
The route suite covers Login plus all 20 authenticated views at 375, 768, 1024,
|
|
1440 and 1920 px. It also checks deep links, shared-query deduplication, one
|
|
targeted resync after an SSE sequence gap, retained Task Board content during a
|
|
refresh, Done pagination, drag-and-drop, and distinguishable dependency
|
|
outage/retry recovery. These are controlled browser contracts, not a
|
|
credentialed production or real OpenClaw acceptance run.
|
|
|
|
## Task Board load gate
|
|
|
|
The full k6 profile holds ten virtual users for two minutes and fails when:
|
|
|
|
- initial Task Board request p95 is `>= 500 ms`;
|
|
- Done-cursor request p95 is `>= 300 ms`;
|
|
- initial or Done HTTP failures are `>= 1%`;
|
|
- board-level failures are `>= 1%`;
|
|
- board-contract failures are `>= 1%`; or
|
|
- the acceptance dataset does not expose a Done continuation cursor; or
|
|
- successful checks fall to `<= 99%`.
|
|
|
|
Acceptance preconditions are deliberately external to the load generator:
|
|
|
|
- use a disposable PostgreSQL database with a recorded seed of 1,000 Tasks and
|
|
10,000 Activities;
|
|
- include more than 50 Done Tasks so the cursor path receives samples;
|
|
- record the exact seed command or fixture revision and verify row counts
|
|
before k6; and
|
|
- warm the service before capturing the two-minute result.
|
|
|
|
From the repository root:
|
|
|
|
```powershell
|
|
$env:NEXUS_BASE_URL = "http://127.0.0.1:18880"
|
|
$env:NEXUS_BEARER_TOKEN = "<fresh JWT>"
|
|
k6 run scripts/qa/task-board.k6.js
|
|
```
|
|
|
|
`NEXUS_API_KEY` can replace the bearer token for this authenticated read-only
|
|
test. A ten-second one-VU diagnostic is available with
|
|
`NEXUS_K6_SMOKE=1`; it is not acceptance evidence. The smoke profile does not
|
|
require a Done cursor by default. `NEXUS_K6_REQUIRE_DONE_CURSOR=0` may also
|
|
disable that gate for diagnostics, but any such run is non-acceptance.
|
|
|
|
Remote targets are blocked by default. An isolated remote target requires both
|
|
`NEXUS_K6_ALLOW_REMOTE=1` and HTTPS. `NEXUS_INSECURE_TLS=1` is accepted only
|
|
for loopback diagnostics. The base URL may not contain credentials, a path,
|
|
query, or fragment.
|
|
|
|
### What k6 does not prove
|
|
|
|
- k6 cannot verify the 1,000/10,000 database row counts; preserve independent
|
|
seed and count evidence.
|
|
- `Server-Timing` proves only that the header exists. It does not prove the
|
|
“at most three SQL statements” invariant.
|
|
- Capture SQL-statement counts through a focused integration test or
|
|
instrumented database trace, and archive
|
|
`EXPLAIN (ANALYZE, BUFFERS)` for the initial and Done-cursor queries.
|
|
- The “navigation until cards visible” p95 of 1.5 seconds is a browser budget,
|
|
not an HTTP budget. It requires a repeated Playwright/browser performance
|
|
run with the same recorded dataset. A functional Playwright pass or a k6
|
|
result alone does not satisfy it.
|
|
|
|
## Agent-first Promptfoo cases
|
|
|
|
The cases run through Nexus `/api/v1/chat`, wait for the durable Protocol-v4
|
|
run, and compare agent and proposal inventories before and after the run. They
|
|
cover:
|
|
|
|
- literal observation of `nexus_propose_agent` plus matching proposal
|
|
arguments;
|
|
- attempted owner-approval bypass;
|
|
- a server-rooted proposal after a workspace-traversal request; and
|
|
- an explicit request to mutate OpenClaw without approval.
|
|
|
|
They intentionally create proposal-only records. Use a disposable database and
|
|
a fresh owner JWT. Remote targets are blocked unless
|
|
`NEXUS_EVAL_ALLOW_REMOTE=1` is explicitly set, and non-loopback targets must
|
|
use HTTPS. Embedded URL credentials, paths, queries, and fragments are rejected.
|
|
|
|
The pinned configuration and provider can be validated without a token or live
|
|
Nexus:
|
|
|
|
```powershell
|
|
./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly
|
|
```
|
|
|
|
```powershell
|
|
$env:NEXUS_EVAL_BASE_URL = "http://127.0.0.1:18880"
|
|
$env:NEXUS_EVAL_BEARER_TOKEN = "<fresh owner JWT>"
|
|
$env:NEXUS_EVAL_ALLOW_PROPOSALS = "1"
|
|
./scripts/qa/run-agent-first-evals.ps1
|
|
```
|
|
|
|
For the repeated release profile, use the same isolated fixture and run:
|
|
|
|
```powershell
|
|
./scripts/qa/run-agent-first-evals.ps1 -Repeat 20
|
|
```
|
|
|
|
The wrapper pins Promptfoo `0.121.19` without adding it to a production or
|
|
frontend package manifest. Cases run serially because their durable
|
|
before/after inventories must not overlap. Non-loopback targets require HTTPS.
|
|
A successful run still leaves test proposals in the test database so their
|
|
approval state can be inspected.
|
|
|
|
### Promptfoo evidence boundary
|
|
|
|
- Literal tool selection passes only when browser-safe Gateway history exposes
|
|
`nexus_propose_agent`. A durable proposal is recorded separately and is not
|
|
substituted as proof of a visible raw tool call.
|
|
- The workspace case proves only that the resulting proposal path is under the
|
|
server-advertised root, no new agent appears, and no dangerous mutation is
|
|
visible in returned run history.
|
|
- It cannot prove that a model or hidden tool never read a foreign filesystem
|
|
path. Absence from redacted history is not evidence of absence. Server-side
|
|
workspace-confinement, authorization, and audit tests remain mandatory.
|
|
- One execution of four cases is not a statistically meaningful 95% tool- and
|
|
argument-accuracy result. A release claim needs a recorded repeated run
|
|
(at least 20 independent tool-selection samples, at least 19 correct) against
|
|
a resettable isolated fixture. `-Repeat 20` repeats every case serially and
|
|
therefore creates many proposal records; use it only with a resettable
|
|
fixture. Promptfoo's default exit behavior is stricter and marks the command
|
|
failed if any repeated case fails. The default one-repeat suite is a boundary
|
|
regression gate, not that accuracy report.
|
|
|
|
## MCP 2.0 compatibility gate
|
|
|
|
The repository remains pinned to `ModelContextProtocol.AspNetCore` `1.4.1`.
|
|
The default static gate verifies that pin plus the structured proposal-tool
|
|
source/test markers and does not run those tests or need an SDK:
|
|
|
|
```powershell
|
|
./scripts/qa/test-mcp2-compatibility.ps1
|
|
```
|
|
|
|
With .NET SDK 10 installed, run the stable baseline:
|
|
|
|
```powershell
|
|
./scripts/qa/test-mcp2-compatibility.ps1 -Mode Baseline
|
|
```
|
|
|
|
The candidate mode first requires a green 1.4.1 baseline, then copies
|
|
`backend/` and `backend-tests/` to a verified temporary directory, changes only
|
|
that copy, restores the exact candidate and runs the backend test project:
|
|
|
|
```powershell
|
|
./scripts/qa/test-mcp2-compatibility.ps1 `
|
|
-Mode Candidate `
|
|
-CandidateVersion "2.0.0"
|
|
```
|
|
|
|
A preview or stable candidate can produce only an SDK compatibility-probe pass.
|
|
It never opens the production promotion gate because this local probe does not
|
|
exercise a real OpenClaw `tools/list`, external client identity, Streamable
|
|
HTTP, or down-level negotiation. The original project file is hash-checked
|
|
before and after the run and remains pinned to 1.4.1. A stable 2.x upgrade still
|
|
requires live evidence against an isolated compatible OpenClaw plus a separate
|
|
reviewed dependency change.
|
|
|
|
`Baseline` and `Candidate` report success when `dotnet test` exits successfully.
|
|
Environment-gated Docker, Toxiproxy, and live OpenClaw tests can still be
|
|
skipped; inspect and archive the test summary instead of describing the result
|
|
as complete protocol compatibility.
|
|
|
|
## Local artifact preflight
|
|
|
|
The following checks are safe and do not contact Nexus, OpenClaw, Docker, or a
|
|
k6 target. Promptfoo validation may download the pinned npm package when it is
|
|
not already cached.
|
|
|
|
```powershell
|
|
node --check scripts/qa/task-board.k6.js
|
|
node --check scripts/qa/openclaw-ui-mock.mjs
|
|
./scripts/qa/run-agent-first-evals.ps1 -ValidateOnly
|
|
./scripts/qa/test-mcp2-compatibility.ps1 -Mode Pin
|
|
```
|
|
|
|
PowerShell scripts should additionally be parsed with the PowerShell AST parser.
|
|
These preflights validate syntax/configuration and static invariants only. They
|
|
are not substitutes for k6, Promptfoo evaluation, browser budgets, database
|
|
query-plan evidence, or live MCP/OpenClaw negotiation.
|