18 KiB
Scanner rework status
Progress on the approved scanner/OCR rework. See ADR-007/008/009/010 in DECISIONS.md for the decisions behind these. For the current live automation runbook, see AUTOMATION_LIVE_SCAN.md.
Current Scanner Status
See scanner-ik-progress-report.md for the full report.
Current status after the 2026-07-09 merge to main:
- The scanner architecture now follows the relevant Inventory Kamera model: 32 artifact targets per page, lookup-derived fields, fast artifact OCR profile, direct detail-fingerprint verification from the OCR capture, page-overlap planning, and batched store work.
- The live runner can compare
currentandik-traineddataengines and rejects runs that are fast but fail miss/review quality thresholds. - The final current-vs-IK-traineddata 100-artifact comparison is now proven for
the current live environment. On 2026-07-08,
npm run scan:goal:compare:validatedpassed with evidence atoutputs/live-soak/2026-07-08T18-38-35/scan-performance-assessment.json.currentwon with100/100parsed,0review,0misses, and378 ms/artifactactive average.ik-traineddatawas rejected at 100 because it parsed97/100, had5review and3misses. The next optional speed target remains3 artifacts/second, which means333 ms/artifactor faster on clean 20-artifact iterations. - The merge-ready default is the visible-inventory path. The app blocks normal
guided Auto-Scan unless Artifact inventory and a visible detail card are
detected.
auto-entry,direct-inventory, andpaimon-menuremain explicit Dev-Control experiments. - Ownership and lock-state proof is no longer theoretical: live captures parsed
equipped characters (
Citlali,Linnea), unlocked artifacts reportedlocked: false, a visibly locked artifact reportedlocked: true, and a bounded auto-scan persisted the locked/equipped state.
Done (implemented, unit-tested, build green)
- OCR eval harness —
src/eval/,npm run eval, gate innpm test. See ocr-eval.md. - C# input/capture sidecar —
native/input-helper/,npm run helper:build. Replaces the PowerShell helper on the same JSON protocol; PowerShell remains a fallback. Verified end-to-end (spawn, runtime, base64 capture). - Layout profiles + OCR preprocessing —
src/lib/layoutProfile.ts(pure geometry, 16:9 detection),src/lib/ocrPreprocess.ts(grayscale + Otsu binarize). main.ts now uses calibrated 16:9 detail/count/grid coordinates first and OCRs an upscaled + binarized copy. - Card-ready gating —
src/lib/cardReadyGate.tsreplaces the fixed 280 ms settle with change+stability polling; robust to animation. - GOOD interop —
src/lib/goodInterop.ts(export + best-effort import for scanned records), Electron file-picker import/export, and store merge. - Rescan-merge —
src/lib/artifactMerge.tscollapses leveled re-scan duplicates by a level-independent identity. - Data staleness warning —
src/lib/dataPackageStatus.ts, surfaced in the Scanner Diagnose data-package line. - Lock detection (experimental) —
src/lib/lockDetection.ts, wired into live capture as a read-onlylockedflag and persisted with scanned records. - Elevated live automation path —
npm run dev:adminnow starts throughscripts/dev-admin.ps1and logs tooutputs/admin-start/admin-dev.log. Live status confirmedisElevated: true,genshinFound: true, andtargetProcess: "GenshinImpact". - Read-only click probe —
/automation/probe-click?index=1verified that the app can focus Genshin, move to a visible inventory tile, click it, and observe a changed detail panel fingerprint (clicked: true,inputBlocked: false,changed: true). - Bounded auto-scan validation —
/scanner/start?limit=2completed live with 2 clicks, 2 verified detail views, 2 parsed artifacts, 2 stored records, 2 review samples, and 0 misses. - Auto-scan OCR performance pass - auto-scan captures now use an artifact OCR mode that skips inventory-count OCR on each tile, keeps equipped-character OCR on the real artifact-read captures, raises the substat crop to catch artifact level, stores automatic review samples without full-screen/inventory screenshots, reads only the tail of large JSONL files, avoids review noise when only level/equipped is missing, starts the scan with an OCR-free preflight capture, prevents repeated startup review reprocessing, omits full-frame and inventory-preview Base64 payloads from tile captures, and applies crop-specific Tesseract page-segmentation/whitelist parameters.
- Visible-page live soak helper -
scripts/live-soak.ps1now drives the dev-control health/status, smart-capture, probe-click, bounded scan, and review-tail endpoints and writes evidence tooutputs/live-soak/. On 2026-07-07 it completed probes at indices 1 and 3 plus scan limits 2, 5, 10, and 20 against the elevated running app. The limit 20 run finisheddonewith 20 attempted, 20 verified, 18 parsed, 18 stored, 1 review, 1 duplicate, 1 miss, and 1 page. - Scroll/page-transition live soak - after the helper and loop fixes,
scripts/live-soak.ps1 -Limits 45 -ProbeIndices 1 -SkipSmartCapturecompleteddoneon 2026-07-07 with 45 attempted, 45 verified, 35 parsed, 35 stored, 9 review, 1 duplicate, 9 misses, and 2 pages. This validates that the scanner can cross from the first visible page into a scrolled page in the live 1920x1080 setup. - Lookup package layer -
scripts/generate-genshin-data.cjsnow emits normalized lookup keys, GOOD keys, piece/set/slot links, aliases, source version metadata, generated time, and validation summary.src/lib/genshinLookup.tsprovides pure matching and validation APIs, and the scanner status/dev-control path exposes lookup validity. Auto-scan preflight blocks when the lookup package is invalid. - Inventory-Kamera-style field split - artifact detail crops now separate name, slot, main-stat label, main-stat value, level, substats, set effects, and footer. OCR uses field-specific PSM/whitelist cleanup, and the parser derives slot/set/main-stat through lookup constraints before falling back to review.
- Paimon-menu auto-entry scaffold - auto-scan supports
scanEntryMode: "paimon-menu"and/scanner/start?entry=paimon-menu&limit=N. The entry sends only read-only navigation (ESC,B, artifact-tab click), then requires a valid lookup, supported layout, and detected artifact grid before the scan loop starts. The existing visible-inventory start remains the fallback/debug path. - OCR benchmark endpoint scaffold -
/scanner/benchmark-ocr?limit=Ncaptures identical artifact crops with the current engine and returns timing/field counts./scanner/benchmark-ocr?engine=comparecan also compare the local Inventory-Kamera-traineddata Tesseract.js path whengenshin_fast_09_04_21.traineddatais present indata/tessdata,work/, orIK_TESSDATA_DIR. The OCR worker pool defaults to four workers and can be tuned withGAA_OCR_WORKERS=1..8. Native Tesseract is still not the default and should only replacetesseract.jsafter the benchmark proves it faster and more accurate on the same crops. - Quality-gated live comparison -
scripts/live-soak.ps1now supports goal runs forcurrent,ik-traineddata, andcompare, writes CSV/JSON summaries, groups results by limit, identifies timing bottlenecks, and rejects winners that miss the requested count, exceed 2% misses, or exceed 15% review.npm run scan:assessment:testverifies this ranking logic without Genshin. The assessment also reportsgoal100Decisionandgoal100.comparisonComplete, so a single-engine 100-artifact run cannot be misread as the final IK comparison. Usenpm run scan:iterate:compare:validated:waitfor the 20-artifact live iteration andnpm run scan:goal:compare:validated:waitfor the final proof when starting directly after UAC. The validator--summaryoutput includes the assessment path and timestamp for reporting. - State-polled guided entry - the guided auto-entry waits for Inventory, artifact grid, and first detail card evidence instead of sleeping the full fixed delay every time. OCR/review/store work still starts only after artifact detail preflight passes.
- Hot-loop speed pass (2026-07-08) - the scan loop no longer performs a
separate card-ready capture before OCR; the artifact OCR capture itself
verifies detail-fingerprint change. Routine click diagnostics and scan stat
publishes are throttled. Auto-scan artifact captures no longer update the
full preview/topbar UI on every tile. Store writes can be batched so the scan
path avoids per-artifact save/reload churn. Auto-scan artifact captures now
use a direct GDI hot path and skip Electron
desktopCapturer.getSources()in the per-artifact loop. - 3/s instrumentation pass (2026-07-08) - artifact hot-path captures omit
the detail-preview payload, and scan stats now split inner capture time from
end-to-end capture roundtrip time. Use
averageCaptureRoundTripMsandaverageCaptureRoundTripOverheadMsin the nextlimit=20live iteration to decide whether the next cut belongs in native capture transport or OCR. - Repeatability and capture-overhead guardrails (2026-07-09) -
scripts/live-soak.ps1now writes capture roundtrip and roundtrip-overhead timing intoscan-performance-assessment.json. The assessment validator can enforce optional speed budgets with--max-active-average-msand--max-capture-roundtrip-overhead-ms, andnpm run scan:repeatability:waitruns 20/45/100 current-engine passes as repeatability evidence without presenting them as an IK comparison.-RepeatabilityRunnow sets those limits inside PowerShell, and the script refuses unsafe limits above 1800 so npm/cmd argument parsing cannot accidentally turn20,45,100into one oversized run. - Distinctive partial piece recovery (2026-07-09) - a live repeatability run
exposed four identical OCR misses where the piece name was read as
Wontiroms Creation pan. The parser now derives a piece only when a long OCR fragment uniquely matches exactly one known artifact piece. This recovered the local case asSharpness That Ceased Upon Wondrous Creation/Disenchantment in Deep Shadowwithout adding a broad fuzzy exception. - Repeatability live pass after parser fix (2026-07-09) -
outputs/live-soak/2026-07-09T09-29-11/scan-performance-assessment.jsoncaptured a clean current-engine 20-artifact run:20/20parsed,0review,0misses,336 ms/artifactactive average,318 msaverage capture roundtrip, and138 msaverage roundtrip overhead. The strict 3 artifacts per second budget still failed by 3 ms (336 msvs333 ms). - 3/s follow-up experiments (2026-07-09) - tested and rejected several
shortcut-style optimizations because live runs got slower or added risk:
skipping Paimon-menu analysis, skipping lock-state as a production shortcut,
reducing the artifact-level crop scale, and raising the OCR worker pool to 6.
The kept low-risk changes are fast-profile OCR crop priority and avoiding a
duplicate DataURL string when the native helper already returns Base64. A
follow-up clean 20-artifact run after payload cleanup reached
351 ms/artifact,331 mscapture roundtrip, and146 msroundtrip overhead, so the next credible 3/s work is native capture transport/roundtrip reduction, not UI recommendation work. - 3/s live attempt (2026-07-08) - the missing-detail-preview review trigger
was fixed and tested. The best clean 20-artifact run reached
7285 ms(364 ms/artifact, about2.75 artifacts/second) with 0 review and 0 misses. The final stable run on2026-07-08-direct-gdi-reviewfixcompleted20/20with 0 review, 0 misses, and7973 mselapsed (399 ms/artifact). Detail region capture, 5 OCR workers, DataURL buffer decode, and substatPSM.SINGLE_COLUMNwere tested and rejected as slower. - Review-to-eval loop (2026-07-08) -
npm run eval:review-candidatesexports the local review queue intooutputs/review-eval-candidates/as a human-labeling worklist. The exporter deduplicates samples, surfaces complete fast-field captures first, marks stale captures, and now surfaces equipped footer OCR pluslocked=true/falsepayload counts for the next ownership/lock validation pass. Its output is deliberately ignored by Git and must not be treated as ground truth until fields are confirmed against the real artifact. Confirmed review labels now have a dedicated corpus file,src/eval/corpus/confirmedReviewCorpus.ts, with tests that reject duplicate ids, empty labels, and unconfirmed entries. The helpernpm run eval:prepare-confirmedgenerates a paste-ready confirmed-case snippet only when explicit expected labels are provided. - Prepared ownership/learning loop (2026-07-08) - fast auto-scan no longer
drops the artifact footer by profile alone; it omits footer OCR only when the
capture option explicitly requests that or when the footer marker is absent.
Parser tests cover noisy equipped names, split
Equipped:/name footers, and one-letter OCR fragments that must stayNot detected. Scanner learning now persists text replacements, field aliases, constrained fixes, crop adjustment proposals, and UI-profile adjustment proposals instead of truncating everything back to text replacements. - Visible-inventory merge guard (2026-07-09) - the normal guided Auto-Scan
start no longer falls back into
auto-entrywhen the artifact detail card is missing. It now blocks and asks the operator to open the Artifact inventory with a visible detail card. The explicitauto-entry,direct-inventory, andpaimon-menuDev-Control modes remain available for targeted experiments, but they are not the merge-ready default path. - Ownership live smoke (2026-07-09) - live artifact detail capture parsed
and stored an equipped footer as
equipped: "Citlali"and the grey lock state aslocked: false. A same-session visible-inventory run with/scanner/start?entry=visible-inventory&limit=20&engine=currentcompleted20/20verified and parsed,19stored,1duplicate,0review, and0misses in8047 mselapsed (402 ms/artifact). - Locked artifact live proof (2026-07-09) - a visibly locked artifact was
selected through a read-only inventory tile click. Smart Capture reported
locked: truewithlockSignal.ratio: 0.14797913950456323over threshold0.06, and/scanner/start?entry=visible-inventory&limit=1&engine=currentpersisted the same artifact withequipped: "Citlali"andlocked: true. Lock detection now decodes the lock crop PNG before measuring active lock pixels because Electron's native bitmap channel order was ambiguous in live captures.
Remaining — needs the live environment or a UI pass
These cannot be finished/validated without Genshin running at the user's resolution or without UI work best tested live:
-
Validate/tune OCR preprocessing on more real captures — confirm invert + threshold + upscale factor help (not hurt) actual Tesseract reads. The text-level eval harness cannot measure image preprocessing.
-
Wire and benchmark native IK-traineddata OCR against the same crop set. The current benchmark can use IK-traineddata through Tesseract.js; native Tesseract integration remains the next implementation step before any engine default changes.
-
Validate explicit entry modes separately from world, direct inventory, and Paimon/menu states with low limits only. These are now Dev-Control experiments, not the normal merge path; the normal Auto-Scan button blocks unless the visible artifact detail card is already present.
-
Repeat locked=true on another page/session if lock behavior changes. The first positive live proof passed on 2026-07-09, including store persistence. Further repeats are useful for confidence but no longer block the merge.
-
3 artifacts/second iteration - not reached yet. The next credible path is either native Tesseract/IK-traineddata integration that materially reduces substat OCR time, or a larger capture pipeline change that avoids full-frame PNG/Base64 transport without hurting safety checks. The target remains
<= 6667 mselapsed for 20 parsed artifacts with 0 misses and no silent OCR review regression. The latest clean 20-artifact repeatability run reached336 ms/artifact, so 3/s remains close but unproven. -
Broader scan soak test — direct-GDI current-engine runs now passed at
20/20,45/45, and100/100with 0 misses. Continue withnpm run scan:repeatability:waitin later sessions to check duplicate rate, scroll behavior, and capture roundtrip timing without changing defaults. -
Repeatability pass — repeat the qualified current-vs-IK-traineddata run in a later live session before making major OCR-engine defaults or speed claims beyond this environment. Current-engine-only repeatability is useful evidence, but it is not an IK parity claim.
Visible-page limits up to 20, scroll/page-transition limit 45, the final 100-artifact current-vs-IK-traineddata comparison, equipped footer live smokes, and one positive locked-artifact persistence proof have passed for the current environment. Remaining soak work is repeatability, OCR corpus growth, additional equipped/locked repeats, and optional 3 artifacts/second speed work.
Grow the eval corpus
Every low-confidence review sample already stores its crops + OCR. Confirm/correct
those via reviewSampleToEvalCase and commit them into src/eval/corpus/ so the
harness keeps measuring real-world accuracy across patches. See
ocr-eval.md.