Skip to content

feat(web): verification states, attempt counter and findings - #54

Merged
stephane-segning merged 1 commit into
mainfrom
claude/nifty-hypatia-letfst
Sep 30, 2026
Merged

stephane-segning merged 1 commit into
mainfrom
claude/nifty-hypatia-letfst

Conversation

@stephane-segning

Copy link
Copy Markdown
Contributor

1. Summary

This PR changes:

  • Badge: a verifying state with its own colour token (contrast about 7:1 light, 9:1 dark) and spoken text. It counts as active, so Cancel is offered while it lasts.
  • Attempt counter: "Attempt 2/3", read aloud as "Attempt 2 of 3". It comes from STATE_SNAPSHOT.job, falls back to Thread.job, and is shown only under a gate.
  • Check card (vymalo.check): one per source, attempt and verification, replaced in place by its id. It shows the status in words with an icon, the source, the attempt, the short sha and the summary.
    • Findings: shown as plain text only, never markdown or HTML. A long finding is cut, with Show more/less; more than five fold behind a control.
    • Stale answer: a card of its own, dashed, muted and marked "Stale".
  • Rework divider (vymalo.rework): "Attempt 2 of 3: sent back with N findings".
  • Failure notice: RUN_ERROR code:"checks_failed" becomes "Checks failed after N attempts".
  • Hold: a hold while verifying reaches the runtime as a pending interrupt.
  • Mock (web/mock/): replays verify-green and verify-red frame for frame, so every golden is covered. It follows the orchestrator's check-card ids check-<attempt>-<verification>-<source>, and takes job.sha only from branch artifacts. Two mock-only scenarios are added: verify-ci and verify-wait.
  • System suite: a gated agent in the Playwright system suite's agents file, and a verification.spec.ts against the real orchestrator.

It solves:


2. Intent

The intent of this PR is:

Slice 3 made the orchestrator verify an agent's work and send it back. This slice lets a person in the chat see it: that the work is being checked, which attempt it is on, what failed and why, and when checks ran out. Findings are text an agent or CI produced, so they are never rendered as markup.


3. Scope

In Scope

  • Web renderers, parsing (parseCheck, parseRework and parseJob return null on malformed input), the badge, the counter and the notice.
  • The mock server and its scenarios; unit, DOM, mock Playwright (with axe) and system Playwright tests.
  • Docs: web/README (a Verification section with a Mermaid pair), the agui.md rendering rules, dev/README "What the chat shows", and the mvp.md row.

Out of Scope

  • The CI result card (slice 8) and the verifier as its own subagent (slice 10). Check cards for an unknown source already draw by name.
  • A flow for sending a message while verifying. The orchestrator allows it; the mock refuses it for now.
  • No Rust, dependency or lockfile changes.

4. Verification

I verified this change by:

  • Running automated tests
  • Running manual tests
  • Checking logs
  • Checking metrics
  • Testing error cases
  • Testing permissions/security behavior
  • Testing rollback or failure behavior, if relevant

Commands run:

cd web
pnpm lint && pnpm check && pnpm typecheck && pnpm test && pnpm build
pnpm exec playwright test                 # mock suite (chromium + mobile), axe light/dark
# system suite against the real orchestrator + fake agent on Postgres
cd .. && node tools/docs-check/check-docs.mjs && node tools/agui-conformance/check.mjs

Results:

unit/DOM: 426 tests in 19 files passed
mock Playwright: 56 passed, 1 skipped (pre-existing skip); Lighthouse passes
system Playwright: 22 passed (4 new verification scenarios: red once -> done 2/3, red -> checks_failed, pass, no gate)
lint/check/typecheck/build clean; docs OK; 22 AG-UI goldens pass
  • XSS-style tests: <script>, markdown, <img onerror> and javascript: links in findings render as text, both at unit level and end to end from the wire.
  • Robustness: a reconnect mid-verification and a fresh page show the same state.
  • Rebase: this is rebased onto main after feat(orchestrator): inbox, watches and timers #53, which changed no web files or goldens.

5. Screenshots / Evidence

Screenshots were taken locally in light and dark for four states: verifying with a pending card, green after a rework, stale and replaced cards, and checks failed. They are not committed. CI runs the web, mock and system suites on this PR.


6. Risk Assessment

Risk level:

  • Low
  • Medium
  • High

Potential risks:

  • Findings are untrusted text. Agent or CI output could carry markup.
  • The mock can drift from the orchestrator's projection.

Mitigation:

  • Findings: rendered only as React text nodes, and the XSS-style cases are tested.
  • Mock drift: golden.test.ts requires the mock's connect stream to equal every golden, including verify-*.

7. AI Usage Declaration

AI was used for:

  • Understanding existing code
  • Generating code
  • Refactoring
  • Generating tests
  • Drafting documentation
  • Reviewing the diff
  • Not used

Human verification:

  • I understand every meaningful change in this PR
  • I checked generated code manually
  • I checked generated tests manually
  • I removed unsupported AI assumptions
  • I accept responsibility for this PR

8. Reviewer Focus

Please focus your review on:

  • Correctness
  • Architecture
  • Security
  • Performance
  • Tests
  • Maintainability
  • Product intent
  • Edge cases

🤖 Generated with Claude Code

https://claude.ai/code/session_01LvSwLecV9s4iZdbQ9HeBEJ


Generated by Claude Code

MVP slice 4 (ADR 0018): the web renders the verification gate the
orchestrator projects since slice 3.

- A "verifying" badge state with its own label, colour token and spoken
  text; it counts as active, so Cancel is offered while the work is checked.
- An attempt counter ("Attempt 2/3") beside the badge, from the STATE_SNAPSHOT
  job (or Thread.job before the stream); shown only when a job is present.
- A vymalo.check card per source and attempt (status in words, source, short
  commit, summary, findings), replaced in place by id; a stale answer is its
  own muted card. Findings and every other string are plain text, long ones
  cut with an expand control, long lists folded.
- A vymalo.rework divider: "Attempt 2 of 3: sent back with N findings".
- RUN_ERROR checks_failed is kept by ThreadAgent and shown as "Checks failed
  after N attempts" instead of the ordinary "This thread is failed".
- The mock replays verify-green and verify-red (golden.test.ts no longer
  skips them) plus verify-pass and the mock-only verify-ci and verify-wait;
  Thread.job is served by the mock.
- Tests: parsing and renderers (malformed payloads draw nothing, unknown fields
  are ignored, markup and markdown in findings stay text), the goldens through
  the runtime, both goldens through the whole ChatShell against the mock
  (state history, counter, order, reconnect mid-verification), mock Playwright
  with axe, and the system suite against the real orchestrator with a gated
  fake agent.
- Docs: web/README.md, agui.md rendering rules, dev/README.md, mvp.md, ADR 0018.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LvSwLecV9s4iZdbQ9HeBEJ
@changeset-bot

changeset-bot Bot commented Sep 30, 2026

Copy link
Copy Markdown

⚠️ No Changeset found

Latest commit: 0333694

Merging this PR will not cause a version bump for any packages. If these changes should not result in a new version, you're good to go. If these changes should result in a version bump, you need to add a changeset.

Click here to learn what changesets are, and how to add one.

Click here if you're a maintainer who wants to add a changeset to this PR

@github-actions

Copy link
Copy Markdown

✅ AI Governance check passed

This PR declares AI usage, references a source of truth, and provides verification evidence. Thank you.

@stephane-segning
stephane-segning marked this pull request as ready for review September 30, 2026 10:34
@stephane-segning
stephane-segning merged commit 4d1adc0 into main Sep 30, 2026
15 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants