# 70 — Calm Agent experience: visual reference and implementation ledger Status: **approved direction; implementation in progress, not shipped as a unified experience**. Updated 2026-09-24. This is the durable work list for the four screens created with the user in this task. The images are design references, not evidence that the UI or example data exists. Do not substitute the older `docs/assets/presentation-2026/` mockups for these screens. ## User decision and product aim Vak is for everyday work across many domains. Most users are not developers. A person should arrive, say or type what they want, work with an Agent, see a clear result, and continue their day. The interface should be calm, attentive, smooth, and useful every day. It should conceal machinery until that machinery helps the user make a decision; it must not remove capabilities. Voice is a primary/default entry alongside typing. Agents have recognizable characters, glyphs, names and personalities, with restrained motion instead of perpetual attention seeking. The user explicitly chose **invited humans in the same Agent conversation and draft**. Human and Agent coworking, inline revision, review, and acceptance are central to this direction. Sandbox work is primarily an Agent/LLM working environment; the user sees the outcome, evidence, preview and decisions. Coding, canvas and technical inspection remain powerful when relevant, without defining Vak's whole identity. Preserve the existing broad capability and presentation system, including the many semantic card/renderer types, while changing their everyday expression. Authoritative architecture still comes from `64-agent-owned-platform.md`. Use `30-output-engineering.md`, `30-render-architecture.md`, `57-adaptive-presentation-runtime.md`, `67-presentation-renderer-guide.md`, `54-task-environments-and-promotion.md`, `66-immersive-artifact-canvas.md`, `38-voice-personality.md`, `49-live-voice.md`, and `69-shared-conversation-coworking.md` for their respective contracts. Read each document's `Status:` before treating it as shipped. ## The four screens to preserve | Screen | Reference | Essential interaction | |---|---|---| | 1. Everyday result | [01-everyday-result.png](../assets/vak-experience-2026/01-everyday-result.png) | An Agent gives a concise outcome first; a result card exposes preview, checks, draft status, review and revision without filling the conversation with internal activity. | | 2. Everyday plan | [02-everyday-plan.png](../assets/vak-experience-2026/02-everyday-plan.png) | A nontechnical plan is a useful, editable result with options and actions; voice and typing are both ready in the composer. This is a universal product, not a coding shell with a friendly theme. | | 3. Review and accept | [03-review-and-accept.png](../assets/vak-experience-2026/03-review-and-accept.png) | Show the actual candidate, exact changes, observed checks, destination and remaining uncertainty. Accept selected changes into the workspace only after review; request changes or keep the draft. | | 4. Conversation and canvas | [04-conversation-and-canvas.png](../assets/vak-experience-2026/04-conversation-and-canvas.png) | The Agent conversation and live artifact are visible together. The person can try, comment, ask for a revision, switch viewport, focus the canvas and move to review without losing context. | These are generated visual proposals. Example names, dates, results, checks, venues, prices and code are illustrative, not app fixtures or claims about runtime evidence. Preserve the **relationships and interaction patterns**, not accidental text or fake proof. A later visual may replace an image only if it covers the same user decisions at least as clearly; record the replacement and why here. Do not silently revert to the old dense workbench or use the old `presentation-2026` mockups as the target. ## Feasibility review — 2026-09-20 **Decision: build these four as anchor journeys, not literal screenshots or an exhaustive screen set.** Their layout and interaction direction fit Vak's architecture. A working implementation must also include a voice-first entry/listening state and a shared-human conversation/review state, which the four images do not depict. Add those companion states without replacing the four anchors. Every example fact, check and action must come from runtime records and available capabilities. | Anchor | Existing support found | Required wiring before it is truthful | |---|---|---| | Everyday result | Agent-owned conversation/session; every projected Agent answer now has a stable result ID; structured/adaptive card renderer; artifact canvas; result-bound candidate review entry point. | Compose one primary result from the semantic timeline, attach observed checks to that exact result/version, and expose review/preview/revise from the same identity. The example “8 passed” cannot be shown without an observed check receipt. | | Everyday plan | Timeline, table, research, recipe, metric, card and other semantic renderers; voice composer; general Agent conversation. | Select a useful plan/option presentation from actual result data. Calendar, venue and other external actions appear only when an installed connector and a scoped action proposal exist. A rendered plan or invented venue is not a completed booking or calendar write. | | Review and accept | Draft file manifest, hash comparison, selectable file review, server-held result-bound candidate lookup and promotion receipt. Export now copies reviewed bytes to a server-owned candidate directory and review reads those bytes through a session-scoped route. | Journal and recover multi-file apply; verify integrated target; expose conflicts, missing checks and separate setup/publish decisions. Current per-file promotion is not crash-safe as a group. | | Conversation and canvas | User-opened split/focused canvas, HTML/dev-server/PDF/image/code previews, viewport controls, Agent steering from the canvas. Review can now open a saved candidate file in Canvas, which reads the saved version by session/candidate/file identity. Compound HTML embeds linked manifest CSS, scripts and media from that same candidate into its sandboxed document. | Support version-bound live previews; add typed comments anchored to artifact/result/version and selected location; isolate interactive preview from the control origin and live data; synchronize conversation, draft and review state. Current feedback is a steering text message with candidate context, not an anchored comment. | Cross-cutting feasibility limits and decisions: 1. **One product, two disclosure levels.** `57-adaptive-presentation-runtime.md` already defines Everyday and Advanced as projections of one conversation/history. Implement the four anchors on that state. Advanced can reveal files, diff, terminal, route, provenance and receipts; it must not fork the conversation or alter authority. 2. **Voice is prominent, never ambient recording.** The governed WebSocket and composer control exist, but microphone use requires an explicit gesture and permission. “Voice as default” means an obvious, ready path with listening/processing/interruption states and text parity. Realtime providers and full desktop/browser coverage remain extensions in `49-live-voice.md`. 3. **Character is a system, not five icons.** Agent identity already carries a name, personality and character. The current `AgentMark.tsx` is an initial mark renderer. The final design needs stable character assignment, readable names, voice/tone expression, accessible identification and behaviour across surfaces. The screenshots' Mira shape is a reference, not yet the implemented glyph. 4. **Shared people need real principals.** `ConversationContext` names an audience, but the browser currently authenticates with one operator cookie. Same-conversation coworking requires invitations, distinct principals, route-wide audience guards, attributed append-only messages/comments, authority ceilings and immediate revocation. A participant chip or “Share” button alone would falsely imply safe collaboration. See `69-shared-conversation-coworking.md`. 5. **Presentation breadth remains.** The current structured renderer registry and declarative primitives cover many result kinds, with a fixture harness. The screenshots show four compositions, not a replacement for the full renderer vocabulary. Add composition/selection and responsive variants; do not reduce Vak to weather, plans or code. 6. **Technical proof follows the result.** Sandbox tests, browser checks, target checks and external receipts are separate evidence. Show “passed,” “ready,” “applied,” or “free shipping” only from the relevant observed result or live artifact. A preview may be unavailable, stopped, stale, or noninteractive; the UI must represent those states. The deep feasibility conclusion is **yes, conditional on the wiring above**. The proposed interaction does not require replacing Vak's Agent loop or card engine. The largest work is binding identities, candidate versions, evidence and authority across existing surfaces. The screens cannot be called implemented until end-to-end tests demonstrate those joins. ## Visual and interaction rules - Use a warm, quiet canvas, readable typography, gentle curves, sparse borders and deliberate contrast. Color and motion communicate state; they do not demand attention while idle. - Give each result a clear headline, useful artifact, evidence and next action. Put operational detail one deliberate step away. Never invent a success metric or show a check as passed before observed evidence exists. - Keep the Agent present through a distinct glyph/character, name and writing voice. Animation is contextual and respects reduced-motion settings. Personality is expressed in helpful behaviour, not decorative chatter. - Keep the composer persistently available for speech and typing. Voice capture, transcription, listening, interruption and spoken responses must have legible states and parity with text work. - Let the same result move through conversation, canvas, review and acceptance using stable identities. Inline feedback identifies what the user is referring to. Revision creates a new version; prior feedback remains attached to the version it addressed. - Let a person open source, files, terminal, logs, checks and execution provenance when the task calls for them. Technical chrome is contextual rather than a permanent first impression. - Support narrow and wide windows, keyboard access, focus management, contrast and screen readers. Avoid fixed card widths and modal-only paths for routine work. ## Implementation ledger `[x]` means code exists in the current working tree; it does **not** mean the four screens match the visual references or that a release has shipped. Keep this section honest as work progresses. ### Work completed in this task so far - [x] Four visual references preserved in this repository under `docs/assets/vak-experience-2026/`. - [x] `AgentMark.tsx` adds five restrained SVG character marks, used in conversation and several Agent-facing surfaces. Idle marks are still; working marks animate subtly and respect reduced motion. This is an initial character vocabulary, not the full personality system. - [x] The conversation exposes a `Review draft` entry point for sandbox artifacts. - [x] Workbench has a candidate review dialog with actual draft/workspace file comparison, selection, inspected-file tracking, feedback to the Agent, keyboard focus handling and apply action. - [x] Artifact Canvas has an in-context revision composer that steers the current Agent conversation about the open draft. - [x] Sandbox artifacts rejoin their durable result turn through the brokered tool-call/execution identity. Artifact review is a typed `OutputItem` action. The client no longer infers produced deliverables from filenames mentioned in Agent prose or attaches session-wide executions to unrelated answers. - [x] Candidate export now generates and stores a server-owned candidate ID/manifest. Promotion loads the saved candidate, applies only selected manifest paths, rejects changed candidate content and workspace conflicts, and records a receipt. The arbitrary browser `POST /sandbox/records` route was removed. - [x] Candidate and promotion records are scoped to the owning session and bind `turn_id`, stable `result_id`, brokered `execution_id` and candidate digest. Export verifies the requested source against durable execution evidence. Session routes recover pending reviews across reconnects and an unbounded turn history without depending on the visible/current turn. The Operations Center is an evidence view; it no longer offers a second acceptance path. - [x] The conversation initially mounts the latest 40 turns and progressively reveals earlier 40-turn windows when the person scrolls upward, preserving the reading position while new turns arrive. The durable append-only history remains the source; the UI no longer mounts an arbitrarily large transcript DOM on first open. - [x] Long conversations have a quiet turn rail at the left edge. New user turns add dashes to a vertically centered stack that grows in both directions; longer histories compress the visual marks to 72, but pointer position and keyboard arrows still address every turn. Hover softly reveals the user's request without a bordered card; click or keyboard selection jumps to that turn. Far jumps mount a bounded window around the target, with adjacent windows loaded as the person scrolls in either direction and the mounted range kept to at most 120 turns. The rail reads the real transcript, including rehydrated turns, and does not introduce a separate history store. A fresh 3.5.1 dev server with a recorded five-turn conversation verified the pointer preview, click, Home/End navigation, active-mark alignment and return to Latest. Larger histories still need a visual endurance pass. - A second browser pass loaded a copied historical ledger in an isolated dev home. Its active branch projected 33 user turns from 71 raw user entries; the rail correctly uses the projected count. A saved replay variant projected 32 turns: a middle-rail click selected turn 16 of 32 and aligned its request 22 px below the pane top; Home and End reached the first and last turns. At 390 px, the 320 px dash stack stayed centered inside the conversation pane. This does not exercise the window shift above 40 rendered turns. Source review found and fixed a compressed-rail defect: when there are more than 72 turns, the active turn may fall between sampled dashes, so the nearest drawn dash now receives the active state. An over-40 real-result browser pass remains open. - [x] Settled turns now compose the Agent answer first, followed by structured cards and artifacts inside one calm result surface. Only observed evidence receipt counts, requested checks, human review state, or non-success status appear in its footer. The surface carries the stable result ID without exposing it as everyday chrome. - [x] Rehydrated sandbox artifacts inherit the result identity of their durable turn. Opening an artifact carries its owning session, result and execution into Canvas; Canvas feedback returns to that owning conversation/result even if the user changes the active conversation while it remains open. - [x] Everyday conversation now presents one quiet Agent working state while reasoning and tools run. It preserves unresolved approvals and failures inline, and offers `View activity` only when there is an exact execution to inspect. This fixes the prior blank interval where a hidden active tool also suppressed the generic activity indicator. - [x] Voice entry is visible before the first text turn and opens the selected Agent conversation before connecting. A final transcript is submitted once by the governed voice socket; the composer records it in prompt history without a second HTTP run. Browser-recognized utterances get distinct IDs, and the server deduplicates final IDs. Server-transcribed audio now starts the Agent turn as well as returning text. This is transport and dispatch wiring; live speech quality and the full voice-first screen still need end-to-end verification. - [x] Result composition now keeps structured cards and artifacts with the answer whether they arrive immediately before or after it in the durable turn. Each settled result offers an `Ask for a change` action that targets its owning session and stable result ID in the composer. Timeline plans accept an explicit `options` array as items, so noncoding choices render as readable entries rather than a JSON blob. This is a working Screen 2 path, not yet a complete plan editing or external action journey. - [x] Explicit plan `options` render as alternatives rather than numbered steps. When a card belongs to one unambiguous result, each alternative offers `Use this`, which prepares an editable, result-bound follow-up in the same Agent conversation and preserves any unsent draft. Travel comparisons with `pros`/`cons` or `left`/`right` use the existing table primitive. Travel comparison tables are presented as quiet Options surfaces without dataset search, row counts or CSV chrome; each row can prepare the same result-bound choice when an owner interaction context exists. - [x] Verification for the option slice: client typecheck and both production builds passed; the local card harness loaded and showed the noncoding alternatives; `git diff --check` passed. The design detector found only three existing layout-transition warnings in `styles.css`, outside this slice. The harness is a visual renderer gallery for card types and combinations, not a pipeline test. A live Agent conversation with an emitted options card remains to be exercised. - [x] The existing candidate review now shows an observed manifest summary (new, changed and unchanged files), the explicit destination, its linked execution/candidate identity, per-file state, and the exact selected-file count on Apply. It states that no candidate-bound check receipts are present instead of implying a build or test passed. This improves Screen 3's decision surface; immutable candidate freezing, transactional apply, integrated verification, and a dedicated route remain open. - [x] Promotion now preflights every selected destination's existing file type and parent shape before writing the first file; later directory or parent-file collisions return a conflict without applying earlier files. Focused sandbox tests cover both cases. This reduces predictable partial applies but does not provide crash recovery, concurrent-editor coordination, or atomic multi-file promotion. - [x] Candidate export now copies manifest-checked draft bytes into a distinct server-owned directory under the durable sessions home. The review dialog reads the saved version through a session/candidate/file route that verifies its hash, and promotion reads that same saved version. Later Agent writes to scratch therefore do not change the reviewed result. A changed or missing saved file still fails its hash check. This is a frozen copy with integrity checks; filesystem-level immutability, candidate-bound test receipts, crash-recoverable promotion, and preview version binding remain open. - [x] Focused sandbox and server tests verify the frozen candidate survives later scratch changes, is promoted from the reviewed bytes, refuses a different session's read, and reports tampering of the saved copy. Client typecheck and the production web bundle pass. This verifies the export/review/apply slice, not the complete four-screen journey. - [x] Review now reopens an existing pending candidate for the selected execution instead of silently exporting a new one each time. An explicit action prepares a newer version; the review surface lists pending versions, and switching versions clears selected-file inspection state. The pending list recovers from durable server records after reconnect. This keeps the decision bound to the visible candidate, while version-linked canvas preview and a dedicated review route remain open. - [x] Review can open its selected saved file in Canvas. Canvas carries the candidate ID, reads text or raw media through session-scoped hash-checked candidate routes, and labels the saved draft version. It refuses a failed candidate read instead of silently falling back to mutable scratch. Saved HTML embeds linked CSS, scripts and media only when each path belongs to the same candidate manifest; live dev-server previews are not yet version-bound. Revision feedback includes the candidate ID in its Agent steering text, but is not yet an attributed, typed comment. - [x] Review and candidate-bound Canvas feedback now use one typed candidate-comment endpoint. The server validates the owning session, saved candidate, manifest path and optional line range; records a `CandidateComment` activity with result/execution/candidate/file anchors; and sends the same text into the governed Agent conversation so model-visible feedback remains reconstructable. An active run receives steering; a settled historical conversation starts a real next turn from its writable append-only ledger. The current authenticated actor is still the single operator, and the UI does not yet expose freeform text selection. - [x] Candidate source Canvas exposes an optional exact line or line-range anchor for revision feedback, with client and server range validation. Whole-file feedback remains one click away. Direct text selection, visible anchored comment history and multi-human attribution remain open. - [x] Candidate comment history is readable from the append-only session ledger plus the active-turn buffer, deduplicated by activity ID and filtered to the saved candidate. Review and Canvas show the same anchored comments and refresh after a new comment. The displayed `You` identity is still the single authenticated operator; distinct participant attribution remains open. - [x] The server now has an append-only conversation audience-grant store with distinct principal ID/display name, conversation and audience scope, explicit capabilities, expiry, hashed bearer tokens and sticky revocation. Raw invitation tokens are never stored. Focused tests cover verification, expiry and revocation. HTTP authentication accepts a participant token only from an Authorization bearer header and attaches a principal distinct from the operator. Its initial fail-closed route policy permits GET access only to that exact conversation's transcript, presentation/results, sandbox receipt list, saved candidate bytes and candidate comment history when the grant has `read`. SSE, messages, comment writes, execution, approvals, promotion, global filesystem and configuration remain forbidden. Invitation UI/browser redemption, attributed writes and immediate live-stream revocation are still open. - [x] Operator-only conversation invitation endpoints now create bounded read grants against the immutable audience recorded in the session header, return the raw bearer token exactly once, list secret-free lifecycle summaries and append sticky revocations. Grant IDs cannot be reused, expired/revoked credentials fail closed, and participants cannot call the lifecycle routes. These grants do not yet provide live presence or writes. - [x] The conversation header now opens a quiet Share dialog for the selected conversation. The owner can create a named read invitation with a bounded expiry, copy its one-time bearer credential, see active/expired/revoked invitations and revoke access. The dialog explicitly states the current read-only boundary; same-conversation writes remain open. Client typecheck and both production builds pass. The design detector reported only three pre-existing CSS layout-transition warnings outside this slice. - [x] A separate `/app/?shared=1` browser entry now accepts the invitation code by paste, holds the participant bearer only in memory, and reads the granted conversation and saved draft list through scoped routes with `credentials: omit`. It mounts the latest 40 visible human/Agent messages and reveals earlier windows on request. A revoked or expired grant clears the view. The invitation code is never placed in the URL. Commenting and live scoped refresh are implemented below; participant-authored Agent turns remain open. - [x] The participant view can now open a saved candidate text or image file through the same session/candidate/hash-checked read routes as the owner's review, and show the version-bound comment history for that file. Its refresh also updates open comment history. Other binary formats and full semantic result rendering remain open; this does not grant comment writes or Agent intervention. - [x] Invitations can now explicitly grant `comment` alongside `read`. The participant view writes an attributed comment on a saved candidate and optional file/line, while the owner and participant see the same append-only history. The server records the grant, actor, result, execution and candidate digest, and returns a comment receipt without starting an Agent turn. Existing owner revision feedback still steers the Agent. Participants cannot promote, approve or execute through this grant. Live delivery, participant messages and explicit conversion of a participant comment into Agent work remain open. - [x] In the owner's saved-draft review, an explicit action can now turn a selected participant comment into an Agent revision request. The server resolves that exact append-only comment, checks its saved candidate digest, and passes its author and file/line anchor through governed steering. A participant `read, comment` grant cannot call this action. The running two-person journey and any broader participant authority remain to be verified. - [x] A focused HTTP journey now creates a saved candidate and live session, admits a distinct participant with `read, comment`, writes an anchored comment, verifies the same attributed record through participant and owner reads, confirms the receipt did not start an intervention, rejects a read-only participant's write, and blocks the first read after revocation. This validates the server boundary, not the browser experience or live delivery. - [x] The participant view now opens a bearer-authenticated, content-free update stream for its exact conversation. Agent events and newly saved candidate comments trigger scoped re-reads; a two-second heartbeat only rechecks the grant. Revocation emits a final revoked signal and closes the stream. The previous periodic refresh remains a reconnect fallback. A focused server test covers heartbeat, comment wakeup, revocation and closure; the full two-browser saved-draft journey remains open. - [x] A running isolated server and the shipped `/app/?shared=1` browser UI now verify participant admission and live revocation end to end. A real scoped invitation opened the conversation as the named participant with comment capability; an operator API revocation returned the page to its invitation entry with the revoked/expired explanation inside the heartbeat window. This browser pass had no saved candidate, so draft opening, comment submission and owner conversion still rely on the HTTP integration test and require a full two-browser artifact journey. - [x] Owner conversion of a participant's saved-draft comment now uses the same conversation dispatch as direct owner feedback. When the Agent is running it steers that run; when the conversation is settled it attempts a governed next turn against the historical ledger. A focused test with a saved candidate, attributed anchored comment and read-only historical session confirms a missing provider is reported as unavailable rather than falsely claiming idle steering was queued. The successful two-browser conversion journey remains open. - [x] The shared participant screen now reads the scoped presentation timeline and renders settled structured and adaptive results with the same registered card implementations as the owner surface. It deliberately supplies no owner control context, so plan discussion, presentation feedback, approval, promotion and execution actions are absent rather than failing after a click. Plain transcript remains available around the cards. A real projected-result browser fixture remains to be exercised. - [x] Agent character is now required frozen session identity alongside Agent ID, revision and name. New sessions preserve the selected profile glyph instead of rereading mutable profile state, and the scoped coworking identity response exposes only ID, name, character and revision. The participant header and assistant turns render that name and glyph. Missing identity or character fails rather than inventing a compatibility default; a focused session test pins the required field. - [x] Invited people can now add attributed messages to the same Agent conversation through an explicit `message` grant. Each message is an append-only user message carrying the verified principal ID, display name and an author-scoped idempotency key; owner and participant transcripts render that identity. Posting never dispatches the Agent, invokes a tool, resolves an approval or changes the workspace. The contribution enters the durable conversation for the next owner-authorized turn. A focused HTTP test covers exact capability admission, attribution, replay idempotency and the absence of a started run. - [x] Shared conversation presence is now derived from recently observed, still-authorized participant streams. It is ephemeral server state with a seven-second lease, never a fabricated invitation status and never a ledger event. The owner header shows quiet named presence through the existing session heartbeat; participants see other currently observed people through the scoped update stream. Vanished tabs age out without disconnect bookkeeping, and revocation stops renewal immediately. - [x] The owner can delegate one currently pending Agent decision to one active invitation. The participant sees only gates assigned to that invitation, reviews the tool, reason and exact arguments, and may allow once or deny. An invitation itself grants no standing approval power. The first valid answer consumes the request and records the verified participant on the resolved approval activity. This route cannot remember a rule, accept files, promote a candidate, change configuration or resolve another gate. Focused authenticated HTTP coverage verifies owner delegation, rejection before delegation and for another gate, attribution, single consumption and the absence of a learned permission file. - [x] A fresh development binary was run against a disposable Vak home and workspace with the locally available `gemma4:e2b-mlx` Ollama model. The model twice returned thinking without a visible answer or tool call; the requested HTML draft was not created. The outcome evaluator correctly marked its deliverable unmet, but the Agent loop incorrectly returned `Completed`. The loop now returns a typed parse failure after its one bounded empty-step retry, with a focused regression test. This run does not close the real two-person draft journey; it identifies a real local-model blocker without substituting a fabricated draft. - [x] The current development binary completed the real two-browser saved-draft journey against the previously Agent-produced `isolated-review.html` candidates. A new scoped participant joined the exact candidate conversation, read both saved versions, opened Version 2, saw its real frozen HTML, and added a line-1 comment. The owner simultaneously showed observed presence and received the attributed comment in open Review with the explicit `Ask Agent to address this` action. Revoking that invitation returned the participant tab to the invitation entry with the revoked explanation immediately. This uses the development server on port 8912 and real session/candidate ledgers; the older installed app on 8901 was not evidence. - [x] A delegated gate now records the owner's exact assignment in the append-only conversation activity before the participant can answer. Repeating the same assignment is idempotent; assigning that live gate to a different invitation returns a conflict. The focused HTTP test checks the recorded delegate identity and refuses reassignment. - [x] The packaged character runtime now includes the canonical indigo-and-saffron Vak songbird plus seven palette-distinct companions, state-driven motion, deliberate local sound previews, and one identity across private and shared surfaces. The brand and extension contract is recorded in `71-agent-character-system.md`. - [x] `69-shared-conversation-coworking.md` records the chosen same-conversation/same-draft multi-human contract and its security boundaries. It is a proposal, not a shipped collaboration feature. - [x] Verification performed for the above slice: client typecheck and production build; nine focused server promotion tests; `git diff --check`. The presentation card harness loaded visually. These checks do not establish full end-to-end usability of the new screens. - [x] The running Screen 1 result now gathers working-file preview, draft review and result-bound revision into one action row sourced from the projected artifact actions and result ID. Structured-only results get the same container. A result-bound preview keeps the recorded path instead of matching another execution's file by basename or borrowing HTML from unrelated assistant text. - [x] A real isolated server with the local Ollama route produced a birthday-lunch plan through the browser. That run revealed a transcript hydration race: the server had returned HTTP 200 with only `run in progress`, which the client read as a transcript. The route now returns 409 and the client retries the settled read. The same run produced three substantive plan drafts; the latest is shown as the primary result and the earlier versions remain in one disclosure. This was a real model result, not a harness fixture. It does not prove artifact preview, candidate review, voice, or shared-human use in that same browser journey. - [x] The former intent explanation card was removed from the everyday composer. Intent resolution remains in the runtime and its explicit inspection surfaces, but the old debugging display no longer demands attention in conversation. - [x] Screen 2's option renderer accepts the real model shape of string `columns` plus positional array `rows`, normalizes it before rendering, and keeps every value visible. A real local-model conversation exposed an additional transport defect: the model supplied valid values but mixed row delimiters and appended one closing brace. The Vak fence reader now repairs only delimiter syntax that the nesting stack proves, preserves every model-supplied key and value, and continues to reject truncated or ambiguous content. Reloading that same durable conversation in the production web bundle rendered all three options without data-grid search/count/download controls. Choosing `Indoor Picnic & Craft Project` prepared the result-bound follow-up `Use “Indoor Picnic & Craft Project” in this plan.` in the real composer without sending it. The harness still covers the malformed regression alongside all registered semantic renderers, but the session and composer interaction are the end-to-end evidence for this screen. - [x] A real artifact request exposed a false model claim: `invitation.html` was absent after a failed `write` call, yet a separate card receipt made the turn appear complete. Named saved-file requests now require a successful matching write/edit receipt for a produced outcome. The result renderer no longer turns unverified file names in Agent prose into Canvas actions; preview actions come from recorded artifacts. This receipt check does not prove that the file still exists or that its contents satisfy the request, so end-to-end artifact verification remains open. - [x] After rebuilding the server, a fresh real-model follow-up asking for `invitation.html` ended without a file or successful write. The new outcome evaluator recorded `primary deliverable: Unknown`, confirming that unrelated work no longer counts as proof of the saved file. The local model did not complete the artifact journey; successful file creation, preview and review still need live verification. - [x] A later real browser run successfully wrote and read `canvas-check.html`. The settled result showed the Agent answer, observed artifact evidence and a result-bound action; opening it rendered the actual HTML in the sandboxed split Canvas with preview/code, device, reload, focus and revision controls. Durable tool projection and sandbox sidecar evidence for the same file now merge into one artifact instead of producing duplicate cards and buttons. Direct workspace writes no longer advertise `Review draft`, because they have no isolated execution root that the candidate freezer can truthfully review; isolated candidate review remains a separate journey. - [x] A real quarantined bash run created `.vak/scratch/.../isolated-review.html`, recorded `Produced`, projected one `Open` and one `Review draft` action, and froze a verified candidate containing exactly `isolated-review.html` with the same result ID. Review actions now require a live isolated execution root containing that artifact; an ordinary workspace write or an expired/missing scratch root cannot advertise review. This first checkpoint established the candidate used by the owner browser pass below. - [x] The owner browser review now survives a server restart and opens from the real conversation result. Sandbox candidate and promotion records live with the shared session data beside frozen candidates and execution evidence, while source and destination confinement resolve from the owning session workspace instead of whichever Core happens to be active. The recovered `isolated-review.html` candidate opened as Version 1 with its one new file already inspected, the exact `/private/tmp/vak-screen1-live` destination, frozen-copy integrity, an honest “No checks attached” state, expandable result/execution/candidate/digest provenance, revision feedback, saved-version Canvas, `Keep as draft`, and an exact `Accept 1 selected file` boundary. Acceptance was deliberately not triggered during visual proof. - [x] The saved candidate now opens in the synchronized conversation and Canvas with a real version label, preview/code modes, desktop/tablet/mobile widths, focus/split modes, clickable source-line anchors, the same durable attributed comment history as review, and a direct return to the exact candidate decision. A restart-recovered historical conversation accepted an anchored line-one comment and persisted it in the append-only ledger. Desktop and narrow Canvas layouts were inspected in the shipped browser UI. The initial conversion to an ordinary Agent turn was unsafe and has been closed pending isolated revision execution. - [x] A two-browser live trial used a scoped one-hour invitation for Asha Test. The guest joined the same Agent conversation, read the saved draft and added line-anchored comments; the owner saw them update in open Review and Canvas through an authenticated, audience-scoped refresh stream. The guest transcript retains its last settled content while a turn is running (the transcript endpoint normally returns 409 then). This proves live shared read/comment synchronization, not safe Agent revision. - [x] The same trial exposed a trust-boundary failure: converting a comment to a normal Agent turn requested direct workspace `edit`; after denial, the Agent used `write` on the workspace path. The test file was removed from the disposable workspace. Candidate comments now remain durable feedback, while Agent revision requests fail closed with 409 until they can execute in an isolated task copy and produce a new reviewable candidate. The UI labels comments honestly. Do not reopen the ordinary-turn path. - [x] The native sandbox now has an explicit strict task-copy profile on Seatbelt and Landlock. It excludes the ordinary workspace-write allowance for host temp trees. A macOS worker test wrote inside a copy under the system temp root and was denied a write to a sibling workspace. This is a containment primitive only; the revision route remains closed until it creates the copy, uses this profile for every effectful tool, and binds the resulting files to a new candidate. - [x] The revision-copy helper now seeds a fresh writable directory only from paths in the frozen candidate manifest, verifies every saved file hash before copying, excludes files added later to the saved directory, and removes a partial copy on failure. Its test preserves Version 1 despite later source edits and rejects tampered saved bytes. The server has not yet attached an Agent run to this copy. - [x] A child Core can now be stamped with an immutable task-copy boundary: its permission mode is capped at workspace-write, the worker uses the strict native sandbox, its tool vocabulary is limited before configured allows are considered, and inherited host-command hooks are disabled. Focused tests cover an owner FullAccess setting, an explicit allow for an out-of-scope tool, and inherited hooks. Candidate revision dispatch now uses that boundary and the parent's pinned worker binary; the child session lives in the same Agent's private session home. - [x] An owner can turn a saved candidate comment into a real isolated Agent revision. The server seeds a verified task copy from the selected frozen version, records a running revision in the candidate ledger and parent conversation, runs a child Agent turn with no answerable approval surface, and freezes changed files as a linked new candidate. The original destination is untouched until an exact candidate is accepted. A scripted-provider integration test exercises a brokered write, Version 2 export and exact promotion. A live local Ollama run revised `isolated-review.html` from a saved owner comment, produced Version 2 with changed bytes, and the rebuilt browser Review showed that version's content and selected label after a server restart. The disposable destination still had no file before acceptance. The live model initially tried an out-of-scope document reader, recovered with the allowed `read`, and repeated an already-applied edit; the final candidate was derived from verified changed bytes, not its prose. Version 2 acceptance is covered by the integration test, not by that browser trial. - [x] Review now keeps its native version selector synchronized with the actual selected candidate and removes the stale action that could regenerate a version from the original scratch execution. Canvas feedback and Review comments both call the same explicit revision route; open Review listens for the resulting candidate and failure state. - [x] Candidate acceptance now has one production path with a shared cross-process workspace lock and a persistent transaction journal outside the Agent-controlled workspace. It saves verified before-images before the first rename, records each file as prepared/applying/applied, rolls an interrupted multi-file import back before retry, and refuses recovery when a later human edit no longer matches either side of the recorded transition. A completed journal is idempotent only while destination hashes still match. Promotion filesystem work runs on a blocking worker, and focused tests cover interruption recovery and preservation of later workspace edits. Target integration checks remain open. - [x] Applied candidates now expose scoped `Undo acceptance` in Workbench, including after a reload. Undo is an authenticated, append-only server action bound to the original conversation and candidate. It takes the same shared workspace lock, restores verified before-images, removes files that acceptance added, journals per-file undo progress across interruption, verifies the resulting state, and refuses to erase later workspace edits. Focused sandbox tests cover changed/new files, replay, conflict preservation and crash resumption; the server promotion test exercises the real route and durable undo record. Target integration checks remain open. - [x] Deletion is now a first-class additive candidate operation. An isolated revision that removes a file with an existing workspace baseline freezes a `Delete` entry rather than failing or manufacturing empty bytes. Review labels and counts the deletion, shows the current content beside the explicit deletion outcome, and requires inspection before acceptance. The promotion journal verifies the baseline, records and observes destination absence, and scoped undo restores the saved original. A focused end-to-end sandbox test covers revision copy, Version 2 freeze, deletion acceptance and undo. Removing a draft-only new file produces no fake workspace deletion. ### Next implementation work: the four screens - [x] Build Screen 1 as the actual everyday Agent conversation: Agent identity and presence, compact user turns, outcome-first response, one primary result, secondary evidence, review/preview/revise controls, and quiet progress. It is driven by `OutputTimeline` and stable result IDs; the real plan, workspace artifact, Canvas and isolated candidate runs above close this anchor. Voice and shared-human companion states remain separate work below. - [x] Build Screen 2 with a real noncoding result journey. The durable Saturday-plan conversation renders a calm Options surface from the Agent's actual response, keeps all three alternatives and comparison values visible, and binds each choice back to its owning result in the everyday composer. No developer workbench is required for the journey, and the behavior is expressed through the existing semantic renderer vocabulary rather than a domain-specific screen. - [ ] Make voice visibly primary in both screens, including listening, processing, interruption and fallback to text. Use the existing voice transport and personality contracts; verify actual speech and transcription flows. - [x] Build Screen 3 as a first-class review surface. The real restart-recovered candidate presents versions, selected and inspected files, a plain change summary, explicit destination, observed-check state, frozen-copy status, expandable provenance, revision feedback, keep-draft behavior, saved-version Canvas and an exact acceptance scope. Current/draft file inspection remains in the surface, with deeper Canvas inspection available contextually. - [x] Build Screen 4 as a synchronized split conversation and canvas: real artifact preview, contextual comments or selections, draft version, responsive viewport controls, focus/split modes and direct path to review. Existing polyglot canvas formats remain available. - [x] Add the companion voice-first entry/listening state, showing microphone permission, capture, transcription, interruption, speaking and text fallback without auto-recording. The composer now exposes the permission transition, keeps interim browser transcription visible without appending every revision to the durable ledger, submits only the final utterance through the existing governed socket, distinguishes listening/processing/speaking, stops and reports playback when new speech begins, and returns directly to text after a surfaced voice failure. Capture still begins only from an explicit press. A real microphone/provider round trip remains required by the broader voice verification item above. - Desktop and browser hosts without the Web Speech recognition API now detect speech onset and a bounded silence tail from captured PCM, emit explicit `speech_started`/`speech_stopped` boundaries, and keep the governed voice socket open for later utterances. Silence is not streamed into the audio budget before speech begins. This closes the previous path where raw audio never became a turn until the entire voice session was stopped; physical microphone/provider verification remains open. - The governed turn executor now sends its completed answer over the voice socket with the originating utterance ID. The client removes control scaffolding and synthesizes that answer through the configured provider route with the owning session. New speech cancels an unfinished synthesis or interrupts playback without cancelling the Agent run. The earlier local-only WebSocket path that spoke the user's own transcript back has been removed. Playback failure leaves the text answer intact and returns capture to listening. A physical microphone/provider round trip remains to be verified. - The Voice entry now reads the effective runtime setting before opening a conversation or requesting microphone access. When voice is disabled, it opens Voice setup with a clear state. `vak doctor` on the current installed runtime reports voice disabled and the optional local TTS backend unconfigured, so no physical microphone/provider playback has been claimed as verified. - A fresh 3.5.1 dev build on an isolated loopback home verified the disabled Voice entry. Its setup path now waits for Settings data to load, scrolls directly to Voice, and focuses the enable switch; previously it landed above 75 presentation definitions. The installed `/Applications/Vak.app` serving port 8901 is 3.5.0 and was not used as evidence for this UI. This verifies the setup path, not microphone or provider playback. - A user screenshot of the real dev Settings pane exposed the Voice providers label and description collapsing into a column of broken words. Settings cards now size controls within the available pane and stack rows when the card itself becomes narrow. A rebuilt 3.5.1 dev server at the same narrow browser width showed the full label, readable description and wrapped provider state without clipping. This is a layout check, not a provider readiness claim. - [x] Add the companion shared-human state to conversation, canvas and review: invited participant identity, presence only when observed, attributed messages/comments, version anchors, individual authority and revocation. - [x] Connect all four through one result/conversation identity and coherent navigation, rather than separate modal islands. Conversation result actions carry their owning conversation and execution into Canvas and Review; saved Canvas versions additionally carry the exact candidate. Returning from Canvas reopens that candidate version rather than whichever draft is newest in the currently focused conversation, and both Canvas and Review have a direct path back to the owning conversation. A browser run against the real two-version `isolated-review.html` conversation verified conversation → working Canvas → conversation and conversation → Version 2 Review → saved Version 2 Canvas → Version 2 Review → conversation, with the same result, candidate and attributed comment throughout. ### Platform wiring required for the experience - [ ] Complete the task-environment contract in `54-task-environments-and-promotion.md`: broader task environment preparation, stronger filesystem protection for frozen candidates, candidate/result/preview/check binding, and the remaining domain verifier adapters. Isolated human-requested revisions, additions/changes/deletions, destination conflict handling, crash-recoverable multi-file apply and scoped undo are wired. Acceptance emits a durable digest for the exact applied workspace state after read-back. Registered JSON, CSV/TSV structural, SVG, PNG/JPEG/GIF/WebP decode, PDF structure, and DOCX/XLSX/PPTX Open XML package verifiers are planned into the frozen candidate, rerun from the applied workspace, and shown individually as passed or failed; unsupported target checks stay explicitly unavailable, and sandbox activity is never presented as target evidence. Additional build ecosystems, application journeys, rendered document inspection and semantic data reconciliation adapters remain open. - Frozen candidate trees are now made read only after their copied bytes pass the manifest hashes. Direct file edits and additions are denied through ordinary process access, executable bits remain usable, and revision/promotion paths continue to rehash every selected file. Cleanup unlocks only the server-owned frozen tree before removal. Broader environment preparation and the remaining domain verifiers stay open. - Failed export and revision persistence paths now use the same sandbox-owned cleanup boundary, including nested read-only directories. A direct sandbox test verifies protected additions fail, cleanup removes the complete tree, and repeated cleanup is idempotent. - Applied PNG, JPEG, GIF and WebP artifacts now pass only after the format decoder reads their pixel data and reports real dimensions; matching a short magic prefix is no longer accepted as image evidence. PDF checks parse the catalog and nonempty page tree with a decompression limit and report the actual page count. This structural evidence does not claim that pages were visually rendered. - DOCX, XLSX and PPTX checks now parse the required content-types, package relationships and document/workbook/presentation XML roots. A ZIP with correctly named but meaningless parts no longer passes package verification; visual page and slide inspection remains a separate open adapter. - Open XML verification continues through each complete required XML part and reports the applied document’s paragraph, worksheet or slide count. Malformed content after a valid root now fails instead of being ignored; these structural counts remain distinct from rendered visual inspection. - Frozen HTML/HTM drafts now receive a browser-shaped structural check before review: an unclosed raw-text element that swallows the body fails instead of looking like a usable page. The saved candidate record carries each observed draft check, and Review shows pass/fail evidence beside the planned target checks. The same verifier runs again from the accepted workspace state. This checks parsed body/script presence, not rendered visual quality or JavaScript behavior. Older saved candidates have no backfilled draft-check evidence and remain explicitly planned-only. - Canvas now presents saved CSV/TSV files as a bounded, horizontally scrollable table with row/column counts, while Source retains the exact file text. Quoted separators, escaped quotes and multiline cells parse without splitting values; inconsistent rows surface an explicit error. This is artifact inspection from the authenticated saved-file route, not an Agent-generated dataset journey or semantic reconciliation claim. - Accepted JavaScript application candidates now offer their declared production build and test scripts as separate target-workspace checks. Each check runs only after the server revalidates the accepted state digest, and its output becomes a durable workspace-check record instead of borrowing earlier sandbox evidence. The installed web bundle was rebuilt in the same change after the server freshness gate found it lagging behind the renderer source. - JavaScript workspace checks and live preview launch now share one package-manager resolver over the accepted project’s recognized `packageManager` declaration or npm, pnpm, Yarn or Bun lockfile. Commands come from fixed script templates rather than executable manifest text, so preview and post-acceptance verification no longer disagree on the project toolchain. - Managed live-preview processes receive the same scrubbed non-secret environment as brokered commands and run in an isolated process group. Persistent startup now crosses a versioned broker request into the pinned Vak worker executable; the worker owns the preview child, applies the configured worker or command sandbox target, and streams its bounded logs back to the server-owned lifecycle. Stop kills and waits for the worker process group, so a preview child cannot retain ambient credentials or survive after Canvas closes. - Preview readiness now distinguishes a missing runtime from a configured port already owned by another process. Start rechecks the port before spawning, a child that exits during readiness is removed and reported with its bounded final output, and completed children no longer remain projected as running. An unrelated listener therefore cannot be mistaken for a successfully started Vak preview. - Saved drafts can now start live previews from server-owned, candidate-bound roots. Launch discovery, executable readiness, process identity, stop and Canvas routing carry the candidate ID; startup revalidates every manifest hash before executing. The Canvas URL opens the selected saved path, so an open preview and its review/comments refer to the same immutable version. - Preview discovery now probes the configured executable against the same effective PATH used after environment scrubbing. The UI shows a concrete `needs setup` state and disables Start when the runtime is absent; startup rechecks readiness to close the stale-probe race. A manifest alone is no longer presented as a runnable environment. - JavaScript framework previews distinguish an available package manager from an actually prepared dependency tree. Declared dependencies without `node_modules` produce `needs preparation`; Canvas offers an explicit preparation action. Preparation copies only manifest-listed candidate files into a server-owned runtime, uses fixed lockfile-aware install arguments with lifecycle scripts disabled and isolated HOME/cache paths, crosses the versioned worker and configured sandbox, caps evidence and time, and revalidates every source hash before recording the candidate digest. Launch uses that prepared runtime only while the digest and source hashes still match. - Every started preview preparation now appends candidate/result/digest-bound `Preparing` and terminal `Ready` or `Failed` records with the fixed command and bounded evidence. Candidate review projects the latest record after reconnect, while the prepared runtime marker remains only a cache validity mechanism rather than the source of review truth. - Partial acceptance now persists only the checks valid for the selected files. An unselected or deleted project manifest cannot contribute its proposed commands; the Workbench reads the promoted check list on first apply and reload. Each check revalidates the accepted digest after execution as well as before it, and records failure with command output if the accepted bytes changed during the run. - [x] Executable workspace checks are separately authorized after acceptance. Cargo, explicit npm test scripts, Go and pytest projects receive an exact planned command visible during review. The command runs only from the post-acceptance action, after applied-state revalidation, through the existing broker and sandbox; every attempt appends pass/fail evidence bound to the applied digest, survives reload, refuses drift and undo, and never installs dependencies or escalates permission. - [x] Give human participants distinct authenticated principals, conversation/audience-scoped invitations, revocation and read/write route guards. Participant credentials remain in the dedicated shared view and are never exchanged for the operator cookie. - [x] Add append-only, attributed human messages and comments anchored to result/candidate versions and optional artifact locations; add explicit conversion of feedback into Agent intervention. Participant actions remain below the owner's authority ceiling; comments and messages do not dispatch work by themselves. - [x] Scope shared transcript, SSE, preview, candidate and artifact reads to the same audience grant. The authentication whitelist still limits participants to the exact conversation routes and capabilities, and a second centralized middleware boundary now resolves the conversation's live durable audience before any allowed handler runs. A valid, unexpired token for the same conversation but a different audience receives `403`; missing conversation audience also fails closed. The focused HTTP test covers valid read, audience mismatch, conversation mismatch, disallowed execution and revocation. Review comments retain actor/grant/candidate/result/execution attribution; acceptance and external publishing remain operator-only decisions. - [x] Complete Agent character/personality settings and presentation across sidebar, conversation, voice, results, channels and compact surfaces. Character remains identifiable through its frozen name, character id and silhouette without relying on animation or color alone. - Existing Agents now have a real identity editor in Fleet Roster for name, packaged character, personality, movement and voice style. Writes update only the exact Shared or workspace layer that owns the Agent and increment its revision; existing conversations retain their frozen identity. Agent creation also seeds its full replacement from the Shared layer instead of copying effective workspace overrides. The web shell now serves every packaged character asset before authentication, with an end-to-end server test; the live build renders all eight atlases without fallback failures. Movement and voice style are additive frozen conversation identity: movement drives header, conversation, results, voice, sidebar and coworking marks, while narration resolves the active conversation's style and personality after explicit user overrides. Channel deliveries carry the exact Agent id, revision, name, character, movement and voice in both answer and document provenance while their bot-scoped transport identity remains authoritative. - [ ] Audit all existing card/renderer types and presentation surfaces against the new design. Make common cards calm and legible; preserve specialist and technical capability behind contextual access. - The server admission vocabulary and desktop renderer registry now have an executable equality check. The audit removed nine client-only aliases that could never pass delivery validation, leaving the 97 canonical semantic types accepted and rendered on both sides. The visual harness identifies itself as renderer-fixture coverage rather than an end-to-end provider/tool test, maps every registered renderer to branch fixtures, and currently renders 545 permutations plus multi-card, malformed-output and character scenarios. The sweep also corrected unknown-status test rows so they no longer show a green pass mark while the summary reports zero verified. Responsive and common-card calmness review remains open in this item. - The common metric grid no longer promotes whichever field happens to arrive first into a full-width accent strip. Values now have equal visual weight, quiet static surfaces, and columns that can shrink to a narrow conversation or Canvas pane without clipping. The result prose or an explicit card title carries the hierarchy; field ordering does not invent it. - A common metric with a `label` heading and several values now renders every value instead of an empty single-value card. A location used as the heading no longer repeats as a metric tile. The renderer harness includes the labeled multi-value shape; the production web build and typecheck pass. This is renderer coverage, not a provider or tool round trip. - Option comparisons now show every choice without the dataset table's 280px internal height cap. Comparison values also no longer acquire green/red status from leading signs or words such as “fail”; those patterns were guesses about meaning, not observed check evidence. Explicit status text remains readable, and dataset search, sorting and export remain available where applicable. - At a 390 px browser viewport, the harness's fixed 340/480 px grid minimums made the fixture page 504 px wide and prevented a genuine phone-width card inspection. Its grids now shrink to their container and long variant labels wrap. The running harness reports 97 types and 545 permutation cards; the stressed `?stress=1` pass has no horizontally overflowing fixture cells at 390 px. This establishes narrow renderer-fixture coverage, not real result delivery or Canvas/Review acceptance. - A visual pass over the live renderer gallery found two misleading common-card states: a test suite with no tests claimed 100% success, and a suite with unnamed outcomes claimed 0% success. Both now report unavailable results without a green success cue. Small non-option tables also say “Table” instead of “Options.” The browser gallery showed the corrected states and client typecheck passed. The remaining packs still need the real-result visual audit described above. - The built-in metric pack's selected tree exposed a content-loss case: `emit_metric_card` accepts several named readings, but a `metric` primitive without `value` showed a dash. The client now turns that tree into its existing metric grid view, and constrained Markdown delivery includes every reading. A ledger → presentation-selection → projection test verifies the selected `seed.metric` tree retains condition, temperature, and humidity and lowers all three; the live browser gallery renders all three from that selected-tree shape. This is a focused real emit/projection proof plus a client render check, not a claim that all 75 packs have completed a real-model visual journey. - The all-built-in conformance check now rejects a selected pack whose real emitter payload lowers to empty text. It initially found 32 of 75 packs (small tables, universal shapes, and terminal) that compiled rich yet disappeared on constrained surfaces. The generic primitive lowerer now prints the bound fields when no specialized text was produced; the 75-pack check passes. This establishes nonempty constrained delivery for one valid emitter shape per pack, not visual acceptance of all shapes or full provider journeys. - [ ] Verify real workflows: everyday question, noncoding plan, document/data result, coding change, live canvas revision, draft review/acceptance, two humans in one conversation, participant revocation, and narrow-screen/keyboard/voice operation. - The running conversation shell at 390 px settles with its sidebar off-canvas and the empty state centered; the first screenshot during the 160 ms close animation briefly showed the sidebar over the content. A separate unauthenticated dev tab exposed repeated identical session-expired notices covering the phone composer. Notice insertion now coalesces identical active messages. Authentication and this empty-state check alone do not establish a completed narrow-screen workflow; the signed-in result, Canvas and Review pass follows below. - A fresh 3.5.1 dev binary on loopback reopened the real `isolated-review.html` conversation and frozen Version 2 at 390 px. The settled result fits the phone width. Review originally gave its four-row scope summary almost the whole dialog and left a 54 px file-comparison window, while the workspace header overlaid its own heading. Phone Review now uses one full-height scroll surface with a pinned heading and decision footer; the saved file comparison, attributed comment, feedback box and exact selected-file decision are reachable by scrolling. Canvas opens from that saved candidate using the authenticated, hash-checked file route and returns to the same Version 2 review. The preview is genuinely blank because this draft has an unclosed `` tag, which swallows its body; the collaborator's line-one comment already calls out that defect. Canvas now explains an empty parsed HTML body and directs the user to Code or revision, where the malformed source is visible. No acceptance or Agent revision was triggered in this visual pass. The broader signed-in workflow matrix remains open. - A fresh isolated 3.5.1 dev server on port 8918 ran a real Google Agent request for `activity.csv`. Its brokered write created the requested three rows, and its read receipt confirmed the bytes. It emitted a metric card showing 60 minutes, then hit the intentionally low four-step limit before a final prose answer. The original settled conversation showed only the user prompt: projection had required a final assistant document and discarded the saved artifact, card and error. The rebuilt client now projects a settled turn when it has an observed artifact, structured card or failure; the result displays the CSV and 60-minute card with a plain step-limit notice. A second omission hid the file action when there was no final result ID; the renderer now uses the observed artifact as its action anchor. In the running signed-in UI, `Open working file` opens the actual `activity.csv` as a three-row, two-column Canvas table with a source tab. This proves a partial real data result and inspection, not a completed final answer, frozen draft review, or acceptance journey. - A capped run now keeps its saved partial work and offers `Continue this task?` in the same Agent conversation. The new turn is explicitly a `follow_up`; it asks the Agent to inspect saved files and receipts, do only outstanding work, and read actual data before calculating. Repeating the original imperative had caused the model to write the file again and then hit the cap. The stop gate may count an earlier successful file write only when the new intent continues the same thread, the prior run stopped at `max_turns`, the workspace file still has the exact bytes logged in the append-only write receipt, and the new turn successfully inspects work. A changed file, different thread, or later completed run cannot borrow that receipt. Focused tests cover these boundaries; a live end-to-end completion after this fix remains to be recorded. The turn cap itself remains bounded and does not silently increase. A fresh `gpt-4.1-mini` run on the dev server wrote `continuation-proof.csv` with 20, 30 and 10, then its same-batch read raced the write and failed; the file was subsequently present with exactly the logged bytes. The Agent now serializes same-path read/write or edit batches in issue order while unrelated paths retain parallel execution. A follow-up read the saved CSV successfully, but the provider repeatedly asked for that same read until the duplicate-call approval guard denied the third request. The Continue prompt now explicitly asks it to use one successful read. A second fresh bounded `gpt-4.1-mini` journey on the rebuilt dev server created `continuation-clean.csv`, read the actual saved 20/30/10 rows in issue order, stopped at the one-step cap, and offered the saved artifact. An explicit follow-up in the same session read the file, emitted a metric of 60 minutes, answered in prose, and finished with both `primary deliverable: Produced` and `completed` receipts. The first follow-up wording accidentally asked for “evidence,” which raised a cited-evidence requirement and marked an otherwise completed answer partial; removing that word in the final wording produced a succeeded answer and completed run. This verifies the real provider, tools, session, projection and API path; the browser presentation of the completed journey was also inspected. The wider four-journey matrix remains open. - Canvas revision now starts a real routed Agent turn in the same conversation, with the selected result as its correction target; a Diff comment uses the same route. The live `continuation-clean.csv` revision added 50 after the earlier 40, read the saved CSV, and completed with the correct 150-minute answer and 23-byte artifact. The first revision exposed redundant writes and reads; the prompt now tells the Agent to inspect saved work before writing, and repeated read calls receive a bounded denial after the fifth identical call while effectful calls retain human approval at the third. Switching Agents exposed two independent rehydration faults: historical Agent headers without the later `character` field could not be read, and the switcher could select an empty default conversation instead of the last viewed task. Header reads now accept the absent historical field while new headers keep writing it; the switcher remembers each Agent’s selected conversation, and transcript loading precedes long-lived event connections. Live browser verification on the fresh UI preview switched from the 150-minute Vak result to Research Analyst’s saved market answer and back, with the correct URL and content after reload. - Development builds now activate all 75 built-in presentation seeds in the effective in-memory library for real result projection and the presentation list. User and workspace activations still take precedence, and the durable store remains unchanged; release builds retain explicit activation. A seed can render only when the result has a compatible validated semantic payload. The existing CSV final answer was prose plus an artifact, so enabling recipes alone does not invent a metric card. Tavily and other tool text is evidence in task details; presentation cards should show only validated user-facing results. Artifact promotion still needs an explicit distinction between user deliverables and files used only to answer internally. - 2026-09-23 live voice evidence: a fresh 3.5.1 dev server in an isolated workspace used the existing encrypted Shared Gemini credential only in process memory. A generated 16 kHz WAV sent through Vak's `/voice/transcribe` returned “Hello, what is 2+2?” with discovered `gemini-2.5-flash-lite` in 4.43 seconds. Vak's `/voice/speak` using discovered `gemini-2.5-flash-preview-tts` returned a valid 24 kHz, 3.01 second WAV in 4.6 seconds; sending that WAV back through `/voice/transcribe` returned “two plus two is four.” The Settings provider row now checks the canonical credential scope, and local backend readiness shows its actual probe result. Gemini Live synthesis on this host timed out after 25 seconds with both `gemini-3.8-live` and `gemini-2.5-flash-native-audio-latest`; the working batch TTS route now dispatches by TTS model class. A physical microphone/browser permission round trip and the Live socket failure diagnosis remain open, so voice UI completion is not claimed from this API test. - Follow-up on the Live timeout: a direct WebSocket probe showed Google sends `setupComplete` as a **binary frame containing JSON**. Vak had ignored binary frames, so the timeout was a protocol parser error. The adapter now reads both text and binary JSON frames. The same live `/voice/speak` request returned a 2.12 second WAV in 5.13 seconds with `gemini-2.5-flash-native-audio-latest`. Its re-transcription did not faithfully match the supplied text, while the batch TTS re-transcription did; batch TTS remains the verified choice for faithful short replies. In the in-app browser, voice activation remained pending at microphone capture; the control now reveals its “Allow microphone” phase instead of showing “Connecting” throughout. A real microphone-to-Agent-to-playback UI run remains open. - Browser entry exposed two more real faults. `createScriptProcessor(320)` was invalid under Web Audio; capture now uses a legal 512-frame buffer and releases the stream if setup fails. The browser then reached Listening, but chose `SpeechRecognition` over Vak's configured provider and failed. That competing branch was removed so captured PCM always enters the governed WebSocket. Finally, raw PCM sent as `audio/pcm` to Gemini produced a plausible but wrong transcript for a known WAV utterance. The WebSocket now wraps 16 kHz mono PCM as WAV before calling Gemini, using the shared audio utility. The browser has reached Listening and Processing with a physical microphone; a verified Agent reply and spoken playback are still open. - The next UI round trip exposed a control-frame mismatch: the client sent `type` while the server's versioned protocol expects `t`. The client control type and calls now agree with the server, and a browser voice utterance reached the Agent conversation. The machine's speaker-to-microphone pickup produced inaccurate transcripts, so that physical audio source is not acceptable evidence of recognition accuracy. The Agent dispatch then failed because the Google function declaration API rejected JSON Schema `type` arrays nested in presentation tools. Google's API accepted the equivalent `anyOf` schema in a live call; Vak now converts these unions at the provider boundary, with a focused regression test. A clean browser microphone-to-accurate-transcript-to-answer-to-playback proof remains open. - With those fixes in a fresh dev build, a controlled 16 kHz PCM utterance sent through the actual `/voice/session` WebSocket produced the final transcript “Hello, what is 2 plus 2?” and a completed governed Agent reply answering 4. A separate typed turn in the running UI also answered 4, verifying that Google accepted the full current tool declarations. Browser microphone capture reaches Listening and dispatches turns, but the machine's speaker-to-microphone pickup generated inaccurate words; physical spoken-input accuracy and browser playback of a completed answer remain open. The working batch TTS endpoint has been independently verified above. - Playback preparation now uses user-facing prose rather than the raw Agent completion. A live completion containing a metric heading, JSON fence and model-authored conversation block produced “2 plus 2: 4 The answer is 4.” for speech; a card-only result uses a short review cue while preserving the actual card on screen. Stopping voice now closes capture without sending a final `speech_stopped` frame, so stopping during an unfinished utterance does not submit it as a new Agent turn. The web build and typecheck pass. This is content and stop-behavior verification, not a claim that audible browser playback was observed. - Gemini synthesis without a model now returns an immediate, actionable configuration error instead of attempting a Live socket with an empty model id. The selected model still comes from the effective voice route; no model catalogue was hardcoded. - Playback receipts now reflect observed audio emission: `/voice/speak` no longer appends a synthetic zero-millisecond success, and the browser sends a completion receipt when its AudioBuffer ends (or an interruption receipt with elapsed milliseconds when stopped). A fresh authenticated voice WebSocket session recorded an actual `VoicePlayback` activity with `emitted_ms=1260` and `interrupted=false`. Audible browser playback is still a separate acceptance check. - The browser now unlocks its Web Audio output context during the explicit Voice click and checks that it reaches `running` before starting a synthesized answer. If the browser blocks output after the provider reply, the text answer remains and the user sees a playback error instead of a false Speaking state. The production web build and typecheck pass. The dev UI on port 8914 currently has voice disabled and no provider credential; this change does not claim an audible browser round trip. - A fresh isolated dev workspace on port 8915 used the existing encrypted Shared OpenAI credential and models discovered for that key: `gpt-4o-mini` for the Agent, `gpt-4o-mini-transcribe` for input, and `gpt-4o-mini-tts` for output. A controlled WAV round trip through the real HTTP voice routes produced a transcript and a 187,244-byte, 3.9-second spoken answer. The browser now also accepts a recording as a real voice input path: it decodes and resamples the file to 16 kHz PCM, sends it through the same governed voice WebSocket, and plays the answer through Web Audio. In the running browser, the recorded question produced “Hello, VAC. What is two plus two?”, the Agent answered “Hello! Two plus two is four.”, and the append-only session recorded `voice_playback` with `emitted_ms=2712` and `interrupted=false` after the AudioBuffer ended. This closes browser audio emission for a controlled recording; acoustic confirmation by a listener and accurate physical-microphone recognition remain open. The live run also exposed and fixed OpenAI's rejection of the MCP tool's top-level `oneOf`, missing WAV framing on WebSocket transcription, the provider name mismatch in Settings, and empty noise transcripts that formerly ended voice sessions. No model catalogue was hardcoded and the configuration was confined to the isolated workspace. - Microphone onset now buffers candidate frames until at least 240 ms of voiced audio is present, then sends the buffered beginning through the same voice socket. A short click/noise burst is discarded after 160 ms of silence instead of leaving an open utterance that never reaches the stop threshold. The production web build and typecheck pass; this specific onset change still needs a physical-microphone browser check before claiming real-room accuracy. - A real browser microphone captured the short speaker-played question as “What is two plus two?” and the governed Agent produced a text answer. The original 650 ms silence tail split a longer sentence across two turns, so capture now allows a 1.5 second natural pause. The isolated test workspace's intentional 1 MiB session cap was reached during prolonged open-mic testing; it was raised to 4 MiB for this test only, with the four-requests-per-minute guard retained. WebSocket tracing then exposed a race in which an older empty noise transcript reopened capture during a newer turn. The client now binds the in-flight input phase to the exact utterance ID, ignores new mic frames while that turn is being prepared, and retains spoken interruption during playback. The production web build and typecheck pass. This room remained noisy: unrelated speech was picked up and dispatched, and no new `voice_playback` receipt was produced by these microphone runs. A quiet-room listener test remains required before claiming the physical microphone-to-audible-answer journey is complete. - **Deferred at the owner's request.** Keep the physical microphone-to-audible-answer check open for a future voice refinement pass. The next pass should use a quiet room or headset, inspect exact utterance-scoped WebSocket control frames and the append-only `voice_playback` receipt, verify that ambient noise cannot dispatch unrelated turns, and confirm interruption with a listener. The controlled recording/browser emission proof above remains valid; it does not substitute for this human-heard check. Continue the shared-workflow and responsive acceptance work now. - 2026-09-23 voice refinement pass (code; no human-heard run yet). The noisy-room dispatch had a structural cause: the only speech gate was a fixed client loudness threshold, and a noisy room clears any fixed threshold. The server now measures every closed utterance against its own quiet floor before it may reach a provider (`vak_voice::vad::SpeechEvidence`); flat noise of any loudness is answered with `discarded: insufficient_speech`. The browser detector shares those constants, calibrates and adapts its floor, learns the room from noise-only utterances, caps an utterance at 30 s, and needs twice the level and duration to interrupt a playing answer, so the answer's own echo does not barge in. The socket protocol now has separate client and server frame types, so a client can no longer author a transcript that starts a turn (that second path had outlived the removed Web Speech branch); utterance ids are strict and never reused; playback receipts now name the utterance whose answer was played instead of an invented id; and error frames use the protocol's `t` tag (the server had been sending `type`). Transcription routing had two diverging copies (the socket accepted `local`, the batch route did not; they defaulted to different provider names); there is now one route over one provider enum (`gemini`, `openai`, `local`), one per-minute budget covering socket utterances too, and no legacy `voice.model` fallback or inert `realtime_model` setting. Settings could not clear a voice provider or model because `null` deserialized as "unchanged"; it now clears. The Pause and Stop audio buttons never enabled while speaking because they read a non-reactive variable; they now track playback. Doctor reports enabled voice without a provider, credential or model pin as a failure with its remedy. Evidence: `crates/vak-server/tests/voice_session.rs` drives a real WebSocket against the real router with a scripted transcriber and model: noise is discarded without reaching the transcriber (a mutation run with the gate removed fails this test), speech becomes one Agent turn whose answer returns on the socket, playback lands in the ledger, a forged client transcript is refused and never logged, and a cross-origin upgrade is refused. Unit tests cover the evidence measure, protocol direction, provider parsing, config patching and the local engines. A voice session with no usable provider route is now refused at connect, before the microphone prompt, with the remedy in the notice; local readiness reports the transcriber and speech engines separately (it previously described only TTS, so a working transcriber read as "not configured"); `/config` rejects unknown provider names. Live UI check on a disposable isolated 3.5.1 server (scratch HOME and data home, local route with a scripted transcriber, no real credentials): Settings shows the three canonical providers with accurate readiness, no Realtime model, and a free-text voice id; enabling voice and choosing `local` persisted, and choosing Inherit cleared the provider; pressing Voice with no provider showed "Voice unavailable", the "Choose a voice provider in Voice settings" notice and Type instead; with `local` selected the socket reached `ready` and the control waited at Allow microphone, where a person must grant access. Not verified: a physical microphone in a real room, audible playback heard by a person, and the adaptive client detector against real speech — the quiet-room/headset listener check above is still the acceptance gate. Committed as `66e20b68`. - Channel voice follow-up. Telegram, Discord and Slack bridges transcribed each voice note through `/voice/transcribe` before `/gateway/inbound`, so a chat still awaiting operator approval could spend a paid provider call, and the model was then told `[audio attachment received; transcription provider is not configured]` even when transcription had succeeded. Per-bot/per-chat voice provider and model pins were stored but never applied. Bridges now only attach the audio; the gateway transcribes it after allowlist admission and request de-duplication, through the chat's workspace ← bot ← chat voice tiers, under the shared per-minute budget. The transcript joins the typed text as the user's message, an untranscribable note reaches the model as `[voice note not transcribed: <reason>]`, spoken words never act as control commands, and each heard note is a ledger transcript activity. `/voice/speak` resolves the chat from the session's gateway binding, so channel replies use that chat's route, voice and persona; the admin Preview now auditions the pinned provider too. `/voice/transcribe` and the unused bot-persona resolver are removed. Bot and chat voice tiers are validated when written, so an unknown provider name is refused at the admin console instead of failing a later voice note. Evidence: `crates/vak-server/tests/voice_channels.rs` posts a voice note through the real gateway router with a scripted transcriber and model — a pending chat's note is refused without reaching the transcriber, and an admitted chat's note reaches the model as "what is two plus two" with no false provider statement; unit tests cover the tier overlay and the model-facing note text. Not verified: a live voice note through a real Telegram/Discord/Slack bot. ### Anchor check after the visual refresh (2026-09-26) Compared in the running web app (dev build of `main` after V3.13, served from `/tmp/vak-screen1-live` with the real review data, and a fresh home for first run) against the four references above and the three doc 75 mockups. Screenshots: `docs/assets/visual-refresh-2026/after/V3.10-*` (1440 × 900 light and dark, 390 × 844 for four screens; no sideways scroll on any). The references are compared for relationships and interaction, not for their example text or the pre-refresh palette. | Screen | Holds | Gaps | |---|---|---| | 1. Everyday result | The answer as text first; result actions (Open, Review changes, Ask for a change); no workbench chatter | The result card is a file chip ("Draft · 146 bytes"), not the reference's card with a preview, a draft status and one primary Review changes; the newest result's actions omit Review changes; a revision request from before V1.12 still shows its raw ids; the header names the Agent by the identity frozen in that conversation ("Vak"), while the sidebar shows today's name | | 2. Everyday plan | The options are a result, each with Use this bound to its owning result | A plain comparison table, not the reference's plan-plus-options composition; this saved conversation ends in a failed follow-up turn | | 3. Review and accept | One sheet: preview, then changes, then the decision (goes to, saved copy, checks), Accept selected and Keep as draft | Single column rather than a decision panel beside the changes; this draft is an HTML file with no automatic check, so the observed-checks row the reference shows has nothing to report | | 4. Conversation and canvas | Conversation beside the draft, Draft preview with the version, device switcher, Work on this together, comments, Review changes | None at the level of relationships; the review draft itself renders blank (its unclosed `<title>`) | Follow-up, 2026-09-26: V3.14 closed screen 1's result card gap. A file result is now the reference's card: a preview, the draft status in words from the server's records ("Draft, version 2, waiting for your review"), Review changes as the one primary action on the newest waiting draft, then Open and Ask for changes, with no byte count (`docs/assets/visual-refresh-2026/after/V3.14-*`). The review conversation's newest result is a file the old revision path wrote straight to the folder, so it truthfully reads "Saved in your folder" and offers no review. V3.15 then put the Agent's character beside its name on its Settings page and in the Agents navigation (`docs/assets/visual-refresh-2026/after/V3.15-*`). V3.16 closed screen 2's gap: a plan step carries typed `options`, and a plan whose step offers them is drawn with the options card beside it, each option with Use this bound to the result; checked with real luna plans (`docs/assets/visual-refresh-2026/after/V3.16-*`). Doc 75's three screens: first run matches its mockup (greeting, Connect card in the greeting, four starters with examples; the header status is a chip rather than a line under the name). Settings has the Everyday, Agents and Advanced structure and names the Agent once; the mockup's Agent page also carried the permissions and personality, which V3.5 placed under Privacy and safety and Your agents, and its Agent character beside the name and in the navigation is missing. The conversation mockup's result card is the gap listed under screen 1. ## Completion bar ### Real everyday presentation check (2026-09-23) Three real Agent turns on the isolated dev server exercised a Saturday-morning timeline, a rainy-day outing comparison, and a vegetarian breakfast recipe. The model called the built-in card tools and the projection produced `plan.timeline`, `table` plus `research.synthesis`, and `recipe.card` respectively. The timeline turn exposed a delivery flaw: two intermediate answer drafts and two timeline revisions were shown alongside a final prose answer. The everyday conversation now shows only the final answer, carries the latest card of each semantic type to that result, and presents the card first. The model's final narration is available as a collapsed Agent note when a card is present. This was verified after rebuilding and restarting the real dev server at port 8918 on the original timeline session; one timeline card appears and no earlier-drafts disclosure appears. The card/narration output contract now makes the card authoritative after a successful card call. Final text is visible beside it only when explicitly marked `Note:` or `Additional note:`; the complete model text remains in the append-only ledger. The same parser controls browser and channel projection. A fresh real Agent run (`01a0ce62-ee9b-78c1-afd8-2692c91312aa`) still wrote repeated unmarked prose despite the stronger prompt, but the reloaded browser showed one complete timeline card and no duplicate schedule or draft disclosure. The comparison session still showed its table and recommendation cards together. Focused tests cover channel suppression, explicit note delivery, and note metadata projection. The next development pass wired emitted card ledger entries through active presentation definitions. Four complete seed mappings are certified for live output: timeline, table, recipe, and research synthesis. The other 71 definitions remain available in the development catalog but cannot displace real card output until their data mapping is complete. The selected tree retains the card's original fallback text and provenance, and historical session projection reloads the effective library instead of dropping the selected pack. The real saved sessions above now project `seed.timeline` with four steps, `seed.table` with two rows, `seed.research-brief` with takeaways and sources, and `seed.recipe` with ingredients and directions. The development UI rendered all four after server restart; recipe controls and the timeline text were visible. Focused seed tests assert content in the compiled children, and the existing projection tests still pass. This verifies those four definitions through persisted real output; it does not certify the remaining 71 definitions or every presentation combination. An everyday UI review then removed the adaptive card's `Show original` toggle and successful-result receipt/check counts from the conversation. Those remain in the result data and explicit Details surfaces. Small tables render as quiet Options with plain column labels; dataset search, sorting affordances, row counts and CSV export stay available for larger data tables. This preserves those functions where they help with data work without putting them on a two-choice family comparison. The next pass promoted the other 71 built-in definitions from catalog-only previews to selectable defaults in every build. All 75 seed definitions now bind the complete validated payload of their emitted shape. Table, timeline, recipe, research, diff and test definitions expand their collections into render nodes; universal shapes retain all free-form fields for the nested-value renderer. A saved explicit user/workspace selection still wins. Conformance tests select all 75 from an empty library and compile each against its real emit-tool payload shape. This proves selection and schema completeness; live human review has so far covered the four everyday examples above, not every one of the 75 visual combinations. The pack library remains user-controlled. Settings has scoped search, grouping, activation/deactivation, and import/export; retired built-in revisions are retained for old results but hidden from the everyday catalog unless selected. Deactivation now persists a scoped suppression, so a built-in default stays off after reload until explicitly activated or reset. Bulk activation can reach all built-ins even before the catalog is opened; bulk deactivation covers them all. Imported definitions are disabled previews rebased to the local user library, making subsequent activation possible without granting it during import. A real dev API round trip deactivated and reactivated `seed.metric`; the catalog showed 75 active packs. The universal family uses one complete nested-value renderer with a pack-specific label rather than an arbitrary primitive that could render an empty specialized view. The current add path is local JSON pack import; a hosted discovery marketplace is still open work. A selected travel-options pack exposed an interaction loss: its compiled table had no choice action, although the ordinary structured result offered one. The adaptive timeline now passes the conversation and result ids into the selected tree and shares the same owner-choice handler with the structured renderer. `seed.travel-options` renders its table as options, so choosing a row prepares a targeted reply in the same conversation. The UI harness displays the selected-tree shape and both buttons; the TypeScript check and desktop/web production builds pass. This is renderer and interaction coverage, not a live Agent travel turn. A string literal in a pack spec could not be used for the table variant because the current internally tagged `SpecValue` newtype cannot serialize that value; the client projects this built-in pack's known variant until the spec value contract is repaired. The pack value contract now serializes static text, number, boolean and JSON literal values with an explicit `value` field; binding values retain their existing wire shape. The travel-options seed uses that contract to carry `variant: options` in its own definition (revision 7). Earlier saved revision 6 trees retain the client-side projection so their choices remain usable. A focused spec round trip and a built-in travel pack serialize/parse/compile check pass. The audit scenario suite now checks current immutable seed revisions and leaves all-pack compilation to the server projection conformance test, which supplies each emitter's actual payload shape; both suites pass. The next live travel run exposed three everyday path gaps. A model-authored `table` card with an Option column was not treated as a selectable choice, a card without a separate result receipt did not get a reply target, and ⌘N reopened an existing Agent conversation. Small option/choice tables now expose row actions, cards fall back to their stable timeline item id for reply routing, and the Agent-open API accepts an explicit new-conversation request with a distinct durable ledger identity. The focused server test for new Agent conversations passes. In a real Agent run on the fresh development build, the Saturday outing result showed Museum and Garden buttons; clicking Museum prepared `Use “Museum” in this plan.` with `Replying to option “Museum”`. Sending that follow-up from a separate preview was refused because an older dev server held the saved conversation lease, so that particular answer round trip is not claimed. A fresh ⌘N run produced a separate conversation and a selectable two-row options card. The presentation tool descriptions now tell the Agent that general cards are static and name the table contract for selectable choices. The same runs showed streamed assistant drafts briefly appearing and then disappearing, and the user message was remounted when the settled presentation arrived. The everyday chat now keeps assistant drafts out of the active turn, shows one working indicator until settlement, keeps the user message mounted, and inserts the final projection in the same turn. A fresh live run displayed only the user request and working state during generation, then the final two-choice card. TypeScript, the production web build, the focused Agent conversation test, and Rust server/core checks pass. Raw streaming and receipts remain available in task details; this review does not yet measure layout stability across every media type and long conversation. This work is complete only when the running desktop and web UI can demonstrate the four journeys with real data and runtime evidence, the same Agent conversation and draft can safely include an invited human, and technical capability remains reachable without dominating everyday use. Update this ledger with tested evidence and screenshots of the **running implementation** as each screen lands. A build passing or a generated mockup alone does not close a checkbox for a screen.