04 · Flow · Voice-first personal computing
Voice that can actually change your day.
Tell Flow you are running late, need to renew your passport, or promised Maya a proposal. It does not open a chatbot. It changes the product — and leaves the result there for you to inspect, edit or undo.
Built with
- React 19
- TypeScript 5.9
- Vite 7
- Tailwind CSS
- Motion
- Zod
- Web Speech API
- Playwright
My scope
Product design, React frontend, voice runtime & deterministic action systems
Product surface
Calendar · Journal · Friends · Memories · Plans · shared voice runtime

The hard part
The hard part is not speech recognition.
The hard part is not turning speech into text. It is letting a sentence change real state without making the product unpredictable. If a target is ambiguous, Flow has to ask. If dinner is protected, it has to stay put. If the result is wrong, undo has to restore the whole transaction.
One request, end to end
“I’m 35 minutes behind. Keep dinner.”
Flow turns that sentence into constraints, works out what can move, protects the dinner anchor, validates the changed schedule, commits one transaction and leaves the calendar editable.
- 01
Hear
Speech becomes input, not authority.
- 02
Resolve
Use route, selection, recent references and the current time scope.
- 03
Clarify
Ask when the target or consequence is genuinely ambiguous.
- 04
Protect
Confirm consequential changes and preserve explicit constraints.
- 05
Transact
Apply typed actions to a draft and validate the result before commit.
- 06
Render
Show the same state in Calendar, Journal, Friends or Memories.
- 07
Recover
Correct, interrupt, undo or redo without inventing a second source of truth.
Behind the interface
Architecture & tools.
Speech is only the front door. A contextual resolver interprets the request against the current route and recent references; the command controller turns that into typed actions; the transaction layer clones, applies and validates shared state before commit; React renders that same state back as ordinary editable UI.
- Interface
React 19 · TypeScript 5.9 · Tailwind CSS · Motion
Render responsive product surfaces and keep voice-created state directly editable with ordinary controls.
- Voice & intent
Web Speech API · contextual intent resolver · conversation context
Turn a transcript into a ranked intent using route, selection, recent references and pending conversational state.
- State & safety
Zod · typed LifeAction transactions · shared LifeDocument · local persistence
Apply changes to a draft, validate invariants, preserve stable identity and commit one deterministic state transition.
- Verification
Vitest · Testing Library · Playwright · command traces
Exercise parser behavior, state invariants, undo/redo, cross-surface journeys and the exact actions produced by a command.
- 01
Hear
Capture a transcript without giving the recognizer authority to mutate product state.
- 02
Resolve
Rank intent with route, selection, recent targets, pending context and the active time scope.
- 03
Transact
Apply typed actions to a draft, validate invariants and stop on clarification, confirmation or conflict.
- 04
Keep editing
Commit one shared document so voice and direct manipulation operate on the same objects.
Ambiguity can stop before mutation, protected calendar anchors remain fixed, command traces record the chosen intent and actual actions, and undo/redo restore complete transactions. Automated recognition verifies the application path; physical microphone acceptance is tracked separately.
Decisions that shaped the product
I built the path from conversation to typed product action: contextual intent resolution, clarification and confirmation, a shared LifeDocument, transaction validation, undo/redo, and the React surfaces that keep every result directly editable. Calendar, Journal, Friends and Memories all operate on that same model.
Speech never writes directly to state.
A transcript is resolved against the current route, selected entity, recent references and time scope. The command controller turns the result into typed actions; only the transaction layer is allowed to commit them.
The tradeoff. There is more machinery than a direct voice-to-handler shortcut, but interpretation mistakes cannot silently become arbitrary state mutations.
Ambiguity is a product state, not a parser failure.
If Flow cannot identify the right event or the action is consequential, it can stop at clarification or confirmation before anything changes. Constraints such as protected time stay part of the transaction.
The tradeoff. Some requests take one extra turn. That is preferable to a confident-looking interface that changed the wrong thing.
Undo restores the transaction, not the animation.
Voice and direct manipulation operate on the same shared document. Undo and redo restore domain history while presentation remains disposable, so motion never becomes a second source of truth.
The tradeoff. History, interruption and stale presentation need their own tests, but recovery stays reliable across surfaces.
Verification, not theatre
The repo is built like a product, not a voice demo.
The interesting failure modes are not visual. They are stale context, the wrong target, a partial transaction, a duplicate command, or an undo that only fixes the screen. Flow has explicit tests around those boundaries.
- 120
- unit & integration test files
- State, parsing, context, recovery and feature behavior in the current Flow repository.
- 32
- Playwright journey specs
- Cross-surface behavior, responsive states, motion, recovery and voice journeys.
- 14
- commands in one voice acceptance journey
- One continuous automated recognition session from navigation through edits, commitments, undo, redo and pause.
Those numbers describe the current repository structure. The automated voice journey injects transcripts into the production action pipeline; physical microphone quality is a separate acceptance gate.
The outcome
Flow treats voice as another way to operate the product, not as a separate assistant mode. An intention can become a plan, one step can become scheduled time, a promise can stay attached to the right person, and the user can keep editing everything with ordinary controls.
The portfolio film drives the real action pipeline with deterministic recognition. That proves the application path, not microphone accuracy or Chrome speech-service quality. The repository keeps physical microphone acceptance as a separate release gate.
Next case study
Mirror AI