← Back to work

Blitzit / Work story / 23 min read

Blitzit 3.0, from the backend up.

I designed and built Blitzit 3.0’s entire backend from scratch, then built the AI and integrations on top of it. One foundation for tasks, shared work, connected tools, and an assistant that can act.

The Blitzit 3.0 workspace. Original interface, captured locally with sample tasks and names.

The Blitzit 3.0 workspace

Original interface, captured locally with sample tasks and names.

A task app seems straightforward until the same task can be changed by a person, a connected service, and an AI assistant. Now a checkbox is part of a larger system. Who is allowed to change it? Which copy is current? What happens if the external service is down? Can the user undo what the assistant just did? These are the questions behind much of my work at Blitzit.

I designed and built the entire Blitzit 3.0 backend from scratch. That includes the initial server, database models, authentication, domain modules, and event bus, then the systems for sync, collaboration, integrations, billing, migration, and AI. My earlier work was on the existing product. For 3.0, I had responsibility for creating the foundation those features would share.

This story starts with that backend architecture, then follows the connected tools and the voice assistant built on top of it. The voice backend is merged; the complete frontend voice experience is in review as of October 2026.

My role

I owned the backend from the first Fastify server and database models to the modular architecture, typed event bus, API and worker roles, authentication, billing, and migration. That foundation now connects 12 providers, shared AI tools, and realtime collaboration. My latest work brings it into Blitzy’s voice experience, custom wake word, and interactive orb.

The whole 3.0 backend. From a blank repo.

A shared foundation for the product: accounts, tasks, collaboration, billing, integrations, and AI, with clear boundaries as each part grows.

I designed and built the backend in Fastify and TypeScript, with domain modules separating routes, validation, and business logic. MongoDB holds product state; a typed event bus connects side effects without putting every concern inside a request handler. Redis and BullMQ support queued work, while separate API and worker roles let background processing run apart from public requests.

I started the 3.0 backend in March 2026. The first commit set up Fastify, TypeScript, MongoDB, Redis, and the event bus. Database models and authentication followed, then the core task and list APIs, client sync, recurrence, and the integration SDK. I was building the backend itself, including the rules that later features would depend on.

I organized the code around product domains: auth, tasks, lists, sharing, billing, scheduling, notifications, integrations, and AI. Each module keeps its routes, validation, and business logic close together. The server assembles those modules with shared plugins for the database, authentication, rate limits, and observability. That makes a change easier to place and review without turning every endpoint into a mixture of unrelated responsibilities.

The event bus connects those responsibilities. When a task changes, other parts of the product may need to sync a provider, update a connected client, reconcile a reminder, or refresh the assistant’s context. The task service should not need to understand every one of those systems. It publishes a typed domain event, and the relevant listeners decide what to do with it. Event types define the payload each listener receives; task events also carry attribution so downstream systems know who caused the change.

For example, task creation saves the task, records its change history and shared-list activity, then emits task.created. The integration listener checks whether the task belongs to a connected provider and whether that connection is usable before arranging outbound sync. Changes that came from an integration are filtered to avoid a loop where a webhook triggers a sync that triggers another webhook. This is where the architecture meets ordinary product behavior: one task should not become an endless chain of repeated updates.

The bus is an in-process event dispatcher. It is not a durable queue or a second database, and I do not treat it as a guarantee that work survives a process crash. MongoDB stores the product state. BullMQ carries queued background work. Socket.IO’s Redis adapter handles cross-process realtime broadcasts. Keeping those jobs separate matters because broadcasting every domain event to every process would cause listeners to repeat side effects. Catching and reporting listener errors helps keep failures visible, but it does not turn an in-memory event into a persisted job.

As background work grew, I split the runtime into API and worker roles. Queue producers and local event listeners remain available in both roles, so an API request can still enqueue work. Consumers and schedulers run in the worker role. That distinction is easy to get wrong: disable a producer and the request can succeed while the follow-up work never starts. I centralized startup and shutdown for this work so the split can be reviewed in one place.

The result is a backend that can support features through the same foundations. Authentication and permissions still matter when the caller is an AI agent. A scheduled change still needs to reach connected clients. A provider update still needs the right source attribution. Building those foundations myself made it possible to carry the same rules into the integration framework, the change journal, and eventually Blitzy’s voice experience.

Talk naturally. Get real work done.

Spoken requests can use the same context and capabilities as typed chat, with the same checks before changing anything.

I separated speech from execution. The realtime voice model handles conversation; the existing reasoning agent receives the final transcript, uses scoped tools, and returns a grounded outcome. I built interruption handling, fresh context per turn, and coordination so acknowledgements and results do not talk over each other.

Blitzy’s real canvas orb. The original orb component rendered locally in a presentation wrapper. Listening state is set for this preview; this is not a recording of a live voice session.

Blitzy’s real canvas orb

The original orb component rendered locally in a presentation wrapper. Listening state is set for this preview; this is not a recording of a live voice session.

“Hey Blitzy.” From wake word to working companion.

A hands-free entry point, with an orb that makes listening, thinking, and acting visible.

I trained a custom wake-word classifier in Colab and integrated local ONNX inference. A short transcription check filters candidate detections before starting a conversation. I built the microphone handoff and session lifecycle, plus an orb driven by actual audio and agent state. It pauses offscreen and respects reduced motion.

A voice companion you can configure. Wake-word, sound and orb appearance controls in the frontend voice implementation. This experience is in review.

A voice companion you can configure

Wake-word, sound and orb appearance controls in the frontend voice implementation. This experience is in review.

Your tools, working together.

Tasks and changes can move between Blitzit and the services people already use, instead of becoming another isolated list.

I built a provider plugin contract for authentication, field mapping, discovery, and sync. Shared OAuth, webhook, queue, and retry infrastructure supports Asana, ClickUp, Notion, Trello, Todoist, TickTick, Linear, GitHub, and Google and Microsoft task/calendar providers. Provider errors retain their meaning, so reconnecting an account is not treated like a temporary outage.

An integration is easy to underestimate if you only look at its happy path. Fetching a task is not the same problem as keeping it correct after a rename, a revoked grant, a partial outage, or a change that comes back through a webhook. A shared provider contract gave me one place to define what a provider supports, while leaving its actual mapping and API behavior explicit.

I did not want each integration to become a small backend of its own. The point of shared infrastructure was to keep ordinary concerns such as retries, ownership, and authentication consistent. It also made differences easier to see. A calendar event and a task are not interchangeable just because both have a title and a date.

One home for connected providers. The real provider catalogue, captured locally. Demo accounts are not connected to external services.

One home for connected providers

The real provider catalogue, captured locally. Demo accounts are not connected to external services.

One set of rules for every AI client.

Blitzy and external AI clients can act on the product through a consistent permission model.

I built the MCP surface with OAuth 2.1 and PKCE, scopes, typed schemas, and read/write annotations. External clients and the in-app agent share tool definitions and the executor, instead of maintaining separate implementations that gradually disagree.

Tool count is useful as a measure of scope, but the more important decision was avoiding two sources of truth. If a tool is available to the in-app assistant and an external client, both should encounter the same validation and ownership checks. Maintaining parallel handlers would make it too easy for one surface to gain a capability without gaining the matching restriction.

The tool schema is only one part of that boundary. The executor still has to enforce it when a call arrives. That is especially important for read-only requests. Telling a model not to write is an instruction; preventing a write in the execution path is an actual product rule.

Context belongs to the work. List-level AI instructions give the assistant context for a particular workspace. The sample list is shared with fictional teammates.

Context belongs to the work

List-level AI instructions give the assistant context for a particular workspace. The sample list is shared with fictional teammates.

AI made a change? You can undo it.

Automation becomes easier to trust when a mistaken action has a clear way back.

I designed a per-user change journal. Human and AI mutations become reversible commits, with before/after state and content checks. Undo checks the current state before restoring anything, so it does not silently overwrite newer edits.

Undo changes the relationship between a person and an assistant. Without it, a direct AI action can feel like handing over control. With it, the person can inspect a result and recover from a mistake. But a useful undo system has to know which state it is restoring and whether something has changed since the original action.

That is why I approached the journal as a sequence of changes rather than a second set of CRUD buttons. Content checks matter because the user may have edited the same task after the assistant did. A blind restore could turn the recovery feature into another source of data loss. A conflict should be visible, not quietly resolved by deleting somebody’s newer work.

Shared work that stays in sync.

People can share lists, see activity, and work across clients with explicit access boundaries.

I built shared-list invitations and granular permissions, realtime delivery, and soft edit locks. Durable writes go through the API; local-first clients receive document updates and checkpoints, with recovery paths when a connection drops. Later work separated change-stream connections and shared database cursors across listeners to reduce contention.

Realtime delivery and durable state have different jobs. A live event makes another person’s edit appear quickly. It should not be the only record that the edit happened. Keeping durable writes in the API and supporting recovery for clients makes the system less dependent on a perfect connection.

The same distinction showed up in database work. A long-lived change stream can consume a connection for a very different amount of time than an ordinary query. Sharing cursors and separating connection pressure helps keep those two workloads from competing in ways that are hard to understand from the interface alone.

Finding work across lists. Search surfaces lists, active tasks and completed work from the local sample workspace.

Finding work across lists

Search surfaces lists, active tasks and completed work from the local sample workspace.

An assistant with useful memory.

The assistant can carry relevant context forward without treating every interaction as a fresh start.

I built memory and personalization machinery that separates remembered facts from derived activity signals. Extraction, deduplication, budgets, decay, and forgetting keep that context bounded. User controls matter as much as recall: remembered information needs a way to be inspected and removed.

Remembered facts and inferred patterns should not be mixed without thought. A fact the user explicitly asks the assistant to remember has a different meaning from a pattern derived from activity. I kept those responsibilities separate so that the system could decide what to include, what to age out, and what to forget.

More context is not automatically better context. It adds cost, can crowd out the current task, and can make an old preference look more certain than it is. Budgets, attribution, deduplication, and user controls are part of making memory useful rather than simply making it large.

Memory the user can inspect. A sample preference in Blitzy’s memory screen, with controls to pin, edit and remove it.

Memory the user can inspect

A sample preference in Blitzy’s memory screen, with controls to pin, edit and remove it.

The platform underneath the features.

Tasks, reminders, billing, and accounts need to keep working even when nobody is talking to an AI.

My work includes the Fastify backend, authentication and sessions, task/list APIs, server-side recurrence and timezone-aware scheduling, notification queues, Stripe and RevenueCat billing, and managed-team billing. These are product systems with failure and recovery paths, not just endpoints around a model.

The move from client-owned scheduling to server-owned scheduling was part of a broader change in responsibility. A recurring task should not depend on a particular client being awake at the right moment. Timezones, reminders, and retries belong in systems that can keep their own state and explain what they have done.

Billing has similarly specific boundaries. Managed-team billing does not mean every member shares the same workspace or data. Those are different product concepts, and the implementation has to preserve that difference. I try to be precise about those distinctions in both the code and the way I describe the work.

Tasks have a place in time. The weekly calendar and unscheduled tasks use the same underlying task records. Sample dates and tasks shown.

Tasks have a place in time

The weekly calendar and unscheduled tasks use the same underlying task records. Sample dates and tasks shown.

Activity without losing your place. The notification inbox alongside the task board. Entries use sample content.

Activity without losing your place

The notification inbox alongside the task board. Entries use sample content.

A new platform without starting users over.

Moving to Blitzit 3.0 should preserve the work people have already put into the product.

I built per-user migration from the legacy platform, with resumable imports, idempotent writes, and delta handling. Existing 3.0 edits take precedence where needed, and recurring task templates are handled separately from generated occurrences. Earlier work moved scheduling and recurrence responsibility from clients into the backend.

A rewrite is much easier when you pretend the existing product is empty. A real migration starts with people who already have lists, recurring work, and their own expectations. Imports need to be safe to resume. Running them again should not duplicate the user’s work, and a later sync should not overwrite an edit made on the new platform.

I treated the old platform as a source to read from, then put safeguards around how that data entered the new one. Recurring templates needed separate treatment from occurrences that had already been generated. These are small distinctions on paper, but they decide whether the new app feels like a continuation or a reset.

Failures that lead to the right next step.

A revoked account, a rate limit, and a provider outage should not all look like the same broken feature.

I worked on bounded retry decisions, integration diagnostics, notification delivery, instrumentation, and database connection pressure. I classify outcomes before retrying and keep durable state as the source of truth. The aim is to make recovery explainable as well as automatic.

Retries are not a substitute for understanding the failure. Retrying a revoked credential does not repair it. Retrying a throttled request immediately can make the problem worse. A provider timeout, an absent resource, and a permission failure need different handling, even if they all interrupt the same action on screen.

The useful engineering work is often making those differences survive the trip from provider response to queue decision to user-facing status. That makes incidents easier to investigate and gives the person using the product a next step they can actually take.

From recorded work to a useful overview. Productivity reporting in the local demo. These figures illustrate the interface, not product-wide usage.

From recorded work to a useful overview

Productivity reporting in the local demo. These figures illustrate the interface, not product-wide usage.

The records behind the totals. Individual time sessions can be inspected separately from summary charts. All entries are sample data.

The records behind the totals

Individual time sessions can be inspected separately from summary charts. All entries are sample data.

Inside Blitzy’s voice experience.

The shared backend is what makes voice useful. The next challenge is keeping the conversation accurate, responsive, and understandable while that work happens.

Blitzy orb design study showing listening, speaking, thinking, acting and muted states
Original orb design study from the implementation repository. These are interface states, not a recording of a live call.

Two models, one owner of the work

It is easy to describe this as a talker and a thinker. That description only helps if their authority is clear. If both can interpret a request, select tools, and decide what succeeded, there are two opportunities to change the meaning of what the user asked.

The implementation evolved through that boundary. Earlier voice work gave the talker a small delegation surface. The latest revision removes business tools from the realtime session entirely. Its tool list is empty, and tool choice is disabled. It receives bounded text to speak, rather than instructions to independently complete a task.

The ordinary agent owns the turn: reading context, choosing tools, applying guards, and producing an outcome. Voice is another way into that agent and another way out. The MCP layer is a related but separate concern: external clients share product tools and the executor, while voice and typed chat share the in-app agent itself.

Use the words the user actually said

The voice adapter prefers the literal final transcript over an interpreted delegation request. This matters because an apparently helpful rewrite can quietly add authority. “What should I do today?” and “Reorganize my tasks for today” concern the same data, but authorize different behavior.

The controller waits for a completed transcription and checks that it contains speech. Failed transcription, empty input, partial text, and raw voice activity events do not start product work. Input IDs are remembered so a repeated completion event does not create a second turn.

This is deliberately a narrower boundary than “the microphone got loud.” Typing, background noise, and unfinished words are not enough to interrupt an answer or execute an action. The boundary cannot make transcription perfect, but it prevents uncertain intermediate events from becoming independent commands.

Context has to be fresh, and absence is not emptiness

A voice conversation can continue while the user changes lists or views. Capturing the current page once at call startup is not enough. The adapter reads client context again for each turn, so references such as “this list” can be resolved against the current product state.

There is a second distinction that is easy to miss: a task collection that has not been loaded is unknown, not empty. The agent should read the actual relevant data before telling someone their day is clear. Missing local context must not become a confident answer about stored tasks.

The same principle applies to scope. If the contextual read fails, that failure should not widen a request about one list into a workspace-wide action. Context is part of correctness, not just a prompt enhancement. It determines which objects a request is allowed to mean.

Advice must not silently become a write

The shared agent classifies read-only turns using the user’s words. For those turns it exposes only tools annotated as reads. There is also a check at execution time that rejects a write, even if an inappropriate call gets proposed.

Both layers matter. Filtering the tool list helps the model choose a suitable next step. The executor check enforces the boundary when a model response does not follow that expectation. A prompt alone cannot substitute for the second layer.

Explicit action requests still use the normal workflow. I am not trying to turn every request into a confirmation dialogue. The aim is to distinguish asking for a recommendation from asking the app to change something, then keep that distinction consistent in voice and text.

Generation finishing is not the same as audio finishing

A particularly important coordination detail is the difference between a model finishing a response and the speaker finishing playback. Generated audio may still be buffered after the response completion event arrives. Starting the next utterance at that point can create overlap or an abrupt cut.

I put spoken output behind one coordinator. It tracks the current owner, response identity, and playback lifecycle. The output lane remains occupied until playback stops or is explicitly cleared. An old completion event cannot release a newer utterance’s lane.

The coordinator also separates activity cues from final outcomes. A result replaces pending progress speech, and owner checks prevent late events from an earlier turn from speaking over the current one. This is ordinary concurrency control applied to a conversation, where the user hears every ordering mistake.

The final sentence needs its own contract

The agent can produce a detailed display response and a short spoken outcome. Those serve different needs. Reading the first streamed sentence aloud is unreliable because it may describe an intention, a partial result, or a step that later failed.

The voice adapter waits for the dedicated completion outcome. It does not invent one by clipping accumulated display text. A missing outcome is represented explicitly rather than treated as successful completion. That makes the boundary testable and gives the surrounding system a real failure to handle.

Speech generation then receives that bounded result with no business tools available. It presents the result of the shared workflow. It does not get another opportunity to select different tasks or silently retry a failed mutation.

Stopping a conversation is also a backend problem

Cancellation crosses several kinds of work: queued turns, an active agent run, tool execution, persistence, and audio playback. Stopping the speaker does not automatically undo a write that already completed. Treating those as one boolean would make the interface misleading.

The voice path uses cancellation signals and ownership checks to discard work that has not started and stop eligible active work. Tool-driven mutations participate in the existing change journal, which gives completed changes a separate undo path. Cancellation and reversal remain different operations.

Session replacement and attachment also need atomic handling when requests reach different server instances. Late callbacks must belong to the session that created them. A new call should not inherit an old call’s cleanup, quota accounting, or final response.

Separate the reusable audio features from the custom name

The runtime pipeline has three model stages. Audio becomes mel-spectrogram frames, those frames become speech embeddings, and a small classifier scores the recent embedding window for the wake phrase. The custom model learns the name; it does not need to learn an entire speech representation from scratch.

The application feeds 16 kHz mono audio through that pipeline. Processing uses 80 ms chunks, keeps context between chunks, and accumulates a window of sixteen 96-value embeddings before asking the classifier for a score. Buffering and sample scaling must match the training assumptions.

I kept the streaming pipeline separate from the runtime that loads ONNX sessions. Its model functions are injected, so chunk boundaries, reset behavior, and feature-window construction can be tested without loading the actual models. This makes audio preprocessing a piece of testable application code rather than an opaque callback.

Colab was a training environment, not the product runtime

The training scripts prepare synthetic speech, background material, room impulse responses, and the model configuration in Colab. The custom classifier is deliberately small: a DNN configuration with a 32-unit layer size. An always-listening feature has a different cost profile from an occasional large-model request.

Getting that environment reproducible required practical compatibility work. The committed setup pins the Torch and torchaudio versions expected by the audio tooling and installs the speech synthesis dependencies used by the sample generator. A notebook that runs once is useful; a script that explains how it ran is much easier to maintain.

Training targets in a configuration are not measured outcomes. A requested false-positive rate or a number of training steps says what the run is aiming for. It does not establish the rate of false wakes in an office, through laptop speakers, or during a meeting. I keep those distinctions explicit when reading the results.

The name was outside the pronunciation dictionary

Blitzy is a product name, and the speech generation path did not have a reliable dictionary entry for it. That matters before a single weight is trained. If the generated positives pronounce the name incorrectly, the model can learn the synthetic mistake very successfully.

The training wrapper adds explicit pronunciations for the intended name and relevant variants. It also normalizes punctuation for dictionary lookup, so a pause after “hey” does not send an otherwise known word down a different pronunciation path.

This is the kind of model work that looks like ordinary data plumbing. It is still central to the result. The classifier only sees the examples it receives. A clean training loop cannot rescue contradictory labels or a generator that is saying a different word.

Evaluate through the pipeline that will run in the app

The evaluation script uses the application’s keyword spotter and ONNX models, rather than only reporting the training framework’s score. It generates speech with macOS voices at different speaking speeds and feeds it through the streaming path. That helps catch mismatches between training-time assumptions and application-time buffering.

The committed evaluation covers 29 synthetic voices at two speeds, with positive phrases and confusable negatives. It is useful for comparing variants and exposing a failure such as “hey busy.” It is not a field study of real accents, microphone distances, room acoustics, or overlapping speakers.

A strong clean-speech result can be encouraging without becoming a claim of universal accuracy. Before treating this as a broadly validated wake system, I would want recordings from real devices and environments, false wakes measured over listening hours, and misses measured across people who did not contribute to tuning.

The second check has an explicit privacy boundary

After the local model produces a candidate, the app transcribes the short audio window that triggered it. The verification hint includes both the intended name and confusable words: Blitzy, busy, Bixby, baby, and buddy. Supplying only the desired spelling would risk nudging the transcriber toward confirming what the first model already guessed.

The normalized transcript is then checked for a word that sounds like the name. This gives the system a second opportunity to reject a near-miss before opening a conversation. It is a filter after local detection, not continuous cloud transcription of the microphone.

That distinction also matters when describing privacy. Continuous keyword spotting runs on the device, but a candidate causes a roughly two-second clip to leave the device for verification before the call starts. Calling the entire feature fully offline would be inaccurate. The implementation boundary should guide the explanation shown to a user.

A timeout is a product decision

The verifier has a 2.5-second timeout. In the current implementation, a timeout or verification error accepts the local wake instead of discarding it. This favors availability: a slow verification request should not make a genuine wake disappear. The tradeoff is that some false candidates can also open a call.

That acceptance is permission to open the voice interface, not permission to execute a business action. Commands still go through the voice input validation and shared agent workflow. A wake-started call that hears no speech also ends automatically after a short inactivity window.

There are other reasonable policies, including requiring the verification to succeed. I would choose between them using observed missed wakes, accidental opens, and the consequences of opening the interface. The important thing is to make the failure behavior deliberate and test it, rather than let a rejected promise decide the experience.

The microphone needs a clear owner

Starting and stopping are asynchronous. The user can disable listening while permission is being requested, or close the companion while a verification request is still in flight. Generation checks let those late completions recognize that they no longer belong to the active listener.

The detector releases only the media tracks it owns. The wake listener and the realtime call must hand off microphone use cleanly, then rearm listening after the call ends. A late verifier must not reopen a companion that the user just closed.

These details are part of the feature, not cleanup to add after training. This article reflects the October 1, 2026 implementation, with the complete frontend voice work still in review. The custom model and training notes establish what I built; real-world acoustic evaluation remains a separate bar for how confidently it can be described.

The animation has a job

A voice assistant has fewer visible clues than a normal interface. When a button changes a task, the user can usually see the button, the task, and the result together. During a spoken request, they may be looking elsewhere. A small animated companion becomes the most immediate sign of what the system thinks is happening.

For Blitzy, I wanted that companion to feel friendly without turning every moment into a performance. The orb carries the brand’s face on a rotating surface of dots. It can blink, listen, think, read, work, and speak. Each behavior is tied to a state in the voice system.

The useful question was not how many effects I could put on a sphere. It was whether someone could tell the difference between “ready for you,” “heard you,” “working,” and “finished speaking.” A pretty animation that blurs those states would make the assistant harder to trust.

Start with a state model, then draw it

The voice model distinguishes off, connecting, listening, user speaking, thinking, acting, speaking, muted, ended, and error. Those are not interchangeable shades of busy. Connecting means there may not be a usable conversation yet. Listening means the session is ready. Acting means product work is actually underway.

I separated the orb’s numeric model from its canvas renderer. The model describes rotation, breathing, audio response, brightness, saturation, a scan band, and a working ring. The renderer turns those values into a frame. That keeps a large amount of the behavior inspectable without mounting the visual component.

It also gives state changes a consistent vocabulary. Rather than unrelated animations competing for attention, each state chooses a target set of parameters. Transitions blend toward those targets. The same face remains present while its behavior changes.

A microphone meter does not prove the assistant heard you

Raw microphone levels respond to everything the device captures. A keyboard, a chair, a fan, or another person can make the meter move. Feeding that value directly into an expressive character suggests recognition that has not actually happened.

The orb model gates input energy by the user-speaking state. Output energy is similarly used only while the assistant is speaking. Listening readiness still has a calm visual treatment, but a random loud sound is not automatically presented as a conversational response.

This is a small but important distinction between a signal and its meaning. Volume is a measurement. A recognized speaking state is the application’s interpretation. The interface should make it clear which one it is showing, especially when a user may rely on the animation instead of reading a transcript.

The sphere turns, but the face keeps looking at you

The surface uses points distributed around a sphere with a Fibonacci construction. That gives an even-looking arrangement without building a large mesh. The canvas renderer projects the points into two dimensions and adjusts their appearance to create depth.

The eyes behave differently from the dots. They remain positioned toward the viewer while the surface rotates beneath them. If the whole face rotated away with the sphere, the companion would repeatedly stop being readable. Keeping the face stable lets the motion add life without losing the expression.

A blink and a small upward gaze during thinking provide personality. They are modest details, and their timing belongs to the visual system rather than the agent prompt. The assistant does not need to generate a new instruction to decide how every eyelid moves.

Smooth transitions should not depend on the refresh rate

Blending a fixed percentage toward a target on every frame makes an animation behave differently at different frame rates. A display that produces more frames completes more blend steps in the same wall-clock time. That can change both the speed and the feel of the character.

The orb blends using elapsed time and an exponential approach to the target. Rotation also advances according to elapsed time. These choices make the state transitions more consistent when frames are uneven or the display refresh rate changes.

There is no need for a physics engine to get that behavior. The underlying model is a small collection of numbers with carefully chosen transitions. Keeping the mechanism simple makes it easier to reason about an abrupt state change, such as speech being interrupted or the microphone becoming muted.

The cheapest frame is the one you do not draw

A desktop app can stay open for hours. A decorative canvas that keeps rendering while hidden can accumulate a real cost even if each frame is cheap. I use viewport intersection and document visibility to stop the animation loop when the orb is offscreen or the document is hidden.

When rendering resumes, timing is reset so the paused interval does not become one giant animation step. The canvas also follows theme changes, and its pixel density is bounded rather than growing without limit on a high-density display.

Reduced motion takes a different path: a still frame for the current state, with redraws when relevant state or appearance changes. The state remains available without requiring continuous motion. I would still measure frame cost and power use on representative hardware before claiming a specific performance result.

What I take from it.

The common thread is control. A person should know when Blitzy is listening, whether an action actually ran, what changed, and how to recover. That requires work across model boundaries, APIs, permissions, persistence, and the interface. None of those layers can provide the whole experience on its own.

This is the kind of engineering I enjoy: a feature that looks simple enough to use immediately, backed by decisions that hold up when the interaction gets complicated. Voice is the newest example. The integration framework, change journal, collaboration, and migration work are the foundation that gives it something useful to do.

Work reviewed through October 2026. Voice backend changes are merged; the complete frontend voice experience is in review. Screenshots use an isolated local demo account.

Keep exploringMaddyCustom: From a custom design to a delivered order. ↗Have a problem like this? Let’s talk ↗
0% · Reading with you