diff --git a/src/content/projects/ai-voice-service-agent.mdx b/src/content/projects/ai-voice-service-agent.mdx index 039f648..69b8255 100644 --- a/src/content/projects/ai-voice-service-agent.mdx +++ b/src/content/projects/ai-voice-service-agent.mdx @@ -1,6 +1,6 @@ --- title: "AI Voice Service Agent Prototype" -description: "A non-production personal learning prototype exploring how a voice agent can answer customer calls, retrieve current company data, use tools to create a job request, and keep final confirmation under human control." +description: "Personal prototype: a voice agent answers a small-business call, reads live company data, drafts a job via tools, and stops until a person confirms." image: /images/projects/ai-voice-service-agent-workflow.jpg imageFit: contain period: "2026–Present" @@ -14,19 +14,17 @@ technologies: - "Structured Data" - "Human-in-the-Loop" achievements: - - "Prototyped a task-oriented voice workflow for inbound small-business customer calls" - - "Grounded conversations in current company information retrieved from real service APIs" - - "Used validated tool calls to turn customer intent into a structured job request" - - "Kept consequential actions human-controlled by requiring operator review before customer confirmation" + - "Task-oriented inbound call flow for a small service business" + - "Answers grounded in current company data from APIs, not model memory" + - "Job requests created only as drafts, via validated tools" + - "Operator review before any confirmation back to the caller" challenges: - - "Maintaining a natural conversation while gathering complete structured job details" - - "Grounding answers in current business data instead of relying on model memory" - - "Mapping conversational intent to safe, validated service actions" - - "Handling uncertainty and escalation without making unsupported commitments" + - "A natural call still has to produce a complete job record" + - "Company facts change; the weights do not" + - "Ambiguous or high-stakes requests should escalate, not bluff" outcomes: - - "Practical learning in voice-agent interaction and tool-use patterns" - - "A sample architecture separating conversation, company-data retrieval, business actions, and human review" - - "A clear path from customer request to operator-reviewed job creation" + - "A sample split between conversation, reads, writes, and human review" + - "A concrete path from phone call to operator-reviewed job" role: "Personal Learning Prototype — Architecture and Experimentation" tags: - "Voice AI" @@ -36,94 +34,58 @@ tags: - "Human-in-the-Loop" category: "ai" publishDate: 2026-08-11 -lastUpdated: 2026-08-11 +lastUpdated: 2026-08-27 --- -## Scope - -This is a **non-production personal learning prototype** exploring an AI voice agent for small service businesses. The sample focuses on an inbound phone-call scenario in which a customer needs information or wants to request work. - -The agent can hold a task-oriented conversation, retrieve current information about the company from services, and use tools to prepare a structured action. For example, it can gather the details of a requested job and create a job request for an operator to review. - -The prototype intentionally keeps final authority with a person. Creating a draft request is useful automation; accepting work, committing to details, or confirming the request to the customer remains subject to human review. - -## Example Customer Journey - -1. A customer calls the company's service number -2. The voice agent answers and identifies what the customer needs -3. It retrieves current company, service, or availability information from APIs when required -4. It asks focused follow-up questions to collect missing job details -5. Validated tool calls create a structured draft job request in the business system -6. A human operator reviews, corrects, and approves the request -7. The approved outcome can then be confirmed to the customer +## Why I built this -This workflow connects conversational AI to real business data and actions while placing a control point before a consequential commitment is made. +I wanted to see whether a voice agent can take an inbound service call — “are you free Thursday, can you quote a repair” — without becoming a chatbot that invents opening hours. The interesting part is not speech-to-text. It is: live data, structured writes, and a human still owning the commitment. -## Grounding in Real Company Data +This is a non-production personal prototype. -A customer-facing agent must not rely only on the model's training data or prompt text. Company details can change, including offered services, coverage areas, availability, customer records, and operating policies. - -The sample therefore treats business services as the source of truth. The agent retrieves relevant data when needed and uses the returned information to answer the customer. If required information is unavailable or ambiguous, the safe response is to clarify or escalate rather than invent an answer. - -## Tool-Based Business Actions - -Conversation becomes operationally useful when it can produce a structured result. Tool calls provide an explicit boundary between the LLM and company systems. - -For a job request, the tool contract can require fields such as: +## Scope -- Customer and contact details -- Requested service -- Location -- Preferred timing -- Description and notes gathered during the call -- Source and conversation reference -- Review status +The sample is one inbound call into a small service business. The agent can: -Validating these inputs before calling a service makes the action testable and prevents arbitrary model output from being written directly into company data. +- Figure out whether the caller wants information or wants work booked +- Fetch current company, service, or availability data from APIs +- Ask only for the missing job fields +- Create a **draft** job request through a validated tool -## Human Review and Confirmation +It cannot accept the job, promise a price or a slot, or confirm back to the caller until an operator says so. Telephony, recording consent, and production IAM are out of scope. -The draft job request is routed to a person before it is treated as confirmed. The operator can review the transcript and structured fields, correct misunderstood details, check operational constraints, and decide whether the company can accept the work. +## Core workflow -This human-in-the-loop step addresses several risks: +1. Caller rings the service number. +2. The agent identifies the ask. +3. It reads from company APIs when the answer has to be true today. +4. It gathers missing job fields with short questions, not a form recitation. +5. A tool call writes a structured draft: contact, service, location, timing, notes, conversation reference, review status. +6. An operator reads transcript + fields, edits, accepts or rejects. +7. Only then does confirmation go to the customer. -- Speech recognition or interpretation errors -- Incomplete customer information -- Unsupported services or locations -- Availability conflicts -- Pricing or contractual commitments -- Requests that need judgement or specialist handling +If the API is down or the request is outside what the business does, the agent clarifies or escalates. It does not fill the gap from training data. -The agent assists with intake and preparation; the business retains control of the decision. +## Design choices -## Architecture Explored +**The model is not the CRM.** Services, coverage, and availability change. Anything the caller will act on is retrieved. “I think they cover that suburb” is a bug. -The prototype separates concerns into components that could evolve independently: +**Writes go through a narrow tool.** The job schema is the contract: required fields, types, source of the conversation. Free-form model output does not land in the business system. That makes the write testable and keeps the LLM off unrestricted APIs. -- **Telephony and audio layer:** receives the call and streams speech -- **Conversation layer:** maintains context and decides whether to ask, retrieve, act, or escalate -- **Company-data tools:** read current information from authorised services -- **Action tools:** validate and submit structured requests -- **Human review queue:** presents the request and supporting context to an operator -- **Confirmation path:** communicates only the approved outcome +**Read and write are different trust boundaries.** Looking up opening hours is not the same operation as creating work. Draft-only writes plus a review queue is the control point for speech errors, incomplete addresses, jobs the company cannot take, and anything contractual. -This separation helps keep model reasoning away from direct, unrestricted access to business systems. +**Escalation is a successful path.** Low confidence, missing data, or a request that needs judgement should leave the happy path. Pretending otherwise is how you get a polite, wrong booking. -## Agentic Engineering Approach +I used coding agents on this prototype the same way I would on a small internal tool: specify the flow and tool schemas, generate, then check against the workflow, authz, and failure cases. Accepting the first generated handler would miss the point. -I used the prototype to practise an AI-assisted engineering workflow as well as voice-agent concepts. Coding agents can support requirements analysis, conversation-flow design, interface definitions, implementation, test generation, and review. Their output is checked against the intended workflow, tool schemas, security boundaries, and failure cases. +## What I learned -The goal is not to accept generated code without scrutiny. Agentic engineering works best when specifications, project context, automated validation, independent review, and human judgement form part of the process. +Voice needs shorter turns and stricter gathering than text chat. People will not sit through eight questions; the agent has to know which fields block a draft. -## What I Learned +Grounding is not a prompt paragraph. If the retrieve step is optional, the model will skip it under time pressure. -- Voice interaction requires concise prompts and progressive information gathering -- Real service data is essential for trustworthy customer answers -- Tool calls should be narrow, authorised, schema-validated, and observable -- Read operations and consequential write operations need different trust boundaries -- A draft-and-review workflow can provide value without granting the model final authority -- Explicit escalation is a feature, not a failure, when confidence or information is insufficient +The review queue is where the prototype earns its keep. Intake automation is useful even when the model is not allowed to speak for the business. -## Production Considerations +## Not production -This prototype demonstrates concepts only. A production implementation would need telephony-provider integration, identity and consent handling, privacy and recording policies, authentication and authorisation for every tool, latency and interruption testing, prompt-injection controls, evaluation against realistic calls, monitoring, fallback handling, and a well-defined operational review process. +A real phone agent still needs a telephony vendor, consent and recording rules, identity, per-tool authn/z, latency and barge-in behaviour, prompt-injection controls, an eval set of real calls, monitoring, and a defined on-call review process. This repo explores the workflow, not that operating environment. diff --git a/src/content/projects/cicd-keyless-delivery.mdx b/src/content/projects/cicd-keyless-delivery.mdx index 9906d15..4078d43 100644 --- a/src/content/projects/cicd-keyless-delivery.mdx +++ b/src/content/projects/cicd-keyless-delivery.mdx @@ -1,6 +1,6 @@ --- title: "Keyless Multi-Environment CI/CD on GCP" -description: "A build-once delivery platform that publishes versioned artifacts from GitHub Actions, promotes them across environments, renders reusable Kustomize overlays, and uses Workload Identity Federation instead of long-lived service-account keys." +description: "Build once to Artifact Registry for GKE and Cloud Run, promote the same image through dev/staging/prod, coordinate Firebase artifacts under the same release, and replace JSON keys with Workload Identity Federation." image: /images/projects/keyless-cicd-gcp.jpg imageFit: contain period: "2026–Present" @@ -19,21 +19,18 @@ technologies: - "Docker" - "IAM" achievements: - - "Standardised development, staging, and production delivery around versioned artifacts from a shared Artifact Registry" - - "Created reusable Kubernetes configuration with common Kustomize bases and service- and environment-specific overlays" - - "Implemented keyless GitHub Actions authentication to GCP through Workload Identity Federation" - - "Removed mounted service-account key files from GKE workloads by adopting workload-specific federated identities" - - "Supported consistent deployment paths for Kubernetes, Cloud Run, and Firebase" + - "Promoted one versioned artifact through development, staging, and production" + - "Shared Kustomize bases with service and environment overlays instead of copied manifests" + - "GitHub Actions reaches GCP through Workload Identity Federation — no service-account keys in secrets" + - "GKE workloads use federated identities instead of mounted JSON key files" challenges: - - "Preventing environment drift across multiple services and deployment targets" - - "Sharing Kubernetes configuration without duplicating manifests for each service and environment" - - "Replacing long-lived credentials without interrupting CI/CD or workload access" - - "Applying least-privilege identity across deployment automation and runtime services" + - "Rebuilding per environment meant staging and production were not the same bits" + - "Copied Kubernetes YAML drifted across services" + - "Long-lived keys in GitHub and in pods were a rotation and leak problem" outcomes: - - "Consistent promotion of the same tested artifact across environments" - - "Reusable and reviewable Kubernetes configuration" - - "Fewer long-lived secrets in GitHub Actions and service workloads" - - "Clearer service- and environment-specific deployment differences" + - "The image that passed staging is the image that runs in production" + - "Environment differences are overlays in git, not hand-edited live manifests" + - "CI and runtime credentials are short-lived and scoped" client: "Zimi Ltd" role: "Senior Cloud Engineer" tags: @@ -45,112 +42,49 @@ tags: - "Cloud Security" category: "infrastructure" publishDate: 2026-08-10 -lastUpdated: 2026-08-11 +lastUpdated: 2026-08-27 --- -## Project Overview +This is the delivery path for the Zimi GCP platform: GitHub Actions builds a versioned image once, stores it in Artifact Registry, and promotes that same digest through development, staging, and production for GKE and Cloud Run. Firebase Hosting and Functions use their own release artifacts and deployment toolchain. Kubernetes config is composed with Kustomize. GitHub and GKE authenticate to GCP with Workload Identity Federation, not JSON keys. -This project standardised how application artifacts move from the main branch into development, staging, and production on Google Cloud Platform. +## Rebuilds, copied YAML, and keys in the cluster -It combined three related improvements: +Building separately for each environment looks convenient until something fails only in production. The git SHA might match; the image often does not. Different build times, different caches, different unnoticed Dockerfile changes. -1. Build once and promote the same versioned artifact between environments. -2. Manage Kubernetes configuration through reusable Kustomize bases and overlays. -3. Replace long-lived CI/CD and workload credentials with Workload Identity Federation. +Kubernetes had a quieter form of the same drift. Each service and environment carried its own copy of probes, labels, ports, and resources. A one-line fix became a scavenger hunt. Intentional prod-only differences were hard to tell from accidents. -The resulting delivery model supports Kubernetes, Cloud Run, and Firebase while keeping environment differences explicit and reducing the number of static credentials that need to be stored and rotated. +Credentials were the third copy. Service-account keys in GitHub secrets and keys mounted into pods are long-lived, easy to over-scope, and painful to rotate. Anyone who has the file has the access. -## The Delivery Problem +## What changed -Building separately for each environment can introduce drift: development, staging, and production may nominally reference the same source revision but run artifacts created at different times or under different conditions. +Main-branch CI builds one versioned container image and publishes it to a shared Artifact Registry. GKE and Cloud Run promote that digest through development, staging, and production instead of rebuilding it. Firebase is a separate path tied to the same source revision and release: Hosting deploys built directory contents, while Functions are packaged and built through the Firebase toolchain. -Kubernetes configuration can also drift when every service and environment maintains a copied set of manifests. Small fixes must then be repeated across many files, and it becomes difficult to distinguish intentional differences from accidental ones. +Kubernetes manifests are layered: -Finally, service-account JSON keys create an avoidable security and operational burden. Keys stored in CI/CD or mounted into containers are long-lived credentials that must be distributed, protected, audited, and rotated. +- A **base** for the common shape: resources, labels, probes, ports. +- A **service overlay** for things that belong to one app. +- An **environment overlay** for dev, staging, or production. -## Build Once, Deploy Many +CI renders the selected composition to plain YAML and deploys that. The override is visible in git; the rendered output can be reviewed before it hits the cluster. -The pipeline uses a main-branch build as the source of a versioned container artifact: +GitHub Actions presents its OIDC identity to GCP and receives short-lived credentials. IAM bindings constrain *which* repository, branch, and workflow can assume the deployer identity, and *what* that identity may do: push images, read config, deploy to the intended project. -1. GitHub Actions checks out and validates the main branch. -2. The pipeline builds a versioned container image once. -3. The image is published to a shared Artifact Registry. -4. Development, staging, and production reference that same artifact version. -5. Runtime-specific jobs deploy to Kubernetes, Cloud Run, or Firebase as required. - -Promoting an immutable artifact means production runs the same build that was already exercised in earlier environments. It also creates a clearer audit trail between a source revision, artifact version, and deployment. - -## Kubernetes Configuration with Kustomize - -Kubernetes manifests were organised into composable Kustomize layers: - -- A **base** contains common resources, labels, probes, ports, and configuration structure. -- **Service overlays** apply changes that belong to an individual service. -- **Environment overlays** apply development, staging, or production differences. -- CI renders the selected composition into standard Kubernetes manifests before deployment. - -This approach keeps common configuration in one place and makes every override visible in source control. Rendered output can be reviewed or validated before it reaches the cluster. - -## Keyless GitHub Actions Authentication - -GitHub Actions authenticates to GCP through Workload Identity Federation. - -Instead of storing a service-account JSON key in GitHub secrets, the workflow presents its GitHub OIDC identity and receives short-lived GCP credentials. IAM conditions and roles restrict which repository, branch, and workflow can assume the deployment identity and what that identity can do. - -The federated credentials support authorised operations such as: - -- Publishing container images to Artifact Registry -- Reading deployment configuration -- Deploying to the intended GCP runtime -- Accessing only the projects and environments assigned to the workflow - -This removes a high-value static credential from CI/CD and reduces key-rotation work. - -## Workload Identity Federation for GKE - -Runtime services were also moved away from mounted service-account key files. - -With Workload Identity Federation for GKE, a Kubernetes service account is associated with an authorised GCP identity. The workload obtains short-lived credentials at runtime and can access only the GCP resources allowed by its IAM roles. - -This allows each service to have a dedicated, least-privilege identity without copying private key files into pods. Application code uses standard GCP authentication rather than knowing where a credential file has been mounted. +GKE workloads dropped mounted key files. A Kubernetes service account is bound to a GCP identity; the pod gets short-lived tokens at runtime and only the IAM roles of that identity. ![GKE workloads using Workload Identity Federation for keyless, short-lived access to GCP resources](../../assets/projects/gke-workload-identity-federation.jpg) -## Artifact Access - -Artifact Registry and Kubernetes access follow the same identity-first principle: platform identities and short-lived credentials are preferred over manually distributed keys. - -The pipeline publishes versioned images under an authorised deployment identity, and the Kubernetes environment retrieves approved artifacts through its configured GCP identity path. - -## Results - -### Delivery consistency +## Why federation instead of JSON keys -- Established a shared Artifact Registry as the source of deployable images -- Promoted the same versioned artifact through development, staging, and production -- Supported Kubernetes, Cloud Run, and Firebase delivery paths -- Made deployments easier to trace back to source and artifact versions +A JSON key works from anywhere that holds it: a laptop, a fork, a log line. Federation ties access to an attested workload identity. The token expires. The condition can say “only this workflow on `main`,” which is a much smaller accident surface than “anyone with `gcp-sa.json`.” -### Maintainability +The same rule at runtime: the pod should not know where a secret file was mounted. Application code uses ADC; GCP decides whether that GKE service account may read the bucket or publish to Pub/Sub. -- Reduced duplicated Kubernetes YAML -- Centralised common configuration in Kustomize bases -- Kept service and environment differences explicit in overlays -- Rendered standard manifests for validation and deployment +The trade-off is setup cost. Federation needs a pool, provider, attribute mapping, and IAM conditions — more moving parts than pasting a key into GitHub. It is also much easier to reason about six months later, when nobody wants to be the person who rotates twenty keys. -### Security +## After -- Replaced static GitHub-to-GCP credentials with short-lived federated access -- Removed mounted service-account key files from GKE workloads -- Applied workload-specific identities and least-privilege IAM -- Reduced credential distribution and rotation responsibilities +Environments promote an immutable artifact instead of rebuilding it. Kubernetes differences live in overlays. GitHub Actions and GKE workloads talk to GCP without long-lived keys in the repo or the cluster. -## Skills Demonstrated +The coordinated release covers Kubernetes and Cloud Run with the promoted container digest, and Firebase with separately built deployment artifacts. Every production deploy traces back to the same source revision and release; only GKE and Cloud Run trace to the shared Artifact Registry digest. -- Multi-environment CI/CD architecture -- Immutable artifact creation and promotion -- GitHub Actions and Artifact Registry -- Kubernetes configuration management with Kustomize -- GCP IAM and Workload Identity Federation -- Keyless workload authentication -- GKE, Cloud Run, and Firebase delivery +Related: [production deployment in the C4 model](/projects/iot-platform-architecture). diff --git a/src/content/projects/gcp-cost-optimization.mdx b/src/content/projects/gcp-cost-optimization.mdx index 8b6c506..31e9294 100644 --- a/src/content/projects/gcp-cost-optimization.mdx +++ b/src/content/projects/gcp-cost-optimization.mdx @@ -1,6 +1,6 @@ --- title: "GCP Database & Logging Cost Optimisation" -description: "A production cost-control initiative using monthly Cloud SQL PostgreSQL partitions, historical-data archival, partition-based retention, and focused Cloud Logging to constrain recurring storage and logging cost growth." +description: "Monthly Cloud SQL partitions, archive-and-drop retention, and quieter logging so time-series storage and Cloud Logging stop growing without bound." image: /images/projects/gcp-storage-cost-optimisation.jpg imageFit: contain period: "2026" @@ -15,19 +15,18 @@ technologies: - "Time-Series Data" - "Data Archival" achievements: - - "Controlled active Cloud SQL storage growth through monthly partitioning, archival, and partition-based retention" - - "Improved recent time-series query paths by allowing PostgreSQL to prune irrelevant historical partitions" - - "Replaced large row-by-row retention deletes with fast removal of expired partitions" - - "Reduced Cloud Logging volume while preserving essential operational and diagnostic context" + - "Stopped unbounded Cloud SQL growth with monthly partitions, archival, and partition drops" + - "Recent time-series queries prune history instead of scanning it" + - "Replaced long row-by-row retention deletes with dropping an expired month" + - "Cut Cloud Logging volume while keeping the logs operators actually use" challenges: - - "Controlling continuously growing time-series storage without slowing recent-data queries" - - "Removing historical data efficiently while preserving an archive" - - "Reducing logging volume without losing information needed to operate and troubleshoot services" + - "Time-series and logs grow every day; deleting them naively is worse than keeping them" + - "Recent operational queries were paying for years of history still sitting in the hot table" + - "Low-value log lines were a recurring ingest-and-retain cost" outcomes: - - "More predictable Cloud SQL storage and Cloud Logging spend" - - "A repeatable archive-before-delete data-retention process" - - "Faster access paths for recent time-series data" - - "Lower Cloud Logging ingestion and retention volume" + - "Active database size and logging volume became a policy, not an accident" + - "Retention is archive-then-drop, not a multi-hour DELETE" + - "Recent telemetry queries hit a smaller working set" client: "Zimi Ltd" role: "Senior Cloud Engineer" tags: @@ -38,91 +37,47 @@ tags: - "Cloud Logging" category: "cloud" publishDate: 2026-08-11 -lastUpdated: 2026-08-11 +lastUpdated: 2026-08-27 --- -## Project Overview +The Zimi platform stores years of device telemetry and control history in Cloud SQL for PostgreSQL, and it sends application logs to Cloud Logging. Both grow every day. I treated that growth as a data-lifecycle problem: what must stay hot, what can be archived, and what should never have been written. -This production GCP cost-control initiative addressed two continuously growing sources of recurring spend: time-series records in Cloud SQL for PostgreSQL and application logs sent to Cloud Logging. +## The storage bill was a data-lifecycle problem -The goal was not simply to delete more data or log less. It was to establish deliberate lifecycle policies that retained the information needed for operations while keeping active storage and logging volume controlled over time. +Time-series tables do not plateau. Every power sample and state change lands in the operational database, so storage, backups, and query plans all get heavier even when nobody is asking about last year. -## The Challenge +Retention as `DELETE FROM … WHERE time < …` made that worse. Removing millions of rows is a long write that generates WAL, dead tuples, bloat, and autovacuum pressure. Ordinary `VACUUM` generally makes that space reusable inside PostgreSQL; it does not return the table's disk allocation to the operating system. Meanwhile the queries operators actually run — last hour, last day, this month — were still scanning a table that contained the whole history. -Time-series systems naturally accumulate data. Keeping all historical records in the operational Cloud SQL database increased storage usage, made retention harder to manage, and caused queries to operate against a larger active dataset than necessary. +Logging had the same shape. Repetitive, low-signal lines were cheap to `console.log` and expensive to ingest and retain. -Large row-by-row deletion jobs were also an inefficient retention mechanism. They could run for a long time, add database load, create substantial write-ahead logging, and require additional cleanup before storage could be reclaimed. +## What changed -Application logging had a similar lifecycle problem. Repetitive or low-value messages increased ingestion and retention volume even though they provided little benefit during monitoring or incident investigation. +I converted the hot time-series tables to monthly PostgreSQL range partitions on Cloud SQL. New months are created before they are needed; inserts land in the current month; older months are physical tables we can manage on their own. -## Monthly PostgreSQL Partitions +The retention path is then boring on purpose: -I organised the time-series tables as monthly PostgreSQL range partitions in Cloud SQL. Each partition represents a clear operational period and can be managed independently from the rest of the table. +1. Archive a month that has left the active window. +2. Check the archive is complete. +3. `DROP` the partition from Cloud SQL. -This supports two important behaviours: +Dropping an expired partition avoids row-by-row deletion and the resulting vacuum overhead. It is fast, but PostgreSQL takes an `ACCESS EXCLUSIVE` lock on the parent table for the operation, so the retention job runs in a controlled window and keeps the lock brief. -- Queries for recent time windows can use partition pruning and avoid scanning unrelated historical partitions. -- Retention can operate on a complete partition rather than deleting millions of individual rows. +On the logging side I went through what the services emitted and kept lines that help with monitoring, incidents, state changes, and audit. Noise that never appeared in those workflows stopped going to Cloud Logging. -Monthly boundaries provide a practical balance: they are fine-grained enough for common recent-data access patterns while remaining straightforward to create, monitor, archive, and remove. +## Why drop a month, not delete rows -## Archive and Retention Workflow +Monthly boundaries match how this data is used. Support and operations almost always ask about *recent* time. Partition pruning lets PostgreSQL ignore months outside the query range, so the hot path stays a few partitions even as the estate ages. -The lifecycle follows a repeatable process: +A month is also a sane ops unit: large enough that we are not managing hundreds of tiny tables, small enough to archive and drop on a schedule. Daily partitions would have been more moving parts; yearly partitions would have left too much history in the active query path. -1. Create upcoming monthly partitions before they are needed. -2. Route new time-series records into the appropriate partition. -3. Monitor partition size, active storage, and query behaviour. -4. Archive partitions after they pass the active retention period. -5. Validate that the archive is complete and usable. -6. Drop the expired partition from the operational database. +The important product decision sits next to the SQL: how long is “active,” where does the archive live, and who is allowed to read it. Partitioning without that agreement just moves the mess. -Dropping a partition is a fast metadata-level operation compared with a large row-by-row `DELETE`. It also makes retention easier to reason about because the unit being archived and removed is explicit. +## After -## Recent-Data Query Performance +Active Cloud SQL size is now governed by the retention window, not by how long the platform has been running. Recent telemetry queries work against a smaller set of partitions. Retention no longer depends on a multi-hour delete job. -Most operational queries focus on recent time windows. Partition pruning allows PostgreSQL to target only the monthly partitions that overlap the requested range. +Cloud Logging ingest and retain less, without dropping the context needed to run the platform. -The result is a smaller working set for common queries and a data model whose query behaviour remains more predictable as historical records accumulate outside the active window. +Approved before-and-after storage, query, and logging figures are not published here. -## Logging Optimisation - -I reviewed application logging to distinguish actionable operational information from repetitive noise. - -The revised approach preserves the context needed for: - -- Production monitoring -- Troubleshooting and incident investigation -- Important state transitions and failures -- Auditability where required - -Messages that did not materially help those activities were removed or adjusted so they were not unnecessarily ingested by Cloud Logging. This lowered logging volume without sacrificing essential diagnostic context. - -## Results - -### Storage and cost control - -- Constrained active Cloud SQL storage and its associated cost growth -- Established a predictable archive-before-delete retention process -- Made historical-data removal faster and operationally simpler -- Reduced Cloud Logging ingestion and retention volume - -### Query and operational behaviour - -- Improved access paths for recent time-series queries through partition pruning -- Avoided large recurring row-deletion jobs -- Made monthly storage growth and retention easier to monitor -- Preserved the logs needed to operate and troubleshoot the platform - -## Skills Demonstrated - -- GCP cost and capacity optimisation -- Cloud SQL for PostgreSQL operations -- PostgreSQL range partitioning and partition pruning -- Time-series data lifecycle design -- Data archival and retention planning -- Production logging and observability design - -## Measurement - -The mechanisms and outcomes are documented here without publishing confidential cost or storage figures. Approved before-and-after storage, query, and logging metrics can be added when they are available for public use. +Related: the same data stores sit in [the C4 model](/projects/iot-platform-architecture) — document snapshot, Redis hot path, partitioned PostgreSQL history, archive. diff --git a/src/content/projects/iot-dashboard.mdx b/src/content/projects/iot-dashboard.mdx index bba529b..a067bea 100644 --- a/src/content/projects/iot-dashboard.mdx +++ b/src/content/projects/iot-dashboard.mdx @@ -1,6 +1,6 @@ --- title: "IoT Network Management Dashboard" -description: "Comprehensive web-based platform for IoT device network management and monitoring with interactive data visualizations for device telemetry, usage patterns, and IoT device health metrics." +description: "React admin console for a 70,000+ device fleet — live status, telemetry charts, and role-aware tools for operators and support, not a spreadsheet of every device." period: "2018-2022" status: "completed" featured: true @@ -18,26 +18,18 @@ technologies: - "Redis" - "Socket.IO" achievements: - - "Real-time monitoring of a 70,000+ device fleet" - - "Interactive data visualizations for complex telemetry data" - - "Comprehensive usage pattern analytics" - - "Device health monitoring with predictive alerts" - - "Admin dashboard with role-based access control" - - "Mobile-responsive design for field technicians" + - "Operator console for a 70,000+ device fleet" + - "Live status and telemetry without polling the whole estate" + - "Role-based access for admins, support, technicians, and customers" challenges: - - "Handling large volumes of real-time data" - - "Creating intuitive visualizations for complex IoT data" - - "Ensuring responsive performance with thousands of devices" - - "Implementing efficient data aggregation and filtering" - - "Managing user permissions and access control" + - "You cannot render 70,000 live rows and call it a dashboard" + - "Telemetry is useful only if you can filter, time-bound, and drill into one network" + - "Support, installers, and customers must not see the same surface" outcomes: - - "Faster operator troubleshooting and better fleet visibility" - - "Improved operational efficiency for support teams" - - "Enhanced customer satisfaction through proactive monitoring" - - "Data-driven insights for product development" - - "Reduced operational costs through automated monitoring" + - "Support can see a silent device before someone drives to the house" + - "Fleet health is a first-class view, not a database query" + - "The same APIs serve the console and the rest of the platform" image: /images/projects/iotdashboard.png - client: "Zimi Ltd" role: "Full Stack Engineer" tags: @@ -48,339 +40,51 @@ tags: - "Admin Panel" category: "web" publishDate: 2022-03-01 -lastUpdated: 2024-02-01 +lastUpdated: 2026-08-27 --- -## Project Overview - -The IoT Network Management Dashboard is a sophisticated web-based platform designed to provide comprehensive monitoring, management, and analytics capabilities for large-scale IoT device networks. Built to handle real-time data from thousands of smart electrical devices, the dashboard serves as the central command center for administrators, support teams, and field technicians. - -## The Challenge - -Managing a network of 70,000+ IoT devices presents unique challenges: - -- **Scale**: Handling real-time data from thousands of devices simultaneously -- **Complexity**: Visualizing complex telemetry data in an intuitive way -- **Performance**: Maintaining responsive UI with continuous data updates -- **Usability**: Creating interfaces suitable for both technical and non-technical users -- **Reliability**: Ensuring 24/7 monitoring capabilities with high availability - -## System Architecture - -The admin console sits in the C4 model as one container of the wider platform — see [Diagramming a Smart-Home IoT Platform with C4](/projects/iot-platform-architecture). - -### Frontend Architecture -Built using modern React patterns with Redux for state management: -- **Component-based Design**: Reusable UI components for consistent experience -- **State Management**: Redux for complex application state handling -- **Real-time Updates**: WebSocket integration for live data streaming -- **Responsive Design**: Mobile-first approach for field technician use - -### Data Visualization Layer -Implemented multiple charting libraries for comprehensive data representation: -- **Chart.js**: Standard charts for metrics and KPIs -- **D3.js**: Custom visualizations for complex IoT data relationships -- **Real-time Charts**: Live updating charts with efficient data handling -- **Interactive Filters**: Dynamic filtering and drill-down capabilities - -### Backend Integration -Seamless integration with IoT platform services: -- **RESTful APIs**: Standard HTTP APIs for data retrieval and device management -- **WebSocket Connections**: Real-time data streaming for live updates -- **Data Aggregation**: Server-side processing for efficient data delivery -- **Caching Layer**: Redis caching for improved performance - -## Key Features - -### Real-time Device Monitoring - -**Device Status Overview** -- Live status indicators for all connected devices -- Network connectivity health monitoring -- Battery level tracking for wireless devices -- Last communication timestamps - -**Interactive Device Map** -```typescript -// Device map component with real-time updates -const DeviceMap: React.FC = () => { - const [devices, setDevices] = useState([]); - const [selectedDevice, setSelectedDevice] = useState(null); - - useEffect(() => { - // WebSocket connection for real-time updates - const socket = io('/device-updates'); - - socket.on('device-status-update', (update: DeviceUpdate) => { - setDevices(prev => updateDeviceInList(prev, update)); - }); - - return () => socket.disconnect(); - }, []); - - return ( -
- - {devices.map(device => ( - setSelectedDevice(device)} - status={device.status} - /> - ))} - - {selectedDevice && ( - - )} -
- ); -}; -``` - -### Advanced Analytics Dashboard - -**Usage Pattern Analysis** -- Historical usage data visualization -- Peak usage time identification -- Energy consumption patterns -- Device utilization metrics - -**Predictive Health Monitoring** -- Device health scoring algorithms -- Predictive failure analysis -- Maintenance scheduling recommendations -- Performance degradation alerts - -### Data Visualization Components - -**Custom Chart Components** -```typescript -// Real-time telemetry chart component -const TelemetryChart: React.FC = ({ deviceId }) => { - const [data, setData] = useState([]); - const chartRef = useRef(); - - useEffect(() => { - const updateChart = (newData: TelemetryData) => { - setData(prev => [...prev.slice(-100), newData]); // Keep last 100 points - - if (chartRef.current) { - chartRef.current.data.datasets[0].data = data; - chartRef.current.update('none'); // Smooth animation - } - }; - - const socket = io('/telemetry'); - socket.on(`device-${deviceId}`, updateChart); - - return () => socket.disconnect(); - }, [deviceId]); - - return ( -
- -
- ); -}; -``` - -**Interactive Filter System** -- Dynamic filtering by device type, location, status -- Date range selection for historical data -- Real-time search with autocomplete -- Custom filter combinations - -### User Management & Access Control - -**Role-based Access Control** -- Admin: Full system access and configuration -- Manager: Dashboard access and basic device management -- Technician: Read-only access with device status information -- Customer: Limited access to their own devices - -**Audit Logging** -- Comprehensive logging of all user actions -- Device interaction history -- System configuration changes -- Security event tracking - -## Technical Implementation - -### Performance Optimization - -**Data Handling Strategies** -- Implemented data virtualization for large device lists -- Used efficient data structures for real-time updates -- Implemented smart caching with Redis -- Optimized database queries with proper indexing - -**UI Optimization** -```typescript -// Memoized device list component for performance -const DeviceList = React.memo(({ devices, filters }) => { - const filteredDevices = useMemo(() => { - return devices.filter(device => - matchesFilters(device, filters) - ); - }, [devices, filters]); - - const virtualizedList = useVirtualization(filteredDevices, { - itemHeight: 60, - containerHeight: 400, - }); - - return ( -
- {virtualizedList.items.map(device => ( - - ))} -
- ); -}); -``` - -### Real-time Data Management - -**WebSocket Integration** -- Efficient connection management for multiple data streams -- Automatic reconnection handling for network interruptions -- Message queuing for offline scenarios -- Bandwidth optimization with data compression - -**State Management** -- Redux store structure optimized for real-time updates -- Normalized data structures for efficient updates -- Middleware for handling asynchronous device operations -- Optimistic updates with rollback capabilities - -### Responsive Design - -**Mobile-First Approach** -- Responsive layouts using Bootstrap grid system -- Touch-friendly interface elements -- Optimized charts for mobile viewing -- Offline capabilities for field technicians - -**Progressive Web App Features** -- Service worker for offline functionality -- Push notifications for critical alerts -- App-like experience on mobile devices -- Background sync capabilities - -## Key Features in Detail - -### Device Health Monitoring - -**Health Score Algorithm** -```typescript -const calculateDeviceHealth = (device: Device): HealthScore => { - const factors = { - connectivity: device.lastSeen < 5 * 60 * 1000 ? 100 : 0, // 5 minutes - battery: device.batteryLevel || 100, - errorRate: Math.max(0, 100 - (device.errorCount / device.totalRequests) * 100), - uptime: (device.uptime / device.totalTime) * 100, - }; - - const weightedScore = - factors.connectivity * 0.3 + - factors.battery * 0.2 + - factors.errorRate * 0.3 + - factors.uptime * 0.2; +The admin console is the human face of the Zimi IoT platform. I built it in React, Redux, and TypeScript so operators, support, and field technicians can see device health, telemetry, and usage without SSHing into anything. It is one container in [the C4 model](/projects/iot-platform-architecture), sitting on the same APIs as the rest of the backend. - return { - score: Math.round(weightedScore), - factors, - status: getHealthStatus(weightedScore), - }; -}; -``` +## Operators cannot stare at 70,000 rows -### Advanced Analytics +A fleet that size is not a table you sort. Most devices are fine most of the time. The job is the exceptions: a gateway that stopped talking, a network with a high error rate, a device whose power draw looks wrong, a customer who says “the switch does nothing.” -**Usage Pattern Recognition** -- Machine learning algorithms for pattern detection -- Anomaly detection for unusual device behavior -- Predictive analytics for maintenance scheduling -- Energy efficiency optimization recommendations +The other audience problem is mixed. Admins configure the system. Support needs one household, fast. Technicians need status in the field, on a phone. Customers should see only their own devices. One UI with no roles would have leaked data or drowned people in the wrong knobs. -**Custom Report Generation** -- Automated report generation for management -- Customizable report templates -- Scheduled report delivery via email -- Export capabilities (PDF, Excel, CSV) +And the data is chatty. Telemetry and connectivity events arrive continuously. If the browser subscribed to everything, it would fall over; if it polled everything, the API would. -## Results & Impact +## What changed -### Operational efficiency -- Faster device troubleshooting for operators -- Proactive monitoring instead of waiting for field reports -- Better visibility into device health and network status +The console is a set of questions, not a fleet dump: -### Business benefits -- Support teams can see telemetry and usage without a device visit -- Automated monitoring and alerts -- Analytics for product and operations decisions -- Real-time visibility into device performance +- Which devices or networks are unhealthy *right now*? +- Where is this device, and when did it last talk? +- What has it done in a chosen window? +- How is this household using power and loads over time? +- Who is allowed to change what? -### Technical achievements -- Interactive Chart.js visualisations of time-series telemetry -- Built for a 70,000+ device fleet -- React/Redux administration portal with role-based access -- Live updates for operators +An **interactive map** sits next to the lists: location plus live status, so support can see a silent gateway in a suburb instead of hunting a serial number. Status itself is last-seen, connectivity, and — for wireless devices — battery. Those three answer “is it even reachable?” before anyone opens a telemetry chart. -## User Experience Design +Lists and the map share the same filters: device type, location, status, and a date range for history. Paginated, and virtualized where a list can still get long. Charts use Chart.js for the usual time-series KPIs; heavier custom views use D3 when a stock chart is the wrong shape. Redis sits behind the API so common lookups are not a database round-trip every time an operator opens a page. -### Interface Design Principles -- **Clarity**: Clean, uncluttered interface focusing on essential information -- **Consistency**: Standardized UI components and interaction patterns -- **Efficiency**: Quick access to frequently used features and information -- **Accessibility**: WCAG 2.1 compliant design for inclusive access +Live updates go over WebSockets for the things that actually change while someone is looking: online/offline, last-seen, battery, in-progress commands. Historical charts load over REST for a time range. That split keeps the socket from becoming a second copy of PostgreSQL. -### Information Architecture -- **Hierarchical Navigation**: Logical grouping of features and functions -- **Contextual Actions**: Relevant actions available based on current context -- **Progressive Disclosure**: Complex information revealed gradually as needed -- **Customizable Dashboards**: User-configurable layouts and widgets +Access is role-based. Admin configuration, support tooling, technician read-only status, and customer-scoped views share components but not data. Actions are audited. -## Lessons Learned +The UI is usable on a phone because a lot of the work happens in a switchboard or a hallway, not at a desk. -### Performance Considerations -- Real-time data requires careful balance between update frequency and performance -- Data visualization libraries need optimization for large datasets -- Browser performance varies significantly with complex charts and animations -- Caching strategies are crucial for responsive user experience +## Live where it matters, aggregate everywhere else -### User Experience Insights -- Field technicians need simplified, task-focused interfaces -- Management users prefer high-level dashboards with drill-down capabilities -- Mobile responsiveness is essential for field operations -- Offline capabilities significantly improve user satisfaction +The tempting design is “stream every register to the browser.” It looks impressive in a demo and dies in production. -### Technical Architecture -- Modular component design enables easy feature additions and modifications -- State management complexity increases with real-time data requirements -- WebSocket connection management requires robust error handling and recovery -- Database optimization is critical for dashboard performance at scale +What support actually needs is an aggregated health picture plus the ability to open *one* network and *one* device. Aggregation and indexing belong on the server. The client holds a window: current filters, a selected entity, a live subscription for that entity. -## Future Enhancements +That also decides the Redux shape. Normalise devices and networks so a status event is one update, not a rewrite of a giant tree. Optimistic command UI is fine; the source of truth is still the backend event that comes back after the gateway reports. -### Advanced Analytics -- Machine learning integration for predictive maintenance -- AI-powered anomaly detection and alerting -- Advanced forecasting for capacity planning -- Integration with external weather and environmental data +I did not put machine-learning “predictive maintenance” in this console. Health is last-seen, errors, connectivity, and the telemetry we already store. Alerts come from those signals. -### Enhanced Visualization -- 3D visualization for complex network topologies -- Augmented reality interface for field technicians -- Voice-controlled dashboard navigation -- Advanced geospatial analysis and mapping +## After -### Integration Capabilities -- Third-party system integrations (ERP, CRM, ITSM) -- API ecosystem for external developers -- Webhook support for custom integrations -- Export capabilities to business intelligence tools +Operators can see fleet and household state without a site visit. Support uses telemetry and usage charts instead of guessing. The same role-aware APIs feed this UI and the rest of the platform. -This comprehensive dashboard platform has become the cornerstone of IoT device management operations, providing stakeholders with the insights and tools necessary to maintain high-performance device networks while optimizing operational efficiency and customer satisfaction. +Related: [migration](/projects/iot-platform-migration) (the backend this talks to), [C4 container view](/projects/iot-platform-architecture), [data lifecycle](/projects/gcp-cost-optimization) for the history those charts read. diff --git a/src/content/projects/iot-platform-architecture.mdx b/src/content/projects/iot-platform-architecture.mdx index d0b777f..50d1cbe 100644 --- a/src/content/projects/iot-platform-architecture.mdx +++ b/src/content/projects/iot-platform-architecture.mdx @@ -1,6 +1,6 @@ --- title: "Diagramming a Smart-Home IoT Platform with C4" -description: "A C4 / LikeC4 model of a production smart-home IoT platform: context, containers, components, flows and deployment for a live estate of 70,000+ devices." +description: "A C4 / LikeC4 model of the production smart-home platform I have owned since 2017: context, containers, flows and deployment for 70,000+ devices." period: "2017–Present" status: "ongoing" featured: true @@ -23,20 +23,16 @@ technologies: - "React" achievements: - "Owned the production backend and cloud platform from design through operations" - - "Zero-downtime migration from a hosted IoT PaaS onto service-oriented GCP" - - "70,000+ devices and 10,000+ accounts at 99.999% uptime and 100+ events/s" - - "Certified Alexa and Google Home fulfilment with OAuth account linking" - "C4 model covering context, containers, components, flows and deployment" + - "70,000+ devices and 10,000+ accounts at 99.999% uptime and 100+ events/s" challenges: - - "Keeping a live device fleet online while changing brokers, stores and voice integrations" - - "One decode path for a mixed electrical fleet with device-class-specific registers" - - "Role-aware APIs that separate unrestricted admin work from ownership-scoped customer access" - - "Keeping container and flow views aligned as the platform grew" + - "Keeping a live fleet online while brokers, stores and voice integrations changed" + - "One decode path for a mixed electrical fleet" + - "Admin work and customer access must not share a permission model" outcomes: - - "A maintainable service-oriented platform that is still the system of record" - - "Reusable library for auth, RBAC, MQTT and domain models across four services" - - "Telemetry, voice, CRM, OTA and onboarding as separate containers with clear contracts" - - "A LikeC4 model covering context, containers, components, flows and deployment" + - "A service-oriented platform that is still the system of record" + - "Shared library for auth, RBAC, MQTT and domain models across services" + - "Telemetry, voice, CRM, OTA and onboarding as separate containers with contracts" image: /images/projects/c4/index.png imageFit: contain gallery: @@ -56,25 +52,31 @@ tags: - "Architecture" category: "iot" publishDate: 2026-08-26 -lastUpdated: 2026-08-26 +lastUpdated: 2026-08-27 --- import LikeC4Embed from '../../components/LikeC4Embed.astro' ## What this model is -This project is a [C4](https://c4model.com/) model of a production smart-home / smart-electrical platform I have owned since 2017. The diagrams are the work: one [LikeC4](https://likec4.dev/) source file, several views you can drill into. +This is a [C4](https://c4model.com/) model of the production smart-home / smart-electrical platform I have owned since 2017. The diagrams *are* the work: one [LikeC4](https://likec4.dev/) source file, several views you can drill into. -The live estate behind the pictures is **70,000+ devices**, **10,000+ accounts**, **100+ events per second** and **99.999% uptime**. The system in the model is the **Smart Device IoT Platform**. Around it sit a **Home Gateway**, a **Platform API**, and the neighbouring services the household actually talks to. +The live estate behind the pictures is **70,000+ devices**, **10,000+ accounts**, **100+ events per second** and **99.999% uptime**. The system in the model is the **Smart Device IoT Platform**. Around it sit a **Home Gateway**, a **Platform API**, and the neighbouring services a household actually talks to. - Live explorer: [nilushan.github.io/nilushan-projects-c4](https://nilushan.github.io/nilushan-projects-c4/) - Source: [github.com/nilushan/nilushan-projects-c4](https://github.com/nilushan/nilushan-projects-c4) -Related write-ups: the [zero-downtime migration](/projects/iot-platform-migration), [voice ecosystem](/projects/voice-control-ecosystem), [admin dashboard](/projects/iot-dashboard) and [C4 / LikeC4 notes](/blog/c4-architecture-diagrams). +How it was built, as a method: [C4 / LikeC4 notes](/blog/c4-architecture-diagrams). + +## How to read it + +Start wide and click in. Context is who is around the platform. Containers are what we deploy. The device pipeline, identity, voice and data views are slices of the same model, not separate drawings. Flows are request-time pictures. Deployment is where those containers run. + +Matter manufacturing is marked **designed, not shipped**. I do not claim a payment gateway, a GraphQL API, or Apple Home as a cloud fulfilment integration I built. ## System context -People around the platform: the **customer**, the **installer** who commissions a household and hands it over, and internal **admin / support / operator** roles. Neighbouring systems are a managed identity provider, an MQTT broker, Google Home, Alexa, mobile push, transactional email and a customer-engagement platform. +People: the **customer**, the **installer** who commissions a household, and internal **admin / support / operator** roles. Neighbours: a managed identity provider, an MQTT broker, Google Home, Alexa, mobile push, transactional email, a customer-engagement platform. -The home side is a **local device mesh** plus a **home gateway**. The gateway is the only thing that speaks MQTT to the cloud. Devices themselves stay on the LAN. +The home side is a **local device mesh** plus a **home gateway**. The gateway is the only thing that speaks MQTT to the cloud. Devices stay on the LAN. ## Core containers -The slide-sized container view. Clients on the left, the API and shared library in the middle, stores on the right. +Clients on the left, API and shared library in the middle, stores on the right. -What this view is careful to show: +Worth noticing: -- a **shared platform library** (RBAC dispatch, typed document access, MQTT broker port, domain models) +- a **shared platform library** — RBAC dispatch, typed document access, MQTT broker port, domain models - an **invitation portal** for installer → customer handover -- **Redis** as cache, pub/sub backplane and delayed-job store — not only a cache -- an **OAuth authorization server** used to link Alexa and Google Home, not a generic “identity box” +- **Redis** as cache, pub/sub backplane *and* delayed-job store +- an **OAuth authorization server** used to link Alexa and Google Home, not a generic identity box ## Device telemetry and control @@ -114,13 +116,9 @@ Unsolicited telemetry and cloud-to-device commands share one decode path. height="34rem" /> -- The gateway publishes state, requests and schedules onto the MQTT broker. -- A **device event bus** fans those messages to pull subscribers. -- A **register decoder** turns a packed binary protocol into typed state. Dimmers, fans, garage doors and sensors expose different registers; every type also reports power. -- Snapshot goes to the document store, history and power to partitioned PostgreSQL, hot state to Redis. -- Redis pub/sub feeds a **WebSocket realtime gateway** so apps do not poll. -- User-visible register changes are **reported** to Alexa and Google Home. -- Gateway requests (provisioning, keys, CRC, schedules) are **request/response**, not fire-and-forget. +The gateway publishes state, requests and schedules onto the MQTT broker. A **device event bus** fans those messages to pull subscribers. A **register decoder** turns a packed binary protocol into typed state — dimmers, fans, garage doors and sensors expose different registers; every type also reports power. + +Snapshot goes to the document store, history and power to partitioned PostgreSQL, hot state to Redis. Redis pub/sub feeds a **WebSocket realtime gateway** so apps do not poll. User-visible register changes are **reported** to Alexa and Google Home. Gateway requests (provisioning, keys, CRC, schedules) are **request/response**, not fire-and-forget. ## Identity and access @@ -186,8 +184,25 @@ The Matter manufacturing cloud — factory registration, catalogues, quotas, ser ![Designed only — Matter manufacturing cloud](../../assets/projects/c4/matter-manufacturing.png) -I also do not claim a payment gateway, a GraphQL API, or Apple Home as a cloud fulfilment integration I built. Apple Home appears only as a local Matter path on the designed landscape. +Apple Home appears only as a local Matter path on that designed-only view. + +## What the diagrams are arguing + +A few boundaries that the pictures are there to keep honest: + +- **The gateway is the cloud-facing device.** Nothing on the LAN speaks MQTT northbound. That keeps the fleet protocol in one place and keeps individual loads off the public internet. +- **One decode path.** Telemetry and commands share a register model. Two decoders would be two device types. +- **Identity is split on purpose.** App JWT + RBAC is not the same as voice OAuth linking. Mixing them would either over-privilege an assistant or make account linking look like a user session. +- **Redis is a backplane, not a cache box.** If you only draw it as cache, you miss pub/sub into WebSockets and the delayed jobs that duration-based device behaviour depends on. +- **History is a lifecycle.** Snapshots are documents; time-series is partitioned SQL that can be archived. Treating both as “the database” is how the hot path inherits last year’s power samples. + +## Related implementation work -## How the model is kept +- [Zero-downtime Xively → GCP migration](/projects/iot-platform-migration) — how this platform got onto GCP +- [Admin dashboard](/projects/iot-dashboard) — the operator console on these APIs +- [Voice assistants](/projects/voice-control-ecosystem) — translators, linking, report-state +- [Partner API](/projects/third-party-api-platform) — OAuth and events for third parties +- [Keyless CI/CD](/projects/cicd-keyless-delivery) — artifact promotion and Workload Identity Federation +- [Database and logging cost](/projects/gcp-cost-optimization) — monthly partitions and archive-and-drop -The LikeC4 source lives in its own repo, [nilushan/nilushan-projects-c4](https://github.com/nilushan/nilushan-projects-c4). Views are projections of one model: context, containers, components, dynamic flows and deployment. Pushes to `main` publish the [interactive explorer](https://nilushan.github.io/nilushan-projects-c4/) on GitHub Pages. +The LikeC4 source lives in [nilushan/nilushan-projects-c4](https://github.com/nilushan/nilushan-projects-c4). Pushes to `main` publish the [interactive explorer](https://nilushan.github.io/nilushan-projects-c4/). diff --git a/src/content/projects/iot-platform-migration.mdx b/src/content/projects/iot-platform-migration.mdx index 239659e..4030f73 100644 --- a/src/content/projects/iot-platform-migration.mdx +++ b/src/content/projects/iot-platform-migration.mdx @@ -1,6 +1,6 @@ --- title: "Enterprise IoT Cloud Platform Migration" -description: "Zero-downtime live migration of a consumer IoT platform from Xively to Google Cloud Platform. The platform now serves 70,000+ devices and 10,000+ user accounts at 99.999% uptime." +description: "Zero-downtime move of a live consumer IoT backend off Xively onto GCP after Google acquired Xively and later discontinued it. Same platform now serving 70,000+ devices and 10,000+ accounts at 99.999% uptime." period: "2017-2020" status: "completed" featured: true @@ -19,28 +19,24 @@ technologies: - "Express" - "Fastify" - "Docker" + - "Cloud Build" + - "Cloud Storage" + - "Cloud Monitoring" + - "Cloud Trace" - "Cloud Functions" achievements: - "Zero-downtime live migration from Xively to GCP" - - "Platform now serving 70,000+ devices and 10,000+ user accounts" - - "Lowered infrastructure cost and request latency" - - "99.999% uptime after the migration" - - "Successful migration of all users, apps, and voice control integrations" - - "Seamless transition with comprehensive documentation" + - "Platform now serving 70,000+ devices and 10,000+ user accounts at 99.999% uptime" + - "Apps, devices, and voice integrations kept working through the cutover" challenges: - - "Migrating live production system with zero downtime" - - "Maintaining data integrity across platforms" - - "Ensuring device connectivity during transition" - - "Coordinating with multiple development teams" - - "Managing complex system dependencies" + - "Devices in the field cannot take a maintenance window" + - "Xively was acquired by Google and later discontinued; the fleet had to leave a dying PaaS" + - "Mobile apps, firmware, and voice skills had to keep talking to something during the move" outcomes: - - "Improved system reliability and performance" - - "Significant cost savings for the organization" - - "Enhanced scalability for future growth" - - "Better monitoring and observability" - - "Streamlined development and deployment processes" + - "We own the backend instead of renting an IoT PaaS" + - "Owning MQTT, identity, and storage meant we were no longer waiting on a vendor roadmap" + - "The GCP system is still the system of record" image: /images/projects/iotmigration.png - client: "Zimi Ltd" role: "Full Stack Developer" tags: @@ -51,131 +47,51 @@ tags: - "Real-time Systems" category: "cloud" publishDate: 2020-12-01 -lastUpdated: 2024-01-15 +lastUpdated: 2026-08-27 --- -## Project Overview - -The Enterprise IoT Cloud Platform Migration was a critical infrastructure project that involved migrating a large-scale IoT system from the Xively platform to Google Cloud Platform (GCP). This project was essential for improving system performance, reducing costs, and ensuring long-term scalability for smart electrical devices. - -## The Challenge - -The existing system was built on the Xively IoT platform, which had limitations in terms of scalability, cost-effectiveness, and customization options. With a live fleet that has since grown to 70,000+ devices and more than 100 events per second, we needed a solution that could: - -- Handle high-volume real-time data processing -- Provide better cost efficiency -- Offer improved latency and reliability -- Support future growth and feature development -- Maintain zero downtime during migration - -## Architecture & Design - -The current platform is documented as a C4 model — context, containers, telemetry pipeline, voice, data and deployment — in [Diagramming a Smart-Home IoT Platform with C4](/projects/iot-platform-architecture). - -### Service-Oriented Architecture -I designed a microservices-based architecture on GCP that emphasized: -- **Scalability**: Auto-scaling components based on demand -- **Reliability**: Redundancy and fault tolerance built-in -- **Security**: End-to-end encryption and secure device communication -- **Maintainability**: Clear separation of concerns and comprehensive documentation -- **Observability**: Detailed monitoring and tracing capabilities +From 2017 I was the full-stack and cloud engineer for Bluekey’s (later Zimi’s) smart-electrical platform. The first backend sat on [Xively](https://en.wikipedia.org/wiki/Xively): device identity, MQTT, a way to ship. Google later acquired Xively and then discontinued it. I designed the GCP replacement, built the services, and moved a live fleet off that PaaS before it went away. That GCP system is still what the field talks to. -### Technology Stack -- **Cloud Platform**: Google Cloud Platform -- **Container Orchestration**: Kubernetes for scalable deployment -- **Database**: Cloud SQL (PostgreSQL) for relational data, Redis for caching -- **Messaging**: PubSub for event-driven communication -- **Authentication**: Firebase Auth for user management -- **Storage**: Firestore for device metadata, Cloud Storage for files -- **Monitoring**: Cloud Monitoring and Trace for observability +## The PaaS under the fleet was going away -## Implementation +This was not a greenfield “let’s try Kubernetes” project. Xively was the IoT contract — broker, device identity, the path from a home gateway to the cloud — and that contract was ending. Staying meant betting a product already in people’s walls on a discontinued platform. -### Phase 1: Research & Planning -- Analyzed existing system architecture and dependencies -- Designed new microservices architecture -- Created comprehensive migration plan -- Set up development and staging environments +Moving meant owning MQTT, identity, and storage ourselves, on a deadline we did not set. **The devices were already installed.** Gateways in homes, mobile apps in pockets, voice skills already linked. A weekend cutover was not an option. Data had to stay consistent. MQTT sessions had to keep working. Alexa and Google Home could not quietly break. -### Phase 2: Development -- Built new backend services using TypeScript and Node.js -- Implemented REST APIs and event-driven services -- Developed device communication protocols using MQTT -- Created admin dashboards with React/Redux -- Established CI/CD pipelines with Docker and Cloud Build +## What changed -### Phase 3: Testing & Validation -- Performed extensive testing in staging environment -- Validated data integrity and system performance -- Conducted load testing with simulated device traffic -- Verified all integrations including voice control systems +I built a service-oriented backend on GCP and moved traffic onto it in batches rather than in one freeze. -### Phase 4: Migration Execution -- Executed zero-downtime migration strategy -- Gradually migrated devices in batches -- Monitored system health throughout the process -- Coordinated with mobile app and firmware development teams +The replacement had to cover everything Xively had been doing for us: -## Technical Highlights +- MQTT between home gateways and the cloud +- Device identity and registry +- APIs for apps and the admin console +- Device metadata and the latest snapshot +- Event fan-out for telemetry and commands +- Auth for customers, installers, and operators +- File storage (firmware images, exports, other blobs) -### Real-time Event Processing -Implemented a robust event processing system capable of handling 100+ events per second from IoT devices, with automatic scaling based on load. +On GCP that became TypeScript/Node services on Kubernetes. **IoT Core** was the device registry and MQTT path that replaced Xively’s broker — not a logo on a slide. Gateways authenticated to that channel; traffic was encrypted. We owned the topic contract and the credentials instead of inheriting a PaaS default. -### Device Communication Security -Established secure MQTT communication channels with proper authentication and encryption, ensuring device data integrity and privacy. +Firestore held document state, PostgreSQL history, Redis hot paths, Pub/Sub events, Cloud Storage files. Firebase Auth handled user identity. Cloud Monitoring and Trace were how we watched the cutover, not a post-launch extra. I put shared MQTT, auth, and domain code in libraries so the new services did not each invent a protocol. -### Monitoring & Observability -Integrated comprehensive monitoring using Cloud Monitoring, custom metrics, and distributed tracing to ensure system health and performance visibility. +The [React/Redux admin console](/projects/iot-dashboard) was built in the same effort: operators needed a place to see the new backend while devices were still moving. CI ran through Docker and Cloud Build. -### Development Efficiency -Created reusable TypeScript/Node.js libraries containing common functionality, reducing duplication and speeding delivery of new features. +Staging ran load against simulated device traffic and the existing voice integrations before any household moved. Cutover was gradual: batches of devices, watching connectivity and error rates, coordinating with the mobile and firmware teams so field software and cloud moved in step. -## Results & Impact +The current shape of that system — containers, telemetry pipeline, voice, data, deployment — is in [the C4 model](/projects/iot-platform-architecture). -The migration project delivered significant improvements across multiple dimensions: +## Cut over in batches without a freeze -**Cost and operations** -- Lowered infrastructure cost through the Xively-to-GCP redesign -- Removed the previous platform licence constraint -- Tightened compute and storage usage +A big-bang DNS flip would have been simpler to describe and harder to survive. If the new broker, the new stores, and the new APIs were even slightly wrong, every silent device would have been a support call. -**Performance** -- Lowered request latency after the cutover -- Improved responsiveness for apps and operators -- Kept real-time device event processing in place at 100+ events per second +Batching meant the blast radius of a mistake was a slice of the estate, not the estate. It also forced the new stack to run *next to* the old one: same app versions, same firmware, two backends, until we were willing to stop the old one. That is slower than a rewrite-and-launch. It is how you move a live IoT product. -**Reliability** -- Maintained 99.999% uptime after the migration -- Switched backends, databases, users, devices and voice integrations of a live system -- Improved disaster-recovery planning +The unglamorous work was the contract with the field: what a gateway publishes, what it expects back, how a missed message is retried. Cloud diagrams do not matter if a dimmer stops reporting. -**Delivery** -- Streamlined development with modern CI/CD -- Faster, repeatable deployments -- Better testing and quality-assurance workflows +## After -## Knowledge Transfer & Documentation +The platform left Xively without a customer-facing outage, before the discontinued PaaS could take the fleet with it. It now serves **70,000+ devices** and **10,000+ accounts**, handles **100+ events/s**, and runs at **99.999% uptime**. Once we owned the stack we could also size compute and storage ourselves, which brought cost and request latency down — that was a consequence of having to leave, not the reason we left. -Recognizing the importance of knowledge sharing, I: -- Documented the entire system architecture with detailed diagrams -- Created comprehensive API documentation -- Conducted knowledge transfer sessions with the development team -- Established best practices for ongoing maintenance and development - -## Lessons Learned - -This project reinforced several key principles: -- **Planning is Critical**: Thorough planning and testing prevented major issues during migration -- **Communication is Key**: Regular coordination with all stakeholders ensured smooth execution -- **Monitoring is Essential**: Comprehensive monitoring helped identify and resolve issues quickly -- **Documentation Matters**: Detailed documentation facilitated team knowledge transfer and future development - -## Future Considerations - -The new architecture positions the platform for future enhancements: -- Capability to handle millions of devices -- Support for advanced analytics and machine learning -- Integration with additional IoT protocols and standards -- Enhanced real-time dashboard capabilities - -This project demonstrates the successful execution of a complex cloud migration while maintaining business continuity and achieving significant improvements in performance, cost, and reliability. +I still operate this backend. Later work on the same system — [voice](/projects/voice-control-ecosystem), [admin dashboard](/projects/iot-dashboard), [partner API](/projects/third-party-api-platform), [CI/CD](/projects/cicd-keyless-delivery), [data lifecycle](/projects/gcp-cost-optimization) — is evolution of this cutover, not a second platform. diff --git a/src/content/projects/smartsmb-agentic-workflow.mdx b/src/content/projects/smartsmb-agentic-workflow.mdx index 16f4ba7..38d9e55 100644 --- a/src/content/projects/smartsmb-agentic-workflow.mdx +++ b/src/content/projects/smartsmb-agentic-workflow.mdx @@ -1,6 +1,6 @@ --- title: "SmartSMB Human-in-the-Loop AI Workflow" -description: "A non-production personal learning project exploring stateful agent orchestration for small businesses: enquiry triage, tool-assisted quoting, deterministic pricing, durable state, and human approval before action." +description: "A learning system for small-business enquiries: LangGraph routes the job, code calculates the quote, PostgreSQL holds state, and quotes above a configurable threshold require approval." period: "2026" status: "completed" featured: false @@ -16,20 +16,18 @@ technologies: - "LLM Tool Calling" - "Human-in-the-Loop" achievements: - - "Built a stateful workflow that routes quotes, complaints, FAQs, and unclear enquiries to specialist paths or human handoff" - - "Kept quote calculations deterministic in code while using the LLM to gather details and select validated tools" - - "Implemented approve, revise, and reject controls so consequential actions remain under human supervision" - - "Persisted graph state in PostgreSQL so paused conversations survive restarts and resume safely" - - "Created a swappable model layer and comprehensive deterministic tests for agent behaviour" + - "Quote / complaint / FAQ / handoff routing in a stateful graph, not a single prompt" + - "Prices calculated in code; the model only gathers fields and picks tools" + - "Approve, revise, reject as graph interrupts — revisions go back through approval" + - "PostgreSQL checkpoints so a paused thread survives a process restart without double-sending" challenges: - - "Combining flexible language understanding with predictable business decisions" - - "Pausing and resuming a multi-step workflow safely across process restarts" - - "Preventing duplicate delivery when a checkpointed workflow is replayed" - - "Allowing human revisions without bypassing the approval gate" + - "Language is messy; prices and sends must not be" + - "A workflow that waits overnight has to resume from disk, not from RAM" + - "Replay after a checkpoint must not email the customer twice" outcomes: - - "A working learning implementation of controlled agentic patterns for small-business workflows" - - "Practical experience with LangGraph state, conditional routing, tools, interrupts, and persistence" - - "A testable architecture where models are swappable and critical calculations remain deterministic" + - "A small, testable example of controlled agentic patterns" + - "Swappable chat models behind the same graph and business rules" + - "Deterministic tests around routing, tools, and approval — not 'the LLM usually does this'" github: "https://github.com/nilushan/langgraph-frontdesk-agent" role: "Personal Project — Architecture and Implementation" tags: @@ -40,84 +38,61 @@ tags: - "Human-in-the-Loop" category: "ai" publishDate: 2026-07-01 -lastUpdated: 2026-08-11 +lastUpdated: 2026-08-27 --- -## Scope and Purpose +## Why I built this -SmartSMB is a **non-production personal learning project** built to understand how agentic and multi-agent patterns can support common small-business workflows without handing uncontrolled authority to an LLM. +I wanted to learn how far a stateful agent can go on a small-business front desk *without* giving the model the company chequebook. The question was not “can an LLM talk like a receptionist.” It was: can it triage an enquiry, fill a quote, and require a person to approve amounts above a configurable threshold. -The sample receives a customer enquiry, identifies its intent, gathers the information needed for a quote, calculates the price in normal application code, and pauses for a person to approve, revise, or reject higher-value work before anything is sent. +This is a personal learning project. It is not running a business. -Although deliberately compact, the project treats the agent as a software system: state is typed, decisions are testable, side effects are idempotent, model providers are replaceable, and human approval is part of the workflow rather than an afterthought. +## Scope -## Workflow +The sample accepts a customer message and routes it: -A customer message can be routed to several specialist paths: +- **Quote** — collect job details, call pricing/availability tools, compute a price in application code, and pause when the amount exceeds a configurable approval threshold +- **Complaint** — dedicated path, not mixed into quoting +- **FAQ** — common information +- **Unclear** — hand off instead of guessing -- **Quote:** gather job details, consult pricing and availability tools, calculate a quote, and request approval when required -- **Complaint:** direct the message to a dedicated complaint path -- **FAQ:** handle common information requests -- **Unclear or low-confidence enquiry:** hand off to a person rather than guessing +What it does not do: production auth, real telephony, a live company knowledge base, or real customer-channel delivery. Quotes at or below the configured threshold follow the automatic send path; larger quotes pause for an operator. Models are swappable; the graph, schemas, and side effects are the system. -The quote path uses a ReAct-style loop in which the model can call validated tools. Graph routing is handled by explicit application logic, and the full workflow state is checkpointed after each step. +Repo: [nilushan/langgraph-frontdesk-agent](https://github.com/nilushan/langgraph-frontdesk-agent). -## Deterministic Decisions Around Probabilistic Models +## Core workflow -The LLM is useful for interpreting an unstructured request and gathering missing details, but it does not get to invent the final price. A dedicated node calculates the amount from structured inputs in code. +A quote thread looks like this: -This boundary demonstrates a pattern I consider important for practical AI systems: +1. Classify the message. +2. If it is a quote, loop: ask for missing fields, call Zod-validated tools (pricing, availability). +3. A **code node** turns those fields into a number. The model does not write the total. +4. Above the configured threshold, LangGraph **interrupts**. An operator approves, revises the structured fields, or rejects; quotes at or below it continue to delivery automatically. +5. A revision cannot jump to the customer. It re-enters approval, with who changed what. +6. Delivery is idempotent on quote identity, then the thread continues. -- Use models where language understanding and flexible reasoning add value -- Use schemas to validate model and tool inputs -- Keep financial rules and other critical decisions deterministic -- Route low-confidence or exceptional cases to a person -- Test the deterministic boundaries independently of the model provider +State is checkpointed in PostgreSQL after each step. Restart the process while someone is thinking; the thread is still there. -## Human-in-the-Loop Approval +## Design choices -Quotes over a configurable threshold cause the LangGraph workflow to interrupt and wait. An operator can: +**Probabilistic where language is, deterministic where money is.** The model is good at “they want the gutters done next Thursday.” It is not the pricing engine. Keeping that split means I can test quotes without mocking prose. -- **Approve** the quote as prepared -- **Revise** structured quote details and send the result back through approval -- **Reject** it and move the conversation to human follow-up +**Approval is a node, not a Slack side-channel.** If human review lives outside the graph, resume logic, audit, and “did we send this” all rot. Interrupts make the wait part of the program. -A revision never goes directly to the customer. It loops back through the approval gate, retaining a record of who changed the quote and what changed. This keeps a person accountable for consequential output while still allowing the agent to do useful preparation work. +**Checkpoints plus idempotent send.** Durable state without a delivery key is how you double-email after a retry. The send step is keyed on the quote. -## Durable State and Safe Resumption +**Tools are schemas, not “the model can call the database.”** Invalid arguments fail at the boundary. Routing between quote/complaint/FAQ is application code, not a vibe. -Conversation and workflow state is stored through a PostgreSQL checkpointer. A thread can pause while waiting for approval, survive an application restart, and resume from the recorded state. +I also used this repo to practise writing the spec and ADRs first, then implementing with coding agents, then treating the diff like production code: tests, failure paths, CI. Generated code that fails those is not a pass. -The delivery step is idempotent, keyed by quote identity, so replaying or resuming a workflow does not send the same approved quote twice. +## What I learned -## Model and Tool Architecture +A long-running front-desk job is a state machine that sometimes calls a model, not a chat log with extra steps. Once you admit that, persistence, interrupts, and idempotency stop being optional. -Agent nodes depend on a common chat-model abstraction rather than a hard-coded provider. This allowed me to explore how the same workflow can operate with different LLMs while preserving the surrounding business rules. +Specialist routes only help if each path has a different contract. Dumping complaints into the quote ReAct loop just makes a worse quote agent. -Tool calls use validated schemas for operations such as pricing and availability lookup. This keeps the tool boundary explicit and makes invalid input fail predictably instead of flowing into business services. +Model portability was easy *after* the graph and the price code stopped importing a vendor SDK. Before that, “swap the LLM” meant swapping the business rules by accident. -## Engineering Approach +## Not production -This project was also an exercise in agentic engineering practices: - -1. Write a specification and architecture decisions before implementation -2. Model workflow state, trust boundaries, and side effects explicitly -3. Use coding agents to assist with planning and implementation -4. Add deterministic unit and integration tests -5. Review diffs, failure paths, and generated behaviour as normal production-oriented code -6. Keep human ownership of architecture and acceptance decisions - -The repository includes architecture diagrams, ADRs, a runbook, CI checks, and a comprehensive automated test suite. - -## What I Learned - -- Stateful orchestration is more reliable than treating every interaction as an isolated prompt -- Human approval should be represented in workflow state and control flow, not handled informally outside the system -- Tool use needs schemas, permissions, and deterministic boundaries -- Durable checkpoints make long-running, interruptible workflows practical -- Multi-agent or specialist routing is most useful when each path has a clear responsibility -- Model portability is easier when orchestration and business logic do not depend directly on a provider - -## Production Considerations - -This is a learning implementation, not a claim of production operation. A production rollout would additionally require organisation-specific security review, privacy controls, threat modelling, model and prompt evaluation, rate limiting, production retrieval for company knowledge, channel integrations, monitoring, and operational ownership. +A real rollout would still need threat modelling, prompt-injection and tool-abuse controls, evaluation on real transcripts, rate limits, retrieval over actual company data, channel integration (email/SMS/phone), and someone on-call for the approval queue. None of that is claimed here. diff --git a/src/content/projects/third-party-api-platform.mdx b/src/content/projects/third-party-api-platform.mdx index 4452c4b..151c356 100644 --- a/src/content/projects/third-party-api-platform.mdx +++ b/src/content/projects/third-party-api-platform.mdx @@ -1,6 +1,6 @@ --- title: "Third-Party Integration API Platform" -description: "Comprehensive RESTful API platform for third-party integrations enabling device information retrieval, control, and real-time status events with OAuth2 security and event-driven architecture." +description: "Partner-facing REST API for Zimi devices: OAuth 2.0 user consent, scoped control, and signed event delivery so integrations do not poll the fleet." period: "2020-Present" status: "ongoing" featured: true @@ -19,24 +19,17 @@ technologies: - "OpenAPI" - "Swagger" achievements: - - "Designed and implemented comprehensive API platform" - - "OAuth2-based security with user authorization" - - "Event-driven architecture avoiding polling" - - "Real-time device status synchronization" - - "Comprehensive API documentation and testing tools" - - "High-performance caching and rate limiting" + - "OAuth 2.0 so a user grants a partner access to their own devices, not the fleet" + - "Device read/control APIs with rate limits and cached hot reads" + - "Webhook events for status changes instead of partner polling" challenges: - - "Designing scalable event-driven API architecture" - - "Implementing secure OAuth2 authorization flows" - - "Managing real-time synchronization across multiple systems" - - "Balancing API performance with security requirements" - - "Creating comprehensive documentation for external developers" + - "Partners needed a stable contract without seeing internal services" + - "Polling a 70,000-device estate from outside would have become our load" + - "Consent, scopes, and device ownership have to be right on every call" outcomes: - - "Enabled ecosystem expansion through partner integrations" - - "Faster partner onboarding against a documented API" - - "OAuth 2.0, rate limiting, caching and webhook delivery" - - "Facilitated integration with major smart home platforms" - - "Created new revenue streams through API partnerships" + - "Alexa, Google Home, and other partners share one authorised API surface" + - "Onboarding is an OpenAPI contract, not a guided tour of the monolith" + - "Internal MQTT and stores stay behind the platform API" client: "Zimi Ltd" role: "Lead API Architect & Developer" tags: @@ -47,517 +40,47 @@ tags: - "Integration Platform" category: "api" publishDate: 2023-01-01 -lastUpdated: 2024-06-01 +lastUpdated: 2026-08-27 image: /images/projects/socketapi.png - --- -## Project Overview - -The Third-Party Integration API Platform is a comprehensive RESTful API system designed to enable external partners and developers to integrate with the Zimi smart home ecosystem. This platform provides secure access to device information, control capabilities, and real-time status updates while maintaining the highest standards of security, performance, and reliability. - -## The Challenge - -Creating a robust API platform for third-party integrations required addressing several complex requirements: - -- **Security**: Implementing OAuth2 with granular user consent and authorization -- **Performance**: High-throughput API capable of handling thousands of concurrent requests -- **Real-time Communication**: Event-driven architecture for instant status updates -- **Scalability**: Platform designed to support hundreds of integration partners -- **Developer Experience**: Comprehensive documentation and testing tools -- **Reliability**: Monitoring, error handling and retries on a platform that runs at 99.999% uptime - -## System Architecture - -### Event-Driven API Design -Built around an event-driven architecture to eliminate polling and ensure real-time synchronization: - -```typescript -// Event-driven API architecture -interface APIEvent { - type: 'device_status_change' | 'device_added' | 'device_removed'; - deviceId: string; - userId: string; - timestamp: Date; - data: any; -} - -class EventBroker { - private pubsub: PubSub; - private subscriptions: Map; - - async publishEvent(event: APIEvent): Promise { - // Publish to internal PubSub for real-time processing - await this.pubsub.publish('api-events', event); - - // Notify subscribed webhooks - await this.notifyWebhooks(event); - } - - private async notifyWebhooks(event: APIEvent): Promise { - const webhooks = this.subscriptions.get(event.userId) || []; - - const notifications = webhooks.map(webhook => - this.sendWebhook(webhook, event) - ); - - await Promise.allSettled(notifications); - } -} -``` - -### Microservices Architecture -- **API Gateway**: Request routing, authentication, and rate limiting -- **Authorization Service**: OAuth2 implementation and token management -- **Device Service**: Device information and control operations -- **Event Service**: Real-time event processing and webhook delivery -- **Documentation Service**: Interactive API documentation and testing - -## OAuth2 Security Implementation - -### Authorization Flow -Implemented comprehensive OAuth2 authorization with user consent management: - -```typescript -class OAuth2Service { - async initiateAuthFlow(clientId: string, scopes: string[], redirectUri: string) { - // Validate client and redirect URI - const client = await this.validateClient(clientId, redirectUri); - - // Generate authorization code - const authCode = await this.generateAuthorizationCode({ - clientId, - scopes, - redirectUri, - expiresAt: new Date(Date.now() + 10 * 60 * 1000) // 10 minutes - }); - - return { - authorizationUrl: `${this.authEndpoint}?${new URLSearchParams({ - response_type: 'code', - client_id: clientId, - scope: scopes.join(' '), - redirect_uri: redirectUri, - code: authCode - })}` - }; - } - - async exchangeCodeForToken(code: string, clientId: string, clientSecret: string) { - // Validate authorization code - const authData = await this.validateAuthorizationCode(code); - - // Verify client credentials - await this.verifyClientCredentials(clientId, clientSecret); - - // Generate access and refresh tokens - const accessToken = await this.generateAccessToken({ - userId: authData.userId, - clientId, - scopes: authData.scopes, - expiresIn: 3600 // 1 hour - }); - - const refreshToken = await this.generateRefreshToken({ - userId: authData.userId, - clientId - }); - - return { - access_token: accessToken, - refresh_token: refreshToken, - token_type: 'Bearer', - expires_in: 3600, - scope: authData.scopes.join(' ') - }; - } -} -``` - -### Granular Permissions -Implemented fine-grained permission system allowing users to control exactly what data and capabilities they share: - -```typescript -interface APIScope { - name: string; - description: string; - resources: string[]; - actions: string[]; -} - -const API_SCOPES: APIScope[] = [ - { - name: 'devices:read', - description: 'Read basic device information', - resources: ['devices'], - actions: ['read'] - }, - { - name: 'devices:control', - description: 'Control device state (turn on/off, brightness, etc.)', - resources: ['devices'], - actions: ['write', 'execute'] - }, - { - name: 'events:subscribe', - description: 'Receive real-time device status updates', - resources: ['events'], - actions: ['subscribe'] - } -]; -``` - -## API Design & Implementation - -### RESTful Endpoints -Designed comprehensive REST API with OpenAPI 3.0 specification: - -```typescript -// Device control endpoint -@Route('/api/v1/devices/{deviceId}/actions') -export class DeviceActionsController { - @Post('/') - @Security('oauth2', ['devices:control']) - public async executeAction( - @Path() deviceId: string, - @Body() action: DeviceAction, - @Request() request: AuthenticatedRequest - ): Promise { - - // Validate user has access to device - await this.validateDeviceAccess(request.user.id, deviceId); - - // Execute action with timeout and error handling - const result = await this.deviceService.executeAction(deviceId, action); - - // Publish event for real-time updates - await this.eventBroker.publishEvent({ - type: 'device_status_change', - deviceId, - userId: request.user.id, - timestamp: new Date(), - data: result - }); - - return result; - } - - private async validateDeviceAccess(userId: string, deviceId: string): Promise { - const hasAccess = await this.authService.userHasDeviceAccess(userId, deviceId); - if (!hasAccess) { - throw new ForbiddenError('Access denied to device'); - } - } -} -``` - -### Real-time Event Delivery -Implemented webhook system for real-time event delivery to partner systems: - -```typescript -class WebhookDeliveryService { - private queue: Queue; - private retryPolicy: RetryPolicy; - - async deliverWebhook(webhook: WebhookSubscription, event: APIEvent): Promise { - const delivery: WebhookDelivery = { - webhookId: webhook.id, - url: webhook.url, - payload: this.formatEventPayload(event), - headers: { - 'Content-Type': 'application/json', - 'X-Webhook-Signature': this.signPayload(webhook.secret, event), - 'X-Event-Type': event.type, - 'X-Delivery-ID': this.generateDeliveryId() - }, - attempts: 0, - maxAttempts: 3 - }; - - await this.queue.add('webhook-delivery', delivery); - } - - private signPayload(secret: string, payload: any): string { - const data = JSON.stringify(payload); - return crypto - .createHmac('sha256', secret) - .update(data) - .digest('hex'); - } - - async processDelivery(delivery: WebhookDelivery): Promise { - try { - const response = await fetch(delivery.url, { - method: 'POST', - headers: delivery.headers, - body: JSON.stringify(delivery.payload), - timeout: 10000 // 10 second timeout - }); - - if (!response.ok) { - throw new Error(`HTTP ${response.status}: ${response.statusText}`); - } - - await this.recordSuccessfulDelivery(delivery); - } catch (error) { - await this.handleDeliveryFailure(delivery, error); - } - } -} -``` - -## Performance & Scalability - -### Caching Strategy -Implemented multi-layer caching for optimal performance: - -```typescript -class CachedDeviceService { - private cache: Redis; - private deviceService: DeviceService; - - async getDevice(deviceId: string, userId: string): Promise { - const cacheKey = `device:${deviceId}:${userId}`; - - // Try L1 cache (Redis) - const cached = await this.cache.get(cacheKey); - if (cached) { - return JSON.parse(cached); - } - - // Fallback to database - const device = await this.deviceService.getDevice(deviceId, userId); - - // Cache for 5 minutes - await this.cache.setex(cacheKey, 300, JSON.stringify(device)); - - return device; - } - - // Cache invalidation on device updates - async updateDevice(deviceId: string, userId: string, updates: DeviceUpdate): Promise { - const device = await this.deviceService.updateDevice(deviceId, userId, updates); - - // Invalidate cache - await this.cache.del(`device:${deviceId}:${userId}`); - - return device; - } -} -``` - -### Rate Limiting -Implemented sophisticated rate limiting to prevent abuse while allowing legitimate usage: - -```typescript -class RateLimiter { - private redis: Redis; - - async checkRateLimit(clientId: string, endpoint: string): Promise { - const key = `rate_limit:${clientId}:${endpoint}`; - const window = 3600; // 1 hour window - const limit = this.getLimitForEndpoint(endpoint); - - const multi = this.redis.multi(); - multi.incr(key); - multi.expire(key, window); - - const [count] = await multi.exec(); - - return { - allowed: count <= limit, - remaining: Math.max(0, limit - count), - resetTime: Date.now() + (window * 1000) - }; - } - - private getLimitForEndpoint(endpoint: string): number { - const limits = { - '/devices': 1000, // 1000 requests per hour - '/devices/actions': 500, // 500 actions per hour - '/events': 100 // 100 webhook subscriptions per hour - }; - - return limits[endpoint] || 100; // Default limit - } -} -``` - -## Developer Experience - -### Interactive Documentation -Created comprehensive API documentation using OpenAPI 3.0 with interactive testing: - -```yaml -# OpenAPI specification excerpt -paths: - /devices/{deviceId}/actions: - post: - summary: Execute device action - description: | - Execute an action on a specific device. Actions include turning devices - on/off, adjusting brightness, changing colors, etc. - parameters: - - name: deviceId - in: path - required: true - schema: - type: string - example: "device_12345" - requestBody: - required: true - content: - application/json: - schema: - $ref: '#/components/schemas/DeviceAction' - examples: - turn_on: - summary: Turn device on - value: - action: "turn_on" - set_brightness: - summary: Set brightness - value: - action: "set_brightness" - parameters: - brightness: 75 - responses: - '200': - description: Action executed successfully - content: - application/json: - schema: - $ref: '#/components/schemas/ActionResult' -``` - -### SDK Development -Created SDKs for popular programming languages to simplify integration: - -```typescript -// TypeScript SDK example -export class ZimiAPI { - private client: APIClient; - - constructor(accessToken: string) { - this.client = new APIClient({ - baseURL: 'https://api.zimi.com/v1', - accessToken - }); - } - - async getDevices(): Promise { - return this.client.get('/devices'); - } - - async controlDevice(deviceId: string, action: DeviceAction): Promise { - return this.client.post(`/devices/${deviceId}/actions`, action); - } - - async subscribeToEvents(webhookUrl: string, events: string[]): Promise { - return this.client.post('/events/subscriptions', { - url: webhookUrl, - events - }); - } -} -``` - -## Monitoring & Analytics - -### Comprehensive Monitoring -Implemented detailed monitoring and analytics for API usage: +Partners needed to read Zimi devices, change them, and hear about state changes. They did not need our MQTT broker, our document store, or our admin paths. I built the partner-facing REST API in TypeScript so third parties talk to a contract: OAuth 2.0, scoped resources, and events. -```typescript -class APIAnalytics { - async recordAPICall(request: APIRequest, response: APIResponse): Promise { - const metrics = { - endpoint: request.endpoint, - method: request.method, - clientId: request.clientId, - userId: request.userId, - responseTime: response.duration, - statusCode: response.statusCode, - timestamp: new Date() - }; +## Partners needed the devices, not the internals - // Store in time-series database for analytics - await this.timeSeriesDB.insert('api_metrics', metrics); +An integration that “just uses the internal API” couples their release cycle to ours and puts every internal handler on the public internet. The other failure mode is API keys that represent the company, not the household — a partner with a leaked key should not be able to reach devices a user never authorised. - // Update real-time counters - await this.redis.incr(`api_calls:${request.clientId}:${this.getCurrentHour()}`); - - // Track error rates - if (response.statusCode >= 400) { - await this.redis.incr(`api_errors:${request.clientId}:${this.getCurrentHour()}`); - } - } +Scale makes the obvious sync strategy lethal. If every partner polls every device they care about, their retry loops become our QPS. Status changes already exist as internal events. The API should forward the ones a subscriber is allowed to see. - async generateUsageReport(clientId: string, period: string): Promise { - const metrics = await this.timeSeriesDB.query(` - SELECT - endpoint, - COUNT(*) as request_count, - AVG(response_time) as avg_response_time, - COUNT(CASE WHEN status_code >= 400 THEN 1 END) as error_count - FROM api_metrics - WHERE client_id = ? AND timestamp >= ? - GROUP BY endpoint - `, [clientId, this.getPeriodStart(period)]); +This surface also had to survive the same production bar as the rest of the platform (99.999% uptime): timeouts, retries, and rate limits as part of the design, not as an afterthought. - return this.formatUsageReport(metrics); - } -} -``` +## What changed -## Results & Impact +The public API is a boundary in front of the platform: -### Business impact -- **Ecosystem**: Alexa, Google Home and other partner integrations on a shared API -- **Partner access**: OAuth 2.0-protected APIs so vendors can use customer and device data they are allowed to use -- **Onboarding**: Documented contracts, OpenAPI and knowledge-transfer material +- **OAuth 2.0 authorisation code** with user consent. The partner never gets a god-mode key for the fleet. +- **Scopes** such as device read, device control, and event subscription. A weather-display integration should not be able to toggle loads. +- **REST** for reads and commands, documented with OpenAPI / Swagger so onboarding is the spec, not a slide deck. +- **Ownership checks** on every call. A valid token is not enough; the user must actually have that device. +- **Redis** for hot device reads and for rate-limit counters per client and route. +- **Pub/Sub → signed webhooks** for `device_status_change` (and related) events, with retries on failed delivery. -### Technical achievements -- **Performance**: Caching, rate limiting and event-driven delivery instead of polling -- **Reliability**: Error handling and retries on the same production platform -- **Security**: OAuth 2.0 authorisation and role-based access +Voice assistants use the same idea — user linking and a narrow fulfilment path — and are covered separately in [the voice write-up](/projects/voice-control-ecosystem) and [the C4 identity/voice views](/projects/iot-platform-architecture). This API is the partner contract around device data and events; it is not a dump of every internal service. -### Developer experience -- **Documentation**: OpenAPI / Swagger for the API surface -- **Contracts**: Clear service boundaries for mobile, web and third-party clients -- **Delivery**: Webhooks so partners do not need to poll +## Push events, don't let them poll -## Lessons Learned +The decision that keeps this API from becoming a load generator is event delivery. -### API Design Principles -- **Developer-First Design**: APIs must be intuitive and well-documented -- **Versioning Strategy**: Plan for API evolution from the beginning -- **Error Handling**: Consistent, informative error responses are crucial -- **Rate Limiting**: Balance between preventing abuse and enabling legitimate use +Internal telemetry already fans out after decode. Once a partner has `events:subscribe` and a webhook, we push the authorised subset, signed so they can reject spoofed callbacks. They do not need a cron job walking `/devices`. -### Security Considerations -- **OAuth2 Complexity**: Proper OAuth2 implementation requires careful attention to detail -- **Token Management**: Secure token storage and refresh flows are critical -- **Webhook Security**: Signature verification prevents webhook spoofing -- **Audit Logging**: Comprehensive logging enables security monitoring and compliance +Commands still go through REST: validate token and scope, check ownership, execute, publish the resulting state event. Partners who issued the command and partners who only listen both see the same change. -### Performance Optimization -- **Caching Strategy**: Multi-layer caching significantly improves performance -- **Database Optimization**: Proper indexing and query optimization at scale -- **Real-time Events**: Event-driven architecture scales better than polling -- **Monitoring**: Proactive monitoring enables quick issue resolution +Caching is conservative. Device snapshots are safe to cache briefly; a command invalidates that key. Rate limits are per client and stricter on control routes than on reads. Exact quotas are a commercial setting, not a number I am publishing here. -## Future Enhancements +I did not bolt GraphQL onto this. The partner questions are known — list devices, act, subscribe — and a stable REST+events contract is easier to support than a query language over the internals. -### Advanced Features -- **GraphQL Support**: More flexible data querying for complex integrations -- **Real-time WebSocket API**: Direct WebSocket connections for real-time applications -- **AI-Enhanced Documentation**: Automated code examples and integration guides -- **Advanced Analytics**: Machine learning-powered usage analytics and optimization +## After -### Ecosystem Expansion -- **Marketplace Integration**: API marketplace for third-party developers -- **Certification Program**: Partner certification for quality assurance -- **Community Platform**: Developer community and support forums -- **Extended Protocol Support**: Support for Matter, Thread, and other emerging standards +External integrations sit on an authorised, documented API instead of on internal handlers. Users grant access per partner. Status moves as events. The MQTT broker and stores stay private. -This API platform has become a cornerstone of the Zimi ecosystem, enabling rapid expansion through partnerships while maintaining the highest standards of security, performance, and developer experience. The event-driven architecture and comprehensive OAuth2 implementation have proven scalable and secure, supporting the company's growth into new markets and use cases. +Related: [C4 identity and access](/projects/iot-platform-architecture), [voice](/projects/voice-control-ecosystem), [migration](/projects/iot-platform-migration). diff --git a/src/content/projects/voice-control-ecosystem.mdx b/src/content/projects/voice-control-ecosystem.mdx index 7ead571..4bd7997 100644 --- a/src/content/projects/voice-control-ecosystem.mdx +++ b/src/content/projects/voice-control-ecosystem.mdx @@ -1,6 +1,6 @@ --- title: "Zimi Smart Home Voice Assistant Integration" -description: "Architected and implemented voice control integrations for Google Home and Amazon Alexa using a unified translation-based architecture with common business logic and platform-specific adapters for consistent smart home device control." +description: "Certified Alexa and Google Home control for Zimi devices from one domain layer — platform-specific translators, OAuth linking, and report-state from live telemetry." period: "2019-2021" status: "completed" featured: true @@ -16,24 +16,17 @@ technologies: - "REST APIs" - "Event-driven Architecture" achievements: - - "Successfully certified for both Google Home and Alexa platforms" - - "Developed unified translation-based architecture for consistent behavior" - - "Implemented real-time state synchronization across voice platforms" - - "Created reusable common business logic reducing maintenance overhead" - - "Achieved seamless OAuth account linking for both platforms" - - "Maintained consistent device control experience across ecosystems" + - "Certified Alexa Smart Home and Google Home fulfilment" + - "One device/command model behind two assistant APIs" + - "Report-state from the telemetry path so assistants are not polling us" challenges: - - "Translating between different API paradigms and request/response formats" - - "Implementing unified business logic while supporting platform-specific requirements" - - "Ensuring real-time state synchronization from Zimi backend to voice platforms" - - "Managing OAuth authentication flows for both Google Home and Alexa" - - "Handling platform-specific device capability reporting and discovery" + - "Alexa and Google disagree on discovery, traits, and state reporting" + - "A wall plate is several virtual endpoints, not one gadget" + - "A voice 'on' has to hit the same MQTT command path the app uses" outcomes: - - "Expanded Zimi's reach to major voice assistant ecosystems" - - "Enhanced user experience with hands-free device control" - - "Established scalable foundation for future voice platform integrations" - - "Reduced code duplication through common business logic architecture" - - "Enabled consistent voice control experience across platforms" + - "Same household devices behave the same from app and from either assistant" + - "Capability lists live in one place, so certification is not two products" + - "Account linking is OAuth; voice never sees customer app JWTs" image: /images/projects/voicecontrol.png gallery: - "/images/projects/zimi-voice-architecture.jpg" @@ -50,339 +43,45 @@ tags: - "Real-time Synchronization" category: "iot" publishDate: 2021-06-01 -lastUpdated: 2024-01-10 +lastUpdated: 2026-08-27 --- -## Project Overview +I built and certified Zimi’s Amazon Alexa and Google Home integrations so the same in-wall devices can be asked to switch, dim, and report state from either assistant. Fulfilment runs on our GCP backend. Assistants never talk MQTT. -The Zimi Smart Home Voice Assistant Integration project implements comprehensive voice control capabilities for Zimi smart electrical devices through both Google Home and Amazon Alexa platforms. The project features a sophisticated backend architecture with unified business logic and platform-specific translation layers to ensure consistent functionality across voice assistant ecosystems while maximizing code reuse and maintainability. +Voice, OAuth linking, and report-state are views in [the C4 model](/projects/iot-platform-architecture). -## The Challenge +## Two assistant APIs, one device model -Integrating Zimi smart devices with multiple voice assistant platforms presented unique architectural challenges: +Google Sync/Query/Execute and Alexa Discover/Control/ReportState are the same three questions in different envelopes: what devices exist, what are they doing, do this. -- **Platform Diversity**: Google Home and Alexa have different API structures, request/response formats, and authentication flows -- **Code Duplication**: Risk of maintaining separate implementations for each platform -- **Real-time Synchronization**: Device state changes from Zimi backend must be instantly reflected across all voice platforms -- **OAuth Complexity**: Secure account linking between voice assistants and Zimi user accounts -- **Consistent Behavior**: Ensuring identical device control experience across different voice ecosystems +If each assistant gets its own stack, those answers drift. A dimmer grows a trait on Google and not on Alexa. Certification becomes two products. Multi-function hardware makes that worse: one physical plate is a switch, a dimmer, a fan, a garage — virtual endpoints with different capabilities. -## Architecture & Design +Auth is a second split. App users have IdP JWTs. Voice needs OAuth account linking, with the assistant holding tokens that must never be treated as an admin session. -Voice fulfilment, OAuth account linking and report-state sit in the C4 model: [Diagramming a Smart-Home IoT Platform with C4](/projects/iot-platform-architecture). +And state has to move both ways. A finger on the wall should update the Google Home graph. A spoken command should go through the same MQTT command path as the mobile app, or “Alexa off” and the app will disagree. -### Unified Translation-Based Architecture +## What changed -The core innovation was designing a translation-based architecture that converts platform-specific requests into a common internal format: +I put a **translation layer** on each assistant and a **shared domain layer** behind both. -- **Translation Layer**: Converts Google Home and Alexa requests into common internal format -- **Common Business Logic**: Unified processing engine handling device operations regardless of source platform -- **Response Translation**: Converts common internal responses back to platform-specific formats -- **Shared Components**: Device state management, user authentication, and command processing logic reused across platforms +Google and Alexa adapters map discovery, query, and execute into a common request: who is the user, which endpoints, which operation. The domain layer checks the linked account, loads devices, and runs the same command/query code the rest of the platform uses. Responses map back into each assistant’s payload shape. -### Authentication Layer +OAuth account linking is the identity bridge. The voice path does not accept the customer app JWT. Tokens and unlinking (Google Disconnect, Alexa disable) are first-class. -**OAuth 2.0 Integration**: Both Google Home and Amazon Alexa authenticate users through OAuth flows to link Zimi user accounts with their respective smart home platforms, enabling secure device access and control through centralized token management. +Unsolicited changes — someone used the wall switch — already flow through the telemetry processor. That processor **reports state** to Google Home Report State and Alexa Change Report. Assistants are not polling us for “is the light on.” -### System Components +Alexa hits a thin wrapper; fulfilment, state, and commands stay on GCP. Device-class mapping (switch, dimmer, fan, outlet, blind, garage) lives in the shared model so both certifications see the same hardware. -**Platform Translation Services** -- Google Home Translator: Handles conversion between Google Actions format and common format -- Alexa Translator: Manages conversion between Alexa Skills format and common format -- Bidirectional Processing: Supports both request translation and response formatting +## Translate at the edge, share the domain -**Common Business Logic Engine** -- Device Manager: Unified device discovery, state management, and command processing -- User Manager: Permission validation and account management across platforms -- State Engine: Centralized device state tracking and command execution +The alternative was two fulfilment services. That duplicates ownership rules, endpoint IDs, and “what does off mean on a fan.” The cheaper-looking option is the one that fails certification when Google adds a trait and Alexa does not. -**Real-time Event System** -- Event Handler: Processes device state changes from Zimi backend -- Event Translator: Converts internal events to platform-specific formats -- Push Service: Distributes real-time updates to Google Home and Alexa platforms +Keeping the assistants at the edge means a Google payload never leaks into command execution. The domain layer speaks *our* operations. Adding a third assistant, if we ever did, is another translator — not another device model. -## Implementation Details +The operational consequence is that report-state is part of telemetry, not part of the HTTP request that handled the last voice command. If we only updated assistants when voice ran, wall-switch changes would be stale until the next spoken query. Users notice that immediately; they call it “broken.” -### Google Home Integration +## After -**Supported APIs** -The Google Home integration implements the complete Actions on Google Smart Home API: +Both platforms certified. A household can link once per assistant and control the same devices the app sees. Capability mapping is shared. State follows the device, not the last voice request. -- **Sync**: Discovers and registers user's Zimi devices with Google Home -- **Query**: Retrieves current device states for status inquiries -- **Execute**: Processes voice commands to control devices -- **Disconnect**: Handles account unlinking and device removal - -**Request Translation Implementation** -```typescript -// Convert Google Home format to common format -const translateGoogleRequest = (googleRequest: GoogleSmartHomeRequest) => { - const { inputs, requestId } = googleRequest; - - for (const input of inputs) { - switch (input.intent) { - case 'action.devices.SYNC': - return createCommonSyncRequest(input, requestId); - case 'action.devices.QUERY': - return createCommonQueryRequest(input, requestId); - case 'action.devices.EXECUTE': - return createCommonExecuteRequest(input, requestId); - } - } -}; - -// Convert common response to Google format -const translateToGoogleResponse = (commonResponse: CommonResponse) => { - return { - requestId: commonResponse.requestId, - payload: formatGooglePayload(commonResponse.data) - }; -}; -``` - -### Amazon Alexa Integration - -**Supported APIs** -The Alexa integration implements the Smart Home Skill API: - -- **Discovery**: Discovers and registers user's Zimi devices with Alexa -- **Control**: Executes voice commands to control devices -- **Query/ReportState**: Retrieves current device states for status inquiries -- **Authorization**: Manages OAuth and account linking - -**Request Translation Implementation** -```typescript -// Convert Alexa format to common format -const translateAlexaRequest = (alexaDirective: AlexaDirective) => { - const { header, endpoint, payload } = alexaDirective; - - switch (header.name) { - case 'Discover': - return createCommonDiscoveryRequest(alexaDirective); - case 'TurnOn': - case 'TurnOff': - return createCommonControlRequest(alexaDirective); - case 'ReportState': - return createCommonQueryRequest(alexaDirective); - } -}; - -// Convert common response to Alexa format -const translateToAlexaResponse = (commonResponse: CommonResponse) => { - return { - event: { - header: createAlexaHeader(commonResponse), - endpoint: formatAlexaEndpoint(commonResponse), - payload: formatAlexaPayload(commonResponse.data) - } - }; -}; -``` - -### Common Business Logic Engine - -The unified processing engine handles all device operations through a standardized interface: - -```typescript -// Common request processor -const processCommonRequest = async (commonRequest: CommonRequest) => { - // Validate user permissions - await validateUserAccess(commonRequest.userId, commonRequest.deviceIds); - - // Process device operations through Zimi backend - const result = await executeDeviceOperation(commonRequest); - - // Return standardized response - return createCommonResponse(result); -}; - -// Device operation execution -const executeDeviceOperation = async (request: CommonRequest) => { - switch (request.operation) { - case 'sync': - return await syncUserDevices(request.userId); - case 'query': - return await queryDeviceStates(request.deviceIds); - case 'execute': - return await executeDeviceCommands(request.commands); - } -}; -``` - -## Real-time State Synchronization - -### Event Processing Pipeline - -The system implements comprehensive real-time state synchronization to keep all platforms updated when device states change: - -**Backend Event Handler Integration** -```typescript -// Process device state changes from Zimi backend event handler -const handleDeviceStateChange = async (deviceEvent: DeviceStateEvent) => { - // Update internal device state - await updateDeviceState(deviceEvent.deviceId, deviceEvent.newState); - - // Translate to platform-specific events - const googleEvent = translateToGoogleStateEvent(deviceEvent); - const alexaEvent = translateToAlexaStateEvent(deviceEvent); - - // Forward to respective platforms - await sendGoogleStateReport(googleEvent); - await sendAlexaChangeReport(alexaEvent); -}; - -// Event translation for Google Home -const translateToGoogleStateEvent = (deviceEvent: DeviceStateEvent) => { - return { - requestId: generateRequestId(), - agentUserId: deviceEvent.userId, - payload: { - devices: { - states: { - [deviceEvent.deviceId]: formatGoogleDeviceState(deviceEvent.newState) - } - } - } - }; -}; - -// Event translation for Alexa -const translateToAlexaStateEvent = (deviceEvent: DeviceStateEvent) => { - return { - event: { - header: { - namespace: getAlexaNamespace(deviceEvent.deviceType), - name: 'ChangeReport', - messageId: generateMessageId() - }, - endpoint: { - endpointId: deviceEvent.deviceId - }, - payload: formatAlexaChangePayload(deviceEvent.newState) - } - }; -}; -``` - -### State Distribution System - -**Push Service Implementation** -- **Google Home**: Uses Report State API to push real-time device state updates -- **Amazon Alexa**: Uses Alexa Event Gateway to send change reports -- **State Synchronization**: Ensures all platforms maintain consistent device state information - -```typescript -// Push state updates to platforms -const pushStateUpdates = async (deviceEvent: DeviceStateEvent) => { - // Send to Google Home Report State API - await axios.post('https://homegraph.googleapis.com/v1/devices:reportStateAndNotification', - googleEvent, { - headers: { 'Authorization': `Bearer ${googleAccessToken}` } - }); - - // Send to Alexa Event Gateway - await axios.post('https://api.amazonalexa.com/v3/events', - alexaEvent, { - headers: { 'Authorization': `Bearer ${alexaAccessToken}` } - }); -}; -``` - -## Technical Highlights - -### Translation Layer Benefits - -**Code Reuse and Maintainability** -- **Unified Logic**: Single implementation of core business logic behind platform-specific adapters -- **Platform Abstraction**: Translation layers isolate platform-specific code, enabling easier updates -- **Consistent Behavior**: Ensures identical functionality across both voice assistant platforms - -**Extensible Architecture** -- **New Platform Support**: Additional voice assistants can be integrated by adding new translation layers -- **Minimal Core Changes**: Common business logic remains unchanged when adding platforms -- **Isolated Updates**: Platform-specific changes don't affect the core system - -### OAuth Authentication Flow - -**Secure Account Linking** -```typescript -// OAuth flow implementation -const handleOAuthCallback = async (platform: 'google' | 'alexa', authCode: string) => { - // Exchange authorization code for tokens - const tokens = await exchangeAuthCode(platform, authCode); - - // Link user accounts - await linkUserAccounts(tokens.userId, platform, tokens); - - // Return success response - return { linked: true, platform }; -}; -``` - -### Performance Optimization - -**Response Time Optimization** -- Implemented response caching for device discovery requests -- Shared business logic so Alexa and Google Home behave consistently -- Used connection pooling for Zimi backend API calls -- Implemented efficient event batching for state updates - -## Integration Flow - -### Voice Command Processing Flow -1. User issues voice command to Google Home or Alexa -2. Platform sends API request to Zimi voice integration service -3. Translation layer converts platform-specific request to common format -4. Common business logic processes the request and interacts with Zimi backend -5. Response is generated in common format -6. Translation layer converts response back to platform-specific format -7. Device action is executed through Zimi backend and ZCC gateways -8. State change confirmation is sent back to the voice platform - -### Real-time State Synchronization Flow -1. Zimi device state changes (detected by ZCC gateway) -2. ZCC gateway reports change to Zimi backend via MQTT -3. Zimi backend event handler processes the state change -4. Event is translated to platform-specific formats (Google/Alexa) -5. Real-time updates are sent to both Google Home and Alexa platforms -6. Voice assistants update their device state cache -7. Users receive accurate status information for future voice queries - -## Results & Impact - -### Technical Achievements -- **Code reuse**: Shared business logic with platform-specific adapters for Alexa and Google Home -- **Maintenance Efficiency**: Single codebase for core functionality reduces development time -- **Real-time Sync**: Sub-second state synchronization across all voice platforms -- **Scalable Design**: Architecture supports easy addition of new voice assistant platforms - -### User Experience Benefits -- **Consistent Behavior**: Identical device control experience across Google Home and Alexa -- **Hands-free Control**: Natural voice commands for all Zimi smart devices -- **Real-time Updates**: Accurate device status reporting across all platforms -- **Seamless Setup**: Simple OAuth-based account linking process - -### Business Impact -- **Market Expansion**: Access to millions of Google Home and Alexa users -- **Platform Coverage**: Comprehensive voice assistant ecosystem support -- **Development Efficiency**: Unified architecture reduces ongoing maintenance costs -- **Future-Ready**: Scalable foundation for additional voice platform integrations - -## Lessons Learned - -### Architecture Insights -- **Translation Pattern**: Converting platform requests to common format enables maximum code reuse -- **Event-Driven Design**: Real-time state synchronization is critical for voice assistant trust -- **OAuth Complexity**: Account linking requires careful handling of multiple authentication flows - -### Platform Considerations -- **API Differences**: Google Home and Alexa have fundamentally different request/response patterns -- **State Reporting**: Each platform has unique requirements for real-time state updates -- **Certification Requirements**: Platform-specific testing and compliance processes require dedicated effort - -### Development Best Practices -- **Common Interface**: Abstracting platform differences through translation layers simplifies maintenance -- **Comprehensive Testing**: Voice interactions require extensive testing across multiple scenarios -- **Error Handling**: Graceful degradation and clear error messages are essential for voice interfaces - -This project successfully established Zimi's presence in the voice assistant ecosystem through an innovative translation-based architecture that maximizes code reuse while providing platform-specific optimizations for Google Home and Amazon Alexa integration. \ No newline at end of file +Related: [partner API](/projects/third-party-api-platform) (OAuth and events for non-voice partners), [migration](/projects/iot-platform-migration), [C4 voice and identity views](/projects/iot-platform-architecture).