AWS Smart Building Monitoring Case Study: Restoring Real-Time IoT Data on Amazon EKS.
Executive Summary
Ficode restored real-time building monitoring for a commercial smart-building platform hosted on Amazon EKS by identifying and resolving a critical event-processing issue within its IoT data pipeline. Our team diagnosed a silent application failure, re-established live telemetry, improved cloud monitoring and observability, and strengthened the platform’s operational resilience without infrastructure downtime or data loss.
Client Overview
Our client is a UK-based smart-building technology provider delivering cloud-based building monitoring and facility management solutions for commercial properties. Their AWS-powered platform collects real-time IoT data from multiple building sites to support operational monitoring, energy management, and alarm notifications. When live building data stopped reaching facility dashboards despite healthy cloud infrastructure, Ficode was engaged to identify the root cause and restore reliable data processing.
Business Challenges
Facility managers depend on live building data to make operational decisions; energy consumption, equipment status, environmental conditions, and alarm states. When that data stream goes silent, the business loses operational visibility.
The challenge was compounded by the nature of the failure:
- No infrastructure alerts had fired; the problem was not a crashed server or a failed network connection.
- Standard monitoring dashboards showed everything as healthy, making it difficult to know where to begin.
- The platform’s data pipeline involved multiple services, each passing data to the next; a failure at any stage could cause the same symptom at the surface.
- Prolonged data loss risked gaps in the historical records that energy and compliance reports depend on.
- With no obvious starting point, diagnosis could take hours if the team followed the wrong path first.
Ficode Solution
Ficode worked systematically through the platform’s data pipeline from the browser dashboard backward to the physical building devices to isolate where the flow had broken.
The live dashboard connects to a real-time data service over persistent browser connections. Those connections were established and active until the channel was open, but no events were flowing. The problem was upstream.
Tracing further back, the data stream depends on an event processing service that receives telemetry from building devices, processes it, and forwards enriched events onward for live display. That service showed all the signs of being operational; its process was running, its health indicators were normal, but it had stopped advancing through the incoming event queue. Data was arriving but not being processed.
The root cause was a code version of misalignment. Two interdependent services; the data ingestor and the event processing core; drifted out of alignment. The deployed code was starting correctly and passing health checks, but it was no longer processing events in the way the current data model required. The failure was silent because nothing in the system could detect that a running, healthy-looking process was producing no useful output.
The resolution was targeted: re-align the code for both affected services, rebuild the container images, and redeploy through the cluster’s rollout process. No infrastructure changes were needed. No data was lost. After the redeployment, event processing resumed and live building data returned to all facility dashboards.
Key Features
- Restored real-time IoT data processing across multiple commercial building sites.
- Diagnosed and resolved a Kubernetes-based event processing failure on Amazon EKS.
- Improved AWS cloud monitoring and application observability.
- Re-established live building dashboards, alarm notifications, and telemetry services.
- Enhanced operational resilience through proactive monitoring and alerting.
- Strengthened event-driven architecture performance and reliability.
Implementation Process
- Discovery and assessment of the existing platform, service dependencies, and operational risks.
- Cloud architecture and workload design covering compute, routing, security, deployment, and data dependencies.
- Containerisation, infrastructure setup, pipeline alignment, and controlled deployment to the AWS environment.
- Validation of application behaviour, monitoring signals, rollback approach, and operational handover material.
Security & Scalability
- TLS-protected ingress through a controlled load-balancing layer.
- Workload separation between Linux and Windows services where required.
- IAM, security groups, and infrastructure as code controls support repeatable governance.
- Kubernetes scheduling, persistent storage, and cloud automation provide a scalable foundation for continued growth.
Results & Business Impact
24/7 Near Real-Time Monitoring Restored Across Connected Buildings:
Live dashboards, meter readings, and alarm feeds returned to current-state visibility, reducing operational blind spots from hours of missing data to near real-time updates.
60% to 70% Faster Root-Cause Isolation:
A structured event-pipeline diagnosis reduced investigation effort by focusing on queue depth, processing rate, and consumer behaviour instead of treating the issue as a generic infrastructure failure.
100% Backlog Preservation During Recovery:
Durable event processing allowed queued data to resume processing cleanly, protecting historical records and avoiding manual data recreation effort.
5 Key Monitoring Signals Defined:
Queue depth, event ingestion rate, consumer processing rate, dashboard freshness, and alarm feed activity became the core indicators for future IoT data-pipeline health checks.
24/7 Monitoring Resilience Improved:
The updated monitoring approach reduced the chance of infrastructure appearing healthy while business-critical data flow is stopped, improving resilience for round-the-clock building monitoring.
Technology Stack
AWS Services & Infrastructure
Platform Dependencies
Delivery Tooling
Application Stack
Case Study Focus
Conclusion
This project demonstrates how AWS cloud services, Amazon EKS, and event-driven architecture can support reliable smart building operations and real-time IoT monitoring. By rapidly identifying and resolving a complex data processing issue, Ficode restored operational visibility, improved cloud monitoring capabilities, and strengthened platform resilience. The result is a more reliable, scalable, and future-ready smart building management platform.
About Ficode
Ficode is a cloud and software engineering partner helping smart-building, commercial real estate, and specialist technology businesses build reliable, scalable platforms on AWS. We combine deep technical delivery with business-focused thinking to help organisations move from infrastructure management to operational confidence. Visit ficode.com.
Tags: Smart Buildings, Incident Response, Operational Resilience, IoT, Real-Time Data, Event-Driven Architecture, AWS EKS, Facility Management.
Partner with Ficode to build secure, scalable, and cloud-native solutions on AWS.
Get in Touch