Skip to content
Sava Stosic, home

Real-Time Industrial Data Platform

Industrial clients needed high-volume operational data ingested, processed, and made visible in near real time. I ran discovery and workshops, designed the architecture, and delivered systems processing 50,000+ events per second.

RoleRequirements discovery, architecture, implementation, delivery planning

Site Data flows through Azure IoT Hub, Azure Functions, and Azure Event Hub. Event Hub feeds both Fabric and a realtime use case.
High-level technical architecture

Outcome

Peak events processed per second

50K+

Peak events processed per second

Sustained ingestion and processing throughput

Contract value of client environments supported

$3M+

Contract value of client environments supported

The business context these systems operated within

Reflects the value of client environments supported, not personal sales attribution.

Context and problem

Industrial operations generate continuous, high-volume telemetry from equipment and process sensors. That data is only useful if it can be ingested reliably, processed at the rate it arrives, and put in front of the people making operational decisions while the decision still matters.

Clients needed reliable ingestion, processing, and visualization of high-volume operational data in an environment they could trust, operate, and afford.

What made it difficult

  • Sustained high event throughput with peaks well above steady-state load
  • Industrial data sources that cannot be paused or reshaped to suit the pipeline
  • Near-real-time expectations for operational visibility, not overnight batch reporting
  • Enterprise security and access requirements on every layer of the environment
  • Cost that has to stay defensible against the operational value delivered

Role and responsibilities

  • Led requirements gathering and technical workshops with client stakeholders
  • Designed the cloud and data architecture
  • Implemented the ingestion, processing, and analytics layers
  • Owned delivery planning and coordination across client and internal teams
  • Translated technical tradeoffs for non-engineering decision-makers

Generalized architecture

Industrial data sources feed an Azure ingestion layer, which passes events to stream processing for validation and enrichment. Processed data lands in operational storage and analytics, which serves near-real-time visualization for operators and decision-makers.
  1. 01Source

    Industrial data sources

    Equipment and process telemetry emitted continuously from the plant environment.

  2. 02Ingest

    Azure ingestion

    Managed ingestion sized for sustained high throughput with headroom for peaks.

  3. 03Process

    Stream processing

    Validation, shaping, and enrichment applied as events arrive rather than after the fact.

  4. 04Store

    Operational storage and analytics

    Time-series and relational stores tuned for the query patterns the client actually runs.

  5. 05Serve

    Near-real-time visualization

    Operational dashboards that put current process state in front of the people acting on it.

Flow: Industrial data sources feed an Azure ingestion layer, which passes events to stream processing for validation and enrichment. Processed data lands in operational storage and analytics, which serves near-real-time visualization for operators and decision-makers.

Key decisions and tradeoffs

  • Managed ingestion over self-managed brokers

    At tens of thousands of events per second, dropped data during a peak was a greater concern than a few milliseconds of latency. Managed Azure ingestion services absorbed burst load without requiring a team on call to scale a broker cluster. The tradeoff was less control over the transport internals.

  • Processing on the way in, not on the way out

    Validating and shaping events at ingest keeps bad data out of the analytics layer, where it is far more expensive to detect and undo. It costs more compute per event, and it means schema changes have to be handled carefully at the boundary.

  • Choosing the store by query pattern, not by default

    Time-series-optimized storage answers operational questions over recent windows far more cheaply than a general-purpose relational store; relational storage still won where the questions were transactional. Splitting them meant accepting more than one store to operate.

  • Generalized, reusable environment patterns

    Client environments differ, but the shape of the problem repeats. Standardizing the environment pattern made delivery predictable across clients. It required upfront design work that a single deployment would not have justified.

Implementation

  • Ran technical workshops to establish data sources, volumes, latency expectations, and the decisions the data needed to support
  • Built the ingestion and processing pipeline and validated it against realistic sustained and peak loads
  • Modelled operational storage around the query patterns identified in discovery
  • Delivered near-real-time dashboards and iterated on them with the operators using them
  • Planned rollout and handover so client teams could run the environment day to day

Results

  • Systems processing 50,000+ events per second in production
  • Operational data made visible to client teams in near real time rather than after the fact
  • Environments supported representing more than $3M in contract value

Lessons and next iteration

  • The workshop is the highest-leverage hour of the project. Volume and latency numbers agreed in a room beat the same numbers discovered during load testing.
  • Throughput headroom is cheaper than an incident. Designing for peak rather than average avoided the class of failure that is hardest to explain to a client.
  • Next iteration: push more validation and schema contracts to the source side, so bad data is caught before it costs pipeline compute.

Technologies

Azure IoT Hub · Azure Event Hubs · Microsoft Fabric · Azure Data Explorer · Time Series Insights · Azure Functions · Azure SQL · Power BI · Python · SQL · KQL

Relevant technologies across this body of work. Not every technology listed was part of a single identical deployment.

Have a system like this that needs designing, delivering, or rescuing?

Email Sava