The dashboard was live until the network blinked.
After reconnect, some users saw a stale snapshot, others received events twice, and nobody knew whether the red status meant a fresh incident or an old message.
Real-time engineering is mostly about what happens between updates.
The label means little without a freshness window. A collaboration cursor may tolerate hundreds of milliseconds. An operational dashboard may tolerate seconds. A background status update may need only a minute. The architecture should match the decision a delayed event could change.
This page remains distinct from general custom software development because it addresses persistent connections, event semantics, state synchronisation, recovery, and load. It also sets a boundary: device connectivity management belongs on the IoT connectivity platform page, while a normal request-response application belongs with web development.
A live workflow with explicit service limits
- Event contract first
- 1
- Freshness, order, delivery, identity, recovery, permissions, and ownership
- Indicative delivery weeks
- 8-12
- After producers, consumers, environments, and test data are ready
- Starting investment
- $20K
- Fixed after concurrency, protocol, infrastructure, and test scope are known
Performance claims require a configuration and workload. We report observed behaviour for the tested environment, traffic model, regions, payloads, dependencies, and failure cases. Product owners still decide the acceptable delay, loss, duplicate, stale state, and degraded experience.
Use real-time engineering when delay changes a user action.
Prefer refresh, polling, or scheduled processing when a slower path is simpler and adequate.
A fit01Users collaborate, monitor, track, dispatch, trade, communicate, or respond to events inside a defined freshness window.
02The team can identify event producers, consumers, ownership, recovery rules, expected concurrency, and degraded behaviour.
03Persistent connections, event fan-out, replay, or synchronisation are material product or operational requirements.
Not a fit01The data changes rarely and a refresh button, short polling interval, or scheduled job meets the user need.
02Upstream data is unreliable and the proposed live interface would only display wrong information sooner.
03The requirement is described as zero latency, infinite scale, perfect delivery, or exactly once without a bounded business definition.
System scope
What one live workflow may include
Define event names, schemas, identifiers, producers, consumers, sequence,
timestamps, snapshots, retention, compatibility, and ownership. Separate facts
from commands and transient presence from durable business state.
Choose WebSocket, SSE, webhook, queue, stream, or polling by direction and
tolerance. Implement authentication, subscriptions, brokers, partitions,
routing, heartbeats, quotas, and backpressure around measured needs.
03Client synchronisation and experience
Handle optimistic updates, conflicts, presence, stale indicators, offline
state, reconnect, missed-event recovery, permissions, and device or browser
lifecycle. Where shared editing needs CRDT or operational transformation, it
is scoped explicitly.
04Reliability and operations
Add idempotency, acknowledgements, retry, dead-letter handling, replay,
reconciliation, dashboards, alerts, tracing, capacity tests, deploy strategy,
incident runbooks, and service limits.
Choose the update model
| Pattern | Use it when |
|---|
| Refresh or polling | Simple request-response updates | Changes are infrequent and small delays or redundant requests are acceptable. |
|---|
| Server-sent events | One-way server updates over HTTP | Browsers mainly receive a stream and do not need bidirectional messages. |
|---|
| WebSocket | Long-lived bidirectional connection | Interactive messaging, collaboration, presence, or low-delay commands justify connection state. |
|---|
| Queue or event stream | Durable backend distribution | Services need buffering, replay, independent consumption, and operational decoupling. |
|---|
When a connection drops, a spinner is not a recovery plan. The interface should tell users when data was last confirmed, whether a request is queued, which edits conflicted, and what can still be trusted. Reconnect storms after an outage need jitter and capacity protection. Slow consumers need batching, dropping, resync, or disconnection rules.
Event-driven backends also need business reconciliation. A message acknowledgement can confirm transport without proving the downstream state is correct. For money, inventory, entitlement, or other important workflows, periodic comparison with the system of record may matter more than an elegant delivery claim.
Delivery
From event contract to observable live workflow
Four phases make the timing, recovery, and operating model testable.
- Phase 1
01Define events and user tolerance
Map producers, consumers, payloads, identity, freshness, ordering,
duplicates, permissions, concurrency, disconnection, degraded behaviour, and
business ownership.
- Phase 2
02Prove protocol and capacity
Choose polling, SSE, WebSocket, webhooks, queues, or streams through a thin
path; test fan-out, reconnect, backpressure, latency, and cost assumptions.
- Phase 3
03Build workflow and recovery
Implement event contracts, state synchronisation, product UI,
acknowledgements, retries, idempotency, replay, observability, security, and
failure handling.
- Phase 4
04Load test release and hand over
Exercise realistic traffic and faults, stage rollout, verify service
objectives, document runbooks and limits, and transfer or support
operations.
Constraints
What the architecture record must settle
- Delivery semantics
- Define where at-most-once, at-least-once, deduplication, ordering, idempotency, replay, and reconciliation apply.
- Connection lifecycle
- Cover authentication refresh, heartbeats, sleep, network changes, proxies, timeouts, reconnect backoff, resume, and stale-state display.
- Capacity and dependency
- Record traffic model, burst, payload, regions, brokers, storage, downstream limits, quotas, load-test environment, and headroom.
- Operations
- Set service objectives, dashboards, alerts, deploy and rollback, incident roles, support hours, retention, privacy, security, and cost ownership.
Scope and price
A focused real-time workflow starts at $20,000.
Start with one event family, producer, consumer experience, service target, recovery path, and operating owner.
The proposal separates engineering from cloud, broker, network, observability, vendor, on-call, and support costs.
Starting investment
Starts at $20,000
A focused release commonly takes 8 to 12 weeks. Collaboration, durable replay, several regions, large connection counts, device fleets, or regulated data add scope.
Tested assumptions
Capacity results name the environment and workload rather than becoming an
unqualified scale promise.
Visible degraded state
Reconnect, stale data, retries, and recovery are designed as product
behaviour, not hidden implementation details.
Choose the wider platform path