If you need to push live data, such as prices, trades, metrics or events, to many browsers at once, a WebSocket layer backed by Redis is one of the most effective architectures available. This guide covers the system design, data flow, scaling approach and the practices that keep a live data platform fast and stable under load. It is based on the patterns we use in production FinTech platforms.
Architecture overview
The system has four parts:
- An ingestion layer that reads from upstream sources (a market data feed, an event bus, your application) and appends normalized events to Redis Streams.
- Redis as the real-time backbone: streams for ordered, replayable event logs and Pub/Sub for lightweight notifications.
- A fan-out layer of WebSocket servers that read from Redis and push events to connected clients.
- Clients (dashboards, trading screens or mobile apps) that subscribe to the channels they care about.
Keeping ingestion and fan-out separate is the most important decision. Ingestion is optimized for throughput and durability; fan-out is optimized for many concurrent subscribers. When they share a process, a burst on one side stalls the other.
Why Redis Stack?
Redis Stack bundled Redis with modules such as RedisJSON, RediSearch, RedisTimeSeries and RedisBloom; since Redis 8, those capabilities ship as part of Redis Open Source. For real-time streaming, the features that matter most are:
- Redis Streams: persistent, ordered event logs with IDs you can resume from, so a reconnecting server or client can catch up instead of missing data.
- Pub/Sub: instant, fire-and-forget fan-out to every subscriber. Ideal for notifications, but messages are lost if nobody is listening, so do not use it as your only record.
- In-memory storage: sub-millisecond reads and writes for the latest state, such as the current snapshot a client needs when it first connects.
- Search and JSON: indexes over live data, useful for screeners and filters, without adding another database.
WebSocket server design
Each WebSocket server keeps a set of connected clients and a background task that reads new events from Redis and pushes them out. A minimal version in Python with redis-py and websockets looks like this:
# Illustrative only: read a Redis Stream and broadcast new entries to WebSocket clients.
import asyncio, json
import redis.asyncio as redis
from websockets.asyncio.server import serve, broadcast
r = redis.Redis(decode_responses=True)
clients = set()
async def handler(ws):
clients.add(ws)
try:
await ws.wait_closed()
finally:
clients.discard(ws)
async def pump(stream: str = "events"):
last_id = "$" # only new entries; persist this to resume after a restart
while True:
resp = await r.xread({stream: last_id}, block=5000, count=500)
for _name, entries in resp or []:
for entry_id, fields in entries:
last_id = entry_id
broadcast(clients, json.dumps(fields))
async def main():
async with serve(handler, "0.0.0.0", 8765):
await pump()
asyncio.run(main())
In production you add authentication on connect, per-channel subscriptions so each client receives only the symbols or topics it asked for, heartbeats to detect dead connections and a snapshot on connect so the client starts with current state.
Handling scale
- Every fan-out server reads every event it needs. Each WebSocket server reads the stream directly with
XREAD, so adding servers adds capacity without coordination. Consumer groups are for splitting work between processing workers, not for fan-out. - Load-balance connections, not messages. Put the WebSocket servers behind a load balancer that supports long-lived connections. Because every server sees the same events, clients can connect to any of them.
- Scale Redis deliberately. Separate hot streams onto their own connection pools, and move to Redis Cluster when one node’s memory or CPU becomes the limit, keeping each stream on a single shard.
- Measure everything. Track connection counts, messages per second, end-to-end delay from ingestion to client, Redis memory and per-client buffer sizes, for example with Prometheus and Grafana.
Performance tips
- Keep payloads small. Serialize only what the client needs, and send deltas instead of full objects where possible.
- Batch on the way in. Pipeline writes to Redis in batches rather than one command per event.
- Bound memory explicitly. Stream entries do not have individual TTLs. Trim with
MAXLEN ~orMINIDon write, and set expiries on cache-style keys. - Handle slow clients. A client on a weak connection should never slow down everyone else. Cap each connection’s send buffer and drop or disconnect clients that fall too far behind.
- Pool connections. Reuse Redis connections across tasks instead of opening one per client.
Seeing it in production
These are the same patterns behind the real-time options flow and dark pool pipeline we built for an options analytics platform, where live data reaches analytics pages over WebSockets during the heaviest market-open bursts. For a deeper look at the ingestion side, read how we ingest the OPRA options firehose with Redis Streams.
Final thoughts
Redis plus WebSockets is one of the most powerful combinations for building real-time platforms. Whether you are building a trading dashboard, a live analytics tool or a collaborative application, separating ingestion from fan-out, bounding every buffer and measuring end-to-end delay will give your users the speed and reliability they expect. If you are planning a platform like this, see our FinTech and real-time market data development services.