How PubNub works
This page describes how the PubNub network operates connection management, scalability, message ordering, client recovery, delivery guarantees, fault tolerance, and more.
Knowing where the boundary falls between what the network guarantees and what your application still has to handle is what lets you design retries, timeouts, and recovery logic deliberately. Otherwise, you discover the gaps once you're already in production.
Where your clients connect
PubNub operates 11 message-routing Points of Presence (PoPs) at the region and data-center level, across North America, Europe, Asia Pacific, and Southern Asia, and each one is a full deployment of the PubNub services at the edge of the network, not a CDN edge node or a caching layer in front of a single origin, so a publish and a subscribe are both served locally.
Clients are routed to their geographically nearest PoP automatically. DNS resolution maps the client's location to the closest healthy PoP, and PubNub SDKs open a connection there without any configuration on your part. There is no region to select in your keyset and no endpoint to choose per user. That means a single build of your application behaves the same in every market you ship to.
Once one PoP accepts a message, it replicates the message to the others so that subscribers anywhere receive it. It's also why a subscriber reconnecting somewhere else finds its data already present. In sequence, your app publishes and DNS geo-routing sends the request to the nearest PoP. That PoP replicates the message to the other PoPs, and each PoP delivers it to its own subscribers.
Clients reach PubNub over ordinary HTTPS. A subscribe request is held open for up to 310 seconds before the client reissues it with an updated timetoken cursor, and PubNub uses pooled, multiplexed HTTP/2 connections on supported origins. PubNub doesn't use WebSocket.
That choice is deliberate. The subscribe loop travels the same path as any other HTTPS request. So it works through corporate proxies, CDNs, and restrictive firewalls without the connection-upgrade failures and fallback paths that WebSocket deployments have to handle. For the transport details, refer to Data transport and delivery.
How fast messages arrive
- Publish-processing latency, the time PubNub takes to accept and acknowledge a publish request, is about 0.5 ms within the same region, so your client isn't blocked waiting on the network before it can continue. This is not end-to-end delivery latency.
- Inside one region, PubNub adds less than 500 µs to a message between ingestion and the moment it's ready for delivery to subscribers, not counting network transit, as measured by PubNub's same-region latency monitoring.
- Measured end-to-end delivery latency, from publisher to subscriber wire to wire and including network transit, averages approximately 57.79 ms, or approximately 64.97 ms across zones. Because clients connect to their nearest PoP, most of that latency budget is spent on genuine distance rather than detours through a single origin region.
- PubNub compresses and streams messages to subscribers, and delivers several messages ready for the same subscriber as one bundle rather than one round trip each. Long-polling over pooled HTTP/2 connections keeps connection count and bandwidth cost low.
Growing without capacity planning
Channels are created implicitly the first time they are used and do not require provisioning, so PubNub places no limit on the number of channels a keyset or a client uses and there is no migration to run when your channel pattern changes.
Fan-out isn't something you size either. The number of subscribers on a PubNub channel is unlimited. So a channel that carries a two-person conversation and a channel that carries a live event broadcast are the same object with the same API.
Capacity follows traffic. Message API systems autoscale per Availability Zone based on the load they're actually seeing. So a launch spike, a viral moment, or a scheduled event doesn't require you to warn PubNub or pre-warm anything.
One client instance carries many subscriptions. It multiplexes subscriptions to any number of channels over a single subscribe connection. Each PubNub client instance uses two TCP sockets: one for subscribe requests and one for all non-subscribe operations. Because that count doesn't grow with your channels, per-device connection and battery cost stay flat as your application's channel count grows.
The order messages arrive in
PubNub assigns every message a server-side timetoken when it accepts the publish. The timetoken records when PubNub accepted the message, which can differ from when the client sent it. History fetched through Message Persistence returns a channel's messages in timetoken order.
A subscriber receives live messages in the order they reach it, and each message carries its timetoken. Arrival order can differ from timetoken order and from one subscriber to another. Sorting a channel's messages by timetoken gives every client the same order. Timetoken order applies within a channel, so each channel's messages sort on their own.
For the reconnect and buffer mechanics that recovery relies on, refer to Connection management.
When a client goes offline
Mobile clients lose connectivity constantly, so PubNub treats connection recovery as the normal case rather than an error path. When an offline client reconnects, its subscribe cursor resumes from the last timetoken it received. For a short gap, the message buffer replays the missed messages. For a longer gap, your application fetches the missed messages from Message Persistence, within what history stores and retains.
Reconnection is automatic. SDKs detect a dropped connection and reestablish it. Where an SDK preserves the subscribe cursor across a reconnect, it resumes from the last timetoken it actually received rather than from a fresh server-issued one. That preserved cursor doesn't expire or get silently discarded while the client is reconnecting. Within the buffer's window, the client asks for exactly the messages it's missing with no bookkeeping in your application code.
Short gaps are covered automatically, and replayed when the client returns.
The subscriber message buffer queues messages for a reconnecting client. It holds 100 messages for up to 16 minutes by default, and discards the oldest first (FIFO) when a burst exceeds that size. Larger buffers, for example 300 or 500 messages, can be provisioned per keyset by PubNub Support.
Longer gaps are covered by stored history. Your application passes the same timetoken to Message Persistence to fetch the stored messages published after the last message it actually received. Message Persistence doesn't recover signals, messages published with storage turned off, or messages whose per-message time to live has passed.
Message Persistence retention is 1 or 7 days on the Free plan, 30 days, 3 months, or 6 months on Starter, and 1 year or Unlimited on Pro, which bounds how far back you can recover and makes recovery from a multi-day client outage a supported operation rather than a special case.
The two mechanisms are independent, and one doesn't fall back to the other, because they answer different absences. The buffer covers a dropped tunnel or a Wi-Fi handoff, and Message Persistence covers a backgrounded app, a flight, or a longer outage. For the full model, including request parameters, refer to Connection management.
What PubNub guarantees about delivery
PubNub states its delivery semantics precisely, so you can decide where you need stronger guarantees instead of discovering it in production.
| Scenario | What PubNub provides | What you add if needed |
|---|---|---|
| Live stream | At-most-once on a stable connection. A burst exceeding the subscriber buffer discards oldest messages first, and a replay after a reconnect can repeat a message | Reconnect with preserved timetoken cursor to replay what was missed (automatic in SDKs), and drop timetokens you've already processed |
| Publish | No server-side retry: the SDK reports success or failure | Retry with backoff; a retry after an ambiguous failure may produce a duplicate |
| Exactly-once processing | Not built in | Idempotency key per message, deduplicated per receiver or centrally in a Function |
What happens when something breaks
Fault tolerance is designed in, so that individual failures are absorbed rather than escalated and the common cases need no human intervention.
Failures are contained by design. Each availability zone is subdivided into independent isolation zones, separate hardware and routing groups that sustain themselves, so a problem inside one doesn't propagate to the others. Health is measured per isolation zone, not just per region.
Recovery is automatic at every layer. Unhealthy endpoints are removed from DNS rotation so new connections avoid them, and replacement capacity is brought up and takes over. If an entire PoP becomes unhealthy, clients are routed to the next-closest PoP on their next request. PubNub SDKs also bypass stale DNS caches, so clients behind networks that ignore short TTLs still move. Because messages are replicated across PoPs, a client that lands somewhere new finds the messages it missed waiting for it and receives them on reconnect.
Message data is protected independently of the routing layer. Recent messages are held in memory on multiple servers within each PoP and synchronized between PoPs, and stored messages are replicated across regions. PubNub assigns every published message a server-side timetoken: a monotonically increasing 17-digit value precise to 100 nanoseconds (10⁻⁷ s), the number of 100-nanosecond intervals since the Unix epoch. That precision is what makes it possible to work out exactly what a reconnecting client still needs.
PubNub commits to a 99.999% uptime SLA on Enterprise plans, which allows about 5 minutes of downtime per year. That figure is backed by per-isolation-zone failure containment, automatic endpoint and capacity recovery, and cross-PoP message replication.
How PubNub runs the network
PubNub owns the infrastructure and its operations so your team doesn't have to. Every service runs across regions, and message API capacity autoscales with traffic rather than being sized ahead of a launch.
Deployments are rolling. New versions are shifted into the network PoP by PoP and zone by zone while healthy capacity continues to serve traffic, so upgrades aren't customer-visible events. Maintenance activity is scheduled per region during that region's lowest-traffic period. In the rare case where an activity requires scheduled downtime, customers are notified in advance.
Operations run continuously, with a globally distributed team and proactive monitoring designed to catch conditions before they reach applications. Service health is published for every region at PubNub Status, including interruptions that no customer reported, and incidents are followed by root-cause analysis that feeds back into the platform.
Operational visibility isn't limited to the status page. PubNub can stream your keyset's service metrics into the monitoring stack you already run, so your real-time traffic appears on the same dashboards as the rest of your system.
Next steps
- What is PubNub? - see how these network properties support PubNub's capabilities.
- Core concepts - channels, messages, User IDs, timetokens, tokens, and memberships.
- Data transport and delivery - the transport layer in detail: long polling, HTTP/2, connection types.
- Connection management - the full reconnection, status listener, and presence timeout model.