Connection management
Learn which parts of a PubNub connection the SDK manages for you, which failures it recovers from on its own, and which gaps your application has to close itself. Knowing where that boundary falls before you ship matters. It separates an app that rides out a tunnel, a Wi-Fi handover, or an expired token from one where a silently disconnected client is indistinguishable from a quiet channel.
One constraint shapes everything below: your code never opens a socket, schedules a keep-alive, or decides when to retry a failed request. The SDK owns the transport. Your application's part is the configuration you choose up front and the reaction to what the SDK reports afterwards.
| Part of the model | What it does |
|---|---|
| Client instance | Holds the sockets, timers, and cursor, and connects lazily on the first API call |
| Subscribe connection | Carries messages and events to the client over a request the network holds open |
| Non-subscribe connection | Serves every other operation, such as publish or history, as a short request and response |
| Status listener | Reports connection state changes as status events rather than as data |
| Reconnection policy | Decides how long the SDK waits between retries, and how many it makes |
| Timetoken cursor | Marks the point in the stream a reconnecting client resumes from |
| Subscriber message buffer | Holds recent messages so a briefly disconnected client can catch up |
| Heartbeat | Tells PubNub the client is still alive, which is what keeps it present to others |
A connection exists only after the first API call
Constructing a PubNub client doesn't reach the network. No TCP socket opens and no TLS handshake happens until the first API call, such as a publish or a subscribe. Until then, the client is a configured object that knows your keys and your User ID and nothing more.
Reuse one client instance
Because the connection and everything attached to it belong to that instance, a PubNub client is meant to be long-lived. Create it once per user or session, and reuse it for every operation. Reuse preserves three things that a fresh instance throws away.
- Connection reuse. Where an SDK supports HTTP keep-alive, a live client keeps a pooled connection warm, so later requests skip TCP and TLS setup. Several SDKs treat keep-alive as opt-in, which is why the behavior belongs in each SDK's configuration reference rather than in a single platform-wide rule.
- Background resources. Each client holds a subscribe loop, reconnection timers, heartbeat timers, and in some languages a thread pool. Most SDKs release those only when you tear the client down explicitly, so abandoned clients keep their threads, timers, and sockets alive until then.
- Continuity. The timetoken cursor and the publish sequence live on the instance. As a result, a replacement client resumes from whatever the network hands it next, rather than from the last event the previous client saw.
One client per User ID on a server
A server process that acts for many people needs a distinct User ID per person. That's an argument for pooling one client per User ID rather than building a new client per incoming request. Pooling keeps the benefits of reuse without collapsing several people into one identity. It also puts a bound on how many clients, and therefore how many sockets and timers, the process holds at once.
Two kinds of connection
A PubNub client opens two kinds of connection, and they differ in how long they stay open, which timeout governs them, and what a failure looks like.
- Subscribe connection, which carries the real-time stream
- Non-subscribe connection, which serves everything else
Subscribe connection
A subscribe connection carries the real-time stream. The SDK sends an HTTP request naming the channels the client wants, and the PubNub network holds that request open instead of answering immediately. A published message arrives as the response to the waiting request, and the SDK immediately issues the next request starting from the timetoken it just received. A subscribe request is held open for up to 310 seconds before the client reissues it with an updated timetoken cursor.
An idle channel therefore produces the same visible pattern as a busy one: a request that waits, returns, and is replaced. This is HTTP long polling. Because it's built from ordinary HTTP request and response pairs, it survives proxies, firewalls, and older client environments that don't support other real-time transports. It also means the subscribe loop, not your code, is what keeps the stream alive. For how subscriptions themselves are created and scoped, refer to Subscribe.
Non-subscribe connection
A non-subscribe connection serves everything else: publish, presence queries, message history, App Context calls, and file operations. Each is a short request answered in the normal way, so it either returns a result or fails.
Non-subscribe requests use their own timeout, separate from the subscribe hold. Both its default value and its configuration name differ across SDKs. So there's no single platform-wide figure to rely on, only the value in the configuration reference for the SDK in use. When that timeout elapses, the SDK abandons the request and reports a timeout, which some SDKs surface as a status category such as PNTimeoutCategory.
Sockets per client and platform limits
The split between the two connection types is also what the operating system sees. Each PubNub client instance uses two TCP sockets: one for subscribe requests and one for all non-subscribe operations. PubNub places no limit on how many client instances you create, though the device or browser may cap concurrent TCP connections, and the browser is the strictest case. Most browsers allow 6 concurrent connections per host name, and between 10 and 60 in total across all host names. That budget belongs to the page, not to PubNub, so every image, API call, and analytics beacon draws on the same pool. Two sockets per client is cheap, and a page that builds a client per component is not. That's the practical reason a single long-lived client matters more in a browser than anywhere else.
The status listener
The status listener is the handler through which an SDK reports the state of its connection, as opposed to the handlers that deliver messages, presence events, and other data. You attach the status listener to the PubNub client itself rather than to an individual subscription or subscription set. That is because the connection it describes is shared by everything that client does.
It exists because subscribe failures are otherwise invisible. A held-open request that returns nothing looks exactly like a quiet channel. So an application that ignores status events can't distinguish "nobody is publishing" from "this client stopped receiving twenty minutes ago". Each status event is the SDK translating a transport outcome, such as a successful handshake, a rejected token, or an exhausted retry sequence. That outcome becomes one of a small number of categories your code can branch on.
Subscribe lifecycle statuses
These are the categories that describe the subscribe lifecycle.
| Status | What it tells you |
|---|---|
Connected | The subscription started and real-time events are arriving. |
Subscription changed | The mix of subscribed channels and channel groups changed after the initial connection. Carries the current lists. |
Disconnected | The workflow stopped after having been connected, at your application's request. |
Disconnected unexpectedly | The client was receiving events, then lost the connection and exhausted its retries. Carries an error describing why. |
Connection error | The connection was never established and the retries are exhausted. Carries an error describing why. |
The subscribe lifecycle begins when your application starts a subscription. If the client connects, it reports Connected. If it never connects and its retries run out, it reports Connection error. From Connected, a change to the subscribed channels or channel groups reports Subscription changed and the client returns to Connected. When your application stops the workflow, the client reports Disconnected. When the client loses the connection and exhausts its retries, it reports Disconnected unexpectedly.
Two failures that mean different things
Two properties of the subscribe lifecycle statuses carry most of their meaning. Disconnected arrives only after a connection existed, so its absence during a failure isn't a missing event.
The two failure categories then differ by history rather than by cause. Connection error means the client never got in, commonly because a token was rejected or the network was unreachable, while Disconnected unexpectedly means it was working and then stopped. An application that treats both the same way will retry the wrong thing, because a rejected token isn't a network problem and no amount of waiting fixes it.
Status names differ across SDKs
The literal category names belong to each SDK rather than to the platform. The JavaScript SDK, for example, emits PNConnectedCategory, PNSubscriptionChangedCategory, PNDisconnectedCategory, PNDisconnectedUnexpectedlyCategory, and PNConnectionErrorCategory, and adds PNNetworkDownCategory and PNNetworkUpCategory in the browser, where the runtime itself can report that connectivity was lost and restored. Other environments have no equivalent signal, and iOS-based SDKs expose a wider set of categories than most.
The set is also deliberately narrow. In the JavaScript SDK's current subscription workflow, intermediate states such as connecting and reconnecting are handled internally and never surfaced. So a retry sequence produces no status events at all until it either succeeds or gives up. Silence between a failure and a terminal status is the SDK working, not the SDK stuck.
For the categories your SDK emits and the fields on each event, refer to its status events reference, for example JavaScript.
To attach the listener and read status events in code, refer to Monitor and respond to connection status changes.
Reconnection policies
Connection recovery in PubNub SDKs is configured rather than programmed. You choose a retry policy when you initialize the client, and the SDK applies it to failed requests on your behalf. There are two shapes: linear, which waits a constant delay between attempts, and exponential, which doubles the delay after each attempt up to a ceiling.
Defaults and jitter
By default, PubNub SDKs retry only subscribe operations, using exponential backoff for up to 6 attempts with delays growing from 2 to 150 seconds. Growing delays are why an unconfigured client keeps trying for minutes rather than seconds before it reports a terminal status.
PubNub SDKs add random jitter of 0.001 to 0.999 seconds to each reconnection retry delay. Jitter matters more than its size suggests, because a regional network event disconnects many clients at the same instant, and without it they would all retry at the same instant too, turning recovery into a synchronized burst.
What a retry policy doesn't cover
Retry coverage isn't uniform. A policy applies to the endpoints it isn't told to skip, and SDKs let you exclude specific endpoints, which turns retrying off for those operations entirely. Certain SDKs also don't retry subscribe operations by default at all, so what you get without configuring anything depends on the SDK you chose.
Some SDKs, including JavaScript and Python, don't enforce upper bounds on the maximum attempt count or delay you set, which makes an over-eager policy your responsibility. Aggressive retries on a mobile client cost battery and data with no compensating benefit, because a network that's down doesn't come back faster for being asked more often.
SDK reconnection retries don't reach the PubNub network unless one succeeds, so PubNub doesn't count them as transactions. So a retry storm shows up as battery and bandwidth cost on the device rather than on your bill.
When the retries run out
The end of a retry sequence is the handover point. When an SDK exhausts its retry attempts it stops on its own and emits Connection error or Disconnected unexpectedly. From that moment nothing further happens until your application reacts.
What survives a reconnect
A reconnect is only useful if the client can say where it left off, and whether the messages it missed are still available to send. Those are two separate mechanisms with two different lifespans.
The timetoken cursor
The client's place in the stream is the job of the timetoken. PubNub assigns every published message a server-side timetoken: a monotonically increasing 17-digit value precise to 100 nanoseconds (10⁻⁷ s), the number of 100-nanosecond intervals since the Unix epoch. The SDK keeps the timetoken of the last response it received and presents it as the cursor on the next subscribe request. That way, the network knows which point in the stream to resume from.
The cursor itself doesn't expire while a client is reconnecting.
The subscriber message buffer
What the cursor points at can expire, and that's what decides how a gap gets recovered.
The subscriber message buffer queues messages for a reconnecting client. It holds 100 messages for up to 16 minutes by default, and discards the oldest first (FIFO) when a burst exceeds that size. Larger buffers, for example 300 or 500 messages, can be provisioned per keyset by PubNub Support.
Within the buffer's size and window, reconnecting with a preserved cursor replays what the client missed and the gap closes itself. Past either boundary, presenting a valid cursor doesn't bring back messages the buffer has already discarded, and the only remaining copy is the one Message Persistence stored.
Buffer replay compared with history replay
Buffer replay and history replay are independent mechanisms rather than a fallback chain. Buffer replay is automatic and bounded, history replay is deliberate and bounded only by your keyset's retention. Nothing in the SDK escalates from one to the other. The one that applies to a given outage is a judgment your application makes, based on how long it was disconnected and how fast its channels publish.
When a client reconnects, the SDK presents the timetoken cursor. If the gap fits within the subscriber message buffer's size and time window, PubNub replays the missed messages automatically. If it doesn't, the outcome depends on Message Persistence. With Message Persistence enabled, your application calls fetchMessages to replay the gap from history. Without it, the missed messages are gone.
Message Persistence also recovers less than everything that was published. Message Persistence doesn't recover signals, messages published with storage turned off, or messages whose per-message time to live has passed. A design that treats history as a complete record of the stream will silently lose exactly those messages.
What other clients see when you disconnect
Connection management has a second audience. Your own client learns about a drop from its status listener, but everyone watching the same channel through Presence learns about it from a heartbeat that stopped arriving. Those two things don't happen at the same moment.
Implicit heartbeats and the presence timeout
Presence reporting applies once Presence is enabled on the keyset in the Admin Portal. It rests on a mechanism you get without configuring anything: every subscribe request doubles as an implicit heartbeat, telling PubNub the client is alive and resetting a server-side timer. PubNub SDKs set presenceTimeout to 300 seconds by default, which is how long PubNub waits without a heartbeat before marking a client offline.
Because the timer, not the socket, is what defines presence, a client that vanishes stays visible as online until the timeout expires. That tolerance is useful, since it stops a brief network blip from broadcasting a leave and a join to every other subscriber. But it also means presence can lag reality by as long as the timeout you configured.
Explicit heartbeats and what they cost
Explicit presence heartbeats are off by default in PubNub SDKs, because heartbeatInterval defaults to 0. Turning them on adds a dedicated loop that pings on its own schedule, which detects a lost client sooner because the timer is reset more often. The cost is real: each heartbeat is a billable API call made by every connected client, so the interval multiplies across a whole user base.
Shortening detection is worth paying for in a trading or dispatch application and usually not in a chat application, where an active user's own subscribe requests already prove they're there. Some SDKs also derive a heartbeat interval from the presence timeout, so a single configuration change can turn into ongoing traffic that was never requested explicitly.
Leaving deliberately
Unsubscribing before going offline, rather than letting the timer expire, is what produces an immediate leave event instead of a delayed timeout. That choice belongs to your application, and it's the same decision a mobile app faces when it moves to the background.
Where automatic recovery ends
Some disconnections are outside the scope of any retry policy, because retrying is the wrong response to them.
An expired Access Manager token
An expired Access Manager token is the clearest case. PubNub rejects requests that present it, and the rejection is deterministic, so retrying reproduces it. Recovery means obtaining a new token and setting it on the client, after which reconnecting is explicit rather than automatic.
This is a genuine division of labour: the retry policy handles transport failures, and your token refresh logic handles authorization failures. An application with no refresh path stays disconnected for as long as the token stays expired, however healthy the network is.
An intentional disconnect
When your application stops the workflow itself, the SDK doesn't treat that as a fault to recover from. Resuming is a separate call rather than a side effect of subscribing again.
A suspended mobile app
When the operating system suspends a backgrounded app, its subscribe loop stops running, which produces a gap the SDK didn't cause and can't observe. Whether the app unsubscribes on the way out is an application decision, not a connection setting. So is whether it replays the gap from history or resumes with the current state on the way back in. For patterns that address them, refer to How to receive messages effectively.
For more information, refer to Core concepts and Data transport and delivery.