---
source_url: https://www.pubnub.com/docs/design-patterns/rate-limiting
title: Rate limiting
updated_at: 2026-09-30T07:20:08.000Z
---

# Rate limiting

## Documentation index

To discover more PubNub resources:

1. Fetch [PubNub's llms.txt](https://www.pubnub.com/llms-full.txt) for a list of available pages in Markdown format.
2. Identify relevant URLs from that index.
3. Fetch the target pages.

Do not assume a path exists, always check the index first.

A channel built for a handful of participants may require some tweaks when thousands join it. [Send messages effectively](https://www.pubnub.com/docs/design-patterns/send-messages-effectively.md#keep-a-very-busy-channel-usable) names the two options for a channel under that kind of load: shard it, or throttle it. This page works through why a very large audience needs one of those strategies at all, and how to choose between them.

## Two levers drive the problem: message frequency and occupancy

A channel gets harder to use as more messages arrive per second, and separately, as more subscribers receive each one. Both matter, and they don't always move together.

Message frequency is about the reader. In a conversation, each additional message competes for the same attention, and past some rate the channel stops functioning as a dialogue at all. An engagement-focused channel, like a shared reaction stream during a live event, tolerates far more volume, because no one there expects to hold a conversation.

Occupancy is about the system. Every message published to a channel gets delivered to every subscriber on it, so a channel's total delivery volume is message frequency multiplied by occupancy. Ten thousand participants receiving the same message stream costs PubNub, and your infrastructure downstream of it, a great deal more than a hundred participants receiving the identical stream. It also raises the read side. A `Here Now` call or a stream of [presence events](https://www.pubnub.com/docs/presence/presence-events.md) behaves differently on a very large channel than on a small one, covered separately in [Occupancy](https://www.pubnub.com/docs/presence/occupancy.md).

The right fix depends on which lever is the actual problem. If the issue is that a channel has stopped being a usable conversation, splitting the audience helps. If the issue is throughput and cost at an occupancy your application still wants to treat as one audience, throttling helps instead.

## Throttle inside a shared channel to keep one audience feeling like one audience

Splitting a large channel into rooms changes the experience: participants no longer see the same stream. Sometimes that shared-audience feeling is the point, such as a live event where everyone should feel part of the same crowd. Throttling keeps everyone on one channel in that case, while controlling how much of the traffic actually gets delivered.

A [Before Publish Function](https://www.pubnub.com/docs/message-processing/serverless/overview.md) makes the throttling decision synchronously, before a message reaches subscribers. A second Function, subscribed to the same channel, measures the actual delivery rate and adjusts the throttle's configuration in response. The two form a feedback loop instead of a fixed cutoff.

The loop works in these steps:

1. A client publishes to the shared channel, and the Before Publish Function decides whether to allow or drop the message based on the current rate configuration.
2. If the Function allows the message, PubNub delivers it to subscribers.
3. If the Function drops the message, it publishes a copy to `analytics-channel` for later analysis instead of delivering it to subscribers.
4. The message sampler subscribes to the channel, counts messages per period, and adjusts the throttle rate that the Function reads on the next publish.

```mermaid
flowchart TB
    PUB["<b>Client publishes</b><br/>to the shared channel"]
    THROTTLE["<b>Before Publish Function</b><br/>throttler: allow or drop<br/>based on current rate config"]
    SUB["<b>Subscribers</b><br/>receive the message"]
    ANALYTICS["<b>analytics-channel</b><br/>receives a copy of<br/>each dropped message"]
    SAMPLER["<b>Message sampler</b><br/>subscribes to the channel,<br/>counts messages per period"]

    PUB --> THROTTLE
    THROTTLE -->|"allowed"| SUB
    THROTTLE -.->|"dropped"| ANALYTICS
    SUB --> SAMPLER
    SAMPLER -.->|"adjusts throttle rate"| THROTTLE

    class THROTTLE emphasis
    class ANALYTICS,SAMPLER muted
```

The throttler reads its current drop rate from the [key-value (KV) store](https://www.pubnub.com/docs/message-processing/serverless/overview.md#key-value-store), a persistent, globally distributed store all Functions on your keyset share. Every instance handling this channel reads the same configuration this way, instead of drifting apart:

```javascript
const db = require('kvstore');
const pubnub = require('pubnub');

// Missing, non-numeric, or unreadable configuration disables throttling.
const readDropRate = () =>
  db.getItem('throttleRate')
    .then((raw) => {
      const rate = Number(raw);
      return Number.isFinite(rate) ? Math.min(Math.max(rate, 0), 1) : 0;
    })
    .catch((error) => {
      console.error('throttleRate unreadable, delivering without throttling:', error);
      return 0;
    });

export default (request) =>
  readDropRate().then((dropRate) => {
    if (Math.random() >= dropRate) {
      return request.ok();
    }

    return pubnub
      .publish({ channel: 'analytics-channel', message: request.message })
      .catch((error) => console.error('Analytics copy failed, message still dropped:', error))
      .then(() => request.abort());
  });
```

The sampler writes `throttleRate` to the KV store as a string between `0` and `1`, the fraction of messages to drop. `getItem()` returns a Promise, so the throttler waits for the value before it decides. A value above `1` counts as `1`, and a value below `0` counts as `0`. If the key is missing, the value isn't a number, or the KV store read fails, the Function logs the problem and delivers every message. This is a choice made in this example. Functions has no built-in default for it, so pick the behavior your event needs. If your event must shed load even without valid configuration, return a fixed fallback rate from `readDropRate()` instead of `0`.

For a dropped message, the Function publishes a copy to `analytics-channel` and waits for that publish to settle before it returns `request.abort()`, which prevents delivery to subscribers. If the analytics copy fails, the Function logs the error and still drops the message. For the full `request` object, the `pubnub` module, and the KV store API, refer to [Functions development](https://www.pubnub.com/docs/message-processing/serverless/development.md).

This throttling is approximate, not an exact cap. Enforcing an exact global rate would require every message to pass through one coordinating component that counts and republishes. That component would become a single point of failure, and it would add its own latency to every message.

The sampler-and-throttler loop above avoids that. Each Function instance makes its own probabilistic decision from shared configuration, so the actual delivered rate tracks the target rate without ever being pinned to it exactly. A brief lag between the sampler noticing a change and the throttler picking it up is expected.

## Channel sharding

Splitting one channel's audience into several separate channels lowers each participant's message volume by reducing how many people share a stream, at the cost of that shared-audience feeling. It only makes sense when a logical grouping already exists, such as:

* Language, so participants converse with others speaking the same language.
* Team or side affiliation, such as home and away sections at a live event.
* Geography, for regionally relevant content.
* A subscription tier, such as a separate room for premium users.

Without one of those groupings, sharding has nothing principled to split on, and dividing an audience by an arbitrary rule, such as random assignment, fragments the community for no benefit. That's the case where throttling a single shared channel, covered above, is the better fit.

Channel sharding here divides an audience so each shard receives less traffic. That's a different problem from the ingress-side sharding [Message aggregation](https://www.pubnub.com/docs/design-patterns/message-aggregation.md#this-is-a-different-sharding-than-an-audience-facing-channel-needs) covers, which fans many publishers into a small, fixed set of channels for a server-side subscriber. Confirm which side of the traffic you're shaping before reaching for either.

## Next steps

* [Send messages effectively](https://www.pubnub.com/docs/design-patterns/send-messages-effectively.md#keep-a-very-busy-channel-usable). Where this choice fits among the rest of your publish-side design.
* [Functions](https://www.pubnub.com/docs/message-processing/serverless/overview.md). Before Publish and After Publish triggers, and what each can do with a message.
* [Functions development](https://www.pubnub.com/docs/message-processing/serverless/development.md). The `request` object, `request.abort()`, and the KV store used above.
* [Occupancy](https://www.pubnub.com/docs/presence/occupancy.md). How a very large channel's presence responses and events change once it turns busy.
* [Message aggregation](https://www.pubnub.com/docs/design-patterns/message-aggregation.md). The other kind of channel sharding: fanning many publishers into a fixed ingress set.
* [Architectural choices](https://www.pubnub.com/docs/design-patterns/architectural-choices.md). Where channel design decisions like this one fit into the rest of your application.

Last updated at: 2026-09-30T07:20:08.000Z
