Keep chat readable when the audience is very large

PubNub's Pub/Sub model delivers every message published to a channel to every subscriber on that channel. On one channel, that means the read rate a fan experiences is the sum of what every other fan publishes, not just their own. This tutorial builds two mechanisms that keep that rate under control: throttling a publisher so it generates less traffic, and sharding a channel so fewer subscribers share each message. The last part of this tutorial reads live occupancy figures to place a fan in the emptiest shard. That reading comes from Presence, so Presence must be enabled on your keyset for it to work.

What you'll build​

fan.js asks Presence for the occupancy of each shard channel and joins the first one with room, such as game.chat.shard-0. It then throttles its own publishes to that shard, dropping any that come less than two seconds apart. A Before Publish Function on PubNub's network counts publishes per client IP and blocks every publish over the limit, even from a client that bypasses fan.js. watch-shards.js reads the same occupancy figures so you can add shards before the next match.

Before you begin​

Make sure you have:

  1. Node.js 22 or later installed.
  2. A PubNub account and your own keyset. If you don't have one, follow Set up your account to create one.
  3. Presence enabled on that keyset.
  4. The Functions developer role, or Account Admin access, for the Put the limit somewhere the fan cannot edit step later on this page.

A keyset is the set of publish, subscribe, and secret keys that identifies your application to the PubNub network. Presence must be enabled on the keyset, and new keysets have it enabled by default. Enable it from your keyset's settings in the Admin Portal, the same place you enable other features in Set up your account.

Set up the project​

Create a new directory and install the PubNub SDK:

mkdir sme-rate-limiting
cd sme-rate-limiting
npm init -y
npm install pubnub

Add "type": "module" to package.json, since this tutorial's code uses import.

Create two files: fan.js, which plays the part of a single fan's client, and watch-shards.js, a separate script that reports on shard occupancy.

Add this to fan.js:

import PubNub from 'pubnub';

const pubnub = new PubNub({
publishKey: 'YOUR_PUBLISH_KEY',
subscribeKey: 'YOUR_SUBSCRIBE_KEY',
userId: 'fan-42',
});

Add the same initialization to watch-shards.js, with userId: 'match-service' instead, since this script represents your own infrastructure rather than one fan's client. Replace YOUR_PUBLISH_KEY and YOUR_SUBSCRIBE_KEY in both files with the keys from your keyset.

Choose a lever before you write code​

Throttling and sharding solve the same problem in different ways, and they are not interchangeable.

Throttling keeps every fan on game.chat, so a fan who throttles their own publishes still sees everyone else's messages. The cost is that some of that fan's own messages get dropped before they reach anyone.

Sharding keeps every message that gets published, but it splits the audience across several channels. A fan in game.chat.shard-0 no longer sees a fan in game.chat.shard-1, even though both are watching the same match.

Reach for throttling first. It requires no decision about how to divide the audience, and it keeps the single-crowd feeling a shared match chat is usually for. Shard only once your audience already splits along a real line, such as language or which team a fan follows, a rule Channel sharding covers in full. Without a line like that, sharding fragments the crowd for no benefit, and throttling is the better fit.

This tutorial builds both, in the order you would actually add them: throttle first, then shard once the audience is large enough that a throttled channel still needs splitting.

Throttle the fan's own client​

Add this to the end of fan.js:

1

This stops fan.js from calling publish() more often than once every minimumMillisecondsBetweenMessages, dropping any call that arrives sooner rather than queuing it. Most fans type slower than any interval you'd set here, so this alone removes traffic that nobody downstream could have read anyway.

It buys you nothing against a fan who bypasses your client. Anyone who edits the JavaScript running in their browser, or writes their own script against your publish key, calls pubnub.publish() directly and skips this check entirely. Client-side throttling is a courtesy your own client extends, not a limit the network enforces.

This code references a shard variable that doesn't exist yet. That's expected. It's defined in Place a fan in a shard that has room, and fan.js isn't complete, or runnable, until you add that code above this block.

Put the limit somewhere the fan cannot edit​

A throttle inside fan.js only limits fans who run your unmodified client. To cap every publisher regardless of what they run, the limit has to live somewhere a fan's own code never executes. That somewhere is a Before Publish Function, which runs on PubNub's network before a message reaches subscribers.

Create a Package and a Function the same way Create a Function does, selecting Before Publish or Fire as the event type and setting the trigger channel to game.chat. Replace the sample code with this:

export default (request) => {
const db = require('kvstore');

const maxMessagesPerWindow = 20;
const windowKey = `game.chat.rate.${request.meta.clientip}.${Math.floor(Date.now() / 60000)}`;

return db.getCounter(windowKey).then((countThisWindow) => {
if (countThisWindow >= maxMessagesPerWindow) {
return request.abort();
}

return db.incrCounter(windowKey).then(() => request.ok());
});
};

require('kvstore') loads the KV store, a persistent store every Function on your keyset shares. So every instance of this Function counts against the same total, regardless of which edge location ran it. getCounter() returns 0 for a key nothing has incremented yet, so the first publish in a window needs no setup. incrCounter() increments atomically, which means two publishes arriving at the same instant both register.

windowKey folds the current minute into the key, so each publisher gets a fresh counter every 60 seconds without you clearing anything. request.meta.clientip stands in for the publisher's identity, because the request object a Before Publish Function receives carries the publishing client's IP address rather than a User ID. Fans behind one IP therefore share a counter.

Reading the counter and incrementing it are two operations, not one, so two publishes that arrive in the same instant can both read the same value and both be allowed. A publisher can slip one or two messages past the cap that way. For a chat limiter that is the right trade, because the goal is a readable channel rather than an exact count.

KV store counters have no time to live (TTL), so one key per publisher per minute accumulates for as long as your keyset lives. Widen the window if that matters to you, or store the count with setItem() and a TTL instead and accept that the count is no longer atomic.

request.abort() blocks delivery to subscribers for a message over the limit. Deploy the Revision to your keyset the same way Create a Function does, or it saves without processing any live traffic.

If you shard game.chat later in this tutorial, target this Function at the wildcard pattern game.chat.* instead of the single channel name, so the same limit applies across every shard. Wildcard targeting needs the Stream Controller add-on with Wildcard Subscribe enabled on your keyset.

Split the chat when one channel is not enough​

Even with the Before Publish Function capping any one fan, a channel's total read volume is still every allowed publish multiplied by every subscriber on it. Past some occupancy, the fix is not a lower cap. It's fewer subscribers per channel.

Split game.chat into game.chat.shard-0, game.chat.shard-1, and as many more as you need, and place each fan in exactly one. A fan in game.chat.shard-0 only ever sees messages published to game.chat.shard-0. They don't see game.chat.shard-1, even though both shards carry the same match.

That's a product decision as much as a technical one. Splitting the chat means no single fan experiences the full crowd, only their own slice of it. Decide that trade-off deliberately for your own match-day chat before you place a single fan in a shard.

Place a fan in a shard that has room​

Add this to fan.js, directly above the throttling code you added in Throttle the fan's own client, not below it. The throttle code publishes to a shard variable, so that variable has to exist first.

1

This calls Here Now across every shard channel and returns the first one under maxFansPerShard, following the pattern in Get online users in a channel. Here Now is a point-in-time snapshot, so a shard could have picked up a few more fans since this call was made — and that's fine here. You're finding a shard with room, not enforcing an exact cap, so a reading that's slightly behind still picks correctly almost every time.

The number of subscribers on a PubNub channel is unlimited, so nothing about the platform sets maxFansPerShard for you. The value here, 1000, is this example's own choice, not a limit read from anywhere, and it's yours to raise or lower based on how full you're willing to let one shard get.

If Presence isn't enabled, or nobody has subscribed to a shard yet, Here Now returns occupancy: 0 for it, the same response an empty channel gives. pickShardForFan then returns that shard, which is correct. An untracked or empty channel and a genuinely empty one look identical, and either way, the shard has room.

Watch the shards during the match​

Add this to watch-shards.js:

1

This calls Here Now across game.chat.shard-0, shard-1, and shard-2, and logs each one's occupancy. Add more channel names to the list as you add more shards.

Use what it shows you to add shards before the next match, not during this one. Shards drain unevenly as fans leave, some empty out fast and others stay fuller, and that's expected. Don't migrate a fan out of a shard mid-match to even things out. Accept the uneven drain, and correct occupancy for the next match by planning more shards from the start.

Run it​

Start the fan's client:

node fan.js

On a fresh keyset with no other fans subscribed yet, you should see:

this fan joins game.chat.shard-0

fan.js picked the first shard, since all of them report occupancy: 0, then published its one throttled test message silently, the same way a successful publish always behaves in this code. Press Ctrl+C to stop it.

In a second terminal, check occupancy across the shards:

node watch-shards.js

You should see:

game.chat.shard-0 holds 0 fans
game.chat.shard-1 holds 0 fans
game.chat.shard-2 holds 0 fans

Every shard reads 0, because Here Now reports subscribers, and nothing in this tutorial subscribed a fan to a shard, only published into one. Run watch-shards.js again anytime during a real match for a fresh read.

What happened​

  1. fan.js throttled its own publishes to at most one every two seconds, dropping anything faster before it ever reached the network.
  2. The Before Publish Function you deployed counted publishes per client IP in one-minute windows, and would have aborted delivery for any publisher over the limit, whether or not that publisher was running your client at all.
  3. fan.js called Here Now across every shard channel and picked the one furthest from its configured cap.
  4. watch-shards.js called Here Now again from a separate script, the same read your own operations process would use to decide when to add shards.

Throttling and the Function both act on message frequency. Sharding acts on occupancy instead, by reducing how many subscribers share each message. Running both in the same tutorial is the point. A very large audience usually needs the frequency lever and the occupancy lever at different times, not one or the other forever.

Next steps​

  • Rate limiting. The full reasoning behind choosing throttling or sharding, including the feedback-loop pattern for an aggregate rate limit instead of a per-publisher one.
  • Functions. Every trigger type, not just Before Publish, and the built-in modules available to each.
  • Occupancy. Why a Here Now response is capped and cached the way it is, and how it changes once a channel gets busy enough to switch to interval mode.
  • Send messages effectively. Where this choice fits among the rest of your publish-side design.

Was this page useful?

Last updated on