Skip to content

[improve][broker] Change the default number of namespace bundles to 32 and make the system namespace bundle count configurable - #26610

Merged
merlimat merged 1 commit into
apache:masterfrom
lhotari:lh-default-bundles-32
Sep 16, 2026
Merged

merlimat merged 1 commit into
apache:masterfrom
lhotari:lh-default-bundles-32

Conversation

@lhotari

@lhotari lhotari commented Sep 16, 2026 •

Copy link
Copy Markdown
Member

Revives #12137 (closed in 2021 because public/default was already created with 16 bundles by initialize-cluster-metadata; the real cause of the imbalance seen there was the load balancer, which #26609 addresses).

Motivation

Bundles are the unit of assignment of topics to brokers. A namespace created through the admin API, by pulsar standalone, or for the function worker gets defaultNumberOfNamespaceBundles=4 bundles, so its topics can never be spread over more than 4 brokers until bundle splitting catches up, and on a small cluster this routinely leaves a broker idle. Meanwhile public/default has always been created with 16 bundles by initialize-cluster-metadata, and the documentation recommends starting with more bundles than brokers (64 for 1000 topics on 16 brokers). The default should be a reasonable setting for most users out of the box rather than something every namespace has to override.

The pulsar/system namespace is a separate case. Its topics are a small, fixed set, and the transaction coordinators are the ones that matter: a coordinator is owned by whichever broker owns the bundle of its transaction_coordinator_assign partition, so the number of system bundles bounds how far the coordinators can spread. Hashing the 16 default partition names (crc32, as ConsistentHashingTopicBundleAssigner does) into candidate bundle counts:

system bundles distinct bundles holding the 16 default coordinators max coordinators in one bundle
4 4 5
16 (previous initialize-cluster-metadata value) 8 3
32 10 2
64 16 (one each) 1

Today the tools create pulsar/system with 16 bundles, but when the extensible load manager or pulsar standalone creates it on start-up it gets defaultNumberOfNamespaceBundles (4) instead.

Modifications

  • defaultNumberOfNamespaceBundles: 4 → 32 (ServiceConfiguration, broker.conf, standalone.conf, the terraform template, faq.md). The value is now a public constant on ServiceConfiguration so that initialize-cluster-metadata (public/default, -bn default: 16 → 32) and initialize-namespace use the same number and a namespace gets the same bundle count whichever way it is created.
  • New defaultNumberOfSystemNamespaceBundles=64 for pulsar/system when the broker (extensible load manager) or standalone creates it, and a matching -sbn / --system-namespace-bundle-number option on initialize-cluster-metadata and initialize-transaction-coordinator-metadata with the same default. The rationale (each default coordinator in its own bundle) is pinned by NamespaceBundlesTest.
  • ServiceUnitStateChannelImpl and PulsarStandalone use the new setting instead of defaultNumberOfNamespaceBundles.

How the concerns raised on #12137 and #854 apply

Two concerns were raised when this default was last discussed:

  1. Broker shutdown time (#854 (comment)): topics in a bundle are closed in parallel but bundles are unloaded one by one, so more bundles means a longer graceful shutdown. This is still how BrokerService.unloadNamespaceBundlesGracefully works for the modular load manager (the extensible load manager transfers ownership of all bundles concurrently in ServiceUnitStateChannelImpl.doCleanup, batched by loadBalancerServiceUnitStateMaxConcurrentOverrides).
  2. Metadata cost (#12137 (comment)): more bundles means more znodes and larger load reports.

In practice neither changes with this PR, because only bundles that have been looked up cost anything. A bundle gets an ownership entry in the metadata store, an entry in the broker's load report and an unload step at shutdown only once a topic in it has been looked up; the 30 bundles of a namespace with two topics are never owned and cost nothing. The cost scales with the number of active bundles, which is bounded by the number of active topics, not by this default. A namespace with enough topics or traffic to occupy all 32 bundles would have been split well past 4 bundles by the default auto-split (loadBalancerAutoBundleSplitEnabled=true, up to loadBalancerNamespaceMaximumBundles=128) anyway, so its unload and metadata footprint is the same with either default; the difference is that its topics are spread from the start instead of after the split cycles. The same holds for pulsar/system: of its 64 bundles, only the ~16 holding coordinator partitions (plus the load balancer's internal topics) are ever owned.

The metadata store side has also moved on since #12137 was discussed. Transparent batching of metadata operations (#13043, 2.10.0, metadataStoreBatchingEnabled=true by default) coalesces the ownership reads and writes that bundle lookups, unloads and load reports generate, so the per-bundle metadata cost of an owned bundle is a fraction of what it was in 2021, and PIP-335 (#22007) added Oxia as a metadata store for clusters whose metadata volume outgrows ZooKeeper.

The one-way nature of splits (bundles can be split but not merged) still applies, but the default only affects newly created namespaces, and both defaultNumberOfNamespaceBundles and --bundles at creation remain available for namespaces that should be smaller.

A possible follow-up that would remove the remaining dependency of shutdown time on the bundle count for the modular load manager is to unload bundles in parallel at shutdown under a bound on the number of topics in flight, as suggested in #854 (comment); that is independent of this change.

Verifying this change

  • Make sure that the change passes the CI checks.

This change added tests and can be verified as follows:

  • NamespaceBundlesTest.testDefaultSystemNamespaceBundlesSpreadTheDefaultTransactionCoordinators pins that the 16 default coordinator partitions hash into 16 distinct bundles with the default system bundle count.
  • ClusterMetadataSetupTest.testSetBundleNumberForDefaultNamespace now covers --system-namespace-bundle-number together with --default-namespace-bundle-number (given / not given) and asserts both namespaces' bundle counts.
  • ExtensibleLoadManagerImplTest.testSystemNamespaceBundleNumber verifies that the channel creates pulsar/system with defaultNumberOfSystemNamespaceBundles (the test configuration sets defaultNumberOfNamespaceBundles=1).
  • PulsarStandaloneTest asserts the bundle count of public/default and the topic's bundle range for 32 bundles.
  • ServiceConfigurationTest checks that broker.conf matches the Java defaults.

Does this pull request potentially affect one of the following parts:

If the box was checked, please highlight the changes

  • Dependencies (add or upgrade a dependency)
  • The public API
  • The schema
  • The default values of configurations
  • The threading model
  • The binary protocol
  • The REST endpoints
  • The admin CLI options
  • The metrics
  • Anything that affects deployment

Default values changed: defaultNumberOfNamespaceBundles (4 → 32); initialize-cluster-metadata --default-namespace-bundle-number default (16 → 32); new defaultNumberOfSystemNamespaceBundles (64) and --system-namespace-bundle-number (64) where pulsar/system previously got 16 from the tools or 4 from the broker/standalone. Existing namespaces are not modified. The pulsar-site load-balancing page and the CLI reference will need the new values.

…2 and make the system namespace bundle count configurable

Revives apache#12137. The modular load manager only moves bundles, so a namespace
with the previous default of 4 bundles cannot spread over more than 4
brokers and, on a small cluster, routinely leaves a broker idle until
bundle splitting catches up. public/default has always been created with
16 bundles by initialize-cluster-metadata, while a namespace created
through the admin API got 4.

- defaultNumberOfNamespaceBundles: 4 -> 32. Unowned bundles cost nothing
  (no ownership entry, no load report entry, no unload at shutdown), so a
  small namespace is not affected; a namespace with many topics spreads
  across up to 32 brokers without waiting for splits.
- initialize-cluster-metadata and initialize-namespace use the same
  default (32, was 16) so a namespace gets the same bundle count however
  it is created.
- New defaultNumberOfSystemNamespaceBundles (64) and a matching
  --system-namespace-bundle-number option on initialize-cluster-metadata
  and initialize-transaction-coordinator-metadata. pulsar/system used to
  get 16 bundles from the tools but defaultNumberOfNamespaceBundles (4)
  when the extensible load manager or standalone created it. A
  transaction coordinator is owned by whichever broker owns the bundle of
  its transaction_coordinator_assign partition; with 16 bundles the 16
  default coordinators hash into only 8 bundles, 64 is the smallest count
  that gives each its own bundle (pinned by a test).

The broker-shutdown concern from apache#854 (bundles unload one by one) still
holds for the modular load manager, but only for owned bundles, and the
extensible load manager transfers ownership concurrently.

Assisted-by: Claude Code (Opus 5)
@lhotari lhotari added this to the 5.0.0 milestone Sep 16, 2026
@merlimat
merlimat merged commit ab58ba4 into apache:master Sep 16, 2026
83 of 85 checks passed
lhotari added a commit to apache/pulsar-site that referenced this pull request Sep 18, 2026
…5.0 load balancing defaults

Follows apache/pulsar#26609 (AvgShedder as the default shedding and
placement strategy) and apache/pulsar#26610 (defaultNumberOfNamespaceBundles
4 -> 32, defaultNumberOfSystemNamespaceBundles=64 and
--system-namespace-bundle-number).

New pages under Administration:
- Namespace bundles: how many bundles each kind of namespace gets and from
  which tool or setting, what an owned bundle costs (and that unused bundles
  cost nothing), how to choose the count, why the system namespace gets 64
  bundles (one per default transaction coordinator), splitting, and the
  effect of bundles on broker restarts.
- Rolling restarts: what happens when a broker stops with either load
  manager, the orphan path when it is killed, draining with
  `pulsar-admin brokers shutdown --max-concurrent-unload-per-sec`,
  pausing shedding and auto-split with dynamic config, waiting for a broker
  to be listed, healthy and reporting before stopping the next one, moving
  bundles with --destinationBroker, and the Helm chart settings
  (gracePeriod, updateStrategy OnDelete, publishNotReadyAddresses).

administration-load-balance is added to the Administration sidebar (it was
only reachable from the architecture overview) and updated: both load
managers named, 32 bundles, an AvgShedder section and the default note, a
TransferShedder section, the complete split-algorithm list, unloading to a
destination broker, the duplicated anti-affinity section replaced by a link
to its own page, and a Related topics section.

Also updated for consistency: concepts-broker-load-balancing-concepts
(defaults, fallback placement, loadBalancerDistributeBundlesEvenlyEnabled),
-migration (AvgShedder), -quick-start (link), deploy-bare-metal and
-multi-cluster (initialize-cluster-metadata bundle options),
administration-upgrade and helm-upgrade (link to the restart procedure).

Assisted-by: Claude Code (Opus 5)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants