Consumer incremental rebalance and static group membership

Kafka static group membership (group.instance.id) and incremental cooperative rebalancing: two ways to cut consumer downtime during Kafka rebalancing.

By Stéphane Derosiaux · October 1, 2026

Learn how to minimize rebalance disruption

Kafka's incremental cooperative rebalancing and static group membership features reduce the disruption caused by consumer group rebalances, improving overall system stability and performance.

What you'll learn:

  • How incremental cooperative rebalancing reduces processing downtime
  • The multi-phase rebalancing process and how it differs from eager rebalancing
  • How to configure static group membership for stable consumer identities
  • Best practices for production deployments

Traditional rebalancing problems

Eager rebalancing (default on the classic protocol)

  • Stop-the-world: All consumers stop processing during rebalance
  • Complete reassignment: All partitions are revoked and reassigned
  • Processing downtime: No messages processed during rebalance period
  • Cascading rebalances: One consumer failure affects entire group

Performance impact

Consumer 1: [P0, P1, P2] → [ ] → [P0, P1]
Consumer 2: [P3, P4, P5] → [ ] → [P2, P3]
Consumer 3: [P6, P7, P8] → [ ] → [P4, P5, P6, P7, P8]

All consumers stop processing during the transition

Eager versus incremental rebalancing comparison

This diagram compares the impact of eager rebalancing versus incremental cooperative rebalancing:

Incremental cooperative rebalancing

How it works (Kafka 2.4+)

  • Minimal disruption: Only affected partitions are reassigned
  • Continued processing: Unaffected partitions continue processing
  • Gradual transition: Rebalance happens in multiple phases
  • Reduced downtime: Significantly shorter processing interruptions

Rebalancing phases

The incremental rebalance happens in two distinct phases, minimizing disruption:

Key improvements over eager rebalancing:

  • ✅ Only ONE consumer stops ONE partition (P2)
  • ✅ Eight out of nine partitions never stop processing
  • ✅ Cascading failures prevented

Configuration

# Enable incremental cooperative rebalancing (available since Kafka 2.4, not the default)
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor

The default partition.assignment.strategy since Kafka 3.0 is RangeAssignor, CooperativeStickyAssignor. The group uses the first strategy every member supports, so default consumers rebalance with RangeAssignor, which is eager. Removing RangeAssignor from the list in one rolling restart switches the group to cooperative rebalancing. RangeAssignor and RoundRobinAssignor are always eager.

Kafka 4.x: the consumer rebalance protocol (KIP-848)

With group.protocol=consumer (GA in Kafka 4.0), the broker computes assignments and rebalances are always incremental. partition.assignment.strategy, session.timeout.ms and heartbeat.interval.ms cannot be set on the client: the consumer throws a ConfigException at startup. Session timeout and heartbeat interval become broker-side group configs (group.consumer.session.timeout.ms, group.consumer.heartbeat.interval.ms, overridable per group). Static membership with group.instance.id still works. The settings on this page apply to the classic protocol. See Kafka consumer rebalancing for both protocols.

Static group membership

Concept

Static group membership allows consumers to maintain stable identities across restarts, preventing unnecessary rebalances during planned maintenance or brief outages.

Benefits

  • Fewer rebalances: Consumer restarts don't trigger rebalances
  • Stable assignments: Partitions stay with the same consumer instance
  • Faster recovery: Consumers can resume processing from where they left off
  • Operational efficiency: Planned maintenance doesn't disrupt other consumers

Configuration

# Assign static member ID to consumer
group.instance.id=consumer-instance-1

# Session timeout longer than one restart (classic protocol only)
session.timeout.ms=120000  # 2 minutes

# Keep the heartbeat interval at its default: members learn about
# a rebalance through heartbeats, so a long interval slows every rebalance
heartbeat.interval.ms=3000

The trade-off: a static member that crashes keeps its partitions unassigned for the full session.timeout.ms before the group rebalances.

Consumer lifecycle

Properties props = new Properties();
props.put("group.id", "my-consumer-group");
props.put("group.instance.id", "consumer-1"); // Static member ID
props.put("session.timeout.ms", "120000");    // 2 minutes, classic protocol only

KafkaConsumer<String, String> consumer = new KafkaConsumer<>(props);

Use cases and benefits

High-availability applications

# Configuration for critical applications
# Must be stable across restarts: a process ID changes on every restart
group.instance.id=${hostname}
session.timeout.ms=120000
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor

# Allows planned restarts without affecting other consumers

Containerized environments

# Kubernetes StatefulSet example: pod names (consumer-0, consumer-1, ...)
# survive restarts, unlike Deployment pod names, which change every time
apiVersion: apps/v1
kind: StatefulSet
spec:
  template:
    spec:
      containers:
      - name: kafka-consumer
        env:
        - name: POD_NAME
          valueFrom:
            fieldRef:
              fieldPath: metadata.name
        - name: GROUP_INSTANCE_ID
          value: "consumer-$(POD_NAME)"
        # Consumer keeps its identity across pod restarts

Stream processing applications

  • State preservation: Local state stores remain associated with specific consumers
  • Reduced reprocessing: Avoid recomputing state after rebalances
  • Consistent partitioning: Same consumer always processes same partitions

Monitor and observe

Key metrics

  • Rebalance frequency: Number of rebalances per time period
  • Rebalance duration: Time taken for rebalance completion
  • Partition assignment stability: How often partitions change owners
  • Consumer lag during rebalance: Processing delay during rebalances

JMX metrics

# Rebalance metrics
kafka.consumer:type=consumer-coordinator-metrics,client-id=*
- rebalance-rate-per-hour
- rebalance-latency-avg
- rebalance-latency-max

# Assignment metrics
kafka.consumer:type=consumer-coordinator-metrics,client-id=*
- assigned-partitions

Configuration best practices

For incremental rebalancing

# Use cooperative assignors
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor

# Optimize for stability
session.timeout.ms=45000
heartbeat.interval.ms=3000
max.poll.interval.ms=300000

For static group membership

# Stable consumer identity
group.instance.id=unique-consumer-id

# Session timeout longer than one restart (classic protocol only)
session.timeout.ms=120000    # 2 minutes
heartbeat.interval.ms=3000

# Prevent accidental timeouts
max.poll.interval.ms=900000  # 15 minutes

Combined configuration

# Best of both worlds
group.instance.id=consumer-${hostname}
partition.assignment.strategy=org.apache.kafka.clients.consumer.CooperativeStickyAssignor
session.timeout.ms=120000
heartbeat.interval.ms=3000
max.poll.interval.ms=600000

Operational considerations

Deployment strategies

  1. Rolling updates: Use static group membership so each restart does not trigger a rebalance (the restarting member's partitions pause until it returns)
  2. Blue-green: Never run two live instances with the same group.instance.id: one of them is fenced (FENCED_INSTANCE_ID on the classic protocol), which is fatal for that consumer
  3. Canary releases: Incremental rebalancing minimizes impact on stable consumers

Maintenance windows

# Planned consumer restart with static membership
# 1. Consumer stops gracefully
# 2. Other consumers continue processing (no rebalance)
# 3. Consumer restarts with same group.instance.id
# 4. Resumes processing assigned partitions

Troubleshooting

Common issues and solutions:

  • Duplicate static IDs: Ensure unique group.instance.id per consumer
  • Long session timeouts: Balance between stability and failure detection
  • Assignment strategy conflicts: Ensure all consumers use compatible assignors

When migrating to incremental rebalancing and static membership, start with incremental rebalancing first, monitor rebalance behavior and performance, then gradually introduce static group membership and test failure scenarios thoroughly.

Static member considerations

  • Static members that don't restart within session.timeout.ms will be removed from the group
  • Ensure unique group.instance.id values to avoid conflicts
  • Plan for scaling scenarios where static IDs need management

Performance impact

Before (eager rebalancing)

Rebalance triggered → All consumers stop → Complete reassignment → Resume processing
Downtime: every partition in the group, until the last member has rejoined

After (incremental + static)

Rebalance triggered → Only affected partitions stop → Minimal reassignment → Resume processing
Downtime: only the partitions that move; restarts within session.timeout.ms cause no rebalance

Measure the actual gain on your own groups with the rebalance-latency-avg and rebalance-rate-per-hour metrics above, before and after the change.

See it in practice with Conduktor

Conduktor Console visualizes consumer group rebalances in real-time, showing which partitions are being reassigned and the rebalance duration. Monitor consumer lag during rebalances to verify that incremental rebalancing is minimizing disruption as expected.

Next steps