Kafka’s native cluster mirroring moves disaster recovery into the broker
KIP-1279 preserves offsets and compressed batches across clusters, replacing MirrorMaker 2’s translation layer with broker-managed state.
Apache Kafka’s proposed native cluster mirroring changes cross-cluster replication from an external data pipeline into a broker responsibility. A new Red Hat Developer architecture walkthrough shows how KIP-1279 preserves record offsets, compressed batches and consumer-group state between clusters.
What changes from MirrorMaker 2
MirrorMaker 2 runs Kafka Connect workers that consume from one cluster and produce into another. The broker-native design instead puts mirror fetcher threads in destination brokers. They use Kafka’s fetch protocol to copy committed record batches into local partition logs without decompressing and recompressing them.
That byte-for-byte path matters operationally. The destination keeps the source offsets—including gaps left by compaction—so consumers can fail over without consulting an offset-translation topic. Metadata discovery, topic configuration, group offsets and access-control lists are also synchronized by broker components rather than a separate Connect deployment.
Red Hat describes three parts inside each destination broker: a metadata manager that drives lifecycle transitions, a coordinator that persists state in a compacted __mirror_state topic, and fetcher threads with per-mirror authentication contexts. The design exposes standard broker JMX metrics and lets existing source-side client quotas govern mirror traffic.
Failover is simpler, but still asynchronous
Operators create and start a mirror through kafka-cluster-mirrors.sh. Stopping it removes fetchers, records the last mirror epoch, advances the destination leader epoch, aborts in-flight transactions and resets producer state before making partitions writable.
This creates a short failover sequence: stop the mirror, redirect producers and consumers, and continue from synchronized group offsets. It does not create zero-recovery-point protection. Red Hat says recovery-point loss still depends on replication lag because records that have not reached the destination at failure time are lost. A separate KIP-1360 proposal is intended to add synchronous mirroring later.
The unclean-leader-election behavior also deserves an explicit runbook decision. With recovery enabled, the destination waits for every assigned replica—not only in-sync replicas—to converge on the source’s truncated log. With the option disabled, which the article identifies as the default, the mirror enters a failed state rather than silently accepting divergence.
The migration path may be the bigger prize
The source says mirroring can read Kafka clusters as old as 2.1. That opens a migration route from an older ZooKeeper-based cluster to a new KRaft cluster without upgrading every intermediate major version in place: mirror into the new cluster, stop producers, wait for lag to reach zero, stop mirroring and redirect clients.
Teams evaluating the design should test the controls around it, not only throughput. Validate authentication isolation for every mirror, alert on replication lag and failed partition states, and rehearse both failover and reverse-mirror failback. Broker integration removes infrastructure, but it also moves the replication failure domain inside Kafka itself.
sources
- Data liberation: Apache Kafka's native cluster mirroringdevelopers.redhat.com
comments · 0