Skip to content
Back to Insights
Data EngineeringBy KE Engineering Team

Kafka Raft (KRaft)

Kafka Raft (KRaft)DATA ENGINEERING cover for Kafka Raft (KRaft)replicasDATA ENGINEERINGKafka Raft (KRaft)// KRAFT · DEDICATED CONTROLLER

First written December 2022, last updated September 2026.

Here are the steps and a reference project for setting up Kafka without ZooKeeper, with a dedicated controller quorum, using the Confluent community-licensed container images. A Grafana dashboard to observe the new metrics is also provided.

Introduction

KRaft moves Kafka's metadata consensus out of ZooKeeper and into Kafka itself. With this change, the role of a Kafka instance can be that of a controller, broker, or both. Getting this configuration stood up required some tweaks to the Confluent cp-kafka image before version 7.5.

If you want a deeper understanding of the design and implementation details, check out Confluent's Kafka Internals course. Specifically, the control-plane section.

Configuration

There are many configuration parameters with Apache Kafka; highlighted here are the ones necessary to build out the cluster with KRaft.

node.id

The property node.id replaces broker.id. Be sure that all identifiers in the cluster are unique across brokers and controllers (docs).

process.roles

A node can be a broker, a controller, or both. Set to broker,controller to enable both the controller and data planes on a node (docs).

controller.listener.names

List the listener names used for the controller. This tells a node which listener to use for that communication. While the property is a list, just like advertised.listeners, the first one is what is used for controller communication (docs).

controller.quorum.voters

A comma-delimited list of voters in the control plane, where a controller is noted as: node_id@hostname:port (docs). Newer Kafka versions support dynamic controller quorums, so a static voters list is no longer the only option.

example:

yaml
KAFKA_CONTROLLER_QUORUM_VOTERS: 10@controller-0:9093,11@controller-1:9093,12@controller-2:9093

Additional Configuration

There are other tuning parameters for the controller plane, see the documentation for details.

Lesson Learned

Do not remove cluster settings from the dedicated controllers, since a controller is the node that performs administration operations, such as creating a topic.

Incorrectly removing these from the controllers caused topics to be created without Apache Kafka defaults.

yaml
KAFKA_DEFAULT_REPLICATION_FACTOR: 3KAFKA_NUM_PARTITIONS: 4

Storage

Another change to setting up Apache Kafka with KRaft is the storage. The storage on each node must be configured, before starting the JVM. This can be done with a kafka-storage command provided as part of Apache Kafka. With cp-kafka 7.5 and later, the container runs kafka-storage on start-up if the storage directory doesn't exist, so only the cluster UUID needs to be provided.

  • Generate a unique UUID for the cluster; you can use kafka-storage random-uuid or another means.
  • Before starting the cluster, format the metadata storage with kafka-storage format.
shell
kafka-storage format -t $KAFKA_CLUSTER_ID -c <server.properties>

Seeing It In Action

If you are interested in seeing all this in action, check out the kafka-raft docker-compose setup in the dev-local project. It's a fully working example with 3 controllers and 4 brokers.

Metrics

If you are going to deploy Kafka with KRaft to production, having visibility into metrics is important. That visibility matters as much as having dedicated controllers, maybe more. The key with dashboards is to ensure they report correctly whether the data and control planes are separate or combined.

Grafana

Building a Grafana dashboard is a multi-step process: extract the metrics, store them in a time-series database (e.g. Prometheus), then visualize the collected metrics in Grafana.

JMX Prometheus Exporter

The KRaft Monitor metrics are defined in the documentation, with an MBean name, such as kafka.server:type=raft-metrics,name=current-state. Using a JMX client, such as jmxterm, shows the MBean name is just kafka.server:type=raft-metrics with each metric an attribute in that bean. This differed from the documentation at the time of writing.

If you deploy Java applications with JMX metrics in containers, we highly recommend jmxterm.

shell
java -jar jmxterm.jar

In the cp containers, the Java process is process 1, but use the command jvms to see all available processes and verify the JVM’s process ID is indeed 1.

shell
$>open 1

Show all the MBeans in the JVM, with beans.

shell
$>beans...kafka.server:type=raft-metrics...

Select a bean and use info to explore the attributes on a bean and get to fetch the current value of an attribute.

shell
$>bean kafka.server:type=raft-metrics#bean is set to kafka.server:type=raft-metrics
shell
$>info# attributes%0   - append-records-rate (double, r)%1   - commit-latency-avg (double, r)%2   - commit-latency-max (double, r)%3   - current-epoch (double, r)%4   - current-leader (double, r)%5   - current-state (double, r)%6   - current-vote (double, r)%7   - election-latency-avg (double, r)%8   - election-latency-max (double, r)%9   - fetch-records-rate (double, r)%10  - high-watermark (double, r)%11  - log-end-epoch (double, r)%12  - log-end-offset (double, r)%13  - number-unknown-voter-connections (double, r)%14  - poll-idle-ratio-avg (double, r)
shell
$>get current-leadercurrent-leader = 10.0;

Using the above, the following properly exposes these from JMX Prometheus Exporter. Since the current-state attribute value is a string, its value needs to be associated with a label to capture it in Prometheus.

yaml
rules:- pattern: "kafka.server<type=raft-metrics><>(current-state): (.+)"  name: kafka_server_raft_metrics  labels:  name: $1  state: $2  value: 1- pattern: "kafka.server<type=raft-metrics><>(.+): (.+)"  name: kafka_server_raft_metrics  labels:  name: $1

Grafana Dashboard

Grafana is great for custom configuration, but that means time and effort are needed to build dashboards. Here is a dashboard around some of those KRaft metrics; it isn't a complete dashboard.

Node Information

With metrics emitting, put them into a Grafana dashboard.

Using current-state, we can see which node is the leader, in addition to capturing node.id. In addition, the dashboard component joins in other data to provide additional information on each node.

Grafana node information panel listing each Kafka node with its id, role, and current KRaft state
Fig. 1: Node Information

Node Counts

Counts of nodes and active controllers are always reassuring, and this uses an existing metric, kafka_controller_kafkacontroller_value{name="ActiveControllerCount",}. This metric is only emitted from a controller, so by counting the existence of this metric you see the number of controllers in the cluster, and by summing the value of the metric you get the actual number of active controllers; alert if this is ever not equal to 1.

Check out the dashboard to see how the other values are calculated, as it is the same as in the ZooKeeper-based installations.

Grafana single-value panels showing node counts and the number of active controllers
Fig. 2: Nodes

Active Controller

To get the node.id of the quorum leader, just find max(kafka_server_raft_metrics{name="current-leader",}). Because scraping of each node is from slightly different times, different values are possible at the time of a change; max is used to make the single value display easy to build.

If a new leader is being elected, it shows up in the voted-leader metric. In a single-value panel this isn't very useful, but it is worth recording as a time series.

Grafana panel showing the node id of the current KRaft quorum leader
Fig. 3: Leader

Full Dashboard

Here is an example dashboard that also captures the fetch rate of the controller metadata.

Full Grafana KRaft dashboard with node information, counts, active controller, and controller metadata fetch rate
Fig. 4: Full Dashboard

Open-Source Tools

At the time of writing, we checked a variety of open-source tools and could not see KRaft metrics in any of them; active-controller information displayed incorrectly. None of the tools we tried supported KRaft yet, but it is important that before you upgrade, you have a proper monitoring and alerting strategy in place.

Before you move to KRaft

Be sure you properly test your monitoring and Apache Kafka support infrastructure as part of your move to KRaft. Also, validate that kafka-raft metrics are captured and confirm that dashboards and tools work when brokers and controllers are on separate nodes.

Working on something like this?

Start a Conversation