Kafka Raft (KRaft)
First written December 2022, last updated September 2026.
Here are the steps and a reference project for setting up Kafka without ZooKeeper, with a dedicated controller quorum, using the Confluent community-licensed container images. A Grafana dashboard to observe the new metrics is also provided.
Introduction
KRaft moves Kafka's metadata consensus out of ZooKeeper and into Kafka itself. With this change, the role of a Kafka instance can be that of a controller, broker, or both. Getting this configuration stood up required some tweaks to the Confluent cp-kafka image before version 7.5.
If you want a deeper understanding of the design and implementation details, check out Confluent's Kafka Internals course. Specifically, the control-plane section.
Configuration
There are many configuration parameters with Apache Kafka; highlighted here are the ones necessary to build out the cluster with KRaft.
node.id
The property node.id replaces broker.id. Be sure that all identifiers in the cluster are unique across brokers and controllers (docs).
process.roles
A node can be a broker, a controller, or both. Set to broker,controller to enable both the controller and data planes on a node (docs).
controller.listener.names
List the listener names used for the controller. This tells a node which listener to use for that communication. While the property is a list, just like advertised.listeners, the first one is what is used for controller communication (docs).
controller.quorum.voters
A comma-delimited list of voters in the control plane, where a controller is noted as: node_id@hostname:port (docs). Newer Kafka versions support dynamic controller quorums, so a static voters list is no longer the only option.
example:
KAFKA_CONTROLLER_QUORUM_VOTERS: 10@controller-0:9093,11@controller-1:9093,12@controller-2:9093Additional Configuration
There are other tuning parameters for the controller plane, see the documentation for details.
Lesson Learned
Do not remove cluster settings from the dedicated controllers, since a controller is the node that performs administration operations, such as creating a topic.
Incorrectly removing these from the controllers caused topics to be created without Apache Kafka defaults.
KAFKA_DEFAULT_REPLICATION_FACTOR: 3KAFKA_NUM_PARTITIONS: 4Storage
Another change to setting up Apache Kafka with KRaft is the storage. The storage on each node must be configured, before starting the JVM. This can be done with a kafka-storage command provided as part of Apache Kafka. With cp-kafka 7.5 and later, the container runs kafka-storage on start-up if the storage directory doesn't exist, so only the cluster UUID needs to be provided.
- Generate a unique UUID for the cluster; you can use
kafka-storage random-uuidor another means. - Before starting the cluster, format the metadata storage with
kafka-storage format.
kafka-storage format -t $KAFKA_CLUSTER_ID -c <server.properties>Seeing It In Action
If you are interested in seeing all this in action, check out the kafka-raft docker-compose setup in the dev-local project. It's a fully working example with 3 controllers and 4 brokers.
Metrics
If you are going to deploy Kafka with KRaft to production, having visibility into metrics is important. That visibility matters as much as having dedicated controllers, maybe more. The key with dashboards is to ensure they report correctly whether the data and control planes are separate or combined.
Grafana
Building a Grafana dashboard is a multi-step process: extract the metrics, store them in a time-series database (e.g. Prometheus), then visualize the collected metrics in Grafana.
JMX Prometheus Exporter
The KRaft Monitor metrics are defined in the documentation, with an MBean name, such as kafka.server:type=raft-metrics,name=current-state. Using a JMX client, such as jmxterm, shows the MBean name is just kafka.server:type=raft-metrics with each metric an attribute in that bean. This differed from the documentation at the time of writing.
If you deploy Java applications with JMX metrics in containers, we highly recommend jmxterm.
java -jar jmxterm.jarIn the cp containers, the Java process is process 1, but use the command jvms to see all available processes and verify the JVM’s process ID is indeed 1.
$>open 1Show all the MBeans in the JVM, with beans.
$>beans...kafka.server:type=raft-metrics...Select a bean and use info to explore the attributes on a bean and get to fetch the current value of an attribute.
$>bean kafka.server:type=raft-metrics#bean is set to kafka.server:type=raft-metrics$>info# attributes%0 - append-records-rate (double, r)%1 - commit-latency-avg (double, r)%2 - commit-latency-max (double, r)%3 - current-epoch (double, r)%4 - current-leader (double, r)%5 - current-state (double, r)%6 - current-vote (double, r)%7 - election-latency-avg (double, r)%8 - election-latency-max (double, r)%9 - fetch-records-rate (double, r)%10 - high-watermark (double, r)%11 - log-end-epoch (double, r)%12 - log-end-offset (double, r)%13 - number-unknown-voter-connections (double, r)%14 - poll-idle-ratio-avg (double, r)$>get current-leadercurrent-leader = 10.0;Using the above, the following properly exposes these from JMX Prometheus Exporter. Since the current-state attribute value is a string, its value needs to be associated with a label to capture it in Prometheus.
rules:- pattern: "kafka.server<type=raft-metrics><>(current-state): (.+)" name: kafka_server_raft_metrics labels: name: $1 state: $2 value: 1- pattern: "kafka.server<type=raft-metrics><>(.+): (.+)" name: kafka_server_raft_metrics labels: name: $1Grafana Dashboard
Grafana is great for custom configuration, but that means time and effort are needed to build dashboards. Here is a dashboard around some of those KRaft metrics; it isn't a complete dashboard.
Node Information
With metrics emitting, put them into a Grafana dashboard.
Using current-state, we can see which node is the leader, in addition to capturing node.id. In addition, the dashboard component joins in other data to provide additional information on each node.

Node Counts
Counts of nodes and active controllers are always reassuring, and this uses an existing metric, kafka_controller_kafkacontroller_value{name="ActiveControllerCount",}. This metric is only emitted from a controller, so by counting the existence of this metric you see the number of controllers in the cluster, and by summing the value of the metric you get the actual number of active controllers; alert if this is ever not equal to 1.
Check out the dashboard to see how the other values are calculated, as it is the same as in the ZooKeeper-based installations.

Active Controller
To get the node.id of the quorum leader, just find max(kafka_server_raft_metrics{name="current-leader",}). Because scraping of each node is from slightly different times, different values are possible at the time of a change; max is used to make the single value display easy to build.
If a new leader is being elected, it shows up in the voted-leader metric. In a single-value panel this isn't very useful, but it is worth recording as a time series.

Full Dashboard
Here is an example dashboard that also captures the fetch rate of the controller metadata.

Open-Source Tools
At the time of writing, we checked a variety of open-source tools and could not see KRaft metrics in any of them; active-controller information displayed incorrectly. None of the tools we tried supported KRaft yet, but it is important that before you upgrade, you have a proper monitoring and alerting strategy in place.
Before you move to KRaft
Be sure you properly test your monitoring and Apache Kafka support infrastructure as part of your move to KRaft. Also, validate that kafka-raft metrics are captured and confirm that dashboards and tools work when brokers and controllers are on separate nodes.
Working on something like this?
Start a Conversation