Skip to content
Back to Insights
Data EngineeringBy KE Engineering Team

Kafka Toolage: Grafana

Kafka Toolage: GrafanaDATA ENGINEERING cover for Kafka Toolage: GrafanacpumemioDATA ENGINEERINGKafka Toolage: Grafana// DASHBOARDS · METRICS · GRAFANA

First written September 2022, last updated September 2026.

Introduction

Using Grafana to monitor an Apache Kafka cluster takes a fair amount of infrastructure, and it has to fit your cluster's security configuration. While there are plenty of examples out there, you can spend a lot of time adjusting dashboards to get a desired setup.

In short

  • JMX metrics are a great way to see the health of an Apache Kafka Cluster.
  • Grafana Dashboards are a great visualization of Kafka components, including applications built with the kafka-clients or kafka-streams Java libraries.
  • A complete example of using Grafana against 4 differently configured clusters is available in kafka-toolage repository.
  • The coupling of JMX Prometheus Exporter configuration, Prometheus scraping, and Grafana queries makes it harder to combine dashboards from others.
  • The dashboards here are based on our experience, with a lot of inspiration from Confluent’s GitHub repository (Apache 2.0 License).

Important Considerations

  • Kafka Monitoring by Grafana isn't a turnkey solution; expect an investment of time from your organization.
  • Grafana’s license changed to AGPLv3, from Apache License 2.0, in April 2021. Review it and make sure this works for your organization.
  • Using JMX Prometheus Exporter exposes broker metrics to a RESTful endpoint; enable security if that endpoint is accessible outside the Kafka tools network.
  • Scraping JMX metrics through the endpoint can impact performance; tune accordingly by only exporting specific metrics and reducing the frequency Prometheus scrapes them.

Infrastructure Needed

  • JMX Prometheus Exporter Agent attached to every monitored JVM
  • JMX Prometheus Exporter Configuration for each service
  • ZooKeeper
  • Brokers
  • Connect
  • Schema Registry
  • Clients (if desired)
  • Prometheus
  • Configuration for each JVM to pull in the data from the agents
  • Grafana
  • Prometheus Integration
  • Dashboards
  • Variables to allow for supporting multiple clusters

It's a long list. Take it one piece at a time.

Software

This is the software and the versions used at the time of this writing.

ComponentContainerLicenseLatest Version
Prometheus JMX ExporterN/AApache License 2.00.16.1¹ (2021-07-14)
PrometheusDockerApache License 2.02.38.0 (2022-08-16)
GrafanaDockerGNU AGPL v3.09.1.4 (2022-09-09)

¹ At the time of writing, the latest version was 0.17.0 (2022-05-23), but it wasn't working because of a dependency issue resulting in an unknown field value. We saw no technical or security reason to be on 0.17.0.

text
noauth-broker-1   | Exception in thread "main" java.lang.reflect.InvocationTargetExceptionnoauth-broker-1   | 	...noauth-broker-1   | 	at java.instrument/sun.instrument.InstrumentationImpl.loadClassAndCallPremain(InstrumentationImpl.java:525)noauth-broker-1   | Caused by: java.lang.NoSuchFieldError: UNKNOWNnoauth-broker-1   | 	at io.prometheus.jmx.JmxCollector$Rule.<init>(JmxCollector.java:57)

The Challenges

  • The Grafana project doesn't provide Kafka dashboards. There are various ones out there, and their configuration is coupled to the configuration of JMX Prometheus Exporter as well as how that is integrated into Prometheus.
  • Typically, clients already use Grafana for other dashboards and want to customize the experience, but that means the organization takes on the maintenance and configuration of the dashboards, which takes time.
  • Some metrics that belong in Grafana dashboards can't be obtained from JMX; this disconnect can lead to challenges.
  • Monitoring clients (producers, consumers, and stream applications) requires exporting and scraping their metrics. While the repository provides Kafka client dashboards, it isn't reviewed here.
  • Also, JMX monitoring of clients is only available in the Java client. JMX is a Java monitoring API and not implemented by librdkafka or other non-Java client libraries.

Setup

Adding Grafana to your Apache Kafka tooling has many steps. Some of these pieces require working with your organization to make sure you comply with any security considerations. While exposing JMX Metrics is a read-only operation, that doesn’t mean security isn't an issue.

Since setup is an important part of considering a tool, we'll walk through various parts of the setup process. Use the example repository for complete details.

JMX Prometheus Exporter

If you are new to JMX Metrics and the JMX Prometheus Exporter, start with a simple configuration to expose all metrics. Eventually, limit the metrics to only those needed to power the dashboards.

Configuration

This configuration exposes all.

config.yml - everything

yaml
lowercaseOutputName: truerules:- pattern: .*

Integration

Add the JVM agent to each component. For Apache Kafka components, add this to your KAFKA_OPTS environment variable.

shell
-javaagent:/jmx_prometheus_javaagent.jar=7071:/config.yml

Validation

Testing is simple.

shell
% curl http://localhost:7071

Adjustment

Trim the configuration to only the metrics you need, replacing .* with specific patterns. This reduces the burden on the services for providing these metrics.

yaml
- pattern: kafka.server<type=BrokerTopicMetrics, name=BytesInPerSec, topic=(.+)><>OneMinuteRate

config.yml - specific rules

yaml
lowercaseOutputName: truerules:  - pattern: kafka.server<type=BrokerTopicMetrics, name=BytesInPerSec, topic=(.+)><>OneMinuteRate  - pattern: kafka.network<type=RequestMetrics, name=(.+), request=(.+)><>999thPercentile  - pattern: kafka.network<type=RequestMetrics, name=(.+), request=(.+)><>99thPercentile  - pattern: kafka.network<type=RequestMetrics, name=(.+), request=(.+)><>95thPercentile

Some JMX metrics can't be scraped without mapping parameters to labels. Informational metrics, such as version strings, would need a mapping to expose the attribute/value as labels and provide a dummy value (Grafana expects numerical metric values). This is an example of accessing the broker version from an app-info metric.

yaml
- pattern: "kafka.server<type=app-info, id=(.+)><>(.*): (.*)"  name: kafka_server_app_info_$2  labels:    kafka_broker: "$1"    kafka_version: "$3"  value: 1

Our recommendation: only rewrite patterns when absolutely necessary.

Prometheus

Prometheus is a time-series database used to capture the scraped metrics. It has to be configured to scrape the data and store it. Getting Prometheus’ configuration right is the key to supporting multiple types of clusters and multiple instances. For this project, a cluster_type label is added to separate out statistics from brokers, connect, and schema-registry. A cluster_id is added to identify each cluster.

Configuration

To pull the JMX Prometheus Exporter endpoint into Prometheus, you need to set up a job as a scraped configuration. Prometheus provides many service-discovery mechanisms. The mechanism you use depends on the infrastructure of your organization. For this setup static_configs is used. To add client metrics, we would use file_sd_configs for demonstration and kubernetes_sd_config for applications deployed within Kubernetes.

Here is the configuration for the noauth brokers. The additional labels make it easier to separate metrics within Grafana. The relabel_configs allows for the removal of the port from the target’s name.

prometheus.yml

yaml
global:scrape_interval: 15s scrape_configs:  - job_name: noauth-kafka    static_configs:      - targets:          - noauth-broker-1:7071          - noauth-broker-2:7071          - noauth-broker-3:7071        labels:          cluster_type: "kafka"          cluster_id: "noauth"    relabel_configs:      - source_labels: [__address__]        regex: '(.*):(.*)'        target_label: instance        replacement: '$1'

Validation

Prometheus provides a nice web interface for inspecting its database. With Prometheus’ built-in autocomplete, you can start typing a metric name and see current metrics that match.

This Prometheus UI is the best way to validate JMX metrics gathering and configuration. Checking the JMX Prometheus Exporter endpoint works, but it won't show the additional labels added through Prometheus, so use the Prometheus UI. There are additional inspections available to confirm your jobs are configured as expected.

Prometheus web UI with the metric name autocomplete list open while typing a Kafka metric
Fig. 1: Prometheus' autocomplete

NOTE

Storage is based on retention and sampling. The default retention is 15 days. To change the retention time, use the command line parameter --storage.tsdb.retention.time.

properties
needed_disk_space = retention_time_seconds * ingested_samples_per_second * bytes_per_sample

Grafana

The setup of Grafana isn't simple, but if the pieces are in place and validated, then it gets easier. Assuming you have Prometheus integrated and variables configured, you can quickly add in a query. If you are like us, however, you will experiment with various panels and tweak their display and size.

Grafana Kafka overview dashboard with cluster and topic panels and a cluster selector variable
Fig. 2: Grafana Dashboard

Configuration

The heart of the configuration is writing a query that aligns with what you found in Prometheus. Use the Prometheus UI to confirm you have it correct. Grafana provides documentation on writing queries. Understanding by and without is key to writing queries that split or combine measurements.

simple query

sum (    kafka_server_brokertopicmetrics_oneminuterate {        name="MessagesInPerSec",        cluster_id =~ "$cluster",        topic = "${topic}",        cluster_type="kafka"    })

more advanced query

sum without(instance,job) (    kafka_server_brokertopicmetrics_oneminuterate {        name="MessagesInPerSec",        topic!~"(_)+confluent.+",        topic!="",        cluster_id=~"${cluster}",        cluster_type="kafka"    })
Grafana panel editor showing a PromQL query against kafka_server_brokertopicmetrics with cluster and topic variables
Fig. 3: Grafana Query

Grafana queries depend on the configuration of JMX Prometheus Exporter and Prometheus.

Showcase

We are quite happy with the Grafana dashboards. The review isn't about them, but about the success in using Grafana to access and display the metrics across the multiple cluster configurations.

4 Clusters

The clusters used for analysis are from the first article in this series, Apache Kafka Monitoring and Management.

clusterlistener type - 9092listener type - 9093schema registrykafka connect
noauthplaintextsslhttphttp
saslsasl_plaintext (plain)sasl_ssl (scram)https / basic authhttps / basic auth
ssl-sslhttps / basic authhttps / basic auth
oauthsasl_plaintext (plain)sasl_ssl (oauth)n/an/a

Cluster Health

A single dashboard can be stood up to give the health of the cluster.

Grafana cluster health dashboard for one Kafka cluster with broker, partition, and throughput panels
Fig. 4: Kafka Cluster

Topic

A single dashboard can provide drill-down into the messages and bytes in and bytes out.

Grafana topic dashboard with messages in, bytes in, and bytes out for one topic
Fig. 5: Kafka Topic

Connections (Listeners)

If a cluster has multiple listeners, the information on each listener can be tracked.

Grafana dashboard showing per-listener connection counts
Fig. 6: Kafka Connections

Kafka Connect

A single dashboard for each Kafka Connect cluster, even if there are multiple clusters for a single Kafka cluster.

Grafana Kafka Connect dashboard with worker, connector, and task status panels
Fig. 7: Kafka Connect

Confluent Schema Registry

Metrics of each Schema Registry are also available.

Grafana Schema Registry dashboard with request rate and latency panels
Fig. 8: Confluent Schema Registry

To see more, check out the demonstration project kafka-toolage. All shown dashboards are available for your evaluation.

Review

Grafana falls primarily under the Monitoring classification as stated in the setup article.

Classification

We classify tools in three areas: monitoring, observation, and administration. Review of tools is based on their purpose/design. Grafana is a monitoring tool, and the review is based on the supported functionality.

functionality
monitoring✔
observation𐄂
administration𐄂

Connectivity

100% successful in connecting to every deployed component.

ClusterServiceProtocol¹result
noauthbrokerJMX (HTTP)✔
noauthschema-registryJMX (HTTP)✔
noauthconnect (a)JMX (HTTP)✔
noauthconnect (b)JMX (HTTP)✔
sslbrokerJMX (HTTP)✔
sslschema-registryJMX (HTTP)✔
sslconnectJMX (HTTP)✔
saslbrokerJMX (HTTP)✔
saslschema-registryJMX (HTTP)✔
saslconnectJMX (HTTP)✔
oauthbrokerJMX (HTTP)✔

While every cluster other than ssl has two Kafka protocols, the distinction isn't applicable for monitoring tools that use JMX.

Deployments Needed

Only a single deployment of Grafana is needed to observe all Kafka clusters and their respective Connect clusters and Schema Registries.

Ratings

These ratings are subjective. Rating Grafana is difficult: we have worked with it for four years, constantly changing queries and dashboards, and it's hard to fully set that experience aside. What we discovered: the dashboards we had already built had errors and were incomplete; they required modifications to support multiple instances.

Rating Classification: 1 difficult/rigid to 10 easy/customizable

classificationratingsummary
setup1¹coupling of dashboards to jmx exporter makes it difficult to borrow from other provided dashboards
success10was able to access all metrics for all components in all clusters
customizable9very customizable, which is great but also means more time to set up and use
completeness8full access to all the metrics, provided all clusters expose their metrics the same way
integration8free to write your own queries, dashboards, and variables
ease-of-use8once dashboards are built, easy to use

¹ When we started this review, our dashboards were already built. The variables and the ability to move between Apache Kafka clusters and Kafka Connect clusters were failing - we had to modify the dashboards to support multiple clusters. While successful, it took additional time.

Conclusion

Monitoring is a vital component of your Kafka cluster and Apache Kafka applications. For open-source options Grafana is the standard. It's unclear how the licensing change impacts enterprise adoption of Grafana.

While Grafana is a big effort to set up and configure, once configured, it is quite stable and provides a rich set of information for your development and operational teams.

Because this setup relies on metrics exposed via JMX, its use is limited to self-managed installations.

Working on something like this?

Start a Conversation