Skip to content
Back to Insights
Data EngineeringBy KE Engineering Team

Apache Kafka: A Practitioner's Reference

Apache Kafka: A Practitioner's ReferenceData Engineering cover for Apache Kafka: A Practitioner's ReferenceLEADERB1B2B3clusterDATA ENGINEERINGApache Kafka: APractitioner'sReference// THE PARTS THAT DO NOT GO STALE

First written August 2022, last updated September 2026.

This reference is based on our experience supporting clients running Apache Kafka and Confluent Community Software. Confluent Enterprise isn't directly covered here.

It's organized the way we teach Kafka: licensing first, because it trips teams up early, then terminology, then the material from which to learn, then the practical setup notes that save a week of debugging.

Components and licensing

It's easy to get confused about licensing across Kafka's components. Everything referenced in this document is free for developers to use, but some of it is under the Confluent Community License rather than Apache 2.0. That distinction matters: a cloud hosting service other than Confluent's can't offer Confluent Community Licensed software (version 5.1 and later) as a service.

ComponentLicense
Apache KafkaOpen Source, Apache 2.0
Kafka Streams APIOpen Source, Apache 2.0
Kafka Connect [1]Open Source, Apache 2.0
Java ClientOpen Source, Apache 2.0
librdkafkaOpen Source, Apache 2.0
Schema RegistryConfluent Community License
Schema Registry Serdes [2]Open Source, Apache 2.0
REST ProxyConfluent Community License
ksqlDB (KSQL)Confluent Community License
  1. While Kafka Connect is Apache 2.0 licensed, connectors may not be. Check the licensing of each connector independently.
  2. The serializers and deserializers live in the same GitHub Schema Registry repository but are licensed separately. See the LICENSE file within the subdirectories for confirmation.

Versions

Confluent Community releases are based on a specific Apache Kafka version, but bug fix releases are scheduled separately. Confluent may ship a bug fix release that doesn't fully align with a bug fix release from Apache Kafka. Read the release notes, especially when you are tracking down a specific issue you need fixed.

Release history and downloads for every version are on the Apache Kafka downloads page. Check there for the current version.

Kafka 3.0 was a major release with breaking changes, including client changes such as making acks=all the default for producers. If you upgrade a client library to 3.x while running against an older cluster, understand what happens in the broker versus what happens in the client. Kafka 4.0 removed ZooKeeper, and with KIP-896 older protocol versions are no longer supported.

Minimum client versions for Kafka 4.0

  • Kafka Java Client 2.1+, November 2018
  • librdkafka 1.8.2, October 2021
  • KafkaJS 1.15.0, November 2020
  • Sarama 1.29.1, June 2021
  • kafka-python 2.0.2, September 2020

Terminology and key components

If you are ramping a development or operations team on Kafka, these terms are the foundation. Knowing them makes every piece of training material easier to follow. We also recommend reading Confluent's introduction material.

We define terms in order of composition. A key is part of a message, so key comes before message. That ordering is imperfect, since some concepts are best described out of order, so consider reading this section more than once.

Key

The key is usually how a partition is selected. Custom partitioning is possible, but it can easily break stream processing that relies on multiple topics with the same partition count keeping matching keys in the same partition number. If a message has no key (it is null), the producer client library assigns a partition. To a broker, a key is just an array of bytes.

Value

The value is usually the heart of the message. Many serializers and deserializers exist to get your message into Kafka. To the brokers it is just an array of bytes.

Message

A message is the event written to Kafka: a key, a value, a timestamp, and optional headers. Treat that information as atomic and immutable.

Offset

The offset is an ever-increasing number within a given partition. Given a topic, partition, and offset, you always reference the same message in the broker. If a topic is compacted, that message may no longer be available, but the coordinates never come to mean something different.

Partition

A partition is the physical unit to which the producer writes and from which the consumer reads. A topic is configured with one or more partitions. Multiple partitions increase throughput and allow Kafka to scale.

A partition as an append-only logA partition drawn as a row of message cells numbered by offset, with new writes appended at the tail.PARTITIONpartition 0append-only, ordered, immutableMSG 0MSG 1MSG 2MSG 3MSG 4MSG 5MSG 6NEXToffset 0offset 1offset 2offset 3offset 4offset 5offset 6offset 7writes append hereoffsets never change; a topic+partition+offset always names the same message// NEW MESSAGES ONLY EVER GO AT THE END
Fig. 1: A partition is an append-only log addressed by offset.

Topic

A topic is the logical unit to which a producer writes and from which a consumer reads. It's a set of partitions and the messages in them, where the key (when provided) keeps messages of the same key in the same partition, achieving loose ordering.

A topic as a set of partitionsA topic box containing three partitions, with messages of the same key routed to the same partition.TOPICPRODUCERwritestopic orderspartition 0k:ak:dk:ak:apartition 1k:bk:bk:epartition 2k:ck:fsame key, same partition: ordering is preserved per key, not across the topic// THE KEY DECIDES THE PARTITION
Fig. 2: A topic is a set of partitions; the key pins ordering per key.

Replication factor

The number of copies made of a topic for reliability and durability. Replication factor is configured at the topic level but implemented at the partition level. The number of replicas is limited by the number of brokers, since a replica is never placed on the same broker as another replica of the same partition.

Replication factor 3 across three brokersThree brokers each holding replicas of three partitions, with a different leader replica on each broker.REPLICATION FACTOR 3broker 1P0 LEADERP1 FOLLOWERP2 FOLLOWERbroker 2P0 FOLLOWERP1 LEADERP2 FOLLOWERbroker 3P0 FOLLOWERP1 FOLLOWERP2 LEADERno two replicas of a partition share a broker; replica count is capped by broker count// LEADERS SPREAD ACROSS BROKERS
Fig. 3: Replication factor 3 across three brokers, one leader per partition.

Broker

The broker is the heart of Kafka. Multiple brokers exist so that work is distributed for performance and replicated for reliability. Each partition of a topic can have a different broker as its leader.

Producer

The producer publishes messages to the broker. The logic for knowing which broker to talk to for a given message is handled by the client, which communicates directly with the leader for that partition.

The producer routes each message to the leader of its partitionA producer client with a partitioner sending messages directly to the leader broker for each partition.PRODUCERPRODUCERclient libraryPARTITIONERhash(key)BROKER 1 - LEADER P0BROKER 2 - LEADER P1BROKER 3 - LEADER P2the client talks directly to the leader broker for each partition; no proxy in the middle// THE KEY'S HASH PICKS THE PARTITION
Fig. 4: The producer routes each message to the leader of its partition.

Consumer group

A set of consumers, identified by a shared group.id, that split a topic's partitions between them.

Consumer

The consumer receives the message. Consumers typically read from a partition on the broker that leads it, though this is configurable. Consumers work together within a consumer group to distribute work. Partitions aren't split between consumers, so the maximum number of active consumers is limited by the partition count of the topic.

Consumer group assignmentThree partitions assigned across two consumers in a single consumer group.CONSUMER GROUPtopic ordersPARTITION 0PARTITION 1PARTITION 2consumer group billingCONSUMER ACONSUMER Bpartitions are never split across consumers, so active consumers are capped by partition count// MORE CONSUMERS THAN PARTITIONS SIT IDLE
Fig. 5: Consumer group assignment; partitions are never split.

Training material

There is a lot of material available for learning Kafka. What follows is what we have used ourselves, including Confluent on-demand development training, virtual instructor-led operations training, and the Udemy courses referenced below.

Confluent

Confluent offers instructor-led and self-paced training. See confluent.io/training for details. The instructor-led course is three days, and its value is the ability to ask in-depth questions throughout. The self-paced development course covers the same material on your own schedule.

If you go down this path, we recommend the following:

  • Have developers take the same course so discussions continue between sessions.
  • Make sure the team is committed. Watching the videos without doing the coursework wastes the investment.
  • Ask the instructor questions. If they don't know an answer, they will usually chase it down during a break.

Tutorials, examples, and community

  • Confluent Developer pages and tutorials.
  • The Confluent Platform demo and the Confluent GitHub examples.
  • The confluentcommunity.slack.com Slack workspace, best for interactive conversation.
  • The Confluent community forum, which exists to make questions and answers searchable in a way Slack isn't.

Recommended Udemy courses

One instructor's Kafka series on Udemy is consistently strong. Wait for a sale, and note that buying one course usually enables a discount code for the others. The courses worth your time:

Conduktor's Kafkademy is also a good free introduction: What is Apache Kafka and the hands-on CLI tutorials.

Books

The following are excellent, and Confluent provides free copies of several of them.

Kafka: The Definitive Guide

Kafka: The Definitive Guide, 2nd Edition (O'Reilly). The first edition is also available. Pair it with the Apache Kafka documentation as your reference set.

Kafka in Action

Kafka in Action, a practical introduction from Manning.

Kafka Streams in Action, 2nd Edition

Kafka Streams in Action, 2nd Edition (Manning). Strong on detail and examples, with care around data governance. If you want a foundation in data modeling for a streaming world, read this one. Its author has contributed to Apache Kafka for years, and the attention to detail shows.

Mastering Kafka Streams and ksqlDB

Mastering Kafka Streams and ksqlDB (O'Reilly). Chapter 5 is the highlight: the windowing and time example is a strong use case and the best introduction to window processing we have found.

Designing Event-Driven Systems

Designing Event-Driven Systems (O'Reilly, free from Confluent). A must-read even if you never use Kafka. Chapter 5 on using events for notification versus using events for state transfer, and chapter 11 on the single writer principle, are both worth the time on their own.

Making Sense of Stream Processing

Making Sense of Stream Processing (O'Reilly, free from Confluent). Less about Kafka specifics, more about distributed systems, concurrency, and transactions. Read it to be a stronger architect.

Kafka Transaction Data Streaming for Dummies

Apache Kafka Transaction Data Streaming for Dummies (Wiley, free from Confluent). Change data capture is central to using Kafka with legacy systems. Chapters 3 and 4 are the useful part.

Gently Down the Stream

Gently Down the Stream, by the author of Mastering Kafka Streams and ksqlDB. Written to explain Kafka to children, and a clear introduction for anyone.

Blogs

The Confluent blog is the easiest way to keep extending your understanding once you have the basics.

Putting Apache Kafka to use

Confluent's founding overview of the stream data platform, in two parts. Part one is a quick read. Part two covers design considerations and makes the case for Avro while pointing out that consistency is what matters. The companion post on Avro is worth pairing with it.

Microservices with Kafka Streams and KSQL

Building a microservices toolkit with Kafka Streams and KSQL covers much of Designing Event-Driven Systems with working code, and is a good use case for interactive queries. The old shortcomings of interactive queries are largely resolved by KIP-535 and KIP-429. The code lives in the Kafka Streams examples repository.

Event sourcing

Operations and performance

Configure Kafka to minimize latency has the best end-to-end latency graphic available. It also makes a point worth internalizing: acks=0, acks=1, and acks=all don't change end-to-end latency. They change when the producer learns the message was accepted.

Podcasts, videos, and meetups

Confluent's Streaming Audio podcast has been on hiatus since 2023, but its back catalog is one of the few programming podcasts worth commute time. The Confluent YouTube channel carries their full video library, and Apache Kafka links the highest-rated Kafka Summit talks by classification. To browse a specific summit, search "kafka summit" at confluent.io/resources. Apache Kafka meetups run by Confluent are recorded.

Talks worth your time

Source code

Have the source downloaded and ready to explore: Apache Kafka and Confluent. It settles arguments. For example, we almost always recommend setting retries to max integer, and the fastest way to make that case is to show the default producer settings used by the Kafka Connect API, which do the same thing. The settings and comments in the source are a great reference for seeing properties configured as a logical unit.

Words of advice: Kafka is constantly evolving

Keep quick access to two versions of the Kafka documentation: the version your training material uses, and the version you have installed for development.

As you work through material, keep the Kafka version of the era in mind. Kafka used to require tools that talked to ZooKeeper directly. If you see --zookeeper <zookeeper>:2181 in documentation or a presentation, it has most likely been replaced with --bootstrap-server <broker>:9092. When the instructions and the command line disagree, compare the docs for your version against the docs from when the material was written.

Installation

There are two installations to think about: one for developers, and one for production or production-like environments.

Developer installation

Developers should have their own local installation, ideally container based. Cloud resources, managed or self-hosted, make it harder for developers to try things independently and introduce unexpected cost.

Our recommendation is to install the Confluent Community Edition on the development machine from Confluent Download, which makes the command line tools available locally, and then run the cluster itself in containers. Confluent's images work well for this, and our setup is published as dev-local: four brokers, components that start and stop independently, a clean slate with docker compose down -v, and the ability to run applications in a container or from the laptop. Local Development for Real-Time Data describes it.

Build your own toolkit using ours or someone else's as a guide. You will spend a lot of time in it, so make it yours.

Production installation

Production recommendations are out of scope here. There are too many variables. Evaluate service offerings alongside self-managed installations, and if you self-manage, plan to train or hire for Kafka operations.

Docker

Docker makes it easy to run multiple brokers on a development machine, which is why we prefer it over a local install. We configure a four-broker cluster for development so replication and direct client-to-broker communication are visible. One broker is fine for learning, but three is the minimum that shows the distributed nature of Kafka.

Kafka needs a fair number of services running. Default Docker settings are often not enough. Four CPUs, 6GB of memory, and 1GB of swap is a reasonable minimum. On a 32GB laptop, consider 8 CPUs, 12GB of memory, and 1 to 2GB of swap. A 64GB Docker disk is usually sufficient. On Apple silicon, use Confluent images 7.2.0 or later for native arm64 support; the RocksDB library also has the native updates Kafka Streams needs.

Docker container addressing

When running in containers, broker addressing and port mapping matter. Clients on the host and clients inside Docker reach the brokers differently. The fix is two advertised listeners per broker, so applications on your laptop use localhost and applications inside containers use the container hostname.

yaml
KAFKA_LISTENER_SECURITY_PROTOCOL_MAP: PLAINTEXT:PLAINTEXT,PLAINTEXT_HOST:PLAINTEXTKAFKA_ADVERTISED_LISTENERS: PLAINTEXT://broker:9092,PLAINTEXT_HOST://localhost:29092

Spring Boot and Gradle

Use the Spring Boot application builder at start.spring.io to scaffold the project: pick the build system, the Spring Boot version, and the components. For Kafka, add Kafka and Kafka Streams, and consider Cloud Stream if you want the Spring Cloud additions. You can also use the kafka-clients library directly from Spring. The Spring integration libraries aren't required, and when working through tutorials it is often better to understand Kafka without them.

If you prefer Gradle over Maven, the Avro Gradle plugin has been easier to work with for generating Java from Avro schemas. With Maven we had to manage compile order between specifications; with Gradle we didn't. Check release notes and use a plugin version compatible with your Avro version.

Security

Security is a discipline of its own with many nuances. Find examples that closely match the configuration you need. kafka-toolage showcases a variety of configurations, including Kafka with SSL client-side authentication.

Kafka Connect

Kafka Connect is a framework for building sources that bring data into Kafka and sinks that pull data out. It distributes tasks and exposes RESTful endpoints for configuring and maintaining workers. Do not overlook it.

Kafka Streams and ksqlDB

Kafka Streams is a Java library built on the consumer and producer APIs for stateful processing, with a state store providing state. A Kafka Streams application is a topology graph, usually constructed with the Streams DSL rather than the low-level APIs. If you are doing stateful processing, explore it. If you already run Spark, Flink, or another stateful engine, work with those teams to decide where each belongs.

ksqlDB is a higher-level language on top of that DSL. It's released as part of Confluent Platform but has its own release cadence, and it is licensed under the Confluent Community License. Check the repository tags for community-tagged versions.

Avro and the Schema Registry

Since Confluent 5.5 the Schema Registry supports JSON and Protobuf in addition to Avro. The Confluent open-source Avro serializer and deserializer are designed around having a schema registry, which is what allows Avro data on a topic without serializing the schema alongside every message.

The schema is to Avro as an XSD is to XML, with one difference that matters: an XML file can be parsed without its XSD, and Avro can't. Avro is strongly typed and the schema enforces it. When serializing there is a specific record and a generic record; comparing to JSON, generic Avro is like JsonNode and specific Avro is like a Java POJO with Jackson annotations.

Avro compatibility comes in forward, backward, and full modes. While iterating on a schema locally you may hit incompatible version errors. The quickest way through is to disable compatibility checks in your development registry.

bash
curl -X PUT http://localhost:8081/config -d '{"compatibility": "NONE"}' -H "Content-Type:application/json"

Do not disable compatibility checks in production.

Use GitHub search to find examples

When we are trying to figure out how to configure something, we search GitHub, selecting Code after the search. Two examples:

text
path:**/*.json org.apache.kafka.connect.transforms.RegexRouter
text
path:**/docker-compose.yml kafka confluent-hub

Apache Kafka and Confluent Community dependencies

Each release bundles specific versions of its dependencies, so check the release notes when planning an upgrade. This is helpful when you are tracking an open source security issue, want to know a library version to see if a feature is available, or want to minimize the dependency graph with your own Java software that uses kafka-client or kafka-streams libraries.

Apache Kafka downloads

Knowing the dependencies of each Kafka and Confluent Community distribution makes CVE audits easier.

Process

A script like the following extracts library versions from a distribution:

for i in $(ls -dr kafka_*.tgz); do  gzcat $i | tar tfv - | grep "/NOTICE$"  gzcat $i | tar tfv - | (echo $i; grep -E 'libs/scala-library|libs/rocksdbjni-|libs/zstd-jni|libs/lz4-java|libs/snappy-java|libs/jackson-core|libs/slf4j-api|libs/slf4j-log4j|libs/zookeeper-[23]' | cut -d/ -f2-; echo "")done
for i in $(ls -dr confluent-community-*.tar); do  gzcat $i | tar tfv - | grep "/README$"  gzcat $i | tar tfv - | (echo $i; grep -E 'kafka/scala-library|kafka/rocksdbjni-|kafka/zstd-jni|kafka/lz4-java|kafka/snappy-java|kafka/jackson-core|libs/slf4j-api|kafka/slf4j-log4j|kafka/zookeeper-[23]|schema-registry/avro|schema-registry/protobuf' | cut -d/ -f2-; echo "")done

Additional effort is spent inspecting release notes and GitHub (Apache's and Confluent's repositories). Knowing your version dependencies is important for understanding the impact of a CVE, available features, and resolving issues.

Working on something like this?

Start a Conversation