Skip to content
Back to Insights
Data EngineeringBy KE Engineering Team

Bite Size Streams: Stream-to-Table Join

Bite Size Streams: Stream-to-Table JoinDATA ENGINEERING cover for Bite Size Streams: Stream-to-Table JoinJOINstreamtableDATA ENGINEERINGBite Size Streams:Stream-to-Table Join// STATEFUL JOIN · KAFKA STREAMS

First written February 2026, last updated September 2026.

A stream-to-table join doesn't always join the record you'd expect. Here are three timing scenarios, including one that surprised us.

The Demo Application

The demo emits events from local processes and windows. This mirrors IoT patterns where device IDs repeat.

Events include an iteration sequence number to show which records join.

The demo app inspects the Kafka Streams topology and builds a browser UI with a panel per topic (the screenshots below), so you can watch records move through each step. You can build your own from the tutorials.

The source code is available on GitHub.

Stream-To-Table Join Algorithm

This tutorial defines the topology in StreamToTableJoin.

Steps

  • toTable - processes are already keyed by processId, so Kafka Streams creates the table.
  • selectKey - windows are keyed by windowId, so they can't be joined to the process table; they must be rekeyed to the processId.
  • join - a join following a selectKey creates a -repartition topic and separate sub-topologies; this allows for co-partitioning. The topology shows this.
  • asString - this function has the stream element on the left-hand-side and the table element on the right-hand-side; the output can be any Java object, just make sure it can be properly serialized. In this topology, the actual join is performed by the underlying code of the Kafka Streams DSL. However, you can write topologies where you interact with state stores directly.

Source Code

java
  protected void build(StreamsBuilder builder) {     KTable<String, OSProcess> processes = builder            .<String, OSProcess>stream(Constants.PROCESSES)            .toTable(Materialized.as("processes-store"));     builder.<String, OSWindow>stream(Constants.WINDOWS)            .selectKey((k, v) -> "" + v.processId())            .join(processes, StreamToTableJoin::asString)            .to(OUTPUT_TOPIC, Produced.with(null, Serdes.String()));  }   private static String asString(OSWindow w, OSProcess p) {    return String.format("pId=%d(%d), wId=%d(%d) %s",            p.processId(),            p.iteration(),            w.windowId(),            w.iteration(),            rectangleToString(w)    );  }   private static String rectangleToString(OSWindow w) {    return String.format("@%d,%d+%dx%d", w.x(), w.y(), w.width(), w.height());  }

Stream-To-Table Join Topology

The topology shows how selectKey and join create a separate sub-topology through repartitioning, enabling co-partitioning by processId.

Kafka Streams topology for the stream to table joinSub-topology 1 reads the windows topic, re-keys it, and writes to a repartition topic. Sub-topology 0 reads that repartition topic and the processes topic, joins the stream against a table backed by processes-store, and writes to the output topic. The store is backed by a changelog topic.TOPOLOGYwindowssub-topology: 1windows-sourcewindows-selectKeywindow-to-process-joiner-repartition-filterwindow-to-process-joiner-repartition-sinkwindow-to-process-joiner-repartitionprocessessub-topology: 0window-to-process-joiner-repartition-sourceprocesses-sourcewindow-to-process-joinerprocesses-toTableoutput-sinkprocesses-storestream-to-table-join-outputs-to-t-join-processes-store-changelog// TWO SUB-TOPOLOGIES JOINED THROUGH A REPARTITION TOPIC
Fig. 1: Topology

Demonstration

Here are examples of the stream-to-table joins in action.

Happy Path

The system produces events as expected.

  • window events are emitted, iteration=N
  • process events are emitted, iteration=N
  • window events are repartitioned to the processId key
  • the join occurs with events that are emitted at the same time (iteration)

This example shows one process with two windows creating two enriched records on the output topic. Notice how the partition changes from 1 to 0 as co-partitioning aligns the events.

Demo UI showing one process and two windows joined in the same iteration, with the partition changing from 1 to 0
Fig. 2: Happy path scenario

The output dialog shows (iteration) for each of the process and window events to provide a clear indication of which events are joined.

Slightly Late-Arriving Processes

This result surprised us. With window events published (and flushed) prior to publishing process events, one might expect window events to join with process events from the previous publishing, but that isn't the case.

  • the app first emitted window events
  • then the publisher flushed buffers, ensuring the Streams application could process those events before the process events arrived
  • the app then emitted process events
  • the application joined process and window events
  • the application joined the window events with the slightly delayed process events from the same iteration

The image highlights the repartition topic’s publishing timestamp, and then shows the result of the join still with the process events from the same iteration.

Demo UI showing window events joined to process events from the same iteration despite the window events being published first
Fig. 3: Slightly late scenario

For time semantics to work, Kafka Streams doesn't (by default) change the event’s timestamp. The demo project adds a producer interceptor that adds a header with the publishing timestamp, which the UI displays (the pop-over highlighted on the timestamp of the window events). In this example, the window event arrives 0.001 seconds after the process event. The window event has the earlier timestamp, but the process event reaches the table first, so the join picks it up.

The rule: a regular table updates in the order records arrive, not in timestamp order, and the repartition hop makes the window events arrive late. So the window stream events join with the later-arriving process events.

Later-Arriving Processes

Another scenario occurs where process events arrive after the window events, this time with a slightly longer delay.

  • window events were emitted
  • publisher flushed buffers
  • slight delay occurred
  • process events were emitted
  • joins occurred with window events and earlier process events

This image highlights the expected behavior. Here, the join occurs with the previous iteration of the process event, since the repartitioned window event arrives prior to the process event.

Demo UI showing window events joined to the previous iteration of process events after a longer delay
Fig. 4: Slightly later scenario

This scenario represents the more expected behavior, where window events with early timestamps join with the previous process events that are already processed. Notice how the timestamp of the repartitioned topic is before the timestamp of the process topic.

When this matters

For most stream-to-table joins, this behavior isn't a concern. The idea is to enrich data with fairly static data, and if that data is slightly more current, it isn't a deal-breaker. However, if your use case can't tolerate this behavior, explore versioned key-value stores, introduced in Kafka 3.5.

Next up: versioned tables, which fix this.

Working on something like this?

Start a Conversation