The short answer
Tiered storage, production-ready since Kafka 3.9, lets Kafka offload data to external storage such as cloud object stores. Its practical effect is to separate how long you retain data from how much broker disk you buy. Brokers still need local capacity for the recent data consumers actually read, and the remote tier is a new dependency to operate.
For most of Kafka's history, retention and broker disk were the same decision. Keep a topic for 30 days and you bought enough local disk on every broker to hold 30 days of it, replicated. Tiered storage, which became production-ready in Kafka 3.9 [1], separates those two decisions.
What it does
Tiered storage lets Kafka offload data to pluggable external storage systems, such as cloud object stores [1]. In practice the most recent data stays on broker disk while older segments move to the remote tier, and consumers reading far back in a topic are served from there.
What it changes
- Retention stops dictating broker disk. Long retention can sit in object storage rather than on every broker, so a topic kept for replay or audit no longer forces a larger cluster.
- Less local data to move. With less held on each broker, there can be less to copy when a broker is replaced or partitions are reassigned.
- New operational controls. Kafka 3.9 added a way to disable tiered storage dynamically per topic, upper bounds on upload and download rates, and visibility of the highest offset stored remotely [1].
What it does not change
- Hot data still needs local capacity. Consumers reading near the head of the log are served from broker disk, so brokers still have to be sized for the data actually in active use.
- It adds a dependency. The remote store is now part of the read path for older data, with its own latency, availability and cost profile to monitor.
- It is a per-topic decision. Topics read only near the head gain little. Topics kept for replay, reprocessing or audit gain the most.
The wider point
Many Kafka cost conversations are retention conversations in disguise. Tiered storage gives teams a way to have the retention argument on its own terms, instead of settling it with a broker purchase order.
Sources
- [1] Apache Software Foundation, Apache Kafka 3.9.0 Release Announcement. Published 6 November 2024.Primary source
Related questions
When did Kafka tiered storage become production-ready?
In Apache Kafka 3.9, announced in November 2024. Tiered storage lets Kafka offload data to pluggable external storage systems such as cloud object stores.
Does Kafka tiered storage reduce broker disk requirements?
For long retention, yes: older data can live in object storage rather than on every broker. Brokers still need enough local capacity for the recent data consumers actively read.