Apache Kafka
Apache Kafka is an event-streaming platform that stores records in partitioned logs and lets consumers process them independently. The skill designs topics, keys and consumer behavior so events can support decoupled services, replay and data pipelines with explicit ordering, retention and failure assumptions.
What it is
Producers append records to topics, which are divided into partitions. Records have an order within a partition, while no single total order is implied across all partitions. Consumers track their progress with offsets, and consumer groups distribute partitions among members for parallel processing. Retained records can be replayed without requiring producers to send them again. Replication supports availability and durability under configured conditions. Kafka is therefore more than a transient message queue, but it does not automatically make every downstream effect exactly once; the producer, processing logic and target system must participate in the relevant guarantees.
What the work involves
The practitioner chooses topic boundaries and record keys from ordering and scaling requirements. They define schemas, retention and consumer offset behavior, then test rebalances, retries and replay. Useful artifacts include an event contract and an operational plan for lag, partition changes and failed records. Consumers should handle duplicate delivery or use appropriate transactional mechanisms where supported. Security and access controls apply to producers and consumers separately, and the team verifies whether retained data can be replayed safely into systems that have already processed earlier versions.
Illustrative example
An inventory system publishes stock changes keyed by product and location. A forecasting consumer processes the retained stream while a separate alerting consumer reacts to low stock. When the forecasting service fails, it resumes from recorded offsets and rebuilds its state. The team checks that replay does not send duplicate customer notifications through another side effect, since storing and replaying events is distinct from making those notifications idempotent.
Limits and common mistakes
Ordering is scoped to partitions, and a poorly chosen key can create hot partitions or scatter related events. Retention limits constrain replay, while lag can make a working consumer operationally stale. Exactly-once claims require precise boundaries and supported integrations. Kafka adds operational complexity, so the choice should be justified by event retention, throughput or decoupling needs rather than assuming every communication between services requires a streaming platform.
Prerequisites
Related skills
- → is an instance of: Event-Driven Architecture
Sources and further reading
- Apache Kafka: introduction
Official event-streaming model, topics, partitions and consumers.
Last updated: 2026-10-10