The SysSoft IoT analytics pipeline processes telemetry from 5,000+ connected devices at 15,000+ records per minute. Apache Kafka is the backbone.
What the documentation does not emphasise: partition count is a scaling decision made at topic creation time. Adding partitions later is possible but disruptive. We partitioned aggressively at creation — 48 partitions for the primary telemetry topic — based on projected throughput at 2× our expected peak.
Consumer group management: each downstream consumer (anomaly detection, database persistence, dashboard aggregation) is a separate consumer group. Each consumer group maintains its own offset, so a slow anomaly detection consumer does not delay database persistence.
The operational lesson: Kafka's consumer lag metric is the most important metric to monitor. Lag growing on any consumer group means a consumer is not keeping up with producer throughput. Alert on it before it becomes an incident.
Kafka is excellent infrastructure. Understand its failure modes before you depend on it.
— Dick Bassey | DevDick | 2025