Apache Kafka vs AWS SQS: Choosing for Production

D
Dick Edidiong Bassey
·

Apache Kafka and AWS SQS both solve the decoupled messaging problem, but with different trade-offs that make one clearly better for specific workloads.

Kafka wins when: you need message replay (consumers can re-read historical messages), you have multiple independent consumer groups reading the same stream, you need message ordering within a partition, or your throughput requirements exceed SQS's effective practical limits.

SQS wins when: you want zero infrastructure to manage (SQS is fully managed), you need dead-letter queue handling with minimal configuration, your messages are short-lived and replay is not needed, and your team does not have Kafka operational expertise.

The SysSoft IoT pipeline uses Kafka: message replay is critical (when a downstream service recovers from an outage, it must be able to re-process missed events), and multiple consumer groups are required.

Internal async jobs in SellTrove use SQS: no replay needed, the team does not want to operate a Kafka cluster, and SQS's at-least-once delivery with visibility timeout is sufficient for job queue semantics.

— Dick Bassey | DevDick | 2024