Why I Moved ML Workloads to AWS EKS

D
Dick Edidiong Bassey
·

When we started running ML inference at Sabivox, we ran models on EC2. It worked. It was also operationally expensive.

We moved to ECS. Better. Then to EKS for the ML workloads specifically: GPU instance scheduling, granular resource quotas for different memory profiles, and HPA integrating with custom metrics (inference queue depth).

EKS is not simple. The learning curve and operational overhead are non-trivial. But for production LLM inference at scale, it was the right answer.

Choose your orchestration layer based on workload characteristics, not marketing.

— Dick Bassey | DevDick | 2024