ONNX Runtime Cut Our Inference Latency by 35%
Converting trained ML models to ONNX format and running inference via ONNX Runtime simplified deployment and cut latency.
Articles on engineering, systems design, startups, and lessons from the trenches.
Converting trained ML models to ONNX format and running inference via ONNX Runtime simplified deployment and cut latency.
The engineers who matter are the ones who ship. Not the ones who design elegant architecture. The ones who deliver.
The architecture that handles 10,000 active storefronts without falling over is not glamorous. It is disciplined.
A postmortem is not a punishment. It is the mechanism by which incidents convert into system improvements.
The difference between mid-level and senior engineer is mostly judgment, not technical knowledge.
Most developers use CloudFront only to serve static assets. Here is how SellTrove uses it as a security layer, API gateway, and dynamic content accelerator.
The books that have most shaped how I think about engineering, architecture, and building systems that last.
They look similar. They are not. Choosing the wrong one creates problems that only appear at scale.
The culture around code review determines whether it produces better engineers or resentful ones.