Open models need infrastructure that makes them practical to run.
Excited to be working on lower-cost, high-performance Kimi K3 serving on AMD and contributing the work upstream to SGLang.
We’re excited to share that AMD has provided Zro with evaluation compute to accelerate Kimi K3 inference serving on AMD MI325X (CDNA3).
Our goal is to push the performance of frontier-scale LLM serving on AMD hardware, and upstream the work so the broader ecosystem can benefit.
Today we added @AMD MI325X as a new serving backend for @zroai_. Kimi K3 is the first model running on it, powered by @sgl_project.
We believe this is the first production support for the official Kimi K3 weights on CDNA3.
With 256 GB of HBM per GPU, the full model fits on a
We believe we’re one of the first non-Day 0 serving platforms to support Kimi K3.
So we cut prices 🤠
We’ve built a full system around this amazing model. Proud of the team! and shout-out to @sgl_project and @AIatAMD !!
Try it: zro.moonmath.ai
Please help us get