AI inference is becoming one of the largest cost centers in technology, but the software stack is not keeping pace with the growing diversity of silicon underneath it. NVIDIA built an extraordinary ecosystem around CUDA. Now AMD, Google TPUs, AWS Trainium, Cerebras, and a growing set of custom accelerators are expanding the number of viable ways to serve AI. The challenge is turning that theoretical choice into production economics.
There is no universally best chip. The right answer depends on the model, traffic pattern, latency target, memory requirements, and serving architecture. Raw hardware specifications are not enough: realizing the economics of a new accelerator still requires scarce engineers to tune kernels and serving systems by hand.
That gap leaves capable silicon underused and AI companies paying more than they should. It also creates the opportunity for a new control layer: software that can continuously discover and implement the best way to run a workload on the best available hardware.
This is the opportunity Wafer was built to capture. We are proud to share that Wing participated in Wafer's $40 million Series A, co-led by our friends at Marathon and Chemistry, with participation from AMD Ventures and Outset Capital. Existing investors Fifty Years and Y Combinator also participated in the round.
AI that optimizes AI infrastructure
At its simplest, Wafer automatically finds and implements the highest-performance, lowest-cost way to serve an AI workload on a given chip.
Underneath that simple idea is a difficult systems problem. Wafer's agents learn a workload's traffic shape and performance constraints, then search across the model, inference engine, kernels, and hardware to find the best production deployment. The system tunes batching, caching, quantization, speculative decoding, memory management, scheduling, and more than 100 serving parameters, while generating custom kernels where the default stack leaves performance behind.
Most inference providers offer a generic deployment, or reserve deep optimization work for their largest customers because manual tuning can take days or weeks. Wafer compresses that cycle into hours. Its team combines unusual depth in GPU kernel engineering, distributed systems, memory management, and scheduling with a full-stack understanding of how models behave under real production traffic.
AMD is the clearest early proof point. In a recent Kimi K3 benchmark, Wafer served the 2.8-trillion-parameter mixture-of-experts model on a single eight-GPU AMD MI355X node at 952 tokens per second aggregate and 118 tokens per second single-stream. B300 was faster in absolute terms at 1,568 tokens per second aggregate and 172 tokens per second single-stream, but under the rental-price assumptions used in Wafer's test, MI355X delivered 48 tokens per second per dollar versus 33 on B300.
The result illustrates why workload-specific optimization matters. Kimi K3 requires more than 1.5 TB of memory for weights alone. The larger HBM footprint of MI355X allowed Wafer to keep the model on one eight-GPU node, while the B200 comparison required 16 GPUs across two nodes. Wafer also resolved two ROCm-path issues without writing a custom kernel. The point is not that AMD always beats NVIDIA. It is that the economic winner can change with model architecture, memory capacity, topology, software, and hardware price - and discovering that optimum is increasingly a software problem.
A technical team moving at commercial speed
Wafer's technical achievement initially caught our attention. The pace of execution made the company impossible to ignore. In roughly three months, Wafer grew to approximately $8 million in annualized revenue. Customers have trusted the platform with production workloads at substantial scale, and the company has won business by delivering the metrics that matter most in inference: throughput, latency, reliability, and cost per token.
That traction is especially striking given the size and age of the team. Co-founders Emilio Andere and Steven Arellano bring experience spanning mathematics, machine learning research, GPU performance, and large-scale AI infrastructure. Around them is a compact group with backgrounds in AWS datacenter software, Stanford computer science, and production inference. They understand the full serving system rather than any single layer of it, and they have shown the persistence to rebuild quickly after their original product did not find traction.
Our relationship began in a way that feels fitting for this team. Emilio saw an analysis I had written about the strategic value of a multi-accelerator architecture. He shared a thoughtful response, agreeing with the core thesis while sharpening the technical and economic assumptions. That exchange led to a meeting at my favorite Turkish coffee shop in San Francisco. The conversation quickly moved from silicon strategy to the kernel-level work required to make it real.
The multi-silicon inference cloud
Wing has spent years studying the infrastructure beneath AI, from data systems to inference and the physical compute stack. Our view is that the next generation of AI companies will treat silicon as a portfolio. Each model and workload can be routed to the hardware offering the best combination of performance, availability, memory, power efficiency, and cost.
The obstacle is software. Every accelerator arrives with a different ecosystem, and many promising chips remain economically stranded because the tooling around them is immature. Wafer is furthest along in turning the multi-silicon thesis into a production service. AMD is the starting point. The same agentic optimization approach can extend to Cerebras, TPUs, Trainium, and emerging accelerators, allowing hardware that others cannot use efficiently to become viable inference capacity.
Serving production workloads creates a potential compounding loop. More traffic produces more verified optimization traces across models, traffic patterns, and hardware. Those traces can improve the agents' ability to search the enormous serving configuration space. Better optimization lowers cost and improves performance, which can attract more workloads and generate more evidence about what works. Over time, that optimization history can become increasingly difficult to reproduce without running the workloads themselves.
As inference scales, the objective is not simply to maximize tokens per second on the fastest chip. It is to maximize useful intelligence per dollar - and increasingly per watt - across a heterogeneous pool of compute.
We believe the winning inference platform will make hardware choice dynamic, workload-specific, and increasingly autonomous. Wafer has the systems expertise, early performance leadership, and exceptional commercial momentum to build it. We are thrilled to partner with Emilio, Steven, and the entire Wafer team as they work toward a world where no viable compute goes to waste.
I write more about the infrastructure powering AI at Data Gravity.




