x
Thank you! Your submission has been received and we will follow up shortly.
Oops! Something went wrong while submitting the form.
Breaking the KV Cache Inference Wall — Data Insights Podcast
Podcast
September 14, 2026

Breaking the KV Cache Inference Wall — Solidigm Data Insights Podcast, Featuring Graid Technology

"You've got a four-lane highway, and to get to your storage, you have to go down to a one-lane bridge."

That's how Garrett McKibben, Senior Director of Technical Marketing at Graid Technology, described the storage bottleneck choking modern AI infrastructure on the latest episode of TechArena's "Data Insights" podcast with Allyson Klein and Jeniece Wnorowski (Solidigm).

The Real Bottleneck

The conversation centered on KV cache, the working memory an inference model holds for every active request. Demand for that memory grows with context length and the number of concurrent users, and teams run out of room fast. McKibben put hard math on it: a Blackwell GPU tops out near 180 gigabytes of HBM at $50 to $80 per gig, and concurrency collapses after three or four users. Buying another Blackwell for every few users just doesn't pencil out.

Moving the Overflow to NVMe

Graid Technology's answer moves that overflow to an NVMe tier and accelerates the IO path on the GPU itself, using roughly 4% of the GPU before releasing those resources back. The company has published testing measuring time to first token when KV cache overflowed to fast storage instead of being recomputed, and in long-context runs, fetching beat recompute.

What's Next

The discussion also covered rack-scale design, agentic workloads that run for hours at a time, and what a team should actually measure to know whether its storage is pulling its weight.

Watch the full episode above.

Learn more about SupremeRAID™ and VROC™ by Graid Technology at graidtech.com.

Learn More

News & Resources

"You've got a four-lane highway, and to get to your storage, you have to go down to a one-lane bridge." Garrett McKibben joins Solidigm on TechArena's Data Insights podcast to unpack why KV cache overflow is choking AI inference.
Graid Technology CEO Leander Yu and VP of Product Development Alven Yen explain how moving RAID processing onto the GPU breaks the bottleneck traditional controllers cannot solve, delivering 36 million GPU-initiated IOPS for AI infrastructure.
Join Graid Technology at Dell Technologies Forum Taipei 2026 and discover how SupremeRAID™ AE transforms local NVMe storage into a high-performance, scalable, and protected extended KV cache—reducing redundant computation and accelerating time to first token (TTFT).