Performance
Session
Performance
Serving LLMs at Scale: The Hidden KV Cache Advantage
Monday Jun 1 / 11:30AM EDT
KV cache is the hidden lever behind inference cost and performance. It directly impacts GPU utilization, throughput, and Time to First Token.
Khawaja Shams
Co-Founder & CEO @Momento, previously @NASA and @Amazon