LLM Inference
Session
LLM Inference
Scaling LLM-Based Ranking Systems for Latency-Critical Search & Recommendation Workloads
Tuesday Jun 2 / 10:20AM EDT
Large Language Models are powerful — but deploying them in latency-critical ranking systems is a fundamentally different problem than building chat applications.
Sundara Raman Ramachandran
Lead Engineer @LinkedIn on LLM Inference Team, Previously Worked on Azure Identity & Authorization and Microsoft Office