LLM Inference

Session LLM Inference

Scaling LLM-Based Ranking Systems for Latency-Critical Search & Recommendation Workloads

Tuesday Jun 2 / 10:20AM EDT

Large Language Models are powerful — but deploying them in latency-critical ranking systems is a fundamentally different problem than building chat applications.

Speaker image - Sundara Raman Ramachandran

Sundara Raman Ramachandran

Lead Engineer @LinkedIn on LLM Inference Team, Previously Worked on Azure Identity & Authorization and Microsoft Office