Events / San Francisco

SF Systems: Research to Practice, from Inference Engines to Multimodal Coding

SF Systems Club talk night at LatchBio's office with two research talks: an LLM inference engine built for cheap high-volume AI-SQL queries, and a multimodal/visual programming system from UC Berkeley.

Thu, Sep 24, 5:30 PMMission Bay, San Francisco, CA

Why it matters

The first talk is a direct look at why general-purpose inference engines waste compute on repetitive, mostly-prefill LLM calls, and what a purpose-built query engine changes about that cost profile.

The tokenmaxxing angle

Very direct: Shreya Shankar's (CMU) talk on 'Quail' covers why general-purpose inference engines are inefficient for AI-SQL workloads that issue one LLM call per row/tuple, and how KV-cache-aware query planning cuts that overhead -- core inference-cost engineering.

From the organizers

Talks confirmed: Shreya Shankar (CMU) on 'Domain Specific LLM Inference for AI-SQL' presenting open-source engine 'Quail', and Parker Zeigler (UC Berkeley) on 'Interactive, Multimodal Programming Systems'; hosted by LatchBio, SF.