Why it matters
On-device inference is the other lever besides model routing for cutting per-request cost, and this is where chipset and small-model vendors compare notes on it.
The tokenmaxxing angle
Running inference on-device via Qualcomm silicon and ZETIC's Melange optimizer sidesteps per-token API billing entirely, an edge-native way to route around cloud inference cost.
From the organizers
Confirmed speakers: ZETIC CEO Yeonseok Danny Kim, Liquid AI's Alan Wong (Head of Silicon Ecosystem Partnerships), and a Qualcomm AI Technology & Product Marketing director, at Hanwha AI Center SF.