Why it matters
AlexNet marked the shift to compute-hungry deep learning that eventually produced today's LLMs — understanding its architecture is useful grounding for why modern inference costs scale the way they do.
The tokenmaxxing angle
Tracing model design from AlexNet's convolutional layers to today's transformer stacks clarifies why parameter count and architecture, not just prompt length, drive the inference-cost curve agent builders optimize against.
From the organizers
Organized by Jason M. and Katharine for PyTorch ATX at Vintage Bookstore and Wine Bar Events; Katharine covers the ImageNet dataset, Jason leads discussion of AlexNet's architecture and results.