Guide

OpenRouter Token Usage Rankings Explained

How to read OpenRouter public model rankings and pricing data without confusing router volume for global model usage.

Updated 2026-05-12openrouter / pricing / open-models
Desk note

OpenRouter-style public rankings are useful because they are visible and model-specific. The risk is scope: a router ranking is not a claim about global usage unless the source explicitly says so.

What the data can show

Public router rankings can show which models are popular on that routing surface and how usage shifts alongside pricing, context, latency, and availability. That is useful directional signal.

  • Use it to compare model momentum on the router.
  • Use it to notice changes worth investigating.

ReceiptsOpenRouter rankingsOpenRouter apps

On this siteOur live leaderboard built on this data

What the data cannot prove

Router rankings are not global model usage unless the source explicitly makes that claim. Treat them as a public slice, not the whole market, and keep the scope visible anywhere the ranking is reused in a model, cost, or adoption argument.

  • Avoid phrases like total global token burn unless sourced.
  • Keep source URL and checked date next to derived claims.

Why pricing matters

A model can be popular because it is cheap enough, fast enough, available through a preferred API, or good enough for a specific workload. Ranking without pricing context can mislead.

  • Compare input and output price separately.
  • Look at context window and provider availability together.

Best use

Use the rankings to compare model cost and router momentum, then validate model choice against your own quality, latency, and acceptance metrics. Public rankings are a map, not your destination.

  • Let public rankings suggest tests.
  • Let your evals and traces decide production routes.
Weekly briefing

The term is moving faster than the definition.

Tokenmaxxing keeps shifting as new receipts land. The weekly briefing tracks who's burning what, and why it matters.

Written by the desk's AI, human-reviewed before send, real numbers only.

Source trail

Current feed records connected to this guide

TechNode source artwork
newsT
news

DeepSeek V4 Flash tops OpenRouter weekly ranking with 7.22 trillion tokens · TechNode

DeepSeek V4 Flash led OpenRouter's July 27 to Aug. 2 usage ranking with 7.22 trillion tokens. Chinese models held all four leading slots, and V4 Flash 0731 plus V4 Pro landed inside the top six.

tokenmaxxingmodel-routerpricing
Read note
SaaStrAI source artwork
long-formS
long-form

20VC x SaaStr This Week : Apple Sues OpenAI, the Token-Maxing Era Begins, and the TAM Question Hanging Over AI Coding

On the 20VC/SaaStr podcast, Jason Lemkin, Harry Stebbings and Rory O'Driscoll call the budget era: ClickHouse has grown its AI bill 60-fold since February, while the best engineers keep 10 to 20 agents busy overnight.

tokenmaxxingai-spendmodel-routing
Read note
Generated Tokenmaxxing editorial thumbnail for Share Of US Models Being Used On OpenRouter Has Collapsed From 70% To 30% Over The Past Year
newsO
news

Share Of US Models Being Used On OpenRouter Has Collapsed From 70% To 30% Over The Past Year

A Bloomberg chart drawing on OpenRouter and Exponential View data puts the combined Google, OpenAI and Anthropic share of tokens routed through the platform near 70% in June 2025 and roughly 30% a year later.

tokenmaxxingmodel-routingscoreboards
Read note
Project layer

Tools that make the guide operational

#1Direct
Routing

LiteLLM

BerriAI/litellm

An OpenAI-compatible gateway and SDK for calling many model providers with budgets, logging, load balancing, guardrails, and cost tracking.

55.8K10.4KSource-available
gatewaycost-trackingrouting
#2Direct
Observability

Langfuse

langfuse/langfuse

Open-source LLM engineering platform for observability, traces, metrics, evals, prompt management, datasets, and playground workflows.

32.7K3.5KSource-available
tracesevalscosts
#5Direct
Evaluation

promptfoo

promptfoo/promptfoo

A CLI and CI workflow for testing prompts, agents, and RAG systems across models, with evals and red-team style checks.

24K2.2KMIT
prompt-evalscirag
Briefing

Fresh source notes each week.

New tokenmaxxing links, model-router signals, agent usage research, and AI cost notes.