Volume leaders by total tokens processed
openai
openai
Token consumption across the gateway over the last 7 days
Models with the lowest Time to First Token (TTFT) on the gateway
Compare the most popular models on DeepToken by token volume
| Rank | Model ID | Tokens (7d)βΌ | Avg Latency | 7d Change |
|---|---|---|---|---|
| 1 | O gpt-4o-mini openai | 1.15T | 550ms | β +15% |
| 2 | O gpt-4o openai | 950B | 850ms | β +10% |
| 3 | G gemini-2.5-flash google | 620B | 450ms | β +23% |
| 4 | A claude-sonnet-4-20250514 anthropic | 620B | 850ms | NEW |
| 5 | G gemini-2.5-pro google | 410B | 1100ms | β +11% |
| 6 | A claude-opus-4-20250514 anthropic | 150B | 2400ms | NEW |
| 7 | G gemini-3.1-flash-image-2k-9x16 google | 10B | 800ms | β - |
| 8 | G gemini-3.1-flash-image-4k google | 10B | 800ms | β - |
| 9 | A claude-opus-4-6 anthropic | 10B | 800ms | β - |
| 10 | A claude-opus-4-6-thinking anthropic | 10B | 800ms | β - |
Rankings are updated every minute using a rolling 7-day request window. Tokens are counted as total processed (input + output).
Latency measures the time to completion, while TTFT represents the Time to First Token. Default seed values are blended with real-time routing metrics to bootstrap new nodes.
Updated Jul 10, 2026, 3:37 PM
View catalog