Skip to content
DeepTokenInference Gateway
HomeDashboardModelsPromptsLeaderboardDocsPricingEnterpriseBlog

    Top Routable Models

    Volume leaders by total tokens processed

    2
    Second Place
    O

    gpt-4o

    openai

    Tokens (7d)950B
    Requests (7d)420M
    Avg Latency850ms
    Avg TTFT120ms
    OpenAI
    ↑ +10%
    1
    Top Model
    O

    gpt-4o-mini

    openai

    Tokens (7d)1.15T
    Requests (7d)890M
    Avg Latency550ms
    Avg TTFT90ms
    OpenAI
    ↑ +15%
    3
    Third Place
    G

    gemini-2.5-flash

    google

    Tokens (7d)620B
    Requests (7d)290M
    Avg Latency450ms
    Avg TTFT80ms
    Google
    ↑ +23%

    Top Models by Token Volume

    Token consumption across the gateway over the last 7 days

    1
    Ogpt-4o-mini
    1.15T
    2
    Ogpt-4o
    950B
    3
    Ggemini-2.5-flash
    620B
    4
    Aclaude-sonnet-4-20250514
    620B
    5
    Ggemini-2.5-pro
    410B
    6
    Aclaude-opus-4-20250514
    150B
    7
    Ggemini-3.1-flash-image-2k-9x16
    10B
    8
    Ggemini-3.1-flash-image-4k
    10B
    9
    Aclaude-opus-4-6
    10B
    10
    Aclaude-opus-4-6-thinking
    10B
    11
    Aclaude-opus-4-7
    10B
    12
    Aclaude-opus-4-8
    10B
    13
    Aclaude-haiku-4-5-20251001
    10B
    14
    Aclaude-sonnet-4-5-20250929
    10B
    15
    Aclaude-sonnet-4-5-20250929-thinking
    10B

    Fastest Models

    Models with the lowest Time to First Token (TTFT) on the gateway

    1G
    gemini-2.5-flashGoogle
    google
    Avg Latency450ms
    Throughput4751 tok/s
    TTFT80ms
    2O
    gpt-4o-miniOpenAI
    openai
    Avg Latency550ms
    Throughput2349 tok/s
    TTFT90ms
    3O
    gpt-4oOpenAI
    openai
    Avg Latency850ms
    Throughput2661 tok/s
    TTFT120ms
    4G
    gemini-3.1-flash-image-2k-9x16Google
    google
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms
    5G
    gemini-3.1-flash-image-4kGoogle
    google
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms
    6A
    claude-opus-4-6Anthropic
    anthropic
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms
    7A
    claude-opus-4-6-thinkingAnthropic
    anthropic
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms
    8A
    claude-opus-4-7Anthropic
    anthropic
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms
    9A
    claude-opus-4-8Anthropic
    anthropic
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms
    10A
    claude-haiku-4-5-20251001Anthropic
    anthropic
    Avg Latency800ms
    Throughput2500 tok/s
    TTFT120ms

    LLM Leaderboard

    Compare the most popular models on DeepToken by token volume

    Showing 10 of 101 models

    RankModel IDFamilyTokens (7d)β–ΌRequests (7d)Avg LatencyAvg TTFT7d Change
    1
    O
    gpt-4o-mini
    openai
    OpenAI1.15T890M550ms90ms↑ +15%
    2
    O
    gpt-4o
    openai
    OpenAI950B420M850ms120ms↑ +10%
    3
    G
    gemini-2.5-flash
    google
    Google620B290M450ms80ms↑ +23%
    4
    A
    claude-sonnet-4-20250514
    anthropic
    Anthropic620B240M850ms140msNEW
    5
    G
    gemini-2.5-pro
    google
    Google410B140M1100ms160ms↑ +11%
    6
    A
    claude-opus-4-20250514
    anthropic
    Anthropic150B60M2400ms350msNEW
    7
    G
    gemini-3.1-flash-image-2k-9x16
    google
    Google10B5M800ms120ms↓ -
    8
    G
    gemini-3.1-flash-image-4k
    google
    Google10B5M800ms120ms↓ -
    9
    A
    claude-opus-4-6
    anthropic
    Anthropic10B5M800ms120ms↓ -
    10
    A
    claude-opus-4-6-thinking
    anthropic
    Anthropic10B5M800ms120ms↓ -

    Market Share (Tokens)

    OpenAI: 48% (2.33T tokens)Google: 32% (1.55T tokens)Anthropic: 19% (920B tokens)DeepSeek: 0.8% (40B tokens)Other: 0.2% (10B tokens)4.85Ttotal tokens
    OpenAI
    2.33T48%
    Google
    1.55T32%
    Anthropic
    920B19%
    DeepSeek
    40B0.8%
    Other
    10B0.2%
    OpenAI: 48% (2.33T tokens)Google: 32% (1.55T tokens)Anthropic: 19% (920B tokens)DeepSeek: 0.8% (40B tokens)Other: 0.2% (10B tokens)
    How this rankings list works

    Rankings are updated every minute using a rolling 7-day request window. Tokens are counted as total processed (input + output).

    Latency measures the time to completion, while TTFT represents the Time to First Token. Default seed values are blended with real-time routing metrics to bootstrap new nodes.

    Updated Jul 10, 2026, 3:37 PM

    View catalog