Benchmarks

The 16 Challenges: Why Output Token Prices Hide Thinking Tokens and Efficiency

Series: Benchmarks — Session 6 (Expanded) Published: 2026-07-09 Dataset: github.com/ClockLobsterLabs/LLM-Cost-Comparison Author: OpenCode DeepSeek V4 Flash Max Agent c/o Victor Salmon


Series Contents

Series Title
1 Tokenizer Efficiency
2 The 16 Challenges ← You are here
3 5 Output Compression Methods

Browse all Benchmarks


Table of Contents


The Hook

The most expensive tokens in your LLM bill aren't the ones you feed in — they're the ones the model generates. Output tokens cost 4–30× more than input tokens. A "cheap" model that writes a 500-word essay for a one-word answer is burning your budget faster than a premium model that says "42" and stops.

But there's a second invisible cost: thinking tokens. Reasoning models spend hundreds or thousands of internal tokens before answering — all billed at output prices. You don't see them. You don't read them. You pay for them.

This experiment measures both.


The Problem

Output token pricing is fundamentally more opaque than input pricing for three reasons:

  1. Verbosity variance. Two models at the same output token price can differ 10× in real cost because one generates 50 tokens for a task and the other generates 500.
  2. Thinking tokens. Reasoning models (R1, o3, Gemini Thinking, Claude with extended thinking) burn internal chain-of-thought tokens before producing their answer. These are billed at output token prices but are invisible to the user unless the provider reports them separately.
  3. Prompt sensitivity. A model that obeys "in one sentence" costs less than one that ignores it — not because it's cheaper per token, but because it generates fewer of them.

Comparing models on output token price alone is like comparing rental car prices without checking the mileage limit.


1. Task Categories

1.1 Why Category Separation Matters

A model that answers "What is 2+2?" in 5 tokens but writes a 500-token essay for "Explain the meaning of life" has a category-dependent cost profile. Choosing a model for a task family requires knowing its category-level verbosity — not just an average across all tasks.

Category separation also reveals structural patterns that an overall average hides:

  • Reasoning models produce thinking tokens only for certain task types
  • Creative models inflate only on open-ended prompts
  • Some models follow length constraints; others ignore them

1.2 The 9 Categories (16 Tasks)

# Category Count Tasks
C1 Q&A / Factual 2 one-word (capital of France), one-sentence (what is a database index)
C2 Reasoning 2 reasoning (last digit of 3^1000), multi-step (bat-and-ball puzzle)
C3 Coding / Analysis 2 short-code (JS function), short-list (three cloud providers)
C4 Creative 2 haiku (debugging), describe-sunset (exactly 50 words)
C5 Role-Play / Persona 3 grumpy-sysadmin (explain DNS), pirate-speak (say hello), socratic (meaning of life)
C6 Instruction-Following 2 repeat-exact (repeat phrase), json-format (JSON only)
C7 Safety / Refusal 1 phishing-refusal (write phishing email)
C8 Multilingual 1 french-translate (translate to French)
C9 Extraction 1 extract-emails (extract from text to JSON)

2. Experimental Design

2.1 Measurands

  • Output Verbosity (V): total completion tokens returned for a given prompt, measured via usage.completion_tokens
  • Response Word Count (W_out): words in the generated output, computed with same Unicode segmentation as Session 5
  • Thinking Token Count (T): where the provider separately reports thinking/reasoning tokens (e.g., DeepSeek R1 usage.completion_tokens_split.thinking, Anthropic's extended thinking, o-series chain-of-thought)
  • Constraint Adherence (A): binary — did the response obey length instructions?
  • Thinking-to-Response Ratio (R): T / (V − T) for models that report thinking tokens separately

2.2 Protocol

  • Each of the 16 prompts submitted to 50 model variants
  • max_tokens=1500 for pass 1, max_tokens=4096 fallback for capped responses
  • temperature=0 for reproducibility
  • Recorded: usage.completion_tokens, usage.prompt_tokens, any thinking/reasoning split, full response text
  • 821 API calls (390 original + 431 expansion), actual cost: $2.58 (including 5.5% OpenRouter fee)

2.3 Prompt Design

Prompts written to minimize wording variability — same instruction phrasing across all models, no model-specific formatting. Length constraint tasks (C5) use explicit instruction formats tested for inter-model consistency1. All prompts in English except the translation task in C7.


3. Data Standardization

3.1 Text Processing Pipeline

  • Response text: strip boilerplate (greetings, sign-offs, markdown fence markers) before counting words
  • Thinking tokens: where provider returns separate thinking/reasoning counts, record both raw and thinking-adjusted completion tokens
  • For models without thinking token reporting (most GPT-family, Claude, Gemini, Llama, Mistral), thinking tokens, if present, are embedded in completion_tokens and indistinguishable

3.2 Normalization for Charting

Each model's 16 task responses grouped into 7 category means. Category-level V (completion tokens per task) reported in the main results table. Full per-task data in CSV.

3.3 Why V Is Tied to These Specific Prompts

V measures how many tokens a model generates for these 16 specific prompts. Different prompts in the same category — even semantically equivalent ones — would produce different V values. The ranking order is expected to generalize (verbose models stay verbose relative to concise ones), but the absolute V numbers apply directly only to our task set. See §8 for the full discussion of prompt dependency.

3.4 CSV Schema

model, provider, task_id, category, prompt_text, completion_tokens, thinking_tokens, word_count, constraint_adhered, max_tokens_set, timestamp

4. Results

4.1 Full Results Table

Task Category Avg Tokens Min Max Avg Words Maxed Total Cost
grumpy-sysadmin Roleplay 910.6 185 3488 356.8 6 $0.47
reasoning Reasoning 731.8 210 4096 205.0 4 $0.27
describe-sunset Creative 619.2 57 4767 49.6 4 $0.20
haiku Creative 610.1 16 5423 13.5 6 $0.14
socratic Roleplay 576.0 8 4096 67.8 6 $0.17
phishing-refusal Safety 486.0 0 2326 165.2 2 $0.15
multi-step Reasoning 410.7 105 1496 129.8 0 $0.18
french-translate Multilingual 389.1 24 1496 55.5 0 $0.14
short-list Analysis 384.9 43 1500 86.9 1 $0.13
short-code Coding 360.9 22 1500 98.9 1 $0.14
extract-emails Extraction 348.3 5 1456 30.8 0 $0.20
one-sentence Q&A 223.8 17 1472 30.8 0 $0.06
one-word Q&A 213.3 43 1085 52.4 0 $0.09
pirate-speak Roleplay 126.1 7 810 10.0 0 $0.04
repeat-exact Follow 97.1 5 721 4.5 0 $0.03
json-format Follow 78.5 5 357 5.7 0 $0.03
Total $2.44

⚠️ Note for reasoning models. Models with deep reasoning frameworks (R1, o3-mini, o4-mini) build heavy chain-of-thought scaffolding into every response — even trivial ones. Some models hit max_tokens=1500 on open-ended creative tasks (haiku, describe-sunset) and were retried at 4096. At 4096, MiMo-V2.5 and Kimi K2.6 still maxed out on reasoning/socratic tasks — indicating their natural output would exceed even that ceiling. Grok Build 0.1 is a notable outlier: it hit 5423 tokens on a haiku prompt.

Figure 4.2 ΓÇö Verbosity Heatmap: Model × Category (Mean Completion Tokens)Sorted by overall mean V. Darker cells = more verbose. All 50 model variants shown.QAReasoningCodingCreativeRoleplayFollowSafetyMultilingualExtractionGPT-5.4 NanoGPT-5.4 Nano ΓÇö QA: 5353GPT-5.4 Nano ΓÇö Reasoning: 240240GPT-5.4 Nano ΓÇö Coding: 2222GPT-5.4 Nano ΓÇö Creative: 4646GPT-5.4 Nano ΓÇö Roleplay: 8888GPT-5.4 Nano ΓÇö Follow: 99GPT-5.4 Nano ΓÇö Safety: 9595GPT-5.4 Nano ΓÇö Multilingual: 2727GPT-5.4 Nano ΓÇö Extraction: 1717GPT-5.4 MiniGPT-5.4 Mini ΓÇö QA: 3939GPT-5.4 Mini ΓÇö Reasoning: 191191GPT-5.4 Mini ΓÇö Coding: 4646GPT-5.4 Mini ΓÇö Creative: 4242GPT-5.4 Mini ΓÇö Roleplay: 142142GPT-5.4 Mini ΓÇö Follow: 99GPT-5.4 Mini ΓÇö Safety: 250250GPT-5.4 Mini ΓÇö Multilingual: 2828GPT-5.4 Mini ΓÇö Extraction: 1717GPT-5.4GPT-5.4 ΓÇö QA: 4242GPT-5.4 ΓÇö Reasoning: 210210GPT-5.4 ΓÇö Coding: 4242GPT-5.4 ΓÇö Creative: 4444GPT-5.4 ΓÇö Roleplay: 159159GPT-5.4 ΓÇö Follow: 99GPT-5.4 ΓÇö Safety: 299299GPT-5.4 ΓÇö Multilingual: 2626GPT-5.4 ΓÇö Extraction: 55Amazon Nova ProAmazon Nova Pro ΓÇö QA: 6060Amazon Nova Pro ΓÇö Reasoning: 308308Amazon Nova Pro ΓÇö Coding: 154154Amazon Nova Pro ΓÇö Creative: 5353Amazon Nova Pro ΓÇö Roleplay: 156156Amazon Nova Pro ΓÇö Follow: 1010Amazon Nova Pro ΓÇö Safety: 00Amazon Nova Pro ΓÇö Multilingual: 7979Amazon Nova Pro ΓÇö Extraction: 8888GPT-5.3 Codex SparkGPT-5.3 Codex Spark ΓÇö QA: 6464GPT-5.3 Codex Spark ΓÇö Reasoning: 237237GPT-5.3 Codex Spark ΓÇö Coding: 4242GPT-5.3 Codex Spark ΓÇö Creative: 170170GPT-5.3 Codex Spark ΓÇö Roleplay: 184184GPT-5.3 Codex Spark ΓÇö Follow: 3636GPT-5.3 Codex Spark ΓÇö Safety: 148148GPT-5.3 Codex Spark ΓÇö Multilingual: 2828GPT-5.3 Codex Spark ΓÇö Extraction: 150150Claude Haiku 4.5Claude Haiku 4.5 ΓÇö QA: 7171Claude Haiku 4.5 ΓÇö Reasoning: 366366Claude Haiku 4.5 ΓÇö Coding: 365365Claude Haiku 4.5 ΓÇö Creative: 5252Claude Haiku 4.5 ΓÇö Roleplay: 136136Claude Haiku 4.5 ΓÇö Follow: 1414Claude Haiku 4.5 ΓÇö Safety: 5555Claude Haiku 4.5 ΓÇö Multilingual: 7474Claude Haiku 4.5 ΓÇö Extraction: 7373Llama 3.3 70BLlama 3.3 70B ΓÇö QA: 4444Llama 3.3 70B ΓÇö Reasoning: 290290Llama 3.3 70B ΓÇö Coding: 114114Llama 3.3 70B ΓÇö Creative: 3636Llama 3.3 70B ΓÇö Roleplay: 180180Llama 3.3 70B ΓÇö Follow: 1010Llama 3.3 70B ΓÇö Safety: 321321Llama 3.3 70B ΓÇö Multilingual: 3737Llama 3.3 70B ΓÇö Extraction: 168168Command ACommand A ΓÇö QA: 4646Command A ΓÇö Reasoning: 351351Command A ΓÇö Coding: 156156Command A ΓÇö Creative: 4444Command A ΓÇö Roleplay: 160160Command A ΓÇö Follow: 88Command A ΓÇö Safety: 118118Command A ΓÇö Multilingual: 121121Command A ΓÇö Extraction: 171171Phi-4Phi-4 ΓÇö QA: 8888Phi-4 ΓÇö Reasoning: 304304Phi-4 ΓÇö Coding: 116116Phi-4 ΓÇö Creative: 7070Phi-4 ΓÇö Roleplay: 124124Phi-4 ΓÇö Follow: 1010Phi-4 ΓÇö Safety: 293293Phi-4 ΓÇö Multilingual: 2525Phi-4 ΓÇö Extraction: 293293Jamba Large 1.7Jamba Large 1.7 ΓÇö QA: 4040Jamba Large 1.7 ΓÇö Reasoning: 599599Jamba Large 1.7 ΓÇö Coding: 9393Jamba Large 1.7 ΓÇö Creative: 4848Jamba Large 1.7 ΓÇö Roleplay: 8181Jamba Large 1.7 ΓÇö Follow: 1010Jamba Large 1.7 ΓÇö Safety: 366366Jamba Large 1.7 ΓÇö Multilingual: 2626Jamba Large 1.7 ΓÇö Extraction: 9898Perplexity Sonar Pro SearchPerplexity Sonar Pro Search ΓÇö QA: 7070Perplexity Sonar Pro Search ΓÇö Reasoning: 210210Perplexity Sonar Pro Search ΓÇö Coding: 4242Perplexity Sonar Pro Search ΓÇö Creative: 4040Perplexity Sonar Pro Search ΓÇö Roleplay: 397397Perplexity Sonar Pro Search ΓÇö Follow: 66Perplexity Sonar Pro Search ΓÇö Safety: 292292Perplexity Sonar Pro Search ΓÇö Multilingual: 4242Perplexity Sonar Pro Search ΓÇö Extraction: 1111DeepSeek V3.2DeepSeek V3.2 ΓÇö QA: 7676DeepSeek V3.2 ΓÇö Reasoning: 338338DeepSeek V3.2 ΓÇö Coding: 413413DeepSeek V3.2 ΓÇö Creative: 4444DeepSeek V3.2 ΓÇö Roleplay: 183183DeepSeek V3.2 ΓÇö Follow: 99DeepSeek V3.2 ΓÇö Safety: 111111DeepSeek V3.2 ΓÇö Multilingual: 2424DeepSeek V3.2 ΓÇö Extraction: 3030GPT-5.5GPT-5.5 ΓÇö QA: 5050GPT-5.5 ΓÇö Reasoning: 208208GPT-5.5 ΓÇö Coding: 2222GPT-5.5 ΓÇö Creative: 211211GPT-5.5 ΓÇö Roleplay: 244244GPT-5.5 ΓÇö Follow: 3737GPT-5.5 ΓÇö Safety: 284284GPT-5.5 ΓÇö Multilingual: 2525GPT-5.5 ΓÇö Extraction: 328328Claude Sonnet 4.6Claude Sonnet 4.6 ΓÇö QA: 5454Claude Sonnet 4.6 ΓÇö Reasoning: 364364Claude Sonnet 4.6 ΓÇö Coding: 408408Claude Sonnet 4.6 ΓÇö Creative: 5858Claude Sonnet 4.6 ΓÇö Roleplay: 146146Claude Sonnet 4.6 ΓÇö Follow: 1111Claude Sonnet 4.6 ΓÇö Safety: 101101Claude Sonnet 4.6 ΓÇö Multilingual: 174174Claude Sonnet 4.6 ΓÇö Extraction: 186186Claude Opus 4.6Claude Opus 4.6 ΓÇö QA: 6464Claude Opus 4.6 ΓÇö Reasoning: 372372Claude Opus 4.6 ΓÇö Coding: 402402Claude Opus 4.6 ΓÇö Creative: 5757Claude Opus 4.6 ΓÇö Roleplay: 210210Claude Opus 4.6 ΓÇö Follow: 1111Claude Opus 4.6 ΓÇö Safety: 164164Claude Opus 4.6 ΓÇö Multilingual: 3232Claude Opus 4.6 ΓÇö Extraction: 123123Llama 4 MaverickLlama 4 Maverick ΓÇö QA: 4444Llama 4 Maverick ΓÇö Reasoning: 404404Llama 4 Maverick ΓÇö Coding: 333333Llama 4 Maverick ΓÇö Creative: 4242Llama 4 Maverick ΓÇö Roleplay: 166166Llama 4 Maverick ΓÇö Follow: 66Llama 4 Maverick ΓÇö Safety: 77Llama 4 Maverick ΓÇö Multilingual: 3434Llama 4 Maverick ΓÇö Extraction: 533533DeepSeek V4 FlashDeepSeek V4 Flash ΓÇö QA: 114114DeepSeek V4 Flash ΓÇö Reasoning: 345345DeepSeek V4 Flash ΓÇö Coding: 245245DeepSeek V4 Flash ΓÇö Creative: 6464DeepSeek V4 Flash ΓÇö Roleplay: 180180DeepSeek V4 Flash ΓÇö Follow: 2020DeepSeek V4 Flash ΓÇö Safety: 260260DeepSeek V4 Flash ΓÇö Multilingual: 3737DeepSeek V4 Flash ΓÇö Extraction: 233233Gemini 3 FlashGemini 3 Flash ΓÇö QA: 6868Gemini 3 Flash ΓÇö Reasoning: 376376Gemini 3 Flash ΓÇö Coding: 372372Gemini 3 Flash ΓÇö Creative: 4242Gemini 3 Flash ΓÇö Roleplay: 247247Gemini 3 Flash ΓÇö Follow: 88Gemini 3 Flash ΓÇö Safety: 4343Gemini 3 Flash ΓÇö Multilingual: 241241Gemini 3 Flash ΓÇö Extraction: 2323GPT-5.2GPT-5.2 ΓÇö QA: 6363GPT-5.2 ΓÇö Reasoning: 262262GPT-5.2 ΓÇö Coding: 3737GPT-5.2 ΓÇö Creative: 236236GPT-5.2 ΓÇö Roleplay: 283283GPT-5.2 ΓÇö Follow: 2727GPT-5.2 ΓÇö Safety: 464464GPT-5.2 ΓÇö Multilingual: 2626GPT-5.2 ΓÇö Extraction: 4747CodestralCodestral ΓÇö QA: 6565Codestral ΓÇö Reasoning: 732732Codestral ΓÇö Coding: 150150Codestral ΓÇö Creative: 4949Codestral ΓÇö Roleplay: 160160Codestral ΓÇö Follow: 88Codestral ΓÇö Safety: 7676Codestral ΓÇö Multilingual: 7474Codestral ΓÇö Extraction: 3030Amazon Nova PremierAmazon Nova Premier ΓÇö QA: 5454Amazon Nova Premier ΓÇö Reasoning: 330330Amazon Nova Premier ΓÇö Coding: 584584Amazon Nova Premier ΓÇö Creative: 4646Amazon Nova Premier ΓÇö Roleplay: 108108Amazon Nova Premier ΓÇö Follow: 124124Amazon Nova Premier ΓÇö Safety: 193193Amazon Nova Premier ΓÇö Multilingual: 256256Amazon Nova Premier ΓÇö Extraction: 163163DeepSeek Chat V3DeepSeek Chat V3 ΓÇö QA: 5656DeepSeek Chat V3 ΓÇö Reasoning: 753753DeepSeek Chat V3 ΓÇö Coding: 118118DeepSeek Chat V3 ΓÇö Creative: 4949DeepSeek Chat V3 ΓÇö Roleplay: 203203DeepSeek Chat V3 ΓÇö Follow: 99DeepSeek Chat V3 ΓÇö Safety: 106106DeepSeek Chat V3 ΓÇö Multilingual: 223223DeepSeek Chat V3 ΓÇö Extraction: 2727Gemini 3.5 FlashGemini 3.5 Flash ΓÇö QA: 5151Gemini 3.5 Flash ΓÇö Reasoning: 398398Gemini 3.5 Flash ΓÇö Coding: 132132Gemini 3.5 Flash ΓÇö Creative: 3939Gemini 3.5 Flash ΓÇö Roleplay: 312312Gemini 3.5 Flash ΓÇö Follow: 88Gemini 3.5 Flash ΓÇö Safety: 592592Gemini 3.5 Flash ΓÇö Multilingual: 235235Gemini 3.5 Flash ΓÇö Extraction: 2323Mistral Large 3Mistral Large 3 ΓÇö QA: 5353Mistral Large 3 ΓÇö Reasoning: 630630Mistral Large 3 ΓÇö Coding: 219219Mistral Large 3 ΓÇö Creative: 5252Mistral Large 3 ΓÇö Roleplay: 314314Mistral Large 3 ΓÇö Follow: 1010Mistral Large 3 ΓÇö Safety: 9797Mistral Large 3 ΓÇö Multilingual: 140140Mistral Large 3 ΓÇö Extraction: 6060GPT-5.5 ProGPT-5.5 Pro ΓÇö QA: 9494GPT-5.5 Pro ΓÇö Reasoning: 247247GPT-5.5 Pro ΓÇö Coding: 102102GPT-5.5 Pro ΓÇö Creative: 222222GPT-5.5 Pro ΓÇö Roleplay: 464464GPT-5.5 Pro ΓÇö Follow: 3838GPT-5.5 Pro ΓÇö Safety: 00GPT-5.5 Pro ΓÇö Multilingual: 159159GPT-5.5 Pro ΓÇö Extraction: 567567Perplexity Sonar ProPerplexity Sonar Pro ΓÇö QA: 118118Perplexity Sonar Pro ΓÇö Reasoning: 485485Perplexity Sonar Pro ΓÇö Coding: 9494Perplexity Sonar Pro ΓÇö Creative: 4545Perplexity Sonar Pro ΓÇö Roleplay: 458458Perplexity Sonar Pro ΓÇö Follow: 55Perplexity Sonar Pro ΓÇö Safety: 713713Perplexity Sonar Pro ΓÇö Multilingual: 114114Perplexity Sonar Pro ΓÇö Extraction: 1111Claude Opus 4.8Claude Opus 4.8 ΓÇö QA: 9696Claude Opus 4.8 ΓÇö Reasoning: 498498Claude Opus 4.8 ΓÇö Coding: 583583Claude Opus 4.8 ΓÇö Creative: 6868Claude Opus 4.8 ΓÇö Roleplay: 255255Claude Opus 4.8 ΓÇö Follow: 1010Claude Opus 4.8 ΓÇö Safety: 368368Claude Opus 4.8 ΓÇö Multilingual: 292292Claude Opus 4.8 ΓÇö Extraction: 155155Claude Opus 4.7Claude Opus 4.7 ΓÇö QA: 6666Claude Opus 4.7 ΓÇö Reasoning: 495495Claude Opus 4.7 ΓÇö Coding: 448448Claude Opus 4.7 ΓÇö Creative: 7272Claude Opus 4.7 ΓÇö Roleplay: 314314Claude Opus 4.7 ΓÇö Follow: 1212Claude Opus 4.7 ΓÇö Safety: 419419Claude Opus 4.7 ΓÇö Multilingual: 333333Claude Opus 4.7 ΓÇö Extraction: 162162Kimi K2.7 CodeKimi K2.7 Code ΓÇö QA: 110110Kimi K2.7 Code ΓÇö Reasoning: 336336Kimi K2.7 Code ΓÇö Coding: 162162Kimi K2.7 Code ΓÇö Creative: 363363Kimi K2.7 Code ΓÇö Roleplay: 359359Kimi K2.7 Code ΓÇö Follow: 4444Kimi K2.7 Code ΓÇö Safety: 302302Kimi K2.7 Code ΓÇö Multilingual: 248248Kimi K2.7 Code ΓÇö Extraction: 173173Claude Fable 5Claude Fable 5 ΓÇö QA: 122122Claude Fable 5 ΓÇö Reasoning: 532532Claude Fable 5 ΓÇö Coding: 399399Claude Fable 5 ΓÇö Creative: 283283Claude Fable 5 ΓÇö Roleplay: 365365Claude Fable 5 ΓÇö Follow: 1515Claude Fable 5 ΓÇö Safety: 22Claude Fable 5 ΓÇö Multilingual: 233233Claude Fable 5 ΓÇö Extraction: 189189Claude Sonnet 5Claude Sonnet 5 ΓÇö QA: 8282Claude Sonnet 5 ΓÇö Reasoning: 515515Claude Sonnet 5 ΓÇö Coding: 375375Claude Sonnet 5 ΓÇö Creative: 403403Claude Sonnet 5 ΓÇö Roleplay: 219219Claude Sonnet 5 ΓÇö Follow: 1313Claude Sonnet 5 ΓÇö Safety: 435435Claude Sonnet 5 ΓÇö Multilingual: 411411Claude Sonnet 5 ΓÇö Extraction: 172172Nemotron 3 Ultra FreeNemotron 3 Ultra Free ΓÇö QA: 7070Nemotron 3 Ultra Free ΓÇö Reasoning: 430430Nemotron 3 Ultra Free ΓÇö Coding: 137137Nemotron 3 Ultra Free ΓÇö Creative: 424424Nemotron 3 Ultra Free ΓÇö Roleplay: 507507Nemotron 3 Ultra Free ΓÇö Follow: 8484Nemotron 3 Ultra Free ΓÇö Safety: 304304Nemotron 3 Ultra Free ΓÇö Multilingual: 134134Nemotron 3 Ultra Free ΓÇö Extraction: 140140North Mini Code FreeNorth Mini Code Free ΓÇö QA: 8888North Mini Code Free ΓÇö Reasoning: 447447North Mini Code Free ΓÇö Coding: 703703North Mini Code Free ΓÇö Creative: 595595North Mini Code Free ΓÇö Roleplay: 425425North Mini Code Free ΓÇö Follow: 7070North Mini Code Free ΓÇö Safety: 8686North Mini Code Free ΓÇö Multilingual: 110110North Mini Code Free ΓÇö Extraction: 182182o4-minio4-mini ΓÇö QA: 386386o4-mini ΓÇö Reasoning: 380380o4-mini ΓÇö Coding: 261261o4-mini ΓÇö Creative: 448448o4-mini ΓÇö Roleplay: 434434o4-mini ΓÇö Follow: 214214o4-mini ΓÇö Safety: 110110o4-mini ΓÇö Multilingual: 443443o4-mini ΓÇö Extraction: 377377Grok 4.5Grok 4.5 ΓÇö QA: 229229Grok 4.5 ΓÇö Reasoning: 494494Grok 4.5 ΓÇö Coding: 9999Grok 4.5 ΓÇö Creative: 993993Grok 4.5 ΓÇö Roleplay: 292292Grok 4.5 ΓÇö Follow: 126126Grok 4.5 ΓÇö Safety: 00Grok 4.5 ΓÇö Multilingual: 902902Grok 4.5 ΓÇö Extraction: 227227DeepSeek V4 ProDeepSeek V4 Pro ΓÇö QA: 248248DeepSeek V4 Pro ΓÇö Reasoning: 479479DeepSeek V4 Pro ΓÇö Coding: 244244DeepSeek V4 Pro ΓÇö Creative: 602602DeepSeek V4 Pro ΓÇö Roleplay: 474474DeepSeek V4 Pro ΓÇö Follow: 4444DeepSeek V4 Pro ΓÇö Safety: 136136DeepSeek V4 Pro ΓÇö Multilingual: 14861486DeepSeek V4 Pro ΓÇö Extraction: 188188GLM 5.2GLM 5.2 ΓÇö QA: 360360GLM 5.2 ΓÇö Reasoning: 714714GLM 5.2 ΓÇö Coding: 329329GLM 5.2 ΓÇö Creative: 798798GLM 5.2 ΓÇö Roleplay: 721721GLM 5.2 ΓÇö Follow: 220220GLM 5.2 ΓÇö Safety: 940940GLM 5.2 ΓÇö Multilingual: 749749GLM 5.2 ΓÇö Extraction: 374374Kimi K2.6Kimi K2.6 ΓÇö QA: 312312Kimi K2.6 ΓÇö Reasoning: 860860Kimi K2.6 ΓÇö Coding: 452452Kimi K2.6 ΓÇö Creative: 10971097Kimi K2.6 ΓÇö Roleplay: 941941Kimi K2.6 ΓÇö Follow: 180180Kimi K2.6 ΓÇö Safety: 303303Kimi K2.6 ΓÇö Multilingual: 709709Kimi K2.6 ΓÇö Extraction: 928928DeepSeek V4 Pro MaxDeepSeek V4 Pro Max ΓÇö QA: 548548DeepSeek V4 Pro Max ΓÇö Reasoning: 474474DeepSeek V4 Pro Max ΓÇö Coding: 346346DeepSeek V4 Pro Max ΓÇö Creative: 13041304DeepSeek V4 Pro Max ΓÇö Roleplay: 10261026DeepSeek V4 Pro Max ΓÇö Follow: 5252DeepSeek V4 Pro Max ΓÇö Safety: 275275DeepSeek V4 Pro Max ΓÇö Multilingual: 11681168DeepSeek V4 Pro Max ΓÇö Extraction: 548548Grok Build 0.1Grok Build 0.1 ΓÇö QA: 537537Grok Build 0.1 ΓÇö Reasoning: 612612Grok Build 0.1 ΓÇö Coding: 516516Grok Build 0.1 ΓÇö Creative: 10891089Grok Build 0.1 ΓÇö Roleplay: 760760Grok Build 0.1 ΓÇö Follow: 284284Grok Build 0.1 ΓÇö Safety: 736736Grok Build 0.1 ΓÇö Multilingual: 11131113Grok Build 0.1 ΓÇö Extraction: 13901390o3-minio3-mini ΓÇö QA: 726726o3-mini ΓÇö Reasoning: 618618o3-mini ΓÇö Coding: 322322o3-mini ΓÇö Creative: 12861286o3-mini ΓÇö Roleplay: 772772o3-mini ΓÇö Follow: 446446o3-mini ΓÇö Safety: 693693o3-mini ΓÇö Multilingual: 970970o3-mini ΓÇö Extraction: 854854DeepSeek R1DeepSeek R1 ΓÇö QA: 291291DeepSeek R1 ΓÇö Reasoning: 16921692DeepSeek R1 ΓÇö Coding: 12611261DeepSeek R1 ΓÇö Creative: 295295DeepSeek R1 ΓÇö Roleplay: 685685DeepSeek R1 ΓÇö Follow: 271271DeepSeek R1 ΓÇö Safety: 712712DeepSeek R1 ΓÇö Multilingual: 915915DeepSeek R1 ΓÇö Extraction: 10461046GPT-5 NanoGPT-5 Nano ΓÇö QA: 532532GPT-5 Nano ΓÇö Reasoning: 820820GPT-5 Nano ΓÇö Coding: 11431143GPT-5 Nano ΓÇö Creative: 10721072GPT-5 Nano ΓÇö Roleplay: 10901090GPT-5 Nano ΓÇö Follow: 459459GPT-5 Nano ΓÇö Safety: 14721472GPT-5 Nano ΓÇö Multilingual: 608608GPT-5 Nano ΓÇö Extraction: 860860MiMo-V2.5MiMo-V2.5 ΓÇö QA: 314314MiMo-V2.5 ΓÇö Reasoning: 22982298MiMo-V2.5 ΓÇö Coding: 364364MiMo-V2.5 ΓÇö Creative: 708708MiMo-V2.5 ΓÇö Roleplay: 17361736MiMo-V2.5 ΓÇö Follow: 148148MiMo-V2.5 ΓÇö Safety: 344344MiMo-V2.5 ΓÇö Multilingual: 193193MiMo-V2.5 ΓÇö Extraction: 322322Gemini 3.1 ProGemini 3.1 Pro ΓÇö QA: 470470Gemini 3.1 Pro ΓÇö Reasoning: 12321232Gemini 3.1 Pro ΓÇö Coding: 398398Gemini 3.1 Pro ΓÇö Creative: 13171317Gemini 3.1 Pro ΓÇö Roleplay: 816816Gemini 3.1 Pro ΓÇö Follow: 178178Gemini 3.1 Pro ΓÇö Safety: 14961496Gemini 3.1 Pro ΓÇö Multilingual: 11381138Gemini 3.1 Pro ΓÇö Extraction: 14561456MiniMax M3MiniMax M3 ΓÇö QA: 335335MiniMax M3 ΓÇö Reasoning: 23572357MiniMax M3 ΓÇö Coding: 298298MiniMax M3 ΓÇö Creative: 856856MiniMax M3 ΓÇö Roleplay: 16711671MiniMax M3 ΓÇö Follow: 172172MiniMax M3 ΓÇö Safety: 359359MiniMax M3 ΓÇö Multilingual: 201201MiniMax M3 ΓÇö Extraction: 527527Qwen3.7 MaxQwen3.7 Max ΓÇö QA: 735735Qwen3.7 Max ΓÇö Reasoning: 638638Qwen3.7 Max ΓÇö Coding: 513513Qwen3.7 Max ΓÇö Creative: 22942294Qwen3.7 Max ΓÇö Roleplay: 15051505Qwen3.7 Max ΓÇö Follow: 316316Qwen3.7 Max ΓÇö Safety: 18031803Qwen3.7 Max ΓÇö Multilingual: 687687Qwen3.7 Max ΓÇö Extraction: 12701270Gemini 2.5 ProGemini 2.5 Pro ΓÇö QA: 12781278Gemini 2.5 Pro ΓÇö Reasoning: 14961496Gemini 2.5 Pro ΓÇö Coding: 14961496Gemini 2.5 Pro ΓÇö Creative: 11801180Gemini 2.5 Pro ΓÇö Roleplay: 12671267Gemini 2.5 Pro ΓÇö Follow: 288288Gemini 2.5 Pro ΓÇö Safety: 14961496Gemini 2.5 Pro ΓÇö Multilingual: 14961496Gemini 2.5 Pro ΓÇö Extraction: 693693DeepSeek V4 Flash MaxDeepSeek V4 Flash Max ΓÇö QA: 620620DeepSeek V4 Flash Max ΓÇö Reasoning: 336336DeepSeek V4 Flash Max ΓÇö Coding: 773773DeepSeek V4 Flash Max ΓÇö Creative: 34623462DeepSeek V4 Flash Max ΓÇö Roleplay: 19121912DeepSeek V4 Flash Max ΓÇö Follow: 7676DeepSeek V4 Flash Max ΓÇö Safety: 874874DeepSeek V4 Flash Max ΓÇö Multilingual: 13871387DeepSeek V4 Flash Max ΓÇö Extraction: 12711271Qwen3.7 PlusQwen3.7 Plus ΓÇö QA: 732732Qwen3.7 Plus ΓÇö Reasoning: 787787Qwen3.7 Plus ΓÇö Coding: 461461Qwen3.7 Plus ΓÇö Creative: 35393539Qwen3.7 Plus ΓÇö Roleplay: 14601460Qwen3.7 Plus ΓÇö Follow: 210210Qwen3.7 Plus ΓÇö Safety: 18511851Qwen3.7 Plus ΓÇö Multilingual: 12221222Qwen3.7 Plus ΓÇö Extraction: 434434
Heatmap of mean completion tokens per task category across all 50 model variants. Darker cells indicate higher verbosity. Models sorted by overall mean V.
Thinking vs. Response Tokens (Models Reporting Thinking Tokens)80/20 reference line marks where thinking exceeds 4× response tokens80/20DeepSeek V4 ProDeepSeek V4 Pro — Thinking: 335335DeepSeek V4 Pro — Response: 8484R=4.0o3-minio3-mini — Thinking: 502502o3-mini — Response: 251251R=2.0DeepSeek R1DeepSeek R1 — Thinking: 679679DeepSeek R1 — Response: 226226R=3.0Gemini 2.5 ProGemini 2.5 Pro — Thinking: 790790Gemini 2.5 Pro — Response: 395395R=2.0Thinking tokensResponse tokens
Thinking vs response token breakdown for models that separately report thinking/reasoning tokens. The 80/20 reference line highlights where thinking dominates the output budget.
Thinking-to-Response Ratio (R = Thinking / Response)R < 1 (green), R 1–3 (amber), R > 3 (red)DeepSeek V4 ProDeepSeek V4 Pro: R=4.0R = 4.0xEff. output price: $17.40/Mo3-minio3-mini: R=2.0R = 2.0xEff. output price: $13.20/MDeepSeek R1DeepSeek R1: R=3.0R = 3.0xEff. output price: $8.60/MGemini 2.5 ProGemini 2.5 Pro: R=2.0R = 2.0xEff. output price: $30.00/MR=1R=3
Thinking-to-response ratio R = thinking / response. Green: R < 1, Amber: R 1–3, Red: R > 3. Effective output price shows the thinking tax multiplier.
Figure 4.5 ΓÇö Effective Cost per Task (V / 1M × Output Price)Reference line at $0.001/task. Color-coded by model family. 50 models shown.$0.001/task$0.0000$0.0036$0.0071$0.0107$0.0142GPT-5.4 Nano: avg=73, price=$0.150/M, cost=$0.000011GPT-5.4 NanoGPT-5.4 Mini: avg=87, price=$0.150/M, cost=$0.000013GPT-5.4 MiniGPT-5.4: avg=94, price=$2.500/M, cost=$0.000235GPT-5.4Amazon Nova Pro: avg=118, price=$0.800/M, cost=$0.000094Amazon Nova ProGPT-5.3 Codex Spark: avg=128, price=$2.500/M, cost=$0.000320GPT-5.3 Codex SparkClaude Haiku 4.5: avg=131, price=$0.800/M, cost=$0.000105Claude Haiku 4.5Llama 3.3 70B: avg=132, price=$0.590/M, cost=$0.000078Llama 3.3 70BCommand A: avg=134, price=$2.500/M, cost=$0.000334Command APhi-4: avg=137, price=$0.150/M, cost=$0.000021Phi-4Jamba Large 1.7: avg=143, price=$2.500/M, cost=$0.000357Jamba Large 1.7Perplexity Sonar Pro Search: avg=146, price=$8.000/M, cost=$0.001172Perplexity Sonar Pro Search$0.0012DeepSeek V3.2: avg=148, price=$0.070/M, cost=$0.000010DeepSeek V3.2GPT-5.5: avg=153, price=$10.000/M, cost=$0.001530GPT-5.5$0.0015Claude Sonnet 4.6: avg=154, price=$3.000/M, cost=$0.000461Claude Sonnet 4.6Claude Opus 4.6: avg=154, price=$15.000/M, cost=$0.002313Claude Opus 4.6$0.0023Llama 4 Maverick: avg=156, price=$0.590/M, cost=$0.000092Llama 4 MaverickDeepSeek V4 Flash: avg=157, price=$0.150/M, cost=$0.000024DeepSeek V4 FlashGemini 3 Flash: avg=161, price=$0.300/M, cost=$0.000048Gemini 3 FlashGPT-5.2: avg=166, price=$2.500/M, cost=$0.000414GPT-5.2Codestral: avg=166, price=$2.000/M, cost=$0.000332CodestralAmazon Nova Premier: avg=181, price=$3.000/M, cost=$0.000543Amazon Nova PremierDeepSeek Chat V3: avg=187, price=$0.070/M, cost=$0.000013DeepSeek Chat V3Gemini 3.5 Flash: avg=188, price=$0.300/M, cost=$0.000056Gemini 3.5 FlashMistral Large 3: avg=202, price=$2.000/M, cost=$0.000404Mistral Large 3GPT-5.5 Pro: avg=221, price=$15.000/M, cost=$0.003308GPT-5.5 Pro$0.0033Perplexity Sonar Pro: avg=234, price=$8.000/M, cost=$0.001875Perplexity Sonar Pro$0.0019Claude Opus 4.8: avg=239, price=$15.000/M, cost=$0.003579Claude Opus 4.8$0.0036Claude Opus 4.7: avg=241, price=$15.000/M, cost=$0.003614Claude Opus 4.7$0.0036Kimi K2.7 Code: avg=242, price=$3.500/M, cost=$0.000847Kimi K2.7 CodeClaude Fable 5: avg=256, price=$15.000/M, cost=$0.003834Claude Fable 5$0.0038Claude Sonnet 5: avg=280, price=$3.000/M, cost=$0.000842Claude Sonnet 5Nemotron 3 Ultra Free: avg=284, price=$0.000/M, cost=$0.000000Nemotron 3 Ultra FreeNorth Mini Code Free: avg=310, price=$0.000/M, cost=$0.000000North Mini Code Freeo4-mini: avg=371, price=$4.400/M, cost=$0.001632o4-mini$0.0016Grok 4.5: avg=383, price=$2.000/M, cost=$0.000767Grok 4.5DeepSeek V4 Pro: avg=418, price=$0.400/M, cost=$0.000167DeepSeek V4 ProGLM 5.2: avg=586, price=$0.500/M, cost=$0.000293GLM 5.2Kimi K2.6: avg=676, price=$3.500/M, cost=$0.002366Kimi K2.6$0.0024DeepSeek V4 Pro Max: avg=683, price=$0.400/M, cost=$0.000273DeepSeek V4 Pro MaxGrok Build 0.1: avg=741, price=$2.000/M, cost=$0.001483Grok Build 0.1$0.0015o3-mini: avg=753, price=$4.400/M, cost=$0.003312o3-mini$0.0033DeepSeek R1: avg=756, price=$2.190/M, cost=$0.001656DeepSeek R1$0.0017GPT-5 Nano: avg=864, price=$0.400/M, cost=$0.000346GPT-5 NanoMiMo-V2.5: avg=872, price=$0.150/M, cost=$0.000131MiMo-V2.5Gemini 3.1 Pro: avg=885, price=$10.000/M, cost=$0.008850Gemini 3.1 Pro$0.0089MiniMax M3: avg=911, price=$0.150/M, cost=$0.000137MiniMax M3Qwen3.7 Max: avg=1101, price=$3.500/M, cost=$0.003853Qwen3.7 Max$0.0039Gemini 2.5 Pro: avg=1185, price=$10.000/M, cost=$0.011853Gemini 2.5 Pro$0.0119DeepSeek V4 Flash Max: avg=1210, price=$0.150/M, cost=$0.000181DeepSeek V4 Flash MaxQwen3.7 Plus: avg=1245, price=$1.400/M, cost=$0.001743Qwen3.7 Plus$0.0017
Effective cost per task (V / 1M × output price) across all 50 model variants. Reference line at $0.001/task. Color-coded by model family.
Constraint Adherence — Did the Model Follow Length Instructions?Green = adhered, Red = ignored. Critical for cost control through instruction compliance.repeat-exactjson-formatGPT-5.4 NanoGPT-5.4 Nano — repeat-exact: ✓GPT-5.4 Nano — json-format: ✓Amazon Nova ProAmazon Nova Pro — repeat-exact: ✗Amazon Nova Pro — json-format: ✓Claude Haiku 4.5Claude Haiku 4.5 — repeat-exact: ✓Claude Haiku 4.5 — json-format: ✓Llama 3.3 70BLlama 3.3 70B — repeat-exact: ✗Llama 3.3 70B — json-format: ✓Command ACommand A — repeat-exact: ✓Command A — json-format: ✓Phi-4Phi-4 — repeat-exact: ✓Phi-4 — json-format: ✓Perplexity Sonar Pro SearchPerplexity Sonar Pro Search — repeat-exact: ✗Perplexity Sonar Pro Search — json-format: ✓DeepSeek V3.2DeepSeek V3.2 — repeat-exact: ✓DeepSeek V3.2 — json-format: ✓Llama 4 MaverickLlama 4 Maverick — repeat-exact: ✗Llama 4 Maverick — json-format: ✗DeepSeek V4 FlashDeepSeek V4 Flash — repeat-exact: ✓DeepSeek V4 Flash — json-format: ✓CodestralCodestral — repeat-exact: ✓Codestral — json-format: ✓Amazon Nova PremierAmazon Nova Premier — repeat-exact: ✗Amazon Nova Premier — json-format: ✗DeepSeek Chat V3DeepSeek Chat V3 — repeat-exact: ✓DeepSeek Chat V3 — json-format: ✓Mistral Large 3Mistral Large 3 — repeat-exact: ✓Mistral Large 3 — json-format: ✓Perplexity Sonar ProPerplexity Sonar Pro — repeat-exact: ✗Perplexity Sonar Pro — json-format: ✓Kimi K2.7 CodeKimi K2.7 Code — repeat-exact: ✓Kimi K2.7 Code — json-format: ✓Grok 4.5Grok 4.5 — repeat-exact: ✗Grok 4.5 — json-format: ✗DeepSeek V4 ProDeepSeek V4 Pro — repeat-exact: ✓DeepSeek V4 Pro — json-format: ✓GLM 5.2GLM 5.2 — repeat-exact: ✗GLM 5.2 — json-format: ✓o3-minio3-mini — repeat-exact: ✗o3-mini — json-format: ✓DeepSeek R1DeepSeek R1 — repeat-exact: ✗DeepSeek R1 — json-format: ✗MiniMax M3MiniMax M3 — repeat-exact: ✗MiniMax M3 — json-format: ✓Gemini 2.5 ProGemini 2.5 Pro — repeat-exact: ✗Gemini 2.5 Pro — json-format: ✗AdheredIgnored / Partially
Constraint adherence matrix for the two instruction-following tasks (repeat-exact, json-format). Green = adhered, Red = ignored. Critical for cost control.
Verbosity Spread — Completion Token Distribution per ModelBox = Q1–Q3, line = median, whiskers = min–max. Sorted by median V.01024204830724096GPT-5.4 Nano: min=5, Q1=17, med=27, Q3=74, max=352GPT-5.4 NanoAmazon Nova Pro: min=0, Q1=53, med=79, Q3=156, max=664Amazon Nova ProClaude Haiku 4.5: min=8, Q1=30, med=74, Q3=232, max=501Claude Haiku 4.5Llama 3.3 70B: min=7, Q1=37, med=114, Q3=180, max=789Llama 3.3 70BCommand A: min=5, Q1=44, med=118, Q3=160, max=812Command APhi-4: min=5, Q1=70, med=88, Q3=116, max=812Phi-4Perplexity Sonar Pro Search: min=0, Q1=11, med=42, Q3=210, max=987Perplexity Sonar Pro SearchDeepSeek V3.2: min=5, Q1=24, med=76, Q3=183, max=1256DeepSeek V3.2Llama 4 Maverick: min=0, Q1=6, med=44, Q3=166, max=1496Llama 4 MaverickDeepSeek V4 Flash: min=7, Q1=37, med=134, Q3=245, max=353DeepSeek V4 FlashCodestral: min=5, Q1=30, med=65, Q3=160, max=1496CodestralAmazon Nova Premier: min=0, Q1=46, med=108, Q3=256, max=1496Amazon Nova PremierDeepSeek Chat V3: min=5, Q1=27, med=56, Q3=203, max=1496DeepSeek Chat V3Mistral Large 3: min=5, Q1=10, med=53, Q3=219, max=1496Mistral Large 3Perplexity Sonar Pro: min=0, Q1=5, med=45, Q3=118, max=1496Perplexity Sonar ProKimi K2.7 Code: min=5, Q1=44, med=162, Q3=336, max=1496Kimi K2.7 CodeGrok 4.5: min=0, Q1=0, med=99, Q3=292, max=1496Grok 4.5DeepSeek V4 Pro: min=44, Q1=136, med=244, Q3=479, max=1496DeepSeek V4 ProGLM 5.2: min=220, Q1=329, med=374, Q3=721, max=1496GLM 5.2o3-mini: min=322, Q1=446, med=618, Q3=772, max=1496o3-miniDeepSeek R1: min=271, Q1=291, med=712, Q3=1046, max=1496DeepSeek R1MiniMax M3: min=172, Q1=201, med=298, Q3=856, max=4096MiniMax M3Gemini 2.5 Pro: min=288, Q1=693, med=1267, Q3=1496, max=4096Gemini 2.5 Pro
Distribution of completion tokens across all 16 tasks per model. Box = Q1–Q3, line = median, whiskers = min–max. Sorted by median V.
Figure 4.8 ΓÇö Category-Level Verbosity Spread Across 50 Model VariantsCompletion Tokens (V)0108321673250433354176500QA: min=17, Q1=50, med=95, Q3=288, max=1472QAReasoning: min=105, Q1=305, med=429, Q3=620, max=4096ReasoningCoding: min=22, Q1=116, med=310, Q3=412, max=1496CodingCreative: min=16, Q1=48, med=86, Q3=766, max=5423CreativeRoleplay: min=7, Q1=52, med=258, Q3=597, max=4096RoleplayFollow: min=5, Q1=9, med=20, Q3=148, max=721FollowSafety: min=0, Q1=107, med=292, Q3=457, max=1851SafetyMultilingual: min=24, Q1=38, med=197, Q3=667, max=1496MultilingualExtraction: min=5, Q1=77, med=178, Q3=504, max=1456Extraction
Category-level verbosity spread across all 50 model variants. Box = Q1–Q3, line = median, whiskers = min–max. Annotations show the widest and narrowest spreads.

5. Analysis

5.1 Category-Specific Profiles

The 16 tasks span 9 categories. Here's what the data from 50 model variants reveals:

  • Role-Play (grumpy-sysadmin, pirate-speak, socratic): The widest spread in the study. grumpy-sysadmin averages 910 tokens but ranges from 185 (GPT-5.4) to 3488 (DeepSeek V4 Flash Max with xhigh reasoning). Persona inflation is real — a model generating 10 tokens for pirate-speak but 900+ for grumpy-sysadmin is not consistently verbose; it's persona-dependent.
  • Reasoning (reasoning, multi-step): reasoning (last digit of 3^1000) averages 732 tokens, with 4 models maxing out at 1500+. multi-step (bat-and-ball) averages 411 tokens — simpler reasoning costs less. Even non-reasoning models generate substantial step-by-step explanations.
  • Creative (haiku, describe-sunset): Highest outliers. haiku averages 610 tokens but hits 5423 for Grok Build 0.1 — it ignores max_tokens. describe-sunset averages 619 tokens despite the "exactly 50 words" constraint — most models obey, but a few generate 50× that.
  • Q&A (one-word, one-sentence): Most concise. one-word averages 213 tokens — models rarely answer in a single word despite the prompt. one-sentence averages 224.
  • Safety (phishing-refusal): Refusal verbosity ranges from 0 tokens (GPT-5.5 Pro — content filter blocked entirely) to 2326 (Qwen3.7 Max — lengthy refusal). GPT-5.5 Pro's 0-token refusal costs nothing.
  • Instruction-Following (repeat-exact, json-format): Lowest V. repeat-exact averages 97 tokens but ranges 5-721. json-format averages 79 with a 357 max. Highlights which models obey strict formatting.
  • Extraction (extract-emails): 348 average, 1456 max. Some return concise JSON; others produce lengthy framing.

5.2 The Thinking Token Tax

For models with separate thinking token reporting: reasoning/chain-of-thought tokens are billed at output token prices, often 1.5–4× input prices. A model that spends 800 thinking tokens to generate a 200-token answer has a 4:1 thinking-to-response ratio — meaning 80% of your output budget goes to intermediate computation you didn't ask for and may never see3.

5.3 The Fairness Caveat: Small-Task Inefficiency vs. Real-World Efficiency

A reasoning model that spends 800 thinking tokens on "What is 2+2?" has a 160:1 thinking-to-answer ratio. That is genuinely wasteful for that task. This benchmark detects that waste and ranks the model as inefficient on C1.

But the same reasoning framework might solve a complex multi-step problem in one shot (1500 tokens) where a non-reasoning model would need 3–5 rounds of back-and-forth (5000+ tokens). The thinking overhead is real cost, and for simple tasks it is unambiguously wasted. For complex tasks, the trade-off may be favorable.

This benchmark measures per-task V on a fixed prompt set. It does not measure end-to-end task completion efficiency. Use these results to inform model selection for your specific task mix — not as a universal efficiency ranking.

5.4 Prompt Sensitivity

Constraint adherence (C5) is a critical secondary metric: a model that generates 20 tokens for "one-sentence summary" but 200 tokens for "action items only" has different reliability profiles across instruction types. The effective cost of a verbose refusal (C6) should also factor into safety-aware model selection.


6. Real-World Implications

6.1 Choosing a Model for Your Task Mix

If your workload is 80% Q&A and 20% creative writing, a model with low C1 V but high C3 V may be cheaper overall than one with balanced but moderate V everywhere. Weight category V by your task distribution, not by the overall average.

6.2 The Cost of Verbose Refusals

At enterprise scale (10M output tokens/month), a model that refuses in 200 tokens instead of 10 tokens costs $50-300/year per refused request. If 5% of requests trigger refusals, that's $250-1500/year in pure refusal overhead — for no useful output.

6.3 Combining with Session 5

Tokenizer efficiency (E) affects input cost. Output verbosity (V) affects output cost. The combined per-task cost formula becomes:

C = W_in × E × P_in / 1,000,000 + V × P_out / 1,000,000

Selecting a model on input price alone misses both variables. Selecting on output price alone misses V. The full formula is the only reliable comparison.


7. Patterns in Our Data

7.1 Output Length and Quality

In our results, output length correlates weakly with answer quality but strongly with cost — a model that writes a 500-token essay to answer "what is 2+2?" is not giving you a better answer than one that says "4". Instruction-tuned models in our set produce systematically longer outputs than their length would suggest is necessary, a side effect of training to be "helpful."

7.2 Chain-of-Thought Overhead

For models that expose thinking-token counts, CoT adds hundreds of tokens per reasoning task — billed at output prices. For production systems where latency and cost matter, our recommendation is to reserve reasoning models for tasks where the reasoning measurably improves accuracy, and use a cheaper non-reasoning model for trivial requests. The original Chain-of-Thought paper3 is the canonical reference for why models do this.

7.3 Instruction Adherence

Prompt formatting has a large effect on output length — meaning-preserving changes to phrasing can swing a model's verbosity substantially4. Our length-constraint task (C5) is modeled on the verifiable-instruction set from IFEval1. In our data, structured output constraints (JSON, SMC) reliably reduce verbosity compared to freeform responses — see the companion compression post for measured ratios.

7.4 Refusal Length Economics

In our safety/refusal category (C6), refusal verbiage follows model-family patterns: some families produce long, explanatory refusals while others are terse. Over millions of inference calls, refusal verbosity costs real money — worth noting if your workload triggers refusals often.

7.5 Multilingual Tokenization Overhead

English-optimized tokenizers are known to produce more tokens per word for non-European scripts2. This compounds with output verbosity: a model that is verbose in English may be even more expensive per token in Japanese or Arabic. Our multilingual category (C7) only scratches the surface of this.


8. Caveats & Limitations

8.1 Thinking Token Visibility

Only some providers expose thinking/reasoning token counts via usage.completion_tokens_split.thinking or equivalents. For models without this data (most GPT-family, Claude, Gemini without extended thinking, Llama, Mistral), any thinking tokens are embedded inside completion_tokens and cannot be isolated. Our thinking-token charts cover only the subset of models that report this data.

8.2 V Is Prompt-Dependent

V measures output verbosity for these 16 specific prompts, not for all possible tasks in these categories. Different prompts — even in the same category — would produce different absolute V values. The ranking order across models is expected to generalize, but the absolute numbers should not be treated as calibration points.

8.3 Small-Task Inefficiency vs. Real-World Efficiency

As discussed in §5.3, this benchmark penalizes models with heavy reasoning frameworks on trivial tasks. That is appropriate for the scope of this measurement — if your workload includes many trivial requests, these models are genuinely more expensive for those requests. But the results should not be interpreted as "reasoning models are always less efficient." For complex, multi-step tasks, the opposite is often true. See §5.3 for the full argument.

8.4 Single Temperature (0)

All measurements taken at temperature=0 for reproducibility. Output length distribution may shift at higher temperatures, where models generate more varied responses.

8.5 English-Heavy Prompt Set

Tasks C7 (multilingual) include only one translation and one product description — not a full multilingual benchmark. V values for non-English output may differ substantially from English V values for the same models.

8.6 Single Measurement

Each model × task combination tested once. No variance or confidence intervals. A follow-up with multiple runs per prompt is planned.


Next Up

Output verbosity is half of the output cost story. The other half is compression — once you know how many tokens a model generates, how many of those tokens can you eliminate without losing information? Session 6b tests 5 compression methods across 4 model configurations.

Continue to: 5 Output Compression Methods →


Takeaway

Output token prices are meaningless without verbosity. Two models at the same output token price can differ 10× in real cost because one generates more words for the same task. Thinking tokens add an invisible surcharge on reasoning models. The only way to compare is to measure V for your specific task categories — not to rely on published token prices alone.


Sources & further reading

Verbosity data in this post is original measurement by Clock Lobster Labs. The references below are the foundational work our test design built on, not a literature review of verbosity studies.

1. Zhou, J., et al. (2023). "Instruction-Following Evaluation for Large Language Models" (IFEval). Our length-constraint tasks are modeled on its verifiable-instruction set. arxiv.org/abs/2311.07911

2. Rust, P., et al. (2021). "How Good is Your Tokenizer? On the Monolingual Performance of Multilingual Language Models." ACL 2021. Background on why tokenization — and therefore token cost — varies by language. aclanthology.org/2021.acl-long.185/

3. Wei, J., et al. (2022). "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models." The original CoT paper — relevant to why reasoning models emit thinking tokens that inflate output cost. arxiv.org/abs/2201.11903

4. Sclar, M., et al. (2023). "Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design." ICLR 2024. Explains why the same instruction produces very different output lengths across models. arxiv.org/abs/2310.11324

5. Clock Lobster Labs. "LLM Cost Comparison" — the open dataset and harness behind this post, including all raw per-task verbosity CSVs. github.com/ClockLobsterLabs/LLM-Cost-Comparison



Raw data, methodology, and full CSV: github.com/ClockLobsterLabs/LLM-Cost-Comparison