THE METHODOLOGY DEPARTMENT
HOW THIS SITE KNOWS
One brand, two tones. The editorial pages are dry and ironic. This page is neutral and precise — it is the record of how every number on the site is produced, labeled and verified.
Classes of evidence
Every surface on this site falls into one of four classes, and the name tells you which one. The word Index is reserved for artifacts that are calculated or based on evidence — never for editorial commentary.
EDITORIAL / EDUCATIONALRankings and verdicts written by a human to illustrate a pattern. Entertaining on purpose. Never presents itself as measured data.
EVIDENCE-BASEDScores derived from a public, versioned rubric with cited primary evidence. Undergoes RFC → freeze → scoring → right of reply before publication.
CALCULATED / REPRODUCIBLECost of a canonical workload computed from verified list prices and declared assumptions. Same input, same output — reproducible.
MEASURED / SAMPLE DISCLOSEDOnly ever derived from real data, with the sample size, cohort and method published. Never invented to keep the page looking alive.
STATUS — Provider Cost Transparency Index: method RFC open. The rubric (5 blocks, 25 binary criteria, ~10 providers) is being reviewed until 2026-08-26. No scores are published until the methodology freezes. A provider can be expensive and score well on transparency; the two axes are unrelated. Zero fabricated scores, by design.
How estimates are made
Costs at list prices are computed as (input × inputPerMTok + output × outputPerMTok) / 1M against the verified public prices on the pricing board. Where a usage file reports cached reads, they are billed at ~10% of the input list price (the prompt-caching convention documented in our prompt-caching guide). Everything is an estimate at list prices — never your invoice, never your negotiated rate.
The burn score and compression in the calculator are heuristics over text (filler, repeated 4-grams, ornament, prompt size vs a baseline). They judge verbosity, not correctness. The compressed prompt is labeled “heuristic — review before use”.
Evidence levels for waste
“Token burn” is spend that did not contribute to the accepted result. We never publish a single universal efficiency score, and we never compare burn between organizations that use different baselines. Instead:
Observational signals from an exported usage file: cache-miss share, output/input ratio, model mix, retries where the file supports it. We do not claim causality a CSV cannot prove.
Attributed causes when telemetry is present and lets us prove what consumed the tokens.
A versioned baseline rerun: burn rate = (observed cost − baseline cost) / observed cost. Requires named baselines and paired runs. Reserved for future paid audits, not built today.
Sources and versioning
Prices are checked daily against public list prices by an automated watchdog (pricing-watch), with sanity guards (a >3× change opens a review instead of applying). Each change is recorded with its date, direction, and before/after in the Price Changes history. The versioning exists so that, months from now, this site can answer: “what did Claude, GPT and Gemini actually cost for workload W in August 2026?” — the historical series is an asset, not a byproduct. Board last changed: 2026-08-04.
Data policy
What we collect: aggregated pageviews and events via PostHog (EU-hosted), used only to understand what people actually use. Newsletter signups send an email to our own list and nothing else.
What we do not collect: the CSV you audit in Audit My Month never leaves your browser — the whole analysis runs locally. Tools never upload prompts, exports or identifiers. No personal data is sold or shared.
Opt-out: append ?tbx-optout=1 to any page URL to disable analytics in your browser (revert with ?tbx-optin=1). Aggregated stats are never reported with sample sizes too small to be meaningful (see Privacy).
Conflicts of interest
TokenBurn Index sells no tokens, accepts no affiliate commissions and takes no money from model providers. Pricing comes from public list prices, not vendor relationships. If a paid audit service ever exists, every piece of data it produces will carry the evidence class and the cohort behind it. When an article has any relationship with a vendor, it is disclosed in the article.
Every feature is judged against one question: does it strengthen the historical data asset, methodological credibility, current useful information, organic discovery, shareability, or repeat usage? A tool that only exists to be a tool is deferred.
Limitations
Estimations are not invoices. Token counts use the o200k tokenizer — exact for GPT models, ~±10% elsewhere. Negotiated rates, discounts and free tiers differ from list prices. Energy disclosure is only ever scored as “is clear information available?”, never as an estimated carbon figure — we do not have a serious methodology for that, so we do not fake one.
Version: methodology v0.1 · Applies to surfaces published from 2026-08-12 onward.