One number in → full cost breakdown across every major frontier API, priced against your local server. Includes input, cached-read, cache-write, and output tiers.
$0.02 · 8×4090 running 70B ≈ $0.10–0.30 · datacenter API-style ≈ $0.50+. Output-heavy workloads cost more per token — model it by splitting the number if needed.cost = (in×pricein + read×pricecache + write×pricewrite + out×priceout) ÷ 1M| Model | In $/M | Cache $/M | Out $/M | Your bill /mo · /day ⇅ | Eff. $/M | Saved vs local |
|---|
MODELS array at the bottom of this file to update.
Long-context surcharges modeled where published: OpenAI ≥272K context ⇒ input & cache rates ×2, output ×1.5
(official long-context columns). Gemini Pro ≥200K estimated at the same pattern — only one tier is published there,
verify if it matters to you. All other providers price flat across context size.