Free tools · Cost calculator

What does a month of AI cost, on-device vs. the API?

Move the sliders. The first number is your month with chanio: everyday questions on your own machine at $0, the hard ones held and sent out once a night. The second is the same month with every question at API prices.

Questions per day30
Tokens in per question (question + the notes it reads)1,500
Tokens out per answer300
Share you’d hold for a frontier model10%
Frontier model for the nightly batch
YOUR MONTH · 1,620,000 TOKENS
$0.54
with chanio: 90% on-device at $0 + 10% held → Claude Sonnet 5
$5.40
every question to Claude Sonnet 5
Claude Sonnet 5 · Anthropic · $2 in / $10 out per 1M$5.40 all · $0.54 held only
Claude Opus 5 · Anthropic · $5 in / $25 out per 1M$13.50 all · $1.35 held only
Claude Haiku 4.5 · Anthropic · $1 in / $5 out per 1M$2.70 all · $0.27 held only
GPT-5.6 Sol · OpenAI · $5 in / $30 out per 1M$14.85 all · $1.49 held only
GPT-5.6 Terra · OpenAI · $2 in / $12 out per 1M$5.94 all · $0.59 held only
GPT-5.6 Luna · OpenAI · $0.2 in / $1.2 out per 1M$0.59 all · $0.06 held only
Gemini 3.6 Flash · Google · $0.375 in / $1.875 out per 1M · introductory price through Dec 31, 2026; then $0.75 / $3.75$1.01 all · $0.10 held only

On-device answers cost $0 in tokens; they cost electricity and a one-time model download. This ignores subscriptions (a $20/mo chat plan is $240 a year regardless of use), caching discounts, and batch-API discounts, which would lower the API columns further. List prices as published in September 2026: Anthropic, OpenAI, Google. Prices move; check the source before you budget.

The meter is the product

This calculator is the same arithmetic chanio does for real, every day, in the meter under every answer: which tier ran, how many tokens, what it cost. The point is not that the API is expensive. It is that most questions never needed it, and you should be able to see which ones did. The founder’s own meter is public.

Questions people ask

How is the cost computed?
Questions per day × 30 days × (input tokens × input price + output tokens × output price) ÷ 1,000,000. The chanio number applies that only to the share you hold for a frontier model; the rest runs on your device at $0 in tokens.
What is a token?
Roughly three quarters of an English word. A typical question with the notes it needs to read is 1,000 to 3,000 tokens in; an answer is 100 to 500 tokens out. Reasoning models bill hidden thinking tokens as output, which this calculator does not model.
Is on-device really $0?
In tokens, yes: nothing is billed. You pay for the electricity, the one-time model download, and a Mac with enough memory. A 3B model on Apple silicon answers a short question in a few seconds.
Why hold anything for a frontier model at all?
Because a small local model is weaker at hard reasoning, current events, and long generation. chanio's honest answer is to route those few questions out, in one nightly batch, and show you the cost, rather than pretend the local model is the frontier.
Are these prices current?
They are the list prices on the providers' pricing pages in September 2026, linked under the table. They change. Google's Gemini 3.6 Flash price is introductory through December 31, 2026. Check the source before you budget.

Related

Which local model fits your Mac?
Pick a model by memory, with the honest trade-offs.
Set it up on your Mac
One line installs the starter and the meter.
Numbers, in the open
Last night's run and this site's visits, counted first-party.