Apple made it free. Kimi made it cheap. Neither one knows who you are.
Two things shipped this month from opposite ends of the market. Almost everyone read them as the same story. They are the opposite story, and the difference is the whole argument.
Start with the facts, because they are better than the takes.
From below: Apple now ships a capable model inside the operating system, on the phone in your pocket and the Mac on your desk. No key. No subscription. No bill. It runs on hardware you already bought, and it will run while you are asleep.
From above: Moonshot released Kimi K3 — roughly 2.8 trillion parameters, open weights, frontier-class results, at $3 per million tokens in and $15 out. Roughly a third of what the top closed models charge.
Free at the bottom. Cheap and un-gatekeepable at the top. Read together, that looks like one trend: intelligence is getting commoditized. True, but useless. The interesting part is that these two facts point in opposite directions, and only one of them is about ownership.
The distinction almost everyone gets wrong
Kimi K3 is open-weight. It will also never run on your phone. Two-point-eight trillion parameters is a datacenter, not a laptop — the clever sparse routing cuts the compute per token, not the weights that have to sit somewhere. “Open” here is a licensing fact. It is not a sovereignty fact.
So K3 is not evidence that you can own the frontier. It is evidence that the rent is collapsing — and that no single company can hold the meter, because you cannot embargo weights that are already downloaded.
Apple’s model is the reverse. It is small. It will never write your hardest paragraph. What it will do — read, sort, derive, connect, compile — is most of the work most of the time, on your own silicon, at zero marginal cost. That is not the frontier collapsing. That is the floor rising into your pocket.
A pincer on the middle
Put the two together and something specific gets crushed. Not the frontier labs — they will keep doing the genuinely hardest work, and they should be paid for it. What gets crushed is the thing most people are actually paying for right now: a monthly subscription to a model that holds your memory hostage inside its own walls.
That product only made sense when the model was scarce. From below, the routine work no longer needs it. From above, the hard work no longer costs what it did. The subscription was never selling intelligence. It was selling continuity — the fact that it remembered you and the alternatives did not.
So what is actually scarce?
Not the model. Two companies just proved from both directions that the model is a commodity. Not the compute; that is a metered utility, and the meter is falling.
The scarce thing is the one asset in the stack that appreciates: the record of how you think. The problems you keep returning to. The decision you already made and the reason you made it. The thread you set down in March and never picked back up.
And here is the tell, the thing that makes this an argument instead of a slogan: neither tier can supply it.
Apple’s assistant builds a memory of you — on your device, which is not the same as in your hands. You cannot open it, read it, correct it, or carry it to another model. It is private and closed. And it indexes your apps: your messages, your mail, your photos. Not the hours you spend actually thinking.
Kimi does not know you exist. It is a magnificent stateless engine you rent by the token — brilliant and completely amnesiac, every single time.
Which is the entire architecture
chanio is not a clever idea about where files should live. It is the shape the market just made obvious:
Your memory lives on your disk. Plain files, in a vault you own, that you can read without us and keep after us.
The nightly work runs on your own device — Apple’s on-device model reads what you keep returning to, derives what you actually care about, and compiles a morning brief while you sleep. Nothing leaves the machine, and it costs nothing, because you already paid for the silicon.
The rare hard step gets rented — the one question a small model should not be trusted with routes out to whatever is best and cheapest that week. Today that might be Kimi at a third of the price. Next quarter it will be something else with a different name. You do not care. Not caring is the entire point of renting.
The model is rented. The mind is yours.
Why this gets stronger, not weaker
Most products are threatened by the model getting better. This one is not, and the asymmetry is worth sitting with. Every quarter, the rented tier gets cheaper and the on-device tier gets more capable. Both moves make the model less of the product and the memory more of it. The thing we are building is the residue of the trend, not a bet against it.
Memory grows. Context shrinks. Cost falls.
The honest part
Three things cut against the comfortable version of this story, and you should have them.
Small models are weakest exactly where it matters. Stanford’s MedHELM evaluation found the leading models score highest on note-writing and communication and lowest on genuine decision support. The machine is strongest where the stakes are lowest. So no, on-device does not do everything — and anyone telling you it does is selling. That weakness is precisely why the loop has two tiers instead of one.
Benchmarks are not capability. Practitioners putting K3 on real codebases report it acing the flashy demos and failing an ordinary debugging task that the closed models one-shotted. Leaderboards measure the thing they measure.
The rent gets cheaper; it does not go to zero. K3 at $3/$15 is cheap against the closed frontier and a real step up from what Chinese models used to cost. The floor is rising even as the ceiling falls. “Intelligence becomes free” is the wrong shape. “Intelligence becomes a utility with a real bill” is the right one — which is exactly why you meter the judgment and own the memory.
What you own at the end
The model you are renting today will be obsolete inside a year. That is not a risk; it is the schedule. The question is what happens to you on that day.
If your memory lives inside the model, you migrate — and you have already read the guides where people hand-move a decade of their own thinking in a hundred-megabyte file. If your memory lives on your disk, you point the new engine at the vault and go back to work.
Everything in this stack is depreciating except the record of your own thinking. Build like that is true, because it is.
Start owning the part that appreciates.
See the whole loop run on Apple’s own on-device model — or bring your existing AI history home first. No account, no key.