SHARECASTERGET THE APP

SHARED EPISODE

Cover art for Kimi K3: The Sticky Note Problem

Kimi K3: The Sticky Note Problem

· 11:44 · 1 PLAY

0:0011:44

A 2.8 trillion parameter model went free on the internet, and the scary part isn't its size, it's the attention trick that makes it cheap to run.

TRANSCRIPT

Okay so... um. There's a file.

It went up on the internet at the end of July. And it's, uh... it's big. Like, embarrassingly big. Nobody publishes the exact download size in a way I trust, so I'm not gonna make one up, but... it's the kind of file where your laptop just, hmm, sort of laughs at you? And it's free. Anyone can grab it. Right now. You, me, a bank, a teenager, the government of, I dunno, somewhere.

And the thing inside that file, um... it's called Kimi K3. It's made by a company in Beijing called Moonshot AI. And it is, according to them, two point eight trillion parameters. Which is, uh... okay, that's the biggest openly released model anybody's put out. They call it the world's first open three-T-class model. Three-T meaning trillion. They rounded up. Marketing, y'know.

And here's the part I actually wanna talk about, because everyone gets this wrong, including me, for like, uh, two weeks. The reason this thing is scary to the big American labs... it's not that it's huge. Honestly? Huge is the boring part. Huge is easy. You just, mm, spend money.

The scary part is that they figured out how to make it cheap to run. And that's, um. Yeah. That's the whole episode, basically. Sorry. I probably should've saved that.

Okay so let me back up, because I skipped like four things.

Moonshot AI. Beijing startup. The founder, uh, the A P reported he's a Pink Floyd guy who got his doctorate in Pittsburgh, which... I love that detail and I have no idea what to do with it. It just, y'know, it lives in my head now. Anyway. They've been shipping these Kimi models for a while. K2, the previous big one, that was a trillion parameters total but only about thirty-two billion actually switched on for any given word. That's a mixture-of-experts thing, we'll, uh, we'll come back to that maybe. Or not.

And then K3 lands. Two point eight trillion. Native vision, so it sees images as part of how it thinks, not bolted on after. A one-million-token context window, which means you can hand it, like... a whole codebase? A stack of legal filings? A novel and its sequels and the fan fiction?

And when it came out, it went to the top of the leaderboard on one specific thing. Front-end coding. Like, building the visible part of a website. That's the Arena leaderboard, where people vote blind on which answer is better. And the guy who runs that leaderboard posted that more results were still coming in and it'd probably stay near the top. So, uh. Grain of salt on that? It's one category, and leaderboards are, hmm, they're vibes with error bars. But still. A free model from a startup, sitting above the paid ones. That's new-ish.

Small honest wobble here, 'cause I said I'd flag this stuff. The dates I can find are a little smeared. There's press coverage from the seventeenth of July, and the public code repository says the twenty-seventh. Could be an announcement then a release, could be a preview, could be somebody's timestamp being weird. I don't know. Um. I'm not going to pretend I know. It's July, 2026. That's the resolution I've got.

Right. So. The actual question.

If you're OpenAI, or Anthropic, or Google... why do you care? Big Chinese model tops one chart. Charts move. You've got distribution, you've got enterprise contracts, you've got, uh, the brand. Somebody publishing weights doesn't take your customers, because, and this is the part people forget, nobody can run a two point eight trillion parameter model. Not at home. Not in a garage. That needs a rack of very expensive machines making a noise like a jet taxiing.

So "open" here doesn't mean what it means for, like, a photo editor you download. It's open the way, um... okay, it's open the way the blueprints for an aircraft carrier being public is open. Congratulations. You still don't have an aircraft carrier.

So that should be fine, right? For the incumbents? Hmm. No. And to explain why not, I have to tell you about the sticky notes.

Okay. Sticky notes.

The way these models work, when they're writing you an answer, they don't just... write. For every single word they produce, they look back at everything in the conversation so far. Every word you typed, every word they already wrote. All of it. Every time.

And to do that fast, they keep a sort of, uh, scratch memory of all that stuff. The jargon is the K V cache. Key-value cache. I like to think of it as sticky notes. Every word that goes past, the model slaps a couple of sticky notes on the desk about it. And then to write the next word, it reads... the whole desk.

You can see the problem. Ten words, fine, cute little desk. But a million tokens of context? The desk is now the size of a parking lot, and the model is walking the parking lot for every single word it emits. That's why long conversations get slow and expensive. That's, uh... that's basically it. That's the bill. The bill is memory, and the bill is the walking.

So the whole game, the actual quiet engineering game underneath all the hype, is: can you stop rereading the entire parking lot?

And Moonshot's answer is a thing called Kimi Delta Attention. K D A. Which they published on before K3, in work they called Kimi Linear, so we can actually look at what it does instead of guessing.

And it's, um... okay it's a version of what's called linear attention, which, the very rough idea is, instead of keeping every sticky note forever, you keep a running summary. A little state. And you update it as you go. Which is way cheaper. But historically it was also, y'know... dumber. You lose stuff. Summaries lose stuff, that's what summaries are.

So the trick, the delta rule part, is that when a new fact arrives, you don't just pile it on. You sort of, mm, subtract the old wrong version and write in the new one. Delta. Change. Like, if I say "my dog's name is Rex" and then later "sorry, actually it's Bruno," you want the memory to overwrite, not to hold both and get confused. Which, uh, older linear attention absolutely got confused.

And Moonshot's addition, the bit they're proud of, is fine-grained gating. Which means, instead of one dial that says "forget a little" for the entire memory... every channel of the memory gets its own dial. So it can forget the color of the room while keeping the name of the person. Selectively. Per, uh, per little slot.

And then, the part I think is actually the wise part... they didn't go all in. They kept some of the old expensive full attention. Three layers of the cheap kind, then one layer of the real thing. Three to one. So the model gets a periodic, um, a periodic look at the actual desk. It's a hybrid. It's a compromise. Engineers, man. Always the compromise.

And the numbers they report from that work: up to seventy-five percent less of that sticky-note memory, and up to six times faster generation at the million-token length.

Okay, receipts on that, because I want to be careful. Those numbers are from the Kimi Linear work, their own benchmarks, their own comparisons. Up to means up to, which is doing a lot of load-bearing work in most sentences it appears in. And K3 says it's built on K D A plus another thing they call Attention Residuals, which, um... I've read the description, and I do not want to explain something I only half understand, so I'm gonna say: I don't know exactly what that one buys them. It's in the announcement. It's part of the recipe. That's as far as I'll go.

But the direction is not ambiguous. They made the memory smaller and the output faster, and then they made the biggest open model anybody's made, and they gave it away.

So. Back to the question. Why does that hurt a big lab?

Because... hmm. Okay. Think about what you're actually buying when you pay for a frontier model. You think you're buying intelligence. You're not, really. You're buying somebody else's data center, amortized. The moat was never the file. The moat was that serving the file at scale, fast, cheap, without setting money on fire, is genuinely hard.

And when the weights are public and the architecture is specifically designed to be cheap to serve... anybody with GPUs can host it. Cloud companies. Inference startups. A bank that doesn't want its documents leaving the building. A government that, uh, would rather not have its stuff going through an American company. And they compete with each other. On price. Immediately. Because they're all selling the exact same model, and there is no brand loyalty in the, y'know, the business of matrix multiplication.

That's a price floor collapsing. And here's the thing about a price floor collapsing: it doesn't have to be better than you. It just has to be good enough, at a price that makes your price look insane. That's the whole DeepSeek shock from last year, except, um, except it's not a one-off anymore. Zhipu put out G L M five point two the month before. Moonshot put out this. It's a cadence now. It's a conveyor belt.

And I'll put my actual opinion on the table, since I've been dancing around it. I think this is a deliberate strategy and I think it's a good one. There's an old, uh, an old business idea, commoditize your complement. If you make money on the thing next to the product, you want the product itself to be free. Give away the razor, sell the blades. Or here: give away the model, and what you're really doing is torching the margins of the people whose entire business is charging for the model. You don't have to win. You just have to make winning worthless.

Is that the plan? I can't read minds and I'm not going to pretend the strategy memo leaked to me. But I'll note that K3 came out right before China's big A I conference in Shanghai, right around Xi Jinping's opening address, and the A P pointed out that timing was, quote unquote, not likely a coincidence. Which, uh. Yeah. Models ship when models ship. Except when they ship on a Thursday for a reason.

Okay. Fairness break. Because I don't want to be the guy who's just, mm, breathlessly hyped at a spreadsheet.

The counterargument is real, and it's roughly this. One, leaderboards are not the job. Topping front-end coding is lovely and it is not the same as being the model a company bets its accounting on. Two, the incumbents have things that don't show up in benchmarks at all, like, uh, reliability, and support contracts, and the ability to call someone at three in the morning. Three, "open weights" is not the same as open source. K3 is released under Moonshot's own license, called the Kimi K3 License, which is a thing that exists specifically because it isn't one of the normal ones. I have not read every line of it and I'm not gonna characterize terms I haven't checked. But when a company writes its own license, it's usually 'cause it wants something the standard ones don't give it.

And four, the one nobody likes saying out loud: nobody outside these companies really knows what any of this costs to train. The claimed efficiency, the claimed budgets, all of it... it's self-reported. Every time. From everyone. Including the American labs. So, uh. Yeah. Grain of salt, industrial size.

But. Here's the frame I keep landing on.

Every technology has a moment where the expensive miracle becomes the boring input. Steel had it. Databases had it. Bandwidth definitely had it. There's always this stretch where the thing is magic and a few people own the magic, and then... hmm. Then somebody publishes how, and the magic becomes plumbing. And plumbing is a worse business. Plumbing is a much, much worse business. But plumbing is what actually gets built into everything.

And the tell, the tell that you're in that moment, is never the flashy demo. It's the boring engineering paper about memory management. It's some person, somewhere, deciding that each channel of a cache should have its own forget gate. That's it. That's the sentence that moves the money.

Which, uh, brings me back to the file.

Somebody downloaded it the night it went up. Probably a lot of somebodies. Sitting there at two in the morning watching a progress bar crawl, filling a drive with the distilled result of, y'know, an enormous amount of electricity and a couple hundred people's year. And nothing stopped them. No sales call, no contract, no waitlist, no "please describe your use case." Just... a download.

That's the shift. Not that a Chinese lab caught up. That the catching-up is now something you can put on a hard drive and keep.

Whether that's the beginning of open models eating the frontier, or just the loudest moment before the big labs pull away again... um. I genuinely don't know. Anyone who tells you they do is selling something. Possibly compute.

But I'd say this. The next time you hear somebody say the moat is the model... ask them how much the desk costs. And if they don't know what you mean by the desk, uh. Well. Then you know more than they do. Which is, hmm. That's a nice feeling, actually.

There's links in the show notes if you wanna go read the actual technical stuff. Some of it's above my head too. That's, y'know. That's kind of the fun part.

LINKS TO CHECK OUT

GET SHARECASTER

Like what you hear? Make your own episode — one sentence is all it takes.

GET THE APP