Kimi K3 Is Here: The First Open 3T-Class AI Model, and What It Means for Your Small Business
On July 16, 2026, Moonshot AI released Kimi K3 and called it the world's first open 3T-class model. It is a 2.8-trillion-parameter system with a 1-million-token context window, it matches or beats Claude Fable 5 and GPT 5.6 Sol on a handful of serious benchmarks, and by July 27, 2026 anyone will be able to download the weights and run it themselves.
That last part is the story. Not the benchmarks.

Key Takeaways:
- It is genuinely frontier-class. 2.8T parameters, a 1M-token context window, and native vision, using a Stable LatentMoE design that activates only 16 of 896 experts per pass.
- It wins some and loses some. K3 beats both Claude Fable 5 and GPT 5.6 Sol on SWE Marathon and BrowseComp, and loses badly to Fable 5 on HLE-Full. Moonshot admits K3 still trails both overall.
- The weights are the news. A model this capable being downloadable means, for the first time, frontier-class AI that no vendor can price up, deprecate, or have restricted out from under you.
- The pricing is mid, not magic. $3.00 per million input tokens and $15.00 per million output tokens on the Kimi API. Cheaper than top-tier proprietary, not free.
- You still should not switch. Your automations are not underpowered. Most of them are unbuilt. That is a different problem, and no new model fixes it.
What Moonshot Actually Shipped
Kimi K3 is built on what Moonshot calls Kimi Delta Attention and Attention Residuals, with a Stable LatentMoE architecture that activates 16 out of 896 experts on any given pass. In plain English: it is enormous, but only a small slice of it fires at once, which is how you get frontier-scale capability without frontier-scale cost per query.
The context window is 1 million tokens. That is roughly a mid-sized company's entire process documentation, or a full codebase, held in the model's head at once. Moonshot's pitch is long-horizon work: it says K3 "can sustain long engineering sessions, navigate massive repositories, and orchestrate terminal tools." The demos back the ambition. In one 48-hour autonomous run, Moonshot reports K3 built, optimized and verified a chip design using open-source EDA tools. In another, it synthesized research by pulling data through 2,800+ web searches and 1,100+ terminal calls.
It is available now through Kimi.com, Kimi Work, Kimi Code and the Kimi API. The full weights are scheduled to land by July 27, 2026.
The Benchmarks, Honestly
Every model launch cherry-picks. Here is the unflattering version alongside the flattering one.
Where Kimi K3 beats both Claude Fable 5 and GPT 5.6 Sol:
- SWE Marathon: 42.0, against 35.0 for Fable 5 and 39.0 for GPT 5.6 Sol. This is the long-running software engineering test, and it is the one that best matches Moonshot's long-horizon claim.
- BrowseComp: 91.2, against 88.0 and 90.4. Real web research, done autonomously.
- Program Bench: 77.8, against 76.8 and 77.6. A narrow win, but a win.
- Automation Bench: 30.8, against 29.1 and 29.7.
- OmniDocBench: 91.1, against 89.8 and 85.8.
Where it loses:
- HLE-Full: 43.5, against 53.3 for Fable 5. That is not a rounding error, that is a gap.
- FrontierSWE: 81.2, against 86.6 for Fable 5.
- Kimi Code Bench 2.0: 72.9, against 76.9 for Fable 5. Moonshot loses on its own coding benchmark.
- GDPval-AA v2: 1668, against 1760 and 1748. This is the agentic economic-work eval, and K3 sits third.
Moonshot says it plainly: K3's overall performance "still trails the most powerful proprietary models, Claude Fable 5 and GPT 5.6 Sol." Credit where it is due. That is a more honest launch post than most.
So the summary is: Kimi K3 is roughly frontier-adjacent. It trades punches. It does not win the fight.
Why the Open Weights Matter More Than the Scores
Here is the part worth your attention, and it has nothing to do with a leaderboard.
Three weeks ago we wrote about Claude Fable and GPT-5.6 getting restricted in the US. The lesson there was that models you rent can vanish on someone else's decision. A policy directive, a pricing change, an acquisition, a deprecation notice. If your business logic is welded to one specific model from one specific vendor, all of that is your problem.
An open 3T-class model changes the shape of that risk permanently. When the weights are public, the model cannot be un-released. It cannot be priced up. It cannot be export-controlled away from you after you have built on it. It sits on disk. It runs on hardware you choose.
That is the first time frontier-adjacent capability has been available on those terms. It is a genuinely big deal for anyone who cares about not being a tenant in their own operations.
It is also, for most small businesses, still theoretical. Running a 2.8T-parameter model is not a thing you do on the office laptop. What actually reaches you is second-order: a credible open model at this tier puts a ceiling on what everyone else can charge, and gives the tools you already use a fallback engine that no one can switch off. You benefit from the competition without ever touching the weights.
The Pricing Is Good, Not Revolutionary
On the Kimi API, K3 runs $0.30 per million tokens for cache-hit input, $3.00 per million for cache-miss input, and $15.00 per million output tokens.
That is meaningfully cheaper than the top proprietary tier, and it is not free. Anyone telling you a new model just made AI cost nothing has not read the price list. The cache-hit rate of $0.30 is the interesting number, because well-built automations reuse the same context constantly, and that is where the real bill lives.
If you want the grounded version of what AI actually costs a small business per month, we broke it down in how much AI automation costs. The short version has not changed because of today's launch.
What This Changes for Your Business This Week
Nothing. And that is the useful answer.
Here is the uncomfortable truth about model launches: the businesses getting real returns from AI are not the ones running the smartest model. They are the ones who bothered to build the boring thing. Lead follow-up that fires in ninety seconds instead of two days. A support inbox that answers the same eleven questions without a human. Data that moves between two apps without anyone retyping it.
None of that was blocked by model intelligence in March, and none of it is unblocked by Kimi K3 in July. Those workflows ran fine on models two generations old. They are not underpowered. They are unbuilt. A 2.8T-parameter model does not fix an automation that does not exist.
What K3 does change is the ceiling and the floor. The ceiling on what your vendors can charge, because there is now credible open competition at the top. And the floor under your automations, because "what if this model goes away" now has a real answer.
So the move is the same one it has always been. Build workflows that do a job rather than workflows that worship a model. Keep the engine swappable. Then when K3, or K4, or whatever lands in October turns out to be cheaper or better at your specific task, you change one line and keep going, instead of rebuilding for three weeks.
That is exactly how we build. If you want to see where to start, our guides on AI agents for small business and how to automate your business with AI are the practical next reads.
Frequently Asked Questions
What is Kimi K3?
Kimi K3 is an AI model released by Moonshot AI on July 16, 2026. It is a 2.8T-parameter model with a 1-million-token context window and native vision, and Moonshot calls it the world's first open 3T-class model. The full weights are scheduled for public release by July 27, 2026, which means anyone can download and run it rather than only renting it through an API.
Is Kimi K3 better than Claude or ChatGPT?
On some benchmarks, yes. Kimi K3 scores 42.0 on SWE Marathon against 35.0 for Claude Fable 5 and 39.0 for GPT 5.6 Sol, and 91.2 on BrowseComp against 88.0 and 90.4. On others it clearly loses, including HLE-Full where it scores 43.5 against Fable 5's 53.3. Moonshot itself says K3's overall performance still trails Claude Fable 5 and GPT 5.6 Sol. It is competitive, not dominant.
How much does Kimi K3 cost?
The Kimi API prices K3 at $0.30 per million tokens for cache-hit input, $3.00 per million tokens for cache-miss input, and $15.00 per million tokens of output. Because the weights are being released openly, you can also self-host it and pay only for the compute you run it on.
Should a small business switch to Kimi K3?
Almost certainly not, and not yet. The automations that actually make small businesses money, such as lead follow-up, customer support and data entry, are not limited by model intelligence. They are limited by whether anyone has built them. The right move is to build workflows that can swap models underneath, so K3 becomes a cheap upgrade later instead of a rebuild.
The model on the front page today is almost never the model that pays your bills. The quiet automation running at 2am, catching the lead you would have missed, is. If you want to know which repetitive tasks in your business are ready to automate right now, on systems built to outlive any single model launch, book a free AI audit call and we will map it out with you.
Sources: Moonshot AI, "Kimi K3" (July 16, 2026). Benchmark figures and pricing quoted as published by Moonshot AI at launch.