Open Source AI Models for Business: Is Kimi K3 the End of the Per-Seat AI Bill?
Kimi K3 launched open-weight. Here's the seats-vs-tokens math: when per-token AI pricing beats fixed seat subscriptions for SMB teams.

When is per-token AI pricing cheaper than per-seat?
Per-token AI is cheaper than per-seat when your team's real monthly work costs less than the fixed seat bill for the same people. Multiply seats by monthly seat price for the subscription side; multiply input, output, and cached tokens by their rates for the usage side. When use is light or mixed, usage usually wins if assumptions stay visible.
Kimi K3, a frontier-class open-weight model that launched this week, is the clearest live test of that math. It just became Duet's Balanced-tier default model.
What just launched, and why does it matter for your AI bill?
Moonshot AI released Kimi K3 on July 16, 2026: a 2.8-trillion-parameter mixture-of-experts model that activates just 16 of its 896 experts per token, with a 1-million-token context window: enough to hold roughly 1,573 pages of text in one pass.
Independent benchmarks (next section) put it near the frontier. But precision matters here: K3 is open-weight, not open source today. Moonshot's own announcement commits to releasing the full model weights by July 27, 2026; for now, it's served only through the API.
That's a real distinction from a model like GLM-5.2, which already ships under a genuine MIT license.
What makes K3 matter for your AI bill isn't the launch itself. It's that frontier-class capability just became available with no per-seat license attached. It's priced instead by the token, the way electricity or bandwidth is.
Duet made K3 the default model on its Balanced tier the same week. Capability plus no seat tax is what makes the rest of this article's math possible.
Are open-weight models actually good enough for real business work?
Yes, for the large majority of drafting, research, ops, and reporting work. Artificial Analysis' independent benchmark ranks Kimi K3 at 57 on its Intelligence Index v4.1: #4 of 189 models tracked, behind only Claude Fable 5 (60) and GPT-5.6 Sol (59), ahead of Claude Opus 4.8 (56). That puts K3 ahead of every other open-weights model on the index.
Even Moonshot admits K3 isn't the top model overall: its own blog says performance "still trails the most powerful proprietary models." That's an honest vendor claim worth taking at face value — K3 doesn't need to be #1 to change your bill.
On the benchmark closest to actual knowledge work, AA-Briefcase, K3 scores an Elo of 1,547, a jump of 732 points over Moonshot's prior model. Only Claude Fable 5 ranks higher — that's the number that matters for drafting or research work.
The pattern holds beyond K3. MIT research on open models found open-weight models now average 89.6% of closed-model performance, closing the gap to a new closed release within roughly 13 weeks. Optimal routing to open models, by MIT's estimate, could save businesses up to ~$25B a year, industry-wide.
Featherless's CEO summed up the practical takeaway: most open models run "N-1" behind the leader, "and N-1 is fine" for the bulk of business work. For more on how we test models against real work instead of invented prompts, see how we evaluate models on real work.
What does your team actually spend on AI seats right now?
Take a real team. Maya runs ops for a 5-person boutique event-services company. Nobody on her team writes code. They draft client proposals, research vendors and competitors, run weekly ops checklists, and put together monthly reports.
Here's what Maya's team pays today, seat by seat:
| Role | Tool/tier | Seats | Price/seat/mo | Subtotal |
|---|---|---|---|---|
| 3 team members | ChatGPT Business (Standard) | 3 | $25 | $75 |
| Ops lead + reporting lead | Claude Team (Premium) | 2 | $125 | $250 |
| Total seat bill | 5 | $325/mo |
Those aren't inflated numbers: ChatGPT Business runs $20–25/user/month; Claude Team runs $20–25 Standard or $100–125 Premium. Two people on Premium for research and reporting is common, not an edge case.
That $325 bills the same in a slow week as a busy one. A seat is a fixed, provisioned-per-human cost, owed whether the person opens the app once or fifty times that month.
What would that same work cost priced by the token?
Now price that exact same month of work by the token, at Kimi K3's real Duet rates: $3 per million input tokens, $15 per million output tokens, $0.30 per million cached input tokens.
Here's Maya's team's actual monthly volume, with every assumption labeled:
| Work type | Volume/mo | Assumption | Tokens |
|---|---|---|---|
| Drafts (emails, proposals, updates) | ~440 | ~700 words in, ~900 words out | 410K in / 527K out |
| Research briefs | ~66 | ~10K words cached background + ~1.5K new, ~1.8K words out | 878K cached / 132K in / 158K out |
| Ops checklists | ~220 | ~300 words in, ~200 words out | 88K in / 59K out |
| Monthly reports | 4 | ~20K words cached prior data, ~2.4K words out | 106K cached / 13K out |
| Total | 630K uncached in / 1.0M cached in / 760K out |
Priced at Kimi K3's rates:
| Token type | Volume | Rate/1M | Cost |
|---|---|---|---|
| Uncached input | 0.63M | $3.00 | $1.89 |
| Cached input | 1.0M | $0.30 | $0.30 |
| Output | 0.76M | $15.00 | $11.40 |
| Total usage bill | ≈ $13.60/mo |
That $13.60 covers the same ~730 pieces of work the $325 seat bill covered. Nothing was rationed. The research briefs and reports lean on caching because they reuse the same client and vendor background repeatedly. That's how agentic work actually runs, not a trick.
Seats vs. tokens: which one wins for your team?
This isn't hypothetical: NPR reported that Lindy.ai switched its traffic to a metered model after its AI bill became "by far" its No. 1 expense, cutting costs roughly 10x.
Put Maya's numbers side by side with the general shape of the decision:
| Per-seat (fixed) | Per-token (usage) | |
|---|---|---|
| What you pay for | A person having access | Work actually run |
| Predictability | Same bill regardless of use | Scales with volume |
| Idle cost | Full price even unused | $0 if no work runs |
| Heavy vs. light user | Same price either way | Heavy user pays more, light pays less |
| Typical price today | ChatGPT Business $20–25/user/mo; Claude Team $20–25 Standard / $100–125 Premium | Kimi K3 in Duet: $3/1M input, $0.30/1M cached, $15/1M output, all at cost |
| Best fit | One person using it constantly | Multiple people, spiky or occasional use |

The rule underneath the table: a seat only earns its keep once a person's real usage would cost more metered than the seat itself. For Maya's team, that threshold sits at roughly 20–25× today's actual volume: tens of millions of output tokens a month, sustained.

For most non-technical work like this, that line rarely gets crossed outside an agent running near-continuously.
Is Kimi K3 the cheapest model to route this work to?
No. Worth saying plainly. At $3/1M input and $15/1M output, Kimi K3 is priced near frontier rates, not floor rates. In Duet's own model picker, GLM-5.2 ($1.40/$4.40/$0.26) and Grok 4.5 ($2/$6) are both cheaper.
K3 costs what it costs because it performs close to what it costs. The Decoder's analysis puts K3's cost-per-task at $0.94, close to GPT-5.6 Sol's $1.04, about half Opus 4.8's $1.80, but still well above GLM-5.2's $0.32 and DeepSeek V4 Pro's $0.04. As the piece frames it, this is the end of super-cheap Chinese AI.
There's a real caveat too. On an adversarial knowledge-probing benchmark (AA-Omniscience), K3's hallucination rate rose to 51%, up from 39% on its predecessor, even as its raw accuracy improved from 33% to 46%.
That's one narrow, adversarial eval, not a claim that K3 gets facts wrong half the time. It's a reason to keep a human reviewing anything client-facing, not a reason to skip the model.
The savings story here was never about finding the cheapest token. It was about killing the seat tax and paying only for real usage.
How do you route work instead of picking one model forever?
The answer isn't picking one model and living with it forever. It's routing: send routine work to a balanced model by default, and send the specific tasks that need more to a frontier model, one click at a time.

Concretely: drafts, research, ops checklists, and reports default to Balanced. A high-stakes deliverable or long-horizon analysis escalates to frontier for that one task, not the whole team.
This is exactly the mechanism a managed multi-model workspace is for. In Duet, Kimi K3 is the Balanced-tier default: the model routine drafts, research, ops checklists, and reports run on by default. When a task actually needs more, you switch the model for that one agent with a single tap, mid-session, without losing context or starting a new subscription.
There's no seat to provision for the switch and no separate bill to reconcile. Every model in the picker, from Kimi K3 to Claude Opus 4.8 to GPT-5.6 Sol, is billed the same way: tokens passed through at cost, no markup, no per-seat tax, unlimited seats for your whole team.
Duet's pricing page shows the live rate for every model in the picker and a sample month's bill itemized to the cent, plus a $5 trial credit with no card required, so you can run your own team's numbers first.
When should you still pay for a frontier model?
Frontier models earn their higher price in three places: the hardest long-horizon reasoning, the highest-stakes client-facing deliverables, and anything where a small accuracy gap costs real money: a legal-adjacent clause, a board-level analysis, a make-or-break pitch.
That's where K3's hallucination caveat from the last section actually matters. On work where being wrong is expensive, route it to Claude Opus 4.8, GPT-5.6 Sol, or Claude Fable 5 instead, and pay the premium on purpose.
Frontier isn't obsolete. It's just not the default for the routine work that fills most of a non-technical team's month. Match the model to the stakes of the task, not to habit or to whichever subscription happens to be open. For a deeper walkthrough of matching models to tasks, see choosing the right AI model for your business.
Frequently asked questions
Is Kimi K3 open source? Not yet, precisely. Kimi K3 is open-weight. Moonshot has committed to releasing the full weights by July 27, 2026, but at launch it's served only through the API. That's different from GLM-5.2, which already ships under a genuine MIT license.
Both are "open" in the sense that matters for cost: no single vendor controls who can serve them.
Is Kimi K3 the cheapest AI model? No. At $3/1M input and $15/1M output, K3 is priced near frontier rates: GLM-5.2 ($1.40/$4.40) and Grok 4.5 ($2/$6) are both cheaper in Duet's own model picker. K3's price reflects near-frontier performance, not a budget position. The savings story in this article isn't about picking the cheapest token.
How much do AI tokens actually cost compared to a seat subscription? It depends entirely on how much you use the seat. A seat costs the same $20–125/month whether one person uses it constantly or barely opens it.
Token pricing charges only for what you actually send and receive. For most non-coding work, that's typically a small fraction of the equivalent seat bill.
What kind of work is safe to route to a balanced open-weight model like Kimi K3? Routine business work: drafting emails and proposals, first-pass research and account briefs, operations checklists, and recurring reports. These are high-volume, lower-stakes tasks where the gap between a #4-ranked model and a #1-ranked model doesn't show up in the output.
When should I still use a frontier model instead? For the hardest long-horizon reasoning, the highest-stakes client-facing work, and anything where a mistake is expensive: legal-adjacent drafts, board-level analysis, a make-or-break pitch. Route those specific tasks to a frontier model and keep everything else on the balanced tier.
How do I actually try this without ripping out my current AI subscriptions? Start by mapping what your team's seats cost per month against what the same work would cost metered. The worked example above is a template.
Duet's pricing page shows Kimi K3's live at-cost rate, unlimited seats, and a $5 trial credit with no card required, so you can run the comparison against your real workload before changing anything.






