Moonshot AI's open-weight flagship — a 1M token context, native vision, and agentic scores that beat models costing far more. Run it on NottoAI without a Moonshot API key.
No credit card · Compare K3 against Claude, GPT and Gemini in one thread
Unofficial tool. Not affiliated with Moonshot AI or Kimi. Prices and scores are gathered from public sources and may lag official changes.
Parameters
Token Context
GPQA Diamond
Per 1M Input Tokens
Capabilities
K3 is not a general chatbot upgrade. It is aimed squarely at long agent runs and large codebases — and it makes real trade-offs to get there.
The largest openly released model to date. Moonshot ships the weights under a Modified MIT license — self-host it, fine-tune it, or run it through an API.
1,048,576 tokens. Moonshot reports 90.4 on a 1M-token evaluation with no context management — long inputs stay coherent instead of degrading at the tail.
Tuned for navigating real codebases: locating the right files, iterating against tests and runtime feedback, and debugging rather than one-shot snippets.
Built for multi-step tool use that runs for hours, not turns. Strong on terminal and browsing benchmarks where most models lose the plot mid-task.
Reads images alongside text — screenshots, diagrams, logs, and design mocks feed straight into the same reasoning pass.
There is no cheap non-thinking mode: K3 reasons on every call, and reasoning_effort accepts only max. Budget output tokens accordingly.
Benchmarks
Reported figures collected from public write-ups in July 2026. Treat them as vendor-adjacent claims, not independent replication — and note where K3 is weak, not just where it wins.
The honest read: K3's agentic and science scores are genuinely frontier-class, its SWE-bench number is competitive rather than dominant, and HLE at 43.5 without tools shows the gap that open-weight models still have on raw breadth of knowledge.
Pricing
Drag the sliders to match your workload and see what K3 costs against the alternatives you are probably weighing it against.
Estimated monthly API spend at list prices, before prompt caching. K3 bills cache hits at $0.30 per 1M input tokens, so a repeated system prompt lands well under these figures in practice. Remember that K3 always reasons — its reasoning tokens bill as output.
Need cache-hit rates and per-call breakdowns? Open the full K3 cost calculator →
The lineup
K3 is the expensive one for a reason. Most teams route the hard 20% of calls to K3 and the rest to a cheaper sibling — all three are available on NottoAI.
| Model | Context | Input / 1M | Output / 1M | Best for |
|---|---|---|---|---|
| Kimi K3moonshotai/kimi-k3 | 1M | $3.00 | $15.00 | Frontier reasoning, big-repo coding, long agent runs |
| Kimi K2.7 Codemoonshotai/kimi-k2.7-code | 262K | $0.82 | $3.75 | Everyday code editing and agent loops on a budget |
| Kimi K2.6moonshotai/kimi-k2.6 | 262K | $0.68 | $3.42 | General-purpose chat and reasoning, cheap |
Load an entire service into the 1M window and let K3 trace call sites, plan the migration, and iterate against your test output instead of guessing at file boundaries.
91.2 on BrowseComp means K3 holds a research thread across dozens of tool calls — the failure mode where an agent forgets its own goal shows up much later.
Feed screenshots, stack traces, and logs together. Native vision plus always-on reasoning turns a pile of artifacts into a hypothesis you can test.
Modified MIT weights make K3 viable where data cannot leave your network — prototype against the API here, then move the same prompts to your own cluster.
Moonshot lists Kimi K3 at $3 per million input tokens and $15 per million output tokens, with cache-hit input billed at $0.30 per million. That puts it in the same bracket as mid-tier Western flagships, and roughly 4× the price of Moonshot's own K2.7 Code endpoint at $0.82 / $3.75.
The API went live on 16 July 2026 at api.moonshot.ai/v1 under the model id kimi-k3. Moonshot said the full weights would follow by 27 July 2026 under a Modified MIT license.
It is open-weight rather than fully open source: the weights ship under a Modified MIT license, but the training data and pipeline are not released. At 2.8 trillion parameters it is the largest openly released model so far, and self-hosting it needs serious hardware.
Reported scores put K3 at 76.8 on SWE-bench Verified and 81.2 on FrontierSWE — competitive with, not clearly ahead of, the top closed models. Its stronger claims are agentic: 88.3 on Terminal-Bench 2.1 and 91.2 on BrowseComp. Its clearest weak spot is HLE, at 43.5 without tools.
1,048,576 tokens — a full 1M. Moonshot reports 90.4 on a 1M-token long-context evaluation run without any context management, which suggests recall holds up rather than collapsing near the limit.
No. Reasoning is always enabled on K3 and the reasoning_effort parameter only accepts max. If you need a cheaper, faster Kimi for simple calls, route those to K2.7 Code or K2.6 instead — both are available here too.
Yes. Sign in to NottoAI and pick Kimi K3 from the model switcher — no Moonshot account, no card, and you can compare it side by side with Claude, GPT and Gemini in the same thread.
Kimi K3 Toolkit
No Moonshot account, no API key, no card. Sign in and switch models mid-conversation to see how K3 stacks up on your own work.
No credit card required · 100 free credits included