K3 is 40% cheaper than Opus 4.8 — but Sonnet 5's introductory $2/$10 currently undercuts K3's $3/$15. What K3 offers instead: weights you can self-host. Claude answers with adjustable effort; K3 reasons at full depth on every call. Here is the honest trade.
Unofficial tool. Not affiliated with Moonshot AI or Kimi. Prices and scores are gathered from public sources and may lag official changes.
| Kimi K3Moonshot | Claude Opus 4.8Anthropic | Claude Sonnet 5Anthropic | |
|---|---|---|---|
| Input / 1M tokens | $3.00 | $5.00 | $2.00 |
| Cached input / 1M | $0.30 | $0.50 | $0.20 |
| Output / 1M tokens | $15.00 | $25.00 | $10.00 |
| Context window | 1,049K | 1,000K | 1,000K |
| Open weights | |||
| Vision input | |||
| Reasoning control | Always max | Adjustable | Adjustable |
| Self-hostable | |||
| Released | Jul 2026 | May 2026 | Jun 2026 |
Prices are vendor list prices, July 2026. Capability rows are factual specs; we deliberately don't print head-to-head benchmark scores here because vendors measure under different harnesses — see the K3 overview for K3's reported numbers with caveats.
Drag to match your traffic. Cache-hit rate matters more than list price for agent workloads.
Cost per 1,000 calls at list prices. All models get the same cache-hit rate (each vendor discounts cached input ~90%). K3 always reasons, so its real output token count runs higher than a model answering directly — drag output up to model that.
The short version: K3 claims Opus-class agentic numbers at Sonnet's list price — but Sonnet 5's introductory $2/$10 (through August 2026) currently beats K3 on both sides of the meter, and K3 makes you pay for reasoning you can't turn off. Heavy agent workloads with warm caches and self-hosting plans favor K3; short interactive traffic and knowledge-heavy work favor Claude. Vendor-reported benchmarks are not apples-to-apples, so test both on your own tasks before committing.
Against Claude Opus 4.8 ($5/$25 per 1M tokens), yes — K3's $3/$15 is 40% cheaper on both sides. Against Claude Sonnet 5 it currently isn't: Sonnet 5 runs at an introductory $2/$10 through August 2026 (list $3/$15, matching K3). Both vendors discount cached input about 90% — $0.30 for K3, $0.50 for Opus 4.8, $0.20 for Sonnet 5 today. And K3 reasons on every call, billing reasoning as output, so on short tasks its effective cost can exceed Claude with reasoning dialed down.
Reported scores put K3 at 76.8 on SWE-bench Verified — competitive with, not clearly ahead of, the top Claude models. K3's stronger claims are agentic (Terminal-Bench 2.1: 88.3, BrowseComp: 91.2). These are vendor-adjacent numbers from July 2026 write-ups, not independent replications; run both on your own repo before deciding.
Effectively yes — K3 offers 1,048,576 tokens and Claude Opus 4.8 and Sonnet 5 both offer 1M. The difference is pricing structure: K3 charges flat rates across the whole window with cache hits at $0.30/1M, while Anthropic prices cache writes and reads separately.
Yes — that's the structural difference. K3's weights ship under a Modified MIT license (full release announced for late July 2026), so it can run in your own infrastructure. Claude is API-only. At 2.8T parameters self-hosting K3 requires serious hardware, but the option changes negotiating leverage and compliance posture.
NottoAI has Kimi K3, Claude Opus 4.8 and Claude Sonnet 5 in the same model switcher — ask a question, switch models mid-thread, and compare answers on your actual work. Free to start, no API keys.
Kimi K3 Toolkit
NottoAI has both models in one thread — ask the same question, switch models, and judge on your own work. Free, no API keys.
No credit card required · 100 free credits included