What Is Kimi K3? Moonshot AI's 2.5T Flagship Model Explained (2026)
Kimi K3 is Moonshot AI's 2.5 trillion parameter MoE model with 1M+ token context. Leaked specs, early benchmarks, estimated pricing, and how to try it today.

Moonshot AI did not plan to announce Kimi K3 this way.
On July 14, 2026, a promotional page on the Kimi Open Platform leaked early — revealing a "K3 launch" top-up bonus starting July 15 at midnight Beijing time. The page was pulled within hours, but by then screenshots were everywhere. The next evening, the official @Kimi_Moonshot account posted what appears to be the model's first public acknowledgment, racking up nearly 6,000 likes overnight. Early testers on X are already comparing it to GPT-5.6 and Fable 5.
If you evaluate open models for production use — or just want to know whether the hype is real — this is the one to watch this week. We cross-referenced the leaked platform page, five independent pre-release reports, and first-day X reactions to separate what is confirmed from what is speculation. This guide covers the specs, early performance signals, estimated pricing, and how to try K3 right now.
Kimi K3 in 30 seconds
Kimi K3 is Moonshot AI's next flagship — a Mixture-of-Experts model with roughly 2.5 trillion total parameters, a reported 1 million+ token native context window, and multimodal input (text, images, audio). It follows the K2 family (K2, K2.5, K2.6, K2.7 Code) and represents a major scale-up: K2.6 had roughly 1T total parameters with 256K context. K3 more than doubles the parameter count and quadruples the context.
Important: As of July 15, 2026, Moonshot AI has not published an official model card, benchmark table, license, or API pricing for K3. The specs below are sourced from the leaked promotional page, independent coverage, and early tester reports. We will update this article as official information becomes available.
What we know about Kimi K3 specs so far
Multiple independent sources converge on the same core numbers. Here is what lines up across the leaked platform page, pre-release coverage, and early access reports:
| Spec | Reported value | Confidence |
|---|---|---|
| Architecture | Mixture-of-Experts (MoE) | High — consistent with K2 family |
| Total parameters | ~2.5 trillion | High — multiple sources agree |
| Active parameters | ~40–100B per token | Medium — estimates vary by source |
| Context window | 1M tokens (some reports say 2M) | High — leaked page + tester reports |
| Input modalities | Text, images, audio | Medium — confirmed by some testers |
| Output | Text | High |
| Languages | 50+ including Chinese, English, Japanese, Korean | Medium |
| Training data | 15T+ tokens | Low — single source |
Kimi-Linear: how 1M context actually works
The context window claim is the one that matters most in practice. Moonshot reportedly uses a technology called Kimi-Linear — a linear attention hybrid with selective state compression — that "reduces KV cache by 75% and boosts decoding throughput 6x for 1M context." If accurate, this means K3 can handle million-token inputs without the extreme VRAM requirements that make other long-context models impractical to serve.
For comparison, most competing open models top out at 128K–256K context. DeepSeek V4 offers 1M but at significantly different pricing. Having practical 1M context at competitive pricing would be a genuine differentiator.
The specs tell half the story. The other half is how much K3 has changed compared to the models it replaces.
From K2 to K3: a generational leap, not an incremental update
Here is how the Kimi model line has evolved:
| Model | Total params | Context | Focus | Pricing (input/output per 1M) |
|---|---|---|---|---|
| K2.6 | ~1T MoE | 256K | Agentic coding, tool use | $0.60 / $2.50 |
| K2.7 Code | ~1T MoE | 256K | Coding with 30% fewer reasoning tokens | $0.95 / $4.00 |
| K3 | ~2.5T MoE | 1M+ | General + agentic + multimodal | TBD |

The K2.7 Code variant was released in June 2026 as a coding specialist, but independent benchmarks were not flattering. On Kimi Code Bench v2 it scored 62.0 versus GPT-5.5's 69.0; on Program Bench, 53.6 versus 69.1. K3 appears to be Moonshot's response: instead of iterating on the coding niche, go bigger across the board.
Numbers on paper are one thing. What matters is whether the model actually performs at the level those numbers imply.
Early performance signals: promising but unverified
No official benchmarks have been published. But early testers and community reactions paint a picture:
- Jun Song (AI researcher) tested K3 against Opus 4.8 on a Flappy Bird coding challenge and called it "significantly better than Opus," claiming "Opus 5 level" — though this is a single anecdotal test, not a benchmark suite.
- BridgeMind noted that rumors place K3 "on par with Claude Opus 4.7" but withheld judgment pending their own BridgeBench evaluation.
- Multiple X users report K3 performing between GPT-5.6 and Fable 5 on coding and reasoning tasks, though systematic independent testing has not been published yet.
- One user described it as a potential "DeepSeek moment" for open-source models.
Rule of thumb: Single-task demos (especially Flappy Bird) are not benchmarks. Wait for Artificial Analysis, Chatbot Arena, or SWE-bench Verified results before making production decisions. Early hype and actual evaluated performance diverge more often than not.
Update: Official benchmarks are now available. See our full breakdown in Kimi K3 Benchmarks: Every Score, Ranked and Explained.
Estimated pricing
Official K3 pricing has not been published. Based on the leaked promotional page and Moonshot's pricing history, estimates from pre-release coverage suggest:
| Estimated range | Context | |
|---|---|---|
| Input / 1M tokens | $0.80–$1.20 | Would undercut GLM-5.2's $1.40 |
| Output / 1M tokens | $3.50–$4.50 | Would undercut GLM-5.2's $4.40 |
| Launch promotion | 10–30% bonus credits | July 15 – August 11 per leaked page |
If these estimates hold, K3 would be priced aggressively — offering 1M context at roughly the same cost as competitors offering 128K. The pricing strategy is consistent with Moonshot's pattern: K2.6 was already one of the cheapest serious models at $0.60/$2.50.
Expert pitfall: Leaked promotional pricing often reflects launch discounts, not steady-state rates. K2.6's $0.60/$2.50 was stable, but K2.7 Code jumped to $0.95/$4.00. Do not lock in capacity commitments based on estimated pricing — wait for the official rate card.
Pricing aside, the fastest way to form your own opinion is to try the model directly.
How to try Kimi K3 right now
Based on early access reports, K3 appears to be rolling out through these channels — listed from zero-friction to most hands-on:
1. Kimi chat (kimi.com) — try it in 60 seconds. Some users report K3 is already available as a model option in the Kimi web interface. Sign up with Google or phone number, no credit card required. Free tier with usage limits. This is the fastest way to test whether K3's quality matches the hype for your use case.
2. Kimi Open Platform API. The leaked promotional page was on Moonshot's API platform (platform.moonshot.cn). If K3 follows the K2 pattern, it will appear as a new model ID in the API shortly after launch.
3. Open weights (expected). Moonshot released K2 and K2.6 as open-weight models. K3 open weights are widely expected but not confirmed. For download details and hardware requirements, see Kimi K3 on Hugging Face. Want to know if your GPU can run K3? Check Kimi K3 System Requirements.
4. Third-party hosts. K2.6 is available through AWS Marketplace, Together AI, and other API aggregators. K3 is now available on OpenRouter — see our Kimi K3 OpenRouter setup guide for pricing and step-by-step instructions.
Kimi K3 vs Inkling, GLM 5.2, and the open-weight field
Without official K3 benchmarks, the fairest comparison uses K2.6 (Moonshot's current shipping model) as a floor and community claims as a ceiling:
| Kimi K2.6 (floor) | Kimi K3 (rumored) | Inkling | GLM 5.2 | |
|---|---|---|---|---|
| Total params | ~1T | ~2.5T | 975B | ~1.5T |
| Context | 256K | 1M+ | 1M (open), 256K (hosted) | 128K |
| SWE-bench Verified | 80.2% | TBD | 77.6% | 80.0% |
| AIME 2026 | 96.4% | TBD | 97.1% | 99.2% |
| Multimodal input | Text | Text, image, audio | Text, image, video, audio | Text |
| Open weights | Yes | Expected | Yes | Yes |
If K3 performs at the level early testers suggest (between GPT-5.6 and Fable 5), it would be the strongest open-weight model available — surpassing both Inkling and GLM 5.2. But "early testers suggest" is not "independent benchmarks confirm."
Rule of thumb: Never swap a production model based on X posts. The minimum bar is one independent evaluation (Artificial Analysis, Chatbot Arena, or SWE-bench Verified) plus your own task-specific test. Every major model launch generates identical "this beats everything" posts; most settle 10–20% below the initial claims within two weeks.
Five things to watch this week
This story is still developing. The key signals to track:
- Official model card and benchmarks. Until Moonshot publishes verified scores on standard benchmarks (AIME, SWE-bench, MMMU, HLE), all performance claims are anecdotal.
- Chatbot Arena and Artificial Analysis rankings. These independent evaluations are the most credible signals. K3 may already be on Arena under the codename "Kivine."
- Open weights release. Whether and when Moonshot releases K3 weights determines whether it joins the open-weight competition or stays API-only.
- Actual pricing. The leaked promotion suggests aggressive pricing, but official rates may differ.
- Context window verification. 1M (or 2M) context is a strong claim. Independent needle-in-haystack and long-context evaluation results will confirm whether it works in practice.
Frequently asked questions
Is Kimi K3 released? As of July 15, 2026, K3 appears to be in a soft launch. The official Kimi account posted what seems to be an acknowledgment, and some users can access it on kimi.com, but no official announcement with specs has been published.
How big is Kimi K3? Roughly 2.5 trillion total parameters in a Mixture-of-Experts architecture. Active parameters per token are estimated at 40–100B depending on the source.
Is Kimi K3 open source? Open weights are confirmed for July 27. Moonshot released K2 and K2.6 as open-weight models, and K3 follows the same pattern. See Is Kimi K3 Open Source? for the full license, timeline, and self-hosting details.
How does it compare to GPT-5.6 or Fable 5? Early testers claim K3 performs between these two models, but no independent benchmarks have verified this. Wait for Artificial Analysis or Chatbot Arena results. For a detailed head-to-head comparison, see Kimi K3 vs Fable.
What does it cost? Official pricing is not published. Estimates suggest $0.80–$1.20 per million input tokens and $3.50–$4.50 per million output tokens, which would make it one of the cheapest frontier-class models available. See our Kimi K3 pricing breakdown for a complete cost analysis.
Can it handle long documents? K3 reportedly supports 1M+ tokens of context using Moonshot's Kimi-Linear technology, which would make it one of the longest-context models available at any price point.
Core summary
If the leaked specs hold up, Kimi K3 is the most significant open-weight release of 2026 so far — a 2.5T MoE with million-token context, multimodal input, and pricing that could undercut every competitor at its tier.
But leaked specs and X posts are not benchmarks. Here is the honest state of play:
- Credible: The parameter count, context window, and pricing range are consistent across multiple independent sources.
- Promising: Early testers report performance between GPT-5.6 and Fable 5, which would make K3 the strongest open model available.
- Unverified: No official model card, no independent benchmark scores, no confirmed pricing.
The concrete next step: go to kimi.com, sign up in 60 seconds, and test K3 on a task you actually care about. Your own evaluation on your own workload is worth more than any leaderboard. We will update this article as official specs and independent benchmarks arrive.
Author
Categories
Seedance 2.0
ByteDance latest video model. Text & image to video, up to 1080p.
Try Seedance 2.0 →Wan Video
Wan 2.7 series — text, image, reference to video & video editing.
Try Wan Video →AI Image Generator
Nano Banana Pro, GPT Image 2 & more. Generate stunning images in seconds.
Try Image Generator →More Posts
Wan 2.7 Audio Guide: Voice Reference, Multi-Character Audio & Audio Cues (2026)
A practical guide to Wan 2.7 audio capabilities: how voice reference works, what audio cues are available, how to assign voices to multiple characters, and how to get synced audio output that matches your video.

How to Use the GPT Image 2 API: A Complete Developer's Guide (2026)
Production-ready GPT Image 2 API guide — authentication, image generation, streaming, multi-turn editing, background removal, pricing, and code samples in Python and Node.js.

Kimi K3 VRAM Requirements: How Much Memory You Actually Need to Run K3
Complete VRAM requirements for Kimi K3: BF16 needs 594 GB, Q4 needs ~350 GB. Detailed per-GPU memory breakdown, KV cache costs, context-length tradeoffs, and practical deployment memory planning.
Newsletter
Join the community
Subscribe to our newsletter for the latest news and updates