Kimi K3 brings a 2.8T open frontier model to long-horizon coding agents

Moonshot AI introduced Kimi K3 on July 16 with 2.8T parameters, native vision, a 1M-token context window, and long-horizon coding workflows; full weights are due July 27.

On July 16, 2026, Moonshot AI introduced Kimi K3 as an open frontier intelligence model. The company lists 2.8 trillion parameters, native vision, and a 1-million-token context window, with long-horizon coding, knowledge work, and deep reasoning as the main use cases. Kimi K3 is available through Kimi, Kimi Work, Kimi Code, and the Kimi API; the full model weights are scheduled for release on July 27.

The engineering story is not only about parameter count. Moonshot says Kimi K3 uses Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), together with Stable LatentMoE, activating 16 of 896 experts per token. The company describes the changes as an approximately 2.5x improvement in overall scaling efficiency over Kimi K2; real costs will still depend on hardware, parallelism, context length, and serving configuration.

The most notable product signal is the emphasis on sustained work. Kimi K3’s coding examples include analyzing and optimizing GPU kernels in a sandbox for up to 24 hours, building a Triton-like MiniTriton compiler, and iterating on game, frontend, and CAD output through a visual loop. These are Moonshot’s internal or product-test narratives, useful for understanding the intended capability but not independent benchmark conclusions.

The knowledge-work demonstrations follow the same pattern. Moonshot describes Kimi K3 reading papers, writing numerical pipelines, validating results, producing interactive HTML dashboards, and using parallel subagents for scientific analysis. One case says the model completed in about two hours work that would normally take one to two weeks; that speed figure is best read as a product-direction claim, not a general production promise.

At the deployment layer, Kimi K3 uses MXFP4 weights and MXFP8 activations and recommends supernode configurations with 64 or more accelerators. Moonshot also says it is contributing a KDA prefill-cache implementation to vLLM. The listed API price is $0.30 per million cache-hit input tokens, $3 per million cache-miss input tokens, and $15 per million output tokens; teams should check the latest platform page before adopting it.

Moonshot also describes important limitations. Kimi K3 is sensitive to how thinking history is passed through. If an agent harness fails to preserve the required history, or if a session switches from another model, generation quality can become unstable. The model can also be overly proactive, making decisions on the user’s behalf when intent is ambiguous or a small execution problem appears. Long execution amplifies context loss and over-broad permissions.

The fact that the full weights are not available today is another key part of the story. A 2.8-trillion total parameter count does not mean every inference activates the same amount of compute, and it does not mean an ordinary team can deploy the model on a normal workstation. Open models create opportunities for audit, fine-tuning, and self-hosting, but they still require the right hardware, quantization, inference framework, agent harness, and security policy.

Kimi K3 moves the open-model competition toward a more concrete question. The comparison is not only a single-turn score; it is whether a model can preserve context, use tools, process visual data, verify outcomes, and stop at the right time across a long task. Once the full weights, technical report, and more independent evaluations arrive, the market will have a firmer basis for judging Kimi K3’s place in frontier models and coding-agent infrastructure.

MODULE.002 //

More insights

Ideas on websites, AI automation, digital marketing, AI news, and VMTS updates.