Kimi K32.8T
Open Frontier Intelligence

Moonshot AI's flagship open model—a 1M-token context window, native visual understanding, long-horizon reasoning, and an architecture built for coding, agents, and knowledge work.

What is Kimi K3?

Kimi K3 is Moonshot AI's most capable flagship model to date. With 2.8 trillion parameters, it's the world's first open-source model in the 3-trillion-parameter class. Its architecture combines Kimi Delta Attention, Attention Residuals, Stable LatentMoE, and native multimodal understanding.

The model targets frontier intelligence scenarios rather than short, isolated prompts. It's designed to remain effective across long-running engineering sessions, large research corpora, screenshots, images, video, terminal tools, and knowledge workflows that require planning and revision.

For readers comparing advanced open models, Kimi K3 stands out through scale, a 1M-token context window, native vision, and a strongly agentic orientation.

2.8T parameters

A frontier-scale open model that extends Moonshot AI's push beyond the trillion-parameter regime.

1M-token context

Designed to work across extensive codebases, documents, research archives, and sustained agent histories.

Native visual understanding

Processes screenshots, images, and video as part of coding, research, creative, and tool-driven workflows.

Who Kimi K3 is for

Kimi K3 is most valuable when a task is long, multimodal, technical, or knowledge-dense—and when the model must keep working through ambiguity instead of returning a single short answer.

01

Engineering teams

Long-horizon coding and repository-scale work for software engineers, infrastructure teams, and technical leads—large repository navigation, terminal and tool orchestration, and visual feedback during implementation.

02

Research & analysis teams

Evidence-heavy knowledge work for analysts, consultants, and scientists—multi-document synthesis, research-to-code workflows, and reports, dashboards, and visualizations across a 1M-token context.

03

Creative & technical teams

Native vision in the loop for teams combining engineering with images, video, and screenshots—visual iteration and debugging, motion and interaction work, and rapid prototyping across visual and technical domains.

The first open model at 2.8T parameters

Kimi K3 marks a sharp expansion of the open-model frontier—from Kimi K2's 1T parameters to 2.8T within 12 months, and the first open-weight model to reach the 3-trillion-parameter class.

The first open model at 2.8T parameters

How Kimi K3 moves information across scale

K3's architecture improves information flow across sequence length and model depth while keeping sparse expert computation efficient. Kimi Delta Attention (KDA) is a hybrid linear attention mechanism for scaling across long sequences; Attention Residuals selectively retrieve representations from earlier blocks; and Stable LatentMoE activates 16 of 896 experts per token—together reaching roughly 2.5× the scaling efficiency of Kimi K2.

How Kimi K3 moves information across scale

Kimi K3 model FAQ

  • Kimi K3 is Moonshot AI's 2.8-trillion-parameter flagship open model, designed for long-horizon coding, reasoning, native visual understanding, and end-to-end knowledge work.