.##....##.########.##......##..######.....########..#######..########.....###....##....##
.###...##.##.......##..##..##.##....##.......##....##.....##.##.....##...##.##....##..##.
.####..##.##.......##..##..##.##.............##....##.....##.##.....##..##...##....####..
.##.##.##.######...##..##..##..######........##....##.....##.##.....##.##.....##....##...
.##..####.##.......##..##..##.......##.......##....##.....##.##.....##.#########....##...
.##...###.##.......##..##..##.##....##.......##....##.....##.##.....##.##.....##....##...
.##....##.########..###..###...######........##.....#######..########..##.....##....##...

All signal, no noise, 24/7.
Built for Humans & AI Agents.

Moonshot AI has recently elevated its Kimi K3 model from a research release to a major player in enterprise cloud infrastructure. The Kimi K3 is an open-weight Mixture-of-Experts (MoE) language and vision model that has captured significant attention in the AI community. Although the Beijing-based laboratory initially released the model on GitHub in July 2026, its integration into commercial platforms, such as Amazon Bedrock, marks a crucial shift, enabling enterprise teams outside of China to deploy the technology.

Model Specifications and Scale

The most immediate and striking detail regarding Kimi K3 is its sheer size. The full checkpoint repository, which resides on Hugging Face, measures approximately 1.56 terabytes (TB) and is distributed across 96 safetensors shards. This makes it one of the largest open-weight models released to date, raising discussion about the practical definition of “open weights” given the massive download size.

Technically, Kimi K3 is described by Moonshot AI as “our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.” This designation details several headline-grabbing specifications:

  • Total Parameters: 2.8 trillion.
  • Context Window: 1,048,576 tokens (or 1 million tokens).
  • Modality: Native multimodal input (vision).

The model’s architecture is built around a Mixture-of-Experts (MoE) framework. Rather than using all 2.8 trillion parameters for every query, the system is designed to activate only a subset. Specifically, for any given token, the model uses 104 billion activated parameters, routed through 896 total experts, with 16 experts selected per forward pass, in addition to two shared experts.

Architectural Innovations for Stability

Moonshot AI names the overall design “Stable LatentMoE.” The development of MoE models at scale has historically faced technical challenges, notably “routing collapse,” where the network overly relies on a small number of experts. To counteract this, Kimi K3 incorporates several unique architectural additions:

  • Kimi Delta Attention (KDA): This is one of two key mechanisms implemented in the 93-layer network (69 layers use KDA, and 24 layers use Gated Multi-head Latent Attention, or MLA).
  • Attention Residuals (AttnRes): This component is intended to prevent signal degradation within the deep network layers, particularly crucial given the model’s extended context window.

The model also integrates visual understanding via a MoonViT-V2 encoder, which contributes roughly 401 million parameters, linking the vision input directly into the language backbone rather than treating it as a separate system.

From Open Source to Enterprise Cloud

The transition of Kimi K3 from a research artifact to a commercially accessible tool is marked by a distinct timeline. The open weights were first published by Moonshot AI in July 2026. The technical specifications were detailed by TechJuice on July 28, 2026.

The ability for developers to test the model locally improved significantly when Unsloth released an updated deployment guide on September 8, 2026. The most significant commercial milestone occurred on September 18, 2026, when the model was listed on Amazon Bedrock, Amazon Web Services’ managed platform for foundation models. This listing is notable because it transitions the model from requiring local management of a 1.56TB repository to being accessible via a managed API endpoint.

Practical Implications of the Size

The 1.56TB size is a key factor shaping how the model will be utilized. Storing and downloading the entire repository is a substantial undertaking, meaning that most enterprises, outside of large, well-funded research labs, will likely interact with Kimi K3 through a hosted API service rather than through local deployment. The availability on Amazon Bedrock addresses this exact operational gap.

Furthermore, Moonshot AI has adopted native low-precision training methods, releasing the weights using MXFP4 weights and MXFP8 activations. This approach differs from common practice—training in high-precision formats like BF16 and then quantizing—and suggests the model was engineered from the outset to be highly deployable and resistant to loss of accuracy from compression.

The direct predecessor, Kimi K2, was replaced by Kimi K3, which Moonshot AI claims delivers a roughly 2.5x improvement in overall scaling efficiency over Kimi K2. This evolution reflects a broader industry trend toward improving the stability and efficiency of sparse activation and long-context retrieval.

While the Bedrock listing is a major development, readers should note that specific details, such as official per-token pricing or standardized benchmark scores (e.g., MMLU, GPQA, AIME, SWE-bench), have not been independently verified by third parties as of the report’s writing.

Max

Written by

Max

Covers AI news, agentic AI, LLMs, and tech developments. When he is not writing, he is comparing open-source models' tokens per second just to see how they hold up.

+ , , ,