AWS Adds 2.8T-Parameter Kimi K3 to Bedrock with Prompt Caching Support
Moonshot AI's massive Kimi K3 model is now available on Amazon Bedrock, bringing a one-million-token context window and explicit prompt caching to enterprise developers.

Key takeaways · 3
- 01
Kimi K3 supports a one-million-token context window for extensive document and codebase analysis.
- 02
The model activates 104 billion parameters during inference out of a total 2.8 trillion.
- 03
Explicit prompt caching requires a minimum of 1,024 tokens and retains data for 30 minutes.
The Bedrock Launch
Moonshot AI's Kimi K3 model is now generally available on Amazon Bedrock through US Geo or Global cross-Region inference profiles, as AWS does not list a single-Region option. [4] The launch introduces a 2.8-trillion-parameter mixture-of-experts model that activates 104 billion parameters during inference. [4] According to Moonshot AI, this model is its most capable offering to date. [2] Kimi K3 supports native vision capabilities and a one-million-token context window, processing image and text inputs to return text outputs. [1][4] The developer reports that Kimi K3 delivers approximately 2.5 times better scaling efficiency than its predecessor, Kimi K2. [1][4]
API and Pricing Mechanics
The model is the first open-weight offering on Amazon Bedrock to support explicit prompt caching. [1] Developers can set explicit cache controls using the Responses and Chat Completions APIs, requiring a minimum of 1,024 tokens for each cache checkpoint. [4] Cached content remains available for at least 30 minutes, which AWS notes can improve cache-hit rates while reducing latency and costs when reusing long prompt prefixes. [4] For Global cross-Region inference on the Standard tier, AWS prices the model at $3 per million input tokens and $15 per million output tokens, while cache reads cost $0.30 per million tokens and cache writes cost $3.75 per million tokens. [4]
Security and Ecosystem Expansion
AWS asserts that Kimi K3 inference data remains within the AWS data boundary and is not shared with the model provider or used to train the underlying model. [2][4] The deployment utilizes zero data retention and zero operator access policies, preventing even AWS staff from accessing prompts and completions. [2][4] The integration builds on a broader expansion of open-weight models on Amazon Bedrock, which has added models from DeepSeek, Google, MiniMax, Mistral AI, NVIDIA, OpenAI, and Qwen since 2025. [2] Kimi K3 developers can also utilize platform-level inference capabilities added to Bedrock in 2026, including tool calling, structured output, reasoning, and response streaming. [2]
What it means
With its massive context window and native vision support, Kimi K3 offers a compelling enterprise alternative to proprietary systems like OpenAI o3 and Claude 3.5 Sonnet for intensive document processing tasks. The introduction of explicit prompt caching gives enterprise developers a crucial mechanism to mitigate the steep costs typically associated with large-context workflows. Furthermore, by utilizing cross-Region inference, AWS ensures the massive compute demands of a 2.8-trillion-parameter model do not bottleneck single availability zones. What the sources don't address: Whether Kimi K3's real-world cache-hit rates and latency improvements validate the added engineering complexity of managing explicit cache checkpoints.
The integration provides enterprise teams with a secure, managed route to a massive 2.8-trillion-parameter model optimized for extended context workflows. Support for explicit prompt caching allows developers to significantly reduce latency and operational costs when building retrieval-augmented generation pipelines.
Why it matters
Put this to work — one session a day, built for your industry.
Create a free account for a daily session — eight questions and one real-work challenge, on the news that affects your role.
Start freeHow this developed
20 September 2026
Event created from source cluster.
20 September 2026
Updated with 2 new sources (enterprise_tech) — now corroborated by 5 sources.
20 September 2026
AWS launches Kimi K3 on Amazon Bedrock with 1M-token context window
20 September 2026
Event created from source cluster.
Sources
- kimi-k3-startet-auf-amazon-bedrock-reasoning-modell-mit-1-mio-token-kontextfenster.htmlSonar Track Ping
- Kimi K3 by Moonshot AI is now generally available on Amazon Bedrock - AWSAWS AI Search
- Introducing Kimi K3 on Amazon Bedrock | Artificial IntelligenceAWS AI Search
- aws-kimi-k3-amazon-bedrockdataphoenix.info
- kimi-k3-integration-on-amazon-bedrock-technical-analysis-2026-09-19explore.n1n.ai