Moonshot Kimi
Connect Moonshot's Kimi models — including Kimi K3 with its 1M-token context — to Fabric Agents. Covers the Moonshot AI and Kimi (Coding) presets, model ids, and regional endpoints.
Moonshot's Kimi models are available as Pi-SDK-backed providers in Fabric Agents. Access is via API key.
Kimi K3 is the current flagship: a 1M-token context window, always-on reasoning, and image input.
Which preset to pick
Fabric Agents ships three Moonshot presets. They differ by endpoint, not by account:
| Preset | Endpoint | Use it for |
|---|---|---|
| Moonshot AI | https://api.moonshot.ai/v1 | The full Kimi catalog on Moonshot's global cloud, including kimi-k3. |
| Moonshot AI (CN) | https://api.moonshot.cn/v1 | The same catalog on Moonshot's Chinese mainland cloud. |
| Kimi (Coding) | https://api.kimi.com/coding | Moonshot's coding-optimised subscription endpoint. |
Pick Moonshot AI (or its CN twin) if you want the full model list and K3's 1M context. Pick Kimi (Coding) if you're on Moonshot's coding plan.
Get an API key
- Sign in at platform.kimi.com (or
platform.moonshot.cnif you're on the Chinese mainland instance). - Create a key under API Keys — it starts with
sk-kimi-. - Top up a balance or confirm you're on Moonshot's free trial — keys with zero quota silently return "quota exceeded" on first send.
Connect in Fabric Agents
- Open Settings → AI → Connections → Add.
- From the provider picker, choose Moonshot AI, Moonshot AI (CN), or Kimi (Coding) (see the table above).
- Paste your key.
- Click Test connection. The model catalog populates — on the Moonshot AI presets you'll see Kimi K3 at the top of the list.
- Save.
The connection shows up in the model picker. Sessions created with this connection are routed through the Pi SDK using Moonshot's coding-optimized inference backend.
Model ids
Moonshot AI / Moonshot AI (CN)
| Id | Context | Max output | Inputs | Notes |
|---|---|---|---|---|
kimi-k3 | 1,048,576 | 131,072 | text, image | Current flagship. Always-on reasoning. |
kimi-k2.6 | 262,144 | 262,144 | text, image | Previous generation. Used as the fast summarization model. |
kimi-k2.7-code | 262,144 | 262,144 | text, image | Code-specialised variant. |
kimi-k2.7-code-highspeed | 262,144 | 262,144 | text, image | Low-latency variant of the above. |
kimi-k2-thinking | 262,144 | 262,144 | text | Reasoning-focused K2 variant. |
The K2 preview ids (kimi-k2-0711-preview, kimi-k2-0905-preview, kimi-k2-turbo-preview, kimi-k2.5) remain in the catalog for existing configurations.
Kimi (Coding)
Moonshot moved off version-stamped ids for this product line — the endpoint exposes names that always point at the current model.
| Id | Maps to | Notes |
|---|---|---|
k3 | Kimi K3 | The coding endpoint's route to the current flagship. |
kimi-for-coding | Current coding model | Unified endpoint. Server-side routing — upgrades automatically. |
kimi-for-coding-highspeed | Current coding model, low-latency | Trades some depth for responsiveness. |
k2p5 | (deprecated) | The former version-stamped name for 2.5. Phased out; use kimi-for-coding. |
Fabric Agents auto-migrates older configs — if you set up Kimi with k2p5 in your model list, the next launch replaces it with kimi-for-coding.
Model capabilities
Kimi models in Fabric Agents:
- Use the
anthropic-messagesAPI family, so tool calling and streaming behave the same as Claude connections. - Accept text and image inputs (except
kimi-k2-thinkingand the K2 preview ids, which are text-only). - Support extended reasoning. The thinking-level selector in the input bar maps to Moonshot's
reasoning_effort. K3 reasons on every turn, so the selector controls depth rather than switching reasoning on and off. - Context windows run from 128 K on the oldest previews up to 1 M on K3; see the model-id tables above for exact figures.
Tier picker defaults
Connecting Moonshot AI (or the CN preset) for the first time resolves the 3-tier defaults to:
| Tier | Default |
|---|---|
| Best | Kimi K3 |
| Balanced | Kimi K3 |
| Fast | Kimi K2.6 |
Kimi (Coding) connections pre-fill k3, kimi-for-coding, and kimi-for-coding-highspeed.
Override any of them in Settings → AI → your connection → Models. Sessions use the Balanced tier by default; pick Best or Fast from the model picker per session.
Regional notes
- Moonshot runs two separate clouds — global (
platform.kimi.com; endpointsapi.moonshot.aiandapi.kimi.com) and China (platform.moonshot.cn; endpointapi.moonshot.cn). Your key is tied to one of them; they don't cross-authenticate. If you've got the wrong region on file,Test connectionreturns 401. Use the Moonshot AI (CN) preset rather than editing the global preset's URL. - The China cloud's rate limits are stricter; for heavy automated use, Moonshot's docs recommend a dedicated enterprise key.
Troubleshooting
No kimi-k3 in the model list — either you're on a build older than v0.12.2, or the connection uses the Kimi (Coding) preset (where the id is k3, not kimi-k3). Upgrade the desktop app, or add a Moonshot AI connection for the full catalog.
Test connection returns 401 — wrong cloud (see Regional notes) or the key was revoked. Cycle it on the Moonshot dashboard and paste the new one.
"model not found" when sending a message — your connection's model list still has k2p5 but Moonshot's API no longer accepts it. Open Settings → AI → Kimi (Coding) → Models, remove k2p5, add kimi-for-coding, save. (The auto-migration in v0.8.12+ handles this on launch; this is only a manual step if you skipped the upgrade.)
Long tool-using sessions degrade / lose context — bump the thinking level for that session. Kimi's reasoning_effort tiers substantially affect chain-of-thought depth on complex agentic tasks.
Related
- LLM Providers overview — other provider setups
- Interactions reference — thinking levels, model picker behaviour
- Interactions reference — thinking levels and the model picker
LLM Providers
Fabric Agents works with every frontier model — Anthropic, OpenAI, Google, Moonshot Kimi, and any OpenAI-compatible endpoint. This page covers setup for each.
Azure AI Foundry
Connect Azure-hosted OpenAI-compatible endpoints to Fabric Agents using Microsoft Entra ID (Azure AD) Bearer-token auth. Resource discovery, deployment selection, and token refresh are handled automatically.