Complex-task reasoning model for long-horizon programming and visual analysis
Kimi K3 is a Kimi series model by Moonshot AI, designed for long-horizon programming, complex reasoning, agents, and knowledge work. It can analyze problems using text and images, generating code, review feedback, or structured results. On this platform, you can use Chat Completions to precisely manage messages and tools, or use AI Chat v2 to build hosted multi-turn workflows.
Choose an available protocol for this model. OpenAI SDK uses a Base URL ending in /v1; Anthropic SDK uses the root URL. See each guide for protocol-specific parameters, tools and response formats.
Specifications and interface features
Clarify capacity, input/output, and invocation methods before selecting a model.
Input methods
Text and images; images use image_url content blocks
Output formats
Text, JSON structured content, and function calls; structured content should be parsed and validated by the application and is not guaranteed to strictly conform to the specified JSON Schema
Reasoning mode
Reasoning is continuously enabled; reasoning_effort: max is recommended but may be omitted
Response methods
Standard JSON responses or SSE streaming output
Dedicated endpoint
POST /kimi/chat/completions; model is kimi-k3
Hosted conversations
/aichat2/conversations; supports continuing chats by conversation ID and conversation management
Vision and reasoning are model capabilities; hosted conversations, file reading, and authorized tool execution are workflow features of the selected platform entry point.
Core Capabilities
Learn what kimi-k3 can bring to your work.
Continuous analysis around engineering constraints
K3 is suited to analyzing requirements, related code, error logs, and test conditions together, rather than generating isolated functions. You can ask it to first identify the scope of impact, then propose fixes and validation steps. Retaining file paths, dependencies, and non-modifiable constraints in the input helps produce engineering deliverables that are easier to review.
Turn visual references into implementation specifications
Combined with screenshots and written requirements, K3 can be used to analyze page layouts, component hierarchies, and interaction intent, then generate UI code or redesign recommendations. It is suitable for starting frontend prototypes from visual references; you should also provide the tech stack, design guidelines, and responsive requirements to avoid asking the model to guess behaviors not shown in the screenshots.
Bring analysis results into tool workflows
By defining fields for structured output, review comments, task lists, and extracted results can be integrated into downstream programs. Function calling allows the model to propose tool names and parameters. Dedicated endpoints are suitable for controlling the execution loop yourself, while AI Chat v2 is suitable for using managed tool workflows; both should retain tool results and failure information.
Use Cases
Start with specific tasks to find where the model can be effective.
Cross-file defect investigation
Provide the observed issue, relevant files, stack traces, and existing tests, and let K3 map the call chain, propose root-cause hypotheses, and deliver modification recommendations, patch drafts, and a regression test checklist. This is suitable for tasks that require connecting multiple modules; continuing to feed back actual test results can help it revise its initial assessment instead of stopping at a single guess.
Screenshot-driven frontend prototypes
Submit page screenshots, component specifications, and the target framework, and ask K3 to output layout descriptions, component code, and interaction check items. Then use browser screenshots to provide feedback on deviations, adjusting spacing, hierarchy, and responsive behavior iteratively. The deliverable is prototype code that can continue to be developed; visual similarity should not be treated directly as passing functional and accessibility acceptance.
Multi-document knowledge organization
Organize related documents into text with titles, sources, and section identifiers, and ask K3 to deliver topic summaries, clause comparisons, or items requiring clarification. When using AI Chat v2, you can also provide file links to enter the reading workflow. Clearly distinguishing source material, derived conclusions, and missing information makes reports easier to trace and review.
How to Choose This Model
Choose based on task complexity, input materials, and expected results.
Choose K3 for Complex Engineering; Allocate Routine Tasks as Needed
When a task requires multi-step analysis, understanding related code, or repeatedly using tool results, K3 is a better fit. K2.6 and K2.5 can serve as alternatives for general conversation and content tasks. Do not directly reuse the Thinking toggle configuration from the K2 series; for short rewrites and simple extraction, compare delivery quality, response wait time, and actual usage before deciding.
Choose an Entry Point Based on State Management
If you already have an OpenAI-style client and need custom tool execution and message history, choose /kimi/chat/completions. If you want to continue conversations through a session ID, read files, and observe tool events, choose AI Chat v2. The former requires the application to maintain the full message history, while the latter can host sessions; internet access and external operations should be designed according to the relevant tools and authorization conditions.
Get Started
From a small-scale task to production integration.
01
Prepare the Task and Materials
Clarify the goal, required inputs, and output requirements, using real business examples as a starting point.
02
Try It in the API Testing Area
Open the trial page, confirm the parameters supported by this entry point, then submit a small-scale task to review the results.
03
Integrate According to the API Documentation
Keep the complete model ID, use the request format specified in the documentation, and confirm billing rules on the Pricing page.
Usage Boundaries
Before formal use, understand output quality and capability limits.
K3 performs reasoning continuously and is not suitable for using disabled thinking as the default strategy for shortening responses. It is recommended to use max or omit reasoning_effort; do not rely on other values to obtain different speed tiers. For long tasks, you still need to control the relevance of materials and set a reasonable output budget for the final answer.
Image understanding is not the same as image generation, and converting screenshots to code does not mean that the page has been run or verified. Blurry text, hidden interactions, and states outside the screenshot should be described separately; generated code needs to be tested in the target environment, especially for responsive layouts, keyboard operation, and business logic.
When a specialized entry point returns a function call, the application needs to execute the tool and return the result; it cannot only save the final text. External reading and writing in AI Chat v2 depend on connections and authorization; unattended write operations also require explicit permission, and a model's proposed action plan does not mean the relevant operations have been completed.
Frequently Asked Questions
Answers to common questions about using kimi-k3.
Can Kimi K3 disable reasoning?
K3 keeps reasoning continuously enabled. It is recommended to set reasoning_effort: max; if omitted, it is used this way as well. Do not copy the thinking switch from K2.6 into K3 requests. If the task is only brief rewriting or classification, consider a model better suited to lightweight tasks rather than relying on disabling reasoning.
Can K3 view screenshots and output webpage code?
You can submit screenshots as image_url image blocks together with framework, style specifications, and interaction requirements for interface analysis and code generation. It outputs text code and suggestions; it does not automatically run the webpage as a result. Verification should be completed through actual rendering, screenshot comparison, and interaction testing.
What content needs to be saved for multi-turn tool calls?
When using a dedicated endpoint, you should save and return the complete assistant message, including the returned reasoning_content, tool_calls, as well as the corresponding call IDs and tool results. Keeping only the final answer will lose execution state. When using a hosted session, continue the task through the same session ID.
How can K3 return results that programs can process?
You can explicitly request JSON output in the prompt and specify field names, types, required fields, and allowed values, with examples; the application should perform JSON parsing, structural validation, and business validation, and set up retries and exception handling for missing fields, type errors, or truncated results. If using response_format to configure the format, first verify the actual effect of the selected mode in kimi-k3 requests, and do not treat the json_schema field as a guarantee of strict Schema compliance.
How should materials be submitted for PDF analysis?
When a file-reading workflow is needed, you can use the AI Chat v2 file_url content block to provide an accessible file link; alternatively, extract the text first and then submit it to a dedicated endpoint for analysis. File reading is an endpoint feature and does not mean the model directly accepts arbitrary formats; the clarity of scanned content also affects analysis.
Model information · Updated: 2026-10-01. For request parameters and billing rules, see the API and pricing sections.