Multimodal reasoning model for complex coding and long-form material analysis
Gemini 3.1 Pro Preview is Google's Pro-level preview model for complex reasoning, software engineering, and multi-step tasks. It improves thinking, Token efficiency, and factual consistency over the Gemini 3 Pro series, making it suitable for analyzing code, long-form materials, and visual content together. You can integrate it into applications using the public request format in this page's API section.
Choose an available protocol for this model. OpenAI SDK uses a Base URL ending in /v1; Anthropic SDK uses the root URL. See each guide for protocol-specific parameters, tools and response formats.
Specifications and interface features
Clarify capacity, inputs and outputs, and calling methods before selecting a model.
Native input limit
1,048,576 Token
Native output limit
65,536 Token
Native input types
Text, images, video, audio, PDF
Output types
Text; does not generate images or audio
Native task capabilities
Thinking, function calling, structured output
Image-and-text calling method
Combine text and image_url in messages; images support public URLs or Base64 data URIs
Integration endpoint
Chat Completions
The native capacity and modalities of Pro Preview are not equivalent to all native capabilities made available by the platform. Chat Completions receives the relevant messages from the application; tool execution and result return are handled according to the current documentation.
Core Capabilities
Learn what gemini-3.1-pro-preview can bring to your work.
Reasoning around engineering problems
This model is optimized for software engineering behavior and multi-step execution, making it suitable for analyzing issues using requirements, code snippets, and error logs rather than merely completing a function. You can have it first identify dependencies and constraints, then propose changes, explain the scope of impact, and organize the tests that need verification, producing deliverables that engineers can readily review.
Bring long materials into a single analytical framework
Its long-input capability is suitable for jointly reading multiple modules, design documents, and discussion records, comparing conditions and conclusions across different passages around the same issue. Combined with image understanding, it can also incorporate screenshots, charts, or interface states into the analysis. When asking questions, clearly specifying the order of materials, the focus, and the output structure is more likely to produce usable results than broadly asking for a summary.
Move from natural language to structured workflows
Native function calling and structured output are suitable for connecting reasoning results to business workflows: for example, classifying issues, extracting to-dos, or proposing the next tool request. Chat Completions provides tools and response_format configuration; actual tool execution is handled by the application, and the model continues responding based on the execution results, making it easier to control each action.
Use Cases
Start with specific tasks to find where the model can be effective.
Code review and fault diagnosis
Provide relevant code, error stacks, expected behavior, and recent changes, and have the model analyze possible failure paths, delivering repair recommendations and a regression test checklist. When cross-file dependencies are involved, you can also provide interface definitions and call sites together, asking it to distinguish observed issues from hypotheses that need verification to avoid surface-level changes only.
Interpreting charts and video content
Submit business questions together with image content blocks to interpret report screenshots, explain interface workflows, or compare design mockups. Video understanding can take input through video links in the Gemini conversation interface; request a content summary, key events, and details that need review. The deliverable is textual analysis, not newly generated video.
Continuously advance materials organization
For complex technical investigations, you can first compare materials and implementations, then propose a verification plan, and finally revise conclusions based on new data. 3.1 Pro Preview is better suited to multi-step problems requiring in-depth thought; record the model, request settings, and evaluation samples, and monitor whether behavior changes after preview version updates.
How to choose this model
Choose based on task complexity, input materials, and expected results.
What to consider when migrating from Gemini 3 Pro
Compared with the Gemini 3 Pro series, the improvements in 3.1 Pro Preview focus on reasoning quality, Token efficiency, factual consistency, and usability for engineering tasks. If a task depends on cross-analysis of long materials, complex coding, or multi-step instructions, test it first. During migration, compare conclusions, modification correctness, and tool requests using the same cases; do not assume a version update is necessarily better for every task.
How to choose between general Pro and CustomTools
This model is suitable for general reasoning, coding, and multimodal analysis. There is also an official gemini-3.1-pro-preview-customtools model, which prioritizes using custom tools and is intended for workflows combining bash and tools; tasks that do not depend on these tools may experience quality fluctuations. Do not use the two names interchangeably, and do not always choose the tool-optimized variant for ordinary analysis tasks.
Start with a specific task
Based on the characteristics of gemini-3.1-pro-preview, first validate small tasks whose results can be checked.
01
Multi-step review of complex technical solutions
You can ask directly: Compare two system designs, explain the trade-offs among consistency, resource overhead, and failure recovery, and provide counterexamples and validation experiments. If materials are missing, list questions first.
02
Prepare inputs that support decisions
Record the preview model and configuration; before launch, use stable regression samples to check behavioral changes and error handling.
03
Then integrate it into your workflow
Use the full model ID gemini-3.1-pro-preview, first confirm the public request format and available parameters on the API page, then connect the application. Retain result parsing, exception handling, and relevant evidence, and evaluate whether it is suitable for continued use with the same set of real samples.
Usage boundaries
Before formal use, understand the output quality and capability scope.
Preview indicates a preview version, not a date-fixed snapshot. When used for long-running code review, automated reporting, or tool workflows, it is recommended to retain a representative test set and, after behavioral changes, retest instruction following, output structure, and key conclusions, rather than judging consistency based solely on the model name.
Multimodal understanding does not equal multimodal generation: this model outputs text, does not generate images or speech, and does not support the Live API. It can natively understand audio and PDFs, but submitted media must use methods supported by the corresponding endpoint; file link blocks cannot be used directly as general attachment fields for Chat Completions.
Function calling and programming capabilities do not mean the model can access repositories or run tests on its own. Applications need to execute tool requests, return results, and check the permissions and allowed scope of operations required by each tool. Generated patches, JSON, and analytical conclusions should still undergo testing, format validation, and verification of key facts.
Frequently Asked Questions
Answers to common questions when using gemini-3.1-pro-preview.
Is gemini-3.1-pro-preview a stable or fixed version?
It is Google's Pro-level preview model, with the invocation ID gemini-3.1-pro-preview and no date-based fixed identifier. It is suitable for evaluating complex reasoning and engineering tasks; if your application depends on a stable output structure, save examples and continuously run regression tests rather than treating Preview as a fixed snapshot.
How do I submit images for analysis?
When calling Chat Completions, write the user message content as an array containing both text and image_url blocks. image_url.url can contain a public image address or a Base64 data URI; state the area and objective to analyze in the question, and the returned text is located in choices[0].message.content.
Can it read PDFs directly?
Prepare the document text, table data, or clear page screenshots relevant to the question, and specify whether you need a summary, comparison, or extraction of specific information. Submit content in the formats supported by the selected public interface; a PDF address cannot be used as image_url. Request that results retain original-text locations, field sources, and unconfirmed items, and verify key numbers against the source materials.
Will it automatically execute code and tools?
Tool workflows should be organized according to the tool definitions and result formats of the selected public interface. The model is responsible for planning, explaining results, and generating invocation suggestions; queries, code execution, and writes are completed by the execution environment provided by the application. Actual completion status should come from tool returns and verification records, not solely from the model's description.
Which entry point should I use for multi-turn follow-up questions?
The application should maintain the Chat Completions messages history, including relevant user and assistant messages, the latest materials, and feedback. For longer tasks, periodically consolidate interim conclusions and conditions that must be retained before continuing to ask questions; continuous conversation does not mean unlimited memory.