Multimodal reasoning model for long documents and complex code
Gemini 2.5 Pro is Google's long-context thinking model, focused on complex programming, mathematics, and scientific problems, and is also suitable for cross-document analysis and large codebase review. It can combine text and images to understand tasks and produce analysis, code, or structured text. Use Chat Completions on this platform, organizing task materials and conversation history to retain in messages.
Choose an available protocol for this model. OpenAI SDK uses a Base URL ending in /v1; Anthropic SDK uses the root URL. See each guide for protocol-specific parameters, tools and response formats.
Specifications and interface features
Clarify capacity, inputs and outputs, and invocation methods before selecting a model.
Native input limit
1,048,576 tokens
Native output limit
65,536 tokens
Native inputs and outputs
Text, image, audio, video, and PDF inputs; text output
Reasoning and structured capabilities
Thinking, function calling, structured output
Version and knowledge cutoff
Stable version gemini-2.5-pro; model updated in June 2025, with knowledge cutoff in January 2025
Chat invocation
Chat Completions; text and image messages, streaming responses
Native capacity and modality descriptions reflect model capabilities; specific input methods, tool execution, and session features are provided separately by the selected platform entry point.
Core Capabilities
Learn what gemini-2.5-pro can bring to your work.
Bring long materials into a single analytical framework
Long context is suitable for organizing specifications, code, reports, and historical discussions at the same time, comparing constraints and conclusions around the same question. Rather than merely requesting a summary, it is better to specify the contradictions, dependencies, and exceptions that need to be checked, and require evidence identified by section or file, so that analysis of long materials becomes a verifiable deliverable.
Focus on code and multi-step reasoning
Gemini 2.5 Pro's reasoning capabilities are designed for code, mathematics, and STEM problems, and can be used to explain algorithms, analyze failure conditions, or derive solutions. When submitting a task, include input conditions, existing attempts, and acceptance criteria, and request conclusions, assumptions, and validation methods separately to facilitate subsequent testing, rather than merely receiving an explanation.
From image-and-text understanding to structured results
Images can be provided together with text questions to explain screenshots, charts, or diagrams, with results still delivered as text. Structured output makes it easier to organize analysis into records with clearly defined fields; function calls are proposed by the model with parameters, and the application executes the corresponding function, then returns the results and checks the returned data against business rules.
Use Cases
Start with specific tasks to find where the model can make an impact.
Cross-file code review
Provide the relevant modules, interface specifications, error logs, and reproduction steps, and ask the model to inspect call chains, boundary conditions, and potential regressions. Deliverables can include an issue list categorized by file, modification recommendations, and test-case drafts. Retain file paths and version information so developers can locate and validate the recommendations rather than accepting changes directly.
Multi-document clause comparison
Label multiple specifications, papers, or contracts by version, and ask 2.5 Pro to compare definitions, experimental conditions, and conclusions, producing item-by-item comparisons and unresolved issues. Long context helps retain the relevant materials; key conclusions should still include sources, and cross-document references should be checked to ensure they point to passages that genuinely support the judgment.
Chart interpretation and technical review
Submit architecture diagrams, product screenshots, or experimental charts together with background information, clearly specifying the relationships, anomalies, or design trade-offs that need to be explained. The model can generate textual analysis, review outlines, and validation checklists. For critical values, first provide clear original images or text data, which helps distinguish observations in the chart from further inferences.
How to choose this model
Choose based on task complexity, input materials, and expected results.
Prioritize complex analysis; weigh trade-offs for simple tasks
When a task requires maintaining constraints across long materials, analyzing code relationships, or completing multi-step reasoning, Gemini 2.5 Pro is a suitable candidate. For short text classification, simple rewriting, or fixed-field extraction, include the Flash series in comparative testing and decide based on actual accuracy, response experience, and usage; there is no need to choose Pro for every task.
Evaluate stable and preview versions separately
gemini-2.5-pro is the stable version code and is not the same version as gemini-3.1-pro-preview. Existing projects can continue evaluating 2.5 Pro around current prompts, structured outputs, and tool workflows; when considering a version switch, run regression tests using the same materials, and do not directly treat the preview version's parameters or performance as capabilities of this model.
Start with a specific task
Based on the characteristics of gemini-2.5-pro, first validate a small task whose results can be checked.
01
Scientific and technical analysis of long materials
You can ask directly: Synthesize several technical materials, compare definitions, experimental conditions, and conclusions, explain the sources of conflicts, and then list questions for further verification. Do not present correlation as causation.
02
Prepare inputs that support judgment
Prepare the relevant text, charts, and questions; retain citations and calculation bases, and verify whether the long context is actually being used.
03
Then connect it to your workflow
Use the full model ID gemini-2.5-pro, first confirm the public request format and available parameters on the API page, then connect your application. Preserve result parsing, exception handling, and supporting evidence, and use the same set of real samples to evaluate whether it is suitable for continued use.
Usage boundaries
Before formal use, understand output quality and capability scope.
Native multimodal input does not mean that every endpoint accepts the same attachments. Do not treat PDF, audio recording, or video links directly as a general-purpose image-field format.
This model outputs text; it does not generate images or speech, nor is it a Live API real-time audio-video model. You can ask it to explain visual materials, write voice-over scripts, or generate drawing instructions, but actual media generation requires the appropriate creative model.
Knowledge is current through January 2025, so recent facts should be supplemented with new materials. Long context also does not mean it automatically finds every detail; for important terms, code conclusions, and mathematical derivations, request explicit evidence and then verify through human review, testing, or calculation.
Frequently Asked Questions
Answers to common questions about using gemini-2.5-pro.
How much material can Gemini 2.5 Pro process at once?
The native input limit is 1,048,576 tokens, and the output limit is 65,536 tokens; they are counted separately. When organizing long materials, retain chapter and file identifiers, and clearly define the scope of analysis; these are the model's native limits and do not mean every call will return the maximum length.
How do I submit screenshots together with a question?
In Chat Completions, write the message content as an array of content blocks, including both text and image_url. The text should describe the location and task to review, while the image provides visual information; read responses from choices, or concatenate delta.content for streaming calls.
Which option should I use to analyze PDFs?
Prepare the document text, table data, or clear page screenshots relevant to the question, and specify whether you need a summary, comparison, or extraction of particular information. Submit according to the content formats supported by the selected public interface; a PDF URL cannot be used as image_url. Require the results to retain original-text locations, field evidence, and unconfirmed items, and verify key figures against the source materials.
Can Gemini 2.5 Pro run code automatically?
The model natively supports code execution and function calling, but generating code or tool parameters in an ordinary conversation does not mean they have already been executed. When using function tools in platform Chat Completions, the application must execute the function and return the result; which operations tools can perform depends on the execution environment and actual authorization, and the platform's default automatic execution cannot be inferred from the vendor's native tool list.
Do I need to resend the entire history for follow-up questions?
The application maintains the Chat Completions messages history, including relevant user and assistant messages, the latest materials, and feedback. For longer tasks, periodically consolidate interim conclusions and conditions that must be retained before continuing with questions; continuous conversation does not mean unlimited memory.