- Improved response quality: Aims to match 2.5 Flash performance.
- Improved instruction following: Targeted improvements to serve as a reliable migration path for complex chatbot and instruction-heavy workflows.
- Improved audio input: Improved audio-input quality for tasks like Automated Speech Recognition (ASR).
- Expanded thinking support: You can control how much reasoning the model performs by choosing from minimal, low, medium, or high thinking levels. This feature lets you balance response quality and speed for your specific use case.
Try in Agent Studio Deploy example app Pricing
| Model ID | gemini-3.1-flash-lite |
|
|---|---|---|
| Modalities |
|
|
| Token limits | Context window | 1,048,576 |
| Maximum output tokens | 65,536 | |
| Capabilities |
|
|
| Tools |
|
|
| Consumption options |
|
|
| Technical specifications | Image |
|
| Text |
|
|
| Video |
|
|
| Audio |
|
|
| Parameter defaults |
|
|
| Supported regions |
|
|
|
||
| Versions |
|
|
| Security controls | Online prediction |
|
| Batch inference |
|
|
| Context caching |
|
|
| See Security controls for more information. | ||
† Listed retirement dates refer to retirement of support in Gemini Enterprise Agent Platform. Models may remain accessible through the Gemini API after these dates have passed. The Gemini API is not a Google Cloud offering and is subject to its own terms of service. For details, see the Gemini API documentation.