**Google’s Latest Gemini Models Prioritise Speed and Cost Efficiency for Enterprise AI Agents**
Google has unveiled a new suite of AI models aimed at enterprise environments, focusing on reducing latency and token costs for autonomous software agents. The release includes Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a restricted variant of Gemini 3.5 Flash designed for cybersecurity tasks. These models represent a shift in optimization, targeting background agents rather than chatbot interfaces, where throughput and efficiency are paramount.
### The Math Behind Gemini 3.6 Flash
Gemini 3.6 Flash focuses heavily on token efficiency. According to Google’s documentation, it produces 17% fewer output tokens than the previous 3.5 Flash version based on Artificial Analysis Index benchmarks. In synthetic tests such as the Datacurve DeepSWE benchmark, token usage dropped by as much as 65%.
The pricing model is set at $1.50 per million input tokens and $7.50 per million output tokens, positioning it for continuous reasoning loops rather than sporadic use. Performance benchmarks show notable gains:
– **DeepSWE success rate:** 49% (up from 37%)
– **MLE Bench score:** 63.9% (up from 49.7%)
– **GDPval-AA v2 score:** 1421 (up from 1349)
### Real-World Integration: Figma, Hebbia, and Harvey
Early adopters include Figma, Harvey, and Hebbia. Figma has integrated 3.6 Flash into its prototyping tools to accelerate design iterations without sacrificing quality. Harvey and Hebbia leverage the model for multimodal document workflows, including parsing financial filings, extracting structured data, and drafting reports.
Google has also embedded a computer-use tool directly into the Gemini API and Gemini Enterprise platforms, eliminating the need for custom intermediary software. According to Google, this boosts the OSWorld-Verified score to 83% and includes enhanced safeguards against misuse.
### A Cheaper Option for High-Volume Agents: Gemini 3.5 Flash-Lite
Gemini 3.5 Flash-Lite is tailored for high-volume, low-latency tasks such as document processing and agentic search. It delivers 350 output tokens per second, the fastest in the 3.5 series. Its pricing is significantly lower at $0.30 per million input tokens and $2.50 per million output tokens, enabling teams to delegate simple tasks to a lightweight model while reserving deeper reasoning for more complex workflows.
Performance improvements are substantial:
– **GDM-MRCR v2 success rate:** 72.2% (up from 60.1%)
– **GDPval-AA v2 score:** Doubled from 642 to 1140
The model also includes the same native computer-use capability as Gemini 3.6 Flash.
### Gemini 3.5 Flash Cyber: Restricted Model for Code Security
Designed for vulnerability management, Gemini 3.5 Flash Cyber validates and remediates code flaws. It is distributed on a restricted pilot basis to governments and vetted partners. Inside Google’s CodeMender security agent, multiple instances of the model operate in parallel, cross-checking findings before generating a final remediation report for human review.
—
## FAQ
**What are Gemini 3.6 Flash and Gemini 3.5 Flash-Lite?**
They are new AI models from Google optimized for enterprise agent workflows. Gemini 3.6 Flash focuses on reasoning and multimodal tasks, while Gemini 3.5 Flash-Lite emphasizes speed and cost efficiency for high-volume background tasks.
**How do these models differ from earlier versions?**
Gemini 3.6 Flash uses 17% fewer output tokens and shows higher benchmark scores. Gemini 3.5 Flash-Lite is the fastest 3.5 variant and offers significantly lower pricing.
**What industries are adopting these models early?**
Figma, Harvey, and Hebbia are integrating these models for prototyping, legal technology, and research workflows respectively.
**Is there a cybersecurity-specific model?**
Yes, Gemini 3.5 Flash Cyber is a restricted model designed to validate and fix code vulnerabilities.
**Where can developers access these models?**
They are available through the Gemini API, Android Studio, Google AI Studio, and the Gemini Enterprise Agent Platform. The new models are also rolling out in the Gemini app and Google Search.
—
## Conclusion
Google’s latest Gemini models reflect a clear shift toward efficiency-focused AI for enterprise use. By decoupling reasoning depth from token volume, teams can now assign lightweight tasks to low-cost models while reserving advanced reasoning for more complex problems. With measurable improvements in speed, cost, and real-world performance, these models strengthen Google’s position in the competitive enterprise AI landscape. As adoption grows, the focus on security, integration, and scalability is likely to drive further innovation in autonomous agent workflows.



