Google launches Gemini 3.6 Flash and 3.5 Flash-Lite for enterprise agents

Gemini 3.6 Flash cuts latency and token costs for enterprise AI agents See how Google’s new models boost speed, efficiency, and production workflows

Google has introduced Gemini 3.6 Flash and Gemini 3.5 FlashLite, two new models aimed at reducing latency and token costs for enterprise AI agents. The company says the models are designed for workloads that run repeatedly in production, where output efficiency and speed matter as much as reasoning quality. According to Google, Gemini 3.6 Flash uses fewer output tokens than its predecessor and is priced at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. Google says the model improved results on several benchmarks, including coding and multimodal tasks, and has already been adopted by companies such as Figma, Harvey, and Hebbia for design, legal, and documentprocessing workflows. Gemini 3.5 FlashLite is positioned as a lowercost option for highvolume tasks such as document processing and agentic search. Google also announced a restricted Gemini 3.5 Flash Cyber model for vulnerability remediation, available only to governments and vetted partners. The models are accessible through Google’s Gemini API, AI Studio, Android Studio, and Gemini Enterprise platforms, with FlashLite also rolling out in Google Search.