Google’s Gemini 3.6 Flash targets enterprise agent token costs

Google’s Gemini 3.6 Flash targets enterprise agent token costs

Google has launched Gemini 3.6 Flash and three.5 Flash-Lite as new workhorses designed to chop latency and token prices for enterprise AI brokers.

The economics of working autonomous software program brokers inside a manufacturing surroundings come all the way down to a hard and fast equation few distributors promote straight. A mannequin must cause via a multi-step activity competently, however each additional token it generates whereas doing so provides price and delay to a workflow which may run 1000’s of occasions an hour.

Groups constructing background brokers moderately than chat interfaces want throughput first and parameter depend second. Google’s reply, introduced this week, splits that trade-off throughout three fashions: Gemini 3.6 Flash for coding and multimodal reasoning, Gemini 3.5 Flash-Lite for high-volume, low-latency work, and a restricted Gemini 3.5 Flash Cyber variant constructed for vulnerability remediation.

The mathematics behind Gemini 3.6 Flash

Google’s developer documentation for 3.6 Flash centres on one determine: 17 % fewer output tokens than the prior 3.5 Flash model, primarily based on measurements from the Synthetic Evaluation Index.

In particular artificial exams, together with the Datacurve DeepSWE benchmark, Google studies drops in token utilization of as much as 65 %. Pricing sits at $1.50/1M enter tokens and $7.50/1M output tokens, positioning the mannequin for reasoning loops that run repeatedly moderately than on-demand.

On DeepSWE, the corporate information a 49 % success fee for 3.6 Flash in opposition to 37 % for its predecessor. On MLE Bench, the rating strikes from 49.7 % to 63.9 %, and on Google’s GDPval-AA v2 check – which makes an attempt to measure real-world data work moderately than coding puzzles – 3.6 Flash scores 1421 in opposition to 1349 for the older mannequin.

Figma, Hebbia, and Harvey put the mannequin to work

Figma has built-in 3.6 Flash into its prototyping infrastructure, and in line with Matt Colyer, the corporate’s Director of Product Engineering, the mannequin offers builders a quicker route via design iterations with out a drop in output high quality.

Authorized expertise platform Harvey and analysis device Hebbia route information via the mannequin for multimodal doc work: ingesting uncooked monetary filings, parsing doc construction, studying embedded charts, and producing draft studies for evaluate.

Google additionally folded a client-side computer-use device straight into the Gemini API and Gemini Enterprise platforms, eradicating the customized middleman software program engineers beforehand constructed to let fashions function on prime of an working system.

The corporate studies an OSWorld-Verified rating of 83.0 %, up from 78.4 %, and says up to date safeguards in opposition to chemical, organic, radiological, and nuclear misuse enhance resistance to jailbreaking with out elevating refusal charges for benign requests.

A less expensive tier for high-volume background brokers

Gemini 3.5 Flash-Lite targets a distinct job: doc processing and agentic search working at quantity moderately than reasoning depth. The Synthetic Evaluation Index measured the mannequin at 350 output tokens per second, the quickest within the 3.5 sequence in line with Google.

Pricing runs at $0.3/1M enter tokens and $2.5/1M output tokens, low-cost sufficient that engineering groups can route easy, high-volume subagent requests to a minimal pondering degree and reserve increased pondering ranges for multi-step work.

On Google’s GDM-MRCR v2 long-context check, Gemini 3.5 Flash-Lite recorded a 72.2 % success fee in opposition to 60.1 % for its predecessor, and its GDPval-AA v2 rating almost doubled, from 642 to 1140. The mannequin carries the identical native computer-use device as 3.6 Flash.

Individually, Google says Gemini 3.5 Professional stays in companion testing forward of a full launch, and pre-training for the subsequent Gemini 4 structure is already underway.

Gemini 3.5 Flash Cyber: A restricted mannequin for patching code

Automated vulnerability scanners now floor flaws quicker than most safety groups can patch them, and that hole is the place Google positions Gemini 3.5 Flash Cyber.

The mannequin is constructed to validate and remediate code vulnerabilities, and Google studies efficiency on the CyberGym benchmark aggressive with frontier fashions (although it hasn’t made these figures public in the identical element as its consumer-facing releases.)

Distribution stays restricted to governments and vetted companions via a pilot programme, a limitation Google frames as a safeguard in opposition to the mannequin producing exploit code for offensive use.

Inside Google’s CodeMender safety agent, a number of cases of three.5 Flash Cyber run in parallel, cross-checking each other’s findings earlier than producing a single remediation report a human reviewer indicators off on.

Engineering groups searching for to combine these new fashions can entry them via the Gemini API by way of Google AI Studio, Android Studio, or the Gemini Enterprise Agent Platform. Shoppers may also entry the brand new fashions within the Gemini app and three.5 Flash-Lite can be rolling out in Google Search.

See additionally: Bristol Myers Squibb buys Nvidia AI system for drug discovery

Need to study extra about AI and massive information from trade leaders? Try AI & Big Data Expo happening in Amsterdam, California, and London. The excellent occasion is a part of TechEx and is co-located with different main expertise occasions together with the Cyber Security & Cloud Expo. Click on here for extra data.

AI Information is powered by TechForge Media. Discover different upcoming enterprise expertise occasions and webinars here.