Google AI Just Released Gemini 3.7 Flash: A Coding and Agent Model at $0.75/1M Input Tokens

Google has launched Gemini 3.7 Flash, the most recent mannequin in its Flash tier, three weeks after Gemini 3.6 Flash. The model card describes it as a refinement of three.6 Flash with algorithmic enhancements to the core reasoning basis — not a brand new pretraining run. It accepts textual content, photos, audio, and video throughout a 1M-token context window, returns as much as 64K output tokens, and helps customizable pondering configurations that commerce high quality towards price and latency. The data cutoff stays at March 2026. The beneficial properties focus in three locations: software program engineering, document-heavy data work, and net growth. The sharper argument is value. Gemini 3.7 Flash ships at $0.75 per 1M enter tokens and $3.75 per 1M output tokens — half the unique 3.6 Flash record charge, and roughly a 3rd the blended price of Claude Sonnet 5 or GPT-5.6 Terra.

Is it Deployable?

Sure, API and enterprise solely. There are not any open weights. Entry runs by hosted surfaces: the Gemini API and Google AI Studio, Google Antigravity, Android Studio, the Gemini Enterprise Agent Platform, and the Gemini Enterprise app. Customers attain it by Gemini Spark on Google AI Professional and Extremely plans.

  • Firm match: Startups and mid-market groups achieve probably the most, as a result of the introductory value makes always-on brokers reasonably priced with no Professional-tier finances. Regulated enterprises get a ruled path by Gemini Enterprise. Groups with data-residency or air-gap necessities are excluded — there may be nothing to self-host.
  • Industries: Google’s personal eval set factors at authorized, monetary providers, biosciences, and enterprise operations. The Harvey LAB-AA, GDP.pdf, and AutomationBench outcomes are the tells.
  • Purposes: Lengthy-running coding brokers, document-heavy back-office automation, UI era from screenshots or design techniques, and PDF-to-structured-data pipelines.

The Benchmark Image

On FrontierCode 1.1 Principal, which measures manufacturing code high quality, Gemini 3.7 Flash scores 43.6% towards 34.4% for 3.6 Flash. On DeepSWE v1.1, a long-horizon software program engineering eval, it reaches 65.3%. On WebDev Arena it posts an Elo of 1588 versus 1538, the highest rating in Google’s comparability desk.

Doc and workflow outcomes transfer additional. GDP.pdf, an professional PDF comprehension eval, goes from 22.0% to 34.0%. AutomationBench, a non-public enterprise workflow set, goes from 17.0% to 30.4% — forward of each Claude Sonnet 5 at 10.7% and GPT-5.6 Terra at 23.6%. Lengthy-context retrieval on GDM-MRCR v2 at 128k reaches 97.0%.

GPT-5.6 Terra is forward on DeepSWE (69.6%), Terminal-bench 2.1 (87.4%), Terminal-bench 3.0 (20.8%), and OSWorld-2.0 (50.2%). On GDPval-AA v2 data work, 3.7 Flash scores 1525 Elo towards 1598 for Sonnet 5 and 1628 for Muse Spark 1.2. CharXiv Reasoning is a regression: 84.5% with out instruments, down from 85.2% for 3.6 Flash. On the Synthetic Evaluation Intelligence Index, 3.7 Flash scores 56, towards 57 for each GPT-5.6 Terra and Muse Spark 1.2.

‘;
h+=’

‘+pts[j].n+’

‘;
}
q.innerHTML=h;
var t=doc.querySelectorAll(“#quad .pt”),l=doc.querySelectorAll(“#quad .ptl”);
for(var m=0;m

Mannequin Blended $/1M Intelligence Index Index per $

“;
var greatest=0;for(var n=0;ngreatest)greatest=r}
for(var o=0;o “+pts[o].n+” $”+pts[o].x.toFixed(2)+” “+pts[o].y+’ ‘+rr.toFixed(1)+”

“}
$(“#vt”).innerHTML=tb;
ping();
}
perform tl(){
var h=””;
for(var i=0;i

‘+TL[i].d+’

‘+TL[i].t+’

‘+TL[i].b+’

Source link

Leave a Reply

Your email address will not be published. Required fields are marked *