Google now offers more than a dozen Gemini models, from text-to-speech voices to Gemini 4 Argon, its new frontier model. If you’re adding AI to your product or to your team’s daily work, you only need to understand a handful of them. This guide covers which ones, what they cost, and what changes for Gemini app users from 9 October 2026.
Two ways to use Gemini
- The Gemini app (web, Android, iOS and Windows). Your team chats with Gemini directly, and you pay per person through a Google AI plan or a Google Workspace plan that includes Gemini.
- The Gemini API (through Google AI Studio, or Gemini Enterprise for larger companies). Your software calls Gemini, and you pay for what you use, by the token. This is how AI features inside apps, websites and back-office systems are built.
Most of this guide is about the second. First, the change that affects the first.
What changes in the Gemini app from 9 October 2026
Google is changing which models each plan can use in the Gemini app on personal Google accounts:
| Plan | Models | Usage limit |
|---|---|---|
| No plan (free) | Flash-Lite only | Standard |
| Google AI Plus | Flash-Lite and Flash | 2x standard |
| Google AI Pro | Flash-Lite, Flash and Pro | 4x standard |
| Google AI Ultra | Flash-Lite, Flash and Pro | 5x or 20x the AI Pro limit |
The change starts on 9 October 2026 for accounts without a subscription. Google AI Plus subscribers will get their date by email. Users can also set an effort level (low, medium or high): higher effort helps on hard tasks but uses more of the limit.
What it means: if your team uses free personal Gemini accounts for work, from 9 October they only get the smallest model. That’s fine for quick drafts and less fine for analysis. If Gemini is part of how your team works, budget for a plan, and keep company data in accounts your company controls. Google’s note covers personal accounts; it doesn’t say how Workspace business plans are affected.
The Gemini models that matter for business software
From Google’s Gemini API model list, these are the ones worth knowing:
Gemini 3.8 Flash: the default choice. Google calls it its “most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows”. It costs $0.75 per million input tokens and $3.75 per million output tokens, an introductory price that runs until 31 December 2026; Google’s standard rates apply after that. Google recommends it, together with 3.5 Flash-Lite, for new projects.
Gemini 3.5 Flash-Lite: for volume. Google’s “fastest, most cost-effective 3.5 model for high-throughput execution”. Use it for simple steps you run thousands of times: sorting emails, tagging support tickets, pulling fields out of forms.
Gemini 3.1 Pro: still in preview. Built for advanced reasoning, complex problem solving and agentic work. Preview models can change, so keep them away from business-critical paths until they’re stable.
Gemini 3.8 Live and Live Extended Thinking: voice. Made for real-time voice agents such as booking lines, first-level support and order-status calls. The Extended Thinking version reasons more before it answers, for conversations that need it.
Gemini 3.5 Transcribe and 3.8 Flash TTS: speech in and out. Transcribe turns speech into text and tells speakers apart, which is useful for meeting notes and call quality checks. The TTS models turn text into natural-sounding voices.
Nano Banana 2 and Nano Banana Pro: images. Image generation and editing for product shots, ad variations and mock-ups.
Gemini 4 Argon: not yet. Google’s new frontier model is limited to cyber defenders for now. Read what Gemini 4 Argon means for your business.
Two housekeeping notes from the same list:
- Gemini 2.5 models are now only available to users who have used them before (they aren’t deprecated). Google points new projects to 3.5 Flash-Lite or 3.8 Flash.
- Google set 1 June 2026 as the shutdown date for Gemini 2.0 Flash. If one of your apps still calls it, move it to a current model now; Google suggests Gemini 3.6 Flash.
A simple way to choose
- Start every new AI feature on Gemini 3.8 Flash.
- Once you can see where the volume is, move simple, repetitive steps to 3.5 Flash-Lite.
- Use the Live models only for real-time voice.
- Move up to a Pro model, or to Argon later, only when your own tests show Flash can’t do the job.
- Keep a model from another company ready as a fallback, such as Claude Sonnet 5.5 or GPT-6.1 Sol.
What it costs: a worked example
Say you add an assistant to your customer portal that answers 10,000 questions a month. Each question sends about 2,000 tokens to the model (the question plus the relevant parts of your help pages) and gets about 500 tokens back.
- Input: 10,000 × 2,000 = 20 million tokens × $0.75 = $15
- Output: 10,000 × 500 = 5 million tokens × $3.75 = $18.75
- Total: about $34 a month on Gemini 3.8 Flash at the introductory price (until 31 December 2026).
From 1 January 2027, Gemini 3.8 Flash moves to $1.50 per million input tokens and $7.50 per million output tokens, so the same assistant costs about $67.50 a month. Budget for that price, not the introductory one.
Real costs depend on how much text you send each time, retries, and any extra features you switch on. Even so, the model is rarely the big number. Building it well is: finding the right data, testing the answers, and handling the questions the assistant shouldn’t answer.
FAQ
Is Gemini free for business use?
The Gemini app has a free tier, but from 9 October 2026 it only includes Flash-Lite on personal accounts. For software you build, the Gemini API is priced per token; check Google’s pricing page for the tier you use.
Which Gemini model is best for coding?
Among the models you can use today, Google positions Gemini 3.8 Flash for long-horizon software engineering. Gemini 4 Argon reports stronger coding results (77.9% on DeepSWE v1.1, according to Google) but isn’t widely available yet.
What’s the difference between Flash and Flash-Lite?
Flash is more capable. Flash-Lite is faster and cheaper. Use Flash where quality matters and Flash-Lite for simple steps at high volume.
Should we wait for Gemini 4?
No. Build on what’s available today and design your software so swapping the model later is a small change.
Need help choosing?
We build AI features into business software and test models on a sample of your real data before you commit. See our AI services or book a free 30-minute call.
Sources
- Google AI for Developers: Gemini models, Gemini API pricing, Gemini deprecations
- Google Gemini Help: Gemini models available by plan
- Google: Introducing Gemini 3.8 Flash and 3.8 Flash Cyber (2 September 2026)
- Google: The latest AI news we announced in September 2026
