On October 9, free Gemini users lose access to Google's stronger models. Google's Gemini support page says anyone without a subscription will be limited to Flash-Lite, a version of which Google called its most cost-efficient model in February 2025.
The same month, OpenAI goes the other way. It says ChatGPT will test clearly labeled image ads in the US later in October, shown while users generate pictures. Free ChatGPT stays open. Advertisers help cover the cost.
Both moves answer one bill. Every chatbot answer costs money to run, and a free user pays nothing for it. Google is cutting what each free answer costs. OpenAI is finding someone else to pay for it.
We cannot see either company's internal cost per answer, and neither one publishes it. Google's public API price list is the best window we have. It explains the downgrade in a few lines of arithmetic.
Every Prompt Runs Up a Compute Bill
Underneath all of this sits inference, the price of running a trained model each time someone sends it a message. That cost repeats on every message from every user, for as long as the product runs.
What one heavy free Gemini user costs at Google's list prices
Somebody always settles that bill. The October moves show four ways it gets paid:
- Cash. Google prices AI Pro at about $19.99 12 a month for more access to its Pro models, such as Gemini 3.1 Pro, and AI Ultra at about $99.99. OpenAI keeps Plus, Pro, Business, Enterprise and Edu free of ads.
- Attention. OpenAI's help center says ad testing started in the US on February 9, 2026 08, for a subset of logged-in adult users on the Free and Go plans.
- Quota. OpenAI says it is testing letting free users turn ads off in exchange for fewer daily free messages and reduced features. Google's reportedly forthcoming effort settings, from low to high, use more of the usage allowance at higher effort.
- Quality. From October 9, Google's free users get one lightweight model.
Google's ladder puts the bill in plain view. AI Plus at $4.99 04 gets Flash-Lite and Flash. AI Pro at $19.99 05 gets Pro, so the bigger model costs about double. Read that way, each tier is just a compute allowance with a monthly price attached.
Our read of the evidence is that cost now sets where those tier lines fall, and quality doesn't. Quality still decides who upgrades. Cost decides who gets the expensive model for free.
Pricing the Free User Two Ways
Google's math is short, so we will run it with you. Google's published API list prices put Gemini 3.5 Flash-Lite at $0.30 18 per million input tokens and $2.50 per million output tokens. Flash costs $0.75 20 and $3.75 13. A token is a small chunk of text, often a piece of a word.
A heavy free user might send 30 25 messages a day, each with 1,000 tokens in and 2,000 tokens out. That comes to 900 messages a month. On Flash, that user costs about $7.43 26 a month at list price. On Flash-Lite, the same user costs about $4.77 27.
Before October 9, that heavy free user on Flash could cost roughly 74% 28 of what an AI Plus subscriber pays. After October 9, the same user on Flash-Lite costs about 48% 29. Google saves $2.66 30 a month on one person who pays it zero dollars. Multiply that across a free user base in the hundreds of millions and the decision makes itself.
One caveat. List prices are not Google's internal costs, and real prompts vary wildly in length. Treat the numbers as direction. And the direction is clear: Flash costs about 56% 31 more than Flash-Lite for this user, and Pro costs more again.
AI Plus now gets rationed the same way. Want Pro reasoning? Google's answer is $19.99 05. The free tier stays as the front door, and each step up the ladder buys a bigger slice of compute.
OpenAI splits the job differently. The free user brings the attention, and the advertiser picks up the check. AI Weekly reports that OpenAI puts ChatGPT's weekly audience at 1.2 billion 09 people. That audience is the asset. Ads turn it into revenue without charging the user.
The new image slot shows OpenAI is serious about building the plumbing. Business Standard reports that OpenAI picks ads using the conversation's topic and the user's past chats. Ask for recipes and you might see a meal kit ad.
The early numbers OpenAI shared are upbeat. Its March 26 update said the pilot showed "no impact on consumer trust metrics" and low ad dismissal rates. Within weeks it announced expansion beyond the US. By August 11, ads were live in the UK, Mexico, Brazil, Japan and South Korea.
Skip the benchmark charts. Neither move tells you which model is smartest. Watch the opt-out trade instead. We think it is the most revealing detail in either announcement, because OpenAI has put an exchange rate on attention: see ads and get more messages, or skip them and get fewer.
What nobody knows yet is whether ad revenue covers a heavy user. Light users see few ads and cost little to serve. The heaviest users, with their long image-heavy sessions, cost the most. AI Weekly notes that OpenAI has not disclosed ad pricing, so nobody outside the company can check whether the ads pay for the heaviest free accounts.
Who settles the bill for each free answer
One heavy free user can nearly match a subscriber.
On Flash, a user sending 30 25 messages a day costs about 74% of the $4.99 04 AI Plus fee at list price. Moving that user to Flash-Lite drops the share to about 48% 29.
Ads need a giant audience to pay off.
OpenAI's ad path rests on a reported 1.2 billion weekly ChatGPT users. OpenAI has not disclosed ad pricing, so nobody outside can check whether ads cover the heaviest image-heavy accounts.
Free tiers should keep getting smarter.
If per-token prices fall 40% a year, five years leaves about 8% of today's price, putting Flash output near $0.29 34 by 2031. The newest model, like Gemini 4 Argon, still starts gated at the top.
2031: Rationing Climbs the Ladder
The dates show how fast this is moving. OpenAI's own updates put ads in the US on February 9, 2026 08. Three more markets followed from March 26, and five more were live by August 11. That is nine markets in about six months, and October adds a new ad format.
The other labs moved toward metering over the same months. Freeainews.com reported that Google cut Gemini Pro API prices 40% 10 on May 13, 2026, and tightened some free-tier quotas the same day. The same outlet reported that Anthropic replaced flat-rate agent access for Claude Pro and Max users with a monthly credit pool on June 15. Shattered.io notes that Google announced Gemini 4 Argon to a small group of cyber defenders on September 30, then cut the free app down to one model days later.
Suppose per-token prices keep falling 40% 33 a year. That rate comes from a single price cut, so hold it loosely. Five years of 40% cuts leaves 0.6 multiplied by itself five times, about 8% of today's price.
At that rate, Flash output at $3.75 34 per million tokens lands near $0.29 by 2031. That is about 12% 35 of what Flash-Lite output costs today. Free users in 2031 could get something stronger than today's Flash for less than today's cheapest model costs to run.
The rationing moves up with the frontier. Argon fits the pattern: the newest model starts scarce and gated, then slides down the ladder as cheaper ones arrive. We expect the free tier to keep getting smarter while the newest model stays behind the highest price. The tier lines will hold. What sits behind them will keep changing.
For a builder, the risk is lopsided. Metering a free tier early costs you a few annoyed users. Leaving it open lets one heavy user cost close to what a subscriber pays, as the 74% 28 figure above shows. OpenAI's ad path depends on an audience of 1.2 billion 09 weekly users. Google's path (a cheap model for free users, a bigger one for paying users) is the one a small product can actually copy.
Put a Price Tag on Your Free Tier
If you run an AI tool with a free plan, you can run Google's October math on your own users in one afternoon. Here are the steps in order.
Step one: export seven days of usage per user. Most model providers return input and output token counts on every call. If you do not store a user ID with each call, add that logging today and come back next week.
Step two: price every user. Multiply each user's input tokens by your model's input price, then their output tokens by the output price. Pull both prices from your provider's pricing page. A spreadsheet handles this fine.
Step three: sort users by monthly cost. Compare the top of the list to your cheapest paid plan. If any free user costs more than half of that plan, you have Google's problem on a smaller scale. Check the median and the maximum. The median is the middle user when you line everyone up by cost, and one user who pastes giant documents will skew the average.
Step four: pick the currency for free users. Route free traffic to your cheapest model, cap daily messages, or add an upgrade that unlocks the bigger model. Ads need a large audience to pay for anything, so skip them until your free base is big enough to attract advertisers.
Step five: test the cheap model on yourself. Run your ten hardest real prompts through it. Some will come back wrong. Write down which ones, because that list becomes the pitch for your paid tier.
Then put it on a calendar. Rerun the spreadsheet every Monday for a month. Expect it to break the first time: a missing user ID, a model you forgot to price, a test account inflating the totals. Fix the gap and run it again. By the fourth Monday you will know what each free user costs you, in dollars, and which currency you want them to pay in.
Put a dollar figure on every free user you serve.
- Export seven days of usage. Pull input and output token counts per call with a user ID attached. If calls carry no user ID, add that logging today and rerun next week.
- Price and sort every user. Multiply tokens by your provider's input and output prices, then rank users by monthly cost. Flag any free user costing more than half of your cheapest paid plan, and check the median and maximum.
- Pick the currency for free users. Route free traffic to your cheapest model, cap daily messages, or sell an upgrade to the bigger model. Run your ten hardest prompts through the cheap model and log the failures as your paid-tier pitch.
Every free answer has a payer, so choose yours on purpose.
Google lowers the cost of each free answer, and OpenAI hands the check to advertisers. Both moves answer the same inference bill, and both put cost in charge of where tier lines fall. A small product cannot summon 1.2 billion 09 weekly users, so Google's path of a cheap model for free users is the one to copy. Run the spreadsheet every Monday for a month, and you will know what each free user costs and how they should pay.
