Google cuts off Meta's Gemini access after over-consumption, forcing token discipline

Google has limited Meta's access to its Gemini AI models after Meta sought more compute than Google could provide, according to a Financial Times report relayed by Reuters. The cap disrupted and delayed several of Meta's internal AI projects that had become dependent on Gemini capacity.
The internal fallout is telling: Meta reportedly encouraged its staff to be more efficient, swapping a culture of 'tokenmaxxing' — consuming as many tokens as possible — for disciplined token-counting. It's a vivid illustration of compute dependency risk, where relying on a rival's model capacity can become an operational liability overnight.
The story dovetails with a broader week-long theme of AI-cost backlash. Reports that Uber burned its entire 2026 AI budget in four months and Gartner's warning that AI-coding tool costs could exceed developer salaries by 2028 are pushing teams toward model routing and cheaper alternatives. Meta being abruptly throttled by Google reframes the question from 'how good is the model' to 'who controls the tap.'
Developers treated the cutoff as a wake-up call, echoing the Mythos/GPT-5.6 access-gating debates: never hard-wire to a single vendor. The recommended posture is abstraction layers, multi-model fallbacks, and mock APIs to survive both government gates and commercial throttling. Skeptics note the irony of Meta — itself a major open-weight Llama proponent — leaning on a closed competitor's API at all. What to watch: whether Meta accelerates its own model self-sufficiency or strikes a new compute deal elsewhere.