Five Tech Giants Hid $1.65 Trillion in AI Debt Buying Models With No Moat

01Google's Gemini 3.6 Flash Cuts Output Tokens 17%, Up to 65% on One Coding Test

For a developer running an agent in production, the bill comes from output tokens, and Google rebuilt its Flash line around shrinking that count. The company released three models at once: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. None is pitched on leaderboard bragging. All three target the cost, latency, and reliability of agents that fire thousands of times a day.

The headline number sits with 3.6 Flash. Google calls it the workhorse of the line, aimed at coding, knowledge work, and multimodal tasks. According to the Artificial Analysis Index, it uses 17% fewer output tokens than 3.5 Flash. On DeepSWE, a coding benchmark from Datacurve, Google says the reduction reaches 65%. Each token also costs less than before, so the savings stack: fewer tokens, cheaper tokens.

Token count matters because agent workflows loop. A model plans, calls a tool, reads the result, then plans again, and each step spends output. Trim the tokens per step by a sixth, and a long-running job costs less and finishes sooner. That math is why Google leads with efficiency instead of raw capability.

3.5 Flash-Lite chases speed rather than savings. Google says it is the fastest, most cost-effective model in the 3.5 class, clocking 350 output tokens per second by the Artificial Analysis Index. The company says it beats earlier Flash-Lite generations on agentic workflows, the class of task where latency compounds across many calls.

The third model breaks the pattern. Gemini 3.5 Flash Cyber is a specialized, cyber-focused model, and Google ships it paired with CodeMender, its code security agent. The company frames security work as orchestrating a model alongside agent infrastructure, and says the pairing reaches competitive performance at the frontier. Google does not spell out what the Cyber model handles on its own.

One model stayed back. Gemini 3.5 Pro is testing with partners, and Google says it will roll out broadly once it is ready. The team says it is already building the next generation. For now, the three shipped models give teams a cheaper workhorse, a faster lightweight option, and a security-specific pairing, with pricing tied to how many output tokens their agents actually burn.

Production agent bills drop up to 65% on one coding benchmarklatency-bound apps get 350 output tokens/sec from Flash-Litesecurity teams gain a dedicated Gemini model via CodeMenderGemini 3.5 Pro stays partner-only with no ship date

02OpenAI Adds Two Financial Operators to Its Boards, Then Turns ChatGPT Into a Small-Business Product

Within days, OpenAI made two moves that share a subject: the company is building the machinery to run at commercial scale, not just to ship research.

The first move points at revenue. OpenAI opened a ChatGPT for Small Businesses program, pitching entrepreneurs on building AI skills, automating tasks, and growing through ChatGPT Work, according to the company. The framing targets the segment least served by enterprise sales teams and most likely to pay per-seat for automation. It reads as a distribution play: reach millions of small merchants directly, then convert usage into recurring spend. A separate advertising entry point inside ChatGPT sits in the background of that same push, though OpenAI has released no detail on how it would work.

The second move points at governance. David Vélez and Robin Vince joined the boards of both the OpenAI Foundation and OpenAI Group PBC, the company said, citing their experience in finance, technology, and governance. Vélez founded a consumer bank. Vince runs one of the oldest custody institutions in global finance. Neither is an AI researcher, and that is the point: the appointments backfill the institutional depth a company needs when it manages large balance sheets, capital structure, and regulatory exposure across two legal entities.

Placed side by side, the two signals describe one buildout. A monetization channel aimed at small businesses supplies the revenue engine. Two directors with banking and custody backgrounds supply the financial and oversight scaffolding to govern that revenue at scale. The research lab is assembling the commercial and financial infrastructure of an operating company.

That shift changes who OpenAI answers to and who it sells to at the same time. Small businesses become a paying customer class rather than free ChatGPT users. Board oversight moves toward the concerns of financial institutions: capital, risk, and controls. The company is now structured as a business that has to earn money and manage the governance of its own size.

Small merchants become a per-seat paying tier, not free usersnew directors bring banking and custody oversight, not AI researchdual-entity structure (Foundation plus PBC) signals capital and compliance buildout ahead

03Five US Tech Giants Hid $1.65 Trillion in AI Debt. A Viral Essay Argues the Models It Buys Have No Moat

The off-balance-sheet debt at five US technology giants reached an estimated $1.65 trillion, according to a Nikkei study. That is up roughly eightfold in about four years as AI spending climbed. The figure exceeds the companies' reported debt, which Nikkei says makes the risk harder for investors to gauge. Meta alone carries about $420 billion in hidden liabilities, nearly triple its transparent debt. The obligations trace to data center leases and GPU supply contracts, the funding behind the closed, capital-heavy model business.

That spending buys a product a widely shared essay argues is barely defensible. Writing on werd.io and climbing to the top of Hacker News with 1,208 points, the author contends AI models have "very little moat beyond what amounts to brand loyalty and superficial switching costs." A user on ChatGPT today can move to Claude tomorrow with little disruption. In engineering the swap is trivial: same prompt, different API.

The real barrier, the essay claims, sits around the model rather than inside it: the enterprise deals, system connectivity, and quality-of-life features that keep customers in place. Vendors can sign those contracts. But the author says there is little long-term technical reason to stay with one provider once a better model appears.

The piece then turns to where that leaves US policy. Export controls on GPUs and rules against sending certain data to Chinese servers mean Chinese firms can train competitive models but cannot run OpenAI-scale centralized services. So they release open weights instead. "Open almost always wins when it comes to infrastructure adoption," the author writes, because open technologies can be hosted anywhere and used without permission.

The two readings do not resolve against each other. Nikkei documents how much capital is committed to the closed route. The essay is one person's argument that the capital is aimed at the wrong layer.

Off-balance-sheet AI debt now exceeds reported debt at five giantsAPI-level model swapping undercuts vendor lock-in claimsenterprise service contracts, not model quality, may decide who keeps customers
04

US floats sanctions on Chinese open AI models over IP theft Treasury Secretary Scott Bessent said the US could sanction Chinese open-weight AI models over alleged intellectual property theft. The threat widens the administration's effort to slow China's AI progress beyond chip export controls. techcrunch.com

05

Court approves Anthropic's $1.5B copyright settlement A judge granted final approval to Anthropic's $1.5 billion settlement with authors over training data. The deal closes one case but sets no precedent on whether training on copyrighted works is legal. techcrunch.com

06

Data center electricity demand set to quadruple by 2035 A new forecast projects data centers will draw four times more electricity by 2035. Facilities built through 2033 alone could consume as much power as India uses today. techcrunch.com

07

OpenAI and Hugging Face disclose security incident during model evaluation OpenAI and Hugging Face published a joint account of a security incident that occurred while evaluating a model. The companies described how they detected and contained it. openai.com

08

Google ships Flash Cyber to undercut Anthropic's Mythos Google launched Gemini 3.5 Flash Cyber, a model built to find and patch software vulnerabilities. Google positions it as a cheaper option than larger security systems like Anthropic's Mythos. theverge.com

09

OpenAI opens advertising in ChatGPT OpenAI launched an advertising product for ChatGPT through a dedicated ads portal. The move adds a revenue stream beyond subscriptions and API fees. ads.openai.com

10

Jack Dorsey launches Buzz to compete with Slack Jack Dorsey introduced Buzz, a workplace group chat platform that places human employees and their AI agents in the same conversations. It targets Slack's team messaging market. techcrunch.com

11

Gritt raises $34M for construction robots, starting with solar Gritt exited stealth with $34 million to build robots that automate the hardest tasks on construction sites. The company will begin with solar plant assembly before expanding. techcrunch.com

12

Halliday's second smart glasses improve the display Halliday released a second-generation pair of smart glasses with a redesigned display. The update addresses the finicky, hard-to-read window that hampered the 2025 original. theverge.com

13

SWE-Pruner Pro prunes coding context from inside the agent Researchers found coding agents already encode which context is relevant when reading tool output. SWE-Pruner Pro uses a small head on the agent's own internal representations to prune, dropping the separate classifier that prior methods required. huggingface.co