Kimi K3 and Qwen 3.8 Match Fable 5, Then Give the Weights Away

01Kimi K3 and Qwen 3.8 claim to match Anthropic's Fable 5, then plan to give the weights away

Within a single week, two Chinese labs shipped models they say reach the frontier. Moonshot released Kimi K3. Alibaba released Qwen 3.8. Both companies claim the models run close to Anthropic's Fable 5, the current top of the closed-model market, and both say they will publish the full model weights publicly in the coming weeks.

The performance claim is not the pressure point. The pricing is. The Verge reported the two labs say their models go toe-to-toe with the best from OpenAI and Anthropic at a fraction of the cost. Neither vendor's parity claim has been independently benchmarked, and the numbers come from the labs themselves. But the structure of the release matters more than any single score: a frontier-class model that anyone can download reorders who pays whom for what.

That reordering is what an analysis on Emerging Trajectories walks through. Foundation models cost enormous sums to train, but once trained, the dominant expense is inference, and inference reduces mostly to electricity and data-center compute. Closed labs have priced access on the assumption that reaching the frontier stays scarce. The post argues Kimi K3 and Qwen 3.8 prove the state-of-the-art frontier is attainable with open weights, and it names Anthropic as the company most exposed, because a downloadable equal erodes its product differentiation.

The strategic argument has now become a political fight inside the U.S. camp. Over the weekend, several current and former advisors to President Trump publicly attacked America's leading AI companies, according to MIT Technology Review. David Sacks, the president's AI and crypto "czar," was among the voices in the exchange. The disagreement is over how Washington should answer cheap Chinese open models: subsidize and protect the domestic closed labs, or treat open weights as the terrain the U.S. has already conceded.

Three signals now point one direction. Two labs hit the claimed frontier. They price at a fraction of incumbents. And the people advising the White House cannot agree on a response. The closed-and-expensive model that defined the American frontier is the thing being tested, and the test is public.

Downloadable frontier weights undercut per-token pricing for OpenAI and AnthropicAnthropic named as most exposed on product differentiationU.S. AI policy advisors split on subsidize-vs-concede as parity claims stay unbenchmarked

02The AI That Made People Twice as Confident and Three Times as Wrong

The pitch inside boardrooms is that judgment can be outsourced. One consultant who says he has led his firm's sales and most of its technical engagements alleges that entire companies are gripped by what he calls "AI psychosis," making decisions no rational argument can dislodge. He reports the same paralysis at banks, hospitals, and government bodies, run by people who either have no plan or see no path other than keeping their heads down.

The measured results point the other way. Researchers from three French and Italian universities gave people access to AI advice and watched their judgment fold. Willingness to say "I don't know" fell from 44% to 3%, and accuracy from 27% to 9%. Confidence moved the opposite direction, climbing from 30% to 76%. "People became much worse, the accuracy was only one third, but they were twice as confident," said Valerio Capraro, associate professor at the University of Milano-Bicocca.

The design ruled out sensible delegation. The team picked questions AI usually gets wrong, such as the colour of a team's uniform in Bend It Like Beckham, and ran a model that reliably failed them. Some participants who would have answered correctly on their own asked the AI and got it wrong. Cash incentives barely moved the needle: accuracy recovered only to 16%, still under the no-AI baseline.

That erosion matters most where AI already sits between people and outcomes. New research covered by MIT Technology Review finds large language models are more likely than humans to form biases when screening job applicants. The models absorb human prejudice from training data, and, according to the researchers, manufacture fresh biases of their own. The résumé filter now carries a failure mode the human recruiter did not.

Highest-risk screening — jobs, loans, care — inherits AI's self-generated biasconfident AI users can't sense when their accuracy has droppedpay-for-accuracy nudges failed to close the confidence gap

03Stop swapping models: Augment's Vinay Perneti says the harness decides the code

Vinay Perneti's pitch inverts the usual AI coding sales script. The Augment Code engineer argues that the model matters less than most developers assume. What decides whether an AI writes good code, in his telling, is the harness wrapped around it: the retrieval, the context, and the tooling that feed the model before it types a line.

He calls his case "Beyond grep." Most coding assistants find relevant code by searching for text, the way a developer runs grep across a repository. Perneti argues that keyword search misses the structure a large codebase actually has. A context-rich harness pulls in the right files and the dependencies, then hands the model a fuller picture. Same weights, different input, different output.

The practical consequence lands on cost and quality. Feed a model thin context and it guesses. Feed it a well-built harness and it has more to reason from. Developers chasing better results by upgrading to the newest model, in Perneti's account, are tuning the wrong variable. The engineering that matters sits in the layer they build around the model, not the model itself.

That layer is getting easier to build. The Model Context Protocol, the standard that lets AI tools plug into outside data and services, is moving to a looser, "stateless" approach to session IDs on the server side. Under the change, MCP servers behave more like ordinary websites, which already handle sessions this way. That lowers the bar for anyone wiring a model into their own tools and data.

Put the two together and the daily job shifts. If the harness sets the ceiling on output quality, and the plumbing to build one keeps getting simpler, the work developers invest in stops being model selection. It becomes context engineering: what to retrieve, when to inject it, and how to keep the pipeline fed. Perneti is selling a product that does exactly that, so read the argument with that in mind. The claim is still testable on any repository large enough to break a grep.

Output quality tracks the harness, not the model versioncontext pipelines become the developer's core work, not prompt tweakssimpler stateless MCP lowers the build cost for custom tool integrations
04

Google builds a custom chip to run Gemini cheaper Alphabet is developing a new AI chip aimed at cutting the cost of running Gemini models, per reports. The effort targets inference efficiency rather than training, extending Google's in-house silicon beyond its existing TPU line. techcrunch.com

05

OpenAI publishes failures from running long-horizon models OpenAI detailed safety risks it observed while deploying models that run extended, multi-step tasks. The post describes specific failure modes and the safeguards added through repeated deployment cycles. openai.com

06

Sony sues Udio over 30,000 songs Sony Music filed suit against AI music generator Udio in a New York court, alleging infringement across more than 30,000 recordings. The list runs from Elvis Presley's "Hound Dog" to Beyoncé's "Say My Name" and Harry Styles' "As It Was." theverge.com

07

Bristol Myers Squibb adds a second NVIDIA SuperPOD The drugmaker is deploying a second NVIDIA DGX SuperPOD, built on the Vera Rubin platform, to expand one of the largest AI clusters in life sciences. BMS uses the cluster for drug research workloads. blogs.nvidia.com

08

YouTube tightens rules on AI-generated low-quality video YouTube updated its monetization policies to define which AI-generated and mass-produced videos cannot earn ad revenue. The rules clarify existing guidelines on repetitive and inauthentic content. techcrunch.com

09

Trump's AI standards director resigns The director of the Center for AI Standards and Innovation stepped down, the latest departure from the role since David Sacks left his AI czar post. The position has turned over repeatedly. techcrunch.com

10

Anthropic opens grants for AI in rare disease research Anthropic is accepting applications for research grants under its AI for Science program, focused on rare diseases. The grants fund scientists applying its models to that work. anthropic.com

11

Xiaomi releases a robotics model trained on 100K hours of real trajectories Xiaomi published Xiaomi-Robotics-1, a vision-language-action model for mobile manipulation trained on over 100,000 hours of real-world data. The team reports it handles unseen environments out of the box and adapts to new tasks with minimal fine-tuning. huggingface.co

12

Adobe adds generative AI to its Indigo camera app Adobe updated Project Indigo, its experimental iPhone camera app, with generative AI editing tools. The features do not rely on Adobe's own Firefly models. theverge.com

13

California and XPRIZE test drones against early-stage wildfires A California-backed XPRIZE competition is trialing drones designed to detect and suppress wildfires before they spread. The program responds to fire seasons that now run most of the year. arstechnica.com