01Kimi K3 Opens 2.8 Trillion Parameters and a Million-Token Context to Developers
Kimi K3 arrived first as a model card with an unusually dense proposition: open weights, native vision, agent tools, and a 1-million-token context window. Moonshot AI released the full model under its Kimi K3 License, giving developers direct access beyond a hosted product interface.
The card drew 1,366 points and 538 comments on Hacker News. That attention preceded the fuller technical report, which supplied the architecture behind the release. Kimi K3 contains 2.8 trillion total parameters, with 104 billion activated for each token. It processes text, images, and video within one model.
Its Mixture-of-Experts system routes each token through 16 of 896 experts. Moonshot combines that design with Stable LatentMoE, Kimi Delta Attention, and Attention Residuals. The company says this package improves overall scaling efficiency by about 2.5 times over Kimi K2. The report attributes the gain to information flow across both sequence length and model depth.
Those choices support the model’s two defining operating ranges. The million-token window targets work across large codebases and long research sessions. Native vision lets the same agent incorporate screenshots, video, designs, and other visual inputs without a separate vision model.
Moonshot says K3 can run long engineering sessions with minimal human oversight, operate terminal tools, and work on tasks including compiler development and GPU kernel optimization. Its model card also lists CAD, chip design, game development, research dashboards, motion design, and video editing. Those remain vendor claims, but the released weights make outside testing possible.
That changes the order of evaluation. Developers can inspect the model card and report, obtain the weights under the published license, and test K3 against their own repositories or multimodal workflows. They can also examine adaptations without routing every experiment through a vendor-controlled API.
The release does not establish deployment cost, hardware requirements, or superiority over closed models. The supplied materials provide no comparable operating data for those questions. They do establish a concrete model-selection option: 104 billion active parameters, native multimodality, agent-oriented tooling, and a million-token context in one downloadable release.
02Windows PCs Can Now Act on Local Files as AI Workers, While Lab Employees Ask Washington to Slow Automation
Perplexity is turning Windows computers into locally run AI agents with access to users’ files and applications. Its Personal Computer product operates as what the company calls a “general-purpose digital worker,” extending a tool first released for Macs in April.
That expansion puts agent software on the world’s most popular desktop operating system. The agent can move beyond generating text and perform work through the applications and data already stored on a computer. Every granted permission therefore expands both its usefulness and the consequences of an error.
Google is pushing the same operating model into production systems. The company announced new Managed Agents capabilities in the Gemini API, including Gemini 3.6 Flash and hooks. Google says the additions help developers build reliable, production-ready agents.
The two releases target different buyers. Perplexity wants an agent working across one person’s computer; Google wants developers embedding agents into products and business processes. Both move AI from answering requests toward executing tasks inside systems that hold real data.
Employees from OpenAI, Anthropic, Google, Meta, Microsoft, Mistral, Thinking Machines, and other labs are asking the US government to intervene as that transition accelerates. Their statement supports a possible slowdown in frontier AI development, or faster work on globally coordinated governance.
The signatories do not represent every employee or an official position shared by their companies. Their statement also creates no policy, pause mechanism, or legal requirement. It asks government to consider restraints while their employers and competitors continue shipping agents with broader operational reach.
That collision is becoming concrete at the permission layer. A chatbot’s bad answer can mislead a user. An agent with file and application access can act on that answer. Managed agents placed inside production services can extend the impact to customers and internal operations.
The next policy fight will concern control rather than conversation: who approves an agent’s permissions, who carries responsibility for failed actions, and which authority can order a pause. For now, product teams are deploying execution capabilities faster than the signatories’ requested governance can become binding.
03Google Raises Its AI Spending Ceiling by $15 Billion as Memory Prices Double
Google raised its capital-spending forecast to as much as $205 billion, up from a previous ceiling of $190 billion. The increase adds up to $15 billion, nearly 8%, before Google reaches the top of its new range.
Even the revised floor is higher. Google now expects at least $195 billion, exceeding its former maximum by $5 billion. The revision arrived during earnings season and unsettled investors, according to The Verge, as spending increased faster than its prior forecast.
That money is moving through three markets at once: corporate budgets, component prices, and semiconductor hiring.
Memory prices have doubled, according to MacRumors. Mac and iPad prices have already risen, while iPhone prices are expected to follow. Those increases put part of the AI buildout’s cost onto buyers who may never purchase an AI subscription.
The pricing pressure reflects demand for the same categories of components needed in data centers and consumer devices. AI operators also face usage costs that conventional software companies can often avoid. Large language models consume paid tokens even when an agent loops, fails, or returns an unusable answer.
Ed Zitron, an AI industry critic interviewed by MacRumors, argues that flat monthly subscriptions obscure those variable costs. That is his assessment, not an established outcome. The doubled memory prices and higher device prices have already occurred.
Labor is shifting inside the chip supply chain as well. MIT Technology Review profiled Lee, a Samsung semiconductor engineer preparing an application to rival SK Hynix after work. He has shared application tips with coworkers who are also considering moves, according to the report.
The account does not establish the scale of departures from Samsung. It does show how competition for AI-related chip production reaches individual engineering teams, not only equipment orders and factory budgets.
Google’s eventual spending, future memory prices, and Samsung’s employee retention will provide the next measurable tests. Each will show whether the infrastructure push keeps absorbing more capital, components, and experienced chip workers.

Amazon Secures a $410 Million Compute Deal With Recursive Superintelligence Recursive Superintelligence will direct much of its budget toward compute for self-improving AI systems. The company aims to automate parts of its own product development. techcrunch.com
Hugging Face Hosts Models That Generate Nonconsensual Nude Images AI Forensics found that seven of Hugging Face’s nine leading image-editing models produced nonconsensual deepfakes. The affected outputs included images of women and children. theverge.com
Insight Partners Invests $200 Million in Bot-Detection Startup Spur Spur Intelligence raised $200 million for technology that distinguishes people from automated web traffic. The product targets companies combating bot abuse and traffic fraud. techcrunch.com
Fish Audio Raises $52 Million for AI Voice Models Fish Audio raised a $52 million seed round to develop voice models for creators and enterprises. Its open-source and hosted products serve eight million users and generate $21 million annually. techcrunch.com
Cursor Cuts Prices and Expands Hiring in India Cursor introduced localized pricing as India became its third-largest market. The coding-tool company also plans to expand local hiring and enterprise sales. techcrunch.com
Runlayer Sues Rippling Over an MCP Gateway Runlayer alleges that Rippling evaluated its MCP gateway before choosing to build a competing product. The startup has filed a lawsuit over the dispute. techcrunch.com
Anthropic Researchers Use Claude to Find Cryptographic Flaws Anthropic researchers prompted Claude Mythos to identify mathematical weaknesses in HAWK and a reduced-round AES variant. The findings do not affect current computer systems. simonwillison.net
OpenAI Documents Scientists Modernizing Research Software With Coding Agents OpenAI published a field report on scientists using coding agents to update scientific computing systems. The report covers software development and research workflows in genomics and other fields. openai.com
JarvisHub Releases an Open Test Harness for Creative AI Agents JarvisHub provides an open harness for agents that produce images, video, audio, interfaces, storyboards, and slides. It tracks drafts, edits, tool actions, evaluations, and human feedback. huggingface.co
Drug Researchers Connect Experimental Results Back to AI Models AI drug-discovery teams are building closed data loops that feed laboratory results into subsequent model training and candidate selection. New drugs still typically require 10 to 15 years. technologyreview.com
Physical AI Developers Test Brain-Wave Data as a Training Signal Researchers are exploring brain-wave recordings alongside multi-camera footage and detailed annotations for physical AI training. The approach targets tasks where video alone cannot capture human intent. techcrunch.com