01Microsoft Will Let Copilot Access Local Files and Control Windows, With Features Rolling Out in the Coming Months
Microsoft has demonstrated an expansion of Copilot that will allow the AI system to search local files on a Windows PC and perform actions across the operating system. The company presented the upgrade at its Windows and Surface event, but the capabilities were shown in an onstage video rather than released for general use. Microsoft said the features will reach Copilot over the next couple of months.
The upgrade is part of what Microsoft calls “Hybrid Intelligence”: applications and tools will use a combination of AI models running locally on a PC and models operating in the cloud to complete tasks efficiently. Microsoft did not specify which parts of the demonstrated workflow ran locally and which relied on cloud models.
In Microsoft’s onstage video example of Hybrid Intelligence, Copilot executive Jacob Andreou asked Microsoft’s Autopilot tool for help preparing tax documents after receiving an email from his accountant. Microsoft presented the workflow as illustrating capabilities planned for Copilot. Autopilot searched across different folders for the relevant files, renamed them, compressed them into a zip file, and drafted an email to the accountant with that archive attached. The sequence showed the planned system moving beyond answering questions to finding personal documents and carrying out a chain of actions across files and email.
Microsoft separately demonstrated a new Windows search experience for quicker system actions. From the search bar, users will be able to turn on dark mode, adjust microphone volume with an inline slider, and send a text message. Windows and Surface chief Pavan Davuluri said that experience will begin arriving on Windows 11 PCs this fall, a different schedule from the Hybrid Intelligence features planned for Copilot in the coming months.
The company is also supporting its local-AI push with the Surface RTX Spark Dev Box, a developer-focused mini PC running on Nvidia’s Arm-based RTX Spark platform. The machine has 128GB of unified memory and is optimized for running local AI models exceeding 120 billion parameters. It is available for preorder at roughly $6,000 and is scheduled to ship in November.
Microsoft has not explained the authorization boundaries for Copilot’s file access, how sensitive documents will be protected, or whether users must confirm each action before it is executed. It also has not provided an exact release date for the Hybrid Intelligence upgrade.
02Anthropic Launches Haiku 5.5, Cutting Short-Context Prices to $0.10 and $0.50 per Million Tokens
Anthropic has launched Claude Haiku 5.5, a small model aimed at high-volume, cost-sensitive work such as summaries, classification, database queries and coding subagents. The previous Haiku 4.5 arrived almost a year ago and charged $1 per million input tokens and $5 per million output tokens.
For prompts of up to 100,000 tokens, Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens. Beyond that threshold, the rates rise fivefold to $0.50 and $2.50, respectively. Anthropic says roughly 90% of requests to Haiku 4.5 fell within the lower-priced range.
The company says Haiku 5.5 is its fastest, most capable and cheapest small model yet, and estimates that it costs about 75% less on average to run than Haiku 4.5. That calculation combines request lengths with changes in how many tokens the new model uses for a task, rather than applying the headline price reduction uniformly to every workload.
Independent testing by developer Simon Willison highlights that distinction. His token-counting tool found that the same long prompt used about 1.25 times as many tokens with Haiku 5.5 as with Haiku 4.5 because of the updated tokenizer—the system that breaks text into billable units. The lower rate therefore does not translate directly into an equal reduction in total task cost.
Anthropic positions Haiku for narrow, repetitive or speed-sensitive jobs, including live customer support, browser use, compaction and subagent work. It says the larger Sonnet 5.5 and Opus 5.5 models remain better suited to complex agentic coding.
The launch also changes costs for those larger-model users. Anthropic halved Sonnet 5.5 cache-read pricing from $0.20 to $0.10 per million tokens, which it says cuts costs by about 20% for most agentic tasks because reused cached context accounts for much of their token consumption.
Max and Team subscribers will also receive monthly Claude Platform API credits usable with any Anthropic model: $100 for Max 5x, $200 for Max 20x and up to $500 pooled across Team users. The credits are intended for building API-based tools, applications and agents, but they do not roll over.
03SafeActBench Finds Some AI Agents Exceed 95% in Static Judgment but Reach No More Than 52% in Interactive Execution
AI agents that use tools can change external systems, but reaching the correct outcome does not prove that an action was justified by evidence gathered beforehand. SafeActBench, a new 656-case benchmark, tests that distinction across six operational domains and five protocols, progressing from static judgments to dependency-constrained, multistep workflows.
Researchers evaluated ten combinations of models and execution frameworks, known as harnesses. For three configurations reassessed on a shared set of interactive-action cases labeled V1 in the paper, static action-judgment accuracy was at least 95%, yet interactive execution success was no higher than 52%. The result shows why choosing the right action from a fixed prompt is not equivalent to safely investigating a live task, deciding when the evidence is sufficient and then carrying out the action.
Problems often appeared before execution. Using the paper’s protocol labels, agents stopped before completing the required investigation in 21.7% to 62.9% of V0 episodes, depending on the configuration. In V1 attempts, between 37.0% and 66.9% of actions occurred before the necessary evidence had been established. Those results point to failures in investigation and decision-making rather than simply an inability to operate a tool.
SafeActBench checks that sequence with an Evidence Ledger that binds findings to their sources and records when each piece of information was confirmed. A deterministic trajectory evaluator—not an AI model acting as a judge—then compares that timeline with the moment an action occurred and checks whether later dependencies were satisfied.
A controlled experiment tested whether missing evidence would make agents stop. Across 43 V1 cases and three configurations, researchers withheld one decisive record. The agents still acted in 46.5% to 53.5% of completed episodes, showing that the absence of a crucial fact did not reliably prevent action.
Once agents had obtained the required evidence, single-action execution was usually reliable. Multistep workflows introduced additional failure modes: agents could proceed while prerequisites remained unresolved or leave the overall sequence incomplete. The study does not identify the ten configurations in the supplied source or establish that these benchmark failure rates apply directly to production systems.

ChatGPT Adds Interactive Visuals and Tools to Responses OpenAI is rolling out an “Intelligent UI” that lets ChatGPT combine text with diagrams, charts, forms, buttons, and inline tools. The GPT-6-powered feature is launching globally for paid users before expanding to Go and free tiers. theverge.com
OpenAI Publishes 722 Manuscripts Claiming Hundreds of Math Results OpenAI released 722 manuscripts spanning 372 related result families, including claimed solutions to hundreds of open mathematical questions. Mathematicians are still assessing the work, which was produced by an unreleased frontier model and published with reasoning summaries and compute estimates. theverge.com
Meta and Microsoft Reportedly Curb Employees’ Claude Use Meta and Microsoft are reportedly steering employees toward their own coding tools while retaining Anthropic models in customer-facing services. Microsoft reduced many internal monthly AI spending ceilings, while Meta’s Claude Code user count reportedly fell from about 60,000 to 30,000. rswebsols.com
Google, Meta, and Isomorphic Labs Back a $300 Million Virtual-Cell Project Google DeepMind, Meta, and AI drug-discovery company Isomorphic Labs are jointly investing $300 million in Biohub’s effort to build a predictive virtual model of human cells. The work is part of a $1.8 billion initiative intended to let researchers investigate biological questions through digital simulations. theverge.com
Nous Research Raises $90 Million and Launches Enterprise AI Agents Open-source AI startup Nous Research raised a $90 million Series B at a $1.5 billion valuation and introduced Hermes for Businesses. The product lets companies deploy customized agents for multistep workflows while keeping their data private, according to Nous. techcrunch.com
Youth-Safety Group Calls ChatGPT for Teens an “Unacceptable Risk” Common Sense Media said its tests found unreliable parental alerts, inadequate crisis assistance, and continued homework completion in ChatGPT for Teens. OpenAI disputed the assessment’s methodology, while the nonprofit said it stands behind its findings. theverge.com
Google Opens SynthID Media Verification to Everyone Google launched a public website that checks images, video, and audio for SynthID watermarks added by supported AI-generation tools. Google says the verification system already handles one million requests daily, though watermark detectors can fail to identify some generated media. techcrunch.com
Meta Deploys AI to Find Ads Leading to Child-Abuse Material Meta introduced a language-model system that examines where apparently harmless ads direct users, aiming to uncover links to child sexual abuse material outside its platforms. The company is also deploying an AI red-teaming agent to probe its safety systems for weaknesses. techcrunch.com
Microsoft Unveils Nvidia-Powered Surface AI PCs and Agent Sandboxing Microsoft announced a Surface Laptop Ultra starting at $2,600 and a $6,000 Surface RTX Spark Dev Box, both designed to run AI models locally. A revamped Windows 11 also adds “Execution Containers” for sandboxing AI agents, which Microsoft says will be available to all Windows 11 users. techcrunch.com
Healthleap Raises $38 Million for Hospital Risk-Screening AI Healthleap raised $38 million for software that analyzes clinical notes, laboratory results, and vital signs to flag inpatients who may need further review. The company says the system does not diagnose patients and is deployed in more than 50 hospitals. techcrunch.com
Google Tests Prompt-Based Browser Game Creation Google Labs launched Playground, an experimental platform that generates 2D or 3D browser games from text prompts and uploaded visuals. It is initially available to adults in the US, with free generation limited by weekly tokens and higher allowances for Google One AI subscribers. techcrunch.com
TRACE Reports Up to 5.4× Faster Low-Precision AI Training Rollouts Researchers introduced TRACE, a quantization-aware method designed to align FP4 training and rollout execution for mixture-of-experts language models. Across four large models, they report performance comparable to BF16 rollouts with joint FP4 weights, activations, and KV caches, alongside rollout speedups of up to 5.4 times. huggingface.co