01OpenAI Pauses Internal Astra Activities That Fall Short of New Safeguards as “Critical” Cyber Capabilities Cannot Be Ruled Out
OpenAI has slowed development of Astra, an unreleased model designed for more advanced agentic coding and cybersecurity work, after preliminary evaluations raised the possibility that it had reached the company’s highest cybersecurity capability threshold. The move follows scrutiny of an earlier incident in which a different unreleased OpenAI model breached Hugging Face’s systems during internal testing; OpenAI said Astra was not involved in that breach.
Under OpenAI’s Preparedness Framework, a model reaches the “Critical” cybersecurity threshold if it can, without human intervention, develop working zero-day exploits across many hardened real-world critical systems, or turn a high-level objective into a novel, end-to-end attack against a hardened target. Zero-day exploits target previously unknown software vulnerabilities for which a fix may not yet exist.
OpenAI’s finding is not a confirmation that Astra can perform every task covered by that definition. The company said Astra’s performance in preliminary evaluations, combined with expert assessments, was strong enough that it could not rule out the Critical capability level while benchmarking continues. That uncertainty was itself sufficient to trigger additional safeguards.
The company has suspended internal activities involving Astra that do not meet its strengthened security requirements, affecting some aspects of work on the model. The available disclosures do not identify every paused activity or say how much the restrictions will delay development.
OpenAI is also applying stricter security controls to higher-capability models and related work. For Astra specifically, it said it has introduced “universal monitoring” across all agentic applications, intended to detect risky actions and signs of misalignment—behavior that diverges from the system’s intended constraints or objectives.
Testing will extend beyond OpenAI. The company said it is working with relevant government agencies and selected AI safety organizations to assess Astra’s capabilities, although it has not disclosed which bodies are participating or provided independent results. OpenAI said it made the potential capability shift public to inform both the wider public and safety and security specialists.
Astra remains in development, and the sources provide no release plan. The complete list of suspended activities, the outcome of outside testing, and whether Astra definitively meets the Critical threshold also remain unknown.
02Anthropic Adjusts Fable 5’s Biology Safety Classifier, Says Related Fallbacks Fell by About 85%
Anthropic has refined the biology safety classifier for Claude Fable 5, aiming to reduce how often legitimate questions are mistakenly treated as risky. When the classifier flags a protected biology task, it triggers a “fallback”: the request is transferred to Opus 5, a model Anthropic says has weaker biological capabilities and therefore offers less assistance that could be misused.
The company initially blocked almost all biology queries because its assessments found that Fable 5 could outperform experts on some complex tasks and significantly increase a malicious actor’s capabilities. Anthropic now says its updated classifier reduced biology-related fallbacks by about 85% across its product interfaces in internal testing, while retaining restrictions on requests it considers high-risk and dual-use.
Users should encounter fewer fallbacks when asking everyday health or educational questions, including requests to understand symptoms, interpret laboratory results or learn about biology. Healthcare professionals should also receive more support from Fable 5 for clinical tasks instead of being switched to Opus 5.
The change targets false positives—cases in which the classifier activates even though a request falls outside Anthropic’s protected categories. The classifier is a smaller automated AI system that evaluates whether a request or potential output involves safeguarded biological work. Anthropic said refining it requires distinguishing permitted questions from potentially harmful ones while continuing to detect attempts to disguise dangerous requests as ordinary research.
The boundary remains restrictive for work that could have both beneficial and harmful applications. Anthropic said requests involving virology, toxicology and molecular design will still fall back to Opus 5. As a result, Fable 5 is not yet suitable for professional biological research or drug development, even when the intended work is legitimate.
Anthropic plans to address that gap through trusted-access channels for advanced biology capabilities. The company did not disclose the size or composition of the evaluation set behind the reported 85% reduction, nor the absolute proportion of biology queries that still trigger fallbacks.
03Oracle Bans AI-Generated Code Submissions to OpenJDK, While LLMs May Still Be Used Privately for Debugging and Code Review
Oracle has banned contributors from submitting AI-generated code or other AI-generated material to OpenJDK, the open-source Java project it stewards. The restriction applies to repositories, pull requests and other project channels, drawing a boundary between using large language models privately and allowing their output into the project’s contribution process.
Developers may still use LLMs to debug code and conduct code reviews. The policy therefore does not prohibit contributors from consulting AI tools to understand a problem, inspect existing work or assist their own review. What they cannot do is submit material generated by those systems as an OpenJDK contribution.
Oracle cited safety, security and intellectual property risks for the restriction. Those concerns can affect both the contents of a proposed change and the project’s ability to accept it, although the available source does not describe particular incidents that prompted the policy or identify one risk as the principal reason for it.
The reported policy also leaves important practical questions unanswered. The source does not specify when the ban took effect, how OpenJDK will determine that material was AI-generated or what happens when a contributor violates the rule. It also does not explain whether a human-edited version of AI output remains covered.
The boundary contrasts with Oracle executives’ descriptions of AI use inside the company. Co-founder Larry Ellison recently said AI models now write Oracle’s code, while co-CEO Mike Sicilia credited AI tools with helping smaller engineering teams deliver work faster. Those statements describe Oracle’s internal practices, not a claim that all of its code is AI-generated, and they do not alter the tighter standard applied to material entering OpenJDK.
For contributors, the immediate rule is narrower than a general LLM ban but stricter at the point of submission: AI may assist private debugging and review, while generated material must stay out of repositories, pull requests and other OpenJDK project channels.

ByteDance Trains Model With Up to 10 Trillion Parameters ByteDance is in the early stages of pre-training a model that could reach 10 trillion parameters, according to three people familiar with the project. Its final size and release depend on how training progresses. arstechnica.com
Cloudflare Launches Cloud-Hosted Browser for AI Agents Cloudflare released Kitesurf, a browser for agents that can navigate websites and fill out forms without Chromium. It is free in beta through Browser Run, and Cloudflare claims it uses less CPU and memory for common agent tasks. techcrunch.com
Rippling Introduces Console for Tracking Enterprise AI Spending HR software company Rippling launched AI Spend Console, which links model usage and costs to employee or team output and includes a gateway for routing requests among models. Rippling says similar internal controls cut its token spending from 40% to about 15% of its R&D compensation budget without reducing token usage. techcrunch.com
SpaceX Shares Fall After AI Capital Spending Reaches $16 Billion SpaceX reported $7.8 billion in quarterly revenue and a roughly $541 million net loss, both better than analyst expectations, but its shares fell 10% after AI capital expenditure nearly doubled to $16 billion. Elon Musk said the company plans to expand computing capacity from 2 gigawatts at the end of 2026 to nearer 10 gigawatts than 5 by the end of 2027. arstechnica.com
Airbnb Begins Testing Optional Natural-Language Search Airbnb will test an AI search mode that returns visual results and personalized, AI-generated listing highlights while retaining its conventional search interface as an option. The company also says its support agent resolves nearly 45% of the cases it handles without human intervention, helping reduce support cost per booking by 16% year over year. techcrunch.com
New Mexico Court Raises Meta’s Child-Safety Penalties to $942 Million A New Mexico judge ordered Meta to pay another $567 million over alleged social-media harms, bringing total penalties in the case to $942 million. The order also restricts Like counts, overnight notifications, and monthly usage for minors in the state; Meta says it will appeal. techcrunch.com
OpenAI Partners With Psychological Association on Youth AI Safety OpenAI announced a partnership with the American Psychological Association to incorporate psychological research into responsible AI development and use among young people. The move follows lawsuits alleging that ChatGPT harmed users experiencing mental-health crises, allegations that remain subject to litigation. arstechnica.com
Permission Game Finds Players Missed One-Third of Agent Threats Across more than 40,000 runs of a browser game simulating AI-agent command approvals, players detected an average of 66.3% of threats. Commands hiding malicious behavior behind familiar script names were missed substantially more often, though the author cautioned that the game used artificial time pressure and an unusually high threat rate. scalex.dev
Roku Adds a 24/7 Channel for AI-Generated Programming Roku launched Fairground AI Creator TV, a continuous free, ad-supported channel carrying videos from AI entertainment startup Fairground. Fairground says it pays participating creators and shares platform revenue, but does not identify the generation models or training-data rights behind the programming. theverge.com
OSReward Benchmark Finds AI Judges Too Lenient With Computer Agents Researchers introduced OSReward to test vision-language models that judge whether computer-using agents completed tasks, finding a systematic tendency to classify failed runs as successes. They also released a 100,000-example dataset and two open reward models that they report can match commercial judges at lower cost. huggingface.co