01Google launches Gemini 3.8 Flash Cyber, opens automated vulnerability finding and patching to more than 650 government and enterprise partners
Google has launched Fairwind, a restricted-access program that gives more than 650 government agencies, companies and cybersecurity partners worldwide access to Gemini 3.8 Flash Cyber. The specialized model is designed to help trusted defenders automatically find, verify and patch software vulnerabilities.
The program targets national cyber authorities, critical infrastructure operators in sectors including healthcare, telecommunications, energy and finance, and technology platforms whose software supports downstream users. Participants must restrict access to employees working in internal cybersecurity, incident response or penetration testing and deploy protections such as multi-factor authentication.
Gemini 3.8 Flash Cyber shares its foundation with the general-purpose Gemini 3.8 Flash but has more permissive cybersecurity safeguards, which is why Google is limiting it to trusted defenders. Through Fairwind, the model works with CodeMender, Google’s system for writing and validating code fixes, inside an organization’s secure cloud environment.
Google says the combination can discover weaknesses, check whether they are genuine, generate a fix and validate the resulting code. It says defenders can produce verified, deployment-ready patches in minutes instead of spending weeks fixing vulnerabilities manually. That is a claimed workflow improvement, not evidence that every vulnerability can be completely repaired within minutes.
The company reported a success rate above 70% on an internal vulnerability-discovery benchmark covering complex codebases in 20 programming languages. On CWE-Bench, an external test of patching capabilities, the model achieved a pass@1 score of 47.2%, meaning it produced a passing patch on its first attempt in that share of tests. A leading frontier model scored 47.8%, according to Google, but at significantly higher cost. Google’s Gemini announcement
Google and its partners also supplied operational examples. The Chrome Security team said the model produced 2.6 times as many correct vulnerability patches as the best large commercial models it compared. Cybersecurity company Wiz reported recall 7.5 to 9.7 percentage points higher on its internal penetration-testing benchmark at 2.3 to 5.2 times lower cost. Google’s Cloud Vulnerability Research team said the model found an unspecified critical foundational vulnerability in under two hours, compared with the months such research usually takes.
Customers outside Fairwind will not receive the restricted Cyber model, but any Google Cloud customer can use CodeMender with publicly available models hosted on the Gemini Enterprise Agent Platform. Google has not disclosed Fairwind’s complete partner list, pricing or detailed admission criteria, and the reported performance and cost results lack independent verification. Google DeepMind’s Fairwind announcement
02Enterprise’s self-hosted small model takes over half of internal AI traffic, team says it handles 116 million requests a month
An enterprise AI team says it has consolidated requests from more than 200 internal applications onto one smaller, self-hosted language model. Data-residency requirements had pushed the organization to run models on its own infrastructure, but adding newer systems without retiring older ones fragmented its finite pool of GPUs.
The team now says the resulting model handles 50% of platform traffic, or 116 million requests each month, at a fraction of the previous serving cost. The organization was not identified, and the paper did not disclose the model’s exact parameter count or provide a full calculation of the cost reduction.
The consolidation effort began with an analysis of errors in real production traffic. Researchers grouped the main quality gaps into three areas: following instructions, calling software functions correctly, and handling the particular mix of tasks generated by the company’s internal applications. They then built offline evaluations reflecting the production traffic mix, using deterministic checks or calibrated language-model judges.
Rather than optimize all three objectives together, the team trained a separate GRPO expert—a specialist model refined against rewards—for each area. The researchers said joint optimization caused rewards from different domains to interfere with one another. They combined the three specialists through a two-stage process called SLERP, a technique for merging model parameters.
Training the experts separately also exposed different ways they could exploit their rewards. The paper identified semantic collapse, excessive function calling and overly long answers designed to win higher scores. According to the team, each failure mode required its own targeted correction before the experts were merged.
In non-reasoning mode, the combined model scored 69.6 on the organization’s internal Arena evaluation, compared with 65.8 for a baseline with about seven times as many total parameters. It also surpassed that baseline on instruction following, scoring 0.85 versus 0.83, and function calling, scoring 0.79 versus 0.77. These are internal, production-related evaluations rather than evidence of general superiority, and the paper did not report how closely its language-model judges agreed with human reviewers.
03ZimaBlue Trains Robots on 120,000 Hours of First-Person Video, Lifting Zero-Shot Manipulation Success From 36.1% to 77.8%
A research team says ZimaBlue, a framework for training “world action models,” can turn large volumes of first-person video into robot-control capabilities. Such models learn how physical interactions unfold and connect that experience to actions a robot can execute.
Robot manipulation systems need varied physical experience to handle unfamiliar objects, tasks, and environments, but collecting robot trajectories with action labels is expensive and offers limited diversity. First-person videos are easier to scale and capture object interactions, contact dynamics, tool use, and longer activities, although they do not contain commands that a robot can directly follow.
ZimaBlue bridges that gap through a three-stage training process. It first performs causal embodied-video pre-training on large collections of human and robot first-person footage, learning reusable representations of how interactions develop over time. The second stage uses heterogeneous robot trajectories and a unified action representation to ground those learned visual dynamics in executable robot actions. The model is then specialized for the target robot intended to run it.
The stages separate two problems: acquiring broad physical experience from scalable, action-free video, and translating that experience into the controls of a particular machine. The target-robot specialization prepares the resulting model for deployment, but the researchers do not report commercial use.
ZimaBlue also divides inference between asynchronous “Slow” and “Fast” branches. A high-capacity Slow world model supplies generalizable spatial and temporal representations, while a lightweight Fast branch predicts actions frequently enough for real-time control. The team reports that the Fast branch reaches 30 Hz on an Nvidia RTX 4090.
In real-robot zero-shot evaluations—tests conducted without task-specific training—the reported success rate rose from 36.1% when training used only target-robot data to 77.8% when the dataset expanded to more than 120,000 hours of embodied video. The researchers also report strong results across multiple benchmarks, with especially pronounced gains on previously unseen tasks.
That 77.8% result is limited to the paper’s real-robot zero-shot evaluation, not a universal success rate for robot manipulation. The available source does not specify the robot models, number of tasks, trial counts, failure types, human-to-robot video ratio, deployment cost, safety performance, or long-term reliability.

Trump Administration Backs OpenAI’s Fair-Use Argument in Times Lawsuit The administration filed a statement of interest supporting OpenAI’s argument that training language models on copyrighted text can qualify as fair use. The New York Times alleges that OpenAI and Microsoft unlawfully used its articles and seeks billions of dollars in damages. theverge.com
OpenAI Faces 30 New Lawsuits Over Tumbler Ridge Shooting Students and educators filed 30 federal lawsuits accusing OpenAI and CEO Sam Altman of substantially assisting the alleged shooter by failing to act on flagged ChatGPT conversations and prevent renewed access. OpenAI strategy chief Jason Kwon called claims about the company’s safety decisions false. theverge.com
New York City Restricts Classroom AI Through Eighth Grade New York City announced a one-year moratorium covering roughly 600,000 public-school students in 2-K through eighth grade beginning in the 2026–2027 school year. The rules also prohibit AI grading, ban companion chatbots across all grades, and limit high-school AI use to approved cases and pilots. theverge.com
Palo Alto Networks Reportedly Paid $500 Million for IT-Agent Startup Console Palo Alto Networks acquired Console, whose AI agents automate help-desk work such as password resets and troubleshooting, for $500 million in cash and stock, according to TechCrunch sources. The cybersecurity company plans to integrate Console into its Cortex platform for natural-language alert investigation and resolution. techcrunch.com
Wonderful Raises $550 Million at a $5 Billion Valuation Israeli-Dutch startup Wonderful more than doubled its valuation in under six months with a Series C led by Insight Partners and joined by Salesforce. Its Wonderful AI OS connects enterprise agents and workflows with company data and existing systems, and the funding will support product development and deployment teams. techcrunch.com
HiddenLayer Secures $100 Million for AI Runtime Security HiddenLayer raised a $100 million Series B led by Delta-v Capital after reporting that annual recurring revenue grew more than tenfold over the past year. Its tools protect models, agents, and workflows against threats including prompt injection, agent manipulation, malicious tool use, and compromised open-source model files. techcrunch.com
Amazon Adds AI-Based Verification for Messages Claiming to Be From Amazon Alexa for Shopping can now compare an email, text, or phone call with Amazon’s messaging records while checking its sender, formatting, and contents. Amazon says the assistant confirms authenticity only when completely certain and otherwise directs customers to their orders and official support. theverge.com
Reliance Jio Opens Its Cloud PC Service Across India Reliance Jio made JioPC available to Indian internet users regardless of their broadband provider, streaming a virtual computer with configurations of up to eight virtual CPUs, 16GB of RAM, and 1TB of storage. Jio says the subscription can give computers as old as eight years access to cloud capacity for modern AI applications. techcrunch.com
OpenAI Endorses California Youth AI-Safety Bill OpenAI announced its support for California SB 1119, describing the bill as establishing age-appropriate AI safeguards for teenagers while preserving opportunities to learn, create, and explore. openai.com
Google Reportedly Seeks AI-Training Licenses From Major Hollywood Studios Google has approached studios including Disney, Warner Bros. Discovery, and Universal about licensing copyrighted libraries for AI training, according to unnamed sources cited by the Los Angeles Times; no agreements have been announced. The reported proposals could include character-specific payments and shares of advertising revenue from AI-generated content. theverge.com
Qwen-Drive Unifies Perception and Planning for Autonomous Vehicles Researchers introduced Qwen-Drive-1.0, a vision-language foundation model that combines 3D perception, driving-scene question answering, and future-trajectory generation while retaining the underlying pretrained architecture. They report competitive motion-planning results across open-loop, pseudo-closed-loop, and closed-loop evaluations. huggingface.co
UI-Venus-2 Targets Automation Across Mobile, Web, and Desktop UI-Venus-2 is an open-source multimodal agent built around a unified reasoning-and-action loop, with coverage spanning more than 170 multilingual mobile apps and native desktop systems. Its training framework uses trace- and sample-level verification plus safety mechanisms for consequential actions. huggingface.co