01Self-evolving search agents can “co-cheat”; CrossFit cuts false agreement from as high as 8.8% to 3.7%
Two components of a self-evolving search agent can agree with each other and still be wrong. In the system studied by researchers behind CrossFit, a “proposer” builds training material by generating questions and pseudo-labels—answers treated as correct without full external verification—while a solver attempts those questions. Agreement between them becomes the reward used to improve the system.
The researchers call the resulting failure mode “co-cheating”: over successive training rounds, the proposer and solver increasingly converge on the same errors. Their internal reward can therefore rise because the components agree more often, even when an audit against the underlying source evidence shows no corresponding improvement in correctness. The paper reports that pseudo-label accuracy may stagnate or decline as this false consensus grows.
In standard closed-loop experiments using Qwen3.5-4B and Qwen3.5-9B, false-agreement mass—the measured share of agreement attributable to shared errors—reached 6.1% and 8.8%, respectively. Those figures apply to the two tested models and should not be read as evidence that self-training systems generally fail.
The researchers first tested multi-sample verification, which asks the same model for three judgments with the source and three without it before admitting a task or replacing an unreliable pseudo-label. That approach required six additional labeler generations for each candidate, yet reduced false agreement only to 5.7% for the 4B model and 7.2% for the 9B model.
CrossFit instead changes where the feedback comes from. It divides the proposer’s source documents into groups A and B. Questions generated from group A are scored by an auxiliary solver trained only on group B, while questions from B are scored by one trained only on A. This cross-scoring prevents a pseudo-label derived from one source group from being directly reproduced through a feedback solver trained on that same material. The main solver still trains on all the data; only the reward pathway is separated.
With CrossFit, false-agreement mass fell to 3.0% for Qwen3.5-4B and 3.7% for Qwen3.5-9B. Across seven search benchmarks, average performance improved by 8.8 and 8.4 points over standard coupled self-evolution, respectively. A replay using identical proposals but source-excluded feedback further reduced false agreement to 0.4% and 0.1%, helping isolate feedback ancestry from changes in the curriculum.
The study does not establish whether the method will transfer to other model families, production search systems, or longer training runs, and it does not provide the auxiliary solvers’ full computational cost.
02SMART Translates Entire Series With Persistent Memory and Dynamic Multi-Agent Routing, Posts Best MQM Scores in All 15 Subtitle Directions
Subtitle translation cannot be handled reliably as a sequence of isolated sentences. Terminology, character references and style must remain consistent across episodes or an entire series, while every line still has to satisfy subtitle-format constraints. Existing single-language-model methods largely work sentence by sentence, and previous multi-agent systems typically follow fixed workflows that do not adjust to scene complexity or production context.
A new research system called SMART reframes the task as a long-form production process. Short for Self-evolving Multi-Agent system for long-foRm subtitle Translation, it builds persistent, series-level memory during a test-time training stage and translates a subset of the script before processing the rest of the series.
For those initial sentences, a dynamic router selects a path through a “Mixture-of-Agents” layer rather than sending every passage through the same fixed sequence. The agents can use tools for terminology verification, subtitle-constraint validation and contextual retrieval, helping the system carry information across scenes and episodes.
SMART then places the proposed translations into a judge-and-refiner loop. The judge scores candidates and supplies written critiques; the system uses that feedback to revise agent prompts and routing policies without retraining the underlying language models. Once this configuration has evolved, SMART applies the updated workflow to the remaining episodes.
The researchers also introduced Subtitle Arena, an evaluation set spanning 14 genres and series ranging from two to 198 episodes, with production years from 1959 to 2023 and 15 target locales. They paired it with SubMQM, a subtitle-specific version of the Multidimensional Quality Metrics framework, covering seven quality dimensions and 19 error categories.
In the paper’s experiments, SMART recorded the best overall MQM score in all 15 Subtitle Arena translation directions. Its average penalty was 6.9% lower than that of the strongest competing agent system. On the public MuSC benchmark, it also produced what the paper describes as the best model result across all four language pairs and received the highest human-evaluation result, scoring 4.50 out of 5.
The source does not report runtime, operating cost, required human editing or performance on real distribution projects, so the quality gains cannot yet be translated into commercial production savings.
03DyRAD Reconstructs Dynamic Driving Radar Scenes, Raising Recovery of Reference-Detected Objects From 26.9% to 90.7% at New Viewpoints
Closed-loop testing of autonomous-driving systems needs sensor observations from routes and viewpoints that a recording vehicle did not actually traverse. Reconstructing those unseen views is particularly difficult for radar because moving objects affect both their apparent position and Doppler—the radar measurement of velocity toward or away from the sensor.
Existing novel-view methods handle only parts of that problem. Approaches designed for dynamic scenes reconstruct range–azimuth tensors but omit Doppler, while methods that render Doppler assume the scene is static. Radar processing creates another complication: a single reflection is spread across multiple data bins. When a reconstruction treats that sensor-generated spread as part of the scene’s physical geometry, it renders the signal incorrectly after the viewpoint moves.
DyRAD addresses these problems by representing a driving scene as static background reflectors plus moving point reflectors tracked over time. It derives the moving reflectors’ velocities from object tracks and projects those velocities onto the radar’s line of sight. The system can therefore render complete range–azimuth–Doppler tensors, while also using the Doppler measurements to supervise the estimated object tracks.
The method separately models the radar’s point-spread function, or PSF: the pattern describing how the signal-processing chain distributes one reflection across nearby bins. DyRAD uses a fixed analytical PSF derived from that processing chain instead of embedding the spread in its representation of the scene. This keeps physical reflectors distinct from effects created by the sensor.
That separation also lets researchers change the simulated radar configuration without reconstructing or refitting the scene. The same modeled reflectors can instead be rendered through specifications associated with another radar configuration, enabling what the researchers call zero-shot sensor-configuration transfer. The source does not provide individual transfer scores or deployment-compute costs.
The researchers evaluated DyRAD on RADIal, Boreas and a synthetic benchmark, covering both recorded-path poses and displaced viewpoints not tested in prior work. On RADIal, DyRAD recovered radar detections for 90.7% of the objects that were detected in the reference data, compared with 26.9% for the strongest baseline. That figure is not an overall object-detection accuracy score, and the study does not establish improved safety or performance on real roads.

Google DeepMind Open-Sources SynthID Bio for Watermarking AI-Designed Proteins SynthID Bio embeds provenance signals in protein sequences and structures to support DNA-synthesis screening and database integrity. DeepMind is releasing its methods, code, model weights, and laboratory data, while noting that resistance to deliberate tampering needs further work. deepmind.google
Barclays Plans Claude Code Access for Most Engineers by 2027 Barclays is expanding Anthropic’s Claude across software development, legacy-system modernization, and operations, with Claude Code adoption expected to reach half its developers by the end of 2026. The bank says a Claude-powered assistant already serves more than 16,000 employees, while another system processes about 120,000 Global Markets emails daily. anthropic.com
Albertsons Adds Safeway Shopping and Checkout to ChatGPT Albertsons is launching a Safeway experience that lets shoppers turn recipes, photos, lists, or meal requests into a cart before proceeding to Safeway checkout. The retailer plans to extend the experience to brands including Albertsons, Vons, Jewel-Osco, Shaw’s, ACME, and Tom Thumb. openai.com
Mid-Harness Raises Terminal-Agent Pass Rate by Verifying Actions Before Execution Mid-Harness samples and checks multiple candidate commands before allowing a terminal agent to execute one, without changing its generator or execution harness. Its researchers report that a GPT-5.6 Sol verifier raised TMAX-9B’s TerminalBench-Lite Pass@1 from 50.00% to 68.03% with eight sampled actions. huggingface.co
AREX-2 Trains Agents to Refine Solutions Across Longer Test-Time Runs AREX-2 uses improvement trajectories from machine-learning and algorithmic-programming tasks to teach an agent based on Qwen3.8-27B how to reflect and iterate over many rounds. The researchers report scores of 81.8 on MLE-bench Lite and 84.0 on BrowseComp, with performance continuing to improve as the round budget increases. huggingface.co
OpenAI Essay Argues AI’s Biggest Contribution May Be Routine Execution In an independently authored essay hosted by OpenAI, Hemanth Asirvatham and Elliott Mokski argue that scientific and economic progress is increasingly constrained by the institutional work required to turn ideas into results. They propose that AI may create substantial value by handling coding, research, coordination, and other execution-heavy tasks. openai.com
The Den Says ChatGPT Work Saves Its Leadership 10–15 Hours Weekly The Den, a small arts-and-food business, says ChatGPT Work reduced the leadership team’s weekly workload by 10–15 hours. It also reports cutting grant-application time by 92% and liquor-license application time by 91%. openai.com