01AutoGUIWorld Generates 79,266 Interaction Samples Without Running Real Software, Lifting a Desktop Agent’s Score From 33.0% to 40.8%
GUI agents need interaction trajectories—records of actions, resulting screen changes, and task progress—to learn how software responds and how to complete multi-step workflows. Expanding those records usually means installing, configuring, and running more applications, with specialized software adding further deployment and runtime costs.
AutoGUIWorld offers another way to produce that training data. The research framework generates simulated sequences of interface interactions without deploying or running the corresponding software, then uses those sequences to fine-tune an agent. In the paper’s experiments, the resulting model performed better on benchmarks covering real desktop and scientific tasks.
The process begins with a structured specification describing the operating-system context, the interface’s visual appearance, and its current state. AutoGUIWorld uses that description to generate an initial GUI scene and creates a task conditioned on what appears in the scene.
A planner then builds the interaction one step at a time. For each step, it specifies an atomic action, such as a single interface operation, along with the visual consequence that the action is expected to produce. Instead of sending that operation to live software, an image generator edits the current screenshot to depict the next observation. Repeating that planner-and-editor cycle creates a continuous synthetic trajectory in which the interface changes after every proposed action.
The framework applies action-grounding checks, which connect actions to spatial locations in the generated interface, and transition-level quality filtering between successive observations. That pipeline produced 79,266 spatially annotated, step-level training samples spanning Ubuntu, Windows, macOS, and Chrome, according to the AutoGUIWorld paper.
The researchers used the generated trajectories to fine-tune Qwen3.5-35B-A3B. On OSWorld, its mean task score increased from 33.0% to 40.8%, a gain of 7.8 percentage points. On ScienceBoard, its task success rate rose from 14.0% to 32.2%, an increase of 18.2 percentage points. The paper presents these results as evidence that synthetic trajectories can improve performance on real desktop and scientific tasks, not as proof that generated interfaces can replace live software environments in every setting.
Important limits remain unspecified. The source does not report the computing cost of generating and filtering the 79,266 samples or the error rate among samples that received no human review. It also does not establish whether the gains extend beyond the tested model and benchmarks.
02Study Finds LLM Post-Training Often Sacrifices Repeated-Sampling Coverage, Proposes a “Sharpening Tax” to Measure the Loss
Post-training can make a large language model more accurate on its first attempt, but a new study finds that the improvement often comes with a less obvious cost: given many attempts, the model may be less likely to uncover the occasional successful path available to its base version. The researchers call that loss the “Sharpening Tax.”
The study examines agentic tasks that require multiple rounds of tool use and interaction, extending a trade-off previously observed in mathematics and coding. In this framing, pass@1 measures whether a model succeeds on one attempt, while pass@K asks whether at least one attempt succeeds when the model is allowed to sample multiple times—essentially, whether trying the same task repeatedly can uncover a workable route.
The researchers found that pretrained models equipped with a lightweight inference harness could already operate as capable agents. Their pass@1 scores were far lower than those of post-trained counterparts. With a sufficiently large test-time budget, however, the base models often achieved better pass@K coverage.
The analysis suggests that post-training pushes tasks toward two extremes: problems that the model consistently solves and problems that it never solves. This sharpening improves single-attempt accuracy, sampling efficiency and consistency, but can eliminate rare successful outputs. The paper does not claim that post-training is harmful overall; it identifies a trade-off between reliable immediate performance and the breadth of tasks solvable through repeated sampling.
Sharpening Tax is the researchers’ diagnostic for quantifying the resulting loss in test-time scalability—the model’s ability to solve additional tasks as more samples are generated. They report that the metric can be estimated from only a few rollouts and correlates well with other measurements.
The team tested 14 base/post-trained model pairs from four model families across three agentic benchmarks, producing 42 model-benchmark cases. The tax appeared in most settings, although the abstract does not specify the exact number or list the individual models and benchmarks.
To reduce the trade-off, the researchers propose posterior-tempered group sampling, or PTGS, a plug-and-play Bayesian sampler that adjusts sampling temperature for each prompt according to its estimated difficulty. Used during reinforcement-learning training in two agentic environments, PTGS produced a smaller Sharpening Tax than a fixed-temperature baseline, improved single-shot accuracy and solved more tasks under repeated sampling. The source does not state its additional training or inference cost.
03Early-Layer 65,537-Parameter LoRA Raises Qwen3-8B Accuracy on 24-Line Reference Chains From 15.5% to 99%
A small adapter inserted early in Qwen3-8B sharply improved its ability to follow long chains of references without producing a visible chain of thought. The test resembles a program such as “K = apple; B = K; D = B; print(D),” but extends the sequence across many lines before asking for the final value.
Researchers found that 13 pretrained models ranging from 0.6 billion to 32 billion parameters could reliably follow only 1.4 to 3.6 lines. Extra passes through pretrained layers offered limited help, suggesting that the models were ending the relevant computation too early rather than simply lacking depth.
The researchers added a task-trained, rank-8 LoRA adapter at layer 14 of Qwen3-8B. LoRA is an adaptation method that trains a small number of additional parameters while leaving the underlying model unchanged. This adapter contained 65,537 parameters, with all base-model weights frozen.
On 24-line chains, exact accuracy rose from 15.5% to 99%. A version trained for longer handled 50 lines in one forward pass, although the source does not report the training-data volume, duration or computing cost.
The proposed mechanism is a relay across layers. The LoRA operates on individual tokens and does not itself transfer information between them. Instead, the researchers say it prompts program lines to carry their chain identities through frozen middle layers, while existing attention heads read progressively farther back through the chain.
Intervention experiments support that account. Blocking every line from attending to its parent line in layers 14 through 22 reduced accuracy to chance. Applying the same block in layers 23 through 29 had little effect, locating the important transfer earlier in the network. Placement was also sensitive: the same LoRA recipe reached 20.5 lines at layer 20 but only 5.2 at layer 21. A measurement on the frozen models predicted the final useful intervention layer within a preregistered tolerance for three of four held-out models.
The result extended beyond synthetic reference chains, but only with separately trained adapters. On MuSiQue multi-hop question answering using gold paragraphs, early-layer LoRAs raised exact-match scores by 9.4 to 17.9 points across three standard models. The study does not show that one LoRA can be reused directly across these tasks.
OpenAI Rolls Out Dot Agent for Computer-Based Work OpenAI’s Dot agent can operate cloud and local computer apps, accept voice instructions, and perform tasks such as editing media or updating websites. The initial rollout targets top-tier subscribers, while hands-on testing found that security checks frequently required human intervention. theverge.com
OpenAI Safety Report Lead Resigns and Criticizes Company Culture David Robinson, who said he led safety reports for major OpenAI launches, resigned after three and a half years and argued that the company’s culture is “broken.” OpenAI said it is strengthening security, third-party evaluations, responsible-task training, and real-time monitoring. techcrunch.com
White House Rebrands AI as “Super Intelligence” President Donald Trump signed an executive order adopting the term “super intelligence” and convened major technology executives to sign an AI safety pledge that he called “morally binding.” Participants reportedly included leaders from Meta, Amazon, xAI, and Anthropic. techcrunch.com
Meta Open-Sources Tools for Building Muse Hardware Meta released code and SDKs that let developers connect its Muse AI agent to devices built with components such as ESP32 boards and Raspberry Pis. The company also made a waitlist available for 5,000 Muse Home Link units designed to control connected household equipment through community-built skills. theverge.com
AWS Stops Using NDAs in Data-Center Approval Talks AWS CEO Matt Garman said Amazon no longer uses nondisclosure agreements with government agencies involved in its data-center projects. The change comes amid local opposition and more than 100 proposed US data-center moratoriums, according to Garman. techcrunch.com
Stability AI Rebuilds Around Licensed Music Tools Stability AI is repositioning itself as an AI toolmaker for music professionals under investor Sean Parker and CEO Prem Akkaraju. The company raised $76 million from investors including Sony, Warner, and Universal, licensed their catalogs for training, and released three audio models plus music-editing software. techcrunch.com
Capcom Plans AI-Assisted Evolution of Its RE Engine Capcom programmer Satoshi Ishida outlined an incremental plan to integrate AI into development workflows and eventually turn the company’s RE Engine into an “AI-generation game engine.” Capcom has previously said it would use AI for production efficiency rather than AI-generated game assets. theverge.com
Circuit Breaker Labs Tests Chatbots for Psychological Harm Circuit Breaker Labs uses simulated users with different ages, cultures, languages, slang, and speech patterns to test whether AI systems respond safely across long conversations. The five-person startup has a working product for high-risk applications including AI coaching, journaling, and mental-health support, but has not named its customers. techcrunch.com
Microsoft Expands Nature-Based Data-Center Projects Microsoft plans to bring its “biomimicry” program—using native vegetation, habitat restoration, and landscape design around data centers—to more than 20 sites in the US and Germany. The company says it will apply the approach to all new US projects and a growing number elsewhere. theverge.com
GraphForge Builds Verifiable Training Workspaces from Real Files GraphForge creates agent-training tasks and evaluation rubrics from evidence graphs grounded in real workspace files. Fine-tuning Qwen3.6-27B on 2,169 generated trajectories improved reported results on GDPVal, Workspace-Bench-Lite, and SpreadsheetBench II. huggingface.co
Distillation Study Questions the Default Preference for On-Policy Data A controlled study across Llama 3 and Qwen 2.5 models found that rollout policy was not consistently the main driver of distillation performance. Token-level KL direction more clearly affected performance and output coverage, while learning rate governed forgetting and update sparsity. huggingface.co