Taobao Live Trains Digital-Avatar Agents to Adapt as Tools and Prompts Change

01Taobao Live’s digital-avatar agent adopts “harness-aware” training; team reports higher GMV and item-page views in online tests

Taobao Live’s digital-avatar streamers must answer product questions, interact with viewers and carry out marketing strategies in real time. That creates competing demands: responses must be fast and accurate, while the instructions and capabilities governing the agent may need frequent updates.

Those capabilities sit in a modular operating framework, or “Harness,” whose skills, hooks, prompts and tools can change independently of the model’s weights. Large models can adapt to new configurations without additional training but are slower, while smaller models offer lower latency but tend to memorize fixed skill names, tool schemas and prompt templates. Taobao Live’s team now says it has deployed a training method designed to make compact models adapt as that framework evolves.

The method, called Harness-Aware Training, exposes the model to many altered versions of the framework without changing the underlying tasks. Its Harness-State Augmentation component varies skill identifiers and content, tool schemas, prompt structures and hook functions. The aim is to prevent the model from relying on configuration-specific surface forms that may disappear after an update.

Training takes place in three stages. First, the team uses trajectories generated by a stronger model to supervise reasoning and tool use across augmented environments. It then applies general on-policy distillation to restore broader capabilities lost during supervised fine-tuning. Finally, reinforcement learning in additional augmented environments is used to improve robustness to changing Harness configurations.

In offline evaluations, the team reported scores of 94.8 on Live-Stream QA, compared with 80.3 for the base model and 93.0 for the strongest general-purpose large language model tested. Harness-Aware Training scored 94.6 on Harness-Variant QA, versus 75.4 for the base model. Fixed-Harness supervised fine-tuning reduced the base model’s IFEval score by 7.7 points, while the configuration-aware approach avoided that decline and reached 83.5.

The optimized system recorded median, or P50, latency of 3.4 seconds and P95 latency of 8.1 seconds on a single NVIDIA H20 GPU, according to the team. It has been deployed in Taobao Live’s digital-avatar service, where the researchers reported positive online A/B-test results for gross merchandise value and item-page views. The source does not disclose the tests’ sample size, duration or the size of either increase.

Taobao Live can update agent skills and tools without necessarily retraining around one fixed configurationcompact models may combine lower latency with greater resilience to operational changesthe reported commercial gains remain unquantified and have not been independently verified.

02OpenExecutive Open-Sources Virtual Executive System: Eight Specialist Agents Share Company Data, Decision Memory and Task Schedules

OpenExecutive, a new open-source system for business advice, presents users with one virtual executive while routing their questions among eight specialist AI agents. Those agents cover strategy, finance, human resources, legal affairs, operations, marketing, product and board communications.

The project is intended to address a practical problem with multi-agent tools: advice from several models can feel fragmented and lose context between conversations. OpenExecutive instead uses Anthropic’s Claude Sonnet 4.6 as an executive orchestrator. It calls the relevant specialists in parallel, then synthesizes their work into a single response without exposing the internal agent structure to the user.

Each specialist can retrieve two kinds of material from ChromaDB, a vector database used to find text by semantic relevance. The first is the project’s built-in collection of business-oriented Markdown documents. The second consists of company documents uploaded by the operator, which are divided into chunks and kept in a separate company_docs collection. Retrieved passages are inserted into the current user turn rather than into the cached system prompt.

OpenExecutive also preserves a record of previous advice. After every response, a background process using Claude Haiku 4.5 extracts decisions, initiatives and recommendations into SQLite. A later session receives relevant history through a past_decisions block, allowing the system to recover what it previously recommended.

Its scheduler can surface follow-ups and other time-sensitive actions. To avoid executing the same task twice, the job runner claims due work with an UPDATE … RETURNING database operation. That protection comes with an important deployment boundary: the project says the API must remain a single instance unless an operator first adds controls around the scheduler, so it cannot be horizontally scaled safely as provided.

The repository includes a FastAPI backend, a Next.js 15 interface and integrations for Slack, email, Telegram, Google Chat and Discord. A first local launch requires Python 3.11 or later, Node 22 or later and an Anthropic API key. Initial setup installs ChromaDB and machine-learning dependencies, then downloads an embedding model of about 90 MB to build the local vector index.

The project provides evaluation scenarios but no independent accuracy assessment, disclosed enterprise adoption figures or comparison with advice from human executives.

Companies testing the system must entrust it with uploaded business documents and stored decision historydevelopers must budget for model API access, local indexing dependencies and a heavier first startupoperators must keep the API single-instance or add scheduler coordination before scaling it out.

03PAWBench Tests 11 Video-Generation Systems: No Model Consistently Reproduces the Probability Distribution of Physical Behaviors

Generating one physically plausible video does not show that a model understands the range of ways an event could unfold. Given the same initial observation and action, a physical process may have several valid outcomes, each occurring with a different probability.

That distinction matters as video generators are increasingly framed as “world models”—systems intended to represent and predict how the world changes. A new study introduces PAWBench to test whether these generators reproduce not only a convincing trajectory, but the full distribution of possible physical behaviors. Across 50 scenarios and 11 current systems, the researchers found that no model consistently recovered the range of valid behaviors while also matching their reference probabilities.

The researchers call this distribution-level requirement “probabilistic alignment.” Traditional video evaluation largely judges individual outputs: whether a generated clip looks credible or depicts a plausible result. Probabilistic alignment instead asks what happens when the model receives the same starting conditions repeatedly. Even if each sampled video appears reasonable on its own, the model remains misaligned if it omits valid outcomes or produces them at the wrong rates.

PAWBench treats a video generator as a stochastic sampler of world dynamics, meaning repeated runs can produce different results from the same setup. Its companion evaluation protocol, PAWEval, converts those repeated video rollouts into an empirical probability distribution over distinct physical behaviors. Researchers can then compare both the behaviors generated and their observed frequencies with the reference distribution.

This exposes a gap that single-video quality scores can miss. A system might generate one correct-looking outcome while systematically favoring it over other possibilities. Conversely, it might cover several valid behaviors without assigning them probabilities that resemble the reference process. The study reports that all 11 tested systems struggled to satisfy both requirements consistently across the benchmark’s 50 scenarios.

The team also tested whether language prompts, initial-noise sampling, or model training could reshape a system’s predictive distribution. The available summary does not identify the systems, provide their individual scores, or quantify how much those interventions helped. The supported conclusion is therefore limited to the models and scenarios evaluated: none consistently matched reference probabilities while covering the valid behavior range.

Developers using video generators as world models need distribution-level tests beyond plausible-looking clipssystems that miss or misweight valid outcomes may give unreliable predictions under uncertaintythe next question is how much prompting, noise selection, or training can close the measured gap.
04

Amazon adds 2 million Nvidia GPUs to AWS expansion Amazon will deploy another 2 million Nvidia Blackwell Ultra, Rubin, and Rubin Ultra GPUs in AWS data centers during 2027 and 2028, following an earlier commitment for more than 1 million GPUs. The expanded partnership also covers Nvidia networking, models, robotics technology, and server CPUs. techcrunch.com

05

Judge overturns Pentagon’s Anthropic blacklist US District Judge Rita Lin ruled that the Pentagon’s designation of Anthropic as a supply-chain risk was unconstitutional retaliation and “arbitrary and capricious.” The dispute followed Anthropic’s refusal to permit its AI systems to be used for mass surveillance of Americans or lethal autonomous weapons. theverge.com

06

OpenAI brings ChatGPT ads to India OpenAI will begin showing ads on ChatGPT’s Free and Go tiers in India, initially working with 50 brands through agency partners WPP and Omnicom. The company also plans to launch a self-service ad manager next month with a minimum daily campaign budget of ₹725. techcrunch.com

07

EPA proposal would shift public-notice decisions for some air permits to states The US Environmental Protection Agency has proposed eliminating the federal public-participation requirement for permits covering certain “minor” pollution sources, including some data centers and power facilities. The comment period has closed, and the agency says it is reviewing more than 4,900 submissions before finalizing the rule. theverge.com

08

Lambda borrows $1 billion to supply Nvidia chips to Microsoft AI cloud provider Lambda raised $1 billion in short-term private debt arranged by JPMorgan Chase to purchase Nvidia GPUs that it plans to lease to Microsoft, Bloomberg reported. The financing follows other large secured loans as Lambda expands customer-specific GPU infrastructure. techcrunch.com

09

Anthropic tests automated researchers for alignment training An Anthropic paper reports that automated research systems improved model performance across 10 benchmarks for misaligned behavior without reducing overall performance. The system searched literature, proposed and tested training methods, and retained successful approaches, though its effectiveness depended on benchmarks accurately representing the intended alignment goals. techcrunch.com

10

OpenAI hires Meta’s regional leader for Asia-Pacific expansion Sandhya Devanathan, Meta’s vice president for India and Southeast Asia, is joining OpenAI to oversee growth, enterprise adoption, partnerships, regulation, and operations across Southeast Asia and Australia. She will be based in Singapore and report to OpenAI’s Asia-Pacific managing director. techcrunch.com

11

Google hires AI researcher Barret Zoph as research vice president Barret Zoph, a Thinking Machines co-founder who subsequently returned briefly to OpenAI, has joined Google as vice president of research. Google said he will contribute reinforcement-learning and post-training expertise to Gemini. techcrunch.com

12

Instinct reaches $2.5 billion valuation while still in private beta Instinct, a personal AI assistant from Spear Street Technology that connects with users’ apps and devices, raised a $250 million Series B co-led by Index Ventures and Benchmark. The round brought its total funding to $350 million, while early access has also prompted concerns about the product’s permissions and terms of use. techcrunch.com

13

OpenAI expands its operations in Brazil OpenAI announced that it is expanding its presence in Brazil and deepening its engagement with developers, businesses, and communities to support AI adoption. openai.com

14

OpenAI and Thailand launch accelerator for 10 AI startups OpenAI and Thailand’s Ministry of Higher Education, Science, Research and Innovation launched an eight-week accelerator for 10 health, wellness, and education startups. The program is intended to help participants turn AI prototypes into products. openai.com