Officials Order 58,000 Exam Retakes After Top Scores Jump Fivefold

01Top scores jumped fivefold. Now 58,000 students must retake an AI-supervised exam

For 58,000 students, completing a remote exam was not enough. They must take it again after an AI-supervised test produced an extraordinary result: top scores increased fivefold.

The retest transfers the cost of the breakdown directly to students. Each must repeat the preparation, scheduling, and examination process, regardless of whether their individual result prompted concern. The available reporting does not establish what caused the surge. It does establish the response: discard the original sitting at scale and start over.

That sequence exposes a missing checkpoint. An abnormal result became visible only after tens of thousands of people had passed through the system. By then, investigators could examine outputs, but students had already relied on the exam as a valid assessment.

A separate research project describes why inspecting AI systems from the user’s position remains difficult. Developers configure system prompts to govern how foundation models behave inside applications. According to the AISPA paper, companies rarely disclose those instructions to users or regulators.

The researchers propose Artificial Intelligence System Prompt Assurance, or AISPA, as a framework for auditing those hidden instructions. Its focus is not the remote-exam incident, and the paper does not identify the cause of the fivefold score increase. Instead, it addresses a related governance problem: outsiders often cannot inspect the rules directing an AI application before deployment.

For exam takers, that opacity limits the evidence available when a result is challenged. A student can see the final score and the demand for a retest. They may not see how the supervisory system classified behavior, handled uncertainty, or escalated anomalies.

Large deployments magnify that information gap. A faulty control affecting one candidate creates an individual dispute. A questionable result across an entire sitting creates 58,000 new examinations.

AISPA places the audit target earlier in that chain. It examines instructions developers configure before users encounter an AI application. The exam incident shows the alternative sequence: detect an extreme statistical shift after deployment, then make the full cohort repeat the process.

Students need appeal channels before scores trigger blanket retestsExam operators face predeployment audits as a procurement thresholdRegulators can demand instruction disclosure for high-stakes AI systems

02Apple asks court to halt OpenAI’s AI device work, cites 11 more ex-employees

Apple is asking a court to stop OpenAI from advancing an AI device or other products allegedly based on Apple technology. The preliminary-injunction request raises the stakes before the trade-secret case reaches a final judgment. Apple also says its investigation points beyond the former employees named in its original complaint.

The new filing requests expedited discovery from senior systems engineer Chang Liu and Chief Hardware Officer Tang Yew Tan, both accused in the case. Apple also seeks material from OpenAI, its foundation, and io, the device startup co-founded by former Apple design chief Jony Ive. These are requests and allegations, not findings that confidential information was taken or used.

Apple says 11 other former employees may have witnessed or otherwise joined the conduct under investigation. Its filing describes one former employee allegedly meeting Liu and OpenAI employee Yu-Ting Peng before Peng’s interview. According to Apple, the group discussed proprietary information about unannounced products.

The company also alleges that another former employee captured screenshots of confidential documents about an unannounced product before an interview. That account supports Apple’s request to examine a broader set of people and records on an accelerated schedule. The filing does not establish that all 11 took data or that OpenAI incorporated Apple secrets into a product.

OpenAI rejects Apple’s version. In a public response titled “Apple is getting this wrong,” the company calls the lawsuit baseless. It says the post corrects claims about its employees and publishes messages documenting what happened. The available summary does not describe those messages, leaving unclear which specific Apple allegations they address.

The two sides are now contesting both the underlying facts and the timing of judicial intervention. Apple wants restrictions and faster evidence gathering before trial. OpenAI publicly disputes the case’s foundation while continuing its hardware push with io. The court must decide whether Apple gets accelerated discovery and product restrictions before the trade-secret claims are adjudicated.

OpenAI hardware teams could face development limits before trialio may owe records under an expedited timetableformer Apple staff could face broader document demands

03AI Speeds Production While Readers Doubt Authors and Developers Retype Code

Automation can cut the time required to produce an image or complete a software feature. It can also transfer work to the person who must trust, review, or maintain the result. Two personal accounts show that transfer appearing in different places: blog readership and software development.

One blogger says AI-generated images now discourage him from reading personal blogs. The images also make him question whether the accompanying text was generated. He expects automation from corporate publishing but views it differently on an independent site, where the author’s individual perspective is the product. He would prefer a crude Paint drawing because it offers clearer evidence of human involvement.

The image may cost almost nothing to generate. The reader pays through extra suspicion, including suspicion directed at text that may be entirely human-written. Decorative content intended to make a post look finished can therefore weaken confidence in the material it accompanies.

A developer describes a parallel problem with coding assistants. One-shot feature generation saves him from tedious implementation work, but leaves him disoriented inside his own project. He calls the resulting gap “cognitive debt”: the software exists, yet he does not fully understand how it works.

Line-by-line review does not eliminate that burden. The developer reports facing hundreds of lines of defensive, poorly commented, and subtly incorrect code. Reviewing that output is work, and it does not provide the same understanding gained by building the feature incrementally. For personal projects, he has responded by manually retyping generated code to reconstruct the reasoning behind it.

That method is not a prescription for every engineering team. It is evidence that generated output and absorbed knowledge are separate deliverables. Teams can preserve the speed benefit by generating smaller sections, then requiring developers to explain, test, and rebuild their understanding before accepting them.

The same accounting applies to publishing. Removing decorative generated images can reduce reader uncertainty without rejecting AI throughout the writing process. In both cases, efficiency claims need to include the human time spent establishing provenance, correctness, and comprehension after generation.

Indie bloggers risk losing readers before the first paragraphCode-review budgets must include time to reconstruct system behaviorTeams need separate metrics for output speed and maintainer understanding
04

Anthropic signs $10 billion cloud deal with Volta Anthropic reportedly signed a $10 billion agreement with AI cloud startup Volta. The deal extends Anthropic’s recent series of cloud partnerships. techcrunch.com

05

Texas halts new data-center projects Texas halted new data-center projects. The governor called for audits as expanding AI infrastructure strained the state’s power supply. techcrunch.com

06

AMD doubles data-center revenue to $6.7 billion AMD reported $6.7 billion in quarterly data-center revenue, up from $5.8 billion in the previous quarter. AI demand pushed annual growth to 107%. theverge.com

07

SpaceX earns $2.6 billion from AI compute SpaceX generated $2.6 billion from AI operations, more than triple the previous year. Compute deals with AI companies made the division its largest revenue source. theverge.com

08

Alibaba releases its largest Qwen model Alibaba released Qwen3.8-Max and made it widely available to users. Alibaba says the model rivals leading systems from Anthropic, OpenAI, and Moonshot AI. theverge.com

09

Trump administration extends AI protectionism to robotics The Trump administration extended its AI protectionist policies to robotics. Humanoid robots still struggle with stable movement and precise hand control. technologyreview.com

10

Nvidia-led alliance proposes defenses against AI agents The Open Secure AI Alliance released its first proposals for defending systems against AI agents one week after forming. Nvidia leads the group, which includes over 120 companies. techcrunch.com

11

OpenAI builds GPT-Live for continuous voice interaction OpenAI built GPT-Live in six months around a turnless speech model. The system supports continuous conversations through a low-latency architecture. openai.com