Today's Key Insights

  • Judge Blocks Pentagon’s Anthropic Supply-Chain Label — The San Francisco ruling gives Anthropic a court victory against the Pentagon’s supply-chain-risk designation, while the company’s second lawsuit in Washington continues.
  • Google DeepMind's Co-Scientist Now Runs Lab Experiments — For Google DeepMind researchers, Co-Scientist's reported expansion connects Gemini-based reasoning to experiment planning and laboratory equipment, but the lack of performance metrics leaves its effect on research speed and quality unknown.
  • Automated Systems Improve on All 10 Misalignment Benchmarks — For developers of automated AI systems, the reported result combines better performance on 10 specific misaligned-behavior benchmarks with no degradation in overall performance.
  • U.S. AI Compute Now Outpaces China by More Than 10x — U.S. AI infrastructure teams are receiving a far larger share of new data-center capacity than their Chinese counterparts, while U.S. chips deliver 2.3–2.7 times more compute per watt.
  • OpenAI Builds Codex Agent That Keeps Working Until Stopped — For OpenAI, the feature would add a persistent operating mode to Codex, letting the coding agent continue proactive work until it is put to sleep.

Top Story

Judge Blocks Pentagon’s Anthropic Supply-Chain Label #

Anthropic won its first court victory against the Pentagon. A federal judge ruled that the Trump administration illegally labeled the AI company a supply-chain risk, and called the Department of Defense’s designation “illegal and baseless.”

The Decoder reported that the Pentagon blacklisted Anthropic in retaliation for the company’s public criticism of government AI policy. The ruling came from a federal court in San Francisco.

Anthropic’s second Pentagon lawsuit continues in Washington, leaving the broader dispute between the AI company and the Defense Department unresolved.

Why it matters: The San Francisco ruling gives Anthropic a court victory against the Pentagon’s supply-chain-risk designation, while the company’s second lawsuit in Washington continues.

Key Takeaways

  • The federal court in San Francisco found the Department of Defense’s designation unlawful.
  • The Decoder said the Pentagon’s blacklist followed Anthropic’s public criticism of government AI policy.
  • Wired described the designation as “illegal and baseless.”

Industry Updates

Google DeepMind's Co-Scientist Now Runs Lab Experiments #

Google DeepMind has expanded Co-Scientist from a hypothesis generator into a research system integrated with laboratory work. According to The Decoder, the Gemini-based multi-agent system now plans experiments, runs laboratory equipment and writes scientific papers.

The reported work spans three disciplines, including materials synthesis and the autonomous development of a medical AI architecture. Co-Scientist is therefore described as moving beyond proposing research ideas to participating in the experimental workflow.

The report provides no performance metrics, publication details or timeline for broader deployment, leaving its effect on research speed and quality unmeasured.

Why it matters: For Google DeepMind researchers, Co-Scientist's reported expansion connects Gemini-based reasoning to experiment planning and laboratory equipment, but the lack of performance metrics leaves its effect on research speed and quality unknown.

Automated Systems Improve on All 10 Misalignment Benchmarks #

Automated systems improved performance on all 10 benchmarks for specific misaligned behaviors without degrading their overall performance, according to TechCrunch AI.

The result ties targeted improvement to stable overall performance across the 10 reported benchmarks.

Why it matters: For developers of automated AI systems, the reported result combines better performance on 10 specific misaligned-behavior benchmarks with no degradation in overall performance.

U.S. AI Compute Now Outpaces China by More Than 10x #

About 70% of new AI data-center watts are now going into the United States, while China gets less than 10%. Next Big Future AI says that reverses the 2022 split, when the U.S. accounted for roughly 45–50% of the world's new compute and China 30–35%.

The source attributes the reversal to export controls and a huge U.S. infrastructure buildout. U.S. AI chips also deliver 2.3 to 2.7 times more compute per watt, giving the U.S. an efficiency advantage alongside its larger share of new data-center capacity.

Why it matters: U.S. AI infrastructure teams are receiving a far larger share of new data-center capacity than their Chinese counterparts, while U.S. chips deliver 2.3–2.7 times more compute per watt.

OpenAI Builds Codex Agent That Keeps Working Until Stopped #

OpenAI is developing a Codex feature that keeps the coding agent working proactively until someone puts it to sleep, according to code reviewed by WIRED.

The code describes the feature as persistent: Codex continues working rather than stopping after a single task, with “put to sleep” serving as the stop condition.

Why it matters: For OpenAI, the feature would add a persistent operating mode to Codex, letting the coding agent continue proactive work until it is put to sleep.

227 Install Commands Pointed Corporate Networks to Unowned Code #

Ars Technica AI found 227 install commands in corporate documents pointing to code nobody owns. The commands involved Claude, Codex and Hermes.

The installations occurred inside corporate networks, according to the report.

Why it matters: Corporate networks are directly implicated because 227 install commands pointed to code nobody owns.

Cheap Models Put 3–20 Overnight Bots Within Reach #

Next Big Future AI describes this month as the start of a “grokbot boom”: one person can use 3–20 background bots that keep working overnight instead of relying on a chatbot only while typing. The article says cheap models make those extra hours of operation affordable.

Next Big Future AI lists GLM-5.3-Flash at $0.15/$0.50 per million tokens, roughly 10 times cheaper than a frontier model. The cost difference supports a clear tradeoff: one chatbot burns tokens during active typing, while 3–20 background bots burn tokens throughout the night.

Why it matters: For a single operator, GLM-5.3-Flash’s $0.15/$0.50-per-million-token pricing makes running 3–20 background bots overnight affordable.