Today's Key Insights

  • Muse Builds Hourly Profiles of Your Personal Network — Meta is turning assistant memory into an evolving dossier about a user’s relationships, not merely a record of prior prompts. Oxford’s Carissa Véliz identifies the governance risk: Muse can accumulate both disclosed facts and inferences about other people, and those inferences may be wrong.
  • Google Restricts Free Gemini Users to Flash-Lite — Google’s power users will face a clearer paywall in October 2026, while OpenAI’s ChatGPT already offers the competing freemium structure Google is adopting.
  • Aleph Benchmark Finds Chinese Models Echo Beijing Doctrine — For the EU, Chinese models are not neutral substitutes: procurement teams must weigh their political outputs against European models’ ability to compete on performance and win broader adoption.
  • AWS Adds Multi-Turn RL for Search Agents to SageMaker — Enterprise search teams now have a third option between paying frontier-model latency and cost and collecting costly expert demonstrations for supervised fine-tuning: specialize a smaller model around their own BM25 and vector-search environment.
  • AI Agents Sound Done—Then Fail Database Checks — Teams putting agents on real records should test 20-run consistency, not just one-shot accuracy: Kimi-K3 covers 476 of 507 tasks once but reliably completes only 68.

Top Story

Muse Builds Hourly Profiles of Your Personal Network #

Meta’s Muse is instructed to create a page for every person in a user’s life—family, partners, friends, colleagues, collaborators, and followed accounts—and update those pages hourly. The profiles can record facts, history, birthdays, recurring threads, closeness, and suggestions for strengthening relationships.

Independent researcher Karan Joshi extracted Muse’s instructions through its regular chat interface and shared them with WIRED. The files say Muse should use only available evidence because invented details are worse than an empty page.

Meta says each user gets a dedicated virtual machine, can wipe memories or disconnect services, and that Muse is designed to seek human confirmation before sending emails or making purchases. The agent also maintains an audit log.

Why it matters: Meta is turning assistant memory into an evolving dossier about a user’s relationships, not merely a record of prior prompts. Oxford’s Carissa Véliz identifies the governance risk: Muse can accumulate both disclosed facts and inferences about other people, and those inferences may be wrong.

Key Takeaways

  • A profile may begin sparse and expand into sections titled Facts, History, The relationship, In common, Open threads, and Strengthening.
  • Each user’s virtual machine is inaccessible to other agents, while Meta says users can delete memories or disconnect external services at any time.
  • Meta says Muse draws context from public information and details users choose to share; researcher Miranda Bogen says it places more emphasis on relationships and personal contacts than rival assistants.

Industry Updates

Google Restricts Free Gemini Users to Flash-Lite #

Starting in October 2026, Google will limit personal-account users without subscriptions to Flash-Lite, Gemini’s smallest model. The free tier will lose Flash and Pro, compared with current access to 3.6 Flash and varying access to 3.1 Pro.

Google’s $4.99-per-month AI Plus plan also loses Pro, leaving subscribers with Flash-Lite and Flash. Access to all three models will require AI Pro at $19.99 per month or AI Ultra at $99.99 or $199.99 per month.

The setup mirrors OpenAI’s ChatGPT freemium model. Casual users may notice little, but Google may also be preparing for Gemini 4 Argon, which could be pricier to run.

Why it matters: Google’s power users will face a clearer paywall in October 2026, while OpenAI’s ChatGPT already offers the competing freemium structure Google is adopting.

Aleph Benchmark Finds Chinese Models Echo Beijing Doctrine #

Alibaba’s Qwen, DeepSeek, and Moonshot AI’s Kimi gave balanced answers to only 17% to 41% of politically sensitive prompts in an Aleph Alpha benchmark. The test covered 967 hand-picked topics, including Tiananmen, Taiwan, and Xinjiang; other responses repeated state doctrine, deflected, or refused.

The pattern appeared beyond China-focused prompts. Asked about U.S. censorship, Qwen 3.6 ended by defending China’s information controls. Nvidia’s Nemotron Cascade 2 showed party-line patterns in 17% of responses; Aleph Alpha linked that result partly to about 3,500 examples generated by DeepSeek and Qwen in its 9.3 million-example training set.

The findings align with China’s requirement that public-facing models reflect “socialist core values.”

Why it matters: For the EU, Chinese models are not neutral substitutes: procurement teams must weigh their political outputs against European models’ ability to compete on performance and win broader adoption.

AWS Adds Multi-Turn RL for Search Agents to SageMaker #

AWS added multi-turn reinforcement learning (MTRL) to SageMaker AI, letting enterprises fine-tune search agents across complete tool-use trajectories instead of scoring one response at a time. In AWS’s example, a Qwen3.6-27B model chooses between BM25 lexical search and vector search while learning from a final-outcome reward.

The managed workflow supports custom rewards and tool loops, asynchronous rollouts, resumable jobs, and serverless per-token execution without GPU-cluster management. MLflow exposes turn-by-turn trajectories and rewards, while evaluation jobs report reward, pass@k, and trajectory metrics before deployment to a SageMaker AI endpoint or Amazon Bedrock. AWS says the approach specializes smaller models for environment-specific behavior, targeting faster, cheaper inference than frontier models.

Why it matters: Enterprise search teams now have a third option between paying frontier-model latency and cost and collecting costly expert demonstrations for supervised fine-tuning: specialize a smaller model around their own BM25 and vector-search environment.

AI Agents Sound Done—Then Fail Database Checks #

AI agents can complete a conversation while leaving the database wrong. Hugging Face’s ThinkingBox benchmark ran 507 stateful business workflows 20 times across 12 language models, checking terminal records and side effects rather than just tool calls or final answers.

Across 121,680 valid trials, 79,853 failed executable checks. Yet 67.24% of those failures terminated cleanly, invoked a state-changing tool, and reported no final tool error. The checks found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%.

Claude Opus 5.5 led single-attempt accuracy at 67.16%, but repeated success was far less reliable. Kimi-K3 solved 93.89% of tasks at least once while succeeding on all 20 attempts for only 13.41%.

Why it matters: Teams putting agents on real records should test 20-run consistency, not just one-shot accuracy: Kimi-K3 covers 476 of 507 tasks once but reliably completes only 68.

Rural Data Centers Could Gain Federal Tax Benefits Next Year #

More than 100 data centers in rural areas could qualify for new federal tax benefits next year. The expanded opportunity-zone program makes eligible rural tracts available for incentives that could lower costs for capital-intensive projects, including hyperscale facilities. A rural address alone is not enough: companies must create a specialized investment vehicle to claim the benefits.

Searchlight Institute identified the potential sites using a conservative database of fewer than 700 planned or under-construction US data centers; other datasets count nearly 1,500 projects in development. Pew found that 67% of planned facilities are in rural areas, compared with 13% of operating centers. Microsoft, Meta, and Amazon denied using the program; Google did not respond. Tax analyst Emily Kraschel said capital investment alone does not guarantee jobs or a local economic boost.

Why it matters: Rural officials could see more than 100 data-center projects qualify next year, but Emily Kraschel’s warning means the tax break may deliver capital investment without the job creation communities expect from a factory.

Only 2% of Consumers Are Buying AI #

Only 2% of consumers are buying AI, according to TechCrunch's Equity podcast, while enterprise customers still appear to provide the sector's biggest money.

The figure landed as the White House gathered Zuckerberg, Bezos, Musk and Anthropic CEO Dario Amodei to sign an AI safety pledge that President Donald Trump called “morally binding.” Trump also signed an executive order rebranding artificial intelligence as “super intelligence,” while Meta and OpenAI put friendlier faces on their AI products.

Why it matters: The 2% figure leaves Meta and OpenAI trying to monetize consumer products in a market where enterprise customers still appear to generate AI's biggest revenue.