Anthropic released Claude Sonnet 5.5, a coding and knowledge-work model that nearly matches Opus 5.5 on several benchmarks while costing up to 30% less per task. Anthropic says output generation is more than 30% faster. The model is available through Amazon Bedrock and Claude Platform on AWS.
On GDPval-AA, Sonnet 5.5 scored 1,844 points versus Opus 5.5's 1,846; on CursorBench 4.0, it scored 55.5% versus Opus's 57.8%. On Terminal-Bench 4.0, Sonnet 5.5 reached 70.6%, compared with 10.3% for Sonnet 5.
Token prices remain $2 per million input tokens and $10 per million output tokens. Anthropic's benchmark and cost claims still await independent testing, and a pre-release bug may have affected some structured-output results.
Why it matters: For AWS Bedrock customers, the practical shift is model routing: Opus 5.5 can handle release debugging and security reviews while Sonnet 5.5 handles SQL generation, document edits and IDE agents at lower per-task cost.
More than 20 AI researchers, including Geoffrey Hinton, Yoshua Bengio and OpenAI research lead Jakub Pachocki, warn that automating AI research could trigger an “intelligence explosion.”
Their paper says AI systems already write most of the code at the companies building them and could automate the entire AI R&D pipeline within years. Progress that normally takes years could happen in months, while society struggles to keep up and control over superhuman AI could slip away.
The authors urge policymakers to gain more visibility into automated AI research. Pachocki separately said no lab has solved alignment well enough to “continue responsibly scaling at maximum speed for much longer.”
Why it matters: OpenAI and policymakers now face a governance problem with a shrinking response window: faster AI R&D could shift power among nations, companies and governments before society can keep up.
MIT researchers Brian Hedden and Manish Raghavan found that using one hiring algorithm across firms can create informational echo chambers that hinder exploration, making it less likely that the best candidates get jobs.
Their result is conditional, not a blanket indictment. Firms using the same algorithm still fill the same number of jobs, while competition for the same candidates can drive up wages. Candidates can also revise and resubmit resumes, weakening one objection about agency.
The researchers’ fix is an ensemble bundling multiple hiring algorithms. In their analysis, it can overcome monoculture’s exploration problem and perform as well as, or better than, a polyculture in which firms use different systems.
Why it matters: For Fortune 500 hiring teams using common screening tools, an ensemble offers a way to overcome monoculture’s exploration problem while matching or outperforming a patchwork of different systems.
Anthropic’s S-1 shows revenue grew twelvefold to nearly $4.6 billion in 2025, while its operating loss widened from $2.98 billion to $8.06 billion. The company spent $7.33 billion on compute and infrastructure—three times the previous year’s figure and more than half its total operating costs.
The filing warns that two customers supplied nearly a quarter of 2025 revenue, while many large accounts lack long-term contracts. About $34 billion of its roughly $42 billion net loss came from a non-cash accounting charge tied to financing that could convert into stock. Backers are targeting a valuation above $2 trillion, and Reuters sources expect an IPO after the U.S. midterm elections in November.
Why it matters: Anthropic’s debut would give OpenAI’s confidentially filed IPO a market benchmark while forcing public investors to decide whether adjusted profits can support $518 billion in planned infrastructure commitments.
Grok 4.7 is now available on Amazon Bedrock, giving AWS customers access to xAI’s model for coding, long-running agents, and knowledge work. It offers a 500,000-token context window and four reasoning levels: low, medium, high, and xhigh.
Bedrock serves the model through cross-Region inference profiles and supports the Responses, Chat Completions, InvokeModel, and Converse APIs. Developers can use the OpenAI SDK to port existing integrations or AWS SDKs to access features such as invocation logging and response streaming.
Artificial Analysis scored Grok 4.7 at 46 on its Intelligence Index, versus 44 for Grok 4.6. At xhigh reasoning effort, Grok 4.7 used roughly 81,000 output tokens per task, compared with about 38,000 for its predecessor.
Why it matters: For Bedrock teams running long-horizon agents, Grok 4.7 turns reasoning effort into a token-budget decision: Artificial Analysis measured about 81,000 output tokens per xhigh task, versus 38,000 for Grok 4.6.
AMD is acquiring World Labs for $8.2 billion, bringing founder Fei-Fei Li to AMD as executive vice president and chief scientist. The deal is expected to close before the end of the year, pending regulatory approval.
World Labs builds deep learning models designed to understand physical reality. Its first product, Marble, creates entertainment experiences and simulated environments for robot training. AMD says workloads from World Labs will shape its chip-making roadmap, while World Labs says scaling its research requires closer collaboration across model research, systems and compute.
The deal gives AMD a stronger route into world models, an area where Nvidia already offers the open-weight Cosmos suite. Synthetic data from these models could help train autonomous vehicles, industrial robots and humanoids as real-world training data remains scarce.
Why it matters: AMD had offered only text- and video-based models publicly; buying World Labs gives it a response to Nvidia’s open-weight Cosmos suite and puts Fei-Fei Li’s physical-world research closer to the chips that run it.