Today's Key Insights

  • OpenAI’s Jalapeño Beats Nvidia’s Blackwell and Rubin in Tests — OpenAI now has benchmark evidence that its in-house inference chip outperforms Nvidia’s Blackwell and Rubin on energy efficiency, latency and interactive-workload performance.
  • Claude Cowork Shares Memory With Claude Chat — For Claude users, shared memory means fewer repeated briefings about projects, preferences, and other context across Claude chat and Cowork.
  • Gemma 4, Llama 3, Mistral: Local Tool-Calling Trade-Offs — Engineering teams choosing among Gemma 4, Llama 3, and Mistral for local deployments get a comparison of each model family's tool-calling trade-offs, but no quantitative basis for selecting a winner.

Top Story

OpenAI’s Jalapeño Beats Nvidia’s Blackwell and Rubin in Tests #

OpenAI’s first in-house inference chip delivered 1.5 to 1.9 times more AI work per watt and 1.7 to 3.6 times lower end-to-end latency than comparison systems. SemiAnalysis tests presented at the Hot Chips conference found Jalapeño ahead of Nvidia’s Blackwell and Rubin in throughput and energy efficiency.

On SemiAnalysis’ InferenceX benchmark, Jalapeño produced more tokens per user and more throughput per kilowatt than the currently available state of the art. For highly interactive workloads, it delivered 2.1 to 4.1 times higher performance.

OpenAI unveiled the chip program in June with Broadcom, building Jalapeño from a blank sheet for large-scale inference.

Why it matters: OpenAI now has benchmark evidence that its in-house inference chip outperforms Nvidia’s Blackwell and Rubin on energy efficiency, latency and interactive-workload performance.

Key Takeaways

  • OpenAI showed Jalapeño at the Hot Chips conference, where the benchmark results were presented
  • OpenAI and Broadcom unveiled the program in June and built the chip from a blank sheet
  • SemiAnalysis’ InferenceX test measured more tokens per user and more throughput per kilowatt for Jalapeño

Industry Updates

Claude Cowork Shares Memory With Claude Chat #

Anthropic is giving Claude shared memory across chat and Cowork, so users no longer have to repeatedly brief the AI on projects, preferences, and other context.

The feature connects Claude's chat experience with Cowork through shared memory, according to TechCrunch AI.

Why it matters: For Claude users, shared memory means fewer repeated briefings about projects, preferences, and other context across Claude chat and Cowork.

Gemma 4, Llama 3, Mistral: Local Tool-Calling Trade-Offs #

Machine Learning Mastery compares how Gemma 4, Llama 3, and Mistral implement tool calling when run locally. The article focuses on the differences among the three model families rather than treating tool calling as a single, identical feature.

The available text does not publish benchmark results, latency figures, pricing, or an overall winner. Instead, it presents trade-offs that engineering teams must weigh against their own local-deployment requirements.

Why it matters: Engineering teams choosing among Gemma 4, Llama 3, and Mistral for local deployments get a comparison of each model family's tool-calling trade-offs, but no quantitative basis for selecting a winner.