Ep 551 Blog 3:00 w/ Justy & Cody

OpenAI and Broadcom unveil LLM Optimized inference chip

OpenAI and Broadcom unveil Jalapeño, a custom AI inference chip designed for LLM workloads, promising substantial performance-per-watt improvements.

Embed this episode

Paste this on any site — the player is a self-contained iframe with no cookies or trackers.

<iframe src="https://sandrise.io/exploring-next/embed/551"
  width="100%" height="180" style="max-width:640px;border:0;border-radius:12px;overflow:hidden"
  title="Exploring Next — Episode 551 audio player"
  loading="lazy" allow="autoplay" referrerpolicy="strict-origin-when-cross-origin"></iframe>
Embed & API docs →
Script Llama 4 Maverick Voice Inworld TTS 1.5 Max

Transcript

Justy Cody, have you seen the news about OpenAI and Broadcom's new inference chip, Jalapeño?

Cody Yeah, I read the announcement. It's a custom ASIC for LLM inference.

Justy Right. And the big claim is it's designed from the ground up for modern LLM inference, not just a tweaked general-purpose accelerator.

Cody Exactly. The article says OpenAI used its models to accelerate parts of the design and optimization process.

Justy Mm-hm. That's the kind of software-hardware co-development we've been hearing about.

Cody The nine-month tape-out cycle is insane. That's faster than most high-performance ASICs.

Justy I know, right? And they're already running ML workloads in the lab at production target frequency and power.

Cody GPT-5.3-Codex-Spark, specifically. That's their new real-time coding model.

Justy Yeah, I saw that. One thousand tokens per second is wild.

Cody So, early testing shows Jalapeño delivering substantially better performance per watt than current state-of-the-art.

Justy That's the real win if it holds up. Performance per watt is critical for data centers.

Cody The architecture reduces data movement and balances compute, memory, and networking resources.

Justy Right, right. That's where the efficiency gains come from.

Cody Broadcom's silicon implementation and networking tech, like Tomahawk, are bringing it to large-scale production.

Justy This is a big step in OpenAI's full-stack strategy. They're not just building models or products, but the infrastructure underneath.

Cody It's a flywheel effect: better infrastructure drives compute efficiency, which enables better models and products.

Justy And the goal is to make advanced AI more broadly available. Inference is where AI reaches people.

Cody Every improvement in cost, speed, and reliability can show up as a faster ChatGPT answer or a more dependable API product.

Justy Exactly. This is about making AI more accessible to more people.

Cody The question is how it stacks up against Nvidia's H100 and upcoming B100 chips.

Justy Yeah, that's the real test. Even a 30% efficiency gain would be significant.

Cody We'll have to wait for the detailed technical report to see the full performance numbers.

Justy Right. But this is a promising start. Who knew OpenAI was getting into the chip game?

Cody Well, they're not just building models anymore.

Justy That's an understatement. Okay, so what's your read on the technicals here?

Cody The technicals look solid. The fact that they're optimizing for LLM inference specifically should give them an edge.

Justy And the partnership with Broadcom is a big deal. They're a major player in the semiconductor space.

Cody Yeah, that gives them the manufacturing and networking muscle to scale this.

Justy Alright, well, I'm excited to see how this plays out. Jalapeño, huh?

Cody Yeah, not the most elegant name, but I guess it sticks.

Justy Well, it's a start. I'll catch you later, Cody.