OpenAI Jalapeño chip targets higher inference efficiency

OpenAI’s Jalapeño chip, built with Broadcom, begins lab testing as part of a new accelerator platform designed for faster, more efficient AI performance.

Lauren Collins ·

OpenAI Jalapeño chip targets higher inference efficiency

The Jalapeño chip, OpenAI’s first purpose-built inference processor developed with Broadcom, has entered early lab testing as the companies expand a multi-generation AI compute platform.

OpenAI said engineering samples are already running machine-learning workloads at target power and frequency, including a model identified as GPT‑5.3‑Codex‑Spark. The companies have not yet published benchmark numbers, but they said initial results indicate materially improved performance per watt versus current leading accelerators.

Custom inference silicon marks a full-stack shift

The announcement positions Jalapeño as the first “Intelligence Processor” in a longer roadmap of custom accelerators and infrastructure components jointly developed by OpenAI and Broadcom (NASDAQ: AVGO). The stated goal is to make advanced AI systems faster, more reliable, and available to more users by tightening integration across hardware and software.

According to the companies, Jalapeño was designed “from scratch” by OpenAI to match how large language models are served in production. The design is informed by OpenAI’s internal plans for models, optimized kernels, serving systems, and product requirements, aiming to better align silicon capabilities with real inference bottlenecks.

Broadcom executives Hock Tan and Charlie Kawwas delivered the first Jalapeño device to OpenAI leadership, including CEO Sam Altman and President Greg Brockman. The handoff was framed as a milestone in OpenAI’s strategy to control more of the technology stack that powers its products.

Architecture focuses on utilization and data movement

OpenAI and Broadcom emphasized an architectural approach that reduces data transfers and seeks a more balanced allocation of compute, memory, and networking. In practical terms, the companies said the chip is built to raise “realized utilization” closer to theoretical peak throughput, a recurring challenge when deploying LLM inference at scale.

The Jalapeño chip is described as flexible enough to support many large language models, not just OpenAI’s own systems. OpenAI said its design choices were guided by inference patterns expected in both current and next-generation models across the wider AI market.

While final performance measurements are still underway, the companies said early testing suggests a substantial gain in performance per watt compared with today’s state-of-the-art. They also said a more detailed technical performance report will be presented in the coming months, signaling that full benchmarking and configuration details have not yet been disclosed.

Broadcom, Celestica support production-scale buildout

Beyond the silicon itself, the partnership includes work to industrialize a broader compute platform. OpenAI cited Broadcom and Celestica as key partners across chip implementation, board design, rack-level system integration, high-performance networking, and scalable manufacturing systems.

Broadcom’s contributions include its networking technologies and Tomahawk networking silicon, which the companies said will help move the overall platform toward large-scale production. High-throughput networking is a critical constraint for inference clusters, where model serving performance can be limited by how quickly data moves between compute, memory, and distributed systems.

The announcement reflects a broader trend in AI infrastructure: as model sizes and inference demand increase, large developers are looking for tighter coupling between software stacks and specialized hardware. By tailoring an accelerator to model-serving workloads—rather than relying solely on general-purpose GPU roadmaps—companies aim to lower operating costs and improve reliability under heavy utilization.

Next steps will hinge on the promised technical report and on whether Jalapeño can sustain its early efficiency claims in broader testing and production configurations. Observers will also watch how quickly the multi-generation platform expands beyond initial samples into deployed systems and whether the “all-LLM” positioning translates into adoption across varied inference workloads.

More stories