Skip to content
    Skip to content

    OpenAI's Jalapeño Chip Arrives: Fast Timeline, Unfinished BenchmarksOpenAI's Jalapeño Chip Arrives: Fast Timeline, Unfinished BenchmarksOpenAI's Jalapeño Chip Arrives: Fast Timeline, Unfinished BenchmarksOpenAI's Jalapeño Chip Arrives: Fast Timeline, Unfinished Benchmarks

    AL
    Aria Lin

    June 25, 2026

    OpenAI and Broadcom delivered the first physical samples of Jalapeño, OpenAI's custom ASIC (Application-Specific Integrated Circuit designed to execute one class of computation faster and more efficiently than a general-purpose chip) for large language model inference, to CEO

    OpenAI's Jalapeño Chip Arrives: Fast Timeline, Unfinished Benchmarks

    OpenAI and Broadcom delivered the first physical samples of Jalapeño, OpenAI's custom ASIC (Application-Specific Integrated Circuit designed to execute one class of computation faster and more efficiently than a general-purpose chip) for large language model inference, to CEO Sam Altman and President Greg Brockman on June 24, 2026. The most revealing detail in the announcement is not the chip itself but what OpenAI still does not know about it: final performance measurements are explicitly described as incomplete, yet Broadcom's CEO has already told Reuters the chip matches Nvidia's Blackwell GPUs and Google's TPUs (Tensor Processing Units, Google's proprietary inference and training accelerators). The gap between the marketing claim and the caveat is the story every AI infrastructure team should read before accepting the announcement at face value.

    Jalapeño is OpenAI's first attempt to own a piece of its own compute stack, reducing dependence on Nvidia GPUs that are, in OpenAI's own framing, "in limited supply." Engineering samples are already running GPT-5.3-Codex-Spark at production target frequency and power in the lab, and OpenAI says deployment at data center scale is targeted before the end of 2026. A detailed technical report is promised, according to OpenAI's announcement, "in the coming months." For AI infrastructure teams, hardware procurement officers, and anyone modeling the long-term cost structure of LLM-at-scale, the arrival of Jalapeño changes the competitive topology of the inference market, even if the chip's actual performance envelope remains unconfirmed.

    What Happened

    wide-angle shot of semiconductor fabrication engineers in cleanroom gowns inspecting early ASIC engineering samples on a lit examination table, cool blue overhead fluorescents, shallow depth of field

    Just nine months after OpenAI announced a partnership with Broadcom, the first engineering samples of Jalapeño arrived in the hands of Altman and Brockman. That timeline, confirmed independently by both OpenAI's blog and reporting on the announcement, is unusually compressed for a ground-up ASIC. The chip was designed "from scratch," according to OpenAI, built around the fundamental requirements of LLM inference (the process of running a trained model to respond to a user request, distinct from the far more compute-intensive training phase) rather than adapted from an existing architecture.

    Broadcom's role in the collaboration was chip implementation: translating OpenAI's architectural requirements into a manufacturable design. A third partner, Celestica, is named in OpenAI's announcement as responsible for board and rack system integration, high-performance networking, and scalable production systems. Celestica's name does not appear in some third-party reporting on the announcement, an omission worth flagging for anyone assessing the full supply chain behind Jalapeño's deployment.

    The engineering samples are not sitting idle. OpenAI confirms that Jalapeño is running ML workloads in the lab at production target frequency and power, including GPT-5.3-Codex-Spark, a specific model version named in OpenAI's official blog. OpenAI describes Jalapeño as "designed with flexibility to work with all LLMs guided by OpenAI's insights into the inference needs of current and future AI models across the industry," framing it not merely as an internal cost-reduction tool but as a platform with potential external applicability. No pricing or licensing terms for third-party use have been disclosed, making that claim difficult to evaluate.

    Why It Matters

    OpenAI's strategic motivation is explicit: reduce reliance on Nvidia GPUs, which the company describes as being "in limited supply." The custom silicon path is, in OpenAI's framing, another way to squeeze out more capacity amid a global compute crunch. For a company running some of the most inference-heavy products in the industry, including ChatGPT and the Codex family, the cost and availability of compute is not an abstract concern.

    OpenAI's performance-per-watt claim deserves careful reading. The company states that "early testing shows that Jalapeño will deliver performance per watt substantially better than current state-of-the-art," then immediately qualifies it: OpenAI is "still measuring final performance." No numeric figure is attached to "substantially better." No independent benchmark data accompanies the claim. The phrase "current state-of-the-art" is undefined in the announcement, leaving open whether the comparison baseline is Nvidia Blackwell, an earlier Nvidia generation, Google's TPU v8t, or something else entirely.

    The strategic framing OpenAI uses is deliberate. OpenAI calls Jalapeño "the first step in a multi-generation compute platform," signaling that this is an infrastructure investment with a multi-year horizon, not a one-cycle experiment. Custom silicon programs at this scale carry substantial non-recurring engineering (NRE) costs that can run into the millions of dollars per design iteration, meaning OpenAI is committing to a capital-intensive roadmap, not hedging with a single chip.

    Competitive Landscape

    tight macro photograph of a custom ASIC chip package freshly unsealed from protective packaging, gold bond wires and silicon die visible under magnification, warm amber light raking across the substrate surface, shallow depth of field, macro lens.

    Broadcom CEO Hock Tan, speaking to Reuters, said Jalapeño "matches the performance of Nvidia's Blackwell chips and Google's Tensor processing units." That claim carries significant market weight and deserves careful framing: it is a paraphrase reported via Reuters, not a verbatim quote extracted from a formal technical disclosure. No independent benchmark data supports it. Broadcom has a commercial interest in the success of Jalapeño, making Tan's characterization closer to a partner endorsement than a third-party assessment.

    The competitive field OpenAI is entering is already occupied by companies with years of custom silicon experience:

      • Nvidia: The incumbent GPU supplier whose Blackwell architecture is Jalapeño's named performance target. OpenAI's entire motivation for building Jalapeño is to reduce dependence on Nvidia's supply-constrained hardware.
      • Google: TPUs have been running inside Google's data centers since 2016, announced publicly at Google I/O in May 2016. Broadcom has co-developed TPUs since their inception, managing ASIC design and fabrication through foundries including TSMC.
      • Microsoft, Meta, and Amazon: All three have launched custom AI chip programs. Amazon's silicon division, originating from the 2015 acquisition of Annapurna Labs, has shipped multiple generations of Trainium and Inferentia, with Trainium2 powering Project Rainier for Anthropic's Claude models.

    Independent analyst commentary specifically on this announcement was not publicly available at publication time.

    The Bigger Picture

    OpenAI's stated deployment ambition is "gigawatt scale" with data center partners. That phrase appears in OpenAI's official announcement without engineering definition. A gigawatt is a unit of electrical power; used as a deployment descriptor without context, it is a scale claim rather than a technical specification. No confirmed watt-hours, rack counts, server counts, or data center locations accompany the statement.

    What the "gigawatt scale" language does signal, alongside the Celestica integration partnership, is that OpenAI is not treating Jalapeño as a lab curiosity. The board-and-rack integration work, high-performance networking, and scalable production systems that Celestica is building are the unglamorous infrastructure that determines whether a chip moves from engineering sample to deployed fleet.

    The broader industry implication is structural. When a hyperscaler designs its own inference chip, the cost structure of running LLMs at scale shifts. If Jalapeño delivers its promised performance-per-watt gains, OpenAI's cost per inference token declines relative to a pure Nvidia-GPU-dependent operation. That margin improvement either flows to OpenAI's bottom line or funds lower API pricing, both of which reshape competitive dynamics for every company building on top of OpenAI's models. Custom silicon is, at its core, a bet that vertical integration of the compute stack produces durable cost advantages -- the same logic that drove Apple's transition to Apple silicon and Google's decade-long investment in TPUs.

    What's Next

    over-the-shoulder medium shot of a technician holding a Jalapeño ASIC engineering sample board toward a bright window, chip package centered in frame, high-key daylight flooding the scene, 85mm lens with soft natural fill.

    OpenAI has stated an end-of-2026 deployment target for Jalapeño, with that timeline appearing in both the official OpenAI announcement and independent reporting. Between now and that deployment window, several significant unknowns remain unresolved.

    OpenAI has promised, according to its official blog, a detailed technical report "in the coming months." That report will be the first opportunity for the broader research and engineering community to evaluate the launch announcement claims against documented specifications. Key figures absent from the current disclosure include process node and die size, transistor count, memory bandwidth and memory configuration, interconnect specifications, and numeric performance benchmarks on standardized workloads.

    The multi-generation platform framing in OpenAI's announcement implies a roadmap beyond Jalapeño, though no dates, names, or specifications for subsequent generations have been provided. OpenAI's flexibility claim -- that Jalapeño is "designed to work with all LLMs" -- also remains unsubstantiated without licensing or access terms. Whether Jalapeño becomes purely internal infrastructure or a platform other organizations can deploy on will materially affect how the industry interprets the broader strategic intent behind the program.

    For software developers and ML engineers building inference pipelines: Jalapeño is an ASIC purpose-built for LLM inference, not training, and it is already running GPT-5.3-Codex-Spark at production target frequency. If OpenAI's performance-per-watt claims hold once the technical report arrives, the practical implication is lower inference latency and cost on OpenAI-hosted endpoints. Engineers evaluating whether to self-host models on Nvidia hardware or route through OpenAI's API will want to track that report closely: the actual memory bandwidth and throughput figures will determine whether Jalapeño meaningfully changes the build-vs.-buy calculus for inference at scale.

    The chip industry has a well-established pattern: performance claims at announcement, benchmark reality at deployment, and market share shift over the following product cycle. OpenAI is claiming Blackwell parity before it has finished measuring its own chip. The technical report, expected within months according to OpenAI's announcement, will be the first real data point. Until then, Jalapeño is a credible strategic move by a company with genuine incentive to break its Nvidia dependency, wrapped in marketing language that the engineering community should treat as provisional. The nine months from partnership to engineering sample is genuinely fast. The gap between "we're still measuring" and "substantially better than state-of-the-art" is genuinely wide. Both things are true simultaneously.

    -- Aria Lin, Enterprise Technology Analyst


    Sources: The Verge · OpenAI (official announcement) · Reuters

    More on Revuzia