IBM took the Scorching Chips 2026 stage to element its future IBM Z and LinuxONE processors and the AI inference acceleration chipset that sits beside it. IBM Distinguished Engineer Christian Zoellin is strolling via a dual-ISA core that natively runs each z/Structure and Arm AArch64, after which coated a second-generation on-chip AI accelerator geared toward bigger enterprise inference workloads. We beforehand acquired to see The IBM z17 Mainframe Brings AI with Telum II and Spyre and The New IBM z17 Telum II Processor Module Cut Open Down to Silicon in cool items.
That is being performed dwell from the session, so please excuse typos.
IBM Z and LinuxONE Twin-ISA Processor and AI Acceleration at Scorching Chips 2026
IBM opened with the historic context for Z, tracing a lineage that runs from out-of-order execution and 64-bit Linux via pervasive encryption, post-quantum safety, and now AI. This arc frames Z because the reliability base onto which IBM is attaching extra fashionable capabilities. One thing like 70% of the world’s transaction quantity goes over IBM Z mainframes. It says that by carry within the Arm ecosystem, it’s bringing in a brand new class of functions to IBM Z mainframes.

Now comes the core of the announcement, a mainframe-grade, dual-ISA processor. This chip carries 11 IBM Z cores at 5.7+ GHz on a 2nm course of, and every core can natively execute each z/Structure and AArch64. IBM pairs these cores with massive 36MB non-public L2 caches joined right into a digital L3/L4, plus a devoted on-chip DPU for I/O acceleration and devoted blocks for AI, compression, cryptography, and type. That is additionally a SMT=2 design and has a 432MB Digital-L3 and three.5GB digital L4 cache. That was not on the slide we acquired, however they confirmed these specs dwell.

A key alternative right here is that IBM applied AArch64 in full {hardware} relatively than via translation. This design makes use of a little-endian Arm implementation alongside big-endian z/Structure, with AArch64 v9.3, SVE and SVE2 help, and a pair of,792 applied AArch64 directions. IBM additionally claims Arm SystemReady compliance, which issues for the way a lot off-the-shelf Arm software program this core can take in. That is completely loopy know-how. Arm software program sees a local Arm processor. Arm runs unmodified, out-of-the-box, and onto a normal Arm platform. IBM stated the 2792 AArch64 directions are greater than twice the Z directions. IBM made a humorous quip about “diminished” in RISC.

Department prediction exhibits how a lot of the prevailing Z core design IBM was capable of reuse for AArch64. Automation consumes Arm’s XML structure descriptions to feed decode, whereas dispatch and problem repurpose register rename for the GR16-31 vary. IBM calls out new management for SVE, new dataflows for FP16, Bfloat16, and crypto, and even non-obvious CISC reuse corresponding to reminiscence copy and clear. I’m sitting right here nonetheless in awe of what IBM is doing right here, this isn’t Z+Arm cores, that is Z and Arm in a single core.

On the software program aspect, IBM exposes the Z accelerators as platform units to Linux on Arm. Crypto, GZIP compression, and the on-chip AI unit floor as Linux units with comparable latency to native Z directions, whereas the Z ISA can nonetheless expose them as directions for s390x workloads.

Reliability carries over to this design with a 99.999999% availability goal. IBM cites error checking throughout arrays, dataflows, and management, clear restoration from transient faults, core sparing for persistent faults, and concurrent restore, plus RAIM reminiscence safety. In case you noticed our video on the Z17, the engineering that IBM did to make its techniques dependable may be very totally different from common objective cloud servers. Truly, NVIDIA has been doing a little comparable issues like changing many cables with PCB to extend reliability of its techniques.

Reasonably than forcing a single ecosystem, IBM positions the cores to coexist via Linux KVM and OpenShift Virtualization. This determine maps s390x Linux and ARM64 Linux alongside IBM z/OS, with logical partitions sharing the identical processor. A thread can run both s390x or Arm software program and the change takes nanoseconds.

Collectively, that coexistence is the pitch for getting one of the best of each software program ecosystems, letting Z clients hold mainframe code whereas tapping the extensive Arm software program base. That’s the reason IBM didn’t simply combine current Arm cores.

IBM’s second-generation AI inference accelerator targets bigger enterprise workloads. This chip packs 16 energetic AI cores plus one redundant core, with FP4 and MXFP4 datatypes that IBM says can ship as much as 4x TOPS. Superb, that is AI with redundancy. It provides 96GB of HBM3e operating as much as about 4TB/s, roughly 20x the reminiscence bandwidth of the present technology, and makes use of PCIe Gen6 as a low-latency peer-to-peer interface.

On the safety and resilience aspect, IBM positions this half for mission-critical AI with confidential computing that protects information and fashions at relaxation, in transit, and in use, together with quantum-safe cryptography. Safe boot and on-chip cryptography mix with a firmware stack tuned for availability and serviceability and tight operational integration with IBM Z.

IBM packs quite a bit into this roadmap, from a local dual-ISA core that merges the mainframe reliability story with the Arm ecosystem to a severe enterprise AI inference half for bigger workloads.
Closing Phrases
I’ve sat via virtually a decade of Scorching Chips displays. This is likely one of the displays I’m virtually in awe of. IBM’s dual-ISA wager is a bid to maintain the mainframe related as Arm software program and AI workloads reshape the info middle, relatively than treating Z as a closed platform. The truth that that they had this concept, and applied it the arduous manner I might have had no idea of two years in the past. After doing the Z17 launch content material, I do know the reply is simply that IBM has engineers that tackle loopy arduous issues. Nonetheless, it’s superior. It is a 2nm core with native AArch64 execution and the second-generation AI accelerator tackle two pressures without delay, software program ecosystem breadth and AI inference density. How a lot of the encompassing Arm software program stack runs unmodified on the core shall be price watching.
