There is no such thing as an AI chip

Public conversation about AI hardware has collapsed onto a single component, and that component is the GPU. It is convenient shorthand, and it is wrong. No model is trained or served by one kind of silicon. An AI compute system is an assembly of general purpose processors, matrix accelerators, memory hierarchies, interconnect fabrics and power management circuits that have to behave as one machine, and the bottleneck moves between them depending on what the system happens to be doing at that moment.
Working out where the bottleneck sits has become the central hardware design exercise of the last few years. It is also why the industry keeps drifting toward purpose built silicon.
Training is not a single workload
A training cycle breaks down into jobs that have almost nothing in common with each other. Dataset preparation is serial, irregular work: decompression, tokenization, shuffling. That work runs on CPUs, along with job orchestration, scheduling, checkpointing and failure recovery, which at cluster scale is a routine event rather than an exception. A CPU in this setting is not there to do math. It is there to keep the accelerators fed.
Then there is the computational core, matrix multiplication, handled by dedicated low precision units. Arithmetic density has grown faster than anything else in the system, to the point where the math units are rarely the constraint. Getting data to them is. Reading an operand from external memory costs far more energy than multiplying it, and the ratio between compute capability and memory bandwidth gets worse with every generation. This is why high bandwidth memory stacked next to the compute die stopped being optional, and why the design of the interface to that memory matters as much as the number of arithmetic units sitting behind it.
Above all of this sits communication. Distributed training lives on collective operations that synchronize gradients across every device in the job. The network decides how much of a system's theoretical compute actually turns into shorter training time. It also shapes the electrical behavior of the whole machine, because a large population of accelerators that stalls together waiting on a synchronization and restarts together produces current transients that no conventional power infrastructure was designed to absorb. That point comes back later, and it matters more than it sounds.
Inference runs the other way around
Serving a model is a different problem, and the industry has only recently internalized how different. Generative inference splits into two phases with nearly opposite requirements. During prefill the model processes the whole prompt in parallel, saturates the arithmetic units and is compute bound. During decode, tokens come out one at a time, and almost all of the time goes into rereading weights and the context cache, so the limit is memory bandwidth and most of the available compute sits idle.
An accelerator built to maximize training throughput does poorly at decode. On a single device that inefficiency is tolerable. Across an inference fleet it turns into the cost line that decides whether a service is viable. Hence the appearance of architectures that split the two phases across different hardware, devices that trade arithmetic capability for on chip memory, and the steady push of workloads toward the edge, where the constraint changes again because there is no active cooling, the thermal budget is tight, and the metric that counts is energy per inference rather than throughput.
Why ASICs are becoming the obvious answer
Programmability has a physical cost you can measure. Instruction fetch and decode, register files, control logic, generic cache hierarchies, datapaths sized for cases a given workload will never exercise. While the workload is still moving, that flexibility is insurance worth paying for. Once the workload settles, and attention based architectures have settled a great deal, flexibility stops being an advantage and becomes silicon that draws power and occupies area without contributing to the result.
A purpose built device takes that cost off the table. It fixes the dataflow instead of reassembling it on every instruction. It sizes on chip memory around the tensors it will actually hold. It implements the numeric formats the workload needs and none of the others. It pulls quantization logic into the datapath rather than wrapping it around the outside. And it balances arithmetic capability against the memory bandwidth genuinely available, instead of inheriting a ratio that somebody chose for a different application. The gains that come out of this are measured in multiples rather than percentages, which explains why every large infrastructure operator now runs an internal silicon program and why an entire generation of companies has built its value proposition around a proprietary accelerator.
There is an economic argument too, and it may be the one that actually tipped the balance. Chiplet disaggregation lets a team concentrate its investment on the die that carries differentiation and treat the rest as infrastructure, reusing I/O, memory interfacing and power management blocks on more mature nodes. The price of admission is no longer a monolithic device on a leading edge node. It is one well designed die packaged alongside others. That has opened the door to companies that a few years ago were priced out of custom silicon entirely.
What has not come down in the same way is the cost of time. Closing timing on a critical path, recovering area without giving up frequency, or working out which combination of synthesis and placement settings actually produces the better result remains an exploration problem measured in weeks of compute and waiting. On a silicon program, weeks are the currency you pay for everything else with. That is why we built Spaceman, which applies AI to compress that exploration inside the EDA flows teams already run, without asking them to change how they work.
So the question is no longer whether a custom device makes sense. It is which parts of the system are worth designing and which are worth buying. That is the question most of the teams we work with arrive with, and it is usually the right one, because the risk in a silicon program rarely concentrates in the datapath. The datapath is the part an architecture team understands better than anyone. The risk almost always shows up somewhere else, and in particular in two places.
Power delivery became a design problem, not a procurement problem
A modern accelerator draws very high current at very low core voltage, and it does so in a spectacularly uneven way. A matrix engine's load is synchronous by construction, so it switches an enormous fraction of the die on and off in the same cycle, and in distributed training that synchrony extends across the entire system. The result is a current profile with transients fast enough to drag the supply rail down before a conventional regulation loop can respond.
The obvious fix is to raise the voltage margin for safety. It is also the most expensive one, because every fraction of a volt of static margin is power burned as heat for the life of the part and clock frequency you have decided not to use. At system scale that margin turns directly into megawatts and into performance left on the floor.
An off the shelf PMIC cannot solve this, because it knows nothing about the load it is feeding. It is sized against an average profile, with control loops tuned for a generic use case and telemetry designed for diagnostics rather than for fine grained regulation. A PMIC designed alongside the die it powers is a different animal. It can be sized against the real transient behavior rather than a datasheet number, close a fast control loop with the adaptive voltage and frequency logic on the die, coordinate the multiple power domains and the power up and power down sequencing of a chiplet based system, and claw back static margin and turn it into efficiency or into clock speed.
This is part of why we do not treat dedicated PMIC work at Move Silicon as a job that sits apart from the accelerator. The load profile of the die is where the specification starts, not something that arrives downstream, and power conversion gets sized together with the on die distribution network and the package that has to house it. Done in sequence rather than in parallel, the result still works. It just works with a margin that somebody is paying for.
The interfaces hold everything else together
The second place risk concentrates is in the links. In an advanced package the function is spread across several dies that have to behave as if they were one, which puts an unusual amount of weight on the memory PHY, on the die to die links and on the interposer microelectronics that carries them.
These are dense mixed signal circuits with tight energy per bit budgets, signal integrity constraints across connections that run through a silicon substrate, and a very close coupling to both power integrity and the thermal behavior of the package. Die to die interconnect standards have made the problem more tractable, but conformance to a standard does not guarantee that a link works on your package, with your floorplan, against the noise your compute engine generates while it switches. Silicon decides that, and if silicon decides it late, the fix costs a mask set and several months. Which is why we treat signal and power integrity analysis as part of designing the interfaces rather than as a check applied to something already drawn.
The underlying point is straightforward. The next generation of AI systems will not be limited by how fast we can multiply matrices. It will be limited by how much energy we can get to the die without wasting it on the way, and how much data we can move between dies without paying too much for it. Teams designing silicon for AI are finding out that those two things sit at the center of the project rather than around its edges.
