The surge in daily AI queries is now outstripping the ability of standard chip designs to keep up, as moving data between separate memory and processing units consumes the bulk of power and time. Hyperscale operators report that memory access alone accounts for around two-thirds of the energy used when serving recommendation models to millions of users. This is not a problem of model size but of sheer volume: every fresh inference requires repeated data fetches across a narrow connection that was never built for billions of operations per day.
Why the classic chip layout is hitting its limit
Traditional processors follow a decades-old pattern in which calculations and storage sit apart. Each time a model needs its trained weights or fresh inputs, data travels across a limited pathway, adding delay and heat. As businesses adopt real-time uses such as personalised product suggestions or on-site image checks, the volume of these transfers rises sharply. The result is higher electricity bills for data centres and slower response times for users, even when the silicon itself still has spare capacity.
Design teams are therefore testing new layouts that place processing elements inside or immediately beside memory arrays. The approach cuts the distance data must travel, lowering both power draw and latency. In return, the chips become more specialised, trading some flexibility for better performance on the narrow task of running already-trained models.
What this means for smaller UK businesses
Most small firms will not buy these chips directly. They will feel the change through the pricing and performance of cloud services they already use. Providers that adopt the new designs should be able to offer inference at lower cost per query or with faster results on the same hardware budget. That could make features such as chat-based customer support or inventory forecasting more affordable for companies with modest IT spend. At the same time, firms locked into older infrastructure may see relative costs rise if their providers delay upgrades.
What to watch next
The first commercial chips built around in-memory or near-memory compute are expected from major foundries within the next two years. Watch for announcements from the large cloud platforms on how they price new inference tiers, and note any early moves by UK data-centre operators to pilot the technology. Those signals will indicate how quickly efficiency gains reach everyday business tools.
