Xiaomi AI Cube: the mini-PC that runs AI without the cloud


Xiaomi has unveiled the Xiaomi AI Cube, a mini-PC designed to run artificial intelligence models entirely locally, with no cloud connection. Presented on 24 August 2026 at the company’s XRING technology conference, the device is built around three in-house chips: the Xring O3, O100 and D100. It claims the ability to run language models ranging from 3 to 120 billion parameters, inside a compact aluminium chassis capable of sustaining 150 watts continuously. One caveat, however: this is currently only a prototype, with no price and no release date announced.

It remains an important signal. After Apple, Nvidia and a handful of start-ups, Xiaomi is positioning itself in the emerging market for personal artificial intelligence machines. Here is what we know.

Xiaomi AI Cube: three in-house chips in one box

jiqizhixin-2092073413116338228-03-1024x768 Xiaomi AI Cube: the mini-PC that runs AI without the cloud

The device’s distinctive feature is its architecture. Where most mini-PCs rely on a single processor, the Xiaomi AI Cube combines three internally designed chips, each with a distinct role.

The Xring O3 acts as the main processor. Built on 3 nm, it carries a ten-core CPU made entirely of performance cores, a 16-core G2-Ultra NX GPU, and a low-power NPU claiming 200 TOPS. TOPS, short for Tera Operations Per Second, measures how many artificial intelligence operations a chip can perform each second. It is the same chip that will power the Xiaomi 18 Fold in September.

The Xring O100 is the dedicated accelerator. Built on 6 nm, it uses three-dimensional wafer-level stacking, which vertically superimposes compute logic and memory. Xiaomi announces 28,672 effective data lines and near-memory compute bandwidth of up to 1.22 TB/s.

The Xring D100, finally, is a 3 nm chip originally designed for intelligent driving. It carries a 20-core CPU, a 16-core NPU and supports up to 160 GB of memory.

Why memory bandwidth is the real issue

One point deserves explanation, because everything else depends on it. When you run a language model locally, raw compute power is not the main limiting factor. It is memory bandwidth, meaning how fast data can move between memory and the compute units.

A 120-billion-parameter model, even compressed, occupies several tens of gigabytes. For every word generated, a significant share of those parameters must be read. If memory cannot keep up, the processor waits, and text generation becomes laborious regardless of the power printed on the spec sheet.

That is exactly what the Xring O100’s near-memory architecture targets. By physically bringing memory cells closer to the compute units, you reduce the distance data travels, therefore latency and power consumption. This approach is currently one of the most active research directions in artificial intelligence hardware acceleration.

A chassis designed to dissipate 150 watts

jiqizhixin-2092073413116338228-02-1024x768 Xiaomi AI Cube: the mini-PC that runs AI without the cloud

The mechanical side is not incidental either. The casing is a unibody in aerospace-grade aluminium, drilled with 33,874 precision CNC cutouts according to Xiaomi. That extensive perforation serves thermal dissipation as much as aesthetics.

The device is designed to sustain 150 watts continuously. That is a significant figure for such a compact format: it implies serious thermal engineering, and suggests Xiaomi is targeting sustained workloads rather than occasional peaks. Running a language model for hours has nothing in common with a three-minute benchmark.

Xiaomi also mentions a toggle between a fast mode and a slow mode, depending on the task at hand. That logic echoes the increasingly common two-speed architectures in artificial intelligence systems, where a light model handles simple queries and a heavier one takes over for complex requests.

A market beginning to take shape

The Xiaomi AI Cube is not landing in a vacuum. The category of desktop machines dedicated to local artificial intelligence has been building for just over a year, driven by a specific demand: developers, researchers and companies wanting to run models without sending their data to a cloud provider.

The motivations are concrete. Confidentiality first, for sectors such as healthcare, law or defence where outsourcing data is problematic. Cost next, since billing every request to a provider quickly becomes heavy at scale. Latency finally, as a local model responds without depending on the network.

Xiaomi’s difference lies in its vertical integration. Where most competitors assemble third-party silicon, the brand uses its own chips end to end. On paper this allows finer hardware-software optimisation and cost control. In practice, everything will depend on the maturity of the software stack, an area where established players enjoy a tooling ecosystem built over years.

Prototype: the word that calls for caution

jiqizhixin-2092073413116338228-01-1024x768 Xiaomi AI Cube: the mini-PC that runs AI without the cloud

This bears repeating clearly, because it is the most important point of the announcement. Xiaomi presents the AI Cube as a prototype. No price has been communicated. No availability date has been announced. No commercial market has been specified.

The device could therefore remain a technical demonstration, intended to showcase Xring chips to investors and partners without ever becoming a commercial product. That is a frequent scenario in the semiconductor industry, where reference platforms serve to prove an architecture’s capabilities.

Were the AI Cube to ship, several unknowns would remain. How much memory in the final configuration? Which operating system? What compatibility with standard frameworks such as PyTorch or llama.cpp? What update policy? Those answers, far more than the quoted TOPS, will determine the machine’s real value.

What to take away

The Xiaomi AI Cube is less a product than a statement of intent. It shows that Xiaomi no longer conceives its Xring chips as a mere smartphone component, but as the foundation of a hardware platform spanning phone, computer and automotive.

For the European user, immediate relevance is nil: no price, no date, no announced market. For the industry observer, however, the signal is clear. Local execution of artificial intelligence models is gradually leaving the laboratory to become a commercial argument, and more and more manufacturers want a share of it.

The real question is no longer whether models will one day run on our personal machines, but at what price and with what ease of use. We will follow this story closely.

Would you run an artificial intelligence model locally at home rather than use an online service? Tell us in the comments what holds you back or motivates you, and find our other articles on on-device artificial intelligence on Wanda-techs.

Read also:

Xiaomi Xring O3: the first mobile chip to surpass 5 million on AnTuTu

Share this content:

Ingénieur passionné et rédacteur web depuis 2018, j'allie mon expertise technique à ma passion pour l'écriture pour partager astuces, actualités et savoirs pratiques avec la communauté.

Post Comment

Vous avez certainement manqué...