Industry & CompetitionAI Products & Platforms

On-Device AI Has No Leader, Because It's Turning from a Selling Point into a Component

In September 2026, Desert Ant Labs, a small European team, launched 18 models capable of running directly on-device, offering a free tier of 100,000 monthly active devices per platform and per model. On launch day, it surged to the front page of Hacker News, sparking lively discussion in the comments. Someone dug up the model card and found that the team’s speech-to-text model, Voz, was a port of NVIDIA Parakeet—a fact the founder soon acknowledged in the comments.

A launch like this makes it easy to feel that on-device AI is kicking off a whole new battle, and I initially had similar expectations myself. But looking past the buzz, a new puzzle emerges: with so much discussion, why can’t we find even a single well-known market leader in on-device AI?

Looking closely at real shipments across the supply chain, the problem actually lies in the label “on-device AI company” itself, not in the market being too small. None of the companies shipping at massive volume and generating real, solid profits classify themselves as an on-device AI company. Yet even within the industry, debates over whether on-device AI has real value have been raging inconclusively for years.

On-device AI is shifting from a keynote headline selling point into a component inside every device that OEMs no longer bother to market separately

A Debate That Never Settles

Discussions around on-device AI have long been split between two diametrically opposed, yet equally emphatic, views. The optimistic camp sees it as the next sure bet, backed by solid commercial figures: Ambarella has shipped roughly 45 million edge AI chips to date; Mobileye’s compute silicon is inside more than 230 million vehicles; and NVIDIA’s Jetson edge computing platform has gathered 2 million developers. These figures come straight from official earnings reports and corporate disclosures, with no inflated fluff.

The pessimistic camp, on the other hand, dismisses it as a marketing gimmick, and real-world user experience feels just as convincing: the AI PCs consumers paid a premium for have dedicated AI chips that sit idle most of the time; the shiny new AI features that smartphone makers push so aggressively are rarely used day-to-day by most buyers, who are unwilling to pay extra for them anyway. In reality, neither side is fabricating facts. The root of the conflict is that everyone uses the single phrase “on-device AI” to refer to three fundamentally different businesses at the same time. Lumping them together inevitably leads to endless disagreement. To make sense of this market, we first need to break down these three separate equations.

On-Device AI Is Actually Three Different Things

Narrow-task intelligence has been running in industrial environments for two to three decades. Defect inspection on factory lines, license plate recognition at access gates, and voice wake-words in noise-canceling headphones were all ubiquitous long before “on-device AI” became a buzzword, relying primarily on classical computer vision and digital signal processing. For the aforementioned Ambarella and Mobileye, the bulk of shipments has always been for these mature workloads, bearing little relation to the large models being discussed today.

At the same time, PC and smartphone giants have leveraged existing hardware refresh cycles to package on-device computing as a fresh selling point. According to forecasts from Gartner, worldwide AI PC shipments were projected to reach 77.8 million units in 2025, accounting for 31% of the global PC market. Yet the core driver behind this massive volume was the major upgrade wave triggered by the end-of-support for Windows 10, having little to do with users actively seeking an AI experience. The share where users genuinely paid a standalone premium for AI capabilities was far smaller: PCs powered by Qualcomm’s flagship debut chip, Snapdragon X, sold only around 720,000 units in Q3 2024, representing just 0.8% of all PCs, according to reporting from TechPowerUp.

As for the new explorations centered around running large models locally, that is a much more recent phenomenon, and it represents the direction teams like Desert Ant are pursuing. However, this layer of exploration is currently being squeezed from both ends—by the cloud above and the silicon platform below.

The term on-device AI actually encompasses three different things: decades of embedded narrow tasks, hardware refresh cycles, and recently emerging local large models

Lumping these three distinct businesses into an umbrella term like “on-device AI” inevitably yields contradictory market forecasts. This divergence is not normal market fluctuation; the crux lies in how data is tracked: some research firms count only dedicated silicon, while others bundle the entire retail price of any machine with even modest intelligent computing capabilities. Relying on such ambiguous aggregate numbers makes it nearly impossible to draw valid conclusions. To truly see who is making money, you have to follow where the physical hardware shipments actually go.

The Real Giants Never Call Themselves On-Device Companies

If you search for an on-device AI leader through the lens of an independent category, you will come up empty-handed. But shift your perspective, and the giants are already everywhere: Apple, Qualcomm, NVIDIA, Google, MediaTek, Intel, and AMD have all deeply integrated AI processing units into their silicon and operating systems. In Q4 2025 smartphone SoC shipments, MediaTek held a 30% share, Apple held 23%, and Qualcomm held 22%, according to figures from Counterpoint. And in AI-capable smartwatches, Apple Watch alone accounted for roughly 90% of shipments. None of these giants would ever claim to be an “on-device AI company,” because to them, local intelligence is simply a standard building block in a complete computing chip.

By contrast, the independent startups that once rode the halo of on-device AI have struggled to survive as standalone commercial brands. Israel’s Hailo saw its valuation drop from $1.2 billion to under $500 million, cutting roughly half its team before Microchip signed an agreement to acquire the company in July 2026; valuation reporting can be found at Calcalist, and the acquisition signing at Microchip. Google’s original Coral edgetpu repository was archived and made read-only in April 2026, as recorded on GitHub. Qualcomm outright acquired Edge Impulse, a TinyML platform. Teams that started out focused entirely on on-device AI ultimately ended up as supporting components inside the ecosystems of major chipmakers.

Desert Ant occupies much the same niche today. The showcase page for its Voz model explicitly states that the solution is optimized for Apple’s Neural Engine and ported from NVIDIA Parakeet TDT 0.6B v3, as seen on the Desert Ant website. Porting an open-source model onto Apple’s hardware architecture and achieving great inference speed certainly requires solid engineering optimization chops. Yet the commercial value is tied entirely to hardware-specific tuning; the model weights themselves offer no moat. On launch day, its core GitHub repository garnered only about 111 stars, and most of its models have fewer than a thousand downloads.

It is easy to raise a counterargument here: mobile computing and cloud services also started out as mere deployment form factors, yet both went on to spawn massive business landscapes. The watershed distinction is that mobile platforms and cloud computing both left an expansive, independent application layer above their foundational hardware and software. In contrast, underlying operating systems and silicon vendors have absorbed and integrated emerging on-device AI capabilities at breakneck speed, leaving virtually no room for an independent software layer in between. The value of on-device AI has not diminished; it simply hasn’t settled at the on-device software layer. Without a standalone application layer to anchor it, on-device technology quickly hit a ceiling in consumer-facing products.

The Consumer Side: Why the Hype Falling Flat Is Real This Time

The lukewarm sentiment reflected in consumer electronics has been remarkably consistent, with multiple cross-channel surveys pointing to similar conclusions. A late-2024 sample survey showed that 73% of surveyed iPhone users and 87% of surveyed Samsung phone users reported seeing little to no value in AI features, while 86.5% of Apple respondents and 94.5% of Samsung respondents explicitly stated they were unwilling to pay extra for AI capabilities, according to data from SellCell. Another consumer tracking study revealed that only 11% of respondents would upgrade their phones for AI features—a 7-percentage-point drop from the previous year, as reported by CNET.

The shift within the PC camp has been even more direct. At its 2026 Build developer conference, Microsoft abandoned its NPU-exclusive policy for Copilot+, allowing associated AI features to run across general-purpose CPUs, GPUs, or other silicon, according to PCMag. In an editorial, Windows Central bluntly pointed out that Copilot in practice relies entirely on cloud services and does not genuinely leverage the NPU inside personal devices.

Even Apple, with its massive investments in on-device hardware R&D, ran into the hard physical limits of terminal compute. The much-anticipated overhaul of Siri was delayed by roughly 18 months from its initial 2024 announcement before Apple ultimately decided to integrate Google’s Gemini models to power the AI-infused Siri, according to reporting by CNBC. The consumer hardware company with the world’s largest installed base demonstrated through its own choices that when dealing with genuinely complex multimodal large model workloads, even premier silicon design capabilities cannot avoid relying on cloud infrastructure in the end. Yet the chill in consumer electronics does not mean on-device technology has nowhere to go; outside the public eye, a completely different industrial logic is quietly delivering commercial value.

The Industrial Side: Real Money, Hidden from View

The frontline where on-device AI generates steady, sustained profits is tucked away in specialized scenarios far removed from public discourse. High-end hearing aids offer a textbook example. The Infinio Sphere hearing aid from Phonak relies on a dedicated on-board compute chip to run a neural network noise reduction model locally. An official technical white paper notes that this approach improves the signal-to-noise ratio by up to 10 dB and boosts speech understanding for wearers by up to 36.8%, as detailed in technical documentation from Phonak. The human ear is unforgiving when it comes to millisecond-level acoustic latency, and environmental noise suppression cannot tolerate the network round-trip of the cloud. The triple hard requirements of ultra-low latency, offline availability, and privacy ensure that compute must remain inside the device right next to the ear canal.

Home cleaning robots operate on the same logic. Roborock states on its official website that all vision and obstacle avoidance computations are processed locally on-device, with no inference data sent to the cloud, as seen on the Roborock website. Inquiries by tech reporters to multiple robot vacuum makers also confirmed that the obstacle photos and maps captured by Roborock vacuums are stored locally on the device, with the cloud serving only as a temporary relay without persistent storage when viewing photos. The brand recorded shipments of 5.8 million units in 2025, proving that this local visual inference architecture is battle-tested production engineering at scale, extending far beyond the realm of trade show concept demos.

In traditional industrial automation and transportation, the commercial scale of local intelligence is even more massive. Schneider Electric, a customer of industrial machine vision leader Cognex, upgraded more than 100 of its factories worldwide into smart production lines, reducing false rejections in visual inspection to roughly 1/70th of previous levels, as documented in a case study from Cognex. In the smart vehicle space, Tesla’s FSD system maintained 1.48 million active users (including both outright purchases and monthly subscribers) in Q2 2026, while Mobileye’s SuperVision assisted driving system has been mass-produced and deployed across production models from Zeekr, Polestar, and others.

These deployment scenarios share the same fundamental logic: local AI compute is the critical component that makes the product viable in the first place. A hearing aid stripped of instantaneous local noise cancellation loses its core utility. A factory production line requiring video frames to stream back to external servers in real time would fail stringent data compliance audits while failing to keep pace with cycle times. While these industrial and professional devices lack the marketing hype of consumer electronics, they form the most solid foundation for edge intelligence. This further clarifies the technical boundaries of on-device AI: its viability rests on irreplaceable, non-negotiable requirements, having little to do with simple cost calculations.

As the Cloud Gets Cheaper, What’s Left for On-Device?

The earliest and loudest argument for on-device AI was saving users inference costs and dodging pay-per-token API bills. Yet that argument is crumbling fast. GPT-4 was priced at $30 per million input tokens in March 2023; by April 2026, Google’s Gemini 3.1 Flash had dropped to $0.10—a decline of approximately 99.7% over three years (comparing list prices across generations and tiers), according to data from price comparison. Facing an ~80% plunge in API pricing from early 2025 to early 2026, the supposed cost advantage of on-device AI is continuously shrinking rather than expanding as once anticipated.

Physical constraints at the hardware level are just as unyielding. Smartphone memory bandwidth typically ranges between 50 and 90 GB/s, whereas data center GPUs offer 2 to 3 TB/s—a 30x to 50x disparity. Autoregressive decoding in large language models is strictly bound by memory bandwidth throughput. Benchmarks show that running a 1.5-billion-parameter small model on a Raspberry Pi 5 would cost just 0.06 cents if routed to a cloud API for the same 12 requests; at a query rate of once every 3 minutes tested in the review, recouping the hardware cost through API savings alone would take roughly two years—without even factoring in electricity costs, as detailed in a review by The DIY Life.

The defensible pillars for on-device compute have ultimately narrowed down to four hard constraints: latency, offline availability, on-device data residency, and legal compliance. These four metrics are genuine engineering constraints that directly sustain industrial real-time control, automotive driver assistance, intelligent hearing aids, and local security systems. Yet together, they point squarely toward highly vertical enterprise and professional hardware markets—they cannot support an all-encompassing, general-purpose consumer assistant.

As cloud inference prices plunge, on-device cost advantages shrink, leaving latency, offline capability, data residency, and compliance as the four remaining hard constraints

For developers and startups entering the on-device space, the commercial opportunity has shifted away from the old playbook of branding standalone models or launching generic developer kits. Instead, it has moved toward engineering delivery capabilities: precision porting, quantization, and fine-tuning of open-source small models tailored to specific silicon and targeted use cases. Desert Ant is practicing precisely this engineering path. The open question is how much commercial premium this specialized engineering capability can command, and how long that premium can last amid rapid hardware iterations—a question no startup has yet convincingly answered.

At this stage of industry evolution, rather than obsessing over the hypothetical total addressable market of on-device AI, a far more pragmatic approach is to pinpoint the specific business scenarios where on-device solutions hold a genuine, expanding advantage over cloud services. Mapping the boundaries of hard constraints in concrete use cases is vastly more meaningful than sketching out abstract, grandiose market projections.