For the past two years, almost all of the capital flowing into AI has gone to the data center. NVIDIA’s quarterly revenue exploded, and HBM and CoWoS owned the headlines every day.
“AI semiconductor” effectively became a synonym for “cloud AI semiconductor.”
Over that same period, Edge AI never got bundled into a single strong investment narrative the way cloud AI did.
The NPU inside your phone, the autonomous driving chip inside your car, the vision AI next to a factory line, the inference engine in the sensor on your wrist.
Some names have already run, but the market is not yet reading Edge AI as a single structural investment map. The reason is that the way revenue shows up is fundamentally different from cloud.
This article starts with what Edge AI actually is, then looks at how this market differs from cloud AI, and which layers capture the money and who captures it. At the end, I lay out the names I am watching.
Disclaimer
It does not recommend buying or selling any specific security, and all judgment and responsibility rest with the reader. The author may hold, or may come to hold, some of the names that appear in this article. Semiconductors are a sector that swings hard with cycles, the macro, and geopolitics. Always do your own research before investing.
Part 1. This Is What Edge AI Is
1. You Are Already Using Edge AI
You ask ChatGPT a question and the answer arrives a few seconds later. The question travels across the internet to a data center, an NVIDIA GPU produces the answer, and it comes back. A massive computer sits far away, connected by a network. This is cloud AI. For the past two years, “AI” has meant roughly this architecture.
Edge AI is the opposite. A small computer sits in your hand, in your car, next to a factory line.
On Google Pixel and the broader Android lineup, some translation and voice processing is gradually moving to local model inference. The photo correction and background removal on Galaxy phones are increasingly handled by the on-device NPU. Cloud dependence and language coverage still vary by feature, but the direction is clear.
Cars are more extreme. Recognizing the vehicle ahead, classifying a pedestrian, deciding whether to brake. None of that can be left to the round-trip latency of a server. Camera data has to be processed in real time by the AI computer inside the vehicle. The same is true for robots. Here, Edge AI is not a convenience feature. It is a safety condition.
Factory sensors follow the same structure. If you keep sending all the vibration, temperature, and sound data to a server, power and communication costs become a problem. Ultra-low-power MCU-based AI filters out only the anomalous patterns near the sensor and sends a signal only when needed. Security cameras, smartwatches, vision AI on logistics lines. Edge AI has already moved into everyday life everywhere.
AI computation happens right where the data is created. Speed does not work, or power does not work, or cost does not work, or privacy does not work. The reason it cannot go to the cloud differs, but the result is the same. It gets processed inside the device. This is Edge AI.
2. Tearing Apart an Edge AI Device
Tear apart a single Edge AI device and you find five kinds of components. Phone, car, factory camera, the structure is identical. Only the size and the performance differ.
The brain (SoC/NPU). The CPU, GPU, and NPU (the dedicated AI engine) sit on one chip. The NPU is the heart of it. AI inference is repeated matrix multiplication, and the NPU is a dedicated engine designed to do exactly this operation fast. Run the same workload on a general-purpose CPU and it is 10 to 100 times slower and burns far more power.
Phones carry Qualcomm Snapdragon or Apple Silicon, and cars carry NVIDIA Drive or Mobileye EyeQ. NPU compute is measured in TOPS (trillions of operations per second). If a phone NPU was 5 TOPS three years ago, it is now above 50 TOPS.
The memory. The AI model’s weights load into memory. The faster and larger the memory, the bigger the AI model that can run.
LPDDR5X is the current mainstream for phone memory, and in the next generation LPDDR6 has emerged as the leading candidate for improving bandwidth per watt. Three years ago the standard was 8GB. Now the trend is moving to 12 to 16GB, and to 24GB configurations at the high end. Quantize a 7-billion-parameter model to INT4 and it is roughly 4GB. Memory capacity directly sets the size of the model you can run.
The storage. This is where the AI model file lives. When a model first loads, it moves from storage into memory. The transition from UFS 3.1 to 4.0/4.1 has more than doubled read speed.
The eyes and ears (sensors). Edge AI starts from sensor data. Cameras, microphones, LiDAR, radar, magnetic sensors, and inertial sensors read the physical world.
A single car carries 8 to 12 cameras, 4 to 6 radar units, and 8 to 12 ultrasonic sensors. The more sensors there are, the heavier the NPU’s compute load and the higher the memory consumption.
The communication (modem/WiFi). This is used when an Edge AI device collaborates with the cloud or when devices exchange data with each other. The phone’s 5G modem, the PC’s WiFi 7, the IoT sensor’s LoRa all fall here.
These five components go into a single device at the same time. And the performance of each one rises every year. This is rising BOM content. Even if the number of devices shipped stays flat, the moment the semiconductor value packed into each device increases, a chip company’s revenue rises. This is the core mechanism of Edge AI investing.
3. Why Now
The technology crossed a threshold. Three years ago, running an LLM on a phone was a demo. Now it is starting to ship in commercial products.
A 10x jump in NPU performance, quantization that lets a model fit into phone memory, and the generational shift in LPDDR bandwidth. The point where these three locked together at once is now. None of them alone was enough. The three crossed their thresholds simultaneously, and on-device LLMs actually started to run.
Capital is looking for where to go next. The debate over whether cloud AI capex is heading toward a peak has begun. At the same time, the device replacement cycle of AI phones and AI PCs has opened. Cars are in the middle of an SDV transition that is exploding SoC performance and memory per ECU.
Big tech’s economic incentive is clear. Cloud inference cost rises linearly as users grow. Moving inference to the device can sharply cut cloud serving cost. By easing the structure where server cost grows alongside the user base, Apple, Google, Samsung, and Microsoft are all moving in the same direction.
That said, this theme does not get priced in all at once the way cloud AI did. Edge AI revenue recognition follows device shipments, and there is a two-to-four-quarter lag before rising BOM content shows up in revenue. An automotive SoC takes 24 to 36 months from design win to mass production. Some layers have already begun to re-rate, and in others revenue recognition is still catching up. Edge AI is less a single-theme buy than a layer-by-layer timing game.
One question remains. How should you read this market?
Part 2. How to Approach Edge AI Investing
Cloud AI Has a Center. Edge AI Does Not.
Cloud AI investing is relatively easy to read. The GPU sits at the center. When a GPU generation turns over, the matching memory, packaging, networking, and power design follow it. Track one thing and the direction of the rest of the chain becomes visible.
In Edge AI, that hierarchy collapses. The king changes from device to device. In phones, memory is king. In cars, safety certification is king. In factories, data capture is king. Knowing the phone’s bottleneck tells you nothing about the car’s. The very type of cost it takes to put AI into a device, the hardware tax, differs by device.
This is what it means to say Edge AI has no center. The type of tax differs, and the tax collector differs. In cloud AI, tracking NVIDIA reveals the whole chain, but in Edge AI, tracking any single company never reveals the whole. So to invest in Edge AI you have to understand each device’s bottleneck and find whoever resolves that bottleneck.
That is why Edge AI has to be viewed layer by layer. Among the five components of a device, the three places where the bottleneck shows up most sharply from an investment view are the memory in phones and PCs, the safety certification in cars and robots, and the sensors that read the physical world. I will go through each layer one at a time: what the bottleneck is, and what kind of company resolves it.






