First 2nm phone chips from Qualcomm, Apple and MediaTek focus on AI
September launches put TSMC's 2nm nanosheet process into phones, with chips that can run AI models of up to 30 billion parameters locally.
In September 2026, Qualcomm, Apple and MediaTek each launched flagship phone processors made on TSMC's 2 nanometre process, TrendForce reported on 6 October. They are the first phone chips to use gate-all-around nanosheet transistors. All three put on-device AI at the centre, and some can run AI models with about 30 billion parameters on the phone itself.
- 5 GHzCPU speed Qualcomm claims for the Extreme chip, a mobile first
- 44%faster graphics than last year, according to Qualcomm
- 2xthe AI processing power of the A19 Pro, Apple says
- 40%less power for always-on AI, MediaTek says
- 10-20%rise in 2nm bookings reported by Economic Daily News
What happened
TrendForce, a Taiwanese market research firm, published a comparison of the three new chip families on 6 October 2026. All three are made by Taiwan Semiconductor Manufacturing Company (TSMC) on its 2 nanometre (2nm) process. TrendForce says this marks the industry's move to gate-all-around (GAA) nanosheet transistors. Each company puts AI first, but each takes a different approach to running it on the phone.
Qualcomm launched the Snapdragon 8 Elite Gen 6 and the higher-end Snapdragon 8 Elite Extreme Gen 6 at its Snapdragon Summit in Maui, Hawaii, which opened on 22 September, according to The Next Web. Qualcomm says the Extreme chip's processor is the first mobile CPU to reach 5 gigahertz. The company says its CPU is 13 percent faster than the last generation, its graphics processor 44 percent faster, and its neural processing unit (NPU) 35 percent faster.
Apple's A20 Pro, which powers the iPhone 18 Pro phones and the iPhone Duo, has a 32-core neural engine. Apple says it gives twice the AI processing power of the A19 Pro, with 50 percent more memory bandwidth. MediaTek's Dimensity 9600 Pro uses two NPUs. MediaTek says its efficient NPU cuts power use by 40 percent for always-on AI tasks, and the chip supports on-device models of up to 30 billion parameters.
The engineering behind it
A transistor is a switch that controls current flowing through a channel. In older planar transistors, the gate that controls the switch sits only on top of the channel. In FinFET transistors, used for about a decade, the channel stands up like a fin and the gate wraps three sides. In gate-all-around nanosheet transistors, the channel is made of thin flat sheets stacked on top of each other, and the gate surrounds each sheet on all four sides.
Wrapping the gate fully around the channel gives better control over the current. This is general device physics: the better the gate controls the channel, the less current leaks when the switch is off. Less leakage means less wasted power and heat. That matters for a phone, which has a small battery and no fan. The name 2nm is a label for a generation of manufacturing process, not the exact size of any one part.
An NPU is a part of the chip designed for the maths of neural networks, mainly large numbers of multiplications and additions. It does this work with far less energy than a general CPU. Running a large AI model also needs memory bandwidth, which is how fast data moves between memory and the processor. Apple's design places memory beside the chip, out of its heat path, and connects the chip to a vapour chamber to spread heat, the company says.
Qualcomm says the Extreme chip can run mixture-of-experts models with more than 30 billion parameters, and handle up to 32,000 tokens of context. In a mixture-of-experts model, only part of the network runs for each piece of text, which saves memory and power. To fit a model on a phone, developers also use quantisation, which stores each number with fewer bits. These are general techniques and not specific to these chips.
Demand and early tests
TrendForce, citing Economic Daily News, says Apple, Nvidia, AMD, Qualcomm and MediaTek have raised their bookings for TSMC's 2nm production by 10 to 20 percent. Industry sources expect TSMC's 2nm capacity to reach about 120,000 wafers a month by the end of 2026, above earlier estimates of 90,000 to 100,000. A wafer is the thin round slice of silicon on which many chips are made at once.
The Next Web reports that a Gizmodo reviewer tested Qualcomm's reference phones at the summit. They beat Apple's A20 Pro by 8 percent on a multi-core benchmark and by about 20 percent on a graphics test. Apple kept a lead of about 10 percent on single-core scores, and the test phones ran hot. The Next Web also notes that Counterpoint Research expects smartphone shipments to fall 14 percent in 2026, as memory shortages push up prices.
What it means in Nepal
The sources do not discuss Nepal, so this section is about skills rather than the local market. Running AI on the phone instead of on a remote server changes how apps are built. Data can stay on the device, and the app can work without a fast internet connection. But the developer must fit the model into limited memory, power and heat. Skills such as model compression, quantisation and measuring energy use become part of app development.
The chips also show why device physics and computer architecture are useful even for engineers who will never work in a chip factory. A phone chip is a set of choices about transistors, cores, memory paths and heat. Engineers who design embedded products, write performance-critical software or choose hardware for a project make better decisions when they understand those trade-offs. Reading launch claims carefully, and waiting for independent tests, is part of that skill.
What to study if this interests you
Electronic Device and Circuits, ENEX 151, in the second semester of BEI and BCT, is where you first study how transistors switch and why leakage matters. Computer Organization and Architecture, ENEX 253 in the fourth semester of BEI and ENCT 303 in the fifth semester of BCT, explains cores, memory bandwidth and processor design. Artificial Intelligence, ENCT 351, in the sixth semester of BCT, covers the neural networks these NPUs run, and the course has a full guide on this site.
Words in this story
- Gate-all-around transistor
- A transistor where the gate surrounds the current channel on every side, giving better control and less leakage.
- NPU
- A neural processing unit, a part of a chip built to run the maths of AI models efficiently.
- Memory bandwidth
- How much data can move between memory and a processor each second.
- Quantisation
- Storing the numbers inside an AI model with fewer bits, so it uses less memory and power.
Where this comes from
- TrendForce, 6 Oct 2026
- The Next Web, 25 Sep 2026
Written in our own words; no sentence is copied from these reports. Researched with AI assistance on 11 October 2026; no member of faculty has reviewed it yet. If you spot a mistake, call 01-5091616 and we will correct it and say so.




