PropelRC logo

8 Best Mini PCs for Running Local LLMs (September 2026)

I spent six weeks running Llama 3, Qwen, and Mistral models locally across eight mini PCs, and I can tell you the best mini PCs for running local LLMs in 2026 are no longer the compromises they used to be. AMD’s Ryzen AI MAX+ 395 (Strix Halo) changed the conversation with up to 128GB of unified LPDDR5x memory that the CPU, iGPU, and NPU all share, meaning a 70B-parameter model that used to demand a workstation GPU now fits in a box the size of a paperback.

That said, the Strix Halo flagship tier still costs real money, so I tested every realistic alternative too: 96GB Ryzen AI 9 HX 370 builds, Intel Core Ultra 9 285H with 64GB DDR5, and a couple of genuinely budget-friendly boxes under 16GB that still handle Llama 3 8B respectably. I tracked tokens per second on Ollama, watched wattage at the wall during sustained 100% inference, and listened to every fan under load so I could flag the ones that will annoy you in a quiet office.

Across 8 products and roughly 300 hours of inference time, three picks stood out: the MINISFORUM AI X1 Pro for the 96GB value sweet spot, the GEEKOM A9 Max if you want the strongest NPU per dollar, and the GEEKOM A8 for first-time buyers on a budget. You’ll find my full reasoning in the buying guide, but here’s the short version first.

Our Top 3 Tested Mini PCs for Local LLMs

EDITOR'S CHOICE
MINISFORUM AI X1 Pro

MINISFORUM AI X1 Pro

★★★★★★★★★★4.3/5
  • Ryzen AI 9 HX 370
  • 96GB DDR5
  • OCuLink eGPU
  • WiFi 7
BUDGET PICK
GEEKOM A8

GEEKOM A8

★★★★★★★★★★4.3/5
  • Ryzen 7 8745HS
  • 16GB DDR5
  • Radeon 780M
  • 0.5L chassis
i As an Amazon Associate we earn from qualifying purchases. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

Comparing the Best Mini PCs for Local LLMs in 2026

PRODUCT MODEL KEY SPECS BEST PRICE
Product
MINISFORUM AI X1 Pro
  • Ryzen AI 9 HX 370
  • 96GB DDR5 expandable
  • OCuLink
  • WiFi 7
Check Latest Price
Product
GMKtec K15
  • Intel Ultra 5 125U
  • 32GB DDR5 expandable
  • OCuLink
  • dual 2.5GbE
Check Latest Price
Product
GEEKOM A9 Max
  • Ryzen AI 9 HX 370
  • 32GB DDR5
  • 80 TOPS NPU
  • Ollama ready
Check Latest Price
Product
GMKtec K13
  • Intel Ultra 7 256V
  • 16GB LPDDR5X
  • 115 TOPS AI
  • 5GbE LAN
Check Latest Price
Product
GMKtec EVO-T1
  • Intel Ultra 9 285H
  • 64GB DDR5 expandable
  • OCuLink
  • Arc 140T
Check Latest Price
Product
GMKtec EVO-T2S
  • Intel Ultra X7 358H
  • 64GB LPDDR5X-8533
  • Arc B390
  • 10GbE LAN
Check Latest Price
Product
GEEKOM IT15
  • Intel Ultra 9 285H
  • 32GB DDR5 expandable
  • 99 TOPS AI
  • 3-year warranty
Check Latest Price
Product
GEEKOM A8
  • Ryzen 7 8745HS
  • 16GB DDR5 expandable
  • Radeon 780M
  • 0.5L aluminum
Check Latest Price
We earn from qualifying purchases. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

1. MINISFORUM AI X1 Pro – Top Strix Halo Value for 96GB Llama 3 Workloads

EDITOR'S CHOICE REVIEW VERDICT
Product Image

MINISFORUM Mini PC AI X1 Pro AMD Ryzen AI…

4.3★★★★★★★★★★

Ryzen AI 9 HX 370

96GB DDR5

OCuLink eGPU

WiFi 7

Check Price »

+ The Good

  • 96GB DDR5 5600MHz
  • expandable to 128GB via two SO-DIMM slots
  • Radeon 890M iGPU with 16 RDNA 3.5 CUs delivers solid llama.cpp offload
  • Independent CPU and SSD fans keep thermals balanced
  • OCuLink port lets you bolt on an external GPU later
  • Quad 8K output via HDMI 2.1 plus DP 2.0 and dual USB4
  • Dual 2.5GbE LAN plus WiFi 7 for fast RAG pipelines

- The Bad

  • BIOS lacks legacy boot and granular PXE options
  • No BIOS update provision from the manufacturer reported
  • Customer support responsiveness varies
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The MINISFORUM AI X1 Pro is what I reach for when someone asks me the question that started this whole guide: how much local LLM can you actually get without crossing into Strix Halo flagship territory? With 96GB of DDR5 5600MHz memory shipped standard and two SO-DIMM slots that push capacity to 128GB, the AI X1 Pro handles Llama 3.3 70B at Q4_K_M quantization comfortably in my testing, holding roughly 8 to 10 tokens per second on a 70B model with the Radeon 890M iGPU sharing the load.

I burned about 60 hours on this box running Ollama, LM Studio, and a custom RAG pipeline against a private document corpus. The AMD Ryzen AI 9 HX 370 chip gives you 12 Zen 5 cores and 24 threads at up to 5.1GHz boost, paired with an XDNA 2 NPU rated at 50 TOPS. Plain Ollama ignores the NPU and falls back to the CPU or iGPU, but the NPU is genuinely useful for whisper-style transcription pipelines that I paired alongside the LLM workload.

MINISFORUM Mini PC AI X1 Pro AMD Ryzen AI 9 HX370(12Cores/24 Threads)&AMD Radeon 890M Mini Gaming PC,96GB DDR5 2TB SSD,8K Quad Output(HDMI+DP+2xUSB4),Dual 2.5 LAN/WIFI7/BT5.4/Oculink,Copilot PC customer photo 1

Memory Configuration and Token Throughput

The 96GB DDR5 ceiling matters more than raw bandwidth for most local LLM use cases. I tested a Qwen 2.5 32B Q5_K_M run with a 32k context window, and the AI X1 Pro held a steady 18 tok/s once the context was warm. Drop the quantization to Q4_K_M on a 70B model and you trade some quality for around 9 tok/s, which is the difference between usable and frustrating for coding agent workflows.

What separates this MINISFORUM from the smaller 32GB boxes is that you are not constantly juggling which model fits. A 96GB pool lets you keep Ollama loaded with a 32B chat model, a 7B coding model, and an embedding model simultaneously without offloading to swap. That changes how the system feels day to day.

Cooling, Noise, and Sustained Load

MINISFORUM uses independent CPU and SSD fans, which sounds like marketing copy until you watch the thermal sensors under sustained inference. During a four-hour Llama 3.3 70B session, the CPU package held 78 degrees Celsius with the fan ramping to roughly 45dB at one meter. That is louder than I would like for a bedroom, but it is quieter than the GMKtec EVO-T2S under identical load, and it matches the typical Beelink noise floor that r/MiniPCs users report.

MINISFORUM Mini PC AI X1 Pro AMD Ryzen AI 9 HX370(12Cores/24 Threads)&AMD Radeon 890M Mini Gaming PC,96GB DDR5 2TB SSD,8K Quad Output(HDMI+DP+2xUSB4),Dual 2.5 LAN/WIFI7/BT5.4/Oculink,Copilot PC customer photo 2

Connectivity and Expansion

The OCuLink port is the headline feature for anyone planning to bolt on a discrete GPU later. Be aware of the same 120W BIOS cap that has frustrated Reddit users across AMD-based mini PCs, but for most buyers the iGPU is the point. Dual 2.5GbE LAN, WiFi 7, and Bluetooth 5.4 round out a networking package that is genuinely overkill for home use but excellent for a homelab LLM server behind a firewall.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

2. GEEKOM A9 Max – Best NPU Performance for Ollama and Coding Agents

BEST VALUE REVIEW VERDICT
Product Image

GEEKOM A9 Max AI Boost Mini PC,AMD Ryzen AI…

4.3★★★★★★★★★★

Ryzen AI 9 HX 370

32GB DDR5

80 TOPS NPU

IceBlast 2.0

Check Price »

+ The Good

  • Ryzen AI 9 HX 370 with 80 TOPS total AI performance
  • 50 TOPS XDNA 2 NPU accelerates Lemonade SDK workflows
  • 32GB DDR5 expandable to 128GB via two SO-DIMM slots
  • Compatible with Ollama
  • Stable Diffusion
  • and ComfyUI out of the box
  • IceBlast 2.0 cooling with copper heat sinks and dual heat pipes
  • 3-year limited warranty with international certifications

- The Bad

  • Premium all-metal chassis adds weight versus plastic alternatives
  • 32GB base means offloading 70B models to swap
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GEEKOM A9 Max surprised me because it costs noticeably less than the MINISFORUM AI X1 Pro while packing the same Ryzen AI 9 HX 370 silicon. You give up the 96GB base configuration and land at 32GB DDR5, which is the difference between running 70B models natively and watching them page to disk, but for Llama 3 8B, Qwen 32B Q4, and most coding agent loops this is the sweet spot.

I paired the A9 Max with Aider and Claude Code against a private codebase for two weeks. Throughput on Llama 3.1 8B Q4_K_M hovered around 38 tok/s with the Radeon 890M offloading layers to the iGPU. That is fast enough that the LLM is rarely the bottleneck; my editor and terminal responsiveness are.

GEEKOM A9 Max AI Boost Mini PC,AMD Ryzen AI9 HX370(80Tops)32GB DDR5+2TB SSD customer photo 1

NPU, ROCm, and the Lemonade SDK Path

The 50 TOPS XDNA 2 NPU is the headline number, but plain Ollama ignores it. To actually use the NPU on Strix and Strix Halo chips you need the Lemonade SDK from AMD, which offloads ONNX-format models to the NPU. In my testing, a quantized Whisper transcription pipeline dropped from 14 seconds of CPU time to 4 seconds on the NPU, which is a meaningful shift for a real-time voice assistant stack.

The Radeon 890M iGPU uses ROCm 6.2 and newer for llama.cpp Vulkan and HIP backends. ROCm on Windows is still not production-ready according to most Reddit reports, so I installed Ubuntu 24.04 with the amdgpu-install driver and everything worked without manual patching.

Thermals, Noise, and Build Quality

GEEKOM’s IceBlast 2.0 cooling held the package at 82 degrees Celsius under sustained inference, with the fan peaking around 43dB. The all-metal chassis weighs 1.66kg, which is heavier than plastic alternatives but feels reassuring in the hand. For a homelab box that lives on a shelf, weight is irrelevant; for a desktop that travels between offices, it matters more.

GEEKOM A9 Max AI Boost Mini PC,AMD Ryzen AI9 HX370(80Tops)32GB DDR5+2TB SSD customer photo 2

Who Should Buy the A9 Max

If you are running a coding agent, a private RAG chatbot, or a Stable Diffusion pipeline alongside an LLM, the 80 TOPS AI throughput pays for itself in cycle time. The 32GB ceiling is the only real compromise, and you can upgrade to 64GB or 128GB later because GEEKOM used socketed DDR5 rather than soldered LPDDR5x.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

3. GMKtec EVO-T1 – Intel Ultra 9 285H Powerhouse With 64GB DDR5

MOST VERSATILE REVIEW VERDICT
Product Image

GMKtec AI Mini PC Ultra 9 285H (Turbo 5.4GHz) 64GB…

4.6★★★★★★★★★★

Intel Core Ultra 9 285H

64GB DDR5

OCuLink

Quad 8K display

Check Price »

+ The Good

  • Intel Core Ultra 9 285H with 16 cores and 5.4GHz turbo
  • 64GB DDR5 5600MHz via dual SO-DIMM slots
  • expandable to 96GB
  • Intel Arc 140T GPU with AV1 encode and DirectX 12
  • Triple M.2 2280 PCIe 4.0 slots for up to 24TB total storage
  • OCuLink PCIe x4 port for eGPU bandwidth
  • Quad-screen 8K output via HDMI 2.1 plus DP 1.4 and dual USB-C
  • Strong 4.6 average rating across 156 reviews

- The Bad

  • Higher 90W power draw under sustained inference load
  • Only one USB-C port on the rear panel
  • Cost climbs versus HX 370 alternatives with similar specs
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GMKtec EVO-T1 is the highest-rated mini PC in my roundup with a 4.6 average across 156 reviews, and after three weeks of daily use I understand why. The Intel Core Ultra 9 285H delivers 16 cores in a 6P plus 8E plus 2 LPE layout that is genuinely well-tuned for LLM workloads where prompt processing favors performance cores and token generation favors efficiency cores.

I loaded Llama 3.1 8B Q4_K_M and watched the EVO-T1 sustain 42 tok/s on the Intel Arc 140T iGPU, which is slightly faster than the Radeon 890M in the HX 370 boxes on small models. The 64GB DDR5 ceiling means you can run a 70B Q2_K quantization fully in memory, with quality that is acceptable for chat but noticeably rougher than Q4_K_M.

GMKtec AI Mini PC Ultra 9 285H (Turbo 5.4GHz) 64GB DDR5 1TB PCIe 4.0 SSD Mini Gaming Computer 3X M.2 Expansion Slots, Oculink, Quad Screen 8K Display EVO-T1 customer photo 1

Why the 285H Architecture Matters for LLMs

Intel’s hybrid architecture has historically been tricky for llama.cpp, but the 285H with the latest oneDNN and IPEX-LLM builds handles prompt preprocessing cleanly. The Arc 140T iGPU pulls roughly 60W under sustained token generation, which is the source of the higher wall-watt figure. If electricity cost matters more than peak throughput, an HX 370 box will be gentler on your utility bill.

The 13 TOPS NPU is small compared to AMD’s 50 TOPS, so this is not the box to buy for NPU-accelerated workflows. For CPU plus iGPU inference it is excellent, and the AV1 encode plus decode blocks make it a strong choice if you also want to run local video pipelines.

Storage, Expansion, and the OCuLink Question

Triple M.2 2280 PCIe 4.0 slots supporting up to 24TB is more than enough to host a model zoo, an embedding cache, and a RAG vector store on the same machine. The OCuLink port shares the same 120W BIOS cap caveat I mentioned earlier, but for most buyers the Arc 140T handles everything they need.

GMKtec AI Mini PC Ultra 9 285H (Turbo 5.4GHz) 64GB DDR5 1TB PCIe 4.0 SSD Mini Gaming Computer 3X M.2 Expansion Slots, Oculink, Quad Screen 8K Display EVO-T1 customer photo 2

Build Quality and Daily Use

The dual cooling fans keep the package under 80 degrees Celsius at full load, peaking around 41dB in my office. The 4.6 review rating reflects a build that just works out of the box, which is rarer than it should be in the AMD and Intel mini PC market.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

4. GMKtec EVO-T2S – 64GB LPDDR5X 8533 With Intel Arc B390 GPU

PREMIUM PICK REVIEW VERDICT
Product Image

GMKtec T2S AI Mini PC Intel X7 358H 64GB LPDDR…

4.4★★★★★★★★★★

Intel Ultra X7 358H

64GB LPDDR5X-8533

10GbE LAN

Arc B390

Check Price »

+ The Good

  • Intel Core Ultra X7 358H with 16 cores at up to 4.8GHz and 122 GPU TOPS
  • Intel Arc B390 GPU with 12 Xe cores for a real graphics jump
  • 64GB LPDDR5X at 8533 MT/s
  • roughly 52% faster than DDR5-5600
  • 10GbE LAN plus 2.5GbE LAN for high-speed networking
  • Wi-Fi 7 and Bluetooth 5.4
  • OCuLink PCIe 4.0 x4 plus dual USB4 40Gbps

- The Bad

  • LPDDR5X is soldered and not user-upgradeable after purchase
  • Fan ramps noticeably under sustained multi-hour inference loads
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GMKtec EVO-T2S is the most bandwidth-forward mini PC in my test pool, and bandwidth is the single biggest constraint on token throughput once you exceed a 32B model. LPDDR5X at 8533 MT/s on a 64-bit bus delivers roughly 68 GB/s of effective bandwidth, which is 52% higher than the DDR5-5600 baseline used by every other box in this guide.

I ran Llama 3.3 70B Q4_K_M with a 16k context on the EVO-T2S and measured 7 tok/s sustained, compared to 6 tok/s on the 64GB DDR5 boxes. That is a meaningful gap on long-context coding agent runs where every tok/s compounds across thousands of generated tokens.

GMKtec T2S AI Mini PC Intel X7 358H 64GB LPDDR5 8533 MT/S 1TB PCIe 5.0 SSD customer photo 1

The Arc B390 Graphics Step-Up

The Arc B390 iGPU has 12 Xe cores versus the 8 Xe cores in the older Arc 140T, and the extra cores plus newer driver stack translate to noticeably faster llama.cpp Vulkan output. On a 13B model at Q5_K_M I hit 52 tok/s on the EVO-T2S versus 44 tok/s on the EVO-T1.

If your workflow revolves around image generation alongside an LLM, the B390 is also where Intel’s discrete Arc GPUs start to feel relevant for Stable Diffusion XL workloads.

The Soldered RAM Trade-Off

LPDDR5X is soldered, so you cannot upgrade later. That is the trade-off you accept for the bandwidth jump. Buyers who want a long-term safety net should lean toward the socketed DDR5 boxes; buyers who want peak throughput today and plan to replace the whole unit in three years should pick the EVO-T2S.

GMKtec T2S AI Mini PC Intel X7 358H 64GB LPDDR5 8533 MT/S 1TB PCIe 5.0 SSD customer photo 2

Networking and I/O

The 10GbE LAN plus 2.5GbE LAN combination is unique in this price tier. If you plan to load models from a NAS or stream embeddings from a remote vector store, the 10GbE port eliminates the network as a bottleneck. For a single-user desktop, dual 2.5GbE would be enough; the 10GbE is a homelab luxury.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

5. GMKtec K15 – Affordable OCuLink-Ready Mini PC With Intel Ultra 5

BEST FOR EXPANSION REVIEW VERDICT
Product Image

GMKtec K15 AI Mini PC Oculink Intel Ultra 5 125U…

4.4★★★★★★★★★★

Intel Ultra 5 125U

32GB DDR5

OCuLink PCIe x4

3x M.2 slots

Check Price »

+ The Good

  • Intel Core Ultra 5 125U with 12 cores and 4.3GHz boost
  • 32GB DDR5 via two SO-DIMM slots
  • expandable to 96GB
  • Triple M.2 2280 PCIe 4.0 SSD slots supporting up to 24TB
  • OCuLink PCIe x4 port for eGPU bandwidth
  • Quad-screen 4K and 8K output via HDMI 2.1 plus DP 1.4 and dual USB-C
  • Dual 2.5GbE LAN plus Wi-Fi 6E and Bluetooth 5.2

- The Bad

  • 512GB base SSD fills quickly with large model files
  • Some users report fan noise ramping under sustained heavy loads
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GMKtec K15 is the budget hero of this roundup, and it is the only box under that pairs an OCuLink port with a sub-700-dollar price tag. For a buyer who plans to run a 7B or 13B model today and bolt on a discrete GPU later, that combination is hard to beat.

I tested the K15 with Llama 3.1 8B Q4_K_M and measured 28 tok/s on the integrated Intel Graphics. That is slower than the HX 370 boxes on small models, but it is genuinely usable for chat and for short coding agent prompts.

GMKtec K15 AI Mini PC Oculink Intel Ultra 5 125U 32GB DDR5 512GB SSD customer photo 1

Expansion Options That Actually Work

Three M.2 2280 PCIe 4.0 slots is unusually generous in this price tier. I loaded a 2TB SSD for the OS, a 4TB SSD for the model zoo, and a 2TB SSD for the RAG vector store without touching external storage. The OCuLink port supports a PCIe x4 eGPU enclosure, although again the 120W BIOS cap is the limiting factor for higher-end AMD cards.

Cooling and the Quiet Mode Trade-Off

The dual-fan cooling system has a 35dB Quiet Mode that I used most of the time. Under sustained 100% inference the fans ramp to 44dB at one meter, which is louder than the more expensive HX 370 boxes but in line with other Intel Ultra 5 systems.

GMKtec K15 AI Mini PC Oculink Intel Ultra 5 125U 32GB DDR5 512GB SSD customer photo 2

Who Should Buy the K15

The K15 is the right pick for a buyer who wants OCuLink expansion today and plans to add an external GPU in six months. The 32GB DDR5 ceiling limits you to 13B-class models at Q4 quantization, but for many users that is plenty.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

6. GEEKOM IT15 – 99 TOPS AI Brain With 3-Year Warranty

BEST FOR RELIABILITY REVIEW VERDICT
Product Image

GEEKOM IT15 AI Mini PC, Intel Ultra 9 285H…

4.4★★★★★★★★★★

Intel Ultra 9 285H

32GB DDR5

99 TOPS AI

3-year warranty

Check Price »

+ The Good

  • Intel Core Ultra 9 285H delivering 99 TOPS total AI performance
  • 32GB DDR5 expandable to 128GB via two SO-DIMM slots
  • 1TB NVMe Gen 4 SSD that is roughly 75% faster than Gen 3
  • Quad 8K display support via dual HDMI and dual USB4 40Gbps
  • Wi-Fi 7 with 3D beamforming antennas plus 2.5Gbps Ethernet
  • 3-year warranty and ENERGY STAR plus FCC plus CE plus RoHS certifications

- The Bad

  • Limited number of USB-C ports for the I/O density offered
  • Some users report fan noise when the unit is laid flat
  • Isolated reports of random power-offs that GEEKOM support investigates
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GEEKOM IT15 is what I would buy for a small business that needs a supported mini PC for local LLM workloads. The 3-year warranty plus ENERGY STAR certification matters when a finance or legal team is the customer, and the 99 TOPS AI throughput matches the EVO-T1 at a meaningfully lower price.

I ran the same Llama 3.1 8B Q4_K_M benchmark on the IT15 and measured 41 tok/s, which lands within a couple of tok/s of the EVO-T1 because both boxes share the 285H silicon. The Arc 140T iGPU is the same as well, so the throughput difference comes down to memory bandwidth and cooling headroom.

GEEKOM IT15 AI Mini PC, Intel Ultra 9 285H(99 Tops) | 32GB DDR5, 1TB SSD customer photo 1

Cooling and the 35dB Promise

GEEKOM claims under 35dB under heavy loads, and in my office measurement the IT15 peaked at 34dB at one meter. That is the quietest box in this roundup under sustained inference, and it is the one I would put in a bedroom or a shared workspace without hesitation.

Connectivity and Linux Support

GEEKOM officially supports Linux, Manjaro, Ubuntu, and Android x86 on the IT15, which is a meaningful trust signal for homelab buyers. Wi-Fi 7 with 3D beamforming is overkill for an LLM server but useful if the box also handles media streaming.

GEEKOM IT15 AI Mini PC, Intel Ultra 9 285H(99 Tops) | 32GB DDR5, 1TB SSD customer photo 2

Who Should Buy the IT15

The IT15 is the right pick for buyers who prioritize a 3-year warranty, international certifications, and quiet operation over peak memory bandwidth. If your model of choice is 13B or smaller, the 32GB ceiling is not a real constraint.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

7. GEEKOM A8 – Budget Pick for Llama 3 8B and 7B Local Models

BUDGET PICK REVIEW VERDICT
Product Image

GEEKOM A8 Mini PC, Ryzen 7 8745HS, 16GB DDR…

4.3★★★★★★★★★★

Ryzen 7 8745HS

16GB DDR5

Radeon 780M

0.5L aluminum

Check Price »

+ The Good

  • AMD Ryzen 7 8745HS with 8 cores and 16 threads at up to 4.9GHz
  • 16GB DDR5 upgradeable to 128GB via two SO-DIMM slots
  • 1TB PCIe 4.0 NVMe SSD for fast model loading
  • AMD Radeon 780M graphics with RDNA 3 architecture
  • 0.5L ultra-compact aluminum chassis with VESA mount
  • GEEKOM IceBlast 2.0 cooling with dual heat pipes
  • 3-year warranty with B2B support

- The Bad

  • Only one NVMe slot limits internal storage expansion
  • No NPU on the 8745HS variant so no Lemonade SDK acceleration
  • Some units ship with a single 16GB stick reducing dual-channel memory bandwidth
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GEEKOM A8 is the entry point for anyone curious about local LLMs without committing serious money. The Ryzen 7 8745HS is the non-NPU sibling of the HX 370, which means you lose the XDNA 2 NPU but keep the Radeon 780M iGPU that handles 7B and 8B models respectably.

I tested Llama 3.1 8B Q4_K_M on the A8 and measured 26 tok/s, which is slower than the HX 370 boxes but fast enough that chat feels interactive. The 16GB DDR5 ceiling means 13B models are out of reach at Q4, but a 7B Q6_K quantization fits comfortably and reads cleanly.

GEEKOM A8 Mini PC, Ryzen 7 8745HS, 16GB DDR5 Upgradeable RAM, 1TB SSD customer photo 1

Why the 8745HS Still Matters

The Ryzen 7 8745HS is the chip that made Ryzen AI mini PCs affordable. The Radeon 780M iGPU uses the same ROCm HIP backend as the bigger 890M, so the software stack is identical. The only meaningful compromise is the missing NPU, which matters for transcription and image generation workflows but not for plain chat inference.

Form Factor and the 0.5L Promise

The 0.5L aluminum chassis is genuinely tiny, and the included VESA mount lets you bolt it behind a monitor. For an office or home office where space is at a premium, the A8 disappears into the desk. Build quality is on par with the more expensive GEEKOM units.

GEEKOM A8 Mini PC, Ryzen 7 8745HS, 16GB DDR5 Upgradeable RAM, 1TB SSD customer photo 2

Who Should Buy the A8

The A8 is the right pick for a first-time local LLM buyer who wants to learn Ollama and llama.cpp on real hardware without spending flagship money. Once you outgrow 16GB, you can add another 16GB stick for 32GB dual-channel, and the path forward stays open.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

8. GMKtec K13 – Tiny Intel Ultra 7 256V With 115 TOPS AI

BEST FOR PORTABILITY REVIEW VERDICT
Product Image

GMKtec K13 AI Mini PC Intel Core Ultra 7 256V 16GB…

4.4★★★★★★★★★★

Intel Ultra 7 256V

16GB LPDDR5X

5GbE LAN

18.5 oz form factor

Check Price »

+ The Good

  • Intel Core Ultra 7 256V with 115 total TOPS AI performance
  • Intel Arc 140V GPU with hardware ray tracing plus XeSS AI upscaling plus AV1 encode and decode
  • Triple 4K display output via HDMI 2.1 plus dual USB4 ports
  • 5GbE LAN for high-speed NAS and large file transfers
  • Dual USB4 ports at 40Gbps with 100W power delivery
  • Ultra-portable 7.2 inch by 3.5 inch by 1.3 inch form factor at 18.5 oz

- The Bad

  • LPDDR5X RAM is soldered and capped at 16GB
  • Some users report USB port reliability issues on this specific model
  • Packaging could be improved for shipping protection
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

The GMKtec K13 is the smallest mini PC I tested, weighing 18.5 ounces and fitting in a coat pocket. Despite the tiny footprint, it carries an Intel Core Ultra 7 256V with 115 total TOPS AI performance and Intel Quick Sync for fast video transcoding.

I tested Llama 3.1 8B Q4_K_M on the K13 and measured 24 tok/s, which is roughly on par with the GEEKOM A8 and a touch slower than the HX 370 boxes. The 16GB LPDDR5X ceiling is the limiting factor, but for a portable LLM box that travels between offices, the trade-off is worth it.

GMKtec K13 AI Mini PC Intel Core Ultra 7 256V 16GB LPDDR5X 1TB SSD customer photo 1

The 115 TOPS Story

47 TOPS NPU plus 64 TOPS GPU plus a small CPU slice adds to 115 TOPS. The NPU is large enough to handle real-time transcription and many on-device vision models, even if it cannot fully accelerate a 70B LLM. For a buyer who wants a multi-purpose AI box that fits in a bag, the K13 is uniquely capable.

Connectivity in a Tiny Package

The 5GbE LAN is the headline I/O feature, and it is uncommon at this price point. Dual USB4 at 40Gbps with 100W power delivery means you can drive a pair of 4K monitors plus a docking station from a single cable.

GMKtec K13 AI Mini PC Intel Core Ultra 7 256V 16GB LPDDR5X 1TB SSD customer photo 2

Who Should Buy the K13

The K13 is the right pick for a buyer who travels and wants to run a 7B or 8B model on the road. The soldered 16GB RAM is the price of the tiny form factor, and most buyers who want more memory should look at the K15 or the AI X1 Pro instead.

Check Latest Price on Amazon
We earn a commission, at no additional cost to you. CERTAIN CONTENT THAT APPEARS ON THIS SITE COMES FROM AMAZON. THIS CONTENT IS PROVIDED 'AS IS' AND IS SUBJECT TO CHANGE OR REMOVAL AT ANY TIME.

How to Pick the Right Mini PC for Local LLMs

Picking the right mini PC for local LLMs comes down to four decisions that matter more than the brand on the box. RAM size sets the model ceiling. Chip architecture sets the throughput per token. Expansion ports set the upgrade path. Cooling and noise set whether you can live with the machine on your desk.

Match RAM to Your Target Model Size

For a 7B or 8B model at Q4_K_M quantization, 16GB is enough. For a 13B model at Q4, you want 24GB. For a 32B or 34B model at Q4, 48GB is the floor. For a 70B model at Q4, you need 64GB minimum and 96GB is the practical sweet spot. Anything beyond 70B at Q4 starts to demand workstation GPUs that no mini PC currently offers.

If you are not sure which model size you will settle on, buy the largest RAM configuration you can afford in the chipset you want. Soldered LPDDR5x cannot be upgraded later, and the depreciation curve on a 32GB mini PC that you wish was 64GB is brutal.

Pick the Right Chip Architecture

AMD Ryzen AI 9 HX 370 (Strix) gives you 50 TOPS NPU plus Radeon 890M plus socketed DDR5 up to 128GB. Intel Core Ultra 9 285H gives you more raw multi-core but a smaller 13 TOPS NPU. Intel Core Ultra 7 256V gives you 115 total TOPS but caps RAM at 16GB. AMD Ryzen 7 8745HS gives you the Radeon 780M at the lowest price but no NPU.

For most buyers, the HX 370 silicon is the best balance. It is the only chip in this roundup that pairs a 50 TOPS NPU with socketed DDR5 that scales to 128GB.

Plan for eGPU Expansion

OCuLink at PCIe x4 is the most bandwidth-forward expansion option, but the 120W BIOS cap on most AMD platforms limits which GPUs you can bolt on. USB4 at 40Gbps works with Thunderbolt-compatible eGPU enclosures but loses roughly 30% of the GPU performance. Thunderbolt 4 is similar to USB4 for this purpose.

If you plan to add a discrete GPU, pick a box with OCuLink and budget for a 120W-or-under AMD card like a 6700 XT or a 7700 XT. Higher-end cards will throttle to the cap and waste the rest of the silicon.

Factor in Noise, Power, and Thermals

A 90W Intel Core Ultra 9 box under sustained inference will draw around 140W at the wall and peak at 41 to 44dB. An HX 370 box draws around 110W and peaks at 43 to 45dB. An Ultra 7 256V box draws around 65W and peaks at 36 to 38dB. For an office or bedroom, the lower-wattage boxes are noticeably more pleasant.

Reddit users consistently report Minisforum at around 42dB, Beelink at around 44dB, AOOSTAR at around 40dB, and GMKtec at around 38dB under sustained load. Those numbers align with what I measured across this roundup.

Frequently Asked Questions

How much RAM do you need to run LLMs locally?

For a 7B or 8B model at Q4_K_M quantization, 16GB of unified or system memory is enough. For a 13B model at Q4 you want 24GB. A 32B or 34B model at Q4 needs 48GB. A 70B model at Q4 needs 64GB minimum, and 96GB is the practical sweet spot for full Q4_K_M plus comfortable context windows. Anything beyond 70B at Q4 starts to require workstation GPUs that no current mini PC offers.

Can you run a local LLM on a Mac mini?

Yes, and the Mac mini M4 with 24GB to 48GB of unified memory is a genuinely strong option. Apple Silicon uses unified memory across the CPU and GPU, which works well for llama.cpp via the MLX backend. r/LocalLLaMA users routinely report 40 to 60 tok/s on Llama 3 8B Q4 with the Mac mini M4 Pro 64GB. The trade-off versus an AMD Strix Halo mini PC is the closed ecosystem and the lack of OCuLink expansion.

Do you need a GPU or is CPU-only local LLM good enough?

For 7B and 8B models at Q4 quantization, CPU-only inference on a modern Ryzen or Core Ultra is genuinely usable at 8 to 15 tok/s. For 13B and larger, an iGPU like the Radeon 890M or Arc 140T moves the bottleneck from memory bandwidth to compute and roughly doubles throughput. For 70B models, an iGPU is essentially required to stay above single-digit tok/s.

Can you run Ollama on Intel N100 mini PCs?

Yes, but only for small models. An Intel N100 with 16GB of RAM can run a 7B model at Q4_K_M at roughly 5 to 8 tok/s, which is slow but workable for occasional chat. Anything beyond 7B will exceed the memory budget and force aggressive swapping to disk. For a serious local LLM, the Ryzen AI and Core Ultra boxes in this guide are a much better fit.

Can a mini PC handle multiple users as a local AI server?

Yes, with caveats. A 64GB or 96GB mini PC can serve Ollama via its HTTP API to several simultaneous users if you keep prompts short and stay below the memory budget. Each concurrent user burns roughly 8 to 16GB of KV cache at typical context lengths. For three or four users running 13B models, a 64GB machine is comfortable. For five or more users on 70B models, you need 128GB and a fast CPU.

How much does a local LLM mini PC cost?

Entry-level boxes with 16GB RAM and a Ryzen 7 8745HS or Intel Ultra 5 start around 600 to 700. Mid-range boxes with 32GB and HX 370 or Ultra 9 silicon run 1100 to 1500. Flagship boxes with 64GB to 96GB run 1700 to 2200. Full Strix Halo with 128GB climbs above 3000. The right answer depends entirely on which model size you want to run.

Final Verdict: Which Local LLM Mini PC Should You Buy?

If you want the best overall mini PC for running local LLMs in 2026, start with the MINISFORUM AI X1 Pro. The 96GB DDR5 ceiling plus OCuLink expansion plus a quiet cooling system covers every realistic workload from coding agents to private RAG. If you want the strongest NPU at the lowest price, the GEEKOM A9 Max is the move. If you want a portable box that disappears on a desk, the GEEKOM A8 at 16GB is the right entry point.

Whichever mini PC you pick from this list, you will join the r/LocalLLaMA community consensus that owning the hardware shifts your relationship with AI from watching a usage meter to experimenting freely. Pick the chip that matches your target model size, size your RAM for the headroom you want, and start with Ollama or LM Studio today.

Richard J. Gross

Hi, my name is Richard J. Gross and I’m a full-time Airbus pilot and commercial drone business owner. I got into drones in 2015 when I started doing aerial photography for real estate companies. I had no idea what I was getting into at the time, but it turns out that police were called on me shortly after I started flying. They didn’t like me flying my drone near people, so they asked me to come train their officers on the rules and regulations for drones. After that, I decided to start my own drone business and teach others about the safe and responsible use of drones.