Can I Run LLM Model Locally?

Find out which AI models your machine can actually run. Check GPU compatibility, VRAM requirements, and expected performance.

Model Rankings

Top models ranked by compatibility for 24 GB VRAM

189 models · 94 excellent · 15 good

LLM models ranked by compatibility and performance
ModelVRAMGrade
Qwen3.8 27B27.8B
Q4_K_M· tok/s·262K ctx·RUNS GREAT
17.4 GBS100
Qwen3.6 27B27.8B
Q4_K_M· tok/s·262K ctx·RUNS GREAT
17.4 GBS100
Gemma 3 27B IT27.4B
Q4_K_M· tok/s·131K ctx·RUNS GREAT
18.1 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
16.1 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.7 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.7 GBS100
Q4_K_M· tok/s·8K ctx·RUNS GREAT
18.0 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
16.1 GBS100
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.7 GBS100
Hy MT2 30B A3B30.1B
Q4_K_M· tok/s·262K ctx·RUNS GREAT
18.4 GBS100

Browse by VRAM

Find the best models for your VRAM tier

Popular Devices

All Hardware →

GPUs, MacBooks, AI boxes, and more — find what runs AI best

Popular Models

View all →

Qwen3 Coder 30B A3B Instruct

Alibaba · 30.5B · runs from 8.8 GB

745.9K 1.2K

Qwen3 Coder 30B A3B Instruct is a code-specialized Mixture of Experts (MoE) model from Alibaba Cloud's Qwen 3 Coder series, with 30 billion total parameters and approximately 3 billion active parameters per forward pass. The MoE architecture allows it to deliver strong coding performance while keeping per-token compute costs low, making it faster at inference than comparably capable dense models. The model is instruction-tuned for programming assistance, code generation, debugging, and software engineering conversation. It requires VRAM proportional to its total 30B parameter count for loading weights, but benefits from efficient inference throughput due to its low active parameter count. Released under the Apache 2.0 license.

ChatCode

Qwen3.6 35B A3B

Alibaba · 36.0B · runs from 10.3 GB

4.4M 2.8K

Qwen3.6 35B A3B is a 36.0B-parameter open language model from Alibaba in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Qwen3.8 27B

Alibaba · 27.8B · runs from 12.6 GB

6.0M 14.0K

Qwen3.8 27B is a 27.8B-parameter open language model from Alibaba in the Qwen 3.8 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Qwen2.5 7B Instruct

Alibaba · 7.6B · runs from 2.7 GB

11.4M 1.6K

Qwen2.5 7B Instruct is a 7.6-billion parameter instruction-tuned model from Alibaba Cloud's Qwen 2.5 series. It supports a 128K token context window and is fine-tuned for conversational AI, instruction following, and general assistant tasks. Its efficient size makes it well-suited for local deployment on consumer GPUs with 8GB or more of VRAM. The model delivers strong performance for its parameter class across reasoning, multilingual understanding, and coding tasks. It benefits from the improved pretraining data and techniques of the Qwen 2.5 generation. Released under the Apache 2.0 license and widely supported by inference frameworks such as llama.cpp, vLLM, and Ollama.

Chat

Qwen3.6 27B

Alibaba · 27.8B · runs from 8.4 GB

5.3M 2.3K

Qwen3.6 27B is a 27.8B-parameter open language model from Alibaba in the Qwen 3.6 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

Gemma 4 26B A4B IT

Google · 25.8B · runs from 7.7 GB

8.3M 1.5K

Gemma 4 26B A4B IT is a 25.8B-parameter open language model from Google in the Gemma 4 family. It supports a context window of up to 262,144 tokens. See its VRAM requirements by quantization and which GPUs and Macs can run it locally below.

Vision

How It Works

Three steps to find your perfect local AI setup

1

Select Your Hardware

Pick your GPU or Apple Silicon device from the dropdown.

2

Check Compatibility

See which models fit in your VRAM with performance grades.

3

Run It

Install via Ollama, LM Studio, or download the GGUF file.