Tools / Local SI

Find the SI model for your machine.

Tell us how much memory you have. Find recent open-weight SI models for chat, code, reasoning, and vision that fit your compute.

38 recent model families · 15 labs · sizes grouped together.

Models that fit your memory

14 model families fit 12 GB VRAM · 4K context · 4-bit

Releases and Hugging Face counts checked 2026-10-04 · refreshed daily. Top rated uses community likes; downloads cover the past month.

Comfortable fitQ4_K_M · GGUF
IBM

Granite 4.2

6.3 GBestimated usage

granite-4.2-8b · selected for your memory

Chat & writingCoding8.8B parameters128K configured context
On Hugging Face 2026-08-07
93 HF likes · 128.4K downloads / month · source model

An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.

5.2 GB weights0.7 GB cache & state0.6 GB runtime
4.5 GB to spareapache-2.0
Other sizes & editions (2)
Comfortable fitQ4_K_M · GGUF
Google

Gemma 4

7.7 GBestimated usage

gemma-4-12B · selected for your memory

Chat & writingCodingReasoningVision11.9B parameters256K configured context
On Hugging Face 2026-05-23
1.6K HF likes · 1.9M downloads / month · source model

A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.

Assumes a runner with sliding-window KV caching. Shared-layer cache savings are not deducted.

6.7 GB weights0.4 GB cache & state0.7 GB runtime
3.1 GB to spareapache-2.0
Other sizes & editions (4)
Comfortable fitQ4_K_M · GGUF
Microsoft

Fara 1.5

6.3 GBestimated usage

Fara1.5-9B · selected for your memory

Chat & writingCodingVision9.0B parameters256K configured context
On Hugging Face 2026-05-12
47 HF likes · 4.6K downloads / month · source model

A model for computer use and screen understanding. Select Vision to include its image encoder.

5.6 GB weights0.2 GB cache & state0.6 GB runtime
4.5 GB to sparemit
Other sizes & editions (2)
Comfortable fitQ4_K_M · GGUF
IBM

Granite 4.1

6.2 GBestimated usage

granite-4.1-8b · selected for your memory

Chat & writingCoding8.8B parameters128K configured context
On Hugging Face 2026-04-06
256 HF likes · 179.3K downloads / month · source model

An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.

5.0 GB weights0.7 GB cache & state0.5 GB runtime
4.6 GB to spareapache-2.0
Other sizes & editions (2)
Comfortable fitQ4_K_M · GGUF
Qwen

Qwen 3.5

6.0 GBestimated usage

Qwen3.5-9B · selected for your memory

Chat & writingCodingReasoningVision9.0B parameters256K configured context
On Hugging Face 2026-02-27
2.1K HF likes · 9M downloads / month · source model

A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.

5.3 GB weights0.2 GB cache & state0.6 GB runtime
4.8 GB to spareapache-2.0
Other sizes & editions (7)
Comfortable fitQ4_K_M · GGUF
Google

Gemma 3

8.1 GBestimated usage

gemma-3-12b · selected for your memory

Chat & writingCodingVision11.8B parameters128K configured context
On Hugging Face 2025-03-01
847 HF likes · 492K downloads / month · source model

A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.

Assumes a runner with sliding-window KV caching. Shared-layer cache savings are not deducted.

6.8 GB weights0.6 GB cache & state0.7 GB runtime
2.7 GB to sparegemma
Other sizes & editions (4)
Comfortable fitUD-Q4_K_M · GGUF
Liquid AI

LFM 2.5

5.6 GBestimated usage

LFM2.5-8B-A1B · selected for your memory

Chat & writingCoding8.5B parameters125K configured context
On Hugging Face 2026-05-28
779 HF likes · 32.4K downloads / month · source model

An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.

5.0 GB weights0.1 GB cache & state0.5 GB runtime
5.2 GB to spareother
Other sizes & editions (4)
Comfortable fitQ4_K_M · GGUF
OpenBMB

MiniCPM o 4.5

5.8 GBestimated usage

MiniCPM-o-4_5

Chat & writingCoding8.2B parameters40K configured context
On Hugging Face 2026-02-03
1.5K HF likes · 735.2K downloads / month · source model

A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.

4.7 GB weights0.6 GB cache & state0.5 GB runtime
5.0 GB to spareapache-2.0
Comfortable fitQ4_K_M · GGUF
OpenBMB

MiniCPM 5

2.2 GBestimated usage

MiniCPM5-2B · selected for your memory

Chat & writingCoding2.5B parameters128K configured context
On Hugging Face 2026-09-06
1.7K HF likes · 1.1M downloads / month · source model

An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.

1.6 GB weights0.2 GB cache & state0.5 GB runtime
8.6 GB to spareapache-2.0
Other sizes & editions (1)
Comfortable fitQ4_K_M · GGUF
Liquid AI

LFM 2.5 Vision

2.2 GBestimated usage

LFM2.5-VL-3B · selected for your memory

Chat & writingCodingVision2.7B parameters32K configured context
On Hugging Face 2026-08-11
214 HF likes · 26.7K downloads / month · source model

A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.

1.6 GB weights0.1 GB cache & state0.5 GB runtime
8.6 GB to spareother
Other sizes & editions (2)
Comfortable fitQ4_K_M · GGUF
NVIDIA

Nemotron 3

3.4 GBestimated usage

Nemotron-3-Nano-4B · selected for your memory

Chat & writingCodingReasoning4.0B parameters256K configured context
On Hugging Face 2026-03-07
126 HF likes · 3.4M downloads / month · source model

An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.

2.8 GB weights0.2 GB cache & state0.5 GB runtime
7.4 GB to spareother
Other sizes & editions (2)
Comfortable fitQ4_K_M · GGUF
Meta

Llama 3.2

2.9 GBestimated usage

Llama-3.2-3B · selected for your memory

Chat & writing3.2B parameters128K configured context
On Hugging Face 2024-09-18
2.7K HF likes · 1.6M downloads / month · source model

An instruction-tuned assistant for conversation, writing, and summarization. Subject to the Llama community license.

1.9 GB weights0.5 GB cache & state0.5 GB runtime
7.9 GB to sparellama3.2
Other sizes & editions (1)
Comfortable fitQ4_K_M · GGUF
OpenBMB

MiniCPM V 4.6

1.1 GBestimated usage

MiniCPM-V-4.6

Chat & writingCodingVision0.8B parameters256K configured context
On Hugging Face 2026-04-13
1.2K HF likes · 287K downloads / month · source model

A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.

0.5 GB weights0.1 GB cache & state0.5 GB runtime
9.7 GB to spareapache-2.0
Tight fitQ4_K_M · GGUF
Microsoft

Phi 4 Reasoning

10.1 GBestimated usage

Phi-4-reasoning-plus

Chat & writingCodingReasoning14.7B parameters32K configured context
On Hugging Face 2025-04-17
348 HF likes · 10K downloads / month · source model

An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.

8.5 GB weights0.8 GB cache & state0.9 GB runtime
0.7 GB to sparemit

What “fits” actually means.

This is a planning estimate for one local conversation, with the entire model in your selected GPUs or Apple unified memory. Vision adds an image encoder and workspace allowance. Speed and actual usage depend on your hardware, runner, and settings.

Download the catalog as JSON ↗
A budget for more than weights

We add the actual GGUF file size, an FP16 conversation cache for your chosen context, and runtime workspace (10% of the weights or 0.5 GB, whichever is larger). We also reserve 10% of dedicated VRAM, at least 1 GB, for the display and other work. Apple Silicon below 96 GB reserves 25% of shared memory, at least 4 GB. At 96 GB and above, it reserves 12.5%, with a 16 GB minimum and 32 GB maximum, for macOS and apps.

“Comfortable fit” leaves another 10% of usable memory or 0.5 GB free. “Tight fit” fits the estimate with less spare room. Values are calculated in GiB (1,024³ bytes), labeled GB to match common hardware specs, and rounded up for display. Actual usage varies by runner and settings.

Context, quantization, and expert models

Choose up to 1M tokens; only models with a verified context limit at least that long qualify. 4-bit weights save memory at a possible quality cost; 8-bit weights use more space. The exact format is shown on each card. For mixture-of-experts models, we count all stored weights.

Attention cache uses 16-bit keys and values, the configured layer count and head dimensions, and your selected context. Sliding-window layers cap their cache at the window size. Hybrid models also include recurrent state at 32-bit precision; latent-attention estimates assume a runner that retains the compressed cache. Models with unverified cache layouts are listed separately without a fit claim. Hugging Face’s cache guide ↗

Vision and image memory

Select Vision to filter for models that understand images. Estimates add the verified image encoder file size and a planning allowance of 1 GB or half the encoder size, whichever is larger, for processing one image. Resolution, crops, and multiple images can need more workspace. Include image tokens in the selected context. Other task settings estimate text use only.

Download both the model and image encoder, and check that your runner supports that model’s vision architecture. These estimates cover one conversation; CPU offloading, video, training, and concurrent users need separate planning. llama.cpp vision setup ↗

Mixing up to four GPU cards

Start with one GeForce card and use “Add another card” to select up to four, including different models or several identical cards. We add their VRAM after reserving 10% or 1 GB per card, whichever is larger. Runtime adds a 0.5 GB allowance per additional card. The vision encoder and its workspace must fit on the largest card.

Estimates distribute model weights and cache by remaining memory and assume sequential layer splitting in one computer. Generation time adds each card’s estimated read time at 60% of its own peak bandwidth. Identical cards increase capacity without multiplying single-token speed. PCIe transfer time, actual layer sizes, and runner support affect results. For connected Macs, choose an Apple chip and use “Add another Mac”. llama.cpp multi-GPU guide ↗

Connecting Macs over Thunderbolt

Choose up to four Macs with separate chip and memory selections. Each Mac reserves memory for its own OS and apps. Four 512 GB Macs provide 2 TB total and 1,920 GB after OS reserves. Distributed runtime workspace is counted separately.

Thunderbolt networking requires software that splits the model; a cable alone does not combine memory. This catalog uses GGUF sizes and sequential layer estimates, which can be used for planning with llama.cpp RPC. EXO uses MLX models; its files, supported architectures, and tensor-parallel performance can differ. Predicted tok/s excludes network overhead and does not model EXO tensor-parallel speedups. llama.cpp distributed inference ↗ · EXO ↗

Four Macs is this matcher’s current limit, not a universal Thunderbolt limit. A fully connected group of four needs six cables and three usable ports per Mac. RDMA needs Thunderbolt 5 Macs and macOS 26.2 or later; older Macs can use IP networking. Other topologies depend on software support. Apple’s connection guide ↗

Predicted tokens per second

The optional speed estimate is for generating one token at a time with a full context at the selected length. We divide 60% of your selected hardware’s published peak memory bandwidth by the estimated active weight bytes plus cache bytes read per token. The 60% factor is a planning assumption, not a measured result for your hardware. Bandwidth uses decimal GB/s.

Dense models read all weights. For expert models, active weights are approximated from publisher active-parameter counts or verified expert dimensions, while the memory-fit estimate still includes every expert. When active weight traffic is unverified, we show a conservative estimate that assumes all weights are read. Models beyond the selected memory budget or with unverified caches explain why tok/s is unavailable. Prompt processing, image encoding, batching, speculative decoding, compute limits, and runner overhead can change actual performance substantially. Hugging Face’s inference guide ↗

Where these recommendations come from

38 recent model families from 15 labs, with 75 size and edition options checked 2026-10-04. We retain releases from the past six months, plus each lab’s latest two or three families, and show each family once, choosing its largest comfortable size, or its largest tight fit. Other sizes stay inside the same card. The default order prioritizes larger comfortable models, with newer releases first among comparable sizes. Duplicate conversions and community fine-tunes are excluded. Daily checks replace duplicate editions without removing recent families just because a newer model appeared. Recommendations pause after seven days without a successful release check.

Download sizes come from the linked Hugging Face repositories; architecture settings come from pinned model configurations. Community GGUF publishers are identified on each card. This is a curated snapshot; dates show when source repositories appeared on Hugging Face. Open the model card for its license and usage instructions, download the precision shown, and set your runner’s context to the length selected here. Get started with LM Studio ↗