Releases and Hugging Face counts checked 2026-10-04 · refreshed daily. Top rated uses community likes; downloads cover the past month.
● Comfortable fitQ4_K_M · GGUF
IBM
Granite 4.2
6.3 GBestimated usage
granite-4.2-8b · selected for your memory
Chat & writingCoding8.8B parameters128K configured context
On Hugging Face 2026-08-07
93 HF likes · 128.4K downloads / month · source model
An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.
5.2 GB weights0.7 GB cache & state0.6 GB runtime
4.5 GB to spareapache-2.0
Other sizes & editions (2)
● Comfortable fitQ4_K_M · GGUF
Google
Gemma 4
7.7 GBestimated usage
gemma-4-12B · selected for your memory
Chat & writingCodingReasoningVision11.9B parameters256K configured context
On Hugging Face 2026-05-23
1.6K HF likes · 1.9M downloads / month · source model
A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.
Assumes a runner with sliding-window KV caching. Shared-layer cache savings are not deducted.
6.7 GB weights0.4 GB cache & state0.7 GB runtime
3.1 GB to spareapache-2.0
Other sizes & editions (4)
● Comfortable fitQ4_K_M · GGUF
Microsoft
Fara 1.5
6.3 GBestimated usage
Fara1.5-9B · selected for your memory
Chat & writingCodingVision9.0B parameters256K configured context
On Hugging Face 2026-05-12
47 HF likes · 4.6K downloads / month · source model
A model for computer use and screen understanding. Select Vision to include its image encoder.
5.6 GB weights0.2 GB cache & state0.6 GB runtime
4.5 GB to sparemit
Other sizes & editions (2)
● Comfortable fitQ4_K_M · GGUF
IBM
Granite 4.1
6.2 GBestimated usage
granite-4.1-8b · selected for your memory
Chat & writingCoding8.8B parameters128K configured context
On Hugging Face 2026-04-06
256 HF likes · 179.3K downloads / month · source model
An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.
5.0 GB weights0.7 GB cache & state0.5 GB runtime
4.6 GB to spareapache-2.0
Other sizes & editions (2)
● Comfortable fitQ4_K_M · GGUF
Qwen
Qwen 3.5
6.0 GBestimated usage
Qwen3.5-9B · selected for your memory
Chat & writingCodingReasoningVision9.0B parameters256K configured context
On Hugging Face 2026-02-27
2.1K HF likes · 9M downloads / month · source model
A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.
5.3 GB weights0.2 GB cache & state0.6 GB runtime
4.8 GB to spareapache-2.0
Other sizes & editions (7)
● Comfortable fitQ4_K_M · GGUF
Google
Gemma 3
8.1 GBestimated usage
gemma-3-12b · selected for your memory
Chat & writingCodingVision11.8B parameters128K configured context
On Hugging Face 2025-03-01
847 HF likes · 492K downloads / month · source model
A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.
Assumes a runner with sliding-window KV caching. Shared-layer cache savings are not deducted.
6.8 GB weights0.6 GB cache & state0.7 GB runtime
2.7 GB to sparegemma
Other sizes & editions (4)
● Comfortable fitUD-Q4_K_M · GGUF
Liquid AI
LFM 2.5
5.6 GBestimated usage
LFM2.5-8B-A1B · selected for your memory
Chat & writingCoding8.5B parameters125K configured context
On Hugging Face 2026-05-28
779 HF likes · 32.4K downloads / month · source model
An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.
5.0 GB weights0.1 GB cache & state0.5 GB runtime
5.2 GB to spareother
Other sizes & editions (4)
● Comfortable fitQ4_K_M · GGUF
OpenBMB
MiniCPM o 4.5
5.8 GBestimated usage
MiniCPM-o-4_5
Chat & writingCoding8.2B parameters40K configured context
On Hugging Face 2026-02-03
1.5K HF likes · 735.2K downloads / month · source model
A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.
4.7 GB weights0.6 GB cache & state0.5 GB runtime
5.0 GB to spareapache-2.0
● Comfortable fitQ4_K_M · GGUF
OpenBMB
MiniCPM 5
2.2 GBestimated usage
MiniCPM5-2B · selected for your memory
Chat & writingCoding2.5B parameters128K configured context
On Hugging Face 2026-09-06
1.7K HF likes · 1.1M downloads / month · source model
An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.
1.6 GB weights0.2 GB cache & state0.5 GB runtime
8.6 GB to spareapache-2.0
Other sizes & editions (1)
● Comfortable fitQ4_K_M · GGUF
Liquid AI
LFM 2.5 Vision
2.2 GBestimated usage
LFM2.5-VL-3B · selected for your memory
Chat & writingCodingVision2.7B parameters32K configured context
On Hugging Face 2026-08-11
214 HF likes · 26.7K downloads / month · source model
A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.
1.6 GB weights0.1 GB cache & state0.5 GB runtime
8.6 GB to spareother
Other sizes & editions (2)
● Comfortable fitQ4_K_M · GGUF
NVIDIA
Nemotron 3
3.4 GBestimated usage
Nemotron-3-Nano-4B · selected for your memory
Chat & writingCodingReasoning4.0B parameters256K configured context
On Hugging Face 2026-03-07
126 HF likes · 3.4M downloads / month · source model
An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.
2.8 GB weights0.2 GB cache & state0.5 GB runtime
7.4 GB to spareother
Other sizes & editions (2)
● Comfortable fitQ4_K_M · GGUF
Meta
Llama 3.2
2.9 GBestimated usage
Llama-3.2-3B · selected for your memory
Chat & writing3.2B parameters128K configured context
On Hugging Face 2024-09-18
2.7K HF likes · 1.6M downloads / month · source model
An instruction-tuned assistant for conversation, writing, and summarization. Subject to the Llama community license.
1.9 GB weights0.5 GB cache & state0.5 GB runtime
7.9 GB to sparellama3.2
Other sizes & editions (1)
● Comfortable fitQ4_K_M · GGUF
OpenBMB
MiniCPM V 4.6
1.1 GBestimated usage
MiniCPM-V-4.6
Chat & writingCodingVision0.8B parameters256K configured context
On Hugging Face 2026-04-13
1.2K HF likes · 287K downloads / month · source model
A model for text and image understanding. Select Vision to include its image encoder in the memory estimate.
0.5 GB weights0.1 GB cache & state0.5 GB runtime
9.7 GB to spareapache-2.0
● Tight fitQ4_K_M · GGUF
Microsoft
Phi 4 Reasoning
10.1 GBestimated usage
Phi-4-reasoning-plus
Chat & writingCodingReasoning14.7B parameters32K configured context
On Hugging Face 2025-04-17
348 HF likes · 10K downloads / month · source model
An open-weight language model. See the publisher’s model card for its capabilities, usage instructions, and license.
8.5 GB weights0.8 GB cache & state0.9 GB runtime
0.7 GB to sparemit