Reference
SI Model Directory
Frontier and notable open SI models — who makes them, when they shipped, and where to try them. Filter by maker, weights, modality, or release year.
Super intelligence (SI) is the term the US government adopted in 2026 for what was called AI; the models below are today's SI systems, not artificial superintelligence (ASI) in the technical sense of a system that outperforms humans at most intellectual tasks.
Every model listed here links to its maker's official announcement or documentation — specs come from those sources, and anything we can't verify is left out. For news about these models, see the SI models coverage hub. Agents and scripts can pull this directory as JSON.
Enter your VRAM. Find local models that fit, with Hugging Face downloads.
No models match those filters.
At a glance
| Model | Maker | Released | Weights | License | Modality |
|---|---|---|---|---|---|
| Claude Sonnet 5.5 | Anthropic | 2026-09-28 | Closed | — | Multimodal |
| Claude Opus 5.5 | Anthropic | 2026-09-22 | Closed | — | Multimodal |
| Claude Fable 5.1 | Anthropic | 2026-09-01 | Closed | — | Multimodal |
| Claude Opus 5 | Anthropic | 2026-07-24 | Closed | — | Multimodal |
| Claude Sonnet 5 | Anthropic | 2026-06-30 | Closed | — | Multimodal |
| Claude Fable 5 | Anthropic | 2026-06-09 | Closed | — | Multimodal |
| Claude Haiku 4.5 | Anthropic | 2025-10-01 | Closed | — | Multimodal |
| GPT-6 Astra | OpenAI | 2026-09-03 | Closed | — | Multimodal |
| GPT-6.1 Sol | OpenAI | 2026-09-29 | Closed | — | Multimodal |
| GPT-6 Sol | OpenAI | 2026-09-22 | Closed | — | Multimodal |
| GPT-6 Luna | OpenAI | 2026-09-22 | Closed | — | Multimodal |
| GPT-5.6 | OpenAI | 2026-07-09 | Closed | — | Multimodal |
| gpt-oss-120b | OpenAI | 2025-08 | Open weights | Apache 2.0 | Text |
| gpt-oss-20b | OpenAI | 2025-08 | Open weights | Apache 2.0 | Text |
| Gemini 4 Argon | Google DeepMind | 2026-09-30 | Closed | — | Multimodal |
| Gemini 3.8 Flash | Google DeepMind | 2026-09-02 | Closed | — | Multimodal |
| Gemini 3.7 Flash | Google DeepMind | 2026-08-13 | Closed | — | Multimodal |
| Gemini 3.5 Flash-Lite | Google DeepMind | 2026-07-21 | Closed | — | Multimodal |
| Gemma 4 | Google DeepMind | 2026-04-02 | Open weights | Apache 2.0 | Multimodal |
| Muse Spark | Meta AI | 2026-04-08 | Closed | — | Multimodal |
| Muse Glimmer 30B | Meta AI | 2026-08-10 | Open weights | Apache 2.0 | Multimodal |
| Llama 4 Maverick | Meta AI | 2025-04-05 | Open weights | Llama 4 Community License | Multimodal |
| Llama 4 Scout | Meta AI | 2025-04-05 | Open weights | Llama 4 Community License | Multimodal |
| DeepSeek-V4.1-Flash | DeepSeek | 2026-09-10 | Open weights | MIT | Multimodal |
| DeepSeek-V4-Pro | DeepSeek | 2026-04 | Open weights | MIT | Multimodal |
| Qwen3.8-Max | Alibaba (Qwen) | 2026-08-03 | Closed | — | Multimodal |
| Qwen3.8-27B | Alibaba (Qwen) | 2026-08 | Open weights | Apache 2.0 | Multimodal |
| Qwen3.8-Flash-Next | Alibaba (Qwen) | 2026-08 | Open weights | Qwen Community License 1.0 | Multimodal |
| Mistral Medium 3.5 | Mistral AI | 2026-04 | Open weights | Modified MIT License | Multimodal |
| Mistral Large 3 | Mistral AI | 2025-12-02 | Open weights | Apache 2.0 | Multimodal |
| Grok 4.7 | SpaceXAI | 2026-09-21 | Closed | — | Multimodal |
| Grok 4.6 | SpaceXAI | 2026-08-12 | Closed | — | Multimodal |
| MiniMax M3 | MiniMax | 2026-06-01 | Open weights | MiniMax Community License | Multimodal |
| Kimi K3 | Moonshot AI | 2026-07-27 | Open weights | Kimi K3 License | Multimodal |
| GLM-5.3 | Z.ai (Zhipu AI) | 2026-08 | Open weights | GLM-5.3 License | Text |
| GLM-5.3-Flash | Z.ai (Zhipu AI) | 2026-08-26 | Open weights | MIT | Multimodal |
| NVIDIA Nemotron 3.5 Lightning | NVIDIA | 2026-08-11 | Open weights | OpenMDW-1.1 | Text |
Anthropic
Claude Sonnet 5.5
The balanced tier in Anthropic's current lineup, positioned as a faster, lower-cost complement to Claude Opus 5.5 and aimed at well-scoped everyday tasks, bug fixes, and producing polished documents, slides, and spreadsheets. Priced unchanged from Claude Sonnet 5 at $2/$10 per million input/output tokens, with cache reads at $0.20 per million; Anthropic says it generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task because it needs fewer tokens for the same work. Carries a 1M-token context window and 128K-token output, with adaptive thinking at a high default effort level. Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.
Claude Opus 5.5
Anthropic's current Opus-tier model and the one its docs recommend for most workloads, priced at $4/$20 per million input/output tokens. Anthropic says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Claude Opus 5. Carries a 1M-token context window and 128K-token output, with adaptive thinking always on at a medium default effort level.
Claude Fable 5.1
Anthropic's top-tier model for demanding reasoning and long-horizon agentic work, extending Claude Fable 5 at the same $10/$50 per million input/output token pricing, with cache reads at 2.5% of the base input price. Has a 1M-token context window and 128K-token output; Anthropic's docs now point to Claude Opus 5.5 for most workloads and to Fable 5.1 when Opus 5.5 at higher effort falls short. The same capabilities with more permissive safeguards are offered as Claude Mythos 5.1 by invitation only.
Claude Opus 5
The prior Opus-tier flagship, positioned near Fable 5's intelligence at roughly half the price, now a legacy model succeeded by Claude Opus 5.5 but still available via the Claude API. Ships with a 1M-token context window.
Claude Sonnet 5
The balanced tier in the Claude 5 family, tuned for the best combination of speed and intelligence across coding and agent workloads, now a legacy model succeeded by Claude Sonnet 5.5 but still available via the Claude API. Carries a 1M-token context window.
Claude Fable 5
The first generally available Mythos-class tier, aimed at long-running agents, now a legacy model succeeded by Claude Fable 5.1 but still available via the Claude API. Priced at $10/$50 per million input/output tokens.
OpenAI
GPT-6 Astra
OpenAI's most capable model, presented as a new model generation and built for complex reasoning, coding, computer use, research, and document creation, with text and image input and a 1.05M-token context window. It is the first OpenAI model rated at the Critical level of cybersecurity capability under its Preparedness Framework, and it rolled out in stages, starting with a limited set of organizations before ChatGPT paid plans and the API.
GPT-6.1 Sol
An upgrade to GPT-6 Sol announced at OpenAI's DevDay 2026, which OpenAI's docs say delivers near-Astra performance at a lower cost for complex coding, computer use, and professional work. Priced at $2/$10 per million input/output tokens for prompts up to 272K tokens, with cached input at $0.10 and cache writes at $2.50 per million. Takes text and image input and returns text, with a 1.05M-token context window (922K maximum input, 128K maximum output).
GPT-6 Sol
The mid-tier GPT-6 reasoning model, built to bring much of GPT-6 Astra's capability into a faster, cheaper model at $2/$10 per million input/output tokens. Takes text and image input with a 1.05M-token context window (922K maximum input, 128K maximum output). OpenAI halved API prices for Sol and Luna against their GPT-5.6 counterparts' promotional pricing. Succeeded by GPT-6.1 Sol on 2026-09-29, and still listed in OpenAI's model catalog with no deprecation notice.
GPT-6 Luna
OpenAI's most cost-efficient GPT-6 model, aimed at focused, high-volume work at $0.10/$0.50 per million input/output tokens. Takes text and image input with a 1.05M-token context window (922K maximum input, 128K maximum output). It is the GPT-6 model offered to ChatGPT Free and Go users in the desktop app, alongside the paid plans.
GPT-5.6
OpenAI's GPT-5.6 family, shipping in three tiers - Luna, Terra, and Sol - from most cost-efficient to most capable, with Sol positioned at launch as OpenAI's strongest coding and vision model. The tiers are not deprecated and remain available in the API, but OpenAI's model catalog now leads with the GPT-6 family of Astra, Sol, and Luna.
gpt-oss-120b
OpenAI's larger open-weight model (about 117B parameters), released under Apache 2.0 for local and self-hosted use with a focus on reasoning tasks.
gpt-oss-20b
OpenAI's smaller open-weight model (about 22B parameters) under Apache 2.0, designed to run efficiently on modest hardware.
Google DeepMind
Gemini 4 Argon
The first model in Google's Gemini 4 generation, announced for complex, long-horizon workflows across real-world software engineering, enterprise knowledge work such as legal and finance, and cyber defense. It raises the output token limit to 1M tokens, up from 64K on earlier Gemini models, which Google says lets it generate hundreds of thousands of tokens in a single trajectory. Announced at an introductory price of $2/$10 per million input/output tokens, rising to $4/$20 afterwards, with cached input at 95% off the input price. At announcement it was rolling out only to a set of trusted cyber defenders through Google's Fairwind Program, with broader access planned starting with paid API customers and Google AI Ultra subscribers; it is not yet listed in the Gemini API model table. Google reports 77.9% on DeepSWE v1.1.
Gemini 3.8 Flash
Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents, and multi-step reasoning, taking text, image, video, audio, and PDF input with a 1M-token context. Stable in the Gemini API and available in Vertex AI and the Gemini app; a separate Gemini 3.8 Flash Cyber variant for vulnerability detection and patching is limited to trusted defenders through Google's Fairwind Program.
Gemini 3.7 Flash
Google's previous Flash workhorse, a high-throughput model tuned for coding and agentic workflows, succeeded by Gemini 3.8 Flash on 2026-09-02. Google says it remains fully supported for efficiency-first workloads.
Gemini 3.5 Flash-Lite
The fastest, lowest-cost model in Google's Gemini 3.5 line, aimed at high-volume, latency-sensitive workloads.
Gemma 4
Google DeepMind's open-weight family (E2B, E4B, 26B MoE, and 31B dense) built from the same research as Gemini 3, now shipped under Apache 2.0. Handles text, images, audio, and video with up to 256K context.
Meta AI
Muse Spark
Meta Superintelligence Labs' proprietary flagship, a multimodal reasoning model that powers the Meta AI assistant. Closed-weight and offered through a paid API; the latest version, Muse Spark 1.3, shipped on 2026-09-02 with improved agentic and coding performance and is available through the Muse Code coding agent and the Meta Model API.
Muse Glimmer 30B
Meta's first open-weight model since Llama 4 and the first open release from Meta Superintelligence Labs. A roughly 29.6B-parameter dense model with text and image input, distilled from the Muse Spark flagship and tuned for on-device agentic use, with a 128K-token context window.
Llama 4 Maverick
Meta's larger open-weight Llama 4 model, a mixture-of-experts design (about 400B total, 17B active parameters) with native text-and-image input.
Llama 4 Scout
The smaller Llama 4 model that fits on a single high-end GPU, with 17B active parameters and a very long context window.
DeepSeek
DeepSeek-V4.1-Flash
The smallest model in DeepSeek's new V4.1 architecture family, a mixture-of-experts model with 552B backbone parameters that activates about 8B per token during prefill and 16B during decode, using a causal encoder-decoder design. MIT-licensed with native image understanding and contexts of up to 1M tokens; it replaced V4-Flash in DeepSeek's API.
DeepSeek-V4-Pro
DeepSeek's largest V4 open-weight model, a 1.6T-parameter mixture-of-experts release under the MIT license with image input. DeepSeek says V4.1-Flash now outperforms it. DeepSeek first said its API would route V4-Pro requests to V4.1-Flash from 2026-09-14 until V4.1-Pro launches, but later said that, in response to user demand, it will keep serving V4-Pro after that date with billing unchanged; the weights remain available.
Alibaba (Qwen)
Qwen3.8-Max
Alibaba's flagship Qwen model, a 2.4T-parameter mixture-of-experts system (about 95B active) with a 1M-token context and native text-plus-vision input. The full multimodal Max is API-only; Alibaba released the underlying text-only base checkpoint (Qwen3.8-2.4T-A95B) as open weights on 2026-08-12 under the Qwen3.8-Max license.
Qwen3.8-27B
An open-weight (Apache 2.0) Qwen model aimed at coding and long-horizon agentic work, with a vision encoder for text, image, and video input and a 256K-token context window (extensible to about 1M). Sized at roughly 27B parameters to run on a single high-end GPU.
Qwen3.8-Flash-Next
An open-weight experimental preview of the architecture Qwen says will underpin Qwen4, released as a mixture-of-experts model with about 125B total parameters and roughly 6B activated per token, alongside a 51B n-gram embedding table and a 4B multi-token-prediction module. Pairs Gated DeltaNet with Qwen Sparse Attention in a hybrid attention design across 48 layers and 512 experts. Takes text, image, and video input and returns text, with a 262,144-token native context window extensible to about 1M tokens.
Mistral AI
Mistral Medium 3.5
Mistral's first flagship merged model, a dense 128B-parameter model that handles instruction-following, reasoning, and coding in a single set of weights, with text and image input and a 256K-token context window. It became the default model in Mistral's assistant, which was renamed from Le Chat to Vibe in August 2026, and replaced Devstral 2 in the Vibe coding agent. The open weights use a modified MIT license that excludes companies with more than $20 million in monthly revenue, which must get a commercial license from Mistral or use its hosted services.
Mistral Large 3
A 675B-parameter sparse mixture-of-experts open-weight model (41B active) from Mistral under Apache 2.0, with an added vision encoder for image understanding. Mistral's docs still list it as a general-purpose open-weight model alongside its newer flagship, Mistral Medium 3.5.
SpaceXAI
Grok 4.7
SpaceXAI's frontier model for coding, agentic tasks, and knowledge work. SpaceXAI says it uses a new, larger base model than Grok 4.6 and was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. Closed-weight, with a 500K-token context window and text-and-image input. It shipped with what SpaceXAI describes as an entirely new safeguard stack and its strongest results to date on refusals and jailbreak resistance.
Grok 4.6
SpaceXAI's 1.5T-parameter frontier model, a refinement of Grok 4.5 through improved fine-tuning and reinforcement learning rather than a scale increase. Closed-weight, and succeeded by Grok 4.7 on 2026-09-21 but still listed in SpaceXAI's model and pricing tables alongside it.
MiniMax
MiniMax M3
MiniMax's latest M-series model for agentic reasoning, tool use, and coding, with about 428B total parameters of which about 23B are activated. Natively multimodal from the first step of training, it takes text, image, and video input, can operate a desktop computer, and has a 1M-token context window, served by a MiniMax Sparse Attention design that MiniMax says speeds up long-context processing. The MiniMax Community License requires separate written authorization from MiniMax for products with more than $20 million in yearly revenue.
Moonshot AI
Kimi K3
Moonshot AI's open-weight mixture-of-experts model with 2.8T total parameters (104B active) and a 1M-token context, among the largest open models released. Multimodal via a native vision encoder.
Z.ai (Zhipu AI)
GLM-5.3
Z.ai's flagship open-weight model, a mixture-of-experts release with about 744B total and 40B active parameters that uses the same base model as GLM-5.2, with all gains coming from post-training. Z.ai highlights its coding and emergent cybersecurity capabilities. The GLM-5.3 License is permissive but requires Model-as-a-Service operators with more than $10 billion in annual revenue to pass a Z.ai security review before commercial use.
GLM-5.3-Flash
Z.ai's open-weight (MIT) model with 320B total parameters and about 18B active per token, pairing mixture-of-experts routing with a hybrid sparse-and-linear attention design. The first natively multimodal model in the GLM-5 series, taking text and image input with a 1M-token context window; previewed anonymously as "Ox Alpha" and served entirely on domestic Chinese chips.
NVIDIA
NVIDIA Nemotron 3.5 Lightning
NVIDIA's open-weight mixture-of-experts model (about 30B total parameters, roughly 3B active) built for high-volume, long-running agent workloads such as coding, security triage, and customer support. Released with weights, training data, and recipes under the permissive OpenMDW-1.1 license.