Reference

SI Model Directory

Frontier and notable open SI models — who makes them, when they shipped, and where to try them. Filter by maker, weights, modality, or release year.

Super intelligence (SI) is the term the US government adopted in 2026 for what was called AI; the models below are today's SI systems, not artificial superintelligence (ASI) in the technical sense of a system that outperforms humans at most intellectual tasks.

Every model listed here links to its maker's official announcement or documentation — specs come from those sources, and anything we can't verify is left out. For news about these models, see the SI models coverage hub. Agents and scripts can pull this directory as JSON.

What can your hardware run?

Enter your VRAM. Find local models that fit, with Hugging Face downloads.

Find your model →

At a glance

ModelMakerReleasedWeightsLicenseModality
Claude Sonnet 5.5 Anthropic 2026-09-28 Closed — Multimodal
Claude Opus 5.5 Anthropic 2026-09-22 Closed — Multimodal
Claude Fable 5.1 Anthropic 2026-09-01 Closed — Multimodal
Claude Opus 5 Anthropic 2026-07-24 Closed — Multimodal
Claude Sonnet 5 Anthropic 2026-06-30 Closed — Multimodal
Claude Fable 5 Anthropic 2026-06-09 Closed — Multimodal
Claude Haiku 4.5 Anthropic 2025-10-01 Closed — Multimodal
GPT-6 Astra OpenAI 2026-09-03 Closed — Multimodal
GPT-6.1 Sol OpenAI 2026-09-29 Closed — Multimodal
GPT-6 Sol OpenAI 2026-09-22 Closed — Multimodal
GPT-6 Luna OpenAI 2026-09-22 Closed — Multimodal
GPT-5.6 OpenAI 2026-07-09 Closed — Multimodal
gpt-oss-120b OpenAI 2025-08 Open weights Apache 2.0 Text
gpt-oss-20b OpenAI 2025-08 Open weights Apache 2.0 Text
Gemini 4 Argon Google DeepMind 2026-09-30 Closed — Multimodal
Gemini 3.8 Flash Google DeepMind 2026-09-02 Closed — Multimodal
Gemini 3.7 Flash Google DeepMind 2026-08-13 Closed — Multimodal
Gemini 3.5 Flash-Lite Google DeepMind 2026-07-21 Closed — Multimodal
Gemma 4 Google DeepMind 2026-04-02 Open weights Apache 2.0 Multimodal
Muse Spark Meta AI 2026-04-08 Closed — Multimodal
Muse Glimmer 30B Meta AI 2026-08-10 Open weights Apache 2.0 Multimodal
Llama 4 Maverick Meta AI 2025-04-05 Open weights Llama 4 Community License Multimodal
Llama 4 Scout Meta AI 2025-04-05 Open weights Llama 4 Community License Multimodal
DeepSeek-V4.1-Flash DeepSeek 2026-09-10 Open weights MIT Multimodal
DeepSeek-V4-Pro DeepSeek 2026-04 Open weights MIT Multimodal
Qwen3.8-Max Alibaba (Qwen) 2026-08-03 Closed — Multimodal
Qwen3.8-27B Alibaba (Qwen) 2026-08 Open weights Apache 2.0 Multimodal
Qwen3.8-Flash-Next Alibaba (Qwen) 2026-08 Open weights Qwen Community License 1.0 Multimodal
Mistral Medium 3.5 Mistral AI 2026-04 Open weights Modified MIT License Multimodal
Mistral Large 3 Mistral AI 2025-12-02 Open weights Apache 2.0 Multimodal
Grok 4.7 SpaceXAI 2026-09-21 Closed — Multimodal
Grok 4.6 SpaceXAI 2026-08-12 Closed — Multimodal
MiniMax M3 MiniMax 2026-06-01 Open weights MiniMax Community License Multimodal
Kimi K3 Moonshot AI 2026-07-27 Open weights Kimi K3 License Multimodal
GLM-5.3 Z.ai (Zhipu AI) 2026-08 Open weights GLM-5.3 License Text
GLM-5.3-Flash Z.ai (Zhipu AI) 2026-08-26 Open weights MIT Multimodal
NVIDIA Nemotron 3.5 Lightning NVIDIA 2026-08-11 Open weights OpenMDW-1.1 Text

Anthropic

ClosedMultimodal

Claude Sonnet 5.5

The balanced tier in Anthropic's current lineup, positioned as a faster, lower-cost complement to Claude Opus 5.5 and aimed at well-scoped everyday tasks, bug fixes, and producing polished documents, slides, and spreadsheets. Priced unchanged from Claude Sonnet 5 at $2/$10 per million input/output tokens, with cache reads at $0.20 per million; Anthropic says it generates output more than 30% faster than Sonnet 5 and costs up to 30% less per task because it needs fewer tokens for the same work. Carries a 1M-token context window and 128K-token output, with adaptive thinking at a high default effort level. Anthropic says Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment.

Released 2026-09-28
AnnouncementTry itDocs
ClosedMultimodal

Claude Opus 5.5

Anthropic's current Opus-tier model and the one its docs recommend for most workloads, priced at $4/$20 per million input/output tokens. Anthropic says it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Claude Opus 5. Carries a 1M-token context window and 128K-token output, with adaptive thinking always on at a medium default effort level.

Released 2026-09-22
AnnouncementTry itDocs
ClosedMultimodal

Claude Fable 5.1

Anthropic's top-tier model for demanding reasoning and long-horizon agentic work, extending Claude Fable 5 at the same $10/$50 per million input/output token pricing, with cache reads at 2.5% of the base input price. Has a 1M-token context window and 128K-token output; Anthropic's docs now point to Claude Opus 5.5 for most workloads and to Fable 5.1 when Opus 5.5 at higher effort falls short. The same capabilities with more permissive safeguards are offered as Claude Mythos 5.1 by invitation only.

Released 2026-09-01
AnnouncementTry itDocs
ClosedMultimodal

Claude Opus 5

The prior Opus-tier flagship, positioned near Fable 5's intelligence at roughly half the price, now a legacy model succeeded by Claude Opus 5.5 but still available via the Claude API. Ships with a 1M-token context window.

Released 2026-07-24
AnnouncementTry itDocs
ClosedMultimodal

Claude Sonnet 5

The balanced tier in the Claude 5 family, tuned for the best combination of speed and intelligence across coding and agent workloads, now a legacy model succeeded by Claude Sonnet 5.5 but still available via the Claude API. Carries a 1M-token context window.

Released 2026-06-30
AnnouncementTry itDocs
ClosedMultimodal

Claude Fable 5

The first generally available Mythos-class tier, aimed at long-running agents, now a legacy model succeeded by Claude Fable 5.1 but still available via the Claude API. Priced at $10/$50 per million input/output tokens.

Released 2026-06-09
AnnouncementDocs
ClosedMultimodal

Claude Haiku 4.5

Anthropic's fastest model, offering near-frontier intelligence at low cost with a 200K-token context window. Supports optional extended thinking.

Released 2025-10-01
Try itDocs

OpenAI

ClosedMultimodal

GPT-6 Astra

OpenAI's most capable model, presented as a new model generation and built for complex reasoning, coding, computer use, research, and document creation, with text and image input and a 1.05M-token context window. It is the first OpenAI model rated at the Critical level of cybersecurity capability under its Preparedness Framework, and it rolled out in stages, starting with a limited set of organizations before ChatGPT paid plans and the API.

Released 2026-09-03
AnnouncementModel cardDocs
ClosedMultimodal

GPT-6.1 Sol

An upgrade to GPT-6 Sol announced at OpenAI's DevDay 2026, which OpenAI's docs say delivers near-Astra performance at a lower cost for complex coding, computer use, and professional work. Priced at $2/$10 per million input/output tokens for prompts up to 272K tokens, with cached input at $0.10 and cache writes at $2.50 per million. Takes text and image input and returns text, with a 1.05M-token context window (922K maximum input, 128K maximum output).

Released 2026-09-29
AnnouncementDocs
ClosedMultimodal

GPT-6 Sol

The mid-tier GPT-6 reasoning model, built to bring much of GPT-6 Astra's capability into a faster, cheaper model at $2/$10 per million input/output tokens. Takes text and image input with a 1.05M-token context window (922K maximum input, 128K maximum output). OpenAI halved API prices for Sol and Luna against their GPT-5.6 counterparts' promotional pricing. Succeeded by GPT-6.1 Sol on 2026-09-29, and still listed in OpenAI's model catalog with no deprecation notice.

Released 2026-09-22
AnnouncementDocs
ClosedMultimodal

GPT-6 Luna

OpenAI's most cost-efficient GPT-6 model, aimed at focused, high-volume work at $0.10/$0.50 per million input/output tokens. Takes text and image input with a 1.05M-token context window (922K maximum input, 128K maximum output). It is the GPT-6 model offered to ChatGPT Free and Go users in the desktop app, alongside the paid plans.

Released 2026-09-22
AnnouncementDocs
ClosedMultimodal

GPT-5.6

OpenAI's GPT-5.6 family, shipping in three tiers - Luna, Terra, and Sol - from most cost-efficient to most capable, with Sol positioned at launch as OpenAI's strongest coding and vision model. The tiers are not deprecated and remain available in the API, but OpenAI's model catalog now leads with the GPT-6 family of Astra, Sol, and Luna.

Released 2026-07-09
AnnouncementDocs
Open weightsText

gpt-oss-120b

OpenAI's larger open-weight model (about 117B parameters), released under Apache 2.0 for local and self-hosted use with a focus on reasoning tasks.

Released 2025-08 Apache 2.0
Weights
Open weightsText

gpt-oss-20b

OpenAI's smaller open-weight model (about 22B parameters) under Apache 2.0, designed to run efficiently on modest hardware.

Released 2025-08 Apache 2.0
Weights

Google DeepMind

ClosedMultimodal

Gemini 4 Argon

The first model in Google's Gemini 4 generation, announced for complex, long-horizon workflows across real-world software engineering, enterprise knowledge work such as legal and finance, and cyber defense. It raises the output token limit to 1M tokens, up from 64K on earlier Gemini models, which Google says lets it generate hundreds of thousands of tokens in a single trajectory. Announced at an introductory price of $2/$10 per million input/output tokens, rising to $4/$20 afterwards, with cached input at 95% off the input price. At announcement it was rolling out only to a set of trusted cyber defenders through Google's Fairwind Program, with broader access planned starting with paid API customers and Google AI Ultra subscribers; it is not yet listed in the Gemini API model table. Google reports 77.9% on DeepSWE v1.1.

Released 2026-09-30
AnnouncementDocs
ClosedMultimodal

Gemini 3.8 Flash

Google's most intelligent Flash model, built for long-horizon software engineering, autonomous agents, and multi-step reasoning, taking text, image, video, audio, and PDF input with a 1M-token context. Stable in the Gemini API and available in Vertex AI and the Gemini app; a separate Gemini 3.8 Flash Cyber variant for vulnerability detection and patching is limited to trusted defenders through Google's Fairwind Program.

Released 2026-09-02
AnnouncementDocs
ClosedMultimodal

Gemini 3.7 Flash

Google's previous Flash workhorse, a high-throughput model tuned for coding and agentic workflows, succeeded by Gemini 3.8 Flash on 2026-09-02. Google says it remains fully supported for efficiency-first workloads.

Released 2026-08-13
AnnouncementDocs
ClosedMultimodal

Gemini 3.5 Flash-Lite

The fastest, lowest-cost model in Google's Gemini 3.5 line, aimed at high-volume, latency-sensitive workloads.

Released 2026-07-21
Docs
Open weightsMultimodal

Gemma 4

Google DeepMind's open-weight family (E2B, E4B, 26B MoE, and 31B dense) built from the same research as Gemini 3, now shipped under Apache 2.0. Handles text, images, audio, and video with up to 256K context.

Released 2026-04-02 Apache 2.0
AnnouncementModel cardDocs

Meta AI

ClosedMultimodal

Muse Spark

Meta Superintelligence Labs' proprietary flagship, a multimodal reasoning model that powers the Meta AI assistant. Closed-weight and offered through a paid API; the latest version, Muse Spark 1.3, shipped on 2026-09-02 with improved agentic and coding performance and is available through the Muse Code coding agent and the Meta Model API.

Released 2026-04-08
Announcement
Open weightsMultimodal

Muse Glimmer 30B

Meta's first open-weight model since Llama 4 and the first open release from Meta Superintelligence Labs. A roughly 29.6B-parameter dense model with text and image input, distilled from the Muse Spark flagship and tuned for on-device agentic use, with a 128K-token context window.

Released 2026-08-10 Apache 2.0
AnnouncementWeights
Open weightsMultimodal

Llama 4 Maverick

Meta's larger open-weight Llama 4 model, a mixture-of-experts design (about 400B total, 17B active parameters) with native text-and-image input.

Released 2025-04-05 Llama 4 Community License
AnnouncementWeights
Open weightsMultimodal

Llama 4 Scout

The smaller Llama 4 model that fits on a single high-end GPU, with 17B active parameters and a very long context window.

Released 2025-04-05 Llama 4 Community License
AnnouncementWeights

DeepSeek

Open weightsMultimodal

DeepSeek-V4.1-Flash

The smallest model in DeepSeek's new V4.1 architecture family, a mixture-of-experts model with 552B backbone parameters that activates about 8B per token during prefill and 16B during decode, using a causal encoder-decoder design. MIT-licensed with native image understanding and contexts of up to 1M tokens; it replaced V4-Flash in DeepSeek's API.

Released 2026-09-10 MIT
AnnouncementWeights
Open weightsMultimodal

DeepSeek-V4-Pro

DeepSeek's largest V4 open-weight model, a 1.6T-parameter mixture-of-experts release under the MIT license with image input. DeepSeek says V4.1-Flash now outperforms it. DeepSeek first said its API would route V4-Pro requests to V4.1-Flash from 2026-09-14 until V4.1-Pro launches, but later said that, in response to user demand, it will keep serving V4-Pro after that date with billing unchanged; the weights remain available.

Released 2026-04 MIT
Weights

Alibaba (Qwen)

ClosedMultimodal

Qwen3.8-Max

Alibaba's flagship Qwen model, a 2.4T-parameter mixture-of-experts system (about 95B active) with a 1M-token context and native text-plus-vision input. The full multimodal Max is API-only; Alibaba released the underlying text-only base checkpoint (Qwen3.8-2.4T-A95B) as open weights on 2026-08-12 under the Qwen3.8-Max license.

Released 2026-08-03
AnnouncementWeights
Open weightsMultimodal

Qwen3.8-27B

An open-weight (Apache 2.0) Qwen model aimed at coding and long-horizon agentic work, with a vision encoder for text, image, and video input and a 256K-token context window (extensible to about 1M). Sized at roughly 27B parameters to run on a single high-end GPU.

Released 2026-08 Apache 2.0
Weights
Open weightsMultimodal

Qwen3.8-Flash-Next

An open-weight experimental preview of the architecture Qwen says will underpin Qwen4, released as a mixture-of-experts model with about 125B total parameters and roughly 6B activated per token, alongside a 51B n-gram embedding table and a 4B multi-token-prediction module. Pairs Gated DeltaNet with Qwen Sparse Attention in a hybrid attention design across 48 layers and 512 experts. Takes text, image, and video input and returns text, with a 262,144-token native context window extensible to about 1M tokens.

Released 2026-08 Qwen Community License 1.0
AnnouncementWeights

Mistral AI

Open weightsMultimodal

Mistral Medium 3.5

Mistral's first flagship merged model, a dense 128B-parameter model that handles instruction-following, reasoning, and coding in a single set of weights, with text and image input and a 256K-token context window. It became the default model in Mistral's assistant, which was renamed from Le Chat to Vibe in August 2026, and replaced Devstral 2 in the Vibe coding agent. The open weights use a modified MIT license that excludes companies with more than $20 million in monthly revenue, which must get a commercial license from Mistral or use its hosted services.

Released 2026-04 Modified MIT License
AnnouncementModel cardWeights
Open weightsMultimodal

Mistral Large 3

A 675B-parameter sparse mixture-of-experts open-weight model (41B active) from Mistral under Apache 2.0, with an added vision encoder for image understanding. Mistral's docs still list it as a general-purpose open-weight model alongside its newer flagship, Mistral Medium 3.5.

Released 2025-12-02 Apache 2.0
Model cardWeights

SpaceXAI

ClosedMultimodal

Grok 4.7

SpaceXAI's frontier model for coding, agentic tasks, and knowledge work. SpaceXAI says it uses a new, larger base model than Grok 4.6 and was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. Closed-weight, with a 500K-token context window and text-and-image input. It shipped with what SpaceXAI describes as an entirely new safeguard stack and its strongest results to date on refusals and jailbreak resistance.

Released 2026-09-21
AnnouncementTry itDocs
ClosedMultimodal

Grok 4.6

SpaceXAI's 1.5T-parameter frontier model, a refinement of Grok 4.5 through improved fine-tuning and reinforcement learning rather than a scale increase. Closed-weight, and succeeded by Grok 4.7 on 2026-09-21 but still listed in SpaceXAI's model and pricing tables alongside it.

Released 2026-08-12
AnnouncementTry it

MiniMax

Open weightsMultimodal

MiniMax M3

MiniMax's latest M-series model for agentic reasoning, tool use, and coding, with about 428B total parameters of which about 23B are activated. Natively multimodal from the first step of training, it takes text, image, and video input, can operate a desktop computer, and has a 1M-token context window, served by a MiniMax Sparse Attention design that MiniMax says speeds up long-context processing. The MiniMax Community License requires separate written authorization from MiniMax for products with more than $20 million in yearly revenue.

Released 2026-06-01 MiniMax Community License
AnnouncementWeights

Moonshot AI

Open weightsMultimodal

Kimi K3

Moonshot AI's open-weight mixture-of-experts model with 2.8T total parameters (104B active) and a 1M-token context, among the largest open models released. Multimodal via a native vision encoder.

Released 2026-07-27 Kimi K3 License
Weights

Z.ai (Zhipu AI)

Open weightsText

GLM-5.3

Z.ai's flagship open-weight model, a mixture-of-experts release with about 744B total and 40B active parameters that uses the same base model as GLM-5.2, with all gains coming from post-training. Z.ai highlights its coding and emergent cybersecurity capabilities. The GLM-5.3 License is permissive but requires Model-as-a-Service operators with more than $10 billion in annual revenue to pass a Z.ai security review before commercial use.

Released 2026-08 GLM-5.3 License
AnnouncementWeights
Open weightsMultimodal

GLM-5.3-Flash

Z.ai's open-weight (MIT) model with 320B total parameters and about 18B active per token, pairing mixture-of-experts routing with a hybrid sparse-and-linear attention design. The first natively multimodal model in the GLM-5 series, taking text and image input with a 1M-token context window; previewed anonymously as "Ox Alpha" and served entirely on domestic Chinese chips.

Released 2026-08-26 MIT
AnnouncementWeights

NVIDIA

Open weightsText

NVIDIA Nemotron 3.5 Lightning

NVIDIA's open-weight mixture-of-experts model (about 30B total parameters, roughly 3B active) built for high-volume, long-running agent workloads such as coding, security triage, and customer support. Released with weights, training data, and recipes under the permissive OpenMDW-1.1 license.

Released 2026-08-11 OpenMDW-1.1
AnnouncementWeights