Super Intelligence (SI) is the name a September 29, 2026 White House executive order gave to what agencies used to call artificial intelligence, and serving SI models is where the field's economics are decided. Training wins the headlines; inference pays the bills. A frontier model is trained once and then serves billions of tokens, so margins, pricing and which SI products are even viable get set at inference time. Almost every optimization that matters is, underneath, a way to move more tokens through fixed hardware without losing more quality than the market will accept. Making sense of the stack starts with one hardware fact.

Why memory bandwidth rules SI serving

Autoregressive decoding produces one token at a time, and every token requires reading all of the model's weights — plus a growing cache — out of GPU memory. For a single request the arithmetic is lopsided: the GPU's compute units mostly sit idle waiting on memory. Decoding is memory-bandwidth-bound, not compute-bound, and that one fact explains most of what follows.

The cache in question is the KV cache. Attention lets each new token look at every earlier one, so instead of recomputing keys and values for the whole sequence at each step, you store them. The KV cache grows linearly with sequence length and batch size and, at long context, ends up dominating memory — often larger than the weights themselves. That is also why long-context serving costs so much: you pay not only for the tokens generated but for the bandwidth to stream an ever-bigger cache on every step. Nearly every inference optimization either shrinks what has to be read (quantization, cache compression) or gets more useful work out of each read (batching, speculative decoding, sparsity).

Quantization

Quantization stores weights — and sometimes activations and the KV cache — in fewer bits than the 16-bit floats used during training. Because decode is memory-bound, halving the bytes per weight roughly doubles throughput and shrinks the memory footprint in proportion. For most deployments it is the single highest-leverage optimization.

Formats split along a few axes. Weight-only methods such as GPTQ and AWQ compress weights to about 4 bits after training while leaving activations at higher precision; AWQ's insight — protect the small share of salient weight channels — makes 4-bit weights nearly lossless for many models and is a common production default. GGUF, the format llama.cpp popularized, provides a range of mixed-precision schemes tuned for CPU and consumer-GPU inference. Activation-aware methods like SmoothQuant, and low-precision floating formats like FP8, quantize weights and activations together, unlocking faster matrix multiplies on hardware with low-precision tensor cores — FP8 on NVIDIA Hopper, and now FP4 on Blackwell, where frontier SI models increasingly ship with quantization-aware training so low precision is built in during training instead of bolted on afterward.

The trade is always accuracy against footprint, and it is not linear. Dropping from 16-bit to 8-bit is close to free; 4-bit is usually a good deal with careful methods; under 4 bits, quality falls off quickly and unevenly — reasoning and long-context behavior tend to break before perplexity visibly does, which is why a model that passes a quick eval can fail quietly inside an SI agent loop. Quantizing the KV cache itself (to 8 or even 4 bits) is a newer lever aimed squarely at long-context memory pressure. The working rule: quantize weights aggressively, activations moderately, and always validate on the real downstream task instead of on loss.

Speculative decoding

Speculative decoding cuts the latency of sequential generation without altering the output distribution. A small, cheap draft model proposes several tokens ahead, and the large target model checks them all in one forward pass. Since a forward pass over a handful of candidates costs about the same as a pass over one — decode is memory-bound, so the weights are being read anyway — the large model verifies several guesses for roughly the price of producing one. Accepted tokens stay; the first rejected token is corrected and generation carries on. Crucially, the accepted output is provably identical in distribution to what the target model would have generated alone: it is a speedup, not an approximation.

Variants differ in where the draft comes from. A separate small model is the classic form; Medusa attaches extra prediction heads to the target model itself; EAGLE and its successors predict at the feature level and, using tree-structured drafting to lift the acceptance rate, are the current state of the art. Realistic speedups are roughly 2-3x on interactive, low-batch paths. The catch: the gain shrinks as batch size rises, because a busy server already uses its memory bandwidth well — speculative decoding pays off most for latency-sensitive, lightly batched workloads, not for maximizing throughput on a saturated GPU.

Batching and continuous batching

If one request wastes the GPU's compute, the answer is to serve many at once. Batching runs multiple sequences together so each costly weight read is spread across many tokens — the most important lever for throughput and cost per token. Naive static batching stalls, though: requests finish at different times, and a batch that waits on its slowest member leaves the GPU idle while short requests sit finished.

Continuous batching (also called in-flight batching) fixes this by scheduling at the token level — the instant one sequence ends, a queued request takes its slot, so the batch stays full. Combined with PagedAttention, which manages the KV cache in non-contiguous pages like virtual memory and removes the 60-80% waste of pre-allocating a contiguous cache per request, it is why modern serving engines reach many times the throughput of a naive generation loop. Beneath it lies a throughput-latency tension: larger batches lower cost per token but raise time-to-first-token for each user, so serving systems tune that balance — and often split the compute-bound prefill phase from the memory-bound decode phase across separate resources so neither starves the other.

MoE at inference

Mixture-of-experts models change the math. A MoE layer contains many expert sub-networks but sends each token to only a few, so the model carries an enormous total parameter count while activating a small slice per token — DeepSeek-V3, for example, has 671 billion parameters but activates roughly 37 billion per token. That separates capability, which follows total parameters, from per-token compute, which follows active parameters — a genuinely favorable trade. The cost shifts to memory and systems: every expert must stay resident in GPU memory even though most sit idle each step, and at scale experts are sharded across GPUs, turning inference into an all-to-all communication problem where routing imbalance can strand hardware. Complementary techniques like DeepSeek's multi-head latent attention — compressing the KV cache into a shared low-rank latent — work alongside MoE to keep long-context serving affordable.

The SI serving stack

The tools that implement all of this have narrowed to a few serious options. vLLM is the widely adopted open default, built on PagedAttention and continuous batching with an OpenAI-compatible API — the practical first pick when you need to serve something today. SGLang is built around RadixAttention, which reuses the KV cache across requests sharing a prefix, and tends to win on multi-turn and heavily structured workloads. TensorRT-LLM compiles models to extract peak throughput from NVIDIA hardware, with FP8 and FP4 support included — fastest on the metal, at the price of a compilation step and hardware lock-in. llama.cpp, with its GGUF format, owns local and edge inference on CPUs and consumer GPUs. The differences matter less than the fact that all four now implement the same core ideas; picking one is mostly about your hardware, your workload shape, and how much you prize peak throughput over portability.

The through-line bears repeating: SI inference optimization is applied economics. Each technique here trades some mix of accuracy, latency, throughput and engineering complexity, and the right mix depends entirely on whether you are tuning a chatbot's time-to-first-token, an agent's cost per completed task, or a local model's ability to run on hardware you own at all. The one constant is the memory wall — until that hardware fact changes, the winning tricks will keep being the ones that read fewer bytes, or get more value from each byte read.