Strata serves a 125B MoE at 94 tokens/s from a 12 GB card and 64 GB of RAM
The byte census in the Strata paper is the whole design in one table. Qwen3.8-Flash-Next stores 34 GB of routed expert weights in its 2-bit file, and any single token reads 0.66 GB of them, about 2%. Everything else in the engine follows from that ratio: if a token touches 2% of the experts, the other 98% do not need to be anywhere near the graphics card, and 24,576 small specialist networks can wait in system RAM until the router asks for ten of them.
The project measures 94 tokens per second at a 4K context on an RTX 5070 with 12 GB of VRAM (engine 0.1.36, Q2_0 file), and a second reference machine with an AMD RX 9070 XT reaches 60 tokens per second with the same 2-bit file. That pair is where the "60 to 94 tokens per second" figure circulating on social media comes from. The machine that produced the higher number also had a Ryzen 5 7600, 64 GB of DDR5-5200 and a 66 GB download on disk, and those parts of the bill are not optional.

Measured throughput on a 12 GB card
The repository publishes measured tables rather than single numbers, and the spread across quantization sizes is the first thing to read. All rows below are from the engine 0.1.26 run on the RTX 5070 machine: output speed on a short chat, prompt speed on a 32K-token prompt.
| File | Download | RAM + VRAM needed | Writes answers | Reads a 32K prompt |
|---|---|---|---|---|
| Q2_0 | 66 GB | 37.6 GB | 93 tokens/s | 2,171 tokens/s |
| IQ2_XS | 68 GB | 39.2 GB | 79 tokens/s | 2,092 tokens/s |
| IQ3_XXS | 76 GB | 47.0 GB | 62 tokens/s | 1,745 tokens/s |
| IQ3_S | 76 GB | 54.8 GB | 53 tokens/s | 1,624 tokens/s |
| Coder (IQ1_M) | 29.6 GB shard 1 | 32 GB | 55 tokens/s | 2,177 tokens/s |
Context length moves the output speed down as the window grows: with Q2_0 the same machine writes 87 tokens/s at 1K, 82 at 32K, 74 at 128K and 60 at 262K. Prompt processing does not degrade, because the sparse attention the model uses reads a bounded set of positions no matter how long the input is. On the AMD card the same Q2_0 file gives 60 tokens/s at a short context, 48 at 128K and 1,160 tokens/s on prompts.

The minimum configuration the maintainer states is 12 GB of VRAM, 32 GB of RAM, and about 80 GB of free disk; 8 GB cards "run, slowly". The RAM column in the table above is the one that decides which file you can load, and 64 GB is what the project recommends for all sizes. Below the floor the failure is explicit rather than graceful: an issue from an RTX 2060 with 6 GB reports that no VRAM was left for the expert cache, which needed at least 241 slots, and the engine named the shortfall in MiB along with the switches that would free them.
Two operational costs are worth pricing in before installing. Strata answers one request at a time by default, and enabling two in parallel slows each answer on a 12 GB card. The first start takes one to three minutes during which the machine can stop responding, because the engine loads 35-55 GB into RAM and pins part of it for the card.
Where the weights live
The paper's Table 1 is a byte census of the 2-bit file, and it is more useful than any marketing sentence about running a large model on a small card.
| Part of the model | Stored | Read per token | Where Strata keeps it |
|---|---|---|---|
| Routed experts (48 layers × 512) | 34.0 GB | 0.66 GB | RAM (all) + VRAM cache (the most-used ~4,500) |
| Mixers, attention, shared experts, routers | 1.8 GB | 1.8 GB | VRAM |
| Gated residual weights | 1.3 GB | 1.3 GB | VRAM |
| Output head | 0.44 GB | 0.44 GB | VRAM |
| N-gram embedding table | 28.8 GB | ≤ 23 KB | SSD, through the OS page cache |
| MTP draft layer | 0.8 GB | per draft | VRAM |
The dense part of the network is about 3.5 GB and stays on the card. The expert tier is split: the GPU holds roughly 4,500 of the 24,576 experts in an adaptive cache, and RAM holds all of them in one page-locked arena. The measured bandwidths on the test machine explain why the split is drawn there: 672 GB/s from VRAM, 41-52 GB/s from DDR5, about 26 GB/s over PCIe 4.0 x16, and 16 row reads per token from the NVMe drive.

The interesting part is how the three tiers run at once. Each layer's router writes the ten chosen expert ids into a small block of pinned host memory, and the CPU spins on that doorbell instead of waiting for a driver call. The experts already in the VRAM cache go to the GPU; a share of the misses is copied over PCIe by the card's copy engine while everything else runs (20% of misses for Q2_0, 55% for the i-quant files, because the CPU needs RAM bandwidth itself); the rest are computed by the CPU cores next to the RAM, and the results go back into mapped memory. One captured CUDA graph covers a whole 48-layer pass, so the host never blocks on a synchronizing call.
The cache learns. It starts from an expert profile recorded on other prompts, then swaps up to 96 experts that the current conversation keeps asking for into the slots of the least-used ones every four rounds. The paper reports a measured hit rate of 0.72 at a 4K context on the 12 GB card, against 0.50 for the profile alone. Every extra gigabyte of VRAM holds roughly 700 more experts, which is why the project estimates an RTX 3090 at 100-140 tokens per second from its 24 GB, and why a faster card with the same capacity would not help much.
Above 64K tokens the KV cache moves to RAM too, with only the most-read 32K tokens resident on the card. That change alone lifted Q2_0 from 50.9 to 62.6 tokens per second at the 262K context, at the cost of about 13.7 KB of RAM per context token.
Six billion active parameters per token
Qwen3.8-Flash-Next is a 125B-parameter model of which 6B run per token, and the architecture is what makes offloading practical. Of its 48 layers, 36 use Gated DeltaNet, a recurrent mixer whose memory is a fixed-size state instead of a growing cache, and the other 12 use Qwen Sparse Attention with two key/value heads attending to at most about 2,048 selected positions. Each layer ends in 512 experts, and the router picks 10 of them. A token therefore touches about 2% of the expert weights, and the 51-billion-parameter n-gram table is addressed by the last three token ids, so its rows are known before they are needed.
Speculation is the other half of the throughput. The model ships a one-layer multi-token-prediction head; the GGUF files do not carry it, so Strata downloads its tensors from the BF16 checkpoint and keeps it in VRAM with a reduced output head over 40,525 likely tokens. Each round sends the last accepted token plus up to three drafts through all 48 layers as one verify window, and the engine keeps every draft the full model agrees with, plus one more token written by the model itself. Average acceptance is 2.4-3.2 tokens per pass. Reading the dense weights is the card's main cost per token and it barely grows when four tokens are processed together, which is where the 1.6-1.8x comes from.
The paper claims this path is exact, meaning token-for-token identical to plain greedy decoding, and says it was tested with forced wrong drafts and rollbacks. That claim is about the verify window, and it has a boundary worth knowing: with the i-quant models, the CPU rounds experts differently when it computes one token than when it computes several, so the number of tokens sharing an expert depends on the drafts. The repository documents the result, that the same prompt at temperature 0 can end in a different but equally good answer, along with the setting that removes the dependency (STRATA_IQ_MT_MIN=1, at 1-3% decode cost on IQ3_S).
Quality is a sizing decision
The 94 tokens per second headline belongs to the 2-bit file. The project's own quality ordering puts Q2_0 at "good" and IQ3_S at "best: matches the full model on the published tests", with 53 tokens per second and a 54.8 GB memory requirement instead of 37.6 GB. Nothing in the published tables claims the fastest file is the best one, and the trade is a quantization decision rather than a cost of the offload itself.
The optional KV cache settings carry measurable quality costs of their own. A 4-bit KV cache halves the context's memory with a Hadamard rotation before rounding, runs about 4% faster at 128K, and raises perplexity by 8-12% on long documents while still passing the needle tests; 8-bit remains the default. A hybrid setting that keeps keys at 8 bits and stores values as rotated 4-bit cuts KV memory by 23%, and on an RTX 3090 running the Coder at a 198K context it moved output from 85 to 99 tokens per second with the same needle results.
Two variants sit outside the quantization ladder. The Coder keeps 256 of each layer's 512 experts, chosen on code, agentic and vision data: its authors report 91.3% of the full model's SWE-bench Verified and 98.7% of LiveCodeBench v6, it needs only 32 GB of RAM, and it is weaker outside code, including CJK text. Swift 1.5 is a fine-tune trained to think for fewer tokens: its authors claim 63% fewer thinking tokens and 1.8x sooner answers, and the repository's own eight-question check had both models correct 8/8, with Swift spending 1,234 output tokens against the original's 2,682.
At the top of the ladder, the experimental 4-bit UD-Q4_K_XL is the closest to full quality and the slowest usable configuration: a 111 GB download whose 77 GB of experts do not fit in RAM, so the engine reads most of them from the SSD and writes 7-8.5 tokens per second on a 64 GB PC. The repository's comparison against llama.cpp on that same file is about agreement rather than speed: 97.5-99% of positions chose the same token on short greedy answers, 90-91% after a 16K prompt.
Reproducibility so far
Every number above is the project's own measurement, which the repository states plainly: the other-GPU tables are labelled estimates with a ±20% band, and the maintainer's own speed report names the engine version, the settings and the prompt for each cell. Independent numbers exist but are thin. The one community benchmark committed to the repository ran an RTX 5090 with a 131K context on IQ2_XS and measured 179.4 tokens/s decode at a 4K prompt, 175.7 at 32K and 165.0 at 128K, with expert-cache hit rates between 97.8% and 99.7% and a host RAM peak of 45.54 GiB. Its author also writes that no same-workload baseline on another engine or GPU was run, and lists long outputs, sampled decoding, thinking, vision, tool use, concurrency and sustained thermal load as untested.
A field report on an AMD RX 7600 XT covers interactive rather than standardized use: 146 requests, 38,490 generated tokens, median 43.3 tokens/s on warm follow-ups, 82.7% cache hits, 76.8% draft acceptance and 13.4 GB/s host-to-device transfers. The AMD backend itself is described as an opt-in source build limited to one GPU, with no images and no calibration, and the RX 9070 XT figures come from the maintainer's own machine.
Read the tables in the repository rather than the paper's abstract for current speed. The abstract reports 45-95 tokens/s and prompts read at 285-571 tokens/s; the tables show prompts at 536-2,653 tokens/s because engine 0.1.13 roughly doubled prompt speed after the paper's numbers were taken. The commit history is also busy enough that a version number matters more than usual here, with 0.1.26, 0.1.30 and 0.1.36 adding prefill kernels, low-RAM modes and decode changes within days of each other, and 360 open issues at the start of October covering everything from HIP crashes on RDNA3 to out-of-memory starts on cards below the VRAM floor.
Whether Strata is twice or six times faster than llama.cpp on the same file is unresolved. No controlled head-to-head with a pinned configuration has been published, and llama.cpp's own MoE offload path is the baseline any such comparison has to hold everything else constant against. What the byte census does establish is narrower and more useful: for this model, on this class of machine, the working set per token is 0.66 GB of experts plus 3.5 GB of dense weights, and an engine that plans around those two numbers will beat one that treats a 66 GB file as a single blob to be paged. If you are considering it, check RAM before the card: 32 GB buys the Coder, 64 GB buys every size, and 12 GB of VRAM is the floor rather than the target.
References
- Niko1221/Strata — repository and README
- Strata — the details (docs/DETAILS.md)
- Running a 125-Billion-Parameter Model on a Normal PC (Strata paper, PDF)
- Community benchmark: RTX 5090, IQ2_XS, 131K context
- Strata — experimental AMD HIP backend (docs/AMD_HIP.md)
- QwenLM/Qwen3.8-Flash-Next — model repository
- ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF — quantizations
- GitHub REST API — repository statistics for Niko1221/Strata