Galactus runs models whose expert banks are much larger than its GPU memory. The machine combines an AMD EPYC 7713, 2 TB of system RAM, and four Radeon Pro V620s with 32 GB of nominal VRAM each. This gives it enough capacity to hold models occupying hundreds of gigabytes, while the GPUs provide a faster memory tier for the parts of the model that can use it.

Capacity explains why the models load. It does not explain their speed. A mixture-of-experts model selects a subset of its routed experts for each token, so two models of similar total size can require quite different amounts of work to generate the next token. The location and representation of those selected weights matter as much as their count.

My working interpretation is that CPU-resident expert traffic accounts for a substantial part of ordinary decode on Galactus. The host’s corrected STREAM results cluster around 148–151 GB/s, but STREAM measures a streaming benchmark rather than a quantized expert kernel. I therefore use that bandwidth as a reference for a cost model, then compare the result with observed token time. This is useful for deciding which changes deserve attention, provided the inferred costs remain distinct from measurements.

The August comparison

The August 15–16 common baseline gives the clearest comparison in the notebook. All five models ran on the 2 TB population after the failed DIMM had been replaced, using ROCm and llama.cpp build 3653e6d6d (10326) with the stock scheduler. The tests used -ngl 99, the exps=CPU tensor override, flash attention, f16 KV, and batch and microbatch sizes of 8192. Prompt processing used 8192 tokens; the separate generation benchmark used 128 tokens.

The override kept routed-expert weights resident in system RAM. During ordinary single-token decode, CPU kernels processed the selected experts, without first copying each expert to a V620. The remaining offloaded path ran across the GPUs. Residency does not establish the computation path for every phase: larger prefill batches can activate operation offload and temporarily stream CPU-resident weights to a GPU. The August command did not disable that facility. Attention, routing, shared experts, and other tensors should be assigned according to the actual allocation and backend support; they should not all be charged to host traffic merely because they contribute to the model’s active parameter count.

Model and export CPU threads Prefill, tokens/s Ordinary decode, tokens/s
MiniMax M2.7, UD-Q5_K_M 64 418.83 ± 24.11 15.18 ± 0.18
Qwen3.5-397B-A17B, UD-Q6_K_XL 64 249.61 ± 19.78 9.37 ± 0.16
DeepSeek-V4-Flash-0731, UD-Q8_K_XL with MXFP4 experts 64 143.54 ± 1.64 10.34 ± 0.10
GLM-5.2, UD-Q4_K_XL 32 95.99 ± 3.36 5.30 ± 0.00
Kimi K2.6, UD-Q8_K_XL with native-INT4 MoE 64 94.23 ± 4.45 5.79 ± 0.01

These are the means and standard deviations reported by llama-bench over two repetitions. A displayed deviation of 0.00 reflects rounding, not perfectly invariant performance. Placement and benchmark dimensions are comparable across rows, while architecture, export, and GLM’s thread count differ. The table describes these configurations rather than isolating an architectural effect or ranking model quality.

Kimi’s 553.71 GiB export was larger than GLM’s 435.19 GiB export, yet it decoded somewhat faster. This is a concrete example of why total file size is insufficient as a performance predictor. The active routed path and the work surrounding it are more useful quantities.

Turning throughput into a cost estimate

For ordinary autoregressive decode, I use the approximation T = C + S, where T is seconds per output token. The streaming term is S = B / W: B is the estimated host-resident expert bytes needed for that token, and W is an assumed effective bandwidth. The residual is C = T - S.

This separates a traffic-dependent term from everything that the estimate leaves unexplained. Those other costs can include attention, shared-expert computation, routing, CPU arithmetic, graph execution, transfers, and synchronization. They can also include inefficiency inside the expert kernel itself. If the kernel achieves less bandwidth than the reference assumes, some of its cost appears in C.

The existing estimates are:

Model Measured time per token Estimated expert-streaming time Remaining time in the approximation
Kimi K2.6 about 173 ms about 101 ms about 72 ms
GLM-5.2 about 189 ms about 91 ms about 98 ms
DeepSeek-V4-Flash about 97 ms about 29 ms about 68 ms
Qwen3.5 about 107 ms about 81–87 ms about 20–26 ms

The measured times are reciprocals of the August decode rates. The streaming terms retain my earlier working estimates; they were not measured by the common-baseline run. The GLM notes and DeepSeek notes record their traffic budgets, while the common-baseline entry records Kimi’s estimate and Qwen’s earlier residual range. Subtracting the estimated terms produces the residuals here. Earlier rounded GLM terms of 96 ms plus 91 ms predicted 5.35 tokens/s, close to the observed 5.30. That agreement is useful as a consistency check, but subtraction cannot independently validate the assumed bandwidth.

Kimi’s estimate, for example, budgets approximately 15 GB of INT4 routed weights per token. Servicing that traffic at about 150 GB/s takes roughly 100 ms. The remaining 72 ms does not establish a 72 ms GPU bottleneck. Profiling would be needed to allocate it to particular operations, and overlapping execution could make a strictly additive model incomplete.

The practical difference between the rows is still informative. Qwen spends most of its modeled token time in the traffic term, so reducing host bytes has a relatively direct path to improving throughput. DeepSeek has a much smaller estimated traffic term, leaving more unexplained time to investigate. A large residual identifies room for investigation; it does not identify a particular optimization or imply that all of that time can be removed.

What a smaller quantization can change

Quantization can reduce both the memory needed to hold a model and the bytes needed to process its selected experts. These are related benefits, but their consequences differ on Galactus. Saving capacity in an already sufficient bank of system RAM does not automatically improve generation. Reducing the CPU-resident expert bytes addresses the traffic term directly, while shrinking other tensors may instead help them fit on a GPU.

The following calculation illustrates the distinction. Take the earlier rounded GLM budget, C = 96 ms and S = 91 ms, for a total of 187 ms or 5.35 tokens/s. Suppose a different export reduces routed-expert traffic by 25%, while kernel efficiency and all remaining costs stay unchanged. The new streaming term is about 68 ms, giving 164 ms per token or 6.10 tokens/s. Throughput rises by approximately 14%, although the streaming portion alone becomes about one-third faster.

This is an assumed traffic reduction, not a measured comparison between GLM exports. Its purpose is to show why applying the reduction to the entire token overstates the gain. If the new representation requires more expensive dequantization, changes tensor placement, or alters the workload, those effects also belong in the comparison.

The export name is insufficient to determine the traffic. A GGUF labeled Q8 can retain native low-bit experts, as the DeepSeek and Kimi rows demonstrate, while other tensors use higher precision. Dynamic quantizations can likewise preserve selected tensor classes at higher precision. The relevant quantity is the storage of the expert tensors actually used by the CPU, including scales, padding, and other representation overhead.

Quality must be assessed alongside that cost. A smaller representation is useful only if it preserves enough of the model’s behavior for the application. Distributional tests such as divergence and top-token agreement can reveal changes on their evaluation corpus, but they do not fully test coding, reasoning, or long-context behavior. I would therefore compare useful exports on both throughput and the tasks for which I intend to run the model. A faster export that no longer provides the needed capability has not solved the original problem.

Verification costs belong to accepted output

Speculative decoding proposes several candidate tokens and asks the target model to verify them together. This can spread some computation and synchronization over more output tokens. On a hybrid MoE system, however, the verification block still has to process the experts selected by its candidate positions.

The important quantity is therefore verification traffic per accepted token. If several positions use the same expert and the execution path reuses its weights, a block can reduce traffic per output token. If the positions select different experts, their union requires more weights. Rejected candidates add work without adding accepted output, and the drafter has a cost of its own.

The August trials found useful shallow configurations. GLM-5.2 reached 6.6 ± 0.3 tokens/s with MTP at a maximum draft depth of two. DeepSeek-V4-Flash reached 14.1 ± 0.7 with DSpark at depth three; depth two was comparable within the observed spread. Deeper settings did not improve these trials. Against the ordinary-decode figures, the increases are approximately 25% and 36%.

These trials used llama-cli, a technical prompt, 256 generated tokens, and greedy decoding, while the ordinary baseline used llama-bench’s generation test. The two speculative repetitions differed by about 9%, despite identical output streams. Their quoted ± values summarize those two rates rather than confidence intervals. The results support the configurations on that workload, but they do not provide a precise matched estimate of the percentage gain or explain the decline at greater depth.

In particular, the fitted C term is not a measured pool of fully amortizable work. Nor is 1/S a universal ceiling on speculative throughput: that ceiling would require the host bytes per accepted token to remain unchanged. Weight reuse could lower that cost, while rejected candidates could raise it. The observed rates establish the benefit here; acceptance and verification traces would be needed to explain it more closely.

Prompt processing is a different phase

Prefill processes many prompt positions together. Tokens selecting the same expert can be grouped, allowing an expert matrix to serve multiple activations. This can amortize weight reads and increase arithmetic work per byte. The GLM scheduler investigation also records operation offload of host-resident experts during large batches. Matrix shapes, expert scheduling, kernel efficiency, and host-to-GPU transfers can consequently matter differently from ordinary single-request decode.

The baseline makes the distinction visible. Qwen processed the 8192-token prompt at about 250 tokens/s while decoding at 9.37. MiniMax reached about 419 versus 15.18, and GLM reached about 96 versus 5.30. A decode traffic estimate cannot explain these prefill rates by simply substituting a larger token count.

Batch width also competes with permanent tensor residency. Earlier partial-expert Qwen placements stopped fitting under the August configuration: the larger microbatch and f16 KV allocations consumed room that the earlier smaller benchmark had left available. Four cards provide 128 GB of nominal aggregate VRAM, but each card has its own allocation limit and must accommodate weights, state, workspace, and transfer buffers. A model’s total GPU budget is therefore insufficient to establish that a particular placement will fit.

This tradeoff affects the experience of using the machine. Faster decode can make a response arrive more quickly once generation begins, while a broad prompt-processing batch can reduce the initial wait. Optimizing one phase without budgeting for the other can produce a configuration that benchmarks well but does not serve the intended workload.

What the cost model contributes

Galactus’s design provides substantial capacity with a smaller GPU tier, and the August measurements show that this arrangement can generate useful output from very large models. Its costs remain model-specific. Host expert traffic explains an important portion of the ordinary-decode budget, while prefill and speculative verification change how that traffic is shared across positions.

The useful next decision follows from that distinction. When host reads dominate, fewer expert bytes, greater weight reuse, or more effective bandwidth address the modeled constraint. When substantial time remains unexplained, kernel and scheduling measurements have more to resolve. Keeping those cases separate makes the model useful for engineering decisions without asking it to predict an unmeasured model’s throughput.