Changing a setting and keeping it when a benchmark improves is a reasonable way to begin performance tuning. The difficulty is explaining the improvement: a setting may change the intended operation, expose a different code path, or interact with an unrecorded default. Without that explanation, it is hard to reproduce the result or decide what still warrants investigation.

I investigated inference on Galactus using bandwidth measurements, an approximate decode model, and controlled comparisons. The server has an EPYC 7713, eight channels of DDR4-2933, and four Radeon Pro V620 GPUs. The July GLM-5.2 and bandwidth tests used its original 1 TB memory population; the early August DeepSeek-V4-Flash tests followed its expansion to 2 TB. Those are separate measurement sets.

Both models use a mixture-of-experts architecture, selecting a subset of expert networks for each token. Many expert weights remain in system RAM because the full weights exceed GPU memory. Reading them contributes substantially to generation, or decode, time. Prompt processing, or prefill, has different opportunities to batch work, so I measured the phases separately.

1. Establish a hardware reference

I measured bandwidth with STREAM, which times Copy, Scale, Add, and Triad operations on large arrays. A thread sweep produced the best rates at 16 threads; additional threads eventually reduced bandwidth. Using every core was therefore unnecessary for this memory workload.

STREAM reports the bytes requested by its array operations. Actual memory traffic can be higher because a cached store may first read the cache line it will overwrite. For the July results, I assumed this write-allocation read for Scale, Add, and Triad, but non-temporal stores without it for Copy:

Operation Reported rate (GB/s) Traffic adjustment Estimated traffic rate (GB/s)
Copy 151.8 ×1 151.8
Scale 103.4 ×1.5 155.1
Add 111.9 ×4/3 149.3
Triad 112.5 ×4/3 150.0

The estimates cluster around 152 GB/s, which I used as a decode-model reference. Their agreement is a consistency check because they share traffic assumptions. Copy’s store behavior was inferred rather than verified in the compiled instructions; an assumed correction exceeding the theoretical peak would question the traffic model without identifying the instructions emitted.

Eight channels at 2,933 million transfers per second and eight bytes per transfer give a theoretical peak of 187.7 GB/s. The reference reached approximately 81% of that peak, suggesting a plausible starting point for software work. A much lower ratio would justify investigating channel population, memory clock, NUMA configuration, and benchmark setup, but would not identify the cause by itself. STREAM’s sequential accesses also differ from inference accesses, so the reference is not a guaranteed application rate.

2. Distinguish fitting from prediction

Exploratory benchmarks can reveal behavior before a model explains it. Once a plausible model exists, writing down an expected result makes the next comparison more informative: disagreement can expose an omitted cost or mistaken assumption.

For the GLM-5.2 hybrid configuration, I used:

time per token ≈ C + (estimated expert bytes per token ÷ bandwidth)

Tensor sizes and active-expert counts gave an estimated 13.77 GB per token, not a measurement of DRAM traffic. At 152 GB/s, reading that amount would take about 91 ms. Subtracting it from the observed decode time left roughly 90 ms. Adding the terms gives about 181 ms per token, or 5.5 tokens per second.

C was therefore a fitted residual that could include synchronization, dense computation, and other costs. It was not an independent GPU measurement, and it need not remain constant across models, placement, context, or builds. The later notebook comparison was:

Configuration Model estimate (tokens/s) Measured decode (tokens/s)
Hybrid, expert weights in system RAM 5.5 5.53
Fitter, with some expert layers resident on GPUs 6.2 6.01
CPU-only, with additional CPU work accounted for 3.9 3.87

This was a retrospective check, not three predictions preceding the measurements. The CPU-only estimate accounted for additional CPU work. Agreement within about 3% supported the approximation for these configurations without independently validating its fitted costs. The Session 3 derivation and Session 5 comparison preserve that chronology.

The project began with spreadsheet estimates and an assumed ±30% planning range, rather than a statistically calibrated confidence interval. Measurements revised the bandwidth assumption and weakened several proposed explanations for the remaining cost. Retaining the estimates alongside the results preserves how the explanation changed.

Speculative decoding changes the accounting. It drafts tokens and verifies them together, potentially producing several accepted tokens per target-model pass, as described in Fast Inference from Transformers via Speculative Decoding. MoE verification can reuse expert weights across positions, select additional experts, and incur work for rejected drafts. The relevant streaming cost is bytes read per accepted token. Holding that cost constant gives a conditional bound, not a universal ceiling.

Early runs reported about 31% improvement with GLM MTP and 45% with DeepSeek DSpark. The DeepSeek sweep remained provisional because acceptance rates were not captured and the fit setting changed. These observations identified useful working configurations without isolating each component’s contribution to the gain.

3. Sweep settings and check what actually ran

A baseline records the model and quantization, software revision, placement, prompt, context, decoding policy, and effective settings. Comparisons hold these conditions constant except for the variable being tested. Repeated runs and observed variation matter especially for small improvements.

Two GLM sweeps showed different responses. Decode peaked at 24 to 32 threads, well below the server’s 64 physical cores:

Threads Decode (tokens/s)
24–32 ~5.5
96 2.76
128 1.29

Higher counts extend into simultaneous multithreading. Contention and scheduling overhead plausibly contributed to the decline, although the curve did not separate them. Using all logical processors was a poor choice here.

The prefill sweep varied the physical microbatch, n_ubatch, with prompt length and logical batch held at 8192 tokens:

Microbatch (n_ubatch) Prefill at 8192 tokens (tokens/s)
512 25.90
1024 41.68
2048 62.64
4096 84.62
8192 104.97

Throughput rose about fourfold, making 8192 the best tested microbatch for this workload. The largest possible microbatch need not remain best after changing context length or placement.

An earlier setup error had concealed this effect. My llama-bench runs requested microbatches larger than 512 but left the prompt at its 512-token default. The benchmark capped the effective microbatch at 512, so the timings described what actually ran while failing to support the intended comparisons. Checking reported effective settings against requested settings exposed the error.

Other interactions made the execution path part of the finding. An operation-offload mode hurt throughput at small microbatches but helped at 8192. Two apparently similar expert-placement approaches used different paths, with one about 17% slower. Testing one could not substitute for testing the other.

The scheduler patch below targeted prefill. A changed decode result would still warrant checking matched conditions and possible regressions, even though decode was not its intended target. Faster execution also needs to preserve the intended computation, with any deliberate approximation identified separately.

4. Investigate the prefill scheduler

GLM-5.2 prefill reached 104.97 tokens/s with an 8192-token prompt and microbatch. The July tests used Unsloth UD-Q4_K_XL, expert weights in system RAM, and llama.cpp build 657e01125 (10001). A separate placement check at a 512-token prompt and microbatch counted 731, 133, 175, and 147 graph splits across the GPUs, totaling 1186. These were assignment counts, not measurements of GPU execution time.

Source inspection identified a possible constraint: reading selected-expert routing information from a GPU back to the host could force synchronization before scheduling continued. I tested the hypothesis that distributing work would provide little benefit while this dependency remained, first changing distribution alone and then adding a routing-readback bypass for large batches.

Configuration Prefill at 8192 tokens (tokens/s)
Stock scheduler 104.97 ± 0.53
Distribution change alone 105.71 ± 0.56
Distribution plus large-batch routing-readback bypass 119.36 ± 0.12

The ± values retain llama-bench’s mean-and-standard-deviation format. Stock and distribution-only commands used two repetitions, making these spreads limited run descriptions rather than confidence intervals.

Distribution balanced the split counts at 285, 300, 294, and 292, while throughput improved only about 0.7%, little relative to the reported variation. Adding the bypass raised throughput 13.7% over stock. The first combined run included a Vulkan build change; later attribution checks found approximately no gain from it. Rolling back a separate pipeline experiment subsequently retained 119.29 ± 0.19 tokens/s. The Session 9 record documents those checks.

Above a large-batch threshold, the bypass treated all experts as used, trading possible copies of unused experts for avoiding the routing synchronization. Its benefit depended on that balance, rather than a guarantee that every large prompt selected every expert.

The source inspection and comparisons supported investigating the readback dependency, but the bypass was not tested without the distribution change. The strongest measured result was therefore the combined patch’s improvement, rather than a uniquely identified contribution from synchronization alone.

Recorded prefill rates progressed from 37.63 to 119.36 tokens/s, about 3.2 times as fast, across different prompt lengths, builds, and settings. That history is not one controlled speedup. The fixed-prompt microbatch ladder and stock-versus-patched comparison better isolate individual effects.

5. Preserve negative and inconclusive results

An unsuccessful experiment needs its conditions and outcome recorded. A completed comparison without a gain, a regression, a crash, and a proposal that never reached its intended execution path answer different questions.

Thread polling and weight repacking completed without useful decode gains. Strict CPU pinning tied the unpinned result at best and hurt some settings. Row-based GPU splitting failed at model load; the pipeline experiment failed allocation and fell back to one copy. The tested vendor math library did not support the routed expert path. These outcomes explained why I stopped pursuing those configurations without measuring every proposed technique’s performance.

Platform checks had similarly specific meanings. The server exposed one NUMA node under NPS1, while host-versus-container STREAM agreement supported the absence of a material container penalty on the memory path. The huge-page snapshots, taken after inference exited, and a missing runtime trace left engagement and potential benefit unresolved.

The provisional DeepSeek sweep illustrates the same distinction: retained throughput remained an observation, while missing acceptance statistics and changing fit state limited its interpretation. A later experiment can revisit any of these ideas when the workload or implementation changes, using the earlier conditions and outcome to explain what the new comparison tests.

That record connects each result to the experiment that produced it. It preserves fitted explanations, controlled comparisons, and unresolved questions in a form that can guide the next investigation without turning assumptions into facts.


Commands, logs, model details, and chronology are retained in the performance-engineering notebook. Speculative comparisons mostly used a single technical prompt with greedy decoding; they do not establish performance across prompts, concurrent requests, or machines. The July 1 TB and August 2 TB results are not a matched hardware comparison.