An inference server can report a good draft acceptance rate and still generate tokens slowly. It can also report fast prompt processing after reusing most of the prompt. Neither result is contradictory, but each becomes misleading if I treat the displayed number as a measurement of work the server did not perform.

A GLM-5.3 request on Galactus provides a useful example. The server reported 5.549 output tokens/s, accepted roughly two-thirds of its draft proposals, and processed a new prompt suffix at 52.84 tokens/s. These figures describe different parts of the request. Reading them together explains more than any one of them, although the retained record does not contain enough configuration detail to turn the request into a reproducible benchmark.

What the request recorded

The request appeared in my September 13 notes with the following values. They are the retained measurements; the calculations below are derived from them.

Quantity Recorded value
Prompt tokens 11,928
Completion tokens 3,251
predicted_ms 585,684.939 ms
Reported generation throughput 5.549 tokens/s
Proposed draft tokens 2,780
Accepted draft tokens 1,860
Cached prompt tokens 8,560
Newly evaluated prompt tokens 3,368
Throughput on the new prompt suffix 52.84 tokens/s

The generation time is approximately 585.7 seconds, or 9.76 minutes. A misplaced decimal would turn it into 58.568 seconds and create an apparent tenfold improvement. Dividing the completion count by the recorded time gives:

3,251 / 585.684939 = 5.5508 tokens/s

That is close to the reported 5.549 tokens/s. The small discrepancy does not affect this interpretation, but I cannot identify its cause from the transcription alone. The original response and the definitions of the counters in the actual server build would be needed to reconcile it exactly.

This also illustrates why I retain the underlying counts and duration. A rate hides the numerator and timing boundary that produced it. Current llama.cpp server documentation distinguishes cache reuse, newly evaluated prompt tokens, and generation timing, but current documentation cannot establish every detail of a historical binary whose revision was not retained.

Acceptance describes useful proposals

Multi-token prediction (MTP) supplies proposals that a target model can verify together. When several proposals survive verification, one cycle can contribute several output tokens. The general attraction of speculative decoding is to obtain more useful output from a target invocation; drafting and verification still have to be paid for. The original speculative decoding paper explains this approach and the conditions under which its algorithm preserves the target distribution. That guarantee should not be assigned automatically to every implementation carrying an MTP label.

The request’s two acceptance ratios answer different questions:

Accepted / proposed drafts = 1,860 / 2,780 = 66.9%
Accepted drafts / completion = 1,860 / 3,251 = 57.2%

The first is the fraction of proposed draft tokens accepted. The second is the fraction of output supplied by accepted drafts. About two-thirds of proposals were useful, and those proposals supplied more than half the completion. This establishes that the draft mechanism contributed substantial output.

It does not establish a 66.9% speedup, or any other speedup. Rejected proposals cost work, and checking a batch of positions can cost more than an ordinary single-token pass. Even useful proposals can fail to save time if their production and verification are expensive. A speedup requires a comparable MTP-off run with the same model, prompt, sampling conditions, placement, and build. That comparison is absent from this request record.

A conditional estimate of verification cycles

Output counts can also suggest how a speculative cycle behaved, provided I state the accounting assumption. Suppose each cycle contributes exactly one non-draft output token, and the first and final cycles introduce no exceptions. Under that convention, subtracting accepted drafts from the completion gives:

Estimated cycles = 3,251 - 1,860 = 1,391
Output per cycle = 3,251 / 1,391 = 2.337
Draft proposals per cycle = 2,780 / 1,391 = 1.999
Cycles per second = 1,391 / 585.684939 = 2.375

The near-two proposal average is consistent with a maximum draft length of two. It does not prove that two was the configured maximum: a larger limit could also produce an average near two, and the assumed convention could differ from the implementation’s counters. An actual cycle count and the launch configuration would settle those questions.

The estimated 2.375 cycles/s is also not a recovered non-MTP decoder rate. A speculative target invocation may evaluate several positions, and the cycle includes draft work, verification, scheduling, and synchronization. Dividing output by an estimated number of cycles describes the output obtained from those cycles; it does not reveal how fast the same target would run without speculation.

Acceptance alone loses another useful detail. A run with regularly accepted two-token drafts can behave differently from one that alternates longer accepted sequences with many failures, even if their overall acceptance percentages match. The distribution of accepted lengths matters because it determines how often the server pays the cycle cost for little output.

Cache reuse changes the prefill question

The prompt accounting reconciles exactly:

8,560 cached tokens + 3,368 newly evaluated tokens = 11,928 prompt tokens

Only the 3,368-token suffix needed new evaluation. The reported 52.84 tokens/s therefore describes that suffix in the presence of an existing cached prefix. It is not a cold-ingestion measurement for all 11,928 tokens. Using the full prompt length as the numerator for a suffix-processing rate would count cached tokens as newly evaluated work.

Caching is useful precisely because it avoids repeating work. In an interactive conversation, reusing a long prefix can reduce the time before generation substantially. For a performance comparison, however, a warm suffix evaluation and a cold full-prompt evaluation need different labels. Otherwise the benefit of reuse can be mistaken for a faster prefill implementation.

The request began with roughly 11.9K prompt tokens and added 3,251 completion tokens, putting the recorded total near 15.2K. Its average generation rate includes that progression; it is not a measurement at one fixed context length.

What the slower output suggests

I had also observed approximately 6.8 output tokens/s under more favorable, shorter-context conditions. The 5.549-token/s request is about 18.4% slower. That is an observation associated with a longer context, not a controlled estimate of the context penalty. The prompts, output, draft acceptance sequence, and other runtime conditions may differ.

Even if acceptance remained stable as throughput fell, I could not assign the slowdown entirely to attention. Draft production, multi-position verification, scheduling, transfers, and the accepted-length distribution can change without producing a large change in the aggregate acceptance percentage. Stable acceptance makes one explanation less likely—a broad collapse in proposal quality—but leaves several costs unresolved.

Galactus has a large host-memory pool and four V620 GPUs, so placement is another plausible influence. CPU-resident experts may execute on the CPU or participate in a GPU offload path; their presence in RAM does not prove that every expert byte crosses PCIe. Likewise, aggregate VRAM does not behave as one uniform memory pool. The execution path determines which bandwidth and synchronization costs matter.

This request does not retain the complete quantization, binary revision, launch command, or time breakdown needed to isolate those costs. I can use it to demonstrate correct accounting, but I cannot use it to rank GLM-5.3 against another model or recover a hardware ceiling.

Reading the result as a whole

The useful conclusion is quite specific. GLM-5.3 produced 3,251 tokens in about 585.7 seconds, with accepted drafts supplying 57.2% of the output and a cached prefix avoiding evaluation of 8,560 prompt tokens. Under an explicit cycle assumption, the counts are consistent with about 2.34 output tokens per cycle. None of those observations establishes net MTP acceleration on its own.

A performance result becomes easier to understand when I keep accepted output, elapsed cost, and actual prompt work separate. Draft acceptance explains how proposals contributed to the completion. Generation timing explains what that completion cost. Cache accounting explains what the server reused before generation began. Together, they describe the request without granting the counters more meaning than the retained evidence supports.