I began this reading program to understand LLM hosting performance, but model papers alone do not explain a deployed system. Generation speed depends on where weights and intermediate state reside, how the runtime schedules work, and which operations it can overlap. Useful task performance adds retrieval, context management, tools, and the model’s ability to recover from mistakes. A faster decoder is valuable, but it does not repair an agent that loses the purpose of its task.

The difficulty is deciding what to read without turning the project into an unbounded bibliography. I want a small foundation that explains recurring problems, followed by extensions chosen for the system being built. In this sense, a canon is a working set of explanations. A paper earns its place by making a mechanism understandable or by changing how I would measure and build a system; inclusion does not imply that it originated every technique it discusses.

For a first pass, I would read eight works: the original Transformer paper, multi-query attention, FlashAttention, PagedAttention, GPTQ, speculative decoding, retrieval-augmented generation, and ReAct. The sequence moves from the model’s computation to its execution, then to the application that uses it. The additional readings below answer questions those eight leave open. They can wait until a project makes the question important.

Establish what the model stores and computes

Attention Is All You Need provides the architectural starting point. Its attention and feed-forward blocks explain how representations pass through a Transformer, while its encoder-decoder structure helps distinguish the original design from decoder-only designs. The paper is useful both for its equations and for the dependencies those equations expose.

For serving, two stores need to remain distinct. Model weights describe the learned computation. The key-value cache retains attention state from earlier positions so that generation can reuse it. Loading a checkpoint and accommodating a long conversation are therefore different memory problems. A model can fit before any requests arrive and still leave too little room for the intended context or concurrency.

Shazeer’s Fast Transformer Decoding: One Write-Head Is All You Need makes cache traffic explicit. Multi-query attention shares keys and values across query heads, reducing the state that incremental decoding must access. Read it after the Transformer paper and ask what changes in the cache, what remains in the computation, and where the quality trade enters.

GQA is the natural extension. Grouped-query attention occupies an intermediate position between one shared key-value head and a separate pair for each query head. The relationship is more useful than memorizing a list of attention acronyms: a deployment decision now has a concrete state-size and quality trade to investigate.

Sparse models introduce another distinction. In Outrageously Large Neural Networks, a learned gate selects a small combination of experts for each example. This permits learned capacity to grow without activating all of it for every input. The original experiments used a different surrounding architecture from current Transformer MoEs, but the conditional-computation argument remains directly relevant.

Total parameters still occupy storage even when only some experts are active. Active parameters indicate part of the work, but do not specify memory transfers, routing overhead, or device communication. The DeepSeek-V2 report is a useful combined example because it discusses both sparse experts and Multi-head Latent Attention, which compresses KV state into a latent representation. These mechanisms reduce different costs.

At this stage I would draw a simple map of a selected model: weights, attention state, routed computation, and communication between devices. The purpose is to identify what must remain available, what each token accesses, and what grows with context. That map gives the serving papers a concrete system to explain.

Read runtime papers as accounts of scarce resources

FlashAttention explains why the number of arithmetic operations is an incomplete performance model. Its exact attention algorithm tiles work to reduce transfers between accelerator memory and on-chip storage. The optimization changes how attention is executed; it does not turn the persistent KV cache into a smaller representation.

This distinction matters when evaluating a long-context claim. Less temporary attention storage, fewer cache bytes, and faster access to those bytes can each help, but they are different interventions. A runtime may improve one while leaving another as the limiting resource.

Orca then moves the discussion from an individual operation to a stream of requests. Iteration-level scheduling allows the batch to change between generation steps. A short request need not remain attached to a longer one for the entire response, and newly arriving work need not wait for every current request to finish.

PagedAttention explains the memory-management side of serving. A request’s KV allocation changes as it generates text, and reserving large contiguous regions can waste capacity through fragmentation or duplication. The paper’s paging approach supports more flexible allocation and sharing. Its throughput results belong to the evaluated systems and workloads; the transferable lesson is the relationship between cache management and feasible batching.

These papers also separate latency from throughput. A scheduler can complete more aggregate work while making a particular interactive request wait longer. A single-user experiment and a saturated server experiment can consequently reach different conclusions about the same optimization.

SGLang is the next extension when requests share prefixes or an application branches through several generations. Its RadixAttention mechanism reuses KV state for common prefixes. Reuse avoids recomputing an unchanged prefix; it is different from summarizing the text or discarding old context.

For a document-analysis application, this suggests a useful experiment: compare an initial request against subsequent requests that reuse the same document prefix. Record the work actually processed and the time to the first output. Otherwise a warm-cache result can be mistaken for the rate at which the system ingests a new document.

Separate numerical compression from efficient execution

Quantization reduces the width used to represent numerical values, but the implementation must still compute with that representation. A small checkpoint may require unpacking, scales, or conversions that change its execution cost. Weight storage, activation precision, and cache precision should therefore be recorded separately.

GPTQ is my first reading in this track because it develops post-training weight quantization using approximate second-order information. Its central question is how to reduce stored width while controlling the error introduced into an already trained model.

Two extensions explain why rounding every value in the same way is insufficient. LLM.int8() isolates important outlier features into a higher-precision computation, while AWQ uses activation statistics to identify and protect salient weight channels through scaling. Their mechanisms differ, but both connect numerical error to the behavior of the network.

KIVI belongs here when cache capacity becomes the problem. It studies different quantization arrangements for keys and values, rather than treating the cache as another copy of the weight matrix. Compressing weights and compressing request state have different consequences for concurrency and context.

The practical comparison needs both timing and quality. If a lower-bit representation moves a model entirely onto the accelerator, its performance benefit includes a placement change. If it introduces errors on the task being performed, preserving an average language-model score is not enough. These papers provide methods to study; they do not certify every local conversion bearing the same nominal bit count.

Understand how speculation changes the unit of work

Leviathan, Kalman, and Matias’s Fast Inference from Transformers via Speculative Decoding explains the draft-and-verify approach. A cheaper approximation proposes several tokens, and the target model checks them together. The paper’s sampling procedure preserves the target distribution while allowing useful work to proceed in parallel.

The important measurement is the time required to produce accepted output. High acceptance can still yield little acceleration if drafting and verification are expensive. Conversely, a modest accepted block can help when it avoids costly serial target invocations. An acceptance percentage and a token rate describe different parts of that result.

Better & Faster Large Language Models via Multi-token Prediction extends the reading into training. It predicts several future tokens through additional heads over a shared trunk. This makes multi-token prediction both an architectural capability and a serving opportunity: the trained model still needs a runtime that turns proposals into efficiently verified output.

I would read newer drafting methods only after this distinction is clear. The questions remain concrete: how proposals are produced, how they are checked, what distribution the procedure implements, and which cost it reduces under the intended batch and context. A longer draft window is a parameter to measure, rather than an optimization simply because it is longer.

Connect the decoder to the work it performs

The first application reading is Lewis and colleagues’ Retrieval-Augmented Generation. It combines learned model parameters with retrieved external memory. This changes where an answer’s evidence can come from and how that evidence can be updated, while introducing retrieval as another possible source of failure.

Suppose an application answers a question about a hardware manual. A poor answer could result from retrieving the wrong section, omitting an important qualification, or misinterpreting the correct passage. Increasing decoder speed addresses none of those directly. The evaluation must preserve enough of the retrieval and generation record to distinguish them.

Lost in the Middle is a useful extension because its experiments separate accepting a long context from reliably using information inside it. The tested models’ performance changed with the location of relevant information. That historical result is a reason to vary evidence placement in a current evaluation, rather than assume a documented context limit guarantees dependable recall.

ReAct adds an action-and-observation loop. The model can act on an environment, inspect what happened, and revise its next step. This makes the interface and the environment part of the system being evaluated. A failed command can supply useful evidence if the agent recognizes and uses the failure; a successful command can still advance an incorrect plan.

Once external material influences privileged actions, security becomes part of this same application problem. Greshake and colleagues’ indirect prompt-injection paper demonstrates how attacker-controlled retrieved data can redirect an LLM application. The resulting boundary concerns authority: which actions the application permits and how it checks them. Formatting a passage as data does not independently enforce that boundary.

Protocols and tool APIs belong in the implementation reading for the selected application. I would use their actual specifications, with a recorded version, rather than infer semantics from superficially compatible endpoints. Their purpose in this program is to explain the interface the model operates through, not to add another complete protocol survey.

Return to training when the system leaves questions unanswered

Serving papers explain the execution of a checkpoint, but not everything about the capability that checkpoint contains. Hoffmann and colleagues’ Training Compute-Optimal Large Language Models connects model size with the training-token budget. Its optimum concerns the training problem studied; a heavily used deployment introduces an additional lifetime inference cost.

Hinton, Vinyals, and Dean’s distillation paper explains how a more cumbersome teacher can supervise a smaller deployable model. The FineWeb report provides a complementary account of dataset filtering and deduplication through documented experiments. Together, these readings prevent architecture alone from becoming the explanation for every capability difference.

Current model reports can then show how several ideas were combined in a release. I would select reports relevant to the actual deployment and keep them separate from the stable foundation. A recent benchmark improvement justifies investigating a model; it does not by itself justify enlarging the core reading program.

The program has served its purpose when a performance result becomes easier to explain. I want to know whether the limiting resource is weight movement, attention state, arithmetic, scheduling, or an application that repeatedly does the wrong work. Each proposed change should identify the resource it saves, the quality it preserves, and the limitation likely to become visible next. That is a more useful outcome than a long list of papers I can recognize but cannot apply.