A local model can match a hosted model on a public aggregate and still be a different tool in practice. The score reflects a task distribution, configuration, harness, and scoring rule. My experience also includes how the system interprets a request, handles corrections, preserves existing work, and recovers after an error. Quantization and execution software add further differences between the published evaluation and the system I actually use.

This makes benchmarks useful for choosing candidates, but insufficient for choosing a complete deployment. I want to know which model and runtime can perform my technical work with acceptable latency and intervention. Loading the largest checkpoint that fits answers a capacity question. It does not establish that the resulting system is the most useful one.

What equal scores establish

An aggregate combines several measurements according to the evaluator’s chosen weights. Two equal scores can result from similar performance throughout the suite, or from substantially different strengths and weaknesses that happen to average to the same value. The latter matters when one of the weak areas is central to the intended workload.

Consider a hypothetical pair of agents that both complete 70% of a repository-task suite. One might perform better on localized fixes and worse on changes crossing several modules. The other might have the opposite pattern. The aggregate establishes equal success on that weighted collection; it does not establish interchangeability for a project dominated by either task type.

The evaluation environment also matters. The original SWE-bench paper made software maintenance a concrete test by pairing repository states with real GitHub issues and evaluating proposed changes. This is a substantial improvement over judging isolated code snippets. It still defines a particular task population and a procedure for assessing success.

A result from a later variant needs that variant’s identity, model configuration, tool harness, and resource budget. A score should not acquire a broader meaning merely because its benchmark name resembles another one. Changes to task selection, tests, effort, or available tools can change what is being measured.

Some differences also depend on how well the task is specified. In its February–March 2026 assessment, METR reported that the public models tested performed worse on the messier subset of its time-horizon tasks. The report also describes saturation that limited this comparison for stronger shared models. The result supports examining ambiguity as a workload feature; it does not establish a permanent gap for every model or professional task.

For my purposes, a benchmark is therefore a sample of evidence. Its value depends on the relationship between the sampled work and the work I intend to perform. A high score is a reason to investigate a candidate, while a small score difference needs component results and uncertainty information before it becomes a useful ordering.

The path to a correct result has a cost

The final state is important, but it does not capture every cost of producing it. An agent might eventually repair a bug after changing unrelated files, misreading a failing test, and requiring several corrections. Another might reach the same result with a smaller patch and little intervention. A scoring rule focused on the final tests could assign them the same outcome.

A hypothetical driver change illustrates the difference. The request is to support a new device identifier while preserving the behavior of existing devices. One agent adds the identifier, traces the initialization path, and checks whether the new device needs a different configuration. Another rewrites initialization, changes shared behavior, and only later restores compatibility after the user identifies a regression. Even if both final patches pass the available tests, the second path consumes review time and creates more opportunity for unnoticed damage.

That example does not imply that every larger patch is worse. A correct implementation may require a substantial change, and a minimal patch can conceal an incomplete understanding. The useful evaluation asks whether the work performed was necessary, whether existing behavior was preserved, and whether the agent understood the evidence produced by its tools.

Failure behavior matters for the same reason. A system that recognizes an unresolved assumption can leave a clear next step. A system that treats the assumption as established can carry it into later edits or analysis. Counting both as failures records the endpoint but leaves out the different recovery costs.

It is tempting to explain long-task reliability by multiplying a small per-step error probability across many decisions. That calculation requires assumptions that rarely hold in an agent workflow. Decisions are related, errors can propagate, and later observations can repair earlier mistakes. I would measure complete-task success and recovery directly rather than infer them from a sequence of supposedly independent turns.

Compare the application that actually runs

A hosted assistant includes more than its model weights. System instructions, tool interfaces, context handling, reasoning settings, and serving behavior affect its output. A local checkpoint also operates inside a surrounding system, with its own template, sampler, quantization, backend, and cache policy. Comparing the products does not automatically identify which layer caused a difference.

For example, a repository agent that receives a concise test failure and an agent that receives thousands of lines of unfiltered output may be operating with the same model but different effective evidence. A context policy that drops a requirement during compaction can create a failure that resembles poor model judgment. Conversely, a better checkpoint may remain useful despite an inefficient interface.

There are consequently two legitimate comparisons. One asks which complete system performs the work better under practical conditions. The other tries to isolate a mechanism by holding the surrounding system constant. The first supports a deployment choice; the second supports an explanation. Conflating them can produce a correct preference and an incorrect diagnosis.

Subjective preference remains useful evidence, provided its scope is clear. I have found compact Qwen models more useful than their size might suggest, and I have observed full GLM-5.3 running materially faster than Kimi K3 on Galactus. Those observations influence which candidates I would test first. They do not establish a general ranking, because task selection and the local configurations remain part of the result.

Quantization changes the evaluated system

An open checkpoint’s published quality may have been measured at a different numerical precision from the local conversion. Reducing stored width can make deployment possible, but it also changes the values used in computation. The relevant question is how the selected representation affects the intended task, rather than whether a family is generally described as tolerant of four bits.

Post-training quantization and quantization-aware training provide different kinds of evidence. GPTQ develops a method for quantizing an already trained model while controlling error. Quantization-aware training incorporates a particular numerical representation or error process into training itself.

The GLM-5 report, for example, documents INT4 quantization-aware training during supervised fine-tuning and a kernel shared with offline quantization. The Kimi K3 card documents MXFP4 weights and MXFP8 activations, with quantization-aware training beginning at the supervised fine-tuning stage. These are model-specific accounts of particular formats.

They do not certify every community conversion described as four-bit. Representations can differ in scales, rounding, tensor treatment, and activation precision. Nor does a technique documented for one release establish an unchanged recipe for every later member of its family.

For a manageable checkpoint, I would compare the intended local format with a higher-precision reference on representative problems. The comparison should examine unsupported inferences and difficult cases, not merely whether the answers retain a similar style. A rare analytical error can matter more than a small average difference when the task requires an argument that another expert will inspect.

Precision can also change placement. If a compressed checkpoint fits entirely on the accelerator while a wider version requires host work, a measured speed improvement includes both the numerical format and the new execution layout. That is a useful practical gain, but an explanation should retain both causes.

Capacity and responsiveness are separate constraints

Galactus has 2 TB of DDR4 and four Radeon Pro V620 GPUs with 128 GB of aggregate VRAM. The host-memory capacity expands the set of checkpoints that can remain loaded. It does not make those checkpoints equally responsive, or make four GPUs behave as a single uniform memory pool.

A deployment must accommodate weights, request state, temporary buffers, and the runtime’s own allocations in the appropriate memory pools. Keeping several models resident also creates competition for resources when they execute concurrently. Availability avoids some loading delay; it does not provide each resident model with independent memory bandwidth.

CPU-resident experts introduce another distinction. A runtime can execute expert computation on the CPU and exchange activations with the GPU, rather than stream every selected expert’s weights to the GPU. These paths have different costs. Host bandwidth, CPU arithmetic, device transfers, attention state, and synchronization can each affect the time a token takes.

Active parameter counts help describe computation, but do not directly state the number of bytes crossing DDR4 on every token. Placement, representation, reuse, and verification batches change that relationship. An architecture with less nominal active work still needs an implementation that realizes the intended advantage.

I encountered a simpler example of runtime sensitivity on a Ryzen AI Max+ 395. Qwen3.8-27B produced roughly three tokens/s in Jan using the selected HIP path, while Vulkan performed much better. Jan was installed as a Flatpak. This identifies the backend and runtime path as a strong suspect, but does not establish that the sandbox caused the problem. Kernel behavior, a fallback, packaging, or configuration could also explain the difference.

The practical choice can be straightforward before the diagnosis is complete: use the path that performs adequately, then inspect logs and matched configurations to explain why the alternative failed. Treating the slow path as proof of inadequate model quality or hardware would collapse several distinct questions into one.

A local stack needs evidence for each role

My initial proposal is to separate inexpensive iteration, longer tool work, and difficult analysis. These roles need not correspond to three permanently resident checkpoints. One model may cover several roles, and changing its effort or tools may sometimes be more useful than handing the task to another model.

The September 13 proposal assigned the following candidates to those roles. These assignments remain unvalidated against my workload; the table records a deployment hypothesis rather than a ranking of model quality.

Proposed role Initial candidate
Fast default Qwen3.8-27B
Efficient agent Qwen3.8-Flash-Next
Serious default GLM-5.3-Flash
Maximum local escalation GLM-5.3
Alternative Kimi K3 or DeepSeek

A compact default earns its role by completing frequent requests quickly enough to support iteration. An agent candidate earns its role by using tools, preserving constraints, and recovering with less intervention on longer tasks. An escalation model earns its additional cost by improving the difficult outcomes, rather than merely producing longer answers.

This also determines the routing signals. A model’s expressed confidence is insufficient by itself. A failed check, an unresolved requirement, a task known to need extensive context, or a repeated inability to make progress provides a more concrete reason to escalate. The handoff should preserve the evidence and unresolved questions; otherwise routing creates another opportunity to lose the task.

A second model can supply a useful alternative interpretation, but agreement does not verify a result. Different models can share training data and make related mistakes. The disagreement is a lead to investigate, while a checked source, reproduced calculation, or executable result provides firmer evidence.

Evaluate replacement on completed work

I would begin with a modest collection of previously completed technical tasks: debugging, repository changes, reverse engineering, document analysis, and research synthesis. Each task should have a preserved input record and an outcome that can be checked. Identical starting records make comparisons more meaningful, while blinding model identity during review reduces the influence of familiarity and presentation.

The evaluation needs to keep several results separate:

Result What it establishes
Correctness and completeness Whether the system met the actual requirements
Evidence and unsupported assumptions Whether its conclusions followed from the available record
Preservation and recovery What it changed unnecessarily and how it handled failure
Human intervention How much correction and review were required
End-to-end latency How long useful completion took, including tools and waiting
Runtime and resource use Which configuration produced the result and at what practical cost

Budgets should be documented as well. Equal token limits do not imply equal computation or equal cost across models, but an unrecorded budget makes the comparison difficult to interpret. A deployment trial can use the settings intended for ordinary operation, while a separate controlled test changes one variable at a time to investigate the mechanism.

A few dozen selected tasks would provide useful evidence about my workload without establishing equivalence across every professional domain. I would retain some tasks outside the selection process and revisit the comparison when the model, conversion, or runtime changes materially.

The replacement decision then becomes concrete. An existing model loses its role when another complete system performs the intended work equally well or better at acceptable latency, intervention, and cost. Public scores help identify that possibility. Local measurements and inspected task outcomes determine whether it has actually been achieved.