· Computing, Research
Frontier LLMs for Expert Analysis: Evidence, Inference, and Inspectable Arguments
For expert technical work, particularly patent infringement analysis, I need more than a model that reaches a plausible answer. I need an argument that another person can inspect: what the record establishes, which steps depend on inference, what alternative interpretations remain possible, and how missing evidence limits the conclusion. A model can arrive at the right endpoint and still leave too much of that work for the reviewer to reconstruct.
Communication is therefore part of the analytical product. A short answer can be sufficient when the inference is simple and the evidence is direct. A more difficult problem may require the model to develop several dependencies before the conclusion becomes useful. The relevant distinction is between an argument whose support is visible and one whose support has merely been asserted.
Model quality depends on the work being requested
Claude Fable 5 and 5.1 have made this distinction concrete for me. In my non-coding work, I have not found them clearly superior to Claude Opus 4.6, and I have sometimes preferred 4.6 for infringement analysis. Fable’s explanations can feel compressed or stunted even when the conclusion appears intelligent. This is an observation about my workload and the responses I have received, rather than evidence that the newer models have generally regressed.
It does establish that a general benchmark improvement is insufficient reason for me to replace a known analytical baseline. An evaluation may measure relevant knowledge, reasoning, or the production of a deliverable without measuring how well a model handles my particular record. The model still needs to interpret technical language, characterize sources accurately, and preserve the distinction between a supported conclusion and a useful hypothesis.
Anthropic’s Fable 5.1 prompting guide documents denser prose and fewer paragraph breaks in some cases. This supports the limited proposition that communication behavior can change between releases. It does not independently establish my observation of abbreviated analysis, or explain why a particular response omitted a necessary step.
Longer output is also an inadequate remedy by itself. A model can repeat a conclusion in several forms, add background that never enters the inference, or produce an elaborate explanation around a false premise. What I need is enough development to make the argument’s dependencies visible.
Evidence must retain its scope
An expert analysis often begins with sources that describe different parts of a system. A manual may describe an intended configuration, source code may describe several supported configurations, and an observed trace may show one particular execution. These sources can reinforce one another, but they do not become interchangeable simply because each mentions the same feature.
For a technical claim limitation, I want the model to identify the implementation detail that matters and the passage or observation supporting it. If the source establishes only a capability, the analysis should not silently convert that into proof of use. If a test establishes behavior in one configuration, the analysis should not silently extend it to every configuration.
This is where visible argument development becomes useful. The reviewer needs to see why the cited fact matters, which part of the requirement it addresses, and what remains to connect that fact to the conclusion. A citation beside a paragraph provides a location to inspect; it does not by itself establish that the source supports every sentence in the paragraph.
The analysis also needs to preserve uncertainty at the point where it enters. A broad qualification at the end cannot repair several earlier paragraphs that treated an unresolved assumption as fact. The narrower wording should accompany the inference so that later conclusions inherit the correct limitation.
A hypothetical controller example
Consider a fictional storage-controller analysis. The technical requirement being examined is: every user-data block is encrypted before it is written to nonvolatile memory. This is a simplified requirement for illustrating the reasoning, not a conclusion about an actual product or patent.
Suppose the record contains three items. A product manual places an encryption engine ahead of flash storage in its write-path diagram. A firmware excerpt calls an encryption function before the storage-write function when a secure-mode flag is enabled, and otherwise passes the original buffer to the write function. A dump from one test unit contains no recognizable plaintext in the region examined, but the test record does not identify a known input block or the unit’s secure-mode setting.
A compressed answer might say that the manual, code, and dump demonstrate encryption before storage. That answer combines relevant evidence, but it conceals the main questions. The diagram describes an intended path. The code describes a conditional path. The dump shows an observation whose cause has not yet been isolated.
The strongest positive evidence is the order of operations in the encrypted branch. If that excerpt represents the tested firmware, and if secure mode was enabled, it supports the conclusion that the buffer on that path was encrypted before the write. The manual is consistent with this interpretation and helps explain the role of the encryption engine.
The code also exposes a limit. It contains a bypass, so the excerpt alone does not establish encryption of every user-data block. That limit might be resolved if another part of the firmware forces secure mode on for all relevant writes. It might instead show that encryption is optional. The record supplied so far does not distinguish those possibilities.
The dump is weaker evidence than it initially appears. An absence of recognizable plaintext is consistent with encryption, but also with compression, an unexamined data layout, or inspecting a region that never contained the chosen data. Without a known input and a reliable mapping to the stored block, the observation cannot isolate encryption as the explanation.
A useful analysis would therefore reach a narrower conclusion: the record documents an encrypted write path and an encryption-capable architecture, but has not established that the relevant unit routes every user-data block through that path. This conclusion preserves the positive evidence instead of replacing it with an undifferentiated statement that everything is unknown.
It would then identify the missing evidence that matters. The analyst needs to connect the excerpt to the deployed firmware, determine how secure mode is set, and inspect whether any other user-data write paths bypass encryption. A controlled write of known data could clarify the interpretation of the dump, especially when paired with a trace or an inspection of the executed path.
These are targeted questions because they arise from specific gaps. They are more useful than asking for complete firmware, all documentation, and every possible test without explaining which uncertainty each item would resolve. They also show how the conclusion could change: evidence that secure mode is invariably enabled and all relevant writes use the encrypted branch strengthens the universal claim; an observed bypass weakens it.
The example illustrates why a model’s final answer needs an argument. The answer can be only a few paragraphs long, but those paragraphs must preserve the conditional facts, the alternative explanations, and the connection between missing evidence and the proposed conclusion.
Alternatives should test the inference
The strongest contrary interpretation is useful because it identifies a place where the argument may fail. In the controller example, the important alternative is that encryption exists but is optional in the relevant configuration. This directly tests the step from a documented encrypted path to a universal statement about actual writes.
An arbitrary alternative adds less value. The mere possibility that every source is fabricated would undermine almost any inquiry, but it does not help distinguish plausible implementations unless the record supplies a reason to suspect fabrication. Alternative explanations should be connected to the evidence and weighed according to what the record supports.
The same principle applies to missing evidence. A model should distinguish an unresolved detail that could reverse the conclusion from one that merely refines an already supported explanation. Otherwise it can respond to every difficult question with a long inventory of uncertainties while avoiding a useful judgment.
For expert work, uncertainty is not the absence of analysis. It is part of the result: this source establishes one fact, that inference depends on another, and the unresolved condition changes the strength or scope of the conclusion. Making those relationships explicit allows the reviewer to decide where additional work is worthwhile.
Inspectable prose is different from a record of internal computation
An explicit explanation should not be treated as a faithful transcript of everything the model did internally. Turpin and colleagues’ study of unfaithful chain-of-thought explanations showed that the tested models could be influenced by features they did not acknowledge in their explanations. The result concerns those experiments and models, but establishes why a plausible rationale alone is insufficient evidence about the process that produced an answer.
This does not make an argument in final prose useless. Its value is that it can be checked independently. The cited passage either supports the stated fact or it does not. The conditional inference either follows under the stated assumptions or it contains a gap. An explanation can be a useful object of review without being an introspective account.
The reviewer should consequently examine both the conclusion and its stated support. An attractive rationale may rationalize a mistaken answer, while a correct answer may contain an invalid argument. Neither should receive full credit in a task whose deliverable is the analysis itself.
This also changes what I ask from the model. I want a defensible account of the evidence, rather than a performance of extensive thinking. The analysis should make the necessary steps available for inspection and leave out background that does not help establish or limit the conclusion.
Evaluate analytical usefulness directly
Opus 4.6 remains my reference because I know its behavior. GLM-5.3 interests me because I have observed it developing substantial explicit arguments, and Kimi K3 belongs among the candidates I would compare. Fable 5 and 5.1 should remain in that comparison as well. A problem with communication may respond to clearer instructions, and a few disappointing responses should not become a permanent judgment about a model.
The useful experiment is a blinded comparison of actual or sanitized technical records. Each candidate should receive the same evidence and the same analytical question. Effort, output limits, and available tools should be recorded because they affect both the result and the practical cost of obtaining it.
Review should separate technical accuracy, characterization of evidence, unsupported inference, treatment of alternatives, recognition of missing facts, and argumentative coherence. A single preference score could otherwise reward a polished explanation whose factual support is weaker.
I would also distinguish an omitted step from an incorrect step. The first may require the reviewer to reconstruct a valid argument; the second may require abandoning or repairing it. Both reduce usefulness, but they create different review costs and may respond differently to a prompt change.
The hypothetical controller case provides several concrete checks. Did the model notice the secure-mode condition? Did it distinguish the manual’s diagram from observed execution? Did it overstate what the dump established? Did it identify evidence that would resolve the universal claim? These questions measure the behavior needed for the task without turning verbosity into the objective.
A comparison should include cases where a favorable conclusion is well supported, where the evidence supports only a narrower claim, and where a tempting premise is wrong. Otherwise the test may reward a model for producing qualifications on every problem, even when a direct conclusion is justified.
I would retain some records outside the initial selection process and examine how much correction each answer requires. The practical result is not simply which draft I prefer, but how much work remains before its argument is accurate, complete, and usable.
For this workload, a model earns its place by preserving the discipline of the analysis. It needs to develop a supported inference, recognize the strongest relevant alternative, and identify the missing fact that actually changes the result. A general benchmark can identify promising candidates. The inspected argument determines whether the candidate is useful here.