University rankings often combine several desirable properties without explaining which question the resulting order answers. A university can be affordable, promote social mobility, provide attentive undergraduate instruction, or produce exceptional scholarship. Each property matters, but the institution that performs best on one need not perform best on the others.

The question that interests me is: Which universities contain the strongest concentrations of intellectual talent and produce the highest-quality scholarship? For a major research university, scholarship should dominate the answer. Such an institution discovers knowledge, assembles exceptional scholars, trains future scholars, and sustains an environment in which difficult intellectual work can occur. Undergraduate education matters partly because students can participate in that environment, with access to people working at the frontier of their disciplines.

This is my chosen definition of excellence for research universities. It does not describe every valuable institutional mission, and it is not an appropriate way to rank a college principally devoted to undergraduate instruction. It does provide a coherent objective against which a measurement system can be assessed. What follows is a proposal for that system, rather than a ranking calculated from a completed dataset.

What does “best” measure?

My objection to the domestic U.S. News Best Colleges ranking concerns the breadth of its claim. A score combining undergraduate outcomes, resources, reputation, and social mobility can help families compare institutions. The title invites readers to interpret that score as a more general judgment of academic quality.

The 2027 edition replaced graduate indebtedness with earnings by major, according to the University of California’s account of the methodology. That changes the economic outcome being rewarded; it does not by itself establish that economic outcomes received a larger total weight. The earlier 2024 revision removed class size, high-school class standing, faculty terminal-degree share, and alumni giving as standalone measures, while adding faculty-research measures and placing greater emphasis on outcomes. The University of Tennessee’s comparison makes both sides of that change visible. The methodology was neither a pure measure of scholarship before the revision nor a complete abandonment of research afterward.

The problem is the objective represented by the composite. Pell-recipient graduation rates can tell us something important about a university’s success in educating lower-income students. They do not directly measure the scholarly distinction of its faculty. Giving them weight makes student opportunity and outcomes part of the publisher’s definition of “best.” My preference for scholarship is also a normative choice; the requirement is to make the choice explicit and select indicators that bear on it.

Suppose A represents a declared set of academic criteria and S a social objective. A score R = αA + βS, with positive weight on both, measures a combination. Even if the attributes are correlated, the score does not answer the narrower question of which institution is academically stronger. There is nothing wrong with publishing the composite, provided its name and interpretation match what was measured.

The distinction also applies to individual accomplishment. A student’s original mathematical research is evidence of academic achievement. An under-resourced background may explain why the accomplishment is especially unusual, or help interpret a record in which opportunities were missing. It does not change the quality of the mathematics. Similarly, a university can add substantial educational value for students who entered with weaker preparation without currently containing the strongest students or producing the strongest research. Those are different achievements, and measuring each separately preserves information that a single score conceals.

Earnings introduce another mismatch. Raw salary partly reflects the labor-market price of a field. An exceptional historian or basic scientist may earn less than an ordinary software engineer or investment banker. Earnings by major addresses part of the composition problem, but even a field-adjusted salary does not directly measure intellectual accomplishment. Where graduate outcomes enter an academic-strength index, I would emphasize distinction within the graduate’s chosen field.

Methodology changes also complicate movements in rank. An institution can move because its performance changed, competitors changed, or the instrument changed. A large movement after a revision does not demonstrate a comparably large change in academic quality. The underlying indicators must support that interpretation.

Useful parts of the existing rankings

International rankings generally take research more seriously, although they answer somewhat different questions. I would use their strongest measures as building blocks rather than adopt one whole methodology.

The 2026 Times Higher Education methodology assigns 29% to research environment and 30% to research quality. Its quality measures include citation impact, the 75th percentile of field-weighted citation impact, top-decile research, and influence through citations from influential papers. These are useful attempts to distinguish exceptional research from volume. However, teaching and research reputation together receive 33%, while international outlook receives 7.5%. Reputation can contain information about quality, but accumulated prestige should not substitute for present accomplishment. International collaboration may improve scholarship; the international composition of a department is not itself scholarly achievement.

The QS methodology gives 30% to academic reputation, 20% to citations per faculty, 15% to employer reputation, 5% to employment outcomes, and 10% to the faculty/student ratio. International measures and sustainability receive the remainder. Citations per faculty address a relevant size problem, but the combined 45% for academic and employer reputation gives surveys a very large role in a ranking of current scholarly strength.

The 2026 Academic Ranking of World Universities is closer to my objective. Its weights are 10% for alumni Nobel and Fields distinctions, 20% for faculty distinctions, 20% for Highly Cited Researchers, 20% for Nature and Science publications, 20% for indexed publications, and 10% for per-capita performance. This emphasizes accomplishment, but the coverage is uneven. Two general-science journals cannot represent the full range of elite scholarship, and Nobel and Fields awards cover few disciplines. ARWU adjusts its treatment of Nature and Science for institutions specializing in humanities and social sciences, but that does not supply equivalent measures of excellence across those fields.

The 2026–27 U.S. News Best Global Universities release describes a research-focused comparison of more than 2,250 institutions. Its top ten—Harvard, MIT, Stanford, Oxford, Cambridge, Tsinghua, Berkeley, Yale, UCL, and Columbia—looks much more like a list of major centers of scholarship than many domestic undergraduate rankings. Research reputation and bibliometrics bring it closer to the question, while still leaving quality, productivity per scholar, and total institutional power insufficiently separated.

The CWTS Leiden Ranking Open Edition provides particularly useful distinctions. It reports both counts and proportions of papers among the top 1%, 5%, 10%, and 50% by citations, normalized to field and publication year, alongside normalized citation scores. Its field normalization distinguishes about 4,500 algorithmically defined fields. This avoids automatically favoring fields with denser citation practices. Its normalized indicators apply to a restricted set of publications suitable for citation analysis, so their coverage should not be mistaken for all scholarship.

Leiden separates output scale from the proportion of high-impact work, but neither quantity alone measures output per researcher. NTU’s reference ranking addresses the staff denominator by dividing four publication and citation indicators by full-time-equivalent academic staff. It uses QS staff data and estimates missing values partly from publication counts. That makes more comparisons possible, but an output-derived denominator is not an independent measurement of research productivity.

China makes the distinction consequential

China’s research growth shows why these choices matter. The 2026 Nature Index academic table, based on 2025 output, places Zhejiang first with a Share of 1,276.85, Harvard second with 1,259.01, Tsinghua third with 1,213.59, and Shanghai Jiao Tong fourth with 1,191.47. The country table records approximately 52,735 for China, 26,006 for the United States, and 4,820 for the United Kingdom. These are rounded Share totals, not counts of researchers or papers.

The results establish enormous research scale within the Index’s selected publications. They do not establish that Zhejiang has a stronger average faculty than Harvard, or that Tsinghua produces more exceptional work per researcher than MIT. Changing the publication slice also changes the apparent order. In Nature and Science alone, Harvard ranks first, Stanford second, MIT third, and Tsinghua eighth. In applied sciences, Tsinghua, Zhejiang, and Shanghai Jiao Tong take the first three positions.

Clarivate’s 2025 Highly Cited Researchers release provides another perspective: 2,670 awards were associated with the United States and 1,406 with mainland China. Harvard had 170, Stanford 141, Tsinghua 91, MIT 85, and Zhejiang 57. These awards have their own coverage and selection limits, but they show why dominance in one aggregate publication measure cannot simply be translated into dominance on every dimension of scholarship.

The more consequential question is whether Chinese universities are becoming comparable with the strongest American and British institutions in exceptional scholarship per active researcher, concentrations of leading faculty, and aggregate research power. National population does not provide the denominator for that institutional comparison. Student enrollment is also a poor substitute for the research workforce. The people producing the work must be measured alongside it.

Scale nevertheless remains a real academic asset. A university with 8,000 excellent researchers can sustain laboratories, doctoral programs, infrastructure, and combinations of expertise that a university with 800 equally good researchers cannot. Erasing scale would lose that capability; allowing scale to dominate would obscure the concentration of excellence. A useful index needs both.

The proposed academic-strength index

I would organize the index around six components:

Component Proposed weight
Research quality and intensity 30%
Faculty scholarly distinction 25%
Aggregate research power 20%
Student caliber 15%
Graduate scholarly and professional distinction 7.5%
Undergraduate intellectual environment 2.5%

Three-quarters of the score concerns faculty and research directly. These weights declare my priorities; they are not estimates of the true contributions to an independently observed quantity called excellence. They also do not finish the design. Each component needs defined indicators, coverage rules, and a common score scale before its weight can have the intended effect.

Research quality and intensity

The core should identify influential contributions relative to the norms of their fields. For suitable disciplines, top-1%, top-5%, and top-10% publication proportions and mean normalized citation scores provide evidence about the quality of an institution’s output. I would give exceptional work more credit than merely above-average work because the objective is to identify scholarship near the frontier.

Quality proportions and research intensity need different denominators. PP(top 10%) already divides highly cited output by total publications. Dividing that proportion again by faculty size would reward a small institution simply for being small. Intensity should instead divide a field-normalized count of exceptional outputs by the research workforce producing them.

A conceptual intensity measure is Q = (w1 P1 + w2 P5 + w3 P10) / F, where the P terms are counts of high-impact publications and F is research FTE over the same period. Top-1% papers also appear in top-5% and top-10% counts. The weights must therefore represent deliberate cumulative bonuses, or the implementation must use disjoint bands. Otherwise some of the emphasis on exceptional work arises from accidental double counting. Any additional influence indicator should remain separate until its scale and contribution are defined.

For illustration, consider the existing scale comparison of 8,000 and 800 researchers. If the larger institution produces 1,000 top-decile contributions and the smaller produces 100 over the same period, each produces 0.125 per researcher. Their intensity is equal in this deliberately simplified example, while the larger institution produces ten times as much exceptional work. Neither result establishes their publication proportions, which also require total output. This is arithmetic to distinguish the measures, not a comparison of actual universities.

The measures must also respect the work of each discipline. Computer science requires serious treatment of conferences; mathematics needs citation windows appropriate to its pace. Humanities scholarship requires books, reviews, disciplinary recognition, and other evidence that a journal-citation formula misses. Creative work likewise needs field-specific assessment. A universal bibliometric formula would make calculation easier by measuring the wrong thing in some fields.

Faculty scholarly distinction

The second component measures the concentration of exceptional scholars. Relevant evidence includes Highly Cited Researchers, academy membership, Nobel Prizes, Fields Medals, Turing Awards, and major disciplinary honors. It should be normalized by field and faculty size, with current membership distinguished from institutional history.

A prizewinner currently on the faculty is evidence about the present community. A prize associated with a scholar who died decades ago is evidence about its history. An old discovery may remain consequential, but its historical affiliation does not supply a current faculty member. Awards with uneven national or disciplinary coverage should not silently penalize communities they rarely recognize.

Awards and citations can also describe the same scholar’s achievement. Their correlation does not make either useless, but a highly cited prizewinner should not appear to provide two independent confirmations of equal strength. The score should disclose the overlap and show what changes when correlated indicators receive less weight. Reputation might supplement accomplishment where expert recognition precedes measurable effects; it should carry the burden of adding information rather than replacing the evidence.

Aggregate research power

This component preserves institutional scale by counting total high-impact contributions. A university producing 1,000 top-decile papers deserves more aggregate-output credit than one producing 100, even when their intensity is equal or the smaller institution’s is greater.

Major grants, research infrastructure, doctoral programs, datasets, and software can also reveal capacity. Grants and infrastructure are inputs rather than completed scholarly achievements, so spending itself should not become excellence. I would report capacity indicators alongside the output score and incorporate them only where their role is defined. This allows the reader to see both what an institution has produced and what it can sustain.

Students, graduates, and undergraduate access

Student caliber receives 15% because a great university is also an intellectual community of unusually capable students. The measures should concern demonstrated ability: comparable examinations, advanced academic accomplishment, competitions, substantial original work, and appropriate evidence for graduate students. International examinations need validated mappings, and results from only the students who chose to submit scores cannot be treated as representative of the whole class.

Acceptance rate should receive no credit. Generating more applications does not itself strengthen the students who enroll. Socioeconomic or demographic background should not independently raise or lower this component, although it can help interpret missing opportunities and unusual records. The quantity being measured is demonstrated intellectual accomplishment.

Graduate distinction receives 7.5%. I would examine field-specific scholarly and professional achievement, major fellowships, distinguished appointments, creative contributions, and progression into demanding research or professional programs. An institution deserves credit for producing people who become exceptionally good at their chosen work, rather than for concentrating graduates in the highest-paying occupations. The interpretation must also distinguish alumni accomplishment from educational value added: selection and subsequent opportunities affect outcomes alongside the university’s contribution.

The final 2.5% concerns undergraduate access to the intellectual environment: work with leading faculty, advanced seminars, independent research, theses, and laboratories. Generic satisfaction is not the intended quantity. Class size can matter through access and participation, but an excellent large lecture is not automatically inferior to a mediocre small seminar. The question is whether capable undergraduates can participate meaningfully in the institution’s academic life.

From indicators to a credible comparison

The largest practical obstacle is measuring the research workforce consistently. Faculty headcount is not research FTE. Teaching duties, clinical appointments, postdoctoral researchers, doctoral students, hospitals, and joint affiliations vary across institutions and national systems. The numerator and denominator must describe the same institutional boundary and period. Counting a hospital’s output while omitting its researchers would inflate intensity. Estimating staff from the output being normalized would weaken the independence of the measure.

Missing denominators should therefore remain visible, through a withheld intensity score or an uncertainty range supported by the available records. Fractional credit for collaborative work, citation windows, self-citation rules, and boundary changes must also be published. These choices alter the quantity being compared, so they belong with the result rather than in a remote methodological appendix.

Subject-level comparisons should precede whole-university aggregation. Field normalization addresses citation customs; it does not determine how much mathematics, medicine, history, or engineering should count in the overall index. Those portfolio weights are another declared judgment. A broad institution and a specialized institute can both be exceptional without being interchangeable intellectual environments.

The component scores should consequently be the primary result. They would show an institution that is intense but small, powerful in aggregate but less exceptional per scholar, distinguished in selected fields, or historically prestigious with weaker current output. Before assigning an overall order, I would vary the weights, component scales, and consequential coverage choices. Institutions whose relative position changes readily should appear as a cluster, rather than receive an ordinal distinction that the evidence cannot sustain.

Separate rankings would still be useful for social mobility, affordability, financial value, teaching, and career outcomes. A university could lead in mobility while placing much lower in scholarship, or lead in scholarship while remaining expensive. Both results could be true, and neither needs to be concealed in an unrestricted “best” score.

The proposed index would answer the narrower question with which I began: where exceptional scholarship, exceptional scholars, and unusually capable students are concentrated. Its contribution would be a transparent separation of quality, intensity, and scale. That would make changes in the global academic system easier to interpret, including whether an apparent rise reflects more research capacity, stronger work per researcher, or a broader concentration of leading scholars.