Research paper · Preprint v1.2

The Morgan dollar grading study

Systematic Differences Among Third-Party-Graded Morgan Dollar Populations: Evidence from a Large Dealer-Sourced Image Corpus. The complete 29-page preprint with methodology, statistics, and supporting analyses.

This is an observational study of current certified-holder populations measured by one common computer-vision instrument. A higher model score for one service at a fixed stated grade is not proof that the other service graded incorrectly, that the model knows the objectively correct grade, or that the difference was caused by the original grading decision.
By Troy LowryPreprint v1.2 · August 22, 202629 pages · PDF

What the study asks

Third-party certification grades are central to the market for Morgan dollars, yet claims that one grading service is systematically stricter than another are usually based on reputation, anecdote, or small crossover samples. This study asks a narrower and directly observable question: whether the current populations of same-grade coins in different services' holders differ visually when measured by one common instrument.

Classic Morgan-dollar images were sourced directly from coin dealers and evaluated with the computer-vision grading system that underlies the CoinAiLyzer application. The model was used as a common measurement instrument, not as ground truth. After identity repair and preregistered scope corrections, the primary cohort contained 121,119 image-level predictions representing 65,092 physical coins, with a fully held-out external cohort of 27,272 additional coins.

What it found

Upper-AU separation

AU-55 & AU-58

PCGS coins scored 0.73 Sheldon points higher than same-grade NGC coins at AU-55, and 0.79 points higher at AU-58—the largest, most repeatable service gap in the study.

Commercial core parity

MS-62 through MS-66

Across the commercial Mint State core, PCGS and NGC populations converged and met the preregistered practical-parity criterion in the primary population.

Within-grade spread

Same label, different coins

The median within-label standard deviation was 0.63 points (PCGS) and 0.65 (NGC); two random same-label coins differed by at least 1.0 modeled point about a quarter of the time.

The held-out external cohort reproduced upper-AU separation and MS-62 through MS-66 convergence, though its largest effect occurred at MS-60 rather than AU-58. Across both cohorts, the stable description is elevated separation around the upper-AU/low-Mint-State boundary, convergence through much of MS-62 to MS-66, and a smaller upper-MS return at MS-67. Same-year, same-mint, same-service, same-grade coins are not visually interchangeable under the common ruler—quantitative support for "buy the coin, not the holder."

How it was tested

  1. Common instrument, out of fold. Five-fold out-of-fold predictions were aggregated at the physical-coin level; the model evaluating a coin had not been trained on that coin, and holders, labels, and backgrounds were removed before grading.
  2. Preregistered analysis. The grade-specific pattern was frozen before the new grader's output was examined, with prespecified measures including score gaps, Cohen's d, probability of superiority, and physical-coin bootstrap confidence intervals.
  3. Extensive controls. Date/mint composition, confidential dealer-source controls, side-specific checks, same-coin old-versus-new replication, and probability-vector diagnostics.
  4. Held-out replication. A fully held-out external cohort of 27,272 physical coins tested whether the confirmatory pattern reproduced.

The paper details the study's limitations, including that coins were not randomly assigned to services, that the observed populations were shaped by submission and crossover behavior, and that weak strike was not independently measured.