The Morgan dollar grading study
Systematic Differences Among Third-Party-Graded Morgan Dollar Populations: Evidence from a Large Dealer-Sourced Image Corpus. The complete 29-page preprint with methodology, statistics, and supporting analyses.
What the study asks
Third-party certification grades are central to the market for Morgan dollars, yet claims that one grading service is systematically stricter than another are usually based on reputation, anecdote, or small crossover samples. This study asks a narrower and directly observable question: whether the current populations of same-grade coins in different services' holders differ visually when measured by one common instrument.
Classic Morgan-dollar images were sourced directly from coin dealers and evaluated with the computer-vision grading system that underlies the CoinAiLyzer application. The model was used as a common measurement instrument, not as ground truth. After identity repair and preregistered scope corrections, the primary cohort contained 121,119 image-level predictions representing 65,092 physical coins, with a fully held-out external cohort of 27,272 additional coins.
What it found
AU-55 & AU-58
PCGS coins scored 0.73 Sheldon points higher than same-grade NGC coins at AU-55, and 0.79 points higher at AU-58—the largest, most repeatable service gap in the study.
MS-62 through MS-66
Across the commercial Mint State core, PCGS and NGC populations converged and met the preregistered practical-parity criterion in the primary population.
Same label, different coins
The median within-label standard deviation was 0.63 points (PCGS) and 0.65 (NGC); two random same-label coins differed by at least 1.0 modeled point about a quarter of the time.
The held-out external cohort reproduced upper-AU separation and MS-62 through MS-66 convergence, though its largest effect occurred at MS-60 rather than AU-58. Across both cohorts, the stable description is elevated separation around the upper-AU/low-Mint-State boundary, convergence through much of MS-62 to MS-66, and a smaller upper-MS return at MS-67. Same-year, same-mint, same-service, same-grade coins are not visually interchangeable under the common ruler—quantitative support for "buy the coin, not the holder."
How it was tested
- Common instrument, out of fold. Five-fold out-of-fold predictions were aggregated at the physical-coin level; the model evaluating a coin had not been trained on that coin, and holders, labels, and backgrounds were removed before grading.
- Preregistered analysis. The grade-specific pattern was frozen before the new grader's output was examined, with prespecified measures including score gaps, Cohen's d, probability of superiority, and physical-coin bootstrap confidence intervals.
- Extensive controls. Date/mint composition, confidential dealer-source controls, side-specific checks, same-coin old-versus-new replication, and probability-vector diagnostics.
- Held-out replication. A fully held-out external cohort of 27,272 physical coins tested whether the confirmatory pattern reproduced.
The paper details the study's limitations, including that coins were not randomly assigned to services, that the observed populations were shaped by submission and crossover behavior, and that weak strike was not independently measured.