
The ACMG/AMP guidelines have anchored germline variant classification since 2015. The vocabulary they introduced, Pathogenic through Benign, is now the shared language of clinical genomics, and the criterion codes (PVS1, PS3, PM2, PP3, BA1) are spoken aloud in variant review meetings every day. What is easy to miss is how much of the framework underneath that vocabulary has been rewritten in the decade since.
The rewriting happened one paper at a time. A ClinGen working group would take a single criterion, ask what evidence actually justifies the strength assigned to it, and publish a revision. Do that for a dozen criteria across ten years and the result is a framework that keeps its original shape while almost every rule inside it has been recalibrated. A lab following the 2015 paper literally today would apply several criteria in ways the field has since abandoned.
This post traces that recalibration: which papers changed which criteria, why computational evidence changed more than anything else, and what to expect from ACMG/AMP v4, the points-based successor now in pilot. For the companion view focused on how these shifts play out in cancer, see Modernized ACMG & Cancer Variant Classification.
What the ACMG Guidelines Actually Specify
Richards et al. published Standards and guidelines for the interpretation of sequence variants in Genetics in Medicine in 2015 [1]. It remains one of the most cited papers in clinical genetics, and it does two things.
First, it defines 28 evidence criteria. Sixteen argue for pathogenicity and are coded by strength: PVS1 (very strong), PS1 through PS4 (strong), PM1 through PM6 (moderate), and PP1 through PP5 (supporting). Twelve argue for benignity: BA1 (stand-alone), BS1 through BS4 (strong), and BP1 through BP7 (supporting). Each criterion names a specific kind of observation, such as a null variant in a gene where loss of function causes disease, or an allele frequency too high for the disorder.
Second, it specifies combining rules that turn a set of applied criteria into one of five classifications: Pathogenic, Likely Pathogenic, Uncertain Significance, Likely Benign, or Benign. One very strong plus one strong criterion yields Pathogenic. Two moderate criteria yield Uncertain Significance. The rules are a lookup table, not a calculation.
That lookup table is the part that has held up. The criteria feeding it are the part that has been rebuilt, and understanding how ACMG classification works in 2026 means knowing which revisions apply.
A Decade of Recalibration: The Papers That Changed the Criteria
ClinGen’s Sequence Variant Interpretation (SVI) working group took on the job of resolving ambiguities in the original criteria. Its output, along with related calibration work, is the reason two labs applying “the ACMG guidelines” in 2026 may be applying substantially different rules depending on which revisions they have adopted.
| Paper | Criteria | What changed |
| Richards et al., 2015 Genetics in Medicine | All | Established the framework: 28 criteria, four pathogenic strength tiers, five classification outcomes. |
| Biesecker & Harrison, 2018 Genetics in Medicine | PP5, BP6 | Recommended retiring the “reputable source” criteria. A classification should rest on the underlying evidence, not on another laboratory’s conclusion about it. |
| Abou Tayoun et al., 2018 Human Mutation | PVS1 | Replaced a single blanket loss-of-function rule with a decision tree. Strength now varies with variant type, exon position, predicted nonsense-mediated decay escape, and whether loss of function is an established mechanism for the gene. |
| Ghosh et al., 2018 Human Mutation | BA1 | Set an explicit allele frequency threshold for stand-alone benign evidence and published a curated exception list of variants that exceed it yet remain pathogenic. |
| Tavtigian et al., 2018 / 2020 Genetics in Medicine / Human Mutation | All | Showed the 2015 combining rules approximate a Bayesian model, then converted the criteria into an additive point system with a natural scale. |
| Brnich et al., 2020 Genome Medicine | PS3, BS3 | Introduced OddsPath calibration for functional assays. An assay’s evidence strength now depends on how well it was validated against known pathogenic and benign controls, not on the assay’s existence. |
| Pejaver et al., 2022 American Journal of Human Genetics | PP3, BP4 | Calibrated missense predictors against curated variant sets and established score thresholds per evidence tier, replacing the consensus-of-multiple-tools approach. |
| Walker et al., 2023 American Journal of Human Genetics | PVS1, PS1, PS3, PM5, PP3, BS3, BP4, BP7 | Unified how splicing evidence enters the framework, covering both computational splice predictions and RNA assay results, and extended PS1 to variants predicted to produce the same splice effect as a known pathogenic variant. |
| Bergquist et al., 2025 Genetics in Medicine | PP3, BP4 | Extended calibration to 13 missense predictors and identified those that reach Strong evidence for pathogenicity and Moderate evidence for benignity. |
Read as a sequence, these papers make one argument repeatedly. The 2015 criteria were written as categorical rules with fixed strengths, and nearly every revision since has replaced a fixed strength with a calibrated, evidence-dependent one. PVS1 stopped being “null variant, very strong” and became a decision tree. PS3 stopped being “functional study shows damaging effect” and became a question about assay validation. PP3 stopped being “several tools agree” and became a score threshold with a measured odds ratio behind it.
PM2 followed the same path in the opposite direction. Absence from population databases was originally moderate evidence; ClinGen’s SVI recommended downgrading it to supporting strength, on the reasoning that rarity is expected for most variants in most genes and therefore says less than the original weighting implied.
From Rules to Points
The most structurally important revision is also the least visible in daily practice. Tavtigian and colleagues showed that the 2015 combining rules, which look like an arbitrary lookup table, behave approximately like a Bayesian model in which each criterion contributes odds of pathogenicity and the strength tiers are related by a consistent multiplier [5].
Once that is established, the criteria can be expressed as points. On the naturally scaled system, a supporting criterion is worth 1 point, moderate 2, strong 4, and very strong 8. Pathogenic criteria contribute positive points and benign criteria negative ones, and the five classifications fall out as score bands.
Two things follow. Evidence that pulls in opposite directions can be combined arithmetically rather than adjudicated by hand, which is the situation the 2015 rules handled least well. And intermediate strengths become expressible: a criterion can be applied at moderate strength even when the guidelines list it as supporting, which is exactly what the calibration papers require. Nearly every gene-specific specification published by a ClinGen Variant Curation Expert Panel depends on that flexibility.
The point system is also the foundation of ACMG/AMP v4, discussed below.
Computational Evidence: The Criterion That Changed Most
No criterion has been revised as thoroughly as PP3/BP4. The 2015 guidelines allowed computational evidence at supporting strength only, and asked for “multiple lines of computational evidence” to agree. In practice that meant running SIFT, PolyPhen-2, MutationTaster, and several others, then counting votes.
The problem with vote counting is that the tools are not independent. Many share training data, and several use the outputs of others as input features. When five correlated predictors agree, the agreement mostly reflects their shared lineage rather than five separate lines of evidence. Consensus felt rigorous and added little information.
Missense Prediction
Pejaver et al. took a different approach [6]. Rather than asking which tools agree, they asked what a given score is actually worth: for each predictor, they estimated the local posterior probability of pathogenicity across the score range, using variants with established classifications as ground truth. That converts a raw score into an evidence strength on the ACMG scale.
The finding that mattered is that a single well-calibrated predictor can support more than the supporting-strength ceiling the 2015 guidelines imposed. Bergquist et al. later extended the calibration to 13 tools and found that several, including REVEL and BayesDel, reach Strong evidence for pathogenicity and Moderate evidence for benignity at appropriate thresholds [7].
ClinGen’s operational recommendation follows directly: choose one calibrated predictor in advance, fixed per laboratory or per gene, and apply its published thresholds. Committing to a tool before seeing the result is the point. It removes the temptation to consult additional predictors when the first one disagrees with expectations, which is the failure mode consensus scoring quietly enabled.
Splice Prediction
The ClinGen SVI Splicing Subgroup did the equivalent work for splicing [8], covering both computational predictions and RNA assay results, and specifying how each enters PVS1, PS1, PS3, PM5, PP3, BS3, BP4, and BP7. Two of its recommendations reshape everyday practice: computational splice evidence should come from a single calibrated tool rather than a panel, and PS1 extends beyond amino acid identity to variants predicted to produce the same splice alteration at the same site as a known pathogenic variant.
CI-SpliceAI is a practical answer to the first recommendation. It is an open-source deep learning model trained on GENCODE annotations that predicts how variants affect RNA splicing [9], and it serves as an open successor to the original SpliceAI without the licensing restrictions that limited clinical adoption of that model. We have covered it in two previous posts:
- Detect Cryptic Splicing Events with the New CI-SpliceAI Annotations in VarSeq
- CI-SpliceAI High Confidence Regions
One consequence deserves attention because it surprises people. Canonical splice donor and acceptor variants should be evaluated through the PVS1 loss-of-function pathway, not through computational splice evidence. Applying both double-counts the same biological observation, and avoiding that kind of double-counting is an explicit design goal of the guideline revisions.
What This Means for Automated Classification
An automated classifier has to implement some or all of these papers recommendations into a comprehensive system. The ACMG Classifier in VarSeq implements the 2015 guidelines, but with the revisions ClinGen paper guidance fully incorporated. Missense evidence comes from a calibrated predictor algorithm rather than a consensus vote. Splice evidence comes from CI-SpliceAI, shipped precomputed for more than 47 million variants on GRCh37 and GRCh38. Population frequency criteria support per-gene, inheritance-aware thresholds, so BA1, BS1, and PM2 can reflect the prevalence and inheritance model of the disorder being tested rather than one global cutoff. Expert-curated interpretations and your laboratory’s own assessment catalog feed the classifier directly, and each classification records where it originated. Alongside the categorical call, the classifier reports a point score on the Tavtigian scale, which is useful for ranking a variant list even though the classification itself still follows the ACMG combining rules.
Configuring all of this, choosing the predictor, setting thresholds, and pointing each evidence category at your preferred annotation sources, is covered step by step in the Starting VSClinical ACMG Guidelines tutorial in our learning site.
ACMG/AMP v4: What Is Coming
The revisions described so far were published piecemeal, each addressing one criterion. The next release consolidates them. ACMG, AMP, the College of American Pathologists (CAP), and ClinGen are jointly developing an updated standard for sequence variant classification, referred to as SVC v4.0. In its versioning, the 2015 Richards paper is v3.
The stated goals are to address the appropriateness of criteria that have not held up (PP5, BP6, and PM2 are named explicitly), resolve ambiguities in how criteria are applied, revisit the strengths assigned to criteria individually and in combination, prevent double-counting of the same underlying observation, and give explicit guidance for combining pathogenic and benign evidence. The working group also intends to fold in ClinGen recommendations as they are published, rather than letting another decade of revisions accumulate outside the standard.
Structurally, v4 makes the Bayesian point system the foundation rather than an overlay. Classification bands are defined directly on the point scale, anchored to probability of pathogenicity:
| Points | Classification |
| ≥ 10 | Pathogenic |
| 6 to 9 | Likely Pathogenic |
| 0 to 5 | Uncertain Significance |
| -3 to -1 | Likely Benign |
| ≤ -4 | Benign |
The benign side sees the larger shift. Richards et al. never defined an explicit probability boundary for Benign, and the Bayesian reformulation inferred one around 0.1%. The v4 draft uses 1% instead and lowers the evidence weight required to reach Likely Benign, which should make confident benign calls easier to reach and reduce the number of variants that linger as uncertain purely for lack of benign evidence.
Beyond scoring, the draft reorganizes evidence codes so that related observations are consolidated rather than counted twice, adds decision trees for evaluating each evidence type, incorporates gene-disease validity into classification, and introduces a way to subdivide variants of uncertain significance by likelihood of pathogenicity. That last change targets a real reporting problem: ACMG has noted that a large share of genetic testing reports contain at least one VUS, and a single undifferentiated category serves neither clinicians nor patients well.
As of mid-2026 the new ACMG SVC v4.0 standard remains in pilot. The pilot is extended to clinical laboratories and industry experts to collect feedback on the experience ofscoring a diverse set of variants using the draft version of the guidelines. Results reported at ACMG in 2026 showed 28 of 30 variants reaching greater than 90% concordance on the three-level scale, and the working group has since opened a wider round to the broader community. Publication is expected in 2027, and the working group has recommended a transition period rather than expecting laboratories to switch validated pipelines overnight.
Frequently Asked Questions About the ACMG Guidelines
What are the ACMG guidelines?
The ACMG guidelines are a joint consensus standard from the American College of Medical Genetics and Genomics and the Association for Molecular Pathology for interpreting germline sequence variants. Published by Richards et al. in 2015, they define evidence criteria and rules for combining them into a five-tier classification. They are the reference standard for germline variant classification in clinical laboratories worldwide.
How many ACMG criteria are there?
There are 28 criteria: 16 supporting pathogenicity (PVS1, PS1-PS4, PM1-PM6, PP1-PP5) and 12 supporting benignity (BA1, BS1-BS4, BP1-BP7). The code prefix indicates direction and strength. P is pathogenic and B is benign; VS is very strong, S is strong, M is moderate, P in the second position is supporting, and A is stand-alone.
What are the five ACMG variant classifications?
Pathogenic, Likely Pathogenic, Uncertain Significance (VUS), Likely Benign, and Benign. “Likely” corresponds to greater than 90% certainty of pathogenicity or benignity, a threshold the guidelines state explicitly so that laboratories apply the term consistently.
What is the difference between the ACMG and AMP guidelines?
The 2015 ACMG/AMP guidelines cover germline variants and classify by pathogenicity. A separate AMP/ASCO/CAP standard, published in 2017, covers somatic variants in cancer and classifies by clinical actionability into four tiers rather than by pathogenicity. The two answer different questions, and a somatic workflow needs both. See our overview of somatic variant analysis for how the tiering system works.
Do the ACMG guidelines cover copy number variants?
Not the 2015 paper. ACMG and ClinGen published a separate technical standard for constitutional copy number variants in 2020 (Riggs et al.), which uses its own point-based scoring system built around dosage sensitivity, gene content, and overlap with established regions. Laboratories reporting both sequence and structural variants apply both standards. Our CNV and structural variant analysis page covers that workflow.
Why do two laboratories classify the same variant differently?
Usually because they have adopted different revisions. One lab may still count agreement among prediction tools for PP3 while another applies calibrated thresholds from a single predictor. One may apply PM2 at moderate strength, another at supporting. Add gene-specific specifications from ClinGen expert panels and differences in which evidence sources each lab subscribes to, and the same variant can reach different classifications while both labs correctly follow “the ACMG guidelines.” Consolidating these divergences is a central motivation for v4.
Can ACMG classification be automated?
Criteria that depend on structured, machine-readable evidence can be evaluated automatically, and that covers a large share of them: population frequency (BA1, BS1, PM2), computational predictions (PP3, BP4), variant type and location (PVS1, PM4, PM1), and prior classifications (PS1, PM5). Criteria that depend on unpublished family data, phenotype nuance, or judgment about assay quality still need a variant curation scientist. Automation is best understood as producing a defensible starting point with its evidence exposed for review, not a final answer.
ACMG v4 FAQ
What is ACMG v4?
ACMG v4, formally SVC v4.0, is the forthcoming update to the sequence variant classification standard, developed jointly by ACMG, AMP, CAP, and ClinGen. It replaces the categorical combining rules of the 2015 guidelines with a Bayesian points system and folds in the criterion-specific revisions ClinGen has published since.
When will ACMG v4 be published?
It has not been published. The standard is in pilot as of mid-2026, and publication is expected in 2027. Earlier estimates pointed to 2026, and the timeline has moved as pilot feedback has been incorporated, so treat any date as provisional until the paper appears.
How is v4 different from the 2015 guidelines?
The largest change is scoring. Evidence is summed on a point scale and classification is read off score bands rather than matched against a combining-rule table, which makes conflicting pathogenic and benign evidence tractable. Beyond that, v4 reorganizes evidence codes to prevent double-counting, adds decision trees for each evidence type, revisits criteria that have not held up in practice (PP5, BP6, PM2), incorporates gene-disease validity, and allows variants of uncertain significance to be subdivided by likelihood of pathogenicity.
Will v4 change existing variant classifications?
Some, yes. The pilot has reported high concordance with current practice for most variants, which suggests the majority of classifications are stable. The changes on the benign side are where movement is most likely: raising the Benign probability boundary to 1% and lowering the evidence needed for Likely Benign should reclassify some variants that currently sit at Uncertain Significance for want of benign evidence. Variants with substantial evidence on both sides are the other group to watch, since they are exactly what the point system handles differently.
Will laboratories need to reclassify their variant catalogs?
Reclassifying every archived variant on publication day is neither expected nor practical, which is why the working group has recommended a transition period. The realistic approach is to apply v4 to new classifications once validated, and to re-examine the existing catalog in priority order: reported variants first, then those whose classification rests on criteria v4 changes most, particularly PP3/BP4, PM2, PP5, and BP6.
What should a laboratory do now to prepare?
Adopt the published ClinGen revisions that v4 consolidates, because they are current best practice regardless of when v4 lands. Specifically: move PP3/BP4 to a single calibrated predictor with published thresholds, adopt the PVS1 decision tree, stop applying PP5 and BP6, and confirm how your pipeline weights PM2. A laboratory current with the ClinGen recommendations will find v4 a change in arithmetic rather than in evidence, which is a far smaller validation exercise.
Conclusion
The ACMG guidelines have proven durable because their structure separates the evidence from the rules for combining it. That separation let the field replace the evidence underneath criterion after criterion without disturbing the five-tier vocabulary clinicians rely on. v4 finally updates the combining rules as well, and it can do so safely because a decade of calibration work established what each piece of evidence is actually worth.
For laboratories, the practical question is not whether to adopt v4 but how current your interpretation of “the ACMG guidelines” is today. If you would like to see how these changes play out in cancer classification specifically, read our companion post, Modernized ACMG & Cancer Variant Classification.
References
- Richards, S., Aziz, N., Bale, S., Bick, D., Das, S., Gastier-Foster, J., … & Rehm, H. L. (2015). Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine, 17(5), 405-424.
- Biesecker, L. G., & Harrison, S. M. (2018). The ACMG/AMP reputable source criteria for the interpretation of sequence variants. Genetics in Medicine, 20(12), 1687-1688.
- Abou Tayoun, A. N., Pesaran, T., DiStefano, M. T., Oza, A., Rehm, H. L., Biesecker, L. G., & Harrison, S. M. (2018). Recommendations for interpreting the loss of function PVS1 ACMG/AMP variant criterion. Human Mutation, 39(11), 1517-1524.
- Ghosh, R., Harrison, S. M., Rehm, H. L., Plon, S. E., & Biesecker, L. G. (2018). Updated recommendation for the benign stand-alone ACMG/AMP criterion. Human Mutation, 39(11), 1525-1530.
- Tavtigian, S. V., Greenblatt, M. S., Harrison, S. M., Nussbaum, R. L., Prabhu, S. A., Boucher, K. M., & Biesecker, L. G. (2018). Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Genetics in Medicine, 20(9), 1054-1060. See also Tavtigian, S. V., Harrison, S. M., Boucher, K. M., & Biesecker, L. G. (2020). Fitting a naturally scaled point system to the ACMG/AMP variant classification guidelines. Human Mutation, 41(10), 1734-1737.
- Pejaver, V., Byrne, A. B., Feng, B. J., Pagel, K. A., Mooney, S. D., Karchin, R., … & Topper, S. (2022). Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3/BP4 criteria. The American Journal of Human Genetics, 109(12), 2163-2177.
- Bergquist, T., Stenton, S. L., Nadeau, E. A. W., Byrne, A. B., Greenblatt, M. S., Harrison, S. M., … & Pejaver, V. (2025). Calibration of additional computational tools expands ClinGen recommendation options for variant classification with PP3/BP4 criteria. Genetics in Medicine, 27(6), 101402.
- Walker, L. C., de la Hoya, M., Wiggins, G. A. R., Lindy, A., Vincent, L. M., Parsons, M. T., … & Spurdle, A. B. (2023). Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: recommendations from the ClinGen SVI Splicing Subgroup. The American Journal of Human Genetics, 110(7), 1046-1067.
- Strauch, Y., Lord, J., Niranjan, M., & Baralle, D. (2022). CI-SpliceAI: improving machine learning predictions of disease causing splicing variants using curated alternative splice sites. PLoS One, 17(6), e0269159.
- Brnich, S. E., Abou Tayoun, A. N., Couch, F. J., Cutting, G. R., Greenblatt, M. S., Heinen, C. D., … & Berg, J. S. (2020). Recommendations for application of the functional evidence PS3/BS3 criterion using the ACMG/AMP sequence variant interpretation framework. Genome Medicine, 12(1), 3.
- Riggs, E. R., Andersen, E. F., Cherry, A. M., Kantarci, S., Kearney, H., Patel, A., … & Martin, C. L. (2020). Technical standards for the interpretation and reporting of constitutional copy-number variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics (ACMG) and the Clinical Genome Resource (ClinGen). Genetics in Medicine, 22(2), 245-257.