Oncogenicity Guidelines for Somatic Variants: Beyond AMP Tiers

· Gabe Rudy · Clinical Genetics
Oncogenicity guidelines for somatic variants: beyond AMP tiers

Ask a germline analyst and a molecular pathologist what a variant means, and you will get answers to two different questions. The germline analyst is asking whether the variant causes an inherited disease. The pathologist is asking whether it is driving this tumor, and whether anything can be done about it. Those questions need different evidence, so they need different frameworks.

For a long time the somatic side lacked one. The 2017 AMP/ASCO/CAP guidelines gave laboratories a way to tier variants by clinical actionability, which is useful for reporting, but actionability is a statement about the evidence connecting a variant to a therapy, not about whether the variant does anything to the cell. A variant can be an unambiguous oncogenic driver and still be Tier III because no drug targets it. That gap is what oncogenicity classification fills.

This post is the somatic companion to our germline deep dive, ACMG Guidelines for Variant Classification: 2015 to v4, which traced how ClinGen rebuilt the ACMG/AMP criteria one paper at a time. Here the subject is how oncogenicity classification works: the frameworks behind it, the code set and point scale the Cancer Classifier uses, and where somatic evidence diverges from germline evidence. For the wider story with worked demo cases and concordance benchmarks, see the webcast summary Modernized ACMG & Cancer Variant Classification.

Oncogenicity Is Not Pathogenicity, and Not Actionability

The single most common source of confusion in somatic interpretation is treating three separate judgments as one. A tumor variant can be scored on all three axes, and the answers do not have to agree.

QuestionFrameworkOutput
Does this variant cause inherited disease?ACMG/AMP 2015 (Richards et al.)Pathogenic → Benign, five tiers
Does this variant drive the tumor?Golden Helix Oncogenicity 2019; ClinGen/CGC/VICC 2022 (Horak et al.)Oncogenic → Benign, five tiers
Can we act on it in this cancer type?AMP/ASCO/CAP 2017 (Li et al.)Tier I → Tier IV, clinical actionability
Three questions, three frameworks. A somatic report often needs answers to all three.

The distinction that matters most in practice is between the second and third rows. Oncogenicity is a biological property of the variant: does it disrupt a tumor suppressor or activate an oncogene. It does not change when a new drug is approved. Actionability is a property of the variant in a specific clinical context, and it moves constantly as trials complete and labels expand. An oncogenic assessment does not change as the drug and trial landscape evolves, but actionability does.

Keeping them separate is what lets a laboratory classify oncogenicity once and revisit actionability on its own schedule. That is the practical argument for adopting the oncogenicity framework even if your reports are organized by AMP tier.

The Frameworks That Define Somatic Classification

FrameworkWhat it definesScope
Li et al., 2017
Journal of Molecular Diagnostics
Four-tier clinical actionability (Tier I strong, Tier II potential, Tier III unknown, Tier IV benign/likely benign) with evidence levels A through D.Reporting and clinical significance, not oncogenicity.
VSClinical AMP, 2019Golden Helix defines Oncogenicity on a five-class scale and presents the point-based scoring system and criteria through webinars and community ClinGen meetings.Biological oncogenicity of somatic variants, cancer-type agnostic.
Horak et al., 2022
Genetics in Medicine
The first community consensus scoring system for oncogenicity, defining point-based oncogenic and benign criteria that produce five classes from Oncogenic to Benign. Validated on 94 variants across 10 cancer genes.Community adoption and expansion of the scoring system.
Burghel et al., 2026
Journal of Medical Genetics
ACGS/SVIG-UK specification of the oncogenicity criteria for UK practice, covering both solid and hematologic tumors.National implementation guidance layered on Horak.
The somatic classification frameworks and what each one answers.

The relationship between them is additive rather than competitive. Li et al. [2] remains the reporting structure most laboratories organize around. Golden Helix Oncogenicity and the Horak scoring system, the ClinGen/CGC/VICC recommendations published by Horak et al. [1], supply the oncogenicity call that Tier assignment assumes but never defines. The SVIG-UK guidelines [3] then do for oncogenicity what ClinGen expert panels do for germline criteria: take the Horak scoring system and specify how to apply it, including for hematologic malignancies, where clonal hematopoiesis and variant allele fraction complicate the picture in ways Horak et al. leave open.

How the Cancer Classifier Scores Oncogenicity

Golden Helix built and productized the first scoring system for Oncogenicity in 2019, before any community consensus existed, and presented that scale and code set to the VICC working group, whose work became the Horak scoring system in 2022. The Golden Helix Cancer Classifier still uses its own criteria and its own point scale rather than porting another framework’s codes. In our recent update to the Cancer Classifier, we have also incorporated refinements from the guidelines published since: criterion strengths are tuned so their relative weights track the Horak scoring system and the SVIG-UK guidelines.

The numeric scale has a different ancestor than you might expect. It follows Sherloc, Invitae’s refinement of the ACMG/AMP criteria [4], rather than the naturally scaled Bayesian point system [5] most germline implementations use. Sherloc’s contribution was to give the categorical strengths a compact, human-legible range where a reviewer can see at a glance how far a variant sits from a boundary. That is why the scale runs from -5 to +5 and why individual criteria are worth small whole numbers.

Evidence sums to a single score on that scale, and the score determines the class:

ScoreClassification
+5 and aboveOncogenic
+3 to +4Likely Oncogenic
-2 to +2Uncertain Significance
-4 to -3Likely Benign
-5 and belowBenign
The Cancer Classifier oncogenicity scale.

The batch Cancer Classifier algorithm also subdivides the uncertain band rather than collapsing it. A net score of +1 or +2 is reported as VUS/Weak Oncogenic, a net score of -1 or -2 as VUS/Weak Benign, and an exact 0 reached by offsetting evidence as VUS/Conflicting. A variant with no criteria applied at all is plain VUS. None of these labels changes the classification, but they let a reviewer sort a variant table by which direction the evidence leans, which is most of what a triage pass needs. The interactive VSClinical AMP interface reports all four cases as Uncertain Significance.

The two directions draw on different evidence. Benign scoring comes from germline population catalogs, in-silico functional and splicing predictions, and previous or clinical evaluations. Oncogenic scoring comes from somatic catalogs, domain and hotspot analysis, and the affinity of the gene for that particular variant type.

Mapping to the Horak Point Scale

When comparing the Cancer Classifier to the Horak scoring system, a general mapping can be done by halving the evidence strengths: Very Strong ±8 becomes ±4, Strong ±4 becomes ±2, Moderate ±2 becomes ±1, and Supporting ±1 stays ±1. The classification thresholds compress the same way, so Likely Oncogenic at +3 and Oncogenic at +5 mirror Horak’s +6 and +10.

One wrinkle is worth knowing, because it explains a scoring behavior that otherwise looks arbitrary. The Cancer Classifier thresholds are symmetric at ±1, ±3, and ±5, but the Horak scoring system reaches its benign tiers sooner: under Horak, a single Supporting benign criterion is already enough for Likely Benign. A strict halving would lose that. So a few benign criteria, notably Silent Variant and the high-frequency band of Population Frequency, are weighted one step beyond the straight conversion, specifically so a single such call still reaches Likely Benign. The goal is to reproduce that behavior on our scale, not to exceed it.

The Criteria

Each criterion is a two-letter code contributing points at one of a few strengths:

CriterionScoreWhat it capturesMaps to
Population Frequency (PF)-4, -3, -1The variant is common in one or more population catalogs.SBVS1, SBS1
Silent Variant (SV)-3Synonymous or non-coding variant with no predicted splicing impact.SBP2
Homozygous in Populations (HP)-2, -1The variant occurs in healthy individuals with a causal genotype.VarSeq extension
In-silico Predictions (IP)-3 to +3Computational evidence supports a deleterious or benign effect.OP1, SBP1
Clinical Evidence (CE)-1, +1, +2Previously classified as a pathogenic variant or pathogenic residue.OS1, OM4
Somatic Catalogs (SC)+1, +2 (+3 by curator)Rate of recurrence of the mutation across somatic catalogs.OS3, OM3, OP3
Splice Predictions (SP)+2Predicted damaging impact on splicing.OP1 (broken out)
Null Variant (NV)+2Loss-of-function variant more than 50 bp upstream of the last exon-exon junction.OVS1 (position half)
Null-Oncogenic Gene (NG)+2Loss-of-function variant in a gene where loss of function is oncogenic.OVS1 (gene half)
In-Frame (IF)+1In-frame insertion or deletion outside a repeat region.OM2
Active Region (AR)+1Occurs in an active binding site domain.OM1
Hotspot Region (HR)+1Occurs in a cancer hotspot region.OS3, OM3, OP3
Nearby Pathogenic (NP)+1Missense variant in a region with multiple pathogenic and no nearby benign variants.OM4
Null Variant Downstream (ND)+1Loss-of-function variant upstream of previously classified pathogenic LoF variants.VarSeq extension
The Cancer Classifier oncogenicity criteria, their point values, and the ClinGen/CGC/VICC codes they correspond to.

Three features of this table explain how somatic evidence behaves.

First, ruling a variant out takes less evidence than ruling it in. Three benign criteria each reach Likely Benign on their own: Population Frequency at -4, Silent Variant at -3, and In-silico Predictions at its benign extreme of -3. On the oncogenic side, only a strong calibrated missense prediction reaches Likely Oncogenic by itself, and every other route there combines at least two criteria. Neither Oncogenic at +5 nor Benign at -5 is reachable from a single criterion. The asymmetry is deliberate, and it matches how somatic interpretation goes. If a variant is common in healthy populations, it is not the driver, and no amount of hotspot proximity changes that. Building the case that a variant is driving a tumor is cumulative, because no single positional signal is sufficient.

Second, the very strong null-variant criterion is split in two. Horak’s OVS1 asks a compound question, whether a null variant sits in a position that triggers nonsense-mediated decay and lies in a bona fide tumor suppressor. The Cancer Classifier decomposes that into Null Variant, which tests the NMD-competent position, and Null-Oncogenic Gene, which tests whether loss of function is an established mechanism in that gene. Each is worth +2, and together they reach +4, past the Likely Oncogenic threshold. Splitting it means a truncating variant in the wrong gene, or in the last exon where it escapes decay, scores only half the evidence instead of all or nothing. Null Variant Downstream then adds +1 when other pathogenic truncations are already known downstream.

Third, computational evidence carries far more weight here than in the Horak scoring system, which folds all computational lines into a single supporting-strength call worth ±1. In-silico Predictions instead spans -3 to +3, the widest range of any criterion in the system, because it inherits the calibrated missense predictor from the germline classifier [7] and can therefore report a strength tier rather than a yes-or-no. Splice prediction is broken out separately at +2 rather than being lumped into the same computational bucket. This is the clearest case where adopting the Horak weights as written would have meant discarding evidence that the field has established as trustworthy.

Where the Cancer Classifier Differs from Horak

Two criteria have no counterpart in the Horak scoring system. Homozygous in Populations borrows the germline ACMG reasoning behind BS2, treating a variant seen in healthy individuals carrying the causal genotype as benign evidence; Horak uses population frequency alone and never looks at genotype state. Null Variant Downstream, described above, is likewise a VarSeq addition.

One criterion is an acknowledged approximation. Horak’s hotspot tiers are defined by per-amino-acid recurrence, at least 50 samples at a residue with at least 10 sharing the exact change for the strong tier. COSMIC does not expose per-amino-acid counts, so Somatic Catalogs uses a two-tier sample-count proxy instead: at least 5 samples earns +1 and at least 35 earns +2, and a curator can raise it to +3 when the literature documents recurrence the count does not reflect. Hotspot Region covers the same ground from a different angle, using the binary hotspot flag from cancerhotspots.org rather than raw counts. Two views of recurrence, scored separately, get closer to the intent of Horak’s hotspot tiers than either would alone.

Finally, two criteria exist but are never scored automatically. Functional Evidence, covering well-established functional studies, and Genetic Etiology, covering whether the variant’s mutational process matches the tumor’s, appear in the interactive workflow and are applied by a curator during an evaluation. They contribute to the same score and the same thresholds, but no database can be look them up. That is the boundary of automation in this framework. Everything a structured annotation can answer is scored for you, and the two criteria that require reading the literature stay with the curator.

Where Somatic Evidence Diverges from Germline

While there is a lot of overlap, the somatic classifier has a couple of interesting differences to review when compared to the ACMG Auto Classifier.

Somatic Recurrence Pulls Opposite of Population Frequency

In germline classification, rarity is evidence for pathogenicity. In somatic classification, the equivalent signal is recurrence: a position mutated over and over across independent tumors is under selection, and selection is the clearest evidence of a driver we have. Population frequency still matters, but it runs the other way. Population Frequency contributes only benign evidence, and at full strength, it is enough on its own to reach Likely Benign: a variant common in germline population catalogs is not driving a tumor.

This is why hotspot catalogs matter to somatic interpretation the way population databases matter to germline. Somatic Catalogs and Hotspot Region both depend on recurrence aggregated across tumors, and the quality of those catalogs directly sets how much evidence you can bring to a given variant. The framework uses two separate criteria for recurrence, scored by independent sources, giving you an idea of how central the signal is.

Mechanism Determines Which Evidence Applies

Germline classification largely treats loss of function as one category. Somatic classification cannot, because oncogenes and tumor suppressors break in opposite directions. A frameshift in TP53 is strong oncogenic evidence; the same frameshift in BRAF is not, because inactivating an oncogene does not drive a tumor. That is the whole reason the null-variant criteria are split the way they are, and why the gene half is checked against somatic and clinical evidence rather than assumed. Activating changes route through the hotspot and active-region criteria instead.

The practical consequence is that gene-level annotation is a prerequisite, not a nicety. Scoring a variant correctly requires knowing what kind of cancer gene you are in before any criterion can be applied.

Computational Evidence Transfers Directly

The one place the two frameworks share evidence directly is in silico prediction. A missense variant’s predicted effect on protein function and a variant’s predicted effect on splicing are the same biological questions regardless of whether the DNA came from blood or tumor. The calibration work done for the germline criteria, including the ClinGen missense calibrations [7] and the splicing recommendations, carries over.

That transfer is worth more than the Horak scoring system assumes. Written when computational evidence meant a panel of uncalibrated tools voting, the Horak criteria cap all of it at supporting strength. Since then the germline field has calibrated individual predictors against known variants and established what a given score is worth, and none of that work is germline-specific: a missense substitution damages a protein the same way in a tumor as in a germline sample. Carrying the calibrated strength tiers across is why In-silico Predictions has the widest range of any criterion here rather than the narrowest.

What This Means for Automated Classification

Almost all of this is machine-evaluable. Recurrence counts, population frequencies, domain overlap, variant consequence, and predictor scores are structured data, and scoring them consistently is exactly what software is good at. The value of writing the criteria down this precisely is that the resulting call is reproducible and every point in it is traceable to a source.

In VSClinical AMP, the computational criteria now run on the same calibrated annotations as the germline classifier. A single calibrated missense predictor feeds In-silico Predictions, and CI-SpliceAI [8] feeds the three criteria that depend on splicing, using a delta score of 0.2 as the disruption cutoff and 0.1 as the benign cutoff: Splice Predictions, Nearby Pathogenic when a variant shares a splice effect with a known pathogenic one, and Silent Variant, where a predicted splice disruption withholds the benign score a synonymous change would otherwise receive. That last one can really make a difference between calling a synonymous variant Likely Benign and quietly dismissing a real splice-disrupting driver. The annotation is covered in Detect Cryptic Splicing Events with the New CI-SpliceAI Annotations in VarSeq.

Clinical Evidence reads exact reference and alternate matches from ClinVar at one star or better, from CIViC, and from any clinical-evidence consortium source you configure, then falls back to same-codon matches for missense and in-frame variants. Population Frequency compares against per-gene frequency thresholds with inheritance-aware defaults, the same approach described for the germline side in Customizing the ACMG Classifier, so a hereditary cancer gene can carry its own cutoff rather than a global one. Each classification records the criteria applied, a one-line reason for each, and separate positive and negative evidence scores, so every point in the final call is traceable to a source.

Configuration, including selecting clinical evidence sources, setting gene-specific frequency thresholds, and choosing the missense predictor, is covered in the Getting Started with VSClinical AMP tutorial. For how oncogenicity fits into a full tumor workflow alongside actionability tiering, see our overview of somatic variant analysis.


Frequently Asked Questions

What does oncogenicity mean?

Oncogenicity is the capacity of a genetic variant to contribute to cancer development, by activating an oncogene or inactivating a tumor suppressor. It is a biological property of the variant. It is separate from whether the variant is clinically actionable, and separate from whether an inherited version of it would cause disease.

What are the five somatic variant classifications?

Oncogenic, Likely Oncogenic, Variant of Uncertain Significance, Likely Benign, and Benign. The Golden Helix Cancer Classifier, the Horak scoring system, and the SVIG-UK guidelines all share these five classes. The class follows from a summed evidence score; on the Cancer Classifier’s -5 to +5 scale, a variant reaching +5 is Oncogenic and one at -5 or below is Benign, with Likely Oncogenic and Likely Benign in the bands either side of an uncertain middle.

What is the difference between oncogenic and pathogenic?

“Pathogenic” is the germline term and means the variant causes or contributes to inherited disease, judged under the ACMG/AMP guidelines [6]. “Oncogenic” means a variant contributes to tumor development, judged under an oncogenicity framework such as the Golden Helix Cancer Classifier or the Horak scoring system [1]. The same DNA change can be evaluated both ways and reach different conclusions, because the frameworks weigh different evidence: population rarity matters for pathogenicity, tumor recurrence matters for oncogenicity.

How is oncogenicity classification different from AMP tiers?

AMP/ASCO/CAP tiers (I through IV) rank a variant by clinical actionability in a given cancer type, based on the strength of evidence linking it to therapy, prognosis, or diagnosis. Oncogenicity asks only whether the variant drives the tumor. A variant can be Oncogenic and still fall in Tier III if no therapy is associated with it. Most laboratories report tiers and use oncogenicity as the biological determination underneath.

What evidence goes into an oncogenicity call?

Six broad kinds. Recurrence across somatic catalogs and cancer hotspots is the signature oncogenic signal. Variant consequence read against the gene’s mechanism decides whether loss of function counts as evidence at all. Position within a functional domain or active binding site adds weight. Calibrated computational and splice predictions contribute across a range of strengths. Curated clinical assertions and prior expert classifications can contribute directly. And germline population frequency contributes benign evidence, strongly enough on its own to rule a variant out.

What is a cancer hotspot and why does it matter so much?

A cancer hotspot is a position recurrently mutated across independent tumors. Recurrence at that scale is unlikely by chance, so it signals positive selection, which is direct evidence the variant confers a growth advantage. The Horak scoring system devotes three of its 17 criteria to hotspot recurrence, at strong, moderate, and supporting strength. The Cancer Classifier captures the same signal through Somatic Catalogs, scored from COSMIC sample counts, and Hotspot Region, scored from the cancerhotspots.org flag. Together they make recurrence one of the most frequently applied oncogenic signals in the system.

Can oncogenicity classification be automated?

Most of it, yes. Hotspot recurrence, population frequency, variant consequence in an annotated tumor suppressor, and computational prediction are all machine-evaluable, and together they account for the bulk of the scoring. Functional-study criteria still need a curator reading the literature. In Golden Helix benchmarking against 219 expert-classified variants drawn from the Horak SOP and the SVIG-UK guidelines, automated oncogenicity calls reached 93% concordance. The benchmark is described in Modernized ACMG & Cancer Variant Classification.


Conclusion

Somatic interpretation spent years without a shared way to say whether a variant actually does anything, which left oncogenicity as an implicit judgment buried inside tier assignment. Making it explicit, quantitative, and auditable is what the last several years of work, ours and the community’s, has been for. The evidence that distinguishes a driver turns out to be its own thing: recurrence across tumors, consequence read against the gene’s mechanism, and functional data, with population frequency inverted into a benign signal.

For laboratories, the useful move is to score oncogenicity once, as a durable biological property, and let actionability move independently underneath it. For the germline side of that shared evidence, and how the ACMG criteria themselves have been rebuilt since 2015, read the companion deep dive ACMG Guidelines for Variant Classification: 2015 to v4. To see both classifiers running off the same annotations, visit the VSClinical product page.


References

  1. Horak, P., Griffith, M., Danos, A. M., Pitel, B. A., Madhavan, S., Liu, X., … & Sonkin, D. (2022). Standards for the classification of pathogenicity of somatic variants in cancer (oncogenicity): joint recommendations of Clinical Genome Resource (ClinGen), Cancer Genomics Consortium (CGC), and Variant Interpretation for Cancer Consortium (VICC). Genetics in Medicine, 24(5), 986-998.
  2. Li, M. M., Datto, M., Duncavage, E. J., Kulkarni, S., Lindeman, N. I., Roy, S., … & Nikiforova, M. N. (2017). Standards and guidelines for the interpretation and reporting of sequence variants in cancer: a joint consensus recommendation of the Association for Molecular Pathology, American Society of Clinical Oncology, and College of American Pathologists. The Journal of Molecular Diagnostics, 19(1), 4-23.
  3. Burghel, G. J., Mason, J., Baker, K., Moloney, K., … & Wragg, C. (2026). Association for Clinical Genomic Science (ACGS) guidelines for the classification of oncogenicity of somatic variants in cancer: recommendations by the UK somatic variant interpretation group (SVIG-UK). Journal of Medical Genetics, 63(3), 147-156.
  4. Nykamp, K., Anderson, M., Powers, M., Garcia, J., Herrera, B., Ho, Y. Y., … & Topper, S. (2017). Sherloc: a comprehensive refinement of the ACMG-AMP variant classification criteria. Genetics in Medicine, 19(10), 1105-1117.
  5. Tavtigian, S. V., Harrison, S. M., Boucher, K. M., & Biesecker, L. G. (2020). Fitting a naturally scaled point system to the ACMG/AMP variant classification guidelines. Human Mutation, 41(10), 1734-1737.
  6. Richards, S., Aziz, N., Bale, S., Bick, D., Das, S., Gastier-Foster, J., … & Rehm, H. L. (2015). Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine, 17(5), 405-424.
  7. Pejaver, V., Byrne, A. B., Feng, B. J., Pagel, K. A., Mooney, S. D., Karchin, R., … & Brenner, S. E. (2022). Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3/BP4 criteria. The American Journal of Human Genetics, 109(12), 2163-2177.
  8. Strauch, Y., Lord, J., Niranjan, M., & Baralle, D. (2022). CI-SpliceAI: improving machine learning predictions of disease causing splicing variants using curated alternative splice sites. PLoS One, 17(6), e0269159.

Leave a comment

Gabe Rudy

About Gabe Rudy

Gabe Rudy is the Vice President of Product and Engineering at Golden Helix, where for over two decades he has led the development of clinically validated software solutions that power precision medicine worldwide. Under his leadership, Golden Helix has delivered a suite of best-in-class tools for genomic analysis, including CNV calling, pharmacogenomics, carrier screening, and somatic variant interpretation. These solutions are designed for flexible deployment across on-premises, private cloud, and managed cloud environments, and are used by organizations ranging from small diagnostic teams to large clinical laboratories and even national-scale genomic initiatives. With a background in Computer Science and graduate work in compiler optimization and high-performance computing, Gabe brings a unique blend of software architecture expertise and deep domain knowledge in genomics. Since 2006, he directed product strategy and engineering at Golden Helix, ensuring the company stays at the forefront of innovation while maintaining the highest standards of usability, scalability, and quality. Gabe is an active participant in the genomics community, regularly presenting on topics such as NGS best practices, variant interpretation workflows, and the integration of AI into clinical diagnostics. His work has supported thousands of labs across the globe in the adoption of robust, intuitive, and clinically actionable bioinformatics workflows. Based in Bozeman, Montana, Gabe balances his passion for advancing precision medicine with family life and a love for the outdoors.

View all posts by Gabe Rudy →