
A transcript’s reference sequence does not always match the genome assembly beneath it. Even a one-base difference can shift downstream coding coordinates, alter a codon, or change a variant’s predicted effect.
VarSeq 3.1.0 accounts for these differences directly. The new RefSeq Genes track records the exact differences between curated RefSeq RNA sequences and the genome assembly. We call these differences Tx Deltas. VarSeq takes these deltas into account during annotation so that HGVS notation, sequence ontology, and effect predictions reflect the true underlying transcript sequence. GenomeBrowse uses the same data to show where the RNA and genome diverge.
This post explains what Tx Deltas are, why they matter, and how VarSeq applies and visualizes them.
Why Curated RNAs and the Genome Can Differ
RefSeq transcripts (NM_ and NR_ records) are based on curated sequences built from mRNA, EST, genomic, and other evidence. Most align exactly to GRCh37 or GRCh38, but some contain differences that affect variant annotation:
- Substitutions: The RNA contains a different base from the genome at a given position. This can reflect a common polymorphism, a difference in the original sequence evidence, or another curation decision.
- RNA-only bases: The transcript contains one or more bases that are absent from the assembly. Every downstream cDNA coordinate shifts by the length of the insertion.
- Genome-only bases: The genome contains bases that the curated transcript omits. Variants on those bases may have no cDNA coordinate for that transcript.
- Frame-correction micro-introns: Unlike the three cases above, this is a difference in the exon model rather than in the sequence. Where an exon does not align co-linearly, the RNA carries a few extra bases relative to the genome, and the gene annotation compensates by excising a very short “intron” of a base or two so that the genomic reading frame stays correct downstream. Those excised bases are present in both the genome and the curated RNA. Only the exon model leaves them out.
These differences can change how a variant is described and interpreted. An RNA-only base shifts every downstream c. coordinate. A substitution can determine whether a variant is synonymous or missense. A variant in sequence omitted from the RNA has no valid cDNA position for that transcript.
Correct annotation therefore requires the curated RNA sequence; a simple extraction of transcript exons from the genome is not sufficient.
Examples include:
- ALMS1
NM_015120.4, which contains a three-base insertion relative to GRCh38 - NKRF
NM_001417890.1, which contains a single-base insertion paired with a micro-intron - FOXO6
NM_001291281.3, which contains a 200-base insertion
For these transcripts, genome-only projection can produce different c. coordinates and, in frameshifting cases, different protein predictions.

How the RefSeq Genes Track Stores Tx Deltas
NCBI publishes transcript-to-genome alignments with each RefSeq annotation release. For VarSeq 3.1.0, we extract and curate every non-identity alignment operation and store the results in two fields on the RefSeq Genes track:
- Tx Genome Delta: An ordered list of differences between the genome and curated RNA. Each entry records the operation type, genomic position, and affected bases.
- Tx Delta Net Indel: The net length difference between the RNA and genome. Use it to identify transcripts with shifted coordinates and see the size of the shift.
Transcripts with identity alignments have empty delta fields and are annotated as before.
The stored operations are:
- Substitution: The RNA contains a different base at this position.
- Insertion: The RNA contains bases absent from the genome.
- Deletion: The genome contains bases absent from the RNA.
- Micro-intron: The exon model excludes a short region retained by the RNA
How Tx Deltas Improve Variant Annotation
VarSeq 3.1.0 applies each Tx Delta before calculating HGVS notation, protein consequences, and variant effects. The resulting annotation follows the curated RefSeq transcript rather than a simple extraction from the genome.
That improves three parts of the result:
- Accurate HGVS coordinates: c. and n. positions include any transcript-to-genome length differences upstream of the variant.
- Correct protein predictions: VarSeq translates the curated RNA sequence, preserving the transcript’s intended codons and reading frame.
- Transcript-specific effect calls: Synonymous, missense, frameshift, and other effects are determined from the transcript sequence the variant is being reported against.
This matters in cases such as the variant X:119605953 -/G in NKRF shown below.

Without accounting for the Tx Delta, the variant appears to be a one-base insertion that would cause a frameshift. However, the curated NKRF transcript already contains this G at that position. By incorporating the Tx Delta into transcript annotation, VarSeq recognizes that the variant restores the curated transcript sequence and correctly reports the synonymous protein consequence NP_001404819.1.Pro15=.
The same genomic variant can also receive different c. and p. descriptions on two transcripts of the same gene when only one contains a Tx Delta. That is expected: each result correctly describes the variant relative to its transcript.
Tx Deltas are also respected when searching or importing HGVS expressions, allowing published transcript coordinates to map to the correct genomic position.
How GenomeBrowse Displays Tx Deltas
GenomeBrowse 3.1.0 reads the same Tx Genome Delta field used during annotation, making transcript-to-genome differences visible at the sequence level.
Codons and amino acids reflect the RNA
At coding-sequence resolution, GenomeBrowse translates the delta-applied RNA. Amino acid labels therefore show the residues encoded by the curated transcript, including residues affected by substitutions.

Codons that contain a delta are shaded
Each codon is drawn across its true genomic footprint. If a codon spans more than three bases in the genome sequence, its box spans the omitted bases and receives distinct shading. Only codons that contain a delta are shaded; downstream codons are translated correctly but remain visually unchanged.

RNA-only bases receive an insertion marker
For RNA-only bases that cannot be drawn on a genomic axis, GenomeBrowse marks their alignment position with a vertical bar, showing where the additional transcript sequence enters.

Transcript details list every difference
The transcript details now show the curated RNA size beside the spliced genomic size and displays the net difference. A new Genome to RNA differences table lists each operation, including its genomic position, type, sequence change, and position in the coding RNA.
Conclusion
Tx Deltas allow VarSeq 3.1.0 to preserve curated RefSeq sequence differences from the source track through annotation and visualization. The result is more accurate HGVS notation, protein predictions, and effect calls for affected transcripts. When a result differs from a genome-only projection, GenomeBrowse shows the transcript-to-genome difference behind it, allowing you to visually inspect each change.
VarSeq 3.1.0 is coming soon. Reach out to our team using the contact link below if you are interested in trying it out.