Non-canonical nucleic acids (NC⋅NAs) describe the set of all strands with nucleic bases apart from adenine (A), guanine (G), cytosine (C), thymine (T) or uracil (U). The non-canonical nucleic building blocks comprise (i) modified RNA bases, (ii) synthetic building blocks and (iii) modified DNA bases. Modified RNA bases are important for RNA signalling and stability, modified DNA bases for genetic regulation, and synthetic building blocks for RNA therapeutics. The name MÜNT⋅RNA sends regards to the painter Gabriele Münter, the expressionist artist on the side of Kandinsky.
Play around with your favourite background! Then, choose non-canonical RNAs or DNAs and compare their folds across structures. Do they explore novel parts of the fold landscape, or of the chemical landscape? How close are they to the native structure? See where in the strand from 5'- to 3'-end they are situated.
Non-canonical RNA or DNA bases (NC⋅NAs) are best found using 1- to 5-letter codes. Here, all modified bases present in the wwPDB Chemical Component Dictionary (867 NC⋅NA compounds) are listed and enriched in which PDB-IDs, which specific chain, which organism, by which method and at which resolution the RNA or DNA strands include these non-canonical building block.
Look for your favorite non-canonical RNA or DNA base, look for your dataset of interest! We provide two sequences: The traditional one-letter sequence (e.g. as used by AlphaFold and .fasta-files) where all non-canonical compound are unspecifically hidden behind 'X'. For improvement, the column 'REF_nc_sequence' provides a 3-letter codes (XXX) bracketed instead of the one-letter X position. Your filtered dataset can be printed or downloaded under attribution to the authors. Attribution to the RSCB PDB mirror (https://pdb101.rcsb.org/more/how-to-cite) and the wwPDB Chemical Component Dictionary (https://www.wwpdb.org/data/ccd) is advised.
| SMILES | ID | Name | Count | PDB ID | Description | Organism | Unmasked Sequence | Masked Sequence (X or parent) | Non-Canonicals in Sequence |
|---|
In the two inset windows, you can compare two chemical structures. A non-canonical RNA or DNA base to a canonical, or two non-canonicals, can be compared, as you wish. If you fetched an ncNA in first tab, the same ncNA will be displayed on the left. You can adjust which two compounds you would like to compare by entering SMILES into the free text fields below each window (OpenEye stereochemical SMILES, to obtain them use the table below).
In the table, SMILES and Tanimoto similarity for RNA or DNA bases in the PDB are given. You can search for an ncNA by its ID. Each column lists the compounds Tanimoto similarities after MORGAN bits (see RDkit) to a canonical deoxyribonuleotide DNA base (DT, DG, DC, DA) or ribonucleotide (U, G, C, A). The table can be sorted by high similarity (value: 1) or low similarity (value: 0). Click on the column header to sort. When downloading the table, take care. It is separated by ;.
Here, one specific RNA or DNA structure can be fetched and analyzed in more detail (see the first tab which PDB IDs are of interest). Default display: 4GXY. Standard positions are displayed gray. Non-standard positions with non-canonical RNA or DNA bases are highlighted red and fully displayed as sticks. Note: If no structure appears, this is due to discontinuity in a chain.