- Sodium is encoded as NA in the chemical component dictionary and worthwhile to remove right at the start of any data analysis.
3-letter codes are the standard for non-canonical amino acids, yet, e.g. histidine derivates A1IQX, A1AZ2 showcase that this is not exclusively the case.
- Peptide nucleic acids (PNAs) are a thing. No agreement on their Ramachandran angles calculation exists, equally β- and γ-amino acids can cause futile angle calculations.
- No Ramachandran angle pairs can be defined for ncAAs at solely first or last position, e.g. histidine amide HIA.
- Utilizing https://www.rcsb.org/ligand/[NAN] is immensely useful to get an overview of whether a component is a ligand, covalently bound or included as true backbone. For example, SAH delivered worse predictability in RareFold, while checking here also shows that it is never a true backbone component.
- The PDB is influenced by reporting bias, and medically relevant proteins are more predominant.
- Post-translational modifications and synthetic inclusions may be analyzed as two different sets.
- BioPython can compute non-canonical amino acid angles when you set the chain up via peptide = builder.build_peptides(structure, aa_only=0).
Cell Biologist:
Go to the Database tab. Enter into the table search fields a modification or functional word of interest, e.g. phospho- or insulin.
The columns Name and Description are of most interest to you. The Chemical properties tab can show you
how the chemical compound looks like, when you search and input the IDs SMILES string.
Structural Biologist:
Use the Ramachandran plot tab for backbone analyses. Use the same tab to upload or fetch a .pdb or .mmcif or .cif file.
Go to the Protein Structure tab for 3D structure and side chain analyses. Non-canonical side chains are highlighted in red.
Protein Design:
Option 1: Go to the Database tab. Enter a broad parent search term in the Name column (e.g. ester, methyl-, ...) or click on the header of the Count column to filter the data for IDs of interest that often occur.
The column Count is of most interest to you. Use these IDs of interest in the Ramachandran plot tab against a canonical background to find candidates that
might enable new backbone structures. Option 2: Go to the Chemical properties tab. Scroll to the bottom to find derivative options to replace your standard amino acid position with.
Medicinal Chemist:
Go to the Chemical properties tab. Scroll to the bottom to find derivative options to replace your standard amino acid position with. Scroll to the
top to sort the database generally by Tanimoto similarity, and find similar candidates for a canonical amino acid to exchange it by. Of most interest are the SMILES to 2D drawing canvases at the top of the
Chemical properties tab: Here you can place SMILES strings and compare the 2D chemistry side-by-side.
Synthetic Chemist:
Go to the Database tab. The column Organism (if synthetic = likely synthesized, then further included in a synthetic peptide e.g. by another pharmaceutical company) is of most interest to you.
You can filter candidates of your interest and go the Chemical properties tab to fetch Tanimoto similarities and SMILES for these. Use the SMILES to populate the chemical structure canvases to see the 2D chemical structure.
If the Database tab shows you organisms, that synthesize a candidate of your interest, this indicates a possible biosynthetic or enzymatic route present.
Pharmaceutical Optimization:
Go to the Chemical properties tab. Scroll to the bottom to find derivative options to replace your standard amino acid position with.
Of most interest are D-amino acids against protease degradation for you. Of second most interest is the Database tab for you: Enter a clinical or function term, e.g. insulin,
in the search field. The database is filtered to components, where at least one PDB ID of their presence, is annotated functionally with the search term (e.g. insulin, insulin-binding, etc.).
Microbiologist:
Go to the Database tab. The column Organism is of most interest to you. Enter a species into the table search field. You find chemical
components which are present in proteins of this organism, hinting at a possible biosynthetic capability of your organism to make this non-canonical amino acid, or modification of your interest.
Especially important for antibiotics research.
Toxicology:
Start in the Database tab and enter, e.g. toxin, into the search field. The column Description filters the dataset down to components, where at least one PDB ID of their presence, is annotated functionally with the term, e.g. toxin.
Additionally, you can upload a .pdb or .mmcif or .cif file in the Ramachandran plot tab. Go then to the Protein structure tab to see your toxic protein or peptide toxin, with the non-standard positions in red. It is displayed engaged with its target,
when both are present in the PDB file to give you insights into binding. Note: For example azetidine (ID: 02A, flower toxin) is teratogenic due to self-inclusion as proline mimic into the collagen chain. We
have very little of such structures resolved in the PDB, hinting at open questions and open data gaps to close in mechanistic toxicology.