Don’t confuse “conserved” with “conservative” – no, I’m not getting political on you. Instead, I’m talking about the makeup of proteins and whether a specific “letter” in a protein’s sequence (an amino acid residue) is constant throughout evolution and/or among different versions of the protein (in which case, we’d call it “conserved”) and, if not, whether a change made to it (a substitution) maintains the general biochemical properties of the amino acid (in which case we’d call it conservative) or substantially alters those properties (in which case we’d call it non-conservative). 

Another technical note: In linking up to form a protein, amino acids join their “amino” and “acid” groups, leaving only a amino acid “residues” – they still have the unique part of the amino acid (the side chain or R-group) and are often still referred to as “amino acids,” but are technically no longer “amino acids.” 

If an “amino acid” is conserved (kept constant) throughout evolution and/or in different versions of the protein, especially if the sequence in the rest of the protein is changing lots, it might imply that that particular amino acid is functionally important – any organisms that randomly acquired mutations causing substitutions of it would have a fitness disadvantage and natural selection would weed them out, preventing that mutation (and corresponding amino acid substitution) from being passed on. 

If, however, an “amino acid” is not conserved, varying greatly across evolutionary time, etc. it is likely less fundamentally functionally important to the protein’s main “purpose.” That doesn’t mean it isn’t important for some things, it just might be in a flexible region that allows for more variation and/or part of a species-specific or even an isoform- (version of the protein even within a species) – specific adaptation.

If an amino acid’s identity is not conserved, its properties still might be. And this gets us to the concept of “conservative” vs. “nonconservative” substitutions. Both involve changes to what amino acid is present at a given location in a protein*, but with different amounts of biochemical drama!

Conservative substitutions swap out one amino acid for one that is biochemically similar: similar size, charge, hydrophobicity (how much water avoids it), functional groups (specific chemical groups such as hydroxyl groups and amino groups that can carry out specific functions), etc. They are therefore unlikely to dramatically alter the shape and/or properties of the protein (though they may enable stepwise evolution of more dramatic changes).

Nonconservative substitutions swap out an amino acid for one that is *not* biochemically similar. 

*Be careful with numbering. When assessing conservation, you want to be sure to compare positions after aligning the sequences (e.g. such as with CLUSTAL, described in more detail below). Don’t assume position “5” in Protein A corresponds to position “5” in Protein B. Instead, Protein B might actually have an extended N-terminus (starting sequence) and therefore “5” in Protein A might correspond to “25” in Protein B! 

Note: we often refer to amino acid “mutations” but mutations actually happen at the level of the DNA (or RNA in the case of RNA viruses) with the instructions for making the protein. This distinction can be important because not all mutations actually cause changes in proteins thanks to things like redundancy in the genetic code (multiple ways to “spell” the same amino acid). Therefore, we instead refer to (or at least we *should* instead refer to) amino acid substitutions when discussing changes in the identity of an amino acid present at a specific location in a protein. 

Mutations that cause conservative substitutions might be evolutionarily neutral, and thus can rapidly accumulate, even if an amino acid is functionally important for a protein – so don’t count a position out as “not important” if it is non-conserved – especially if the changes in it are typically conservative. 

Nonconservative substitutions have greater potential to really alter the properties of a protein – for better or worse. Their impact depends in large part on the importance of the amino acid in the protein’s structure and/or function (for example, a nonconservative substitution might have a minimal effect if it occurs in a floppy loopy region but a large effect if it occurs in the structural core of a protein). 

Some nonconservative substitutions become “standard” for a given version of a protein, but sometimes they’re instead one-offs that cause diseases.

When scientists suspect that a genetic mutation is causing a disease, they might search the patient’s DNA for substitution-causing “mutations.” And they’ll likely find a lot when comparing to a “reference genome” – most of these mutations will have limited effect (and we often refer to them as “polymorphisms” rather than “mutations” in part to get across this point that variation is natural and typically healthy!). Often, then, they’ll try to home in on genes for potentially-relevant proteins and, within those, mutations that are predicted to cause non-conservative substitutions, especially if they occur at positions that are highly-conserved!

Tools for assessing conservation:

CLUSTAL Omega (CLUSTAL Ω/Ο) is one of many MSA (Multiple Sequence Alignment) Tools for aligning and comparing protein sequences. There are other, sometimes more powerful, tools including some that incorporate structural conservation, etc. but CLUSTAL and its formatting remain a mainstay. 

Symbols under the sequences indicate degree of conservation at the above position

* indicates fully conserved: “all” residues at that position in the alignment are identical

• indicates strongly conserved within residue type: “all” residues at that position in the alignment share similar biophysical properties

• indicates weakly/semi-conserved within residue type: some residues at that position in the alignment share similar biophysical properties

lack of symbol indicates non-conserved

Note: it’s surprisingly hard to find details, so hopefully I explained it all correctly…

CLUSTAL coloring scheme is often used to portray information about the similarity of aligned protein sequences at the amino acid level (in tools like JalView). CLUSTAL coloring colors each conserved or semi-conserved* amino acid residue in a sequence based on its biophysical properties

  • light blue: hydrophobic (nonpolar)
  • red: positively-charged (basic) (yes, that’s weird and confusing because blue is typically used for this but it is what it is!)
  • purple: negatively-charged (acidic)
  • green: polar, uncharged
  • salmon: cysteine
  • orange: glycine
  • ugly yellow: proline
  • teal: aromatic
  • white: unconserved 

*Whether or not a residue is colored depends on the percentage of amino acids at that position with similar biophysical properties. Prolines, cysteines, and glycines, are colored even if not conserved.

Note: Some of the thresholds are more complicated than others and allow for some leniency in classification if stricter limits are surpassed (e.g., a K or an R is considered conserved enough to color red if >60% of the residues at that position are K or R, OR iif > 85% are K, R, or Q. https://www.jalview.org/help/html/colourSchemes/clustal.html

With CLUSTAL, you have to start with a list of sequences. You can get that by first using a tool like BLAST. 

Another tool for assessing conservation is ConSurf, which (if the server is functional again in the future at least) finds related proteins, aligns them, then gives you a color-coded sequence and structural map showing you the relative conservation of amino acid residues in each part of the protein. You can get a similar outcome by doing a BLAST search, making an alignment, and then opening that alignment in a ChimeraX session file with the structure you want to map it to. 

From Dr. Chris Berndsen, James Madison University: 

“To map your alignment conservation scores onto the 3D structure, ChimeraX can color by the conservation.

1.     In Tools menu, make sure Command Line Interface has a check next to it.

2.     In the command line, paste: 

color byattr seq_conservation palette cyanmaroon range -1,1 novalue silver key true

3.     In the log on the right side of the screen, it will indicate the seq_conservation range. Adjust the range in the pasted command above and repaste it into the command line. 

ChimeraX will repaint the 3D structure from cyan (variable) through white to maroon (highly conserved), giving you an immediate visual map of which surface patches and buried regions are most constrained.”

Huge thanks to Dr. Berndsen, the Malate Dehydrogenase CUREs community (MCC), and PyBMB

More on CLUSTAL: https://thebumblingbiochemist.com/365-days-of-science/proteinalignments/ &  https://youtube.com/shorts/RUgtc378oqE

If you’re looking to create and/or visualize (interactively or for figures) protein alignments, here are a few free resources I recommend.

Thanks to all who provide these resources!!!!!

More on databases and when & how to use them here: https://bit.ly/databases_guide  ; YouTube: https://youtu.be/ZyLOWqZazgc 

Leave a Reply

Your email address will not be published. Required fields are marked *