We often think about proteins as being these beautifully-shaped things, composed of spiraling alpha-helixes and accordion-like sheets of beta-strands. And a lot of proteins *are* like this. But a lot of them are not! And even ones that are mostly full of those “structured” regions often possess parts that don’t have a set structure and are instead more loosely-goosey, spaghetti like. We call these parts Intrinsically Disordered Regions (IDRs) and we call proteins that are mainly IDR-ful Intrinsically Disordered Proteins (IDPs). Both can play really important roles (serving as scaffolds to bring molecules together, promoting “phase separation” where liquid turns into gooey membraneless compartments, etc.). So it’s helpful to be able to predict the presence of IDRs in proteins. Thankfully, as the “intrinsic” in the name suggests, we should be able to predict them based on the sequence since the lack of defined structure is intrinsic to them. And even more thankfully, there are software programs, including a main one called IUPred that can do this for us. Here’s a practical look at what it is, how it works, and some examples of it in action.
YouTube: https://youtu.be/1bXzOgUZ3PQ
But first, a little more background (some of this is adapted from a previous related post on protein flexibility and conformation changes, which also has more information about IDR functions, but most of it is new and focused on prediction). That other post: blog form: https://bit.ly/proteinflexibility ; YouTube: https://youtu.be/iDonUd_cX8w
Proteins like to move around – we call their different “poses” conformations and we call their shape-shifts conformational changes. Some parts of proteins like to move around more than others. And some proteins have lots of parts that really really like to move around. When regions move around so much that they have no real “set structure” we call them IDRs.
It’s important that we have tools to predict IDRs because sometimes the only way we know there’s an IDR is because we can’t see it! As someone trained in structural biology (the field that determines the 3D shapes that molecules take) I appreciate this all too well! If something has no set structure there’s no structure to see (even though the molecules are physically there).
Basically, their constant wiggling makes it so that, at any given time, the IDR in each copy of the protein will have a slightly – or incredibly – different conformation (shape). This is a problem for structure-solving because the techniques we use (techniques like x-ray crystallography and cryo-EM) rely on using info from lots and lots and lots of copies of the protein. And if they’re all different, you’ve got problems! All the different orientations in essence cancel out their own signal. Therefore, if you look at a structural model of a protein (which scientists make based on experimental signal), IDRs will be missing (or displayed as dashed or dotted lines). Much more on this here: https://bit.ly/crystalstructuremodels
Sometimes, IDRs are only “sometimes” IDRs – binding to a binding partner might cause them to snap into a set shape, or at least stabilize them in a pose. So we can sometimes get to see them if we include those partners. But remember that you’re just catching a single snapshot of one of its forms.
Note: this can also be a reason why you might have a problem expressing (making) some eukaryotic (e.g animal or plant) proteins in bacteria and/or overexpressing them in any type of cell because they don’t have or don’t have enough of partners that usually stabilize them. Removing IDRs can be a way to get them to express better – but might effect their activity.
In the case of crystallography, sometimes IDRs can actually prevent us from getting any crystals and therefore we can’t even figure out the structure of the structured parts! Crystallography relies on all the copies “freezing” in the exact same pose and the wiggliness of the IDRs can prevent them from doing this. So, sometimes we actually chop them off in order to get proteins to crystallize (thankfully IDRs are often located at or near the termini (start or end of the protein) so we can just truncate the genetic recipe for the protein and then ask cells to make it for us to purify). But that relies on being able to predict their presence. Which brings us back to disorder prediction tools.
One of the main ones, and the one I focus on here and in the video is called IUPred (which is currently at IUPred3). The “IU” part stands for Intrinsically Unstructured and what it does is it takes the amino acid sequence of a protein and predicts* what regions of a protein are intrinsically disordered (don’t have a defined structure).
IUPred scores are probabilities of amino acids being in an Intrinsically Disordered Region (IDR) of a protein. Scores range from 0 (almost definitely structured) to 1 (almost definitely disordered). They are calculated* based on the amino acid identity & the surrounding protein context. Then, IUPred plots take the scores that were calculated for each amino acid residue in a protein and plot them along the protein’s length (N-terminus to C-terminus). Regions above the midpoint (0.5) line are likely disordered, and regions below the line are likely structured.
The plot of “well-structured” proteins will mostly be below the line. The plot of intrinsically disordered proteins (IDPs) will mostly be above the line. Many proteins have regions that are well-structured (below the line) interspersed with intrinsically disordered regions (IDRs) (above the line).
*it uses an “energy estimation algorithm” based on which amino acids are happy together in structured protein regions. The exact algorithm it uses is too mathy for my brain, but from best I can surmise, it uses information from solved structures of proteins to create a pairwise matrix of info about how comfortable amino acids are near one another inside of structured proteins. The more comfy, the lower the free energy. The less comfy, the higher the free energy. It can then take that info and apply it to the sequence of the protein (one amino acid at a time) to calculate the estimated free energy. If it’s low, that’s a good sign it likely is structured and if it’s high, it’s likely disordered. At least that’s my best understanding of it and hopefully I explained it okay! Now, let me get back to the stuff *I* am more comfy with!
One thing I’m pretty comfy with is the stuff I researched in grad school, RNAi (RNA interference), which is a way cells control levels of specific mRNAs (and thus regulate the protein made based on their instructions) by using small RNAs such as miRNAs or siRNAs to direct an RNA induced silencing complex (RISC) to them. I talk a lot about it in a lot of posts so I won’t go through it here, other than to use some of the players involved as examples.
Much more on microRNA (miRNA)-mediated RNA interference (RNAi) here: http://bit.ly/microRNARNAi.
As I mentioned briefly at the top, proteins with lots of IDR’s can serve as scaffolds holding other proteins and/or nucleic acids together. A great example of this is GW182-family proteins (e.g. TNRC6A-C) helping connect the RISC complex (in which the core protein, Ago, is sequence-specifically bound to a target mRNA) to “generic” de-capping and de-tailing complexes that can repress whichever mRNA target RISC is bound to.
If you stick the protein accession code for human TNRC6A into IUPred3, you get a plot where the plotted line is almost entirely above the axis, indicating high likelihood of disorder. UniProt page: https://www.uniprot.org/uniprotkb/Q8NDV7/entry
GW182 is basically like a big old spaghetti-like thing, but many proteins have regions within them that are disordered even though most of the protein has strong secondary structure (things like “set” α-helixes & β-strands – more here: https://bit.ly/proteinstructure ). Flexible linker regions often connect more “rock-hard” structural domains, and when proteins undergo shape-shifts (conformational changes), they often involve hinge-like motions taking place between the domains, with most of the “actual” changing happening in the linker regions.
An example of a protein like this is Ago2 – our main Ago protein. If you check it out in IUPred you see mostly below the midline stuff (predicted structures) with some linkery bits and the N-terminus above it (predicted disordered). And if you go look at crystal structures of Ago (e.g. PDB 4f3t) you can “see” that those regions are missing in the structure.
UniProt page: https://www.uniprot.org/uniprotkb/Q9UKV8/entry; PDB page: https://www.rcsb.org/structure/4f3t
You can also see in both the IUPred plot and missingness in the structure that there’s a stretch near the C-terminus (~820) that’s disordered. And the little pushpin like symbols on the PTM (post-translational modification) track below the plot tell you that there are 5 phosphorylation sites there. It’s very common for PTM sites to be located in IDRs. And this cluster of them in particular is very familiar to me because it was the core of my thesis work! And my publication showing how phosphorylation of it regulates Ago’s activity. I will link to that (open-access so free for anyone to read) paper at the end.
You can also see other annotation tracks below the plot including one that highlights regions that have experimental evidence of disorder. There’s a whole database called DesProt that collates experimentally-collected info on IDRs and IDPs.
And it’s growing as scientists realize more and more ways IDRs can do cool things. In addition to serving as scaffolding, providing numerous binding sites, and allowing sturdier domains to move in relationship to one another, IDR’s can facilitate the formation of molecular condensates – kinda like gooey globs inside of cells that hold functionally-related things together to keep them from diffusing away from one another.
In addition to just looking for the absence of distinct signal in crystallography or cryo-EM data, there are other techniques scientists can use to find out things about IDRs. NMR, for one. NMR can give you info about flexible things, but it only works for small proteins and you need a lot of it. What you end up getting is an ensemble of some of the various conformations (looks like an overlay of a bunch of different shapes, kinda like a picture taken of something in motion).
Another technique is HDX-MS (Hydrogen eXchange – Mass Spectrometry) which uses deuterium (heavy hydrogen) to label more flexible & dynamic regions of proteins. I did it in grad school and have a post on it. http://bit.ly/hdxmassspec
There are also a bunch of functional assays (experiments) you can do to find out more about what IDRs do. And other ways to look at how they move. For example, one technique that is sometimes used to study conformational changes is FRET labeling – for example, sticking a fluorophore on one part of a protein and a quencher on another. When those parts of the protein get close you’ll loose signal, but in a more open conformation you’ll see shining. More on fluorescence and FRET here: http://bit.ly/fretandfluorescence
Now here are a bunch of references and links if you want to know more.
I didn’t talk about it here, but I show p53 as an example:
- the corresponding UniProt accession code (for the human version) is P04637: https://www.uniprot.org/uniprotkb/P04637/entry
- And here’s a relevant article: Wells, M., Tidow, H., Rutherford, T. J., Markwick, P., Jensen, M. R., Mylonas, E., Svergun, D. I., Blackledge, M., & Fersht, A. R. (2008). Structure of tumor suppressor p53 and its intrinsically disordered N-terminal transactivation domain. Proceedings of the National Academy of Sciences of the United States of America, 105(15), 5762–5767. https://doi.org/10.1073/pnas.0801353105
DisProt (the server with experimentally-determined IDR info):
- https://disprot.org/
- Quaglia et al., DisProt in 2022: improved quality and accessibility of protein intrinsic disorder annotation, Nucleic Acids Research, Volume 50, Issue D1, 7 January 2022, Pages D480–D487, https://doi.org/10.1093/nar/gkab1082
IUPred (intrinsic disorder prediction software):
- https://iupred3.elte.hu/
- official paper reference for IUPred3: Dosztányi Z. (2018). Prediction of protein disorder based on IUPred. Protein science : a publication of the Protein Society, 27(1), 331–340. https://doi.org/10.1002/pro.3334
- Nice explanation of how it works; Gábor Erdős, Mátyás Pajkos, Zsuzsanna Dosztányi, IUPred3: prediction of protein disorder enhanced with unambiguous experimental annotation and visualization of evolutionary conservation, Nucleic Acids Research, Volume 49, Issue W1, 2 July 2021, Pages W297–W303, https://doi.org/10.1093/nar/gkab408
- This paper is for IUPred2, but it has nice examples of how to use IUPred: Erdős, G., & Dosztányi, Z. (2020). Analyzing Protein Disorder with IUPred2A. Current protocols in bioinformatics, 70(1), e99. https://doi.org/10.1002/cpbi.99
Reviews comparing different disorder prediction methods:
- Lang, B., Babu, M.M. A community effort to bring structure to disorder. Nat Methods 18, 454–455 (2021). https://doi.org/10.1038/s41592-021-01123-5
- Necci, M., Piovesan, D., CAID Predictors. et al. Critical assessment of protein intrinsic disorder prediction. Nat Methods 18, 472–481 (2021). https://doi.org/10.1038/s41592-021-01117-3
Here are links to some relevant posts & videos of mine for background and/or further information:
- more about the PDB and structures: blog: https://bit.ly/pdbstructures ; full video: https://youtu.be/NgXwP7gGPyA short video: https://youtu.be/1uKC08Z_lYQ
- more on UniProt, ProtParam, etc. blog: https://bit.ly/uniprotprotparam; YouTube: https://youtu.be/6oBsTykEeGI
- more about amino acids and proteins: https://bit.ly/aminoacidsposts & https://www.youtube.com/playlist?list=PLUWsCDtjESrFQoCEsEmZX6NxnwlHzjHZ6
- If you need a refresher on X-ray crystallography, start here: http://bit.ly/xraycrystallography2
- and If you want more structural biology content:
- this page on my blog that has links to all my structural biology posts: https://bit.ly/structural_biology
- and here’s a link to my YouTube structural biology playlist: https://youtube.com/playlist?list=PLUWsCDtjESrGhwVxsRbTJdL-BEsN60RCs
And, finally, here’s my paper I talk about if you were interested: Brianna Bibel, Elad Elkayam, Steve Silletti, Elizabeth A Komives, Leemor Joshua-Tor (2022) Target binding triggers hierarchical phosphorylation of human Argonaute-2 to promote target release eLife 11:e76908 https://doi.org/10.7554/eLife.76908
note: in the video I mention that IDRs tend to be enriched in charged amino acids, but that’s no longer thought to be the case – instead, they commonly have lots of proline, serine and glycine. Apologies!
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: https://thebumblingbiochemist.com
























