Scientists (myself definitely included) sometimes get sloppy with our language – including when it comes to complementary DNA (cDNA), which is a DNA copy of messenger RNA (mRNA), which is an edited copy of a gene. For example, often we say we “insert a gene” into a plasmid to get cells (often bacteria) to make a protein for us when what we really are referring to is sticking in the cDNA (and only the coding sequence (CDS) portion of it). We also often talk about “measuring mRNA levels” to see what proteins cells are probably making when we’re really measuring cDNA levels (and assuming they fairly represent the mRNA present). Most of the time, when we’re just talking about things, these nuances don’t matter that much. But sometimes it’s really important to keep in mind the distinctions…
link to video in case embed isn’t working: https://youtu.be/x13Wx9E3H6I
A key time is when you’re trying to do cloning for recombinant protein expression (that thing I mentioned up above). If you stuck the actual “gene” in the plasmid (which would likely be impossible anyway because it would be too long) and put that plasmid into cells, those cells would make jibberish. Because that DNA has a bunch of regulatory information in regions called introns. These introns normally get removed from immature mRNA in a process called splicing so that the protein making machinery (ribosomes) don’t even see them. But the cells aren’t going to be able to splice the plasmid, so the introns would stay in. They’d interrupt the parts with the protein instructions (the exons) and the ribosome would try to translate them (read them and piece together amino acids based on the sequence). This wouldn’t end well…
So, instead of sticking in the gene we stick in the cDNA. But which one? Multiple mRNAs (and thus cDNAs – and protein versions (isoforms)) can be made from the same gene thanks to alternative splicing (you can skip over exons etc.). more here: https://bit.ly/altsplicing ; YouTube: https://youtu.be/lKl87g66Rrk
You need to make sure you stick in the one you want. This might take some sleuthing…
If there’s a protein I’m interested in, I like going to UniProt, which is a database with tons of information about “every” protein, and searching for it. If you scroll down or click on the sequences link you might see that it shows multiple isoforms formed by alternative splicing, each with a name starting with a P. One will be designated “canonical” which is usually, but not always, the main one and thus your safest bet if you want to study it – but be sure to look into the isoforms before diving in!
If you scroll down you will see a part with links to databases. There will be a Consensus CDS (CCDS) link for each of those isoforms. If you click on them it will take you to that isoform’s entry in a database of validated and agreed upon cDNAs. There you will find the cDNA and protein sequences. Now you have the sequence you need – or at least the sequence you can compare available templates to.
To find plasmids containing that cDNA, you can search addgene or DNASU or find a paper that made one and see if you can get some from them. Often people don’t specify the isoform used, so you’ll have to check the sequence provided (and the sequence after you sequence the plasmid to confirm) against the cDNA sequences. Instead of searching directly against the cDNA sequences, you can use a tool like Expasy translate to get the corresponding protein sequence which you can then compare to the isoforms you see in UniProt (I find it much simpler to think in protein land!)
more on UniProt, ProtParam, etc. https://bit.ly/uniprotprotparam ; YouTube: https://youtu.be/6oBsTykEeGI
What if you can’t find the cDNA you want? These days, you might be able to have it synthesized by a company like Twist or IDT. Alternatively, you can go fishing in a cDNA library, which is basically a collection of plasmids containing cDNAs of “all” the mRNAs present in some sample. You can make and use a labeled probe complementary to a sequence in the cDNA to find the plasmid containing that cDNA. Then you can subclone it – stick the cDNA in a different plasmid. This will work even if you don’t know the whole sequence of the cDNA, such as for some obscure species without good sequencing data.
cDNA libraries can be made from different cells or tissues to compare what’s being made where. This is possible because that library should theoretically at least be a good representation of the transcriptome (what mRNA transcripts are present in the sample). And unlike those transcripts, these libraries are more stable and “renewable” since the bacteria can just keep making copies of them. Conventional cDNA libraries have been used to solve many historical medical mysteries, such as cystic fibrosis. More on that here: http://bit.ly/cysticfibrosisscience
But in terms of seeing what’s being expressed, these days it’s much more common to use a technique like RNA seq (RNA sequencing) or RT-qPCR. Although it’s called RNA seq, very rarely (but becoming more common) is RNA itself actually being sequenced. Instead, cDNA is. mRNAs or total RNA is reverse transcribed to get cDNA (more on this below) which is ligated (stitched to) end adapters that allow it to be sequenced. This will tell you “all” that’s present – but if there are only a few things you’re interested in, it’s much simpler to just search for their cDNAs directly using RT-qPCR, where you make lots of copies of it (if it’s present) and use fluorescence to measure the copies as they’re made. The more you start with, the faster the signal will rise so it will tell you about how many copies there were.
Here’s some more detail on this, adapted from a much longer post of it… blog form: http://bit.ly/rtrtqpcrprimer; YouTube: https://youtu.be/kp4ZX2lOr6w
the first step in RT-qPCR (after you isolate the RNA) is making DNA copies of the RNA copies of the DNA recipes through REVERSE TRANSCRIPTION. Normal transcription goes DNA->RNA. REVERSE transcription goes RNA->DNA. It uses a different polymerase (instead of the usual DNA-RNA or DNA-DNA Pols you need an RNA-DNA Pol – we call such Pols reverse transcriptase) – and we call the DNA copies of the mature mRNAs complementary DNA (cDNA)
The reverse transcriptase can make DNA copies of RNA, but it still has the limitation of needing a double-stranded starting platform – so you need to provide primers for it.
Usually you want to measure multiple mRNAs. Even if you’re only interested in levels of one, you need to normalize it to something so you make sure that if you see twice as many of jt under a set of conditions it isn’t just because you had RNA from twice as many cells.
Traditionally this is done by comparing levels of “housekeeping genes” which are recipes that are made at pretty constant levels under all conditions
Since you want to count multiple things, you usually start by stabilizing and reverse-transcribing all the mRNA and/or all the RNA (mRNA or otherwise). To just RT the mRNA you can take advantage of that generic poly-A tail we saw earlier. Since A pairs with T, you can use a short stretch (usually 15) of DNA T’s (an oligo(dT)) as an all-mRNA-specific primer. It’ll latch onto the poly-A tail to provide a starting point for the reverse transcriptase.
If you provide a “normal” oligo-dT primer, it can latch on anywhere along the poly-A tail, but if you use an “anchored oligo-dT” which ends (3’ end) with a letter other than T (a G, C, or A that acts as an anchor) – it can only latch onto the part closest to the end of the unique stuff (binds at the 5’ end of the poly(A) tail.
To illustrate: imagine you have an mRNA that’s unique part ends in a C
blahblahblahCAAAAAAAAAAAAAAAAAAAA
If you use and un-anchored oligo(dT) like TTTTTTTT that can bind anywhere along the stretch of As and serve as a primer for the reverse transcriptase. So you can get
<——————————TTTTTTTT
blahblahblahCAAAAAAAAAAAAAAAAAAAA
or
<——————————————TTTTTTTT
blahblahblahCAAAAAAAAAAAAAAAAAAAA
etc. But if you use anchored oligo(dT)s where you have a mix of ATTTTTTTT, CTTTTTTTT, & GTTTTTTTT, only the G-version can bind and it can only bind in one spot
<——————GTTTTTTTT
blahblahblahCAAAAAAAAAAAAAAAAAAAA
oligo-dT primers are good if you don’t have much RNA to start with – but there are some disadvantages – like it can sometimes prime internal poly(A) sites (if an RNA happens to have a stretch of As before the end the primer can stick to the center of the recipe “thinking it’s the tail end” so you get a truncated (end-lobbed-off) version of the recipe, not the whole thing – and it can also bind to RNAs that aren’t mRNAs they just happen to have a lot of As (like some rRNAs (ribosomal RNAs)).
You can also have the “opposite problem” – some mRNAs like the ones for histone proteins are weird and un-tailed. And if you want to look at expression of things things like tRNA, rRNA, & noncoding RNAs, oligo oligo(dT)s won’t help you.
If you want to reverse transcribe total RNA (not just the mRNA but also tRNA, rRNA, noncoding RNA, etc.) you can use random oligo-dTs. These are short (6-9 nt) random sequences that are short enough that they’re likely to be found in lots of genes and you have enough variety that all the genes are likely to find several matches. So you end up with pieces of the recipe copied, not the full-length thing.
Even if you’re thing is tailed & even if you use anchored oligo(dT)s so you’re as close to the unique stuff as possible, mRNAs can be REALLY LONG and if you want to detect something further towards an mRNA’s start, random oligos or a mix of random & oligo-dT might be a better option.
If you only have a few you want to test & you want high sensitivity (be able to detect tiny amounts) you can use sequence specific primers that are designed to mach the mRNA you want to look for and are longer so those sequence aren’t found other places (unlike the shorter random ones)
After you make the cDNA copies of “everything” or at least all the mRNAs, you want to make copies ONLY of the one you’re interested in. So instead of aiming for genericness, you want your primers to be super specific. And you’ll need primerS now Because now we want to amplify – with reverse transcription we only made 1 strand of cDNA, but now we want to make lots of copies. So we need to make a second strand from that first strand and then we can use those strands as templates for the other strands so you can make more and more and more…
Since you’re now just copying DNA to DNA, you can use a “normal” DNA Pol. And just like in normal PCR, qPCR is performed in cycles of temperature changes – melt (heat up to separate strands) → anneal (cool down to let primers bind & Pol latch on) → extend (let Pol lay down complementary track) → repeat. More on PCR: http://bit.ly/pcrtrain & https://youtu.be/GZSLfECgW3Q
More on gene terminology: Gene-related jargon: exons, introns, UTRs, pre-mRNA, mature mRNA, CDS (coding sequence), & cDNA
more about all sorts of things: #365DaysOfScience All (with topics listed) 👉 http://bit.ly/2OllAB0 or search blog: https://thebumblingbiochemist.com
















