When you order a cDNA plasmid (the DNA copy of the mature messenger RNA (mRNA) instructions for making a protein) from a repository (e.g. addgene, DNASU), it will likely come in some sort of generic cloning plasmid. And you will want to “subclone” it into an expression plasmid if you want to use it to make protein from it (i.e. move it out of the one plasmid and into another).
There can be a lot of options and weird nomenclature and stuff when it comes to plasmids, but the key things they differ in (which we’ll get into in more depth after the overview) are usually:
- copy number – how many copies of the plasmid will each cell host (high’s good for cloning, lower for expression)
- promoter usage – how will you tell the cells to use an RNA Polymerase to make mRNA copies of the gene (and subsequently protein from those mRNA instructions) (T7 (w or w/o lac), tac, T5)
- inducible expression of the RNA Pol? (e.g. induce T7 expression with IPTG)
- selection markers (typically antibiotic resistance genes for Amp, Kan, Strep, etc.)
- potentially secretion signals
- epitope tags (His, Flag, etc., which can be at the N-terminus (start of the protein) or C-terminus (end of the protein))
- often with protease cleavage sites (TEV, HRV3C, thrombin, etc.) for removal from protein
- restriction enzyme cut sites for cloning with restriction enzymes (often there are multiple of them in a region called a multiple cloning site (MCS))
- the different letters after a plasmid number (e.g. pET28a vs pET28b vs. pET28c) usually refer to what reading frame the inserted protein will be read in with respect to the cut site
In terms of reading frames, which plasmid you start with is less of an issue if you’re using a PCR-based cloning strategy (such as SLIC). PCR-based strategies are also good because they let you do “scarless” cloning – you don’t have extra letters on the ends of your proteins that come from having some of the MCS still there. PCR-based strategies are also great because you can easily clone in different tags and things, and do things like swap the tag from the N- to the C-terminus or vise versa, which can sometimes make a difference in terms of tag accessibility, protein folding, etc.. More on cloning methods here: http://bit.ly/molecularcloningguide
now some more details
Copy Number:
We’re demanding lots of protein BUT the protein factories will only increase the supply if we make those demands known! (and they have enough resources to meet them). How many flyers should we put up? More ORIGINal content on PLASMID COPY NUMBER from the bumbling biochemist!
Note: analogy’s not perfect and I wrote this next little part years ago so it repeats a little
We designed a flyer for this cool protein (a circular piece of DNA called a PLASMID containing the gene for that protein) & we put it into bacterial host cells. The host cells have all the machinery we need to make copies of the flyer (replicate the plasmid) &make the protein (translate it). But the host cells also have to make all their own stuff, so we have to convince them to work on our stuff too. They’ll only supply our protein if there’s adequate demand. But the demand doesn’t come from the flyers themselves, it comes from the people seeing the flyers & calling the posted number to request the product.
The products bought are made “on demand” & to some extent, the more flyers there are, the more likely it is that people will see them & call. & the more people that call the better the chances the factory will make more.BUT you don’t want to waste all of your energy making copies if that would lead you to less energy available for answering the calls & making the product!
High copy number plasmids are like companies that focus on making lots of copies of the flyer. This is useful in cloning where the flyer itself is your “product” – bc you’re then going to isolate those flyers & distribute them in other cells (or send them out for editing, etc.)
BUT, when it comes to “answering calls” they have to compete for the host’s call center operators (ribosomes & tRNAs). The more of the flyer, the more likely the calls that come in will be for that flyer.
BUT on the other hand, you don’t want to waste all your metabolic (molecule-building) energy answering calls, leaving less to make the product.
You can use a lower copy number (fewer flyers) but more “effective” flyers in terms of getting people to call & convincing the factory to make more. BUT if you try to crank out too much product, quality control goes out the window – the factory can’t keep up & the proteins misfold & aggregate -> clump together into insoluble “inclusion bodies”
So for cloning we use a high copy number plasmid because we don’t care if people call – in fact, we’d rather they didn’t because we don’t want to waste energy making things we don’t want. But then for expression, we don’t need as high a copy number (don’t need as strong of an ORI) but we do need lots of mRNA made (need a strong PROMOTER) & more (properly-folded) protein made
We can get the host cells to make more copies of the flyer by using a “relaxed” ORIGIN OF REPLICATION (ORI) (the sequence in the DNA that tells DNA Pol to unzip the double-stranded DNA & copy each to give you an identical copy. Plasmids have their own ORI so they can replicate independently of the host (the host only replicates right before dividing, and the plasmids don’t want to wait). So host cells can hold lots of copies of the plasmid, and the average # of plasmid copies per cell (COPY NUMBER) depends on the plasmid, and especially on its ORI. Some produce lots of copies, others just a few because their ORIs are regulated differently -> “relaxed” (which are positively regulated by RNA) will often give you lots of copies, “stringent” (which are positively regulated by availability of host expression proteins) just a few
It comes down to a balance between + regulatory factors (make more copies!) and – regulatory factors (make fewer copies!) which depends on the sequence of the ORI (what copying suggestions you put on the flyer) & what factors like to bind it (what workers in the cellular factories recognize & read those instructions). And those regulatory factors can be really picky! Just 2 mutations in the pMB1 ORI gives you the pUC ORI which makes ~700 copies/cell as opposed to 20.
high copy number plasmids (usually good for cloning, not good for protein expression as we’ll get into in a second) include: pUC (~500-700 copies), pBluescript (~300-500), pGEM (~300-500)
low copy number plasmids (usually not good for cloning, better for protein expression) include: pET, pGEX, and pBR322 (the parent from which those other two are derived) – all of these have ~15-20 copies/cell
Numbers from: https://blog.addgene.org/plasmid-101-origin-of-replication
So, we can get them to make more copies of the plasmid, but that’s not all that matters.
Having more copies burdens the bacteria – this causes them to grow slower and they might be overtaken by plasmid-less cells that escaped selection. If your protein is toxic to the cell, it’s better to have more factories and have each factory make less than to try to have each factory make lots. Even if you don’t have “too many” flyers, you can plaster as many copies of a flyer up on the walls as you want, but if you forget to include the phone number no one will call!
How often people call (how frequently it’s transcribed into mRNA & then translated into protein) depends on its PROMOTER STRENGTH.
Promoters
the PROMOTER is where RNA Pol* will start making mRNA from. It’s not just our protein advertised on our flyer -> The flyer advertises other products as well (it has multiple genes) – when the plasmid gets copied, the whole thing gets copied – but each product has its own number to call (different promoters).
*The promoter is where RNA Pol will get going, but *which* RNA Pol? The endogenous one (the bacteria’s own) or an exogenous one (one that you’re introducing)? Some plasmids rely on the bacteria’s RNA Pol. Examples include plasmids with a T5 promoter. T5 is a bacteriophage, but the T5 promoter is recognized by the E. coli RNA Pol.
examples of plasmids with T5 promoters are: pQE-1 & pQE-60
Most of the plasmids you work with however will use a T7 promoter, which only is recognized by T7 RNA polymerase, *not* by the E. coli host one. This lowers endogenous competition for it, which is great for making lots of transcripts (and subsequently lots of protein from those transcript instructions) and it also allows you to control when your gene gets expressed, by controlling when T7 gets expressed.
It’s like we write the phone number for our protein in invisible ink and then, when we’re ready, turn on the UV light & reveal the message! And demand soars! Though you still have the issue of how persuasive those callers are, which opens up a whole can of translational regulatory worms…
examples of plasmids using T7 promoters include the pET series and pBluescript series
Usually, the T7 promoter is provided from the bacterial strain you transform the plasmid into when you want it expressed. So you have expression cells and cloning cells as well as plasmids that are better for cloning vs expression.
Strains with (DE3) in their name have the “λDE3 lysogen” in them which is just a sequence from a phage with the instructions for making T7 RNA polymerase under control of a lac promoter. In this setup, the T& RNAP instructions are “muted” by a lac operon. This is a sequence upstream of the gene that gets bound by a lac repressor protein, preventing transcription (and subsequent translation) of the polymerase – until you add IPTG (which mimics allolactose, which is a molecule formed when bacteria have lots of lactose and therefore want to turn on the expression of lactose-metabolizing machinery). This allows you to control expression of T7 and, consequently, expression of your gene that’s under the T7 promoter control.
Sometimes, the gene you’re expressing from the T7 promoter will also have a lac promoter in front of them – this makes expression even tighter, so you get less “leaky expression” from small amounts of T7 that get made even without induction (i.e. T7 that is constitutively expressed at low levels).
If low level expression is a big problem, you can also use pLys cells, which express a low level of a lysozyme that inhibits the low level of T7 that’s getting made when you don’t want it. This can be useful if your protein is toxic to the cells or something.
Some plasmids (e.g. pGEX & pMAL) have a “tac” promoter which is a hybrid of a Trp & a lac promoter – both of which use the bacteria’s own RNA Pol. So, you don’t need to provide a separate polymerase, so you don’t need fancy cells, but can still induce expression. But because this expression induction is direct – no T7 intermediary – you have more chances for leakage. So it’s not good to use this if your protein is toxic to the cells.
Other things
plasmids will often have a bunch of notation with things like Δsomething (Δ is delta and it means something is missing, missing things might also be indicated with a superscripted – sign)
- for example, recA– cells are deficient in an E. coli repair system – this system carries out homologous recombination, which can shuffle things around, which you might not want
- an endA mutation makes cells endonuclease I deficient – they don’t make a nonspecific endonuclease (nucleic acid cutter) that’s E. coli normally keep in their periplasmic space – this mutation can keep your plasmid from getting degraded
another thing to look for is the potential for blue-white screening: https://bit.ly/bluewhitescreening
some bacterial plasmids are designed for blue-white screening, including: pGEM-T, pUC18 and pUC19, & pBluescript
but the host cells need to be compatible with it too – some that are: XL1-Blue, DH5α, DH10B, JM109, STBL4, JM110, & Top10
Finding cDNA
When looking for a cDNA clone, the first place I normally look is Addgene. Although it sounds like it would be some commercial entity just out there to make a profit, it actually serves as a nonprofit plasmid repository – labs can send a sample of their plasmids to and Addgene will propagate them (make more copies) and distribute them to the public for a minimal fee ($75/plasmid when I just ordered some).
Addgene is a great place to start because, if authors have deposited the plasmids they used for protein expression, and that expression was in the system you want (e.g bacterial expression not a mammalian expression vector) then you won’t even have to subcclone! and the plasmid might be optimized for good expression – or at least you know it *should* work). If you can only find the gene cloned into a vector for a different expression system, don’t worry – you can just subcclone it into one you want – more work, but shouldn’t be an issue – and definitely not worth paying commercial companies an arm and a leg to get a version in the plasmid you want.
note: Addgene also has a really great educational blog – I suggest checking out their Plasmids 101 guide https://www.addgene.org/educational-resources/ebooks/
Another source which you can turn to if Addgene turns up short is DNASU, which is a depository that has plasmids containing cDNAs for “all” genes, even those that people haven’t worked with already. These are generally in generic cloning vectors but you can easily subcclone them out. Note: you might see different splice versions or alternative transcripts for the same gene, so you might need to do a little looking in UniProt, etc. to make sure you get the version you want. Note 2: some plasmids are marked “fusion” and others “closed” – “fusion” ones don’t have stop codons – these must be supplied by the vector you’re cloning into, “closed” ones do have stop codons
Addgene: https://www.addgene.org/
DNASU: https://dnasu.org/DNASU/Home.do
Final notes
I’ve mainly used pET vectors, chiefly pET28 vectors. I’ve always been curious whether these were really optimal for expression or just what people started using (in the 1980s starting with Studier et al.) and then people used them because others used them and companies (Novagen & Invitrogen) starting making lots of variants of them. So people started buying them and… When I was looking into the matter, I came across this paper from Patrick Shilling et. al describing an optimized pET vector – apparently the original one, though great, had some design flaws including a truncated T7 promoter (compared to the “consensus” one that’s naturally used) and a suboptimal translation initiation region (TIR). They restored the T7 promoter to its full length & tested out a bunch of TIR sequences and found a version that was best for expression of the reporter gene they were using in their test (superfolder GFP). And they deposited this plasmid (pET28a T7pCONS TIR-2 sfGFP) in Addgene if anyone wants to test it out (you can also try cloning the changes into other pET vectors): https://www.addgene.org/154464/
And here’s that paper: Shilling, P.J., Mirzadeh, K., Cumming, A.J. et al. Improved designs for pET expression plasmids increase protein production yield in Escherichia coli. Commun Biol 3, 214 (2020). https://doi.org/10.1038/s42003-020-0939-8
As well as the original Studier et al. paper: Rosenberg, A. H., Lade, B. N., Chui, D. S., Lin, S. W., Dunn, J. J., & Studier, F. W. (1987). Vectors for selective expression of cloned DNAs by T7 RNA polymerase. Gene, 56(1), 125–135. https://doi.org/10.1016/0378-1119(87)90165-x
for more information:
GenScript vector list – a good table of common vectors used for mammalian, bacterial, insect, & yeast expression: https://www.genscript.com/expression-vector-selection-guide.html
Plasmids 101: Origin of Replication, Kendall Morgan, 2020 https://blog.addgene.org/plasmid-101-origin-of-replication
Plasmids 101: Stringent Regulation of Replication, Jason Niehaus, 2015 https://blog.addgene.org/plasmids-101-stringent-regulation-of-replication


















