Codon optimizer
Rewrites a protein or a coding sequence in the codons your expression host uses.
Rewrites a protein or a coding sequence in the codons your expression host uses.
Paste either a protein sequence or a coding DNA sequence. DNA is translated in frame from its first base with the standard genetic code, and the protein is then written back into DNA one amino acid at a time using the codon usage table of the chosen host. Two rules are on offer:
With the checkbox on, the finished sequence is then scanned for the recognition sites of fourteen common cloning enzymes and for runs of seven or more identical bases, and each one is removed by the cheapest synonymous change in the codons that overlap it.
L counts only the codons with a synonymous choice, so methionine, tryptophan and the stop codon are left out. CAI is reported for the optimised sequence, and also for the input when the input was DNA, so you can see how far the gene had to move.
"Try an example" loads the 720 nt coding sequence of EGFP, a gene written for mammalian cells, on the E. coli tab with "Most used codon". EGFP encodes 239 residues and a stop. Its CAI for E. coli is 0.696; the optimised gene reaches 0.995, with122 of the 240 codons changed and the GC content falling from 61.5 % to 49.4 %. Writing every amino acid as its top codon creates two NdeI sites and two runs of seven and eight T, and removing those takes 4 further codon changes, which is what costs the last 0.005 of CAI. Switch the tab to Human and only 27 codons change and CAI goes from 0.959 to 1.000, because EGFP was already written for human cells.
It is not worth doing when the native gene already fits the host, when the protein needs its native translation rhythm to fold, or when the problem is toxicity, inclusion bodies or proteolysis, none of which codons can fix.
The Codon Usage Database at the Kazusa DNA Research Institute (Nakamura, Gojobori and Ikemura, Nucleic Acids Res 2000), compiled from GenBank coding sequences:Escherichia coli W3110 (1,372,057 codons), Saccharomyces cerevisiae(6,534,504 codons) and Homo sapiens (40,662,582 codons). The same tables drive thecodon usage table and therare codon analysis.
For the reverse direction, a degenerate back-translation that covers every codon rather than one, use reverse translate. To check the protein a DNA input encodes, use the DNA to protein translator. To design a silent change at one position only, usemutate for digest.
Rewriting a gene so that it encodes the same protein but uses the codons the expression host reads best. Only synonymous positions change, so the amino acid sequence is untouched. It is used when a gene from one organism is expressed in another and translation is slow, stalls on rare codons, or gives truncated product.
It gives the highest codon adaptation index, and it is what most commercial genes do, but it is not always best. Filling a gene with one codon per amino acid can drain the matching tRNA pool, remove the slow patches that help some proteins fold, and create long runs of the same base that are hard to synthesise. "Match host usage" spreads the codons in the proportions the host itself uses, which is the safer choice for large or difficult proteins.
The codon adaptation index of Sharp and Li (1987). Each codon gets a weight w equal to its frequency divided by the frequency of the most used codon for the same amino acid, and CAI is the geometric mean of those weights over the gene. Methionine, tryptophan and stop codons are left out because they have no synonymous choice. CAI runs from just above 0 to 1, and a highly expressed native gene is typically 0.7 to 0.9.
The recognition sites of fourteen common cloning enzymes: EcoRI, BamHI, HindIII, XhoI, NdeI, XbaI, SalI, NotI, PstI, KpnI, SacI, NcoI, BglII and SpeI. A site is removed by the synonymous change that keeps the highest codon frequency. A site that overlaps the start codon cannot always be removed, and any that remain are reported.
No. Codon choice is one of several limits on expression, alongside mRNA folding near the ribosome binding site, transcript stability, plasmid copy number, induction conditions and the protein itself. Optimisation helps most when the native gene is full of codons the host uses rarely, which the rare codon analysis will tell you before you order anything.