Reverse translate
Protein to DNA. The codons your host prefers, or one degenerate sequence.
Protein to DNA. The codons your host prefers, or one degenerate sequence.
It turns a protein sequence back into DNA. Because 61 codons encode 20 amino acids, the answer is not unique: every residue except methionine and tryptophan has a choice of codons, so a protein of length n has up to 6n coding sequences. This tool gives the two answers that are actually useful. Most likely codon writes one concrete sequence using the codon the chosen host uses most often for each residue, which is where gene synthesis and codon optimisation start. Degenerate IUPAC writes one sequence with ambiguity codes that covers every possible coding sequence, which is how degenerate PCR primers are designed from a peptide.
"Try an example" loads the 18-residue signal peptide MKWVTFISLLFLFSSAYS. ForE. coli the most likely codons give 54 nt at 44 % GC:
For human codon usage the same peptide becomes ATGAAGTGGGTGACCTTCATCAGCCTGCTGTTCCTGTTCAGCAGCGCCTACAGC at 56 % GC, and for yeast ATGAAATGGGTTACTTTTATTTCTTTGTTGTTTTTGTTTTCTTCTGCTTATTCT at 26 % GC. The same peptide in degenerate form is:
That one sequence stands for 206,158,430,208 oligonucleotides, which is why a degenerate primer is designed from five or six well-chosen residues and never from a whole peptide.
Frequencies are per thousand codons from the Kazusa Codon Usage Database. The percentage is the share of that codon among the synonymous codons for the residue.
| Amino acid | E. coli | Yeast | Human | Degenerate |
|---|---|---|---|---|
| Ala (A) | GCG 36 % | GCT 38 % | GCC 40 % | GCN |
| Arg (R) | CGC 40 % | AGA 48 % | AGA 22 % | MGN |
| Asn (N) | AAC 55 % | AAT 59 % | AAC 53 % | AAY |
| Asp (D) | GAT 63 % | GAT 65 % | GAC 54 % | GAY |
| Cys (C) | TGC 56 % | TGT 63 % | TGC 54 % | TGY |
| Gln (Q) | CAG 65 % | CAA 69 % | CAG 74 % | CAR |
| Glu (E) | GAA 69 % | GAA 70 % | GAG 58 % | GAR |
| Gly (G) | GGC 41 % | GGT 47 % | GGC 34 % | GGN |
| His (H) | CAT 57 % | CAT 64 % | CAC 58 % | CAY |
| Ile (I) | ATT 51 % | ATT 46 % | ATC 47 % | ATH |
| Leu (L) | CTG 50 % | TTG 29 % | CTG 40 % | YTN |
| Lys (K) | AAA 77 % | AAA 58 % | AAG 57 % | AAR |
| Met (M) | ATG | ATG | ATG | ATG |
| Phe (F) | TTT 57 % | TTT 59 % | TTC 54 % | TTY |
| Pro (P) | CCG 53 % | CCA 42 % | CCC 32 % | CCN |
| Ser (S) | AGC 28 % | TCT 26 % | AGC 24 % | WSN |
| Thr (T) | ACC 44 % | ACT 35 % | ACC 36 % | ACN |
| Trp (W) | TGG | TGG | TGG | TGG |
| Tyr (Y) | TAT 57 % | TAT 56 % | TAC 56 % | TAY |
| Val (V) | GTG 37 % | GTT 39 % | GTG 46 % | GTN |
| Stop (*) | TAA 65 % | TAA 48 % | TGA 47 % | TRR |
Three residues have six codons in two separate families. Their minimal degenerate codons reach beyond the residue: MGN also covers Ser AGT and AGC, YTN also covers Phe TTT and TTC, WSN covers sixteen codons of which only six are Ser, and TRR also covers Trp TGG. The option above the result switches them to CGN, CTN and TCN, which stay inside one family and code for nothing else. Ambiguous residues are handled too: B (Asx) becomes RAY, Z (Glx) SAR, J (Xle) HTN and X NNN.
Check the result before ordering it. The rare codon analysisflags codons the host reads slowly, GC content should sit near the host average, and DNA to protein translation confirms the sequence reads back to the protein you started with. For a degenerate sequence, theIUPAC expander lists the individual oligonucleotides and counts them, and the codon table gives the full genetic code.
Working out a DNA sequence that would encode a given protein. The genetic code is redundant, so a protein of 100 residues has billions of possible coding sequences. A reverse translator picks one of them: either the codon each host uses most often, or one degenerate sequence written with IUPAC codes that covers them all.
The Codon Usage Database at the Kazusa DNA Research Institute, compiled from GenBank coding sequences: Escherichia coli W3110 (a K-12 strain), Saccharomyces cerevisiae and Homo sapiens. The same tables are used by the rare codon analysis tool on this site.
Not always. Using one codon everywhere can drain a single tRNA pool, remove the pauses that help folding, and create long repeats that are hard to synthesise. Commercial optimisation instead samples codons in proportion to their frequency and avoids unwanted restriction sites, hairpins and rare-codon runs. The most frequent codon is a sound starting point and a fair worst case to check.
They have six codons each, spread over two codon families, so no single degenerate codon covers exactly those six. MGN also covers the two Ser codons AGT and AGC, YTN also covers the two Phe codons, and WSN covers ten codons that are not Ser. Switch the option to CGN, CTN and TCN to stay inside one family, which is exact but misses AGA, AGG, TTA, TTG, AGT and AGC.
Take the least degenerate stretch of the protein, usually five or six residues rich in Met, Trp, Cys, Asp, Glu, Phe, His, Lys, Asn, Gln and Tyr, and order the degenerate sequence for that stretch. Avoid Arg, Leu and Ser. The IUPAC expander lists every oligonucleotide in the mixture and the exact count.