Reverse translate

Protein to DNA. The codons your host prefers, or one degenerate sequence.

Result

What does reverse translation do?

It turns a protein sequence back into DNA. Because 61 codons encode 20 amino acids, the answer is not unique: every residue except methionine and tryptophan has a choice of codons, so a protein of length n has up to 6n coding sequences. This tool gives the two answers that are actually useful. Most likely codon writes one concrete sequence using the codon the chosen host uses most often for each residue, which is where gene synthesis and codon optimisation start. Degenerate IUPAC writes one sequence with ambiguity codes that covers every possible coding sequence, which is how degenerate PCR primers are designed from a peptide.

Most likely codon = the synonymous codon with the highest frequency in the host
Degenerate codon = the IUPAC code that covers all synonymous codons of the residue

Worked example

"Try an example" loads the 18-residue signal peptide MKWVTFISLLFLFSSAYS. ForE. coli the most likely codons give 54 nt at 44 % GC:

ATGAAATGGGTGACCTTTATTAGCCTGCTGTTTCTGTTTAGCAGCGCGTATAGC

For human codon usage the same peptide becomes ATGAAGTGGGTGACCTTCATCAGCCTGCTGTTCCTGTTCAGCAGCGCCTACAGC at 56 % GC, and for yeast ATGAAATGGGTTACTTTTATTTCTTTGTTGTTTTTGTTTTCTTCTGCTTATTCT at 26 % GC. The same peptide in degenerate form is:

ATGAARTGGGTNACNTTYATHWSNYTNYTNTTYYTNTTYWSNWSNGCNTAYWSN

That one sequence stands for 206,158,430,208 oligonucleotides, which is why a degenerate primer is designed from five or six well-chosen residues and never from a whole peptide.

Codon choice by host

Frequencies are per thousand codons from the Kazusa Codon Usage Database. The percentage is the share of that codon among the synonymous codons for the residue.

Amino acidE. coliYeastHumanDegenerate
Ala (A)GCG 36 %GCT 38 %GCC 40 %GCN
Arg (R)CGC 40 %AGA 48 %AGA 22 %MGN
Asn (N)AAC 55 %AAT 59 %AAC 53 %AAY
Asp (D)GAT 63 %GAT 65 %GAC 54 %GAY
Cys (C)TGC 56 %TGT 63 %TGC 54 %TGY
Gln (Q)CAG 65 %CAA 69 %CAG 74 %CAR
Glu (E)GAA 69 %GAA 70 %GAG 58 %GAR
Gly (G)GGC 41 %GGT 47 %GGC 34 %GGN
His (H)CAT 57 %CAT 64 %CAC 58 %CAY
Ile (I)ATT 51 %ATT 46 %ATC 47 %ATH
Leu (L)CTG 50 %TTG 29 %CTG 40 %YTN
Lys (K)AAA 77 %AAA 58 %AAG 57 %AAR
Met (M)ATGATGATGATG
Phe (F)TTT 57 %TTT 59 %TTC 54 %TTY
Pro (P)CCG 53 %CCA 42 %CCC 32 %CCN
Ser (S)AGC 28 %TCT 26 %AGC 24 %WSN
Thr (T)ACC 44 %ACT 35 %ACC 36 %ACN
Trp (W)TGGTGGTGGTGG
Tyr (Y)TAT 57 %TAT 56 %TAC 56 %TAY
Val (V)GTG 37 %GTT 39 %GTG 46 %GTN
Stop (*)TAA 65 %TAA 48 %TGA 47 %TRR

Three residues have six codons in two separate families. Their minimal degenerate codons reach beyond the residue: MGN also covers Ser AGT and AGC, YTN also covers Phe TTT and TTC, WSN covers sixteen codons of which only six are Ser, and TRR also covers Trp TGG. The option above the result switches them to CGN, CTN and TCN, which stay inside one family and code for nothing else. Ambiguous residues are handled too: B (Asx) becomes RAY, Z (Glx) SAR, J (Xle) HTN and X NNN.

After back translating

Check the result before ordering it. The rare codon analysisflags codons the host reads slowly, GC content should sit near the host average, and DNA to protein translation confirms the sequence reads back to the protein you started with. For a degenerate sequence, theIUPAC expander lists the individual oligonucleotides and counts them, and the codon table gives the full genetic code.

Frequently asked questions

What is reverse translation?

Working out a DNA sequence that would encode a given protein. The genetic code is redundant, so a protein of 100 residues has billions of possible coding sequences. A reverse translator picks one of them: either the codon each host uses most often, or one degenerate sequence written with IUPAC codes that covers them all.

Which codon usage table is used?

The Codon Usage Database at the Kazusa DNA Research Institute, compiled from GenBank coding sequences: Escherichia coli W3110 (a K-12 strain), Saccharomyces cerevisiae and Homo sapiens. The same tables are used by the rare codon analysis tool on this site.

Is the most frequent codon the best choice for expression?

Not always. Using one codon everywhere can drain a single tRNA pool, remove the pauses that help folding, and create long repeats that are hard to synthesise. Commercial optimisation instead samples codons in proportion to their frequency and avoids unwanted restriction sites, hairpins and rare-codon runs. The most frequent codon is a sound starting point and a fair worst case to check.

Why are Arg, Leu and Ser written as MGN, YTN and WSN?

They have six codons each, spread over two codon families, so no single degenerate codon covers exactly those six. MGN also covers the two Ser codons AGT and AGC, YTN also covers the two Phe codons, and WSN covers ten codons that are not Ser. Switch the option to CGN, CTN and TCN to stay inside one family, which is exact but misses AGA, AGG, TTA, TTG, AGT and AGC.

How do I use this for a degenerate primer?

Take the least degenerate stretch of the protein, usually five or six residues rich in Met, Trp, Cys, Asp, Glu, Phe, His, Lys, Asn, Gln and Tyr, and order the degenerate sequence for that stretch. Avoid Arg, Leu and Ser. The IUPAC expander lists every oligonucleotide in the mixture and the exact count.