GC content calculator
GC and AT percentage, base composition and a sliding-window plot. Paste DNA, RNA or FASTA.
GC and AT percentage, base composition and a sliding-window plot. Paste DNA, RNA or FASTA.
GC content is the share of bases in a DNA or RNA sequence that are guanine or cytosine, written as a percentage. AT content is the rest. Because a G–C pair is held by three hydrogen bonds and an A–T pair by two, and because G–C pairs stack more strongly, GC content sets how stable a double helix is and therefore its melting temperature. It also varies widely between organisms and along a single genome, which makes it one of the first things to look at in any new sequence.
Worked example. A 200 bp fragment with 45 G and 57 C contains 102 GC bases, so its GC content is 102 ÷ 200 × 100 = 51.0 % and its AT content is 49.0 %. The calculator counts S (G or C) as GC and W (A or T) as AT; every other IUPAC ambiguity code is reported separately and excluded from the percentage. Whitespace, line breaks and the digits of GenBank-style numbering are ignored, and U in RNA is counted with the AT bases.
| Organism | Genome GC |
|---|---|
| Plasmodium falciparum | ~19 % |
| Arabidopsis thaliana | ~36 % |
| Saccharomyces cerevisiae | ~38 % |
| Homo sapiens | ~41 % |
| Drosophila melanogaster | ~42 % |
| Escherichia coli K-12 | ~51 % |
| Mycobacterium tuberculosis | ~66 % |
| Streptomyces coelicolor | ~72 % |
Bacterial genomes span the whole range from below 20 % to above 75 %. Vertebrate genomes sit near 40 % overall but are a mosaic of long GC-rich and GC-poor domains, the isochores, which is what the sliding-window plot reveals.
The plot slides a window of fixed width along the sequence and reports the GC content of each window, so local composition becomes visible where the single overall percentage hides it. A narrow window, 20–50 bp, shows individual GC-rich motifs and primer sites; a wide window, 500 bp or more, smooths noise and shows isochores, islands and transferred regions. The dashed line is the mean for the whole sequence. Choose the window with the control above the plot, or leave it on Auto, which picks a width suited to the sequence length.
GC content asks how much G and C a region holds. GC skew asks whether G or C dominates on one strand, (G − C) ÷ (G + C), and it changes sign at the origin and terminus of replication in bacterial chromosomes. Use the GC skew calculator for that question. For length, molecular weight and GC in one report, see theDNA and RNA length calculator.
Count the G and C bases, divide by the total number of bases, and multiply by 100. A 200 bp sequence with 45 G and 57 C has a GC content of (45 + 57) / 200 × 100 = 51%. AT content is the remainder, 49%.
It depends on the organism. The human genome averages about 41%, E. coli about 51%, yeast about 38%, Plasmodium falciparum only about 19% and Streptomyces over 70%. Within a genome, genes and CpG islands are usually GC-richer than their surroundings.
G–C pairs have three hydrogen bonds and stack more strongly than A–T pairs, so GC-rich DNA melts at a higher temperature. Primers are usually designed at 40–60% GC, and templates above about 65% GC often need additives such as DMSO or betaine to amplify.
S (G or C) counts as GC and W (A or T) counts as AT. All other IUPAC codes (N, R, Y, K, M, B, D, H, V) are reported separately and left out of the percentage, so an N-rich draft assembly does not distort the result.
Yes. U is counted with A and T as an AT base, so the GC percentage of an RNA sequence is computed the same way as for DNA.