Sequence length calculator
DNA, RNA or protein, plain text or FASTA. Counted as you type, nothing leaves the browser.
DNA, RNA or protein, plain text or FASTA. Counted as you type, nothing leaves the browser.
The length is the number of residues in the sequence: every base of a DNA or RNA sequence, or every amino acid of a protein. Everything that is not a residue is ignored before counting: spaces, tabs, line breaks, and the position numbers that GenBank, EMBL and many text editors put at the start of each line. Gap characters (- and .) and the stop sign (*) are also left out and reported separately, so a sequence copied out of an alignment gives its true length.
A sequence pasted as text is a single strand, so its length is reported in nucleotides. For double-stranded DNA such as a plasmid, PCR product or genomic fragment, the same number is the length in base pairs: one base on the top strand pairs with one on the bottom. Protein sequences are reported in amino acids (aa). If a nucleotide length divides by three the number of codons is shown too, which is a quick check that a coding sequence is complete.
Each record is classified from its letters. A sequence made only of A, C, G, T and the IUPAC ambiguity codes is DNA; the same with U instead of T is RNA; anything containing letters that only occur in proteins, such as E, F, I, L, P or Q, is protein. Short sequences that fit both alphabets, for example ACGT, are treated as DNA.
Paste a whole FASTA file. Each record is listed by name with its type, length and GC content, and the summary gives the total length, the average, and the shortest and longest record. Text before the first > header is counted as one unnamed record.
GC content is the share of G and C bases in a nucleotide sequence. It affects melting temperature, secondary structure and how easy a template is to amplify or sequence. It is shown for DNA and RNA records here; for molecular weight, base composition and copy-number conversions use the DNA and RNA length calculator.
Paste the sequence into the box above. Whitespace, line breaks and the position numbers found in GenBank-style text are ignored, so the count is the number of bases only. The length appears as you type.
The count is the number of characters in the sequence you paste, which is nucleotides (nt) for a single strand. For double-stranded DNA the same number is the length in base pairs (bp).
Yes. Protein sequences are detected automatically from their letters and reported in amino acids (aa). Stop signs (*) and gap characters are not counted.
Yes. Paste a multi-record FASTA file and each record is listed with its name, type, length and GC content, followed by the total, average, shortest and longest.
Yes. Every letter is one position, so N, the other IUPAC ambiguity codes and X in a protein each add 1 to the length. Gap characters (- and .) and stop signs (*) are left out and reported separately, and digits and whitespace are ignored.