GC skew calculator
GC skew in sliding windows, the cumulative curve, and the origin and terminus it predicts.
GC skew in sliding windows, the cumulative curve, and the origin and terminus it predicts.
Whether one strand of DNA carries more guanine than cytosine. Over a whole genome the two bases are almost equally common, as base pairing requires, but on a single strand they are not, and the imbalance changes sign at two points on a bacterial chromosome. Those two points are the origin and the terminus of replication.
The windowed skew shows local strand composition and is noisy. The cumulative curve adds +1 for every G and −1 for every C from the start of the sequence, which turns the sign change into a clear V shape: the curve falls while the forward strand is C-rich and rises while it is G-rich. Its lowest point is the predicted origin and its highest point the predictedterminus. The cumulative curve uses no window at all, so the prediction does not depend on the window and step you choose.
"Try an example" loads a simulated 20 kb chromosome whose origin was placed at 4,000 bp and whose terminus was placed at 14,000 bp: the forward strand is G-rich between them and C-rich outside them, with a skew of about 0.15 and the rest of the sequence random. With the automatic settings, a 500 bp window and a 100 bp step, the tool reports the cumulative minimum at3,998 bp and the cumulative maximum at 13,959 bp, within a few tens of bases of the truth, over 196 windows. The window plot swings between roughly −0.28 and +0.30 while the cumulative curve traces a clean V. Real chromosomes are messier: in E. coli K-12 the cumulative minimum sits within a few kilobases of oriC.
During replication the leading strand is synthesised continuously while the lagging strand is made in Okazaki fragments, so its template spends much longer single-stranded. Single-stranded cytosine deaminates to uracil far faster than paired cytosine, and the repaired product is thymine, so cytosines are lost from the lagging-strand template. The result, seen from the forward strand of the published sequence, is a G excess between origin and terminus and a C excess on the way back. Transcription-coupled mutation and the strand bias of coding genes, most of which point away from the origin, add to the same signal.
The window is the stretch of sequence each skew value is measured over, and the step is how far the window moves between values. A small window shows local detail and a lot of noise, a large one smooths the curve and blurs the switch. Left empty, the window is about 2 % of the sequence and the step a fifth of the window, which gives roughly 250 points. The statistics use every window; the plot is thinned to at most 600 points so that a multi-megabase genome stays responsive.
GC skew asks which of G and C dominates a strand. For how much G and C there is in the first place, use the GC content calculator, and for length, composition and weight in one report, sequence statistics. TheCpG island finder answers the related question of where CG dinucleotides cluster.
The imbalance between guanine and cytosine on one strand of DNA, (G − C) ÷ (G + C), measured in a sliding window. A value of 0 means G and C are equally common, a positive value means the strand is G-rich, and a negative value means it is C-rich.
Replication is asymmetric. The leading strand is copied continuously and the lagging strand in short fragments, so the two strands spend different amounts of time single-stranded and suffer different mutation rates, above all the deamination of cytosine. In most bacteria this leaves the leading strand G-rich. On the sequence you hold, the skew is positive from the origin to the terminus and negative from the terminus back round to the origin, so the cumulative skew falls to its minimum at the origin and rises to its maximum at the terminus.
Large enough that noise averages out, small enough to place the switch. For a bacterial chromosome of a few megabases, windows of 10 to 50 kb with a step of a fifth of the window work well. Leave both fields empty and the tool scales them to the sequence, about 2 % of its length. The predicted origin and terminus come from the cumulative curve, which does not depend on the window at all.
On a set of contigs, no: the switch points can only be read from a complete, correctly oriented chromosome. Many plasmids and some bacteria, including cyanobacteria and several endosymbionts, show weak or irregular skew, and archaea often have several origins. Treat a prediction from anything under a megabase with care.
The running sum of the window skews along the sequence. Because the skew changes sign at the origin and the terminus, the cumulative curve has its minimum at the origin of replication and its maximum at the terminus. It is much smoother than the raw skew, which is why origin prediction reads it rather than the noisy window values. A bacterial chromosome of several megabases is analysed in well under a second.