1 / 316100%
Sorghum accessions and genome size determination
We included five Sorghum bicolor (B35, SC56, RTx430, Shanqui red, Tx7000)
accession, in addition to four S. propinquum accessions to facilitate interspecific
comparisons. S. propinquum accession PI653737 was obtained from the USDA
Agricultural Research Service Plant Genetic Resources Conservation Unit (Grifffin, GA)
(henceforth referred to as S. propinquum_USDA), and an unnamed S. propinquum
accession (henceforth referred to as S. propinquum_BR) was provided courtesy of Dr.
William Rooney, Texas A&M University, College Station, TX. Sequences for two
additional accessions, S. propinquum 369-1 and S. propinquum 369-2, were downloaded
from the short read archive (Mace et al. 2013). Plants were grown in the WVU
Department of Biology greenhouse under normal conditions. Leaves were flash frozen in
liquid nitrogen and stored at -80C. Nuclear DNA content was determined via flow
cytometry using chicken erythrocyte nuclei (CEN) as an internal standard with the
nuclear DNA content of 2.5 picograms per 2C, performed in triplicate, at the Flow
Cytometry Core Lab, Benaroya Research Institute at Virginia Mason (Seattle, WA)
(Table 1).
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
DNA extraction and Illumina sequencing
Frozen leaf tissue (~15 g) for all S. bicolor accessions, S. propinquum_USDA, and S.
propinquum_BR was ground to a fine powder using liquid nitrogen and suspended in
sucrose extraction buffer (SEB), in which 1/20th volume of 10% Triton X-100 solution
was added to lyse chloroplasts and mitochondria. The resulting homogenate was double
filtered to remove other cellular debris and nuclei were isolated using centrifugation.
DNA was extracted from isolated nuclei using the Promega Wizard Genomic DNA
Purification Kit (Madison, WI) following the manufacturer’s instructions. For all
accessions except S. propinquum_USDA, Illumina libraries were constructed and
sequenced at the Georgia Genomics Facility (University of Georgia, Athens, GA) using
the Illumina HiSeq 2000 (2 x 100 bp, ~500 bp insert size). S. propinquum_USDA was
sequenced using the Illumina MiSeq (2 × 150 bp, ~500 bp insert size) at the West Virginia
University Genomics Core Facility, Morgantown, WV. S. propinquum accessions 369-1
and 369-2 were sequenced by Mace et al. (2013) using the Illumina HiSeq 2000 (2 x 90
bp with 500 bp insert size).
Copy number estimates for newly sequenced Sorghum accessions
Short read sequence data for the five S. bicolor and four S. propinquum genomes were
trimmed for quality (-q 28) using sickle v1.33 (Joshi, Fass, 2011, available at
https://github.com/najoshi/sickle). Reads that contained any N’s and/or that were under
95 bp (85 bp for S. propinquum accessions 369-1 and 369-2) were removed using a
custom Perl script (available upon request). To estimate the copy number of
LTRretrotransposon families in our newly sequenced accessions, and in addition to the
two published S. propinquum accessions, we subsampled 7.5 million reads from each
filtered sequenced library. The reads were uniquely mapped to the 5’ LTRs mined from
the BTx623 reference genome using Bowtie2, allowing one mismatch within the entire
read
(Langmead and Salzberg, 2012). The total number of reads that strictly aligned to each 5’
LTR was extracted from the BAM output file and the copy number of each element was
estimated using the equation describe above. The copy number of each element was also
estimated using the de novo approach described above.
A two-sample t-test was performed to compare the estimates obtained from both
approaches for each retrotransposon family. This test was also performed for each
retrotransposon family between S. bicolor and S. propinquum to determine statistically
significant interspecific difference among TE families.
Results Identification and characterization of LTR-retrotransposons in the Sorghum
BTx623 reference genome
Using LTRharvest, 12,530 LTR-retrotransposons were mined from the Sorghum bicolor
reference genome (BTx623, v1.0 Paterson et al. 2009). These elements were filtered for
false positives using the PGSB Poaceae database as a reference, and 313 false positives
(mostly ribosomal repeats) were removed from the dataset. The remaining 12,217 intact
LTR-retrotransposons were then grouped into ~210 families based on 5’ LTR sequence
similarity (BLASTN, e-value 1e-10) and visualized as a network in Cytoscape v.3.0
(Figure 1). As observed in most grasses, Gypsy-like elements were the most abundant TE
sequences in the Sorghum genome (~9,000), followed by ~2,200 unclassified and ~1,100
Copia-like elements. The first and largest cluster in the LTRretrotransposon network
(Figure 1) consists of 7,801 5’ LTR sequences belonging to the Gypsy superfamily. These
sequences were further divided into various families such as Onap, Retrosor6, Leviathan,
Tekay-like elements, and RLX-CRM. The smallest clusters in the network contained only
two sequences, mostly of unknown classification. By using the 5’ LTR sequences as a
reference database, we used RepeatMasker to identify an additional 8,240 (presumably
solo) LTRs from the reference genome. The ratio of soloLTR to intact elements is
estimated at 0.67:1, similar to previously published findings (Baidouri and Panaud,
2013). Overall, a total of 32,674 LTRs (including both LTRs from the intact elements and
solo LTRs) were identified in the Sorghum reference genome assembly.
After grouping the elements into families, we estimated the insertional timing of each
element in the genome based on the sequence divergence of the 5' and 3' LTRs of each
individual full-length retrotransposon (Figure 2A). The estimated insertion times ranged
from 0 to 5.8 mya. The average insertion age is ~1.6 mya. As shown in Figure 2B, the
majority of LTRs in the “Retrosor-6” and “2-Unknown” clusters share a high degree of
sequence similarity (100 - 97.41 %) and were inserted into the genome within the last 1
mya. In addition, more than 80% of the LTRs in the “Retrosor-1”, “4-Gypsy” and
“5Unknown” clusters inserted less than 1 mya (Online Resource 3). In contrast, the
LTRs in the Onap cluster share less sequence similarity (92.2 to 85 %) and the estimated
insertion age of these elements falls within the range of 3 to 5.8 mya, indicating that these
insertions are older in origin, and predate the S. bicolor – S. propinquum divergence,
which occurred approximately 1-2 mya (Paterson 2008; Figure 2B & Online Resource
3).
In silico development of methods for copy number estimation
We performed an in silico analysis of the simulated short read dataset generated from the
reference genome to develop a framework for accurate copy number estimation. After
sequence simulation using DWGSIM, the reads were mapped back to the Sorghum
reference genome and the total number of reads that strictly aligned at each of the
previously identified LTRs was recorded. From the total read counts for each LTR, we
estimated the copy number via a probability statistic that takes into account the total
number of reads, number of reads that map to a target, length of the target sequence, and
genome size (Hawkins et al. 2006). We identified a total of 32,674 LTR sequences in the
Sorghum bicolor reference genome. The estimated copy number (LTRs= 32,912 ± 167)
via our statistic is strikingly similar to the actual annotated numbers, indicating that the
equation is remarkably accurate at estimating the copy numbers of repetitive sequences
from the short-read sequence data using this approach (Online Resource 1).
Further, we performed reference-based in silico copy number estimation for the largest
families found in the BTx623 reference genome. Among all families, Onap, a Gypsy
LTR-retrotransposon was most abundant (6,744 ± 64) followed by Retrosor-6 and
2_Unknown, estimated at 5,588 ± 48 and 4,551 ± 51, respectively. As with total LTR
number, these family-level estimates correlate with the annotated numbers identified via
LTR-harvest and RepeatMasker, demonstrating that the equation works well, even at
more refined levels. For example, we identified 5,852 full-length and 303 solo Onap
elements from LTR-harvest and RepeatMasker (total = 6,155; estimate = 6,744 ± 64).
Similarly, estimates correlate with annotated copy numbers for the 5,567 Retrosor-6
(estimate = 5,588 ± 48) and 4,543 2-Unknown (4,551 ± 51) elements. To evaluate
possible sampling effects, we repeated this analysis for 100 independently subsampled in
silico datasets. The average estimated copy number (LTRs = 33,834 ± 169) from the
resampling analysis is similar to the actual annotated number (32,674) as well as the
initial estimated copy number (32,912 ± 167) indicating the robustness of the equation in
capturing the TE landscape from short read datasets.
After verifying data repeatability using our statistic, we performed an additional
analysis to determine its usefulness when employing de novo methods for
characterization of the repetitive fraction of the genome. Copy numbers for each
LTRretrotransposon family were therefore estimated from the consensus sequences of the
largest RepeatExplorer clusters using the same statistical equation. Our de novo analysis
resulted in an estimated 69,558 ± 197 total LTRs in the reference genome, almost double
the total number of LTRs identified by LTRharvest and RepeatMasker (32,674). Among
all families, Onap, Retrosor-6 and 2-Unknown are estimated to be the most abundant
families, with Retrosor-6 estimated at 12,717 ± 69 followed by 2_Unknown (9,799 ± 68)
and Onap (8,636 ± 51). With the exception of Onap, the de novo estimates from the
simulated dataset are in agreement with most of the reference-based and all of the de
novo estimates for the real short-read sequence data (see below).
Reference-based copy number estimation in Sorghum
We estimated the copy numbers of various LTR-retrotransposon families from nine
Sorghum accessions (five S. bicolor and four S. propinquum) using the referencebased
approach (Online Resource 2). The estimated total copy number of LTRs is ~73,000 in S.
bicolor and ~50,000 in S. propinquum accessions (Table 2). The same ten families
contribute the highest number of copies to the total TE fraction in all of the nine Sorghum
accessions (Figure 3). Onap, the most abundant retrotransposon in all accessions, varies
significantly in copy number among species (Online resource 3). Specifically, Onap
copy number is similar among all accessions of S. bicolor (average 25,692 ± 100) and S.
propinquum_USDA (27,908 ± 108), but varies approximately three fold compared to S.
propinquum_BR (9,673 ± 61), S. propinquum 369-1 (9,055 ± 59) and S. propinquum 369-
2 (11,086 ± 65). Indeed, we found that the estimated copy numbers for many families in
S. propinquum_USDA are more similar to that of the S. bicolor accessions than the other
S. propinquum genomes (Figure 3). For example, 4-Gypsy in S. propinquum_BR (103 ±
8 copies), S. propinquum 369-1 (85 ± 7 copies) and S. propinquum 369-2 (100 ± 7 copies)
is composed of twelve-fold fewer copies compared to S. propinquum_USDA (~1,357 ±
27 copies), and approximately seven-fold fewer copies than the five S. bicolor genomes
(on average 775 ± 20 copies). To determine whether these observed differences are
statistically significant, we performed a two-sample t-test for each retrotransposon family
between S. bicolor and S. propinquum (Online resource 3). The copy numbers for five
families (Retrosor-6, RLX-CRM, Giepum,
11_Unclassified and 12_Unclassified) were significantly different (P < 0.05) between S.
bicolor and S. propinquum; however, when we included S. propinquum USDA with the S.
bicolor accessions, 14 families (Online resource 3) showed statistically significant
differences in copy number.
Variation in copy number using a de novo approach
To identify genome-specific repetitive sequences that were not identified by the
reference-based approach, we performed a de novo analysis using RepeatExplorer
(Novak et al. 2013). As this method is not restricted to LTR-retrotransposons, we also
estimated the proportions of a broader range of types of repetitive DNA. The largest
cluster in all accessions consisted of satellite repeats. The number of reads in the satellite
cluster was variable within and between genomes of S. bicolor and S. propinquum; B35,
RTx430, Shanqui red, S. propinquum_BR, S. propinquum 369-1 and S. propinquum 3692
contained ~178,000 satellite-associated reads of 137-274 bp that occupy ~22 Mb of the
genome, whereas SC56, Tx7000, and S. propinquum_USDA contained ~94,000
satelliteassociated reads of 21-68 bp which occupy ~2 Mb of the genome. Except for the
first one or two largest clusters (satellite and/or DNA transposons), all Sorghum genomes
contained a greater number of Gypsy-like LTR-retrotransposon clusters than Copia-like
elements (Online Resource 4).
The total LTR-retrotransposon copy number estimates for S. bicolor via de novo analysis
are similar to estimates from the reference-based method (Table 2). For S. propinquum,
de novo methods result in significantly higher copy number estimates (~ 75,900 ± 260),
with the exception of the estimate for S. propinquum USDA (68,949 ± 270).
Nevertheless, both methods indicate that Onap and Retrosor-6 are the two largest LTR-
retrotransposon families. For de novo estimates, Retrosor-6 is estimated at a higher copy
number than Onap in B35 (11,999 ± 83), SC56 (11,615 ± 83) and RTx430 (11,335 ± 86),
whereas Onap is estimated at higher copy number than Retrosor-6 in Tx7000
(14,910 ± 81), S. red (12,366 ± 68) and S. propinquum USDA (11,874 ± 65) (Figure 4).
Results for the other prevalent families, such as 2-unknown, Leviathan, and RLX-CRM,
while comparable to copy number estimates from the reference-based approach, are
significantly different in some cases (Figures 3 and 4). Interestingly, 2_Unknown is
estimated at higher copy number than either Onap or Retrosor-6 in S. propinquum_BR
(11,947 ± 80), S. propinquum 369-1 (12,240 ± 83) and S. propinquum 369-2 (13,869 ±
92), in concordance with the estimates via the reference-based method (Figures 3 and 4).
We performed a two-sample t-test for each retrotransposon family to determine
statistically significant differences in copy number between S. bicolor and S. propinquum.
There were no significant differences in copy number between S. bicolor and S.
propinquum; however, when we include S. propinquum USDA as one of the S. bicolor
accession, the copy numbers for seven families (2_unknown, Keama, RLX-CRM, Tekay,
5th_Unknown, Retrosor-1 and 12th_Unknown) were significantly different between the
two species (Online Resource 3). The results for five of these families (Keama, RLX-
CRM, 5th_Unknown, Retrosor-1 and 12th_Unknown) correlate with that from the
reference-based t-test comparisons.
Discussion The LTR-retrotransposon landscape in the S. bicolor BTx623 reference
genome
We identified 12,217 intact LTR-retrotransposons from the Sorghum bicolor reference
genome and classified these sequences into ~210 different families based on sequence
similarity of the 5’ LTRs (Figure 1). With the exception of the largest cluster in the
network, all other clusters were composed of sequences from distinct families. The
largest cluster contains 7,801 5’ LTR sequences that belong to several families including
Onap, Retrosor-6, Leviathan, Tekay, RLX-CRM, and Kaema. Sequence similarity among
these families indicates a deep evolutionary connection between the LTRretrotransposons
in the cluster, which can be explained by consecutive sequence evolution of TE families
(Khan 2006, Cordaux and Batzer 2010). For instance, Onap retrotransposons share less
sequence similarity between their 5’ and 3’ LTRs relative to the other families in the first
cluster, suggesting this family is the oldest (3 to 5.8 mya) and could therefore be the
progenitor of the younger related sequences, such as those belonging to Retrosor-6 and
Leviathan (Figure 2A and B).
Interestingly, the LTR-retrotransposon families of greatest abundance in S. bicolor differ
significantly from those in Zea mays. Huck, the most abundant family in maize and a few
other grasses, is found in very low copy number (2 copies) in S. bicolor (Peterson et al.
2002). Even at an e-value cutoff of 1e-05, we could not identify additional copies of this
element. The same is true for Ji and Opie, which are abundant in maize, but found at low
frequency (Ji - 62 copies, Opie – 0 copies) in Sorghum. In contrast, Onap, Retrosor-6, and
Leviathan are highly repetitive in Sorghum, but found in very low copy number in maize,
indicating activation of different TE families over very short evolutionary timescales
(Estep et al. 2013). In other words, although Sorghum and maize diverged only 12 mya
(Swigoňová et al. 2004), each species has undergone independent activation of lineage-
specific TEs.
Though several studies report recent LTR-retrotransposon bursts in related grasses (Piegu
et al. 2006, Bennetzen et al. 2012, Senerchia et al. 2013), we did not detect a large
amount of recent activity in the Sorghum reference genome. Indeed, very few intact LTR-
retrotransposons share 100% sequence similarity between 5’ to 3’ LTRs (Figure 2A).
Retrosor-1 is one of the few families that contained a higher proportion of sequences with
at least 99% similarity (70 out of 99) between 5' and 3' LTRs of the same element,
indicative of relatively recent amplification and insertion (< 1 mya). This family consists
of very few copies, having little effect on genome size. Nevertheless, most of these recent
insertions are located within 5-10 kb of protein coding genes (data not shown), and
therefore carry the potential to induce possible functional consequences on neighboring
genes (i.e, loss-of-function or altered gene expression), such as has been observed for tb1,
ZmCCT, and ZmRAP2.7 in maize (Salvi et al. 2007; Studer et al. 2011; Yang et al. 2013).
One possible explanation for low levels of detected recent transposon activity could be
artifactual in nature. The Sorghum reference genome assembly, consisting of ~730 Mb, is
~100 Mb less than the flow cytometry measurements for the S. bicolor genomes included
in this study (Table 2), suggesting a significant portion of the genome sequence is
missing. Since recently transposed elements share high sequence similarity, reads
belonging to these elements would collapse in the assembly and appear as a single or
small number of repeats, rather than a number of dispersed repeats. This would also
explain why the estimated copy numbers from the Illumina data are significantly higher
than that from the in silico analyses (see below).
Methods to estimate the repetitive fraction using short read sequences
We aimed to develop a method to accurately estimate repetitive sequence copy number
from short-read sequence data, and to use this method to detect inter- and intraspecific
variation in TE abundance for five accessions of S. bicolor and four accessions of S.
propinquum. To this end, we performed an in silico analysis to test the accuracy of our
statistical equation (Hawkins et al. 2006) using both reference-based and de novo
approaches. We annotated a total of 32,674 LTRs from the reference genome, and the
estimated copy number for both the total LTRs (32,912 ± 167) and for LTRs from
individual families (Onap = 6,744 ± 64; retrosor6 = 5,588 ± 48) from our referencebased
in silico analysis was strikingly similar to the annotated number.
The estimated total number of LTRs from the de novo in silico analysis, however, was
much higher. This discrepancy can easily be explained by the fact that, for the
referenced-based in silico analysis, we focused specifically on the number of reads that
mapped to precisely defined genomic locations, namely the 32,674 bioinformatically
identified LTRs. Reads that map to unidentified LTRs would be excluded from the 32,674
regions of interest, drastically reducing the mathematically estimated copy number. In
addition, our mathematical equation required the entire read to map within the first and
last nucleotides of an LTR, which would rarely occur by chance for shorter LTRs. These
same sequence reads would, however, be included in the in silico de novo assembly,
providing a clear explanation for the increased copy number via that method. Indeed, the
de novo in silico estimates correlate strongly with most of the reference-based and all of
the de novo estimates for the real short-read sequence data. Importantly, the referenced-
based in silico analysis was not designed to determine the actual number of LTRs in the
reference genome, but rather to verify the accuracy of our statistical equation.
Repeat diversity is most accurately characterized from short-read data via a
combination of reference-based and de novo approaches
Comparative characterization of TE diversity and abundance in the newly sequenced
accessions using both reference-based and de novo approaches suggests that the former is
primarily suitable when estimating within-species variation while the latter is more
reliable for more evolutionarily distant comparisons. We initially expected reference-
based mapping to efficiently describe TE content in S. propinquum, given that S. bicolor
and S. propinquum diverged as little as 1-2 million years ago (Paterson 2008), but this
was not the case. Although the reference-based and de novo estimates for the total
number of LTRs are in strong agreement for S. bicolor (and S. propinquum_USDA),
suggesting that either approach will accurately characterize within-species diversity, the
reference-based estimates for total LTRs in S. propinquum were considerably low (Table
2). At the individual family level, we note that the reference-based estimate for the most
ancient and abundant TE family (Onap- Figure 2B) in S. propinquum is unexpectedly
low, given its genome size and in comparison to the estimated copy numbers in the other
genomes (Figure 3). Onap, which has accumulated near gene-poor regions and is
composed of much older sequences based on molecular clock dating (Figure 2B), has
likely been retained due to limited selection pressure and recombination suppression in
this part of the genome. As most Onap insertions predate the S. bicolor - S. propinquum
divergence and have therefore accumulated a large number of lineage-specific mutations,
these sequences were more easily identified in S. bicolor accessions using referencebased
methods in comparison to that for S. propinquum. In addition, we suspect that de novo
assembly is ineffective at accurately estimating copy numbers from short read data for
older TE families due to difficulties with assembling reads that contain a larger number of
polymorphisms. We conclude that, although S. bicolor and S. propinquum diverged
recently, the substitution rates for LTR-retrotransposons in Sorghum are sufficiently high
to prevent interspecific comparisons of repetitive content via shared sequence similarity
alone (reference-based mapping), particularly for older sequences. Therefore, caution
should be used when performing interspecific reference-based TE annotation, even
among closely related species, as the results will likely underestimate the actual number
of copies in the genome, especially for sequences of more ancient origin, or those that are
undergoing accelerated rates of diversification.
Comparative analyses reveal detectable repeat variation over short evolutionary
timescales
Our comparative analyses using copy number estimates from the Illumina data reveal
small but detectable variation in LTR-retrotransposon content and copy number among
accessions of S. bicolor, and larger variation between S. propinquum_USDA and the
other three S. propinquum accessions. This result was expected, as the S. propinquum
accessions include the individuals with both the smallest (833 Mb) and largest (902 Mb)
genomes. We anticipated that this size disparity would be associated with recent
transpositional activity in S. propinquum_USDA; however, we could not detect
significantly elevated copy numbers for any specific element in the S.
propinquum_USDA genome. We also estimated the copy numbers for satellite repeats,
but again could not identify large differences that would explain the genome size
disparity. There are two possible explanations for this observation: 1) Genome size
variation among the S. propinquum accessions is due to the accumulation of a small
number of TE copies from a large number of families in S. propinquum_USDA, and
would therefore be undetectable in our analysis, or 2) the excess nuclear content in S.
propinquum_USDA is composed of older decaying TE sequences that can no longer be
identified at 80% sequence similarity, and/or that have been more effectively removed
from the other Sorghum genomes. From comparisons between our two approaches, we
suspect the most likely explanation is the latter. The LTR copy numbers for the largest
retrotransposon family, Onap, are much higher for S. propinquum_USDA than for other
S. propinquum genomes using the reference-based method (Figure 3), but drops
significantly using the de novo approach (Figure 4) suggesting that there are a greater
number of older sequences in the S. propinquum_USDA genome.
Alternatively, we observed that the repeat profile for S. propinquum_USDA is
surprisingly more similar to that of the S. bicolor genomes than S. propinquum, and
unlike the results for the S. propinquum accessions, reference-based methods of analysis
appear to effectively characterize the S. propinquum_USDA content. This could be
Students also viewed