- S Kongsstovu Í, SO Mikalsen, EÍ Homrum, JA Jacobsen, P Flicek, HA Dahl. Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly. Sci Rep 2019;9(1):17716. doi:10.1038/s41598-019-54151-9
[BibTeX] [Abstract]
Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.
@Article{31776409, author = {Í Kongsstovu S and Mikalsen SO and Homrum EÍ and Jacobsen JA and Flicek P and Dahl HA}, title = {Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly}, journal = {Sci Rep}, volume = {9}, number = {1}, pages = {17716}, year = {2019}, doi = {10.1038/s41598-019-54151-9}, abstract = {Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.},}
Description
Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.
Full details are provided in our open access publication in Scientific Reports
Data access
The genome assemblies and sequencing reads were submitted to the European Nucleotide Archive and are available with the following accession numbers:
| Assembly | Accession number |
|---|---|
| A1 | GCA_902175115.1 |
| A2 | GCA_902175115.2 |
| A3 | GCA_902175115.3 |
| Sequencing reads | Accession number |
|---|---|
| Paired-end library | ERR2853087 |
| Mate-pair 4500 library | ERR2853088 , ERR2853089 , ERR2853090 |
| Mate-pair 7000 library | ERR2853091 , ERR2853092 , ERR2853093 |
| Linked (10x Genomics) reads | ERR2853094 |
| MinION reads | ERR2809166 , ERR2809167 , ERR2809168 , ERR2809169 |
Tables 2 and 6 show the QUAST results that were of most interest or relevance. The full QUAST results are available for the reference and no reference cases.
For connexin analysis the connexin sequences were aligned to the assemblies. Alignments are available for the A1 , A2 , A3 and draft assemblies.
The features identified using FRC bam are available here in gff format for the A1 , A2 and A3 assemblies. All reference data features are available here .
Whole genome alignments were generated using the web tool D-Genies. Only the alignment between A3 and the draft assembly (Figure 2) are represented in the article. The files are available for alignments between A1 and A2 , A2 and A3 , and A3 and the draft assemblies .