Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly

  • S Kongsstovu Í, SO Mikalsen, EÍ Homrum, JA Jacobsen, P Flicek, HA Dahl. Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly. Sci Rep 2019;9(1):17716. doi:10.1038/s41598-019-54151-9
    [BibTeX] [Abstract]

    Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.

    @Article{31776409,
    author = {Í Kongsstovu S and Mikalsen SO and Homrum EÍ and Jacobsen JA and Flicek P and Dahl HA},
    title = {Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly},
    journal = {Sci Rep},
    volume = {9},
    number = {1},
    pages = {17716},
    year = {2019},
    doi = {10.1038/s41598-019-54151-9},
    abstract = {Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.},}

Description

Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.

Full details are provided in our  open access publication in Scientific Reports

Data access

The genome assemblies and sequencing reads were submitted to the European Nucleotide Archive and are available with the following accession numbers:

Assembly Accession number
A1 GCA_902175115.1
A2 GCA_902175115.2
A3 GCA_902175115.3
Sequencing reads Accession number
Paired-end library ERR2853087
Mate-pair 4500 library ERR2853088ERR2853089ERR2853090
Mate-pair 7000 library ERR2853091ERR2853092ERR2853093
Linked (10x Genomics) reads ERR2853094
MinION reads ERR2809166ERR2809167ERR2809168ERR2809169

Tables 2 and 6 show the QUAST results that were of most interest or relevance. The full QUAST results are available for the  reference  and  no reference  cases.

For connexin analysis the connexin sequences were aligned to the assemblies. Alignments are available for the  A1A2A3  and  draft  assemblies.

The features identified using FRC bam  are available here in gff format for the  A1A2  and  A3  assemblies. All reference data features are available  here .

Whole genome alignments were generated using the web tool D-Genies. Only the alignment between A3 and the draft assembly (Figure 2) are represented in the article. The files are available for alignments between  A1 and A2A2 and A3 , and  A3 and the draft assemblies .