See also Google Scholar and ORCID.
- D Thybert, M Roller, FCP Navarro, I Fiddes, I Streeter, C Feig, D Martin-Galvez, M Kolmogorov, V Janoušek, W Akanni, B Aken, S Aldridge, V Chakrapani, W Chow, L Clarke, C Cummins, A Doran, M Dunn, L Goodstadt, K Howe, M Howell, AA Josselin, RC Karn, CM Laukaitis, L Jingtao, F Martin, M Muffato, S Nachtweide, MA Quail, C Sisu, M Stanke, K Stefflova, C Van Oosterhout, F Veyrunes, B Ward, F Yang, G Yazdanifar, A Zadissa, DJ Adams, A Brazma, M Gerstein, B Paten, S Pham, TM Keane, DT Odom, P Flicek. Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes. Genome Res 2018;28(4):448–459. doi:10.1101/gr.234096.117
[BibTeX] [Abstract]
Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 and 6 million yr ago, but that are absent in the Hominidae. Hominidae show between four- and sevenfold lower rates of nucleotide change and feature turnover in both neutral and functional sequences, suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. Recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli, which resulted in thousands of novel, species-specific CTCF binding sites. Our results show that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.
@Article{29563166, author = {Thybert D and Roller M and Navarro FCP and Fiddes I and Streeter I and Feig C and Martin-Galvez D and Kolmogorov M and Janoušek V and Akanni W and Aken B and Aldridge S and Chakrapani V and Chow W and Clarke L and Cummins C and Doran A and Dunn M and Goodstadt L and Howe K and Howell M and Josselin AA and Karn RC and Laukaitis CM and Jingtao L and Martin F and Muffato M and Nachtweide S and Quail MA and Sisu C and Stanke M and Stefflova K and Van Oosterhout C and Veyrunes F and Ward B and Yang F and Yazdanifar G and Zadissa A and Adams DJ and Brazma A and Gerstein M and Paten B and Pham S and Keane TM and Odom DT and Flicek P}, title = {Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes}, journal = {Genome Res}, volume = {28}, number = {4}, pages = {448--459}, year = {2018}, doi = {10.1101/gr.234096.117}, howpublished = {Advanced online publication: 21 March 2018}, note = {First posted as a preprint: 2 July 2017}, abstract = {Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 and 6 million yr ago, but that are absent in the Hominidae. Hominidae show between four- and sevenfold lower rates of nucleotide change and feature turnover in both neutral and functional sequences, suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. Recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli, which resulted in thousands of novel, species-specific CTCF binding sites. Our results show that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.},} - R Beekman, V Chapaprieta, N Russiñol, R Vilarrasa-Blasi, N Verdaguer-Dot, JHA Martens, M Duran-Ferrer, M Kulis, F Serra, BM Javierre, SW Wingett, G Clot, AC Queirós, G Castellano, J Blanc, M Gut, A Merkel, S Heath, A Vlasova, S Ullrich, E Palumbo, A Enjuanes, D Martín-García, S Beà, M Pinyol, M Aymerich, R Royo, M Puiggros, D Torrents, A Datta, E Lowy, M Kostadima, M Roller, L Clarke, P Flicek, X Agirre, F Prosper, T Baumann, J Delgado, A López-Guillermo, P Fraser, ML Yaspo, R Guigó, R Siebert, MA Martí-Renom, XS Puente, C López-Otín, I Gut, HG Stunnenberg, E Campo, JI Martin-Subero. The reference epigenome and regulatory chromatin landscape of chronic lymphocytic leukemia. Nat Med 2018;24(6):868–880. doi:10.1038/s41591-018-0028-4
[BibTeX] [Abstract]
Chronic lymphocytic leukemia (CLL) is a frequent hematological neoplasm in which underlying epigenetic alterations are only partially understood. Here, we analyze the reference epigenome of seven primary CLLs and the regulatory chromatin landscape of 107 primary cases in the context of normal B cell differentiation. We identify that the CLL chromatin landscape is largely influenced by distinct dynamics during normal B cell maturation. Beyond this, we define extensive catalogues of regulatory elements de novo reprogrammed in CLL as a whole and in its major clinico-biological subtypes classified by IGHV somatic hypermutation levels. We uncover that IGHV-unmutated CLLs harbor more active and open chromatin than IGHV-mutated cases. Furthermore, we show that de novo active regions in CLL are enriched for NFAT, FOX and TCF/LEF transcription factor family binding sites. Although most genetic alterations are not associated with consistent epigenetic profiles, CLLs with MYD88 mutations and trisomy 12 show distinct chromatin configurations. Furthermore, we observe that non-coding mutations in IGHV-mutated CLLs are enriched in H3K27ac-associated regulatory elements outside accessible chromatin. Overall, this study provides an integrative portrait of the CLL epigenome, identifies extensive networks of altered regulatory elements and sheds light on the relationship between the genetic and epigenetic architecture of the disease.
@Article{29785028, author = {Beekman R and Chapaprieta V and Russiñol N and Vilarrasa-Blasi R and Verdaguer-Dot N and Martens JHA and Duran-Ferrer M and Kulis M and Serra F and Javierre BM and Wingett SW and Clot G and Queirós AC and Castellano G and Blanc J and Gut M and Merkel A and Heath S and Vlasova A and Ullrich S and Palumbo E and Enjuanes A and Martín-García D and Beà S and Pinyol M and Aymerich M and Royo R and Puiggros M and Torrents D and Datta A and Lowy E and Kostadima M and Roller M and Clarke L and Flicek P and Agirre X and Prosper F and Baumann T and Delgado J and López-Guillermo A and Fraser P and Yaspo ML and Guigó R and Siebert R and Martí-Renom MA and Puente XS and López-Otín C and Gut I and Stunnenberg HG and Campo E and Martin-Subero JI}, title = {The reference epigenome and regulatory chromatin landscape of chronic lymphocytic leukemia}, journal = {Nat Med}, volume = {24}, number = {6}, pages = {868--880}, year = {2018}, doi = {10.1038/s41591-018-0028-4}, howpublished = {Advanced online publication: 21 May 2018}, abstract = {Chronic lymphocytic leukemia (CLL) is a frequent hematological neoplasm in which underlying epigenetic alterations are only partially understood. Here, we analyze the reference epigenome of seven primary CLLs and the regulatory chromatin landscape of 107 primary cases in the context of normal B cell differentiation. We identify that the CLL chromatin landscape is largely influenced by distinct dynamics during normal B cell maturation. Beyond this, we define extensive catalogues of regulatory elements de novo reprogrammed in CLL as a whole and in its major clinico-biological subtypes classified by IGHV somatic hypermutation levels. We uncover that IGHV-unmutated CLLs harbor more active and open chromatin than IGHV-mutated cases. Furthermore, we show that de novo active regions in CLL are enriched for NFAT, FOX and TCF/LEF transcription factor family binding sites. Although most genetic alterations are not associated with consistent epigenetic profiles, CLLs with MYD88 mutations and trisomy 12 show distinct chromatin configurations. Furthermore, we observe that non-coding mutations in IGHV-mutated CLLs are enriched in H3K27ac-associated regulatory elements outside accessible chromatin. Overall, this study provides an integrative portrait of the CLL epigenome, identifies extensive networks of altered regulatory elements and sheds light on the relationship between the genetic and epigenetic architecture of the disease.},} - M Roller, E Stamper, D Villar, O Izuogu, F Martin, AM Redmond, R Ramachanderan, L Harewood, DT Odom, P Flicek. LINE retrotransposons characterize mammalian tissue-specific and evolutionarily dynamic regulatory regions. Genome Biol 2021;22(1):62. doi:10.1186/s13059-021-02260-y 10.1038/s41576-019-0173-8 10.1038/nature14217 10.1038/ng.3884 10.1038/s41588-019-0494-8 10.1038/nature12787 10.1371/journal.pbio.1000384 10.1038/nature09033 10.1016/j.molcel.2011.12.021 10.1038/s41467-018-06544-z 10.1038/nature10532 10.1038/s41586-019-1338-5 10.1126/science.1230612 10.1016/j.cell.2015.01.006 10.1093/gbe/evz134 10.1016/j.scr.2019.101456 10.1038/s41559-017-0447-5 10.1038/nature13985 10.1126/science.1246426 10.1016/j.cels.2018.01.002 10.1038/nature13182 10.1101/gr.190546.115 10.1186/s13059-018-1577-z 10.1371/journal.pgen.1003504 10.1101/gr.216150.116 10.1186/s12864-018-4850-3 10.1101/gr.218149.116 10.1101/gr.235747.118 10.1007/s10577-017-9570-z 10.1093/gbe/evv005 10.1126/science.aac7247 10.1093/nar/gkq132 10.1093/molbev/msy143 10.7554/eLife.13926 10.1093/nar/gkw1067 10.1074/jbc.273.2.891 10.1371/journal.pgen.1008036 10.1371/journal.pone.0027513 10.1093/nar/gkz1138 10.1038/nature11243 10.1016/j.cell.2005.01.001 10.1073/pnas.1016071107 10.1016/j.molcel.2013.01.038 10.1038/s41576-019-0128-0 10.1016/j.stem.2015.02.013 10.1073/pnas.1507125112 10.1038/nn.4229 10.1016/j.celrep.2013.05.031 10.1016/j.cell.2019.12.015 10.1038/nature07829 10.1038/ng.3286 10.1186/s13059-018-1432-2 10.1093/gbe/evx194 10.1101/gr.234096.117 10.1093/oxfordjournals.molbev.a003768 10.1038/nrg3802 10.1093/molbev/msx219 10.1016/j.ymeth.2009.03.001 10.1186/gb-2013-14-11-r124 10.1093/bioinformatics/btp324 10.1093/bioinformatics/btp352 10.1101/gr.136184.111 10.1186/gb-2008-9-9-r137 10.1038/ng1966 10.1038/nature09692 10.1038/nature01080 10.1038/ncb1076 10.1038/nature03877 10.1038/ng.154 10.1093/nar/gky955 10.1093/bioinformatics/bty191 10.1093/nar/gkg006 10.1093/nar/gkj112 10.1016/S0022-2836(05)80360-2 10.1186/1748-7188-6-26 10.1093/bioinformatics/btt509 10.1093/bioinformatics/btu170 10.1093/bioinformatics/bts635 10.1038/nbt.1621 10.1093/nar/gkv1189 10.1093/nar/gkw257 10.1093/bioinformatics/btt737 10.1093/bib/bbs017 10.1093/bioinformatics/btx364 10.1186/s13059-014-0550-8 10.1007/978-3-319-24277-4 10.1101/gr.09
[BibTeX] [Abstract]
BACKGROUND: To investigate the mechanisms driving regulatory evolution across tissues, we experimentally mapped promoters, enhancers, and gene expression in the liver, brain, muscle, and testis from ten diverse mammals. RESULTS: The regulatory landscape around genes included both tissue-shared and tissue-specific regulatory regions, where tissue-specific promoters and enhancers evolved most rapidly. Genomic regions switching between promoters and enhancers were more common across species, and less common across tissues within a single species. Long Interspersed Nuclear Elements (LINEs) played recurrent evolutionary roles: LINE L1s were associated with tissue-specific regulatory regions, whereas more ancient LINE L2s were associated with tissue-shared regulatory regions and with those switching between promoter and enhancer signatures across species. CONCLUSIONS: Our analyses of the tissue-specificity and evolutionary stability among promoters and enhancers reveal how specific LINE families have helped shape the dynamic mammalian regulome.
@Article{33602314, author = {Roller M and Stamper E and Villar D and Izuogu O and Martin F and Redmond AM and Ramachanderan R and Harewood L and Odom DT and Flicek P}, title = {LINE retrotransposons characterize mammalian tissue-specific and evolutionarily dynamic regulatory regions}, journal = {Genome Biol}, volume = {22}, number = {1}, pages = {62}, year = {2021}, doi = {10.1186/s13059-021-02260-y 10.1038/s41576-019-0173-8 10.1038/nature14217 10.1038/ng.3884 10.1038/s41588-019-0494-8 10.1038/nature12787 10.1371/journal.pbio.1000384 10.1038/nature09033 10.1016/j.molcel.2011.12.021 10.1038/s41467-018-06544-z 10.1038/nature10532 10.1038/s41586-019-1338-5 10.1126/science.1230612 10.1016/j.cell.2015.01.006 10.1093/gbe/evz134 10.1016/j.scr.2019.101456 10.1038/s41559-017-0447-5 10.1038/nature13985 10.1126/science.1246426 10.1016/j.cels.2018.01.002 10.1038/nature13182 10.1101/gr.190546.115 10.1186/s13059-018-1577-z 10.1371/journal.pgen.1003504 10.1101/gr.216150.116 10.1186/s12864-018-4850-3 10.1101/gr.218149.116 10.1101/gr.235747.118 10.1007/s10577-017-9570-z 10.1093/gbe/evv005 10.1126/science.aac7247 10.1093/nar/gkq132 10.1093/molbev/msy143 10.7554/eLife.13926 10.1093/nar/gkw1067 10.1074/jbc.273.2.891 10.1371/journal.pgen.1008036 10.1371/journal.pone.0027513 10.1093/nar/gkz1138 10.1038/nature11243 10.1016/j.cell.2005.01.001 10.1073/pnas.1016071107 10.1016/j.molcel.2013.01.038 10.1038/s41576-019-0128-0 10.1016/j.stem.2015.02.013 10.1073/pnas.1507125112 10.1038/nn.4229 10.1016/j.celrep.2013.05.031 10.1016/j.cell.2019.12.015 10.1038/nature07829 10.1038/ng.3286 10.1186/s13059-018-1432-2 10.1093/gbe/evx194 10.1101/gr.234096.117 10.1093/oxfordjournals.molbev.a003768 10.1038/nrg3802 10.1093/molbev/msx219 10.1016/j.ymeth.2009.03.001 10.1186/gb-2013-14-11-r124 10.1093/bioinformatics/btp324 10.1093/bioinformatics/btp352 10.1101/gr.136184.111 10.1186/gb-2008-9-9-r137 10.1038/ng1966 10.1038/nature09692 10.1038/nature01080 10.1038/ncb1076 10.1038/nature03877 10.1038/ng.154 10.1093/nar/gky955 10.1093/bioinformatics/bty191 10.1093/nar/gkg006 10.1093/nar/gkj112 10.1016/S0022-2836(05)80360-2 10.1186/1748-7188-6-26 10.1093/bioinformatics/btt509 10.1093/bioinformatics/btu170 10.1093/bioinformatics/bts635 10.1038/nbt.1621 10.1093/nar/gkv1189 10.1093/nar/gkw257 10.1093/bioinformatics/btt737 10.1093/bib/bbs017 10.1093/bioinformatics/btx364 10.1186/s13059-014-0550-8 10.1007/978-3-319-24277-4 10.1101/gr.09}, note = {First posted as a preprint: 31 May 2020}, abstract = {BACKGROUND: To investigate the mechanisms driving regulatory evolution across tissues, we experimentally mapped promoters, enhancers, and gene expression in the liver, brain, muscle, and testis from ten diverse mammals. RESULTS: The regulatory landscape around genes included both tissue-shared and tissue-specific regulatory regions, where tissue-specific promoters and enhancers evolved most rapidly. Genomic regions switching between promoters and enhancers were more common across species, and less common across tissues within a single species. Long Interspersed Nuclear Elements (LINEs) played recurrent evolutionary roles: LINE L1s were associated with tissue-specific regulatory regions, whereas more ancient LINE L2s were associated with tissue-shared regulatory regions and with those switching between promoter and enhancer signatures across species. CONCLUSIONS: Our analyses of the tissue-specificity and evolutionary stability among promoters and enhancers reveal how specific LINE families have helped shape the dynamic mammalian regulome.},} - M Rimoldi, N Wang, J Zhang, D Villar, DT Odom, J Taipale, P Flicek, M Roller. DNA methylation patterns of transcription factor binding regions characterize their functional and evolutionary contexts. Genome Biol 2024;25(1):146. doi:10.1186/s13059-024-03218-6
[BibTeX] [Abstract]
BACKGROUND: DNA methylation is an important epigenetic modification which has numerous roles in modulating genome function. Its levels are spatially correlated across the genome, typically high in repressed regions but low in transcription factor (TF) binding sites and active regulatory regions. However, the mechanisms establishing genome-wide and TF binding site methylation patterns are still unclear. RESULTS: Here we use a comparative approach to investigate the association of DNA methylation to TF binding evolution in mammals. Specifically, we experimentally profile DNA methylation and combine this with published occupancy profiles of five distinct TFs (CTCF, CEBPA, HNF4A, ONECUT1, FOXA1) in the liver of five mammalian species (human, macaque, mouse, rat, dog). TF binding sites are lowly methylated, but they often also have intermediate methylation levels. Furthermore, biding sites are influenced by the methylation status of CpGs in their wider binding regions even when CpGs are absent from the core binding motif. Employing a classification and clustering approach, we extract distinct and species-conserved patterns of DNA methylation levels at TF binding regions. CEBPA, HNF4A, ONECUT1, and FOXA1 share the same methylation patterns, while CTCF’s differ. These patterns characterize alternative functions and chromatin landscapes of TF-bound regions. Leveraging our phylogenetic framework, we find DNA methylation gain upon evolutionary loss of TF occupancy, indicating coordinated evolution. Furthermore, each methylation pattern has its own evolutionary trajectory reflecting its genomic contexts. CONCLUSIONS: Our epigenomic analyses indicate a role for DNA methylation in TF binding changes across species including that specific DNA methylation profiles characterize TF binding and are associated with their regulatory activity, chromatin contexts, and evolutionary trajectories.
@Article{38844976, author = {Rimoldi M and Wang N and Zhang J and Villar D and Odom DT and Taipale J and Flicek P and Roller M}, title = {DNA methylation patterns of transcription factor binding regions characterize their functional and evolutionary contexts}, journal = {Genome Biol}, volume = {25}, number = {1}, pages = {146}, year = {2024}, doi = {10.1186/s13059-024-03218-6}, note = {First posted as a preprint: 21 July 2022}, abstract = {BACKGROUND: DNA methylation is an important epigenetic modification which has numerous roles in modulating genome function. Its levels are spatially correlated across the genome, typically high in repressed regions but low in transcription factor (TF) binding sites and active regulatory regions. However, the mechanisms establishing genome-wide and TF binding site methylation patterns are still unclear. RESULTS: Here we use a comparative approach to investigate the association of DNA methylation to TF binding evolution in mammals. Specifically, we experimentally profile DNA methylation and combine this with published occupancy profiles of five distinct TFs (CTCF, CEBPA, HNF4A, ONECUT1, FOXA1) in the liver of five mammalian species (human, macaque, mouse, rat, dog). TF binding sites are lowly methylated, but they often also have intermediate methylation levels. Furthermore, biding sites are influenced by the methylation status of CpGs in their wider binding regions even when CpGs are absent from the core binding motif. Employing a classification and clustering approach, we extract distinct and species-conserved patterns of DNA methylation levels at TF binding regions. CEBPA, HNF4A, ONECUT1, and FOXA1 share the same methylation patterns, while CTCF's differ. These patterns characterize alternative functions and chromatin landscapes of TF-bound regions. Leveraging our phylogenetic framework, we find DNA methylation gain upon evolutionary loss of TF occupancy, indicating coordinated evolution. Furthermore, each methylation pattern has its own evolutionary trajectory reflecting its genomic contexts. CONCLUSIONS: Our epigenomic analyses indicate a role for DNA methylation in TF binding changes across species including that specific DNA methylation profiles characterize TF binding and are associated with their regulatory activity, chromatin contexts, and evolutionary trajectories.},}
- B Ballester, A Medina-Rivera, D Schmidt, M Gonzàlez-Porta, M Carlucci, X Chen, K Chessman, AJ Faure, AP Funnell, A Goncalves, C Kutter, M Lukk, S Menon, WM McLaren, K Stefflova, S Watt, MT Weirauch, M Crossley, JC Marioni, DT Odom, P Flicek, MD Wilson. Multi-species, multi-transcription factor binding highlights conserved control of tissue-specific biological pathways. Elife 2014;3:e02626. doi:10.7554/eLife.02626
[BibTeX] [Abstract]
As exome sequencing gives way to genome sequencing, the need to interpret the function of regulatory DNA becomes increasingly important. To test whether evolutionary conservation of cis-regulatory modules (CRMs) gives insight into human gene regulation, we determined transcription factor (TF) binding locations of four liver-essential TFs in liver tissue from human, macaque, mouse, rat, and dog. Approximately, two thirds of the TF-bound regions fell into CRMs. Less than half of the human CRMs were found as a CRM in the orthologous region of a second species. Shared CRMs were associated with liver pathways and disease loci identified by genome-wide association studies. Recurrent rare human disease causing mutations at the promoters of several blood coagulation and lipid metabolism genes were also identified within CRMs shared in multiple species. This suggests that multi-species analyses of experimentally determined combinatorial TF binding will help identify genomic regions critical for tissue-specific gene control
@Article{25279814, author = {Ballester B and Medina-Rivera A and Schmidt D and Gonzàlez-Porta M and Carlucci M and Chen X and Chessman K and Faure AJ and Funnell AP and Goncalves A and Kutter C and Lukk M and Menon S and McLaren WM and Stefflova K and Watt S and Weirauch MT and Crossley M and Marioni JC and Odom DT and Flicek P and Wilson MD}, title = {Multi-species, multi-transcription factor binding highlights conserved control of tissue-specific biological pathways}, journal = {Elife}, volume = {3}, pages = {e02626}, year = {2014}, doi = {10.7554/eLife.02626}, abstract = {As exome sequencing gives way to genome sequencing, the need to interpret the function of regulatory DNA becomes increasingly important. To test whether evolutionary conservation of cis-regulatory modules (CRMs) gives insight into human gene regulation, we determined transcription factor (TF) binding locations of four liver-essential TFs in liver tissue from human, macaque, mouse, rat, and dog. Approximately, two thirds of the TF-bound regions fell into CRMs. Less than half of the human CRMs were found as a CRM in the orthologous region of a second species. Shared CRMs were associated with liver pathways and disease loci identified by genome-wide association studies. Recurrent rare human disease causing mutations at the promoters of several blood coagulation and lipid metabolism genes were also identified within CRMs shared in multiple species. This suggests that multi-species analyses of experimentally determined combinatorial TF binding will help identify genomic regions critical for tissue-specific gene control},} - APW Funnell, MD Wilson, B Ballester, KS Mak, J Burdach, N Magan, RCM Pearson, FP Lemaigre, KM Stowell, DT Odom, P Flicek, M Crossley. A CpG Mutational Hotspot in a ONECUT Binding Site Accounts for the Prevalent Variant of Hemophilia B Leyden. Am J Hum Genet 2013;92(3):460–467. doi:10.1016/j.ajhg.2013.02.003
[BibTeX] [Abstract]
Hemophilia B, or the “royal disease,” arises from mutations in coagulation factor IX (F9). Mutations within the F9 promoter are associated with a remarkable hemophilia B subtype, termed hemophilia B Leyden, in which symptoms ameliorate after puberty. Mutations at the -5/-6 site (nucleotides -5 and -6 relative to the transcription start site, designated +1) account for the majority of Leyden cases and have been postulated to disrupt the binding of a transcriptional activator, the identity of which has remained elusive for more than 20 years. Here, we show that ONECUT transcription factors (ONECUT1 and ONECUT2) bind to the -5/-6 site. The various hemophilia B Leyden mutations that have been reported in this site inhibit ONECUT binding to varying degrees, which correlate well with their associated clinical severities. In addition, expression of F9 is crucially dependent on ONECUT factors in vivo, and as such, mice deficient in ONECUT1, ONECUT2, or both exhibit depleted levels of F9. Taken together, our findings establish ONECUT transcription factors as the missing hemophilia B Leyden regulators that operate through the -5/-6 site
@Article{23472758, author = {Funnell APW and Wilson MD and Ballester B and Mak KS and Burdach J and Magan N and Pearson RCM and Lemaigre FP and Stowell KM and Odom DT and Flicek P and Crossley M}, title = {A CpG Mutational Hotspot in a ONECUT Binding Site Accounts for the Prevalent Variant of Hemophilia B Leyden}, journal = {Am J Hum Genet}, volume = {92}, number = {3}, pages = {460--467}, year = {2013}, doi = {10.1016/j.ajhg.2013.02.003}, abstract = {Hemophilia B, or the ``royal disease,'' arises from mutations in coagulation factor IX (F9). Mutations within the F9 promoter are associated with a remarkable hemophilia B subtype, termed hemophilia B Leyden, in which symptoms ameliorate after puberty. Mutations at the -5/-6 site (nucleotides -5 and -6 relative to the transcription start site, designated +1) account for the majority of Leyden cases and have been postulated to disrupt the binding of a transcriptional activator, the identity of which has remained elusive for more than 20 years. Here, we show that ONECUT transcription factors (ONECUT1 and ONECUT2) bind to the -5/-6 site. The various hemophilia B Leyden mutations that have been reported in this site inhibit ONECUT binding to varying degrees, which correlate well with their associated clinical severities. In addition, expression of F9 is crucially dependent on ONECUT factors in vivo, and as such, mice deficient in ONECUT1, ONECUT2, or both exhibit depleted levels of F9. Taken together, our findings establish ONECUT transcription factors as the missing hemophilia B Leyden regulators that operate through the -5/-6 site},} - D Schmidt, PC Schwalie, MD Wilson, B Ballester, A Gonçalves, C Kutter, GD Brown, A Marshall, P Flicek, DT Odom. Waves of Retrotransposon Expansion Remodel Genome Organization and CTCF Binding in Multiple Mammalian Lineages. Cell 2012;148(1-2):335–348. doi:10.1016/j.cell.2011.11.058
[BibTeX] [Abstract]
CTCF-binding locations represent regulatory sequences that are highly constrained over the course of evolution. To gain insight into how these DNA elements are conserved and spread through the genome, we defined the full spectrum of CTCF-binding sites, including a 33/34-mer motif, and identified over five thousand highly conserved, robust, and tissue-independent CTCF-binding locations by comparing ChIP-seq data from six mammals. Our data indicate that activation of retroelements has produced species-specific expansions of CTCF binding in rodents, dogs, and opossum, which often functionally serve as chromatin and transcriptional insulators. We discovered fossilized repeat elements flanking deeply conserved CTCF-binding regions, indicating that similar retrotransposon expansions occurred hundreds of millions of years ago. Repeat-driven dispersal of CTCF binding is a fundamental, ancient, and still highly active mechanism of genome evolution in mammalian lineages. PAPERCLIP
@Article{22244452, author = {Schmidt D and Schwalie PC and Wilson MD and Ballester B and Gonçalves A and Kutter C and Brown GD and Marshall A and Flicek P and Odom DT}, title = {Waves of Retrotransposon Expansion Remodel Genome Organization and CTCF Binding in Multiple Mammalian Lineages}, journal = {Cell}, volume = {148}, number = {1-2}, pages = {335--348}, year = {2012}, doi = {10.1016/j.cell.2011.11.058}, abstract = {CTCF-binding locations represent regulatory sequences that are highly constrained over the course of evolution. To gain insight into how these DNA elements are conserved and spread through the genome, we defined the full spectrum of CTCF-binding sites, including a 33/34-mer motif, and identified over five thousand highly conserved, robust, and tissue-independent CTCF-binding locations by comparing ChIP-seq data from six mammals. Our data indicate that activation of retroelements has produced species-specific expansions of CTCF binding in rodents, dogs, and opossum, which often functionally serve as chromatin and transcriptional insulators. We discovered fossilized repeat elements flanking deeply conserved CTCF-binding regions, indicating that similar retrotransposon expansions occurred hundreds of millions of years ago. Repeat-driven dispersal of CTCF binding is a fundamental, ancient, and still highly active mechanism of genome evolution in mammalian lineages. PAPERCLIP},} - B Ballester, N Johnson, G Proctor, P Flicek. Consistent annotation of gene expression arrays. BMC Genomics 2010;11(1):294. doi:10.1186/1471-2164-11-294
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: Gene expression arrays are valuable and widely used tools for biomedical research. Today’s commercial arrays attempt to measure the expression level of all of the genes in the genome. Effectively translating the results from the microarray into a biological interpretation requires an accurate mapping between the probesets on the array and the genes that they are targeting. Although major array manufacturers provide annotations of their gene expression arrays, the methods used by various manufacturers are different and the annotations are difficult to keep up to date in the rapidly changing world of biological sequence databases. RESULTS: We have created a consistent microarray annotation protocol applicable to all of the major array manufacturers. We constantly keep our annotations updated with the latest Ensembl Gene predictions, and thus cross-referenced with a large number of external biomedical sequence database identifiers. We show that these annotations are accurate and address in detail reasons for the minority of probesets that cannot be annotated. Annotations are publicly accessible through the Ensembl Genome Browser and programmatically through the Ensembl Application Programming Interface. They are also seamlessly integrated into the BioMart data-mining tool and the biomaRt package of BioConductor. CONCLUSIONS: Consistent, accurate and updated gene expression array annotations remain critical for biological research. Our annotations facilitate accurate biological interpretation of gene expression profiles.
@Article{20459806, author = {Ballester B and Johnson N and Proctor G and Flicek P}, title = {Consistent annotation of gene expression arrays}, journal = {BMC Genomics}, volume = {11}, number = {1}, pages = {294}, year = {2010}, doi = {10.1186/1471-2164-11-294}, abstract = {ABSTRACT: BACKGROUND: Gene expression arrays are valuable and widely used tools for biomedical research. Today's commercial arrays attempt to measure the expression level of all of the genes in the genome. Effectively translating the results from the microarray into a biological interpretation requires an accurate mapping between the probesets on the array and the genes that they are targeting. Although major array manufacturers provide annotations of their gene expression arrays, the methods used by various manufacturers are different and the annotations are difficult to keep up to date in the rapidly changing world of biological sequence databases. RESULTS: We have created a consistent microarray annotation protocol applicable to all of the major array manufacturers. We constantly keep our annotations updated with the latest Ensembl Gene predictions, and thus cross-referenced with a large number of external biomedical sequence database identifiers. We show that these annotations are accurate and address in detail reasons for the minority of probesets that cannot be annotated. Annotations are publicly accessible through the Ensembl Genome Browser and programmatically through the Ensembl Application Programming Interface. They are also seamlessly integrated into the BioMart data-mining tool and the biomaRt package of BioConductor. CONCLUSIONS: Consistent, accurate and updated gene expression array annotations remain critical for biological research. Our annotations facilitate accurate biological interpretation of gene expression profiles.},} - D Schmidt, MD Wilson, B Ballester, PC Schwalie, GD Brown, A Marshall, C Kutter, S Watt, CP Martinez-Jimenez, S Mackay, I Talianidis, P Flicek, DT Odom. Five-Vertebrate ChIP-seq Reveals the Evolutionary Dynamics of Transcription Factor Binding. Science 2010;328(5981):1036–1040. doi:10.1126/science.1186176
[BibTeX] [Abstract]
Transcription factors (TFs) direct gene expression by binding to DNA regulatory regions. To explore the evolution of gene regulation, we experimentally determined the genome-wide occupancy of two TFs, CEBPA and HNF4A, in livers of five vertebrates. Although each TF displays highly conserved DNA binding preferences, most binding is species-specific, and aligned binding events present in all five species are rare. Regions near genes with expression levels dependent on a TF are often bound by the TF in multiple species, yet show no enhanced DNA sequence constraint. Binding divergence between species can be largely explained by sequence changes to the bound motifs. Among the binding events lost in one lineage, only half are recovered by another binding event within 10 kilobases. Our results reveal large interspecies differences in transcriptional regulation and provide insight into their evolution.
@Article{20378774, author = {Schmidt D and Wilson MD and Ballester B and Schwalie PC and Brown GD and Marshall A and Kutter C and Watt S and Martinez-Jimenez CP and Mackay S and Talianidis I and Flicek P and Odom DT}, title = {Five-Vertebrate ChIP-seq Reveals the Evolutionary Dynamics of Transcription Factor Binding}, journal = {Science}, volume = {328}, number = {5981}, pages = {1036--1040}, year = {2010}, doi = {10.1126/science.1186176}, abstract = {Transcription factors (TFs) direct gene expression by binding to DNA regulatory regions. To explore the evolution of gene regulation, we experimentally determined the genome-wide occupancy of two TFs, CEBPA and HNF4A, in livers of five vertebrates. Although each TF displays highly conserved DNA binding preferences, most binding is species-specific, and aligned binding events present in all five species are rare. Regions near genes with expression levels dependent on a TF are often bound by the TF in multiple species, yet show no enhanced DNA sequence constraint. Binding divergence between species can be largely explained by sequence changes to the bound motifs. Among the binding events lost in one lineage, only half are recovered by another binding event within 10 kilobases. Our results reveal large interspecies differences in transcriptional regulation and provide insight into their evolution.},} - M Carlile, D Swan, K Jackson, K Preston-Fayers, B Ballester, P Flicek, A Werner. Strand selective generation of endo-siRNAs from the Na/phosphate transporter gene Slc34a1 in murine tissues. Nucleic Acids Res 2009;37(7):2274–2282. doi:10.1093/nar/gkp088
[BibTeX] [Abstract]
Natural antisense transcripts (NATs) are important regulators of gene expression. Recently, a link between antisense transcription and the formation of endo-siRNAs has emerged. We investigated the bi-directionally transcribed Na/phosphate cotransporter gene (Slc34a1) under the aspect of endo-siRNA processing. Mouse Slc34a1 produces an antisense transcript that represents an alternative splice product of the Pfn3 gene located downstream of Slc34a1. The antisense transcript is prominently found in testis and in kidney. Co-expression of in vitro synthesized sense/antisense transcripts in Xenopus oocytes indicated processing of the overlapping transcripts into endo-siRNAs in the nucleus. Truncation experiments revealed that an overlap of at least 29 base-pairs is required to induce processing. We detected endo-siRNAs in mouse tissues that co express Slc34a1 sense/antisense transcripts by northern blotting. The orientation of endo-siRNAs was tissue specific in mouse kidney and testis. In kidney where the Na/phosphate cotransporter fulfils its physiological function endo-siRNAs complementary to the NAT were detected, in testis both orientations were found. Considering the wide spread expression of NATs and the gene silencing potential of endo-siRNAs we hypothesized a genome-wide link between antisense transcription and monoallelic expression. Significant correlation between random imprinting and antisense transcription could indeed be established. Our findings suggest a novel, more general role for NATs in gene regulation.
@Article{19237395, author = {Carlile M and Swan D and Jackson K and Preston-Fayers K and Ballester B and Flicek P and Werner A}, title = {Strand selective generation of endo-siRNAs from the Na/phosphate transporter gene Slc34a1 in murine tissues}, journal = {Nucleic Acids Res}, volume = {37}, number = {7}, pages = {2274--2282}, year = {2009}, doi = {10.1093/nar/gkp088}, howpublished = {Advanced online publication: 23 February 2009}, abstract = {Natural antisense transcripts (NATs) are important regulators of gene expression. Recently, a link between antisense transcription and the formation of endo-siRNAs has emerged. We investigated the bi-directionally transcribed Na/phosphate cotransporter gene (Slc34a1) under the aspect of endo-siRNA processing. Mouse Slc34a1 produces an antisense transcript that represents an alternative splice product of the Pfn3 gene located downstream of Slc34a1. The antisense transcript is prominently found in testis and in kidney. Co-expression of in vitro synthesized sense/antisense transcripts in Xenopus oocytes indicated processing of the overlapping transcripts into endo-siRNAs in the nucleus. Truncation experiments revealed that an overlap of at least 29 base-pairs is required to induce processing. We detected endo-siRNAs in mouse tissues that co express Slc34a1 sense/antisense transcripts by northern blotting. The orientation of endo-siRNAs was tissue specific in mouse kidney and testis. In kidney where the Na/phosphate cotransporter fulfils its physiological function endo-siRNAs complementary to the NAT were detected, in testis both orientations were found. Considering the wide spread expression of NATs and the gene silencing potential of endo-siRNAs we hypothesized a genome-wide link between antisense transcription and monoallelic expression. Significant correlation between random imprinting and antisense transcription could indeed be established. Our findings suggest a novel, more general role for NATs in gene regulation.},}
- SJ Aitken, X Ibarra-Soria, E Kentepozidou, P Flicek, C Feig, JC Marioni, DT Odom. CTCF maintains regulatory homeostasis of cancer pathways. Genome Biol 2018;19(1):106. doi:10.1186/s13059-018-1484-3
[BibTeX] [Abstract]
BACKGROUND: CTCF binding to DNA helps partition the mammalian genome into discrete structural and regulatory domains. Complete removal of CTCF from mammalian cells causes catastrophic genome dysregulation, likely due to widespread collapse of 3D chromatin looping and alterations to inter- and intra-TAD interactions within the nucleus. In contrast, Ctcf hemizygous mice with lifelong reduction of CTCF expression are viable, albeit with increased cancer incidence. Here, we exploit chronic Ctcf hemizygosity to reveal its homeostatic roles in maintaining genome function and integrity. RESULTS: We find that Ctcf hemizygous cells show modest but robust changes in almost a thousand sites of genomic CTCF occupancy; these are enriched for lower affinity binding events with weaker evolutionary conservation across the mouse lineage. Furthermore, we observe dysregulation of the expression of several hundred genes, which are concentrated in cancer-related pathways, and are caused by changes in transcriptional regulation. Chromatin structure is preserved but some loop interactions are destabilized; these are often found around differentially expressed genes and their enhancers. Importantly, the transcriptional alterations identified in vitro are recapitulated in mouse tumors and also in human cancers. CONCLUSIONS: This multi-dimensional genomic and epigenomic profiling of a Ctcf hemizygous mouse model system shows that chronic depletion of CTCF dysregulates steady-state gene expression by subtly altering transcriptional regulation, changes which can also be observed in primary tumors.
@Article{30086769, author = {Aitken SJ and Ibarra-Soria X and Kentepozidou E and Flicek P and Feig C and Marioni JC and Odom DT}, title = {CTCF maintains regulatory homeostasis of cancer pathways}, journal = {Genome Biol}, volume = {19}, number = {1}, pages = {106}, year = {2018}, doi = {10.1186/s13059-018-1484-3}, abstract = {BACKGROUND: CTCF binding to DNA helps partition the mammalian genome into discrete structural and regulatory domains. Complete removal of CTCF from mammalian cells causes catastrophic genome dysregulation, likely due to widespread collapse of 3D chromatin looping and alterations to inter- and intra-TAD interactions within the nucleus. In contrast, Ctcf hemizygous mice with lifelong reduction of CTCF expression are viable, albeit with increased cancer incidence. Here, we exploit chronic Ctcf hemizygosity to reveal its homeostatic roles in maintaining genome function and integrity. RESULTS: We find that Ctcf hemizygous cells show modest but robust changes in almost a thousand sites of genomic CTCF occupancy; these are enriched for lower affinity binding events with weaker evolutionary conservation across the mouse lineage. Furthermore, we observe dysregulation of the expression of several hundred genes, which are concentrated in cancer-related pathways, and are caused by changes in transcriptional regulation. Chromatin structure is preserved but some loop interactions are destabilized; these are often found around differentially expressed genes and their enhancers. Importantly, the transcriptional alterations identified in vitro are recapitulated in mouse tumors and also in human cancers. CONCLUSIONS: This multi-dimensional genomic and epigenomic profiling of a Ctcf hemizygous mouse model system shows that chronic depletion of CTCF dysregulates steady-state gene expression by subtly altering transcriptional regulation, changes which can also be observed in primary tumors.},} - E Kentepozidou, SJ Aitken, C Feig, K Stefflova, X Ibarra-Soria, DT Odom, M Roller, P Flicek. Clustered CTCF binding is an evolutionary mechanism to maintain topologically associating domains. Genome Biol 2020;21(1):5. doi:10.1186/s13059-019-1894-x
[BibTeX] [Abstract]
BACKGROUND: CTCF binding contributes to the establishment of a higher-order genome structure by demarcating the boundaries of large-scale topologically associating domains (TADs). However, despite the importance and conservation of TADs, the role of CTCF binding in their evolution and stability remains elusive. RESULTS: We carry out an experimental and computational study that exploits the natural genetic variation across five closely related species to assess how CTCF binding patterns stably fixed by evolution in each species contribute to the establishment and evolutionary dynamics of TAD boundaries. We perform CTCF ChIP-seq in multiple mouse species to create genome-wide binding profiles and associate them with TAD boundaries. Our analyses reveal that CTCF binding is maintained at TAD boundaries by a balance of selective constraints and dynamic evolutionary processes. Regardless of their conservation across species, CTCF binding sites at TAD boundaries are subject to stronger sequence and functional constraints compared to other CTCF sites. TAD boundaries frequently harbor dynamically evolving clusters containing both evolutionarily old and young CTCF sites as a result of the repeated acquisition of new species-specific sites close to conserved ones. The overwhelming majority of clustered CTCF sites colocalize with cohesin and are significantly closer to gene transcription start sites than nonclustered CTCF sites, suggesting that CTCF clusters particularly contribute to cohesin stabilization and transcriptional regulation. CONCLUSIONS: Dynamic conservation of CTCF site clusters is an apparently important feature of CTCF binding evolution that is critical to the functional stability of a higher-order chromatin structure.
@Article{31910870, author = {Kentepozidou E and Aitken SJ and Feig C and Stefflova K and Ibarra-Soria X and Odom DT and Roller M and Flicek P}, title = {Clustered CTCF binding is an evolutionary mechanism to maintain topologically associating domains}, journal = {Genome Biol}, volume = {21}, number = {1}, pages = {5}, year = {2020}, doi = {10.1186/s13059-019-1894-x}, note = {First posted as a preprint: 12 June 2019}, abstract = {BACKGROUND: CTCF binding contributes to the establishment of a higher-order genome structure by demarcating the boundaries of large-scale topologically associating domains (TADs). However, despite the importance and conservation of TADs, the role of CTCF binding in their evolution and stability remains elusive. RESULTS: We carry out an experimental and computational study that exploits the natural genetic variation across five closely related species to assess how CTCF binding patterns stably fixed by evolution in each species contribute to the establishment and evolutionary dynamics of TAD boundaries. We perform CTCF ChIP-seq in multiple mouse species to create genome-wide binding profiles and associate them with TAD boundaries. Our analyses reveal that CTCF binding is maintained at TAD boundaries by a balance of selective constraints and dynamic evolutionary processes. Regardless of their conservation across species, CTCF binding sites at TAD boundaries are subject to stronger sequence and functional constraints compared to other CTCF sites. TAD boundaries frequently harbor dynamically evolving clusters containing both evolutionarily old and young CTCF sites as a result of the repeated acquisition of new species-specific sites close to conserved ones. The overwhelming majority of clustered CTCF sites colocalize with cohesin and are significantly closer to gene transcription start sites than nonclustered CTCF sites, suggesting that CTCF clusters particularly contribute to cohesin stabilization and transcriptional regulation. CONCLUSIONS: Dynamic conservation of CTCF site clusters is an apparently important feature of CTCF binding evolution that is critical to the functional stability of a higher-order chromatin structure.},} - SJ Aitken, CJ Anderson, F Connor, O Pich, V Sundaram, C Feig, TF Rayner, M Lukk, S Aitken, J Luft, E Kentepozidou, C Arnedo-Pac, SV Beentjes, SE Davies, RM Drews, A Ewing, VB Kaiser, A Khamseh, E López-Arribillaga, AM Redmond, J Santoyo-Lopez, I Sentís, L Talmane, AD Yates, CEC Liver, CA Semple, N López-Bigas, P Flicek, DT Odom, MS Taylor. Pervasive lesion segregation shapes cancer genome evolution. Nature 2020;583(7815):265–270. doi:10.1038/s41586-020-2435-1
[BibTeX] [Abstract]
Cancers arise through the acquisition of oncogenic mutations and grow by clonal expansion\textsuperscript{1,2}. Here we reveal that most mutagenic DNA lesions are not resolved into a mutated DNA base pair within a single cell cycle. Instead, DNA lesions segregate, unrepaired, into daughter cells for multiple cell generations, resulting in the chromosome-scale phasing of subsequent mutations. We characterize this process in mutagen-induced mouse liver tumours and show that DNA replication across persisting lesions can produce multiple alternative alleles in successive cell divisions, thereby generating both multiallelic and combinatorial genetic diversity. The phasing of lesions enables accurate measurement of strand-biased repair processes, quantification of oncogenic selection and fine mapping of sister-chromatid-exchange events. Finally, we demonstrate that lesion segregation is a unifying property of exogenous mutagens, including UV light and chemotherapy agents in human cells and tumours, which has profound implications for the evolution and adaptation of cancer genomes.
@Article{32581361, author = {Aitken SJ and Anderson CJ and Connor F and Pich O and Sundaram V and Feig C and Rayner TF and Lukk M and Aitken S and Luft J and Kentepozidou E and Arnedo-Pac C and Beentjes SV and Davies SE and Drews RM and Ewing A and Kaiser VB and Khamseh A and López-Arribillaga E and Redmond AM and Santoyo-Lopez J and Sentís I and Talmane L and Yates AD and Liver CEC and Semple CA and López-Bigas N and Flicek P and Odom DT and Taylor MS}, title = {Pervasive lesion segregation shapes cancer genome evolution}, journal = {Nature}, volume = {583}, number = {7815}, pages = {265--270}, year = {2020}, doi = {10.1038/s41586-020-2435-1}, howpublished = {Advanced online publication: 24 June 2020}, note = {First posted as a preprint: 8 December 2019}, abstract = {Cancers arise through the acquisition of oncogenic mutations and grow by clonal expansion\textsuperscript{1,2}. Here we reveal that most mutagenic DNA lesions are not resolved into a mutated DNA base pair within a single cell cycle. Instead, DNA lesions segregate, unrepaired, into daughter cells for multiple cell generations, resulting in the chromosome-scale phasing of subsequent mutations. We characterize this process in mutagen-induced mouse liver tumours and show that DNA replication across persisting lesions can produce multiple alternative alleles in successive cell divisions, thereby generating both multiallelic and combinatorial genetic diversity. The phasing of lesions enables accurate measurement of strand-biased repair processes, quantification of oncogenic selection and fine mapping of sister-chromatid-exchange events. Finally, we demonstrate that lesion segregation is a unifying property of exogenous mutagens, including UV light and chemotherapy agents in human cells and tumours, which has profound implications for the evolution and adaptation of cancer genomes.},} - SJ Aitken, F Connor, C Feig, TF Rayner, M Lukk, J Luft, S Aitken, C Arnedo-Pac, JF Hayes, MD Nicholson, A Ewing, V Sundaram, JC Verburg, J Connelly, CJ Anderson, M Behm, S Campbell, M Daunesse, VB Kaiser, E Kentepozidou, O Pich, AM Redmond, J Santoyo-Lopez, I Sentís, L Talmane, Liver Cancer Evolution Consortium, P Flicek, N López-Bigas, CA Semple, MS Taylor, DT Odom. Genetic background sets the trajectory of experimental cancer evolution. Nature 2026. doi:10.1038/s41586-026-10821-z
[BibTeX] [Abstract]
Human cancers are heterogeneous\textsuperscript{1}. Dissecting how germline genetic variation and environmental factors shape tumour evolution using human datasets is limited by inherent diversity in genetic backgrounds\textsuperscript{2} and environmental exposures\textsuperscript{3-5}. Here, to overcome these limitations, we re-ran early tumour evolution hundreds of times in diverged inbred mouse strains, generating matched histology and whole-genome and transcriptome sequences. The sex, environment and carcinogenic exposures were all controlled, and the study design allowed us to capture genetic variation comparable with that observed across human populations while exploiting the nested hierarchical structure of strain-litter-animal-tumour relationships. Our analyses reveal that epistatic interactions between genetic background and acquired somatic mutations result in population-specific disease progression, including choice of driver mutations, occurrence of whole-genome duplication and subclonal selection dynamics that mirror both cancer susceptibility and tumour growth rate. Even modest genetic divergence, comparable with that found across human ancestry groups, can strikingly alter selection pressures during cancer development to shape both cancer risk and the trajectory of tumour evolution.
@Article{42486977, author = {Aitken SJ and Connor F and Feig C and Rayner TF and Lukk M and Luft J and Aitken S and Arnedo-Pac C and Hayes JF and Nicholson MD and Ewing A and Sundaram V and Verburg JC and Connelly J and Anderson CJ and Behm M and Campbell S and Daunesse M and Kaiser VB and Kentepozidou E and Pich O and Redmond AM and Santoyo-Lopez J and Sentís I and Talmane L and {Liver Cancer Evolution Consortium} and Flicek P and López-Bigas N and Semple CA and Taylor MS and Odom DT}, title = {Genetic background sets the trajectory of experimental cancer evolution}, journal = {Nature}, year = {2026}, doi = {10.1038/s41586-026-10821-z}, howpublished = {Advanced online publication: 22 July 2026}, note = {First posted as a preprint: 15 January 2015}, abstract = {Human cancers are heterogeneous\textsuperscript{1}. Dissecting how germline genetic variation and environmental factors shape tumour evolution using human datasets is limited by inherent diversity in genetic backgrounds\textsuperscript{2} and environmental exposures\textsuperscript{3-5}. Here, to overcome these limitations, we re-ran early tumour evolution hundreds of times in diverged inbred mouse strains, generating matched histology and whole-genome and transcriptome sequences. The sex, environment and carcinogenic exposures were all controlled, and the study design allowed us to capture genetic variation comparable with that observed across human populations while exploiting the nested hierarchical structure of strain-litter-animal-tumour relationships. Our analyses reveal that epistatic interactions between genetic background and acquired somatic mutations result in population-specific disease progression, including choice of driver mutations, occurrence of whole-genome duplication and subclonal selection dynamics that mirror both cancer susceptibility and tumour growth rate. Even modest genetic divergence, comparable with that found across human ancestry groups, can strikingly alter selection pressures during cancer development to shape both cancer risk and the trajectory of tumour evolution.},}
- M Roller, E Stamper, D Villar, O Izuogu, F Martin, AM Redmond, R Ramachanderan, L Harewood, DT Odom, P Flicek. LINE retrotransposons characterize mammalian tissue-specific and evolutionarily dynamic regulatory regions. Genome Biol 2021;22(1):62. doi:10.1186/s13059-021-02260-y 10.1038/s41576-019-0173-8 10.1038/nature14217 10.1038/ng.3884 10.1038/s41588-019-0494-8 10.1038/nature12787 10.1371/journal.pbio.1000384 10.1038/nature09033 10.1016/j.molcel.2011.12.021 10.1038/s41467-018-06544-z 10.1038/nature10532 10.1038/s41586-019-1338-5 10.1126/science.1230612 10.1016/j.cell.2015.01.006 10.1093/gbe/evz134 10.1016/j.scr.2019.101456 10.1038/s41559-017-0447-5 10.1038/nature13985 10.1126/science.1246426 10.1016/j.cels.2018.01.002 10.1038/nature13182 10.1101/gr.190546.115 10.1186/s13059-018-1577-z 10.1371/journal.pgen.1003504 10.1101/gr.216150.116 10.1186/s12864-018-4850-3 10.1101/gr.218149.116 10.1101/gr.235747.118 10.1007/s10577-017-9570-z 10.1093/gbe/evv005 10.1126/science.aac7247 10.1093/nar/gkq132 10.1093/molbev/msy143 10.7554/eLife.13926 10.1093/nar/gkw1067 10.1074/jbc.273.2.891 10.1371/journal.pgen.1008036 10.1371/journal.pone.0027513 10.1093/nar/gkz1138 10.1038/nature11243 10.1016/j.cell.2005.01.001 10.1073/pnas.1016071107 10.1016/j.molcel.2013.01.038 10.1038/s41576-019-0128-0 10.1016/j.stem.2015.02.013 10.1073/pnas.1507125112 10.1038/nn.4229 10.1016/j.celrep.2013.05.031 10.1016/j.cell.2019.12.015 10.1038/nature07829 10.1038/ng.3286 10.1186/s13059-018-1432-2 10.1093/gbe/evx194 10.1101/gr.234096.117 10.1093/oxfordjournals.molbev.a003768 10.1038/nrg3802 10.1093/molbev/msx219 10.1016/j.ymeth.2009.03.001 10.1186/gb-2013-14-11-r124 10.1093/bioinformatics/btp324 10.1093/bioinformatics/btp352 10.1101/gr.136184.111 10.1186/gb-2008-9-9-r137 10.1038/ng1966 10.1038/nature09692 10.1038/nature01080 10.1038/ncb1076 10.1038/nature03877 10.1038/ng.154 10.1093/nar/gky955 10.1093/bioinformatics/bty191 10.1093/nar/gkg006 10.1093/nar/gkj112 10.1016/S0022-2836(05)80360-2 10.1186/1748-7188-6-26 10.1093/bioinformatics/btt509 10.1093/bioinformatics/btu170 10.1093/bioinformatics/bts635 10.1038/nbt.1621 10.1093/nar/gkv1189 10.1093/nar/gkw257 10.1093/bioinformatics/btt737 10.1093/bib/bbs017 10.1093/bioinformatics/btx364 10.1186/s13059-014-0550-8 10.1007/978-3-319-24277-4 10.1101/gr.09
[BibTeX] [Abstract]
BACKGROUND: To investigate the mechanisms driving regulatory evolution across tissues, we experimentally mapped promoters, enhancers, and gene expression in the liver, brain, muscle, and testis from ten diverse mammals. RESULTS: The regulatory landscape around genes included both tissue-shared and tissue-specific regulatory regions, where tissue-specific promoters and enhancers evolved most rapidly. Genomic regions switching between promoters and enhancers were more common across species, and less common across tissues within a single species. Long Interspersed Nuclear Elements (LINEs) played recurrent evolutionary roles: LINE L1s were associated with tissue-specific regulatory regions, whereas more ancient LINE L2s were associated with tissue-shared regulatory regions and with those switching between promoter and enhancer signatures across species. CONCLUSIONS: Our analyses of the tissue-specificity and evolutionary stability among promoters and enhancers reveal how specific LINE families have helped shape the dynamic mammalian regulome.
@Article{33602314, author = {Roller M and Stamper E and Villar D and Izuogu O and Martin F and Redmond AM and Ramachanderan R and Harewood L and Odom DT and Flicek P}, title = {LINE retrotransposons characterize mammalian tissue-specific and evolutionarily dynamic regulatory regions}, journal = {Genome Biol}, volume = {22}, number = {1}, pages = {62}, year = {2021}, doi = {10.1186/s13059-021-02260-y 10.1038/s41576-019-0173-8 10.1038/nature14217 10.1038/ng.3884 10.1038/s41588-019-0494-8 10.1038/nature12787 10.1371/journal.pbio.1000384 10.1038/nature09033 10.1016/j.molcel.2011.12.021 10.1038/s41467-018-06544-z 10.1038/nature10532 10.1038/s41586-019-1338-5 10.1126/science.1230612 10.1016/j.cell.2015.01.006 10.1093/gbe/evz134 10.1016/j.scr.2019.101456 10.1038/s41559-017-0447-5 10.1038/nature13985 10.1126/science.1246426 10.1016/j.cels.2018.01.002 10.1038/nature13182 10.1101/gr.190546.115 10.1186/s13059-018-1577-z 10.1371/journal.pgen.1003504 10.1101/gr.216150.116 10.1186/s12864-018-4850-3 10.1101/gr.218149.116 10.1101/gr.235747.118 10.1007/s10577-017-9570-z 10.1093/gbe/evv005 10.1126/science.aac7247 10.1093/nar/gkq132 10.1093/molbev/msy143 10.7554/eLife.13926 10.1093/nar/gkw1067 10.1074/jbc.273.2.891 10.1371/journal.pgen.1008036 10.1371/journal.pone.0027513 10.1093/nar/gkz1138 10.1038/nature11243 10.1016/j.cell.2005.01.001 10.1073/pnas.1016071107 10.1016/j.molcel.2013.01.038 10.1038/s41576-019-0128-0 10.1016/j.stem.2015.02.013 10.1073/pnas.1507125112 10.1038/nn.4229 10.1016/j.celrep.2013.05.031 10.1016/j.cell.2019.12.015 10.1038/nature07829 10.1038/ng.3286 10.1186/s13059-018-1432-2 10.1093/gbe/evx194 10.1101/gr.234096.117 10.1093/oxfordjournals.molbev.a003768 10.1038/nrg3802 10.1093/molbev/msx219 10.1016/j.ymeth.2009.03.001 10.1186/gb-2013-14-11-r124 10.1093/bioinformatics/btp324 10.1093/bioinformatics/btp352 10.1101/gr.136184.111 10.1186/gb-2008-9-9-r137 10.1038/ng1966 10.1038/nature09692 10.1038/nature01080 10.1038/ncb1076 10.1038/nature03877 10.1038/ng.154 10.1093/nar/gky955 10.1093/bioinformatics/bty191 10.1093/nar/gkg006 10.1093/nar/gkj112 10.1016/S0022-2836(05)80360-2 10.1186/1748-7188-6-26 10.1093/bioinformatics/btt509 10.1093/bioinformatics/btu170 10.1093/bioinformatics/bts635 10.1038/nbt.1621 10.1093/nar/gkv1189 10.1093/nar/gkw257 10.1093/bioinformatics/btt737 10.1093/bib/bbs017 10.1093/bioinformatics/btx364 10.1186/s13059-014-0550-8 10.1007/978-3-319-24277-4 10.1101/gr.09}, note = {First posted as a preprint: 31 May 2020}, abstract = {BACKGROUND: To investigate the mechanisms driving regulatory evolution across tissues, we experimentally mapped promoters, enhancers, and gene expression in the liver, brain, muscle, and testis from ten diverse mammals. RESULTS: The regulatory landscape around genes included both tissue-shared and tissue-specific regulatory regions, where tissue-specific promoters and enhancers evolved most rapidly. Genomic regions switching between promoters and enhancers were more common across species, and less common across tissues within a single species. Long Interspersed Nuclear Elements (LINEs) played recurrent evolutionary roles: LINE L1s were associated with tissue-specific regulatory regions, whereas more ancient LINE L2s were associated with tissue-shared regulatory regions and with those switching between promoter and enhancer signatures across species. CONCLUSIONS: Our analyses of the tissue-specificity and evolutionary stability among promoters and enhancers reveal how specific LINE families have helped shape the dynamic mammalian regulome.},}
- D Azazi, JM Mudge, DT Odom, P Flicek. Functional signatures of evolutionarily young CTCF binding sites. BMC Biol 2020;18(1):132. doi:10.1186/s12915-020-00863-8
[BibTeX] [Abstract]
BACKGROUND: The introduction of novel CTCF binding sites in gene regulatory regions in the rodent lineage is partly the effect of transposable element expansion, particularly in the murine lineage. The exact mechanism and functional impact of evolutionarily novel CTCF binding sites are not yet fully understood. We investigated the impact of novel subspecies-specific CTCF binding sites in two Mus genus subspecies, Mus musculus domesticus and Mus musculus castaneus, that diverged 0.5 million years ago. RESULTS: CTCF binding site evolution is influenced by the action of the B2-B4 family of transposable elements independently in both lineages, leading to the proliferation of novel CTCF binding sites. A subset of evolutionarily young sites may harbour transcriptional functionality as evidenced by the stability of their binding across multiple tissues in M. musculus domesticus (BL6), while overall the distance of subspecies-specific CTCF binding to the nearest transcription start sites and/or topologically associated domains (TADs) is largely similar to musculus-common CTCF sites. Remarkably, we discovered a recurrent regulatory architecture consisting of a CTCF binding site and an interferon gene that appears to have been tandemly duplicated to create a 15-gene cluster on chromosome 4, thus forming a novel BL6 specific immune locus in which CTCF may play a regulatory role. CONCLUSIONS: Our results demonstrate that thousands of CTCF binding sites show multiple functional signatures rapidly after incorporation into the genome.
@Article{32988407, author = {Azazi D and Mudge JM and Odom DT and Flicek P}, title = {Functional signatures of evolutionarily young CTCF binding sites}, journal = {BMC Biol}, volume = {18}, number = {1}, pages = {132}, year = {2020}, doi = {10.1186/s12915-020-00863-8}, note = {First posted as a preprint: 31 January 2020}, abstract = {BACKGROUND: The introduction of novel CTCF binding sites in gene regulatory regions in the rodent lineage is partly the effect of transposable element expansion, particularly in the murine lineage. The exact mechanism and functional impact of evolutionarily novel CTCF binding sites are not yet fully understood. We investigated the impact of novel subspecies-specific CTCF binding sites in two Mus genus subspecies, Mus musculus domesticus and Mus musculus castaneus, that diverged 0.5 million years ago. RESULTS: CTCF binding site evolution is influenced by the action of the B2-B4 family of transposable elements independently in both lineages, leading to the proliferation of novel CTCF binding sites. A subset of evolutionarily young sites may harbour transcriptional functionality as evidenced by the stability of their binding across multiple tissues in M. musculus domesticus (BL6), while overall the distance of subspecies-specific CTCF binding to the nearest transcription start sites and/or topologically associated domains (TADs) is largely similar to musculus-common CTCF sites. Remarkably, we discovered a recurrent regulatory architecture consisting of a CTCF binding site and an interferon gene that appears to have been tandemly duplicated to create a 15-gene cluster on chromosome 4, thus forming a novel BL6 specific immune locus in which CTCF may play a regulatory role. CONCLUSIONS: Our results demonstrate that thousands of CTCF binding sites show multiple functional signatures rapidly after incorporation into the genome.},}
- D Schmidt, MD Wilson, B Ballester, PC Schwalie, GD Brown, A Marshall, C Kutter, S Watt, CP Martinez-Jimenez, S Mackay, I Talianidis, P Flicek, DT Odom. Five-Vertebrate ChIP-seq Reveals the Evolutionary Dynamics of Transcription Factor Binding. Science 2010;328(5981):1036–1040. doi:10.1126/science.1186176
[BibTeX] [Abstract]
Transcription factors (TFs) direct gene expression by binding to DNA regulatory regions. To explore the evolution of gene regulation, we experimentally determined the genome-wide occupancy of two TFs, CEBPA and HNF4A, in livers of five vertebrates. Although each TF displays highly conserved DNA binding preferences, most binding is species-specific, and aligned binding events present in all five species are rare. Regions near genes with expression levels dependent on a TF are often bound by the TF in multiple species, yet show no enhanced DNA sequence constraint. Binding divergence between species can be largely explained by sequence changes to the bound motifs. Among the binding events lost in one lineage, only half are recovered by another binding event within 10 kilobases. Our results reveal large interspecies differences in transcriptional regulation and provide insight into their evolution.
@Article{20378774, author = {Schmidt D and Wilson MD and Ballester B and Schwalie PC and Brown GD and Marshall A and Kutter C and Watt S and Martinez-Jimenez CP and Mackay S and Talianidis I and Flicek P and Odom DT}, title = {Five-Vertebrate ChIP-seq Reveals the Evolutionary Dynamics of Transcription Factor Binding}, journal = {Science}, volume = {328}, number = {5981}, pages = {1036--1040}, year = {2010}, doi = {10.1126/science.1186176}, abstract = {Transcription factors (TFs) direct gene expression by binding to DNA regulatory regions. To explore the evolution of gene regulation, we experimentally determined the genome-wide occupancy of two TFs, CEBPA and HNF4A, in livers of five vertebrates. Although each TF displays highly conserved DNA binding preferences, most binding is species-specific, and aligned binding events present in all five species are rare. Regions near genes with expression levels dependent on a TF are often bound by the TF in multiple species, yet show no enhanced DNA sequence constraint. Binding divergence between species can be largely explained by sequence changes to the bound motifs. Among the binding events lost in one lineage, only half are recovered by another binding event within 10 kilobases. Our results reveal large interspecies differences in transcriptional regulation and provide insight into their evolution.},} - AJ Faure, D Schmidt, S Watt, PC Schwalie, MD Wilson, H Xu, RG Ramsay, DT Odom, P Flicek. Cohesin regulates tissue-specific expression by stabilizing highly occupied cis-regulatory modules. Genome Res 2012;22(11):2163–2175. doi:10.1101/gr.136507.111
[BibTeX] [Abstract]
The cohesin protein complex contributes to transcriptional regulation in a CTCF-independent manner by colocalizing with master regulators at tissue-specific loci. The regulation of transcription involves the concerted action of multiple transcription factors (TFs) and cohesin’s role in this context of combinatorial TF binding remains unexplored. To investigate cohesin-non-CTCF (CNC) binding events in vivo we mapped cohesin and CTCF, as well as a collection of tissue-specific and ubiquitous transcriptional regulators using ChIP-seq in primary mouse liver. We observe a positive correlation between the number of distinct TFs bound and the presence of CNC sites. In contrast to regions of the genome where cohesin and CTCF colocalize, CNC sites coincide with the binding of master regulators and enhancer-markers and are significantly associated with liver-specific expressed genes. We also show that cohesin presence partially explains the commonly observed discrepancy between TF motif score and ChIP signal. Evidence from these statistical analyses in wild-type cells, and comparisons to maps of TF binding in Rad21-cohesin haploinsufficient mouse liver, suggests that cohesin helps to stabilize large protein-DNA complexes. Finally, we observe that the presence of mirrored CTCF binding events at promoters and their nearby cohesin-bound enhancers is associated with elevated expression levels
@Article{22780989, author = {Faure AJ and Schmidt D and Watt S and Schwalie PC and Wilson MD and Xu H and Ramsay RG and Odom DT and Flicek P}, title = {Cohesin regulates tissue-specific expression by stabilizing highly occupied cis-regulatory modules}, journal = {Genome Res}, volume = {22}, number = {11}, pages = {2163--2175}, year = {2012}, doi = {10.1101/gr.136507.111}, howpublished = {Advanced online publication: 10 July 2012}, abstract = {The cohesin protein complex contributes to transcriptional regulation in a CTCF-independent manner by colocalizing with master regulators at tissue-specific loci. The regulation of transcription involves the concerted action of multiple transcription factors (TFs) and cohesin's role in this context of combinatorial TF binding remains unexplored. To investigate cohesin-non-CTCF (CNC) binding events in vivo we mapped cohesin and CTCF, as well as a collection of tissue-specific and ubiquitous transcriptional regulators using ChIP-seq in primary mouse liver. We observe a positive correlation between the number of distinct TFs bound and the presence of CNC sites. In contrast to regions of the genome where cohesin and CTCF colocalize, CNC sites coincide with the binding of master regulators and enhancer-markers and are significantly associated with liver-specific expressed genes. We also show that cohesin presence partially explains the commonly observed discrepancy between TF motif score and ChIP signal. Evidence from these statistical analyses in wild-type cells, and comparisons to maps of TF binding in Rad21-cohesin haploinsufficient mouse liver, suggests that cohesin helps to stabilize large protein-DNA complexes. Finally, we observe that the presence of mirrored CTCF binding events at promoters and their nearby cohesin-bound enhancers is associated with elevated expression levels},} - A Scally, JY Dutheil, LW Hillier, GE Jordan, I Goodhead, J Herrero, A Hobolth, T Lappalainen, T Mailund, T Marques-Bonet, S McCarthy, SH Montgomery, PC Schwalie, YA Tang, MC Ward, Y Xue, B Yngvadottir, C Alkan, LN Andersen, Q Ayub, EV Ball, K Beal, BJ Bradley, Y Chen, CM Clee, S Fitzgerald, TA Graves, Y Gu, P Heath, A Heger, E Karakoc, A Kolb-Kokocinski, GK Laird, G Lunter, S Meader, M Mort, JC Mullikin, K Munch, TD O’Connor, AD Phillips, J Prado-Martinez, AS Rogers, S Sajjadian, D Schmidt, K Shaw, JT Simpson, PD Stenson, DJ Turner, L Vigilant, AJ Vilella, W Whitener, B Zhu, DN Cooper, de P Jong, ET Dermitzakis, EE Eichler, P Flicek, N Goldman, NI Mundy, Z Ning, DT Odom, CP Ponting, MA Quail, OA Ryder, SM Searle, WC Warren, RK Wilson, MH Schierup, J Rogers, C Tyler-Smith, R Durbin. Insights into hominid evolution from the gorilla genome sequence. Nature 2012;483(7388):169–175. doi:10.1038/nature10842
[BibTeX] [Abstract]
Gorillas are humans’ closest living relatives after chimpanzees, and are of comparable importance for the study of human origins and evolution. Here we present the assembly and analysis of a genome sequence for the western lowland gorilla, and compare the whole genomes of all extant great ape genera. We propose a synthesis of genetic and fossil evidence consistent with placing the human-chimpanzee and human-chimpanzee-gorilla speciation events at approximately 6 and 10 million years ago. In 30\% of the genome, gorilla is closer to human or chimpanzee than the latter are to each other; this is rarer around coding genes, indicating pervasive selection throughout great ape evolution, and has functional consequences in gene expression. A comparison of protein coding genes reveals approximately 500 genes showing accelerated evolution on each of the gorilla, human and chimpanzee lineages, and evidence for parallel acceleration, particularly of genes involved in hearing. We also compare the western and eastern gorilla species, estimating an average sequence divergence time 1.75 million years ago, but with evidence for more recent genetic exchange and a population bottleneck in the eastern species. The use of the genome sequence in these and future analyses will promote a deeper understanding of great ape biology and evolution
@Article{22398555, author = {Scally A and Dutheil JY and Hillier LW and Jordan GE and Goodhead I and Herrero J and Hobolth A and Lappalainen T and Mailund T and Marques-Bonet T and McCarthy S and Montgomery SH and Schwalie PC and Tang YA and Ward MC and Xue Y and Yngvadottir B and Alkan C and Andersen LN and Ayub Q and Ball EV and Beal K and Bradley BJ and Chen Y and Clee CM and Fitzgerald S and Graves TA and Gu Y and Heath P and Heger A and Karakoc E and Kolb-Kokocinski A and Laird GK and Lunter G and Meader S and Mort M and Mullikin JC and Munch K and O'Connor TD and Phillips AD and Prado-Martinez J and Rogers AS and Sajjadian S and Schmidt D and Shaw K and Simpson JT and Stenson PD and Turner DJ and Vigilant L and Vilella AJ and Whitener W and Zhu B and Cooper DN and de Jong P and Dermitzakis ET and Eichler EE and Flicek P and Goldman N and Mundy NI and Ning Z and Odom DT and Ponting CP and Quail MA and Ryder OA and Searle SM and Warren WC and Wilson RK and Schierup MH and Rogers J and Tyler-Smith C and Durbin R}, title = {Insights into hominid evolution from the gorilla genome sequence}, journal = {Nature}, volume = {483}, number = {7388}, pages = {169--175}, year = {2012}, doi = {10.1038/nature10842}, abstract = {Gorillas are humans' closest living relatives after chimpanzees, and are of comparable importance for the study of human origins and evolution. Here we present the assembly and analysis of a genome sequence for the western lowland gorilla, and compare the whole genomes of all extant great ape genera. We propose a synthesis of genetic and fossil evidence consistent with placing the human-chimpanzee and human-chimpanzee-gorilla speciation events at approximately 6 and 10 million years ago. In 30\% of the genome, gorilla is closer to human or chimpanzee than the latter are to each other; this is rarer around coding genes, indicating pervasive selection throughout great ape evolution, and has functional consequences in gene expression. A comparison of protein coding genes reveals approximately 500 genes showing accelerated evolution on each of the gorilla, human and chimpanzee lineages, and evidence for parallel acceleration, particularly of genes involved in hearing. We also compare the western and eastern gorilla species, estimating an average sequence divergence time 1.75 million years ago, but with evidence for more recent genetic exchange and a population bottleneck in the eastern species. The use of the genome sequence in these and future analyses will promote a deeper understanding of great ape biology and evolution},} - D Schmidt, PC Schwalie, MD Wilson, B Ballester, A Gonçalves, C Kutter, GD Brown, A Marshall, P Flicek, DT Odom. Waves of Retrotransposon Expansion Remodel Genome Organization and CTCF Binding in Multiple Mammalian Lineages. Cell 2012;148(1-2):335–348. doi:10.1016/j.cell.2011.11.058
[BibTeX] [Abstract]
CTCF-binding locations represent regulatory sequences that are highly constrained over the course of evolution. To gain insight into how these DNA elements are conserved and spread through the genome, we defined the full spectrum of CTCF-binding sites, including a 33/34-mer motif, and identified over five thousand highly conserved, robust, and tissue-independent CTCF-binding locations by comparing ChIP-seq data from six mammals. Our data indicate that activation of retroelements has produced species-specific expansions of CTCF binding in rodents, dogs, and opossum, which often functionally serve as chromatin and transcriptional insulators. We discovered fossilized repeat elements flanking deeply conserved CTCF-binding regions, indicating that similar retrotransposon expansions occurred hundreds of millions of years ago. Repeat-driven dispersal of CTCF binding is a fundamental, ancient, and still highly active mechanism of genome evolution in mammalian lineages. PAPERCLIP
@Article{22244452, author = {Schmidt D and Schwalie PC and Wilson MD and Ballester B and Gonçalves A and Kutter C and Brown GD and Marshall A and Flicek P and Odom DT}, title = {Waves of Retrotransposon Expansion Remodel Genome Organization and CTCF Binding in Multiple Mammalian Lineages}, journal = {Cell}, volume = {148}, number = {1-2}, pages = {335--348}, year = {2012}, doi = {10.1016/j.cell.2011.11.058}, abstract = {CTCF-binding locations represent regulatory sequences that are highly constrained over the course of evolution. To gain insight into how these DNA elements are conserved and spread through the genome, we defined the full spectrum of CTCF-binding sites, including a 33/34-mer motif, and identified over five thousand highly conserved, robust, and tissue-independent CTCF-binding locations by comparing ChIP-seq data from six mammals. Our data indicate that activation of retroelements has produced species-specific expansions of CTCF binding in rodents, dogs, and opossum, which often functionally serve as chromatin and transcriptional insulators. We discovered fossilized repeat elements flanking deeply conserved CTCF-binding regions, indicating that similar retrotransposon expansions occurred hundreds of millions of years ago. Repeat-driven dispersal of CTCF binding is a fundamental, ancient, and still highly active mechanism of genome evolution in mammalian lineages. PAPERCLIP},} - T Schauer, PC Schwalie, A Handley, CE Margulies, P Flicek, AG Ladurner. CAST-ChIP Maps Cell-Type-Specific Chromatin States in the Drosophila Central Nervous System. Cell Rep 2013;5(1):271–282. doi:10.1016/j.celrep.2013.09.001
[BibTeX] [Abstract]
Chromatin organization and gene activity are responsive to developmental and environmental cues. Although many genes are transcribed throughout development and across cell types, much of gene regulation is highly cell-type specific. To readily track chromatin features at the resolution of cell types within complex tissues, we developed and validated chromatin affinity purification from specific cell types by chromatin immunoprecipitation (CAST-ChIP), a broadly applicable biochemical procedure. RNA polymerase II (Pol II) CAST-ChIP identifies ∼1,500 neuronal and glia-specific genes in differentiated cells within the adult Drosophila brain. In contrast, the histone H2A.Z is distributed similarly across cell types and throughout development, marking cell-type-invariant Pol II-bound regions. Our study identifies H2A.Z as an active chromatin signature that is refractory to changes across cell fates. Thus, CAST-ChIP powerfully identifies cell-type-specific as well as cell-type-invariant chromatin states, enabling the systematic dissection of chromatin structure and gene regulation within complex tissues such as the brain
@Article{24095734, author = {Schauer T and Schwalie PC and Handley A and Margulies CE and Flicek P and Ladurner AG}, title = {CAST-ChIP Maps Cell-Type-Specific Chromatin States in the Drosophila Central Nervous System}, journal = {Cell Rep}, volume = {5}, number = {1}, pages = {271--282}, year = {2013}, doi = {10.1016/j.celrep.2013.09.001}, howpublished = {Advanced online publication: 3 October 2013}, abstract = {Chromatin organization and gene activity are responsive to developmental and environmental cues. Although many genes are transcribed throughout development and across cell types, much of gene regulation is highly cell-type specific. To readily track chromatin features at the resolution of cell types within complex tissues, we developed and validated chromatin affinity purification from specific cell types by chromatin immunoprecipitation (CAST-ChIP), a broadly applicable biochemical procedure. RNA polymerase II (Pol II) CAST-ChIP identifies ∼1,500 neuronal and glia-specific genes in differentiated cells within the adult Drosophila brain. In contrast, the histone H2A.Z is distributed similarly across cell types and throughout development, marking cell-type-invariant Pol II-bound regions. Our study identifies H2A.Z as an active chromatin signature that is refractory to changes across cell fates. Thus, CAST-ChIP powerfully identifies cell-type-specific as well as cell-type-invariant chromatin states, enabling the systematic dissection of chromatin structure and gene regulation within complex tissues such as the brain},} - PC Schwalie, MC Ward, CE Cain, AJ Faure, Y Gilad, DT Odom, P Flicek. Co-binding by YY1 identifies the transcriptionally active, highly conserved set of CTCF-bound regions in primate genomes. Genome Biol 2013;14(12):R148. doi:10.1186/gb-2013-14-12-r148
[BibTeX] [Abstract]
\textbf{BACKGROUND:} The genomic binding of CTCF is highly conserved across mammals, but the mechanisms that underlie its stability are poorly understood. One transcription factor known to functionally interact with CTCF in the context of X-chromosome inactivation is the ubiquitously expressed YY1. Because combinatorial transcription factor binding can contribute to the evolutionary stabilization of regulatory regions, we tested whether YY1 and CTCF co-binding could in part account for conservation of CTCF binding.
\textbf{RESULTS:} Combined analysis of CTCF and YY1 binding in lymphoblastoid cell lines from seven primates, as well as in mouse and human livers, reveals extensive genome-wide co-localization specifically at evolutionarily stable CTCF-bound regions. CTCF-YY1 co-bound regions resemble regions bound by YY1 alone, as they enrich for co-bound transcription factors, RNA polymerase II and active histone marks. Although these highly conserved, transcriptionally active CTCF-YY1 co-bound regions are often promoter-proximal, gene-distal sites show similar molecular features.
\textbf{CONCLUSIONS:} Our results reveal that these two ubiquitously expressed, multi-functional zinc-finger proteins collaborate in functionally active regions to stabilize one another’s genome-wide binding across primate evolution@Article{24380390, author = {Schwalie PC and Ward MC and Cain CE and Faure AJ and Gilad Y and Odom DT and Flicek P}, title = {Co-binding by YY1 identifies the transcriptionally active, highly conserved set of CTCF-bound regions in primate genomes}, journal = {Genome Biol}, volume = {14}, number = {12}, pages = {R148}, year = {2013}, doi = {10.1186/gb-2013-14-12-r148}, abstract = {\textbf{BACKGROUND:} The genomic binding of CTCF is highly conserved across mammals, but the mechanisms that underlie its stability are poorly understood. One transcription factor known to functionally interact with CTCF in the context of X-chromosome inactivation is the ubiquitously expressed YY1. Because combinatorial transcription factor binding can contribute to the evolutionary stabilization of regulatory regions, we tested whether YY1 and CTCF co-binding could in part account for conservation of CTCF binding.
\textbf{RESULTS:} Combined analysis of CTCF and YY1 binding in lymphoblastoid cell lines from seven primates, as well as in mouse and human livers, reveals extensive genome-wide co-localization specifically at evolutionarily stable CTCF-bound regions. CTCF-YY1 co-bound regions resemble regions bound by YY1 alone, as they enrich for co-bound transcription factors, RNA polymerase II and active histone marks. Although these highly conserved, transcriptionally active CTCF-YY1 co-bound regions are often promoter-proximal, gene-distal sites show similar molecular features.
\textbf{CONCLUSIONS:} Our results reveal that these two ubiquitously expressed, multi-functional zinc-finger proteins collaborate in functionally active regions to stabilize one another's genome-wide binding across primate evolution},} - MC Ward, MD Wilson, NL Barbosa-Morais, D Schmidt, R Stark, Q Pan, PC Schwalie, S Menon, M Lukk, S Watt, D Thybert, C Kutter, K Kirschner, P Flicek, BJ Blencowe, DT Odom. Latent regulatory potential of human-specific repetitive elements. Mol Cell 2013;49(2):262–272. doi:10.1016/j.molcel.2012.11.013
[BibTeX] [Abstract]
At least half of the human genome is derived from repetitive elements, which are often lineage specific and silenced by a variety of genetic and epigenetic mechanisms. Using a transchromosomic mouse strain that transmits an almost complete single copy of human chromosome 21 via the female germline, we show that a heterologous regulatory environment can transcriptionally activate transposon-derived human regulatory regions. In the mouse nucleus, hundreds of locations on human chromosome 21 newly associate with activating histone modifications in both somatic and germline tissues, and influence the gene expression of nearby transcripts. These regions are enriched with primate and human lineage-specific transposable elements, and their activation corresponds to changes in DNA methylation at CpG dinucleotides. This study reveals the latent regulatory potential of the repetitive human genome and illustrates the species specificity of mechanisms that control it
@Article{23246434, author = {Ward MC and Wilson MD and Barbosa-Morais NL and Schmidt D and Stark R and Pan Q and Schwalie PC and Menon S and Lukk M and Watt S and Thybert D and Kutter C and Kirschner K and Flicek P and Blencowe BJ and Odom DT}, title = {Latent regulatory potential of human-specific repetitive elements}, journal = {Mol Cell}, volume = {49}, number = {2}, pages = {262--272}, year = {2013}, doi = {10.1016/j.molcel.2012.11.013}, abstract = {At least half of the human genome is derived from repetitive elements, which are often lineage specific and silenced by a variety of genetic and epigenetic mechanisms. Using a transchromosomic mouse strain that transmits an almost complete single copy of human chromosome 21 via the female germline, we show that a heterologous regulatory environment can transcriptionally activate transposon-derived human regulatory regions. In the mouse nucleus, hundreds of locations on human chromosome 21 newly associate with activating histone modifications in both somatic and germline tissues, and influence the gene expression of nearby transcripts. These regions are enriched with primate and human lineage-specific transposable elements, and their activation corresponds to changes in DNA methylation at CpG dinucleotides. This study reveals the latent regulatory potential of the repetitive human genome and illustrates the species specificity of mechanisms that control it},}
- T Rensch, D Villar, J Horvath, DT Odom, P Flicek. Mitochondrial heteroplasmy in vertebrates using ChIP-sequencing data. Genome Biol 2016;17(1):139. doi:10.1186/s13059-016-0996-y
[BibTeX] [Abstract]
\textbf{BACKGROUND:} Mitochondrial heteroplasmy, the presence of more than one mitochondrial DNA (mtDNA) variant in a cell or individual, is not as uncommon as previously thought. It is mostly due to the high mutation rate of the mtDNA and limited repair mechanisms present in the mitochondrion. Motivated by mitochondrial diseases, much focus has been placed into studying this phenomenon in human samples and in medical contexts. To place these results in an evolutionary context and to explore general principles of heteroplasmy, we describe an integrated cross-species evaluation of heteroplasmy in mammals that exploits previously reported NGS data. Focusing on ChIP-seq experiments, we developed a novel approach to detect heteroplasmy from the concomitant mitochondrial DNA fraction sequenced in these experiments.
\textbf{RESULTS:} We first demonstrate that the sequencing coverage of mtDNA in ChIP-seq experiments is sufficient for heteroplasmy detection. We then describe a novel detection method for accurate detection of heteroplasmies, which also accounts for the error rate of NGS technology. Applying this method to 79 individuals from 16 species resulted in 107 heteroplasmic positions present in a total of 45 individuals. Further analysis revealed that the majority of detected heteroplasmies occur in intergenic regions.
\textbf{CONCLUSION:} In addition to documenting the prevalence of mtDNA in ChIP-seq data, the results of our mitochondrial heteroplasmy detection method suggest that mitochondrial heteroplasmies identified across vertebrates share similar characteristics as found for human heteroplasmies. Although largely consistent with previous studies in individual vertebrates, our integrated cross-species analysis provides valuable insights into the evolutionary dynamics of mitochondrial heteroplasmy@Article{27349964, author = {Rensch T and Villar D and Horvath J and Odom DT and Flicek P}, title = {Mitochondrial heteroplasmy in vertebrates using ChIP-sequencing data}, journal = {Genome Biol}, volume = {17}, number = {1}, pages = {139}, year = {2016}, doi = {10.1186/s13059-016-0996-y}, abstract = {\textbf{BACKGROUND:} Mitochondrial heteroplasmy, the presence of more than one mitochondrial DNA (mtDNA) variant in a cell or individual, is not as uncommon as previously thought. It is mostly due to the high mutation rate of the mtDNA and limited repair mechanisms present in the mitochondrion. Motivated by mitochondrial diseases, much focus has been placed into studying this phenomenon in human samples and in medical contexts. To place these results in an evolutionary context and to explore general principles of heteroplasmy, we describe an integrated cross-species evaluation of heteroplasmy in mammals that exploits previously reported NGS data. Focusing on ChIP-seq experiments, we developed a novel approach to detect heteroplasmy from the concomitant mitochondrial DNA fraction sequenced in these experiments.
\textbf{RESULTS:} We first demonstrate that the sequencing coverage of mtDNA in ChIP-seq experiments is sufficient for heteroplasmy detection. We then describe a novel detection method for accurate detection of heteroplasmies, which also accounts for the error rate of NGS technology. Applying this method to 79 individuals from 16 species resulted in 107 heteroplasmic positions present in a total of 45 individuals. Further analysis revealed that the majority of detected heteroplasmies occur in intergenic regions.
\textbf{CONCLUSION:} In addition to documenting the prevalence of mtDNA in ChIP-seq data, the results of our mitochondrial heteroplasmy detection method suggest that mitochondrial heteroplasmies identified across vertebrates share similar characteristics as found for human heteroplasmies. Although largely consistent with previous studies in individual vertebrates, our integrated cross-species analysis provides valuable insights into the evolutionary dynamics of mitochondrial heteroplasmy},}
2026
- SJ Aitken, F Connor, C Feig, TF Rayner, M Lukk, J Luft, S Aitken, C Arnedo-Pac, JF Hayes, MD Nicholson, A Ewing, V Sundaram, JC Verburg, J Connelly, CJ Anderson, M Behm, S Campbell, M Daunesse, VB Kaiser, E Kentepozidou, O Pich, AM Redmond, J Santoyo-Lopez, I Sentís, L Talmane, Liver Cancer Evolution Consortium, P Flicek, N López-Bigas, CA Semple, MS Taylor, DT Odom. Genetic background sets the trajectory of experimental cancer evolution. Nature 2026. doi:10.1038/s41586-026-10821-z
[BibTeX] [Abstract]
Human cancers are heterogeneous\textsuperscript{1}. Dissecting how germline genetic variation and environmental factors shape tumour evolution using human datasets is limited by inherent diversity in genetic backgrounds\textsuperscript{2} and environmental exposures\textsuperscript{3-5}. Here, to overcome these limitations, we re-ran early tumour evolution hundreds of times in diverged inbred mouse strains, generating matched histology and whole-genome and transcriptome sequences. The sex, environment and carcinogenic exposures were all controlled, and the study design allowed us to capture genetic variation comparable with that observed across human populations while exploiting the nested hierarchical structure of strain-litter-animal-tumour relationships. Our analyses reveal that epistatic interactions between genetic background and acquired somatic mutations result in population-specific disease progression, including choice of driver mutations, occurrence of whole-genome duplication and subclonal selection dynamics that mirror both cancer susceptibility and tumour growth rate. Even modest genetic divergence, comparable with that found across human ancestry groups, can strikingly alter selection pressures during cancer development to shape both cancer risk and the trajectory of tumour evolution.
@Article{42486977, author = {Aitken SJ and Connor F and Feig C and Rayner TF and Lukk M and Luft J and Aitken S and Arnedo-Pac C and Hayes JF and Nicholson MD and Ewing A and Sundaram V and Verburg JC and Connelly J and Anderson CJ and Behm M and Campbell S and Daunesse M and Kaiser VB and Kentepozidou E and Pich O and Redmond AM and Santoyo-Lopez J and Sentís I and Talmane L and {Liver Cancer Evolution Consortium} and Flicek P and López-Bigas N and Semple CA and Taylor MS and Odom DT}, title = {Genetic background sets the trajectory of experimental cancer evolution}, journal = {Nature}, year = {2026}, doi = {10.1038/s41586-026-10821-z}, howpublished = {Advanced online publication: 22 July 2026}, note = {First posted as a preprint: 15 January 2015}, abstract = {Human cancers are heterogeneous\textsuperscript{1}. Dissecting how germline genetic variation and environmental factors shape tumour evolution using human datasets is limited by inherent diversity in genetic backgrounds\textsuperscript{2} and environmental exposures\textsuperscript{3-5}. Here, to overcome these limitations, we re-ran early tumour evolution hundreds of times in diverged inbred mouse strains, generating matched histology and whole-genome and transcriptome sequences. The sex, environment and carcinogenic exposures were all controlled, and the study design allowed us to capture genetic variation comparable with that observed across human populations while exploiting the nested hierarchical structure of strain-litter-animal-tumour relationships. Our analyses reveal that epistatic interactions between genetic background and acquired somatic mutations result in population-specific disease progression, including choice of driver mutations, occurrence of whole-genome duplication and subclonal selection dynamics that mirror both cancer susceptibility and tumour growth rate. Even modest genetic divergence, comparable with that found across human ancestry groups, can strikingly alter selection pressures during cancer development to shape both cancer risk and the trajectory of tumour evolution.},}
2025
- P Flicek. Mapping the genomes of Earth’s interconnected biodiversity. Front Sci 2025;3:1690100. doi:10.3389/fsci.2025.1690100
[BibTeX]@Article{, author = {Flicek P}, title = {Mapping the genomes of Earth's interconnected biodiversity}, journal = {Front Sci}, volume = {3}, pages = {1690100}, year = {2025}, doi = {10.3389/fsci.2025.1690100}, }
2024
- M Rimoldi, N Wang, J Zhang, D Villar, DT Odom, J Taipale, P Flicek, M Roller. DNA methylation patterns of transcription factor binding regions characterize their functional and evolutionary contexts. Genome Biol 2024;25(1):146. doi:10.1186/s13059-024-03218-6
[BibTeX] [Abstract]
BACKGROUND: DNA methylation is an important epigenetic modification which has numerous roles in modulating genome function. Its levels are spatially correlated across the genome, typically high in repressed regions but low in transcription factor (TF) binding sites and active regulatory regions. However, the mechanisms establishing genome-wide and TF binding site methylation patterns are still unclear. RESULTS: Here we use a comparative approach to investigate the association of DNA methylation to TF binding evolution in mammals. Specifically, we experimentally profile DNA methylation and combine this with published occupancy profiles of five distinct TFs (CTCF, CEBPA, HNF4A, ONECUT1, FOXA1) in the liver of five mammalian species (human, macaque, mouse, rat, dog). TF binding sites are lowly methylated, but they often also have intermediate methylation levels. Furthermore, biding sites are influenced by the methylation status of CpGs in their wider binding regions even when CpGs are absent from the core binding motif. Employing a classification and clustering approach, we extract distinct and species-conserved patterns of DNA methylation levels at TF binding regions. CEBPA, HNF4A, ONECUT1, and FOXA1 share the same methylation patterns, while CTCF’s differ. These patterns characterize alternative functions and chromatin landscapes of TF-bound regions. Leveraging our phylogenetic framework, we find DNA methylation gain upon evolutionary loss of TF occupancy, indicating coordinated evolution. Furthermore, each methylation pattern has its own evolutionary trajectory reflecting its genomic contexts. CONCLUSIONS: Our epigenomic analyses indicate a role for DNA methylation in TF binding changes across species including that specific DNA methylation profiles characterize TF binding and are associated with their regulatory activity, chromatin contexts, and evolutionary trajectories.
@Article{38844976, author = {Rimoldi M and Wang N and Zhang J and Villar D and Odom DT and Taipale J and Flicek P and Roller M}, title = {DNA methylation patterns of transcription factor binding regions characterize their functional and evolutionary contexts}, journal = {Genome Biol}, volume = {25}, number = {1}, pages = {146}, year = {2024}, doi = {10.1186/s13059-024-03218-6}, note = {First posted as a preprint: 21 July 2022}, abstract = {BACKGROUND: DNA methylation is an important epigenetic modification which has numerous roles in modulating genome function. Its levels are spatially correlated across the genome, typically high in repressed regions but low in transcription factor (TF) binding sites and active regulatory regions. However, the mechanisms establishing genome-wide and TF binding site methylation patterns are still unclear. RESULTS: Here we use a comparative approach to investigate the association of DNA methylation to TF binding evolution in mammals. Specifically, we experimentally profile DNA methylation and combine this with published occupancy profiles of five distinct TFs (CTCF, CEBPA, HNF4A, ONECUT1, FOXA1) in the liver of five mammalian species (human, macaque, mouse, rat, dog). TF binding sites are lowly methylated, but they often also have intermediate methylation levels. Furthermore, biding sites are influenced by the methylation status of CpGs in their wider binding regions even when CpGs are absent from the core binding motif. Employing a classification and clustering approach, we extract distinct and species-conserved patterns of DNA methylation levels at TF binding regions. CEBPA, HNF4A, ONECUT1, and FOXA1 share the same methylation patterns, while CTCF's differ. These patterns characterize alternative functions and chromatin landscapes of TF-bound regions. Leveraging our phylogenetic framework, we find DNA methylation gain upon evolutionary loss of TF occupancy, indicating coordinated evolution. Furthermore, each methylation pattern has its own evolutionary trajectory reflecting its genomic contexts. CONCLUSIONS: Our epigenomic analyses indicate a role for DNA methylation in TF binding changes across species including that specific DNA methylation profiles characterize TF binding and are associated with their regulatory activity, chromatin contexts, and evolutionary trajectories.},} - B Benjelloun, K Leempoel, F Boyer, S Stucki, I Streeter, P Orozco-terWengel, FJ Alberto, B Servin, F Biscarini, A Alberti, S Engelen, A Stella, L Colli, E Coissac, MW Bruford, P Ajmone-Marsan, R Negrini, L Clarke, P Flicek, A Chikhi, S Joost, P Taberlet, F Pompanon. Multiple genomic solutions for local adaptation in two closely related species (sheep and goats) facing the same climatic constraints. Mol Ecol 2024;33(20):e17257. doi:10.1111/mec.17257
[BibTeX] [Abstract]
The question of how local adaptation takes place remains a fundamental question in evolutionary biology. The variation of allele frequencies in genes under selection over environmental gradients remains mainly theoretical and its empirical assessment would help understanding how adaptation happens over environmental clines. To bring new insights to this issue we set up a broad framework which aimed to compare the adaptive trajectories over environmental clines in two domesticated mammal species co-distributed in diversified landscapes. We sequenced the genomes of 160 sheep and 161 goats extensively managed along environmental gradients, including temperature, rainfall, seasonality and altitude, to identify genes and biological processes shaping local adaptation. Allele frequencies at putatively adaptive loci were rarely found to vary gradually along environmental gradients, but rather displayed a discontinuous shift at the extremities of environmental clines. Of the 430 candidate adaptive genes identified, only 6 were orthologous between sheep and goats and those responded differently to environmental pressures, suggesting different putative mechanisms involved in local adaptation in these two closely related species. Interestingly, the genomes of the 2 species were impacted differently by the environment, genes related to signatures of selection were most related to altitude, slope and rainfall seasonality for sheep, and summer temperature and spring rainfall for goats. The diversity of candidate adaptive pathways may result from a high number of biological functions involved in the adaptations to multiple eco-climatic gradients, and a differential role of climatic drivers on the two species, despite their co-distribution along the same environmental gradients. This study describes empirical examples of clinal variation in putatively adaptive alleles with different patterns in allele frequency distributions over continuous environmental gradients, thus showing the diversity of genetic responses in adaptive landscapes and opening new horizons for understanding genomics of adaptation in mammalian species and beyond.
@Article{38149334, author = {Benjelloun B and Leempoel K and Boyer F and Stucki S and Streeter I and Orozco-terWengel P and Alberto FJ and Servin B and Biscarini F and Alberti A and Engelen S and Stella A and Colli L and Coissac E and Bruford MW and Ajmone-Marsan P and Negrini R and Clarke L and Flicek P and Chikhi A and Joost S and Taberlet P and Pompanon F}, title = {Multiple genomic solutions for local adaptation in two closely related species (sheep and goats) facing the same climatic constraints}, journal = {Mol Ecol}, volume = {33}, number = {20}, pages = {e17257}, year = {2024}, doi = {10.1111/mec.17257}, howpublished = {Advanced online publication: 27 December 2023}, note = {First posted as a preprint: 18 November 2021}, abstract = {The question of how local adaptation takes place remains a fundamental question in evolutionary biology. The variation of allele frequencies in genes under selection over environmental gradients remains mainly theoretical and its empirical assessment would help understanding how adaptation happens over environmental clines. To bring new insights to this issue we set up a broad framework which aimed to compare the adaptive trajectories over environmental clines in two domesticated mammal species co-distributed in diversified landscapes. We sequenced the genomes of 160 sheep and 161 goats extensively managed along environmental gradients, including temperature, rainfall, seasonality and altitude, to identify genes and biological processes shaping local adaptation. Allele frequencies at putatively adaptive loci were rarely found to vary gradually along environmental gradients, but rather displayed a discontinuous shift at the extremities of environmental clines. Of the 430 candidate adaptive genes identified, only 6 were orthologous between sheep and goats and those responded differently to environmental pressures, suggesting different putative mechanisms involved in local adaptation in these two closely related species. Interestingly, the genomes of the 2 species were impacted differently by the environment, genes related to signatures of selection were most related to altitude, slope and rainfall seasonality for sheep, and summer temperature and spring rainfall for goats. The diversity of candidate adaptive pathways may result from a high number of biological functions involved in the adaptations to multiple eco-climatic gradients, and a differential role of climatic drivers on the two species, despite their co-distribution along the same environmental gradients. This study describes empirical examples of clinal variation in putatively adaptive alleles with different patterns in allele frequency distributions over continuous environmental gradients, thus showing the diversity of genetic responses in adaptive landscapes and opening new horizons for understanding genomics of adaptation in mammalian species and beyond.},}
2023
- WW Liao, M Asri, J Ebler, D Doerr, M Haukness, G Hickey, S Lu, JK Lucas, J Monlong, HJ Abel, S Buonaiuto, XH Chang, H Cheng, J Chu, V Colonna, JM Eizenga, X Feng, C Fischer, RS Fulton, S Garg, C Groza, A Guarracino, WT Harvey, S Heumos, K Howe, M Jain, TY Lu, C Markello, FJ Martin, MW Mitchell, KM Munson, MN Mwaniki, AM Novak, HE Olsen, T Pesout, D Porubsky, P Prins, JA Sibbesen, J Sirén, C Tomlinson, F Villani, MR Vollger, LL Antonacci-Fulton, G Baid, CA Baker, A Belyaeva, K Billis, A Carroll, PC Chang, S Cody, DE Cook, RM Cook-Deegan, OE Cornejo, M Diekhans, P Ebert, S Fairley, O Fedrigo, AL Felsenfeld, G Formenti, A Frankish, Y Gao, NA Garrison, CG Giron, RE Green, L Haggerty, K Hoekzema, T Hourlier, HP Ji, EE Kenny, BA Koenig, A Kolesnikov, JO Korbel, J Kordosky, S Koren, H Lee, AP Lewis, H Magalhães, S Marco-Sola, P Marijon, A McCartney, J McDaniel, J Mountcastle, M Nattestad, S Nurk, ND Olson, AB Popejoy, D Puiu, M Rautiainen, AA Regier, A Rhie, S Sacco, AD Sanders, VA Schneider, BI Schultz, K Shafin, MW Smith, HJ Sofia, AN Abou Tayoun, F Thibaud-Nissen, FF Tricomi, J Wagner, B Walenz, JMD Wood, AV Zimin, G Bourque, MJP Chaisson, P Flicek, AM Phillippy, JM Zook, EE Eichler, D Haussler, T Wang, ED Jarvis, KH Miga, E Garrison, T Marschall, IM Hall, H Li, B Paten. A draft human pangenome reference. Nature 2023;617(7960):312–324. doi:10.1038/s41586-023-05896-x
[BibTeX] [Abstract]
Here the Human Pangenome Reference Consortium presents a first draft of the human pangenome reference. The pangenome contains 47 phased, diploid assemblies from a cohort of genetically diverse individuals\textsuperscript{1}. These assemblies cover more than 99\% of the expected sequence in each genome and are more than 99\% accurate at the structural and base pair levels. Based on alignments of the assemblies, we generate a draft pangenome that captures known variants and haplotypes and reveals new alleles at structurally complex loci. We also add 119 million base pairs of euchromatic polymorphic sequences and 1,115 gene duplications relative to the existing reference GRCh38. Roughly 90 million of the additional base pairs are derived from structural variation. Using our draft pangenome to analyse short-read data reduced small variant discovery errors by 34\% and increased the number of structural variants detected per haplotype by 104\% compared with GRCh38-based workflows, which enabled the typing of the vast majority of structural variant alleles per sample.
@Article{37165242, author = {Liao WW and Asri M and Ebler J and Doerr D and Haukness M and Hickey G and Lu S and Lucas JK and Monlong J and Abel HJ and Buonaiuto S and Chang XH and Cheng H and Chu J and Colonna V and Eizenga JM and Feng X and Fischer C and Fulton RS and Garg S and Groza C and Guarracino A and Harvey WT and Heumos S and Howe K and Jain M and Lu TY and Markello C and Martin FJ and Mitchell MW and Munson KM and Mwaniki MN and Novak AM and Olsen HE and Pesout T and Porubsky D and Prins P and Sibbesen JA and Sirén J and Tomlinson C and Villani F and Vollger MR and Antonacci-Fulton LL and Baid G and Baker CA and Belyaeva A and Billis K and Carroll A and Chang PC and Cody S and Cook DE and Cook-Deegan RM and Cornejo OE and Diekhans M and Ebert P and Fairley S and Fedrigo O and Felsenfeld AL and Formenti G and Frankish A and Gao Y and Garrison NA and Giron CG and Green RE and Haggerty L and Hoekzema K and Hourlier T and Ji HP and Kenny EE and Koenig BA and Kolesnikov A and Korbel JO and Kordosky J and Koren S and Lee H and Lewis AP and Magalhães H and Marco-Sola S and Marijon P and McCartney A and McDaniel J and Mountcastle J and Nattestad M and Nurk S and Olson ND and Popejoy AB and Puiu D and Rautiainen M and Regier AA and Rhie A and Sacco S and Sanders AD and Schneider VA and Schultz BI and Shafin K and Smith MW and Sofia HJ and Abou Tayoun AN and Thibaud-Nissen F and Tricomi FF and Wagner J and Walenz B and Wood JMD and Zimin AV and Bourque G and Chaisson MJP and Flicek P and Phillippy AM and Zook JM and Eichler EE and Haussler D and Wang T and Jarvis ED and Miga KH and Garrison E and Marschall T and Hall IM and Li H and Paten B}, title = {A draft human pangenome reference}, journal = {Nature}, volume = {617}, number = {7960}, pages = {312--324}, year = {2023}, doi = {10.1038/s41586-023-05896-x}, howpublished = {Advanced online publication: 10 May 2023}, note = {First posted as a preprint: 9 July 2002}, abstract = {Here the Human Pangenome Reference Consortium presents a first draft of the human pangenome reference. The pangenome contains 47 phased, diploid assemblies from a cohort of genetically diverse individuals\textsuperscript{1}. These assemblies cover more than 99\% of the expected sequence in each genome and are more than 99\% accurate at the structural and base pair levels. Based on alignments of the assemblies, we generate a draft pangenome that captures known variants and haplotypes and reveals new alleles at structurally complex loci. We also add 119 million base pairs of euchromatic polymorphic sequences and 1,115 gene duplications relative to the existing reference GRCh38. Roughly 90 million of the additional base pairs are derived from structural variation. Using our draft pangenome to analyse short-read data reduced small variant discovery errors by 34\% and increased the number of structural variants detected per haplotype by 104\% compared with GRCh38-based workflows, which enabled the typing of the vast majority of structural variant alleles per sample.},} - J Argentin, D Bolser, PJ Kersey, P Flicek. Comparative analysis of repeat content in plant genomes, large and small. Front Plant Sci 2023;14:1103035. doi:10.3389/fpls.2023.1103035
[BibTeX] [Abstract]
The DNA Features pipeline is the analysis pipeline at EMBL-EBI that annotates repeat elements, including transposable elements. With Ensembl’s goal to stay at the cutting edge of genome annotation, we proved that this pipeline needed an update. We then created a new analysis that allowed the Ensembl database to store the repeat classification from the PGSB repeat classification (Recat). This new dataset was then fetched using Perl scripts and used to prove that the pipeline modification induced a gain in sensitivity. Finally, we performed a comparative analysis of transposable element distribution in all plant species available, raising new questions about transposable elements in certain branches of the taxonomic tree.
@Article{37521909, author = {Argentin J and Bolser D and Kersey PJ and Flicek P}, title = {Comparative analysis of repeat content in plant genomes, large and small}, journal = {Front Plant Sci}, volume = {14}, pages = {1103035}, year = {2023}, doi = {10.3389/fpls.2023.1103035}, abstract = {The DNA Features pipeline is the analysis pipeline at EMBL-EBI that annotates repeat elements, including transposable elements. With Ensembl's goal to stay at the cutting edge of genome annotation, we proved that this pipeline needed an update. We then created a new analysis that allowed the Ensembl database to store the repeat classification from the PGSB repeat classification (Recat). This new dataset was then fetched using Perl scripts and used to prove that the pipeline modification induced a gain in sensitivity. Finally, we performed a comparative analysis of transposable element distribution in all plant species available, raising new questions about transposable elements in certain branches of the taxonomic tree.},} - FJ Martin, MR Amode, A Aneja, O Austine-Orimoloye, AG Azov, I Barnes, A Becker, R Bennett, A Berry, J Bhai, SK Bhurji, A Bignell, S Boddu, PR Branco Lins, L Brooks, SB Ramaraju, M Charkhchi, A Cockburn, L Da Rin Fiorretto, C Davidson, K Dodiya, S Donaldson, B El Houdaigui, T El Naboulsi, R Fatima, CG Giron, T Genez, GS Ghattaoraya, JG Martinez, C Guijarro, M Hardy, Z Hollis, T Hourlier, T Hunt, M Kay, V Kaykala, T Le, D Lemos, D Marques-Coelho, JC Marugán, GA Merino, LP Mirabueno, A Mushtaq, SN Hossain, DN Ogeh, MP Sakthivel, A Parker, M Perry, I Piližota, I Prosovetskaia, JG Pérez-Silva, AIA Salam, N Saraiva-Agostinho, H Schuilenburg, D Sheppard, S Sinha, B Sipos, W Stark, E Steed, R Sukumaran, D Sumathipala, MM Suner, L Surapaneni, K Sutinen, M Szpak, FF Tricomi, D Urbina-Gómez, A Veidenberg, TA Walsh, B Walts, E Wass, N Willhoft, J Allen, J Alvarez-Jarreta, M Chakiachvili, B Flint, S Giorgetti, L Haggerty, GR Ilsley, JE Loveland, B Moore, JM Mudge, J Tate, D Thybert, SJ Trevanion, A Winterbottom, A Frankish, SE Hunt, M Ruffier, F Cunningham, S Dyer, RD Finn, KL Howe, PW Harrison, AD Yates, P Flicek. Ensembl 2023. Nucleic Acids Res 2023;51(D1):D933–D941. doi:10.1093/nar/gkac958
[BibTeX] [Abstract]
Ensembl (https://www.ensembl.org) has produced high-quality genomic resources for vertebrates and model organisms for more than twenty years. During that time, our resources, services and tools have continually evolved in line with both the publicly available genome data and the downstream research and applications that utilise the Ensembl platform. In recent years we have witnessed a dramatic shift in the genomic landscape. There has been a large increase in the number of high-quality reference genomes through global biodiversity initiatives. In parallel, there have been major advances towards pangenome representations of higher species, where many alternative genome assemblies representing different breeds, cultivars, strains and haplotypes are now available. In order to support these efforts and accelerate downstream research, it is our goal at Ensembl to create high-quality annotations, tools and services for species across the tree of life. Here, we report our resources for popular reference genomes, the dramatic growth of our annotations (including haplotypes from the first human pangenome graphs), updates to the Ensembl Variant Effect Predictor (VEP), interactive protein structure predictions from AlphaFold DB, and the beta release of our new website.
@Article{36318249, author = {Martin FJ and Amode MR and Aneja A and Austine-Orimoloye O and Azov AG and Barnes I and Becker A and Bennett R and Berry A and Bhai J and Bhurji SK and Bignell A and Boddu S and Branco Lins PR and Brooks L and Ramaraju SB and Charkhchi M and Cockburn A and Da Rin Fiorretto L and Davidson C and Dodiya K and Donaldson S and El Houdaigui B and El Naboulsi T and Fatima R and Giron CG and Genez T and Ghattaoraya GS and Martinez JG and Guijarro C and Hardy M and Hollis Z and Hourlier T and Hunt T and Kay M and Kaykala V and Le T and Lemos D and Marques-Coelho D and Marugán JC and Merino GA and Mirabueno LP and Mushtaq A and Hossain SN and Ogeh DN and Sakthivel MP and Parker A and Perry M and Piližota I and Prosovetskaia I and Pérez-Silva JG and Salam AIA and Saraiva-Agostinho N and Schuilenburg H and Sheppard D and Sinha S and Sipos B and Stark W and Steed E and Sukumaran R and Sumathipala D and Suner MM and Surapaneni L and Sutinen K and Szpak M and Tricomi FF and Urbina-Gómez D and Veidenberg A and Walsh TA and Walts B and Wass E and Willhoft N and Allen J and Alvarez-Jarreta J and Chakiachvili M and Flint B and Giorgetti S and Haggerty L and Ilsley GR and Loveland JE and Moore B and Mudge JM and Tate J and Thybert D and Trevanion SJ and Winterbottom A and Frankish A and Hunt SE and Ruffier M and Cunningham F and Dyer S and Finn RD and Howe KL and Harrison PW and Yates AD and Flicek P}, title = {Ensembl 2023}, journal = {Nucleic Acids Res}, volume = {51}, number = {D1}, pages = {D933--D941}, year = {2023}, doi = {10.1093/nar/gkac958}, abstract = {Ensembl (https://www.ensembl.org) has produced high-quality genomic resources for vertebrates and model organisms for more than twenty years. During that time, our resources, services and tools have continually evolved in line with both the publicly available genome data and the downstream research and applications that utilise the Ensembl platform. In recent years we have witnessed a dramatic shift in the genomic landscape. There has been a large increase in the number of high-quality reference genomes through global biodiversity initiatives. In parallel, there have been major advances towards pangenome representations of higher species, where many alternative genome assemblies representing different breeds, cultivars, strains and haplotypes are now available. In order to support these efforts and accelerate downstream research, it is our goal at Ensembl to create high-quality annotations, tools and services for species across the tree of life. Here, we report our resources for popular reference genomes, the dramatic growth of our annotations (including haplotypes from the first human pangenome graphs), updates to the Ensembl Variant Effect Predictor (VEP), interactive protein structure predictions from AlphaFold DB, and the beta release of our new website.},} - A Frankish, S Carbonell-Sala, M Diekhans, I Jungreis, JE Loveland, JM Mudge, C Sisu, JC Wright, C Arnan, I Barnes, A Banerjee, R Bennett, A Berry, A Bignell, C Boix, F Calvet, D Cerdán-Vélez, F Cunningham, C Davidson, S Donaldson, C Dursun, R Fatima, S Giorgetti, CG Giron, JM Gonzalez, M Hardy, PW Harrison, T Hourlier, Z Hollis, T Hunt, B James, Y Jiang, R Johnson, M Kay, J Lagarde, FJ Martin, LM Gómez, S Nair, P Ni, F Pozo, V Ramalingam, M Ruffier, BM Schmitt, JM Schreiber, E Steed, MM Suner, D Sumathipala, I Sycheva, B Uszczynska-Ratajczak, E Wass, YT Yang, A Yates, Z Zafrulla, JS Choudhary, M Gerstein, R Guigo, TJP Hubbard, M Kellis, A Kundaje, B Paten, ML Tress, P Flicek. GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res 2023;51(D1):D942–D949. doi:10.1093/nar/gkac1071
[BibTeX] [Abstract]
GENCODE produces high quality gene and transcript annotation for the human and mouse genomes. All GENCODE annotation is supported by experimental data and serves as a reference for genome biology and clinical genomics. The GENCODE consortium generates targeted experimental data, develops bioinformatic tools and carries out analyses that, along with externally produced data and methods, support the identification and annotation of transcript structures and the determination of their function. Here, we present an update on the annotation of human and mouse genes, including developments in the tools, data, analyses and major collaborations which underpin this progress. For example, we report the creation of a set of non-canonical ORFs identified in GENCODE transcripts, the LRGASP collaboration to assess the use of long transcriptomic data to build transcript models, the progress in collaborations with RefSeq and UniProt to increase convergence in the annotation of human and mouse protein-coding genes, the propagation of GENCODE across the human pan-genome and the development of new tools to support annotation of regulatory features by GENCODE. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.
@Article{36420896, author = {Frankish A and Carbonell-Sala S and Diekhans M and Jungreis I and Loveland JE and Mudge JM and Sisu C and Wright JC and Arnan C and Barnes I and Banerjee A and Bennett R and Berry A and Bignell A and Boix C and Calvet F and Cerdán-Vélez D and Cunningham F and Davidson C and Donaldson S and Dursun C and Fatima R and Giorgetti S and Giron CG and Gonzalez JM and Hardy M and Harrison PW and Hourlier T and Hollis Z and Hunt T and James B and Jiang Y and Johnson R and Kay M and Lagarde J and Martin FJ and Gómez LM and Nair S and Ni P and Pozo F and Ramalingam V and Ruffier M and Schmitt BM and Schreiber JM and Steed E and Suner MM and Sumathipala D and Sycheva I and Uszczynska-Ratajczak B and Wass E and Yang YT and Yates A and Zafrulla Z and Choudhary JS and Gerstein M and Guigo R and Hubbard TJP and Kellis M and Kundaje A and Paten B and Tress ML and Flicek P}, title = {GENCODE: reference annotation for the human and mouse genomes in 2023}, journal = {Nucleic Acids Res}, volume = {51}, number = {D1}, pages = {D942--D949}, year = {2023}, doi = {10.1093/nar/gkac1071}, abstract = {GENCODE produces high quality gene and transcript annotation for the human and mouse genomes. All GENCODE annotation is supported by experimental data and serves as a reference for genome biology and clinical genomics. The GENCODE consortium generates targeted experimental data, develops bioinformatic tools and carries out analyses that, along with externally produced data and methods, support the identification and annotation of transcript structures and the determination of their function. Here, we present an update on the annotation of human and mouse genes, including developments in the tools, data, analyses and major collaborations which underpin this progress. For example, we report the creation of a set of non-canonical ORFs identified in GENCODE transcripts, the LRGASP collaboration to assess the use of long transcriptomic data to build transcript models, the progress in collaborations with RefSeq and UniProt to increase convergence in the annotation of human and mouse protein-coding genes, the propagation of GENCODE across the human pan-genome and the development of new tools to support annotation of regulatory features by GENCODE. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.},} - B Contreras-Moreira, S Saraf, G Naamati, AM Casas, SS Amberkar, P Flicek, AR Jones, S Dyer. GET_PANGENES: calling pangenes from plant genome alignments confirms presence-absence variation. Genome Biol 2023;24(1):223. doi:10.1186/s13059-023-03071-z
[BibTeX] [Abstract]
Crop pangenomes made from individual cultivar assemblies promise easy access to conserved genes, but genome content variability and inconsistent identifiers hamper their exploration. To address this, we define pangenes, which summarize a species coding potential and link back to original annotations. The protocol get_pangenes performs whole genome alignments (WGA) to call syntenic gene models based on coordinate overlaps. A benchmark with small and large plant genomes shows that pangenes recapitulate phylogeny-based orthologies and produce complete soft-core gene sets. Moreover, WGAs support lift-over and help confirm gene presence-absence variation. Source code and documentation: https://github.com/Ensembl/plant-scripts .
@Article{37798615, author = {Contreras-Moreira B and Saraf S and Naamati G and Casas AM and Amberkar SS and Flicek P and Jones AR and Dyer S}, title = {GET_PANGENES: calling pangenes from plant genome alignments confirms presence-absence variation}, journal = {Genome Biol}, volume = {24}, number = {1}, pages = {223}, year = {2023}, doi = {10.1186/s13059-023-03071-z}, note = {First posted as a preprint: 3 January 2023}, abstract = {Crop pangenomes made from individual cultivar assemblies promise easy access to conserved genes, but genome content variability and inconsistent identifiers hamper their exploration. To address this, we define pangenes, which summarize a species coding potential and link back to original annotations. The protocol get_pangenes performs whole genome alignments (WGA) to call syntenic gene models based on coordinate overlaps. A benchmark with small and large plant genomes shows that pangenes recapitulate phylogeny-based orthologies and produce complete soft-core gene sets. Moreover, WGAs support lift-over and help confirm gene presence-absence variation. Source code and documentation: https://github.com/Ensembl/plant-scripts .},} - M Vromman, J Anckaert, S Bortoluzzi, A Buratin, CY Chen, Q Chu, TJ Chuang, R Dehghannasiri, C Dieterich, X Dong, P Flicek, E Gaffo, W Gu, C He, S Hoffmann, O Izuogu, MS Jackson, T Jakobi, EC Lai, J Nuytens, J Salzman, M Santibanez-Koref, P Stadler, O Thas, E Vanden Eynde, K Verniers, G Wen, J Westholm, L Yang, CY Ye, N Yigit, GH Yuan, J Zhang, F Zhao, J Vandesompele, PJ Volders. Large-scale benchmarking of circRNA detection tools reveals large differences in sensitivity but not in precision. Nat Methods 2023;20(8):1159–1169. doi:10.1038/s41592-023-01944-6
[BibTeX] [Abstract]
The detection of circular RNA molecules (circRNAs) is typically based on short-read RNA sequencing data processed using computational tools. Numerous such tools have been developed, but a systematic comparison with orthogonal validation is missing. Here, we set up a circRNA detection tool benchmarking study, in which 16 tools detected more than 315,000 unique circRNAs in three deeply sequenced human cell types. Next, 1,516 predicted circRNAs were validated using three orthogonal methods. Generally, tool-specific precision is high and similar (median of 98.8\%, 96.3\% and 95.5\% for qPCR, RNase R and amplicon sequencing, respectively) whereas the sensitivity and number of predicted circRNAs (ranging from 1,372 to 58,032) are the most significant differentiators. Of note, precision values are lower when evaluating low-abundance circRNAs. We also show that the tools can be used complementarily to increase detection sensitivity. Finally, we offer recommendations for future circRNA detection and validation.
@Article{37443337, author = {Vromman M and Anckaert J and Bortoluzzi S and Buratin A and Chen CY and Chu Q and Chuang TJ and Dehghannasiri R and Dieterich C and Dong X and Flicek P and Gaffo E and Gu W and He C and Hoffmann S and Izuogu O and Jackson MS and Jakobi T and Lai EC and Nuytens J and Salzman J and Santibanez-Koref M and Stadler P and Thas O and Vanden Eynde E and Verniers K and Wen G and Westholm J and Yang L and Ye CY and Yigit N and Yuan GH and Zhang J and Zhao F and Vandesompele J and Volders PJ}, title = {Large-scale benchmarking of circRNA detection tools reveals large differences in sensitivity but not in precision}, journal = {Nat Methods}, volume = {20}, number = {8}, pages = {1159--1169}, year = {2023}, doi = {10.1038/s41592-023-01944-6}, note = {First posted as a preprint: 6 December 2002}, abstract = {The detection of circular RNA molecules (circRNAs) is typically based on short-read RNA sequencing data processed using computational tools. Numerous such tools have been developed, but a systematic comparison with orthogonal validation is missing. Here, we set up a circRNA detection tool benchmarking study, in which 16 tools detected more than 315,000 unique circRNAs in three deeply sequenced human cell types. Next, 1,516 predicted circRNAs were validated using three orthogonal methods. Generally, tool-specific precision is high and similar (median of 98.8\%, 96.3\% and 95.5\% for qPCR, RNase R and amplicon sequencing, respectively) whereas the sensitivity and number of predicted circRNAs (ranging from 1,372 to 58,032) are the most significant differentiators. Of note, precision values are lower when evaluating low-abundance circRNAs. We also show that the tools can be used complementarily to increase detection sensitivity. Finally, we offer recommendations for future circRNA detection and validation.},} - B Benjelloun, K Leempoel, F Boyer, S Stucki, I Streeter, P Orozco-terWengel, FJ Alberto, B Servin, F Biscarini, A Alberti, S Engelen, A Stella, L Colli, E Coissac, MW Bruford, P Ajmone-Marsan, R Negrini, L Clarke, P Flicek, A Chikhi, S Joost, P Taberlet, F Pompanon. Multiple genomic solutions for local adaptation in two closely related species (sheep and goats) facing the same climatic constraints. Mol Ecol 2023:e17257. doi:10.1111/mec.17257
[BibTeX] [Abstract]
The question of how local adaptation takes place remains a fundamental question in evolutionary biology. The variation of allele frequencies in genes under selection over environmental gradients remains mainly theoretical and its empirical assessment would help understanding how adaptation happens over environmental clines. To bring new insights to this issue we set up a broad framework which aimed to compare the adaptive trajectories over environmental clines in two domesticated mammal species co-distributed in diversified landscapes. We sequenced the genomes of 160 sheep and 161 goats extensively managed along environmental gradients, including temperature, rainfall, seasonality and altitude, to identify genes and biological processes shaping local adaptation. Allele frequencies at putatively adaptive loci were rarely found to vary gradually along environmental gradients, but rather displayed a discontinuous shift at the extremities of environmental clines. Of the 430 candidate adaptive genes identified, only 6 were orthologous between sheep and goats and those responded differently to environmental pressures, suggesting different putative mechanisms involved in local adaptation in these two closely related species. Interestingly, the genomes of the 2 species were impacted differently by the environment, genes related to signatures of selection were most related to altitude, slope and rainfall seasonality for sheep, and summer temperature and spring rainfall for goats. The diversity of candidate adaptive pathways may result from a high number of biological functions involved in the adaptations to multiple eco-climatic gradients, and a differential role of climatic drivers on the two species, despite their co-distribution along the same environmental gradients. This study describes empirical examples of clinal variation in putatively adaptive alleles with different patterns in allele frequency distributions over continuous environmental gradients, thus showing the diversity of genetic responses in adaptive landscapes and opening new horizons for understanding genomics of adaptation in mammalian species and beyond.
@Article{38149334, author = {Benjelloun B and Leempoel K and Boyer F and Stucki S and Streeter I and Orozco-terWengel P and Alberto FJ and Servin B and Biscarini F and Alberti A and Engelen S and Stella A and Colli L and Coissac E and Bruford MW and Ajmone-Marsan P and Negrini R and Clarke L and Flicek P and Chikhi A and Joost S and Taberlet P and Pompanon F}, title = {Multiple genomic solutions for local adaptation in two closely related species (sheep and goats) facing the same climatic constraints}, journal = {Mol Ecol}, pages = {e17257}, year = {2023}, doi = {10.1111/mec.17257}, note = {First posted as a preprint: 19 December 2021}, abstract = {The question of how local adaptation takes place remains a fundamental question in evolutionary biology. The variation of allele frequencies in genes under selection over environmental gradients remains mainly theoretical and its empirical assessment would help understanding how adaptation happens over environmental clines. To bring new insights to this issue we set up a broad framework which aimed to compare the adaptive trajectories over environmental clines in two domesticated mammal species co-distributed in diversified landscapes. We sequenced the genomes of 160 sheep and 161 goats extensively managed along environmental gradients, including temperature, rainfall, seasonality and altitude, to identify genes and biological processes shaping local adaptation. Allele frequencies at putatively adaptive loci were rarely found to vary gradually along environmental gradients, but rather displayed a discontinuous shift at the extremities of environmental clines. Of the 430 candidate adaptive genes identified, only 6 were orthologous between sheep and goats and those responded differently to environmental pressures, suggesting different putative mechanisms involved in local adaptation in these two closely related species. Interestingly, the genomes of the 2 species were impacted differently by the environment, genes related to signatures of selection were most related to altitude, slope and rainfall seasonality for sheep, and summer temperature and spring rainfall for goats. The diversity of candidate adaptive pathways may result from a high number of biological functions involved in the adaptations to multiple eco-climatic gradients, and a differential role of climatic drivers on the two species, despite their co-distribution along the same environmental gradients. This study describes empirical examples of clinal variation in putatively adaptive alleles with different patterns in allele frequency distributions over continuous environmental gradients, thus showing the diversity of genetic responses in adaptive landscapes and opening new horizons for understanding genomics of adaptation in mammalian species and beyond.},} - A Rhie, S Nurk, M Cechova, SJ Hoyt, DJ Taylor, N Altemose, PW Hook, S Koren, M Rautiainen, IA Alexandrov, J Allen, M Asri, AV Bzikadze, NC Chen, CS Chin, M Diekhans, P Flicek, G Formenti, A Fungtammasan, C Garcia Giron, E Garrison, A Gershman, JL Gerton, PGS Grady, A Guarracino, L Haggerty, R Halabian, NF Hansen, R Harris, GA Hartley, WT Harvey, M Haukness, J Heinz, T Hourlier, RM Hubley, SE Hunt, S Hwang, M Jain, RK Kesharwani, AP Lewis, H Li, GA Logsdon, JK Lucas, W Makalowski, C Markovic, FJ Martin, AM Mc Cartney, RC McCoy, J McDaniel, BM McNulty, P Medvedev, A Mikheenko, KM Munson, TD Murphy, HE Olsen, ND Olson, LF Paulin, D Porubsky, T Potapova, F Ryabov, SL Salzberg, MEG Sauria, FJ Sedlazeck, K Shafin, VA Shepelev, A Shumate, JM Storer, L Surapaneni, AM Taravella Oill, F Thibaud-Nissen, W Timp, M Tomaszkiewicz, MR Vollger, BP Walenz, AC Watwood, MH Weissensteiner, AM Wenger, MA Wilson, S Zarate, Y Zhu, JM Zook, EE Eichler, RJ O’Neill, MC Schatz, KH Miga, KD Makova, AM Phillippy. The complete sequence of a human Y chromosome. Nature 2023;621(7978):344–354. doi:10.1038/s41586-023-06457-y
[BibTeX] [Abstract]
The human Y chromosome has been notoriously difficult to sequence and assemble because of its complex repeat structure that includes long palindromes, tandem repeats and segmental duplications\textsuperscript{1-3}. As a result, more than half of the Y chromosome is missing from the GRCh38 reference sequence and it remains the last human chromosome to be finished\textsuperscript{4,5}. Here, the Telomere-to-Telomere (T2T) consortium presents the complete 62,460,029-base-pair sequence of a human Y chromosome from the HG002 genome (T2T-Y) that corrects multiple errors in GRCh38-Y and adds over 30 million base pairs of sequence to the reference, showing the complete ampliconic structures of gene families TSPY, DAZ and RBMY; 41 additional protein-coding genes, mostly from the TSPY family; and an alternating pattern of human satellite 1 and 3 blocks in the heterochromatic Yq12 region. We have combined T2T-Y with a previous assembly of the CHM13 genome\textsuperscript{4} and mapped available population variation, clinical variants and functional genomics data to produce a complete and comprehensive reference sequence for all 24 human chromosomes.
@Article{37612512, author = {Rhie A and Nurk S and Cechova M and Hoyt SJ and Taylor DJ and Altemose N and Hook PW and Koren S and Rautiainen M and Alexandrov IA and Allen J and Asri M and Bzikadze AV and Chen NC and Chin CS and Diekhans M and Flicek P and Formenti G and Fungtammasan A and Garcia Giron C and Garrison E and Gershman A and Gerton JL and Grady PGS and Guarracino A and Haggerty L and Halabian R and Hansen NF and Harris R and Hartley GA and Harvey WT and Haukness M and Heinz J and Hourlier T and Hubley RM and Hunt SE and Hwang S and Jain M and Kesharwani RK and Lewis AP and Li H and Logsdon GA and Lucas JK and Makalowski W and Markovic C and Martin FJ and Mc Cartney AM and McCoy RC and McDaniel J and McNulty BM and Medvedev P and Mikheenko A and Munson KM and Murphy TD and Olsen HE and Olson ND and Paulin LF and Porubsky D and Potapova T and Ryabov F and Salzberg SL and Sauria MEG and Sedlazeck FJ and Shafin K and Shepelev VA and Shumate A and Storer JM and Surapaneni L and Taravella Oill AM and Thibaud-Nissen F and Timp W and Tomaszkiewicz M and Vollger MR and Walenz BP and Watwood AC and Weissensteiner MH and Wenger AM and Wilson MA and Zarate S and Zhu Y and Zook JM and Eichler EE and O'Neill RJ and Schatz MC and Miga KH and Makova KD and Phillippy AM}, title = {The complete sequence of a human Y chromosome}, journal = {Nature}, volume = {621}, number = {7978}, pages = {344--354}, year = {2023}, doi = {10.1038/s41586-023-06457-y}, howpublished = {Advanced online publication: 23 August 2023}, note = {First posted as a preprint: 1 December 2002}, abstract = {The human Y chromosome has been notoriously difficult to sequence and assemble because of its complex repeat structure that includes long palindromes, tandem repeats and segmental duplications\textsuperscript{1-3}. As a result, more than half of the Y chromosome is missing from the GRCh38 reference sequence and it remains the last human chromosome to be finished\textsuperscript{4,5}. Here, the Telomere-to-Telomere (T2T) consortium presents the complete 62,460,029-base-pair sequence of a human Y chromosome from the HG002 genome (T2T-Y) that corrects multiple errors in GRCh38-Y and adds over 30 million base pairs of sequence to the reference, showing the complete ampliconic structures of gene families TSPY, DAZ and RBMY; 41 additional protein-coding genes, mostly from the TSPY family; and an alternating pattern of human satellite 1 and 3 blocks in the heterochromatic Yq12 region. We have combined T2T-Y with a previous assembly of the CHM13 genome\textsuperscript{4} and mapped available population variation, clinical variants and functional genomics data to produce a complete and comprehensive reference sequence for all 24 human chromosomes.},} - DJ Barker, G Maccari, X Georgiou, MA Cooper, P Flicek, J Robinson, SGE Marsh. The IPD-IMGT/HLA Database. Nucleic Acids Res 2023;51(D1):D1053–D1060. doi:10.1093/nar/gkac1011
[BibTeX] [Abstract]
It is 24 years since the IPD-IMGT/HLA Database, http://www.ebi.ac.uk/ipd/imgt/hla/, was first released, providing the HLA community with a searchable repository of highly curated HLA sequences. The database now contains over 35 000 alleles of the human Major Histocompatibility Complex (MHC) named by the WHO Nomenclature Committee for Factors of the HLA System. This complex contains the most polymorphic genes in the human genome and is now considered hyperpolymorphic. The IPD-IMGT/HLA Database provides a stable and user-friendly repository for this information. Uptake of Next Generation Sequencing technology in recent years has driven an increase in the number of alleles and the length of sequences submitted. As the size of the database has grown the traditional methods of accessing and presenting this data have been challenged, in response, we have developed a suite of tools providing an enhanced user experience to our traditional web-based users while creating new programmatic access for our bioinformatics user base. This suite of tools is powered by the IPD-API, an Application Programming Interface (API), providing scalable and flexible access to the database. The IPD-API provides a stable platform for our future development allowing us to meet the future challenges of the HLA field and needs of the community.
@Article{36350643, author = {Barker DJ and Maccari G and Georgiou X and Cooper MA and Flicek P and Robinson J and Marsh SGE}, title = {The IPD-IMGT/HLA Database}, journal = {Nucleic Acids Res}, volume = {51}, number = {D1}, pages = {D1053--D1060}, year = {2023}, doi = {10.1093/nar/gkac1011}, abstract = {It is 24 years since the IPD-IMGT/HLA Database, http://www.ebi.ac.uk/ipd/imgt/hla/, was first released, providing the HLA community with a searchable repository of highly curated HLA sequences. The database now contains over 35 000 alleles of the human Major Histocompatibility Complex (MHC) named by the WHO Nomenclature Committee for Factors of the HLA System. This complex contains the most polymorphic genes in the human genome and is now considered hyperpolymorphic. The IPD-IMGT/HLA Database provides a stable and user-friendly repository for this information. Uptake of Next Generation Sequencing technology in recent years has driven an increase in the number of alleles and the length of sequences submitted. As the size of the database has grown the traditional methods of accessing and presenting this data have been challenged, in response, we have developed a suite of tools providing an enhanced user experience to our traditional web-based users while creating new programmatic access for our bioinformatics user base. This suite of tools is powered by the IPD-API, an Application Programming Interface (API), providing scalable and flexible access to the database. The IPD-API provides a stable platform for our future development allowing us to meet the future challenges of the HLA field and needs of the community.},}
2022
- J Morales, S Pujar, JE Loveland, A Astashyn, R Bennett, A Berry, E Cox, C Davidson, O Ermolaeva, CM Farrell, R Fatima, L Gil, T Goldfarb, JM Gonzalez, D Haddad, M Hardy, T Hunt, J Jackson, VS Joardar, M Kay, VK Kodali, KM McGarvey, A McMahon, JM Mudge, DN Murphy, MR Murphy, B Rajput, SH Rangwala, LD Riddick, F Thibaud-Nissen, G Threadgold, AR Vatsan, C Wallin, D Webb, P Flicek, E Birney, KD Pruitt, A Frankish, F Cunningham, TD Murphy. A joint NCBI and EMBL-EBI transcript set for clinical genomics and research. Nature 2022;604(7905):310–315. doi:10.1038/s41586-022-04558-8
[BibTeX] [Abstract]
Comprehensive genome annotation is essential to understand the impact of clinically relevant variants. However, the absence of a standard for clinical reporting and browser display complicates the process of consistent interpretation and reporting. To address these challenges, Ensembl/GENCODE\textsuperscript{1} and RefSeq\textsuperscript{2} launched a joint initiative, the Matched Annotation from NCBI and EMBL-EBI (MANE) collaboration, to converge on human gene and transcript annotation and to jointly define a high-value set of transcripts and corresponding proteins. Here, we describe the MANE transcript sets for use as universal standards for variant reporting and browser display. The MANE Select set identifies a representative transcript for each human protein-coding gene, whereas the MANE Plus Clinical set provides additional transcripts at loci where the Select transcripts alone are not sufficient to report all currently known clinical variants. Each MANE transcript represents an exact match between the exonic sequences of an Ensembl/GENCODE transcript and its counterpart in RefSeq such that the identifiers can be used synonymously. We have now released MANE Select transcripts for 97\% of human protein-coding genes, including all American College of Medical Genetics and Genomics Secondary Findings list v3.0 (ref. \textsuperscript{3}) genes. MANE transcripts are accessible from major genome browsers and key resources. Widespread adoption of these transcript sets will increase the consistency of reporting, facilitate the exchange of data regardless of the annotation source and help to streamline clinical interpretation.
@Article{35388217, author = {Morales J and Pujar S and Loveland JE and Astashyn A and Bennett R and Berry A and Cox E and Davidson C and Ermolaeva O and Farrell CM and Fatima R and Gil L and Goldfarb T and Gonzalez JM and Haddad D and Hardy M and Hunt T and Jackson J and Joardar VS and Kay M and Kodali VK and McGarvey KM and McMahon A and Mudge JM and Murphy DN and Murphy MR and Rajput B and Rangwala SH and Riddick LD and Thibaud-Nissen F and Threadgold G and Vatsan AR and Wallin C and Webb D and Flicek P and Birney E and Pruitt KD and Frankish A and Cunningham F and Murphy TD}, title = {A joint NCBI and EMBL-EBI transcript set for clinical genomics and research}, journal = {Nature}, volume = {604}, number = {7905}, pages = {310--315}, year = {2022}, doi = {10.1038/s41586-022-04558-8}, howpublished = {Advanced online publication: 6 April 2022}, abstract = {Comprehensive genome annotation is essential to understand the impact of clinically relevant variants. However, the absence of a standard for clinical reporting and browser display complicates the process of consistent interpretation and reporting. To address these challenges, Ensembl/GENCODE\textsuperscript{1} and RefSeq\textsuperscript{2} launched a joint initiative, the Matched Annotation from NCBI and EMBL-EBI (MANE) collaboration, to converge on human gene and transcript annotation and to jointly define a high-value set of transcripts and corresponding proteins. Here, we describe the MANE transcript sets for use as universal standards for variant reporting and browser display. The MANE Select set identifies a representative transcript for each human protein-coding gene, whereas the MANE Plus Clinical set provides additional transcripts at loci where the Select transcripts alone are not sufficient to report all currently known clinical variants. Each MANE transcript represents an exact match between the exonic sequences of an Ensembl/GENCODE transcript and its counterpart in RefSeq such that the identifiers can be used synonymously. We have now released MANE Select transcripts for 97\% of human protein-coding genes, including all American College of Medical Genetics and Genomics Secondary Findings list v3.0 (ref. \textsuperscript{3}) genes. MANE transcripts are accessible from major genome browsers and key resources. Widespread adoption of these transcript sets will increase the consistency of reporting, facilitate the exchange of data regardless of the annotation source and help to streamline clinical interpretation.},} - K Higgins, BA Moore, Z Berberovic, HA Adissu, M Eskandarian, AM Flenniken, A Shao, DM Imai, D Clary, L Lanoue, S Newbigging, LMJ Nutter, DJ Adams, F Bosch, RE Braun, SDM Brown, ME Dickinson, M Dobbie, P Flicek, X Gao, S Galande, A Grobler, JD Heaney, Y Herault, de MH Angelis, HG Chin, F Mammano, C Qin, T Shiroishi, R Sedlacek, JK Seong, Y Xu, C IMPC, KCK Lloyd, C McKerlie, A Moshiri. Analysis of genome-wide knockout mouse database identifies candidate ciliopathy genes. Sci Rep 2022;12(1):20791. doi:10.1038/s41598-022-19710-7
[BibTeX] [Abstract]
We searched a database of single-gene knockout (KO) mice produced by the International Mouse Phenotyping Consortium (IMPC) to identify candidate ciliopathy genes. We first screened for phenotypes in mouse lines with both ocular and renal or reproductive trait abnormalities. The STRING protein interaction tool was used to identify interactions between known cilia gene products and those encoded by the genes in individual knockout mouse strains in order to generate a list of From this list, 32 genes encoded proteins predicted to interact with known ciliopathy proteins. Of these, 25 had no previously described roles in ciliary pathobiology. Histological and morphological evidence of phenotypes found in ciliopathies in knockout mouse lines are presented as examples (genes Abi2, Wdr62, Ap4e1, Dync1li1, and Prkab1). Phenotyping data and descriptions generated on IMPC mouse line are useful for mechanistic studies, target discovery, rare disease diagnosis, and preclinical therapeutic development trials. Here we demonstrate the effective use of the IMPC phenotype data to uncover genes with no previous role in ciliary biology, which may be clinically relevant for identification of novel disease genes implicated in ciliopathies.
@Article{36456625, author = {Higgins K and Moore BA and Berberovic Z and Adissu HA and Eskandarian M and Flenniken AM and Shao A and Imai DM and Clary D and Lanoue L and Newbigging S and Nutter LMJ and Adams DJ and Bosch F and Braun RE and Brown SDM and Dickinson ME and Dobbie M and Flicek P and Gao X and Galande S and Grobler A and Heaney JD and Herault Y and de Angelis MH and Chin HG and Mammano F and Qin C and Shiroishi T and Sedlacek R and Seong JK and Xu Y and IMPC C and Lloyd KCK and McKerlie C and Moshiri A}, title = {Analysis of genome-wide knockout mouse database identifies candidate ciliopathy genes}, journal = {Sci Rep}, volume = {12}, number = {1}, pages = {20791}, year = {2022}, doi = {10.1038/s41598-022-19710-7}, abstract = {We searched a database of single-gene knockout (KO) mice produced by the International Mouse Phenotyping Consortium (IMPC) to identify candidate ciliopathy genes. We first screened for phenotypes in mouse lines with both ocular and renal or reproductive trait abnormalities. The STRING protein interaction tool was used to identify interactions between known cilia gene products and those encoded by the genes in individual knockout mouse strains in order to generate a list of From this list, 32 genes encoded proteins predicted to interact with known ciliopathy proteins. Of these, 25 had no previously described roles in ciliary pathobiology. Histological and morphological evidence of phenotypes found in ciliopathies in knockout mouse lines are presented as examples (genes Abi2, Wdr62, Ap4e1, Dync1li1, and Prkab1). Phenotyping data and descriptions generated on IMPC mouse line are useful for mechanistic studies, target discovery, rare disease diagnosis, and preclinical therapeutic development trials. Here we demonstrate the effective use of the IMPC phenotype data to uncover genes with no previous role in ciliary biology, which may be clinically relevant for identification of novel disease genes implicated in ciliopathies.},} - SE Hunt, B Moore, RM Amode, IM Armean, D Lemos, A Mushtaq, A Parton, H Schuilenburg, M Szpak, A Thormann, E Perry, SJ Trevanion, P Flicek, AD Yates, F Cunningham. Annotating and prioritizing genomic variants using the Ensembl Variant Effect Predictor-A tutorial. Hum Mutat 2022;43(8):986–997. doi:10.1002/humu.24298
[BibTeX] [Abstract]
The Ensembl Variant Effect Predictor (VEP) is a freely available, open-source tool for the annotation and filtering of genomic variants. It predicts variant molecular consequences using the Ensembl/GENCODE or RefSeq gene sets. It also reports phenotype associations from databases such as ClinVar, allele frequencies from studies including gnomAD, and predictions of deleteriousness from tools such as Sorting Intolerant From Tolerant and Combined Annotation Dependent Depletion. Ensembl VEP includes filtering options to customize variant prioritization. It is well supported and updated roughly quarterly to incorporate the latest gene, variant, and phenotype association information. Ensembl VEP analysis can be performed using a highly configurable, extensible command-line tool, a Representational State Transfer application programming interface, and a user-friendly web interface. These access methods are designed to suit different levels of bioinformatics experience and meet different needs in terms of data size, visualization, and flexibility. In this tutorial, we will describe performing variant annotation using the Ensembl VEP web tool, which enables sophisticated analysis through a simple interface.
@Article{34816521, author = {Hunt SE and Moore B and Amode RM and Armean IM and Lemos D and Mushtaq A and Parton A and Schuilenburg H and Szpak M and Thormann A and Perry E and Trevanion SJ and Flicek P and Yates AD and Cunningham F}, title = {Annotating and prioritizing genomic variants using the Ensembl Variant Effect Predictor-A tutorial}, journal = {Hum Mutat}, volume = {43}, number = {8}, pages = {986--997}, year = {2022}, doi = {10.1002/humu.24298}, howpublished = {Advanced online publication: 24 November 2021}, abstract = {The Ensembl Variant Effect Predictor (VEP) is a freely available, open-source tool for the annotation and filtering of genomic variants. It predicts variant molecular consequences using the Ensembl/GENCODE or RefSeq gene sets. It also reports phenotype associations from databases such as ClinVar, allele frequencies from studies including gnomAD, and predictions of deleteriousness from tools such as Sorting Intolerant From Tolerant and Combined Annotation Dependent Depletion. Ensembl VEP includes filtering options to customize variant prioritization. It is well supported and updated roughly quarterly to incorporate the latest gene, variant, and phenotype association information. Ensembl VEP analysis can be performed using a highly configurable, extensible command-line tool, a Representational State Transfer application programming interface, and a user-friendly web interface. These access methods are designed to suit different levels of bioinformatics experience and meet different needs in terms of data size, visualization, and flexibility. In this tutorial, we will describe performing variant annotation using the Ensembl VEP web tool, which enables sophisticated analysis through a simple interface.},} - S Kongsstovu í, S-O Mikalsen, EÍ Homrum, JA Jacobsen, TD Als, H Gislason, P Flicek, EE Nielsen, HA Dahl. Atlantic herring (Clupea harengus) population structure in the Northeast Atlantic Ocean. Fisheries Research 2022;249:106231. doi:10.1016/j.fishres.2022.106231
[BibTeX] [Abstract]
The Atlantic herring Clupea harengus L has a vast geographical distribution and a complex population structure with a few very large migratory units and many small local populations. Each population has its own spawning ground and/or time, thereby maintaining their genetic integrity. Several herring populations migrate between common feeding grounds and over-wintering areas resulting in frequent mixing of populations. Thus, many herring fisheries are based on mixed populations of different demographic status. In order to avoid overexploitation of weak populations and to conserve biodiversity, understanding the population structure and population mixing is important for maintaining biologically sustainable herring fisheries. The aim of this study was to investigate the genetic population structure of herring in the Faroese and surrounding waters, and to develop genetic markers for distinguishing between four herring management units (often called stocks), namely the Norwegian spring-spawning herring (NSSH), Icelandic summer-spawning herring (ISSH), North Sea autumn-spawning herring (NSAH), and Faroese autumn-spawning herring (FASH). Herring from the four stocks were sequenced at low coverage, and single nucleotide polymorphisms (SNPs) were called and used for population structure analysis and individual assignment. An ancestry-informative SNP panel with 118 SNPs was developed and tested on 240 individuals. The results showed that all four stocks appeared to be genetically differentiated populations, but at lower levels of differentiation between FASH and ISSH than the other two populations. Overall assignment rate with the SNP panel was 80.7\%, and agreement between the genetic and traditional visual assignment was 75.5\%. The NSAH and NSSH samples had the highest assignment rate (100\% and 98.3\%, respectively) and highest agreement between traditional and genetic assignment methods (96.6\% and 94.9\%, respectively). The FASH and ISSH samples had substantially lower assignment rates (72.9\% and 51.7\%, respectively) and agreement between traditional and genetic methods (39.5\% and 48.4\%, respectively)
@Article{, author = {í Kongsstovu S and Mikalsen S-O and Homrum EÍ and Jacobsen JA and Als TD and Gislason H and Flicek P and Nielsen EE and Dahl HA}, title = {Atlantic herring (Clupea harengus) population structure in the Northeast Atlantic Ocean}, journal = {Fisheries Research}, volume = {249}, pages = {106231}, year = {2022}, doi = {10.1016/j.fishres.2022.106231}, howpublished = {Advanced online publication: 21 January 2022}, abstract = {The Atlantic herring Clupea harengus L has a vast geographical distribution and a complex population structure with a few very large migratory units and many small local populations. Each population has its own spawning ground and/or time, thereby maintaining their genetic integrity. Several herring populations migrate between common feeding grounds and over-wintering areas resulting in frequent mixing of populations. Thus, many herring fisheries are based on mixed populations of different demographic status. In order to avoid overexploitation of weak populations and to conserve biodiversity, understanding the population structure and population mixing is important for maintaining biologically sustainable herring fisheries. The aim of this study was to investigate the genetic population structure of herring in the Faroese and surrounding waters, and to develop genetic markers for distinguishing between four herring management units (often called stocks), namely the Norwegian spring-spawning herring (NSSH), Icelandic summer-spawning herring (ISSH), North Sea autumn-spawning herring (NSAH), and Faroese autumn-spawning herring (FASH). Herring from the four stocks were sequenced at low coverage, and single nucleotide polymorphisms (SNPs) were called and used for population structure analysis and individual assignment. An ancestry-informative SNP panel with 118 SNPs was developed and tested on 240 individuals. The results showed that all four stocks appeared to be genetically differentiated populations, but at lower levels of differentiation between FASH and ISSH than the other two populations. Overall assignment rate with the SNP panel was 80.7\%, and agreement between the genetic and traditional visual assignment was 75.5\%. The NSAH and NSSH samples had the highest assignment rate (100\% and 98.3\%, respectively) and highest agreement between traditional and genetic assignment methods (96.6\% and 94.9\%, respectively). The FASH and ISSH samples had substantially lower assignment rates (72.9\% and 51.7\%, respectively) and agreement between traditional and genetic methods (39.5\% and 48.4\%, respectively)},} - F Cunningham, JE Allen, J Allen, J Alvarez-Jarreta, MR Amode, IM Armean, O Austine-Orimoloye, AG Azov, I Barnes, R Bennett, A Berry, J Bhai, A Bignell, K Billis, S Boddu, L Brooks, M Charkhchi, C Cummins, L Da Rin Fioretto, C Davidson, K Dodiya, S Donaldson, B El Houdaigui, T El Naboulsi, R Fatima, CG Giron, T Genez, JG Martinez, C Guijarro-Clarke, A Gymer, M Hardy, Z Hollis, T Hourlier, T Hunt, T Juettemann, V Kaikala, M Kay, I Lavidas, T Le, D Lemos, JC Marugán, S Mohanan, A Mushtaq, M Naven, DN Ogeh, A Parker, A Parton, M Perry, I Piližota, I Prosovetskaia, MP Sakthivel, AIA Salam, BM Schmitt, H Schuilenburg, D Sheppard, JG Pérez-Silva, W Stark, E Steed, K Sutinen, R Sukumaran, D Sumathipala, MM Suner, M Szpak, A Thormann, FF Tricomi, D Urbina-Gómez, A Veidenberg, TA Walsh, B Walts, N Willhoft, A Winterbottom, E Wass, M Chakiachvili, B Flint, A Frankish, S Giorgetti, L Haggerty, SE Hunt, GR IIsley, JE Loveland, FJ Martin, B Moore, JM Mudge, M Muffato, E Perry, M Ruffier, J Tate, D Thybert, SJ Trevanion, S Dyer, PW Harrison, KL Howe, AD Yates, DR Zerbino, P Flicek. Ensembl 2022. Nucleic Acids Res 2022;50(D1):D988–D995. doi:10.1093/nar/gkab1049
[BibTeX] [Abstract]
Ensembl (https://www.ensembl.org) is unique in its flexible infrastructure for access to genomic data and annotation. It has been designed to efficiently deliver annotation at scale for all eukaryotic life, and it also provides deep comprehensive annotation for key species. Genomes representing a greater diversity of species are increasingly being sequenced. In response, we have focussed our recent efforts on expediting the annotation of new assemblies. Here, we report the release of the greatest annual number of newly annotated genomes in the history of Ensembl via our dedicated Ensembl Rapid Release platform (http://rapid.ensembl.org). We have also developed a new method to generate comparative analyses at scale for these assemblies and, for the first time, we have annotated non-vertebrate eukaryotes. Meanwhile, we continually improve, extend and update the annotation for our high-value reference vertebrate genomes and report the details here. We have a range of specific software tools for specific tasks, such as the Ensembl Variant Effect Predictor (VEP) and the newly developed interface for the Variant Recoder. All Ensembl data, software and tools are freely available for download and are accessible programmatically.
@Article{34791404, author = {Cunningham F and Allen JE and Allen J and Alvarez-Jarreta J and Amode MR and Armean IM and Austine-Orimoloye O and Azov AG and Barnes I and Bennett R and Berry A and Bhai J and Bignell A and Billis K and Boddu S and Brooks L and Charkhchi M and Cummins C and Da Rin Fioretto L and Davidson C and Dodiya K and Donaldson S and El Houdaigui B and El Naboulsi T and Fatima R and Giron CG and Genez T and Martinez JG and Guijarro-Clarke C and Gymer A and Hardy M and Hollis Z and Hourlier T and Hunt T and Juettemann T and Kaikala V and Kay M and Lavidas I and Le T and Lemos D and Marugán JC and Mohanan S and Mushtaq A and Naven M and Ogeh DN and Parker A and Parton A and Perry M and Piližota I and Prosovetskaia I and Sakthivel MP and Salam AIA and Schmitt BM and Schuilenburg H and Sheppard D and Pérez-Silva JG and Stark W and Steed E and Sutinen K and Sukumaran R and Sumathipala D and Suner MM and Szpak M and Thormann A and Tricomi FF and Urbina-Gómez D and Veidenberg A and Walsh TA and Walts B and Willhoft N and Winterbottom A and Wass E and Chakiachvili M and Flint B and Frankish A and Giorgetti S and Haggerty L and Hunt SE and IIsley GR and Loveland JE and Martin FJ and Moore B and Mudge JM and Muffato M and Perry E and Ruffier M and Tate J and Thybert D and Trevanion SJ and Dyer S and Harrison PW and Howe KL and Yates AD and Zerbino DR and Flicek P}, title = {Ensembl 2022}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D988--D995}, year = {2022}, doi = {10.1093/nar/gkab1049}, howpublished = {Advanced online publication: 17 November 2021}, abstract = {Ensembl (https://www.ensembl.org) is unique in its flexible infrastructure for access to genomic data and annotation. It has been designed to efficiently deliver annotation at scale for all eukaryotic life, and it also provides deep comprehensive annotation for key species. Genomes representing a greater diversity of species are increasingly being sequenced. In response, we have focussed our recent efforts on expediting the annotation of new assemblies. Here, we report the release of the greatest annual number of newly annotated genomes in the history of Ensembl via our dedicated Ensembl Rapid Release platform (http://rapid.ensembl.org). We have also developed a new method to generate comparative analyses at scale for these assemblies and, for the first time, we have annotated non-vertebrate eukaryotes. Meanwhile, we continually improve, extend and update the annotation for our high-value reference vertebrate genomes and report the details here. We have a range of specific software tools for specific tasks, such as the Ensembl Variant Effect Predictor (VEP) and the newly developed interface for the Variant Recoder. All Ensembl data, software and tools are freely available for download and are accessible programmatically.},} - AD Yates, J Allen, RM Amode, AG Azov, M Barba, A Becerra, J Bhai, LI Campbell, M Carbajo Martinez, M Chakiachvili, K Chougule, M Christensen, B Contreras-Moreira, A Cuzick, L Da Rin Fioretto, P Davis, NH De Silva, S Diamantakis, S Dyer, J Elser, CV Filippi, A Gall, D Grigoriadis, C Guijarro-Clarke, P Gupta, KE Hammond-Kosack, KL Howe, P Jaiswal, V Kaikala, V Kumar, S Kumari, N Langridge, T Le, M Luypaert, GL Maslen, T Maurel, B Moore, M Muffato, A Mushtaq, G Naamati, S Naithani, A Olson, A Parker, M Paulini, H Pedro, E Perry, J Preece, M Quinton-Tulloch, F Rodgers, M Rosello, M Ruffier, J Seager, V Sitnik, M Szpak, J Tate, MK Tello-Ruiz, SJ Trevanion, M Urban, D Ware, S Wei, G Williams, A Winterbottom, M Zarowiecki, RD Finn, P Flicek. Ensembl Genomes 2022: an expanding genome resource for non-vertebrates. Nucleic Acids Res 2022;50(D1):D996–D1003. doi:10.1093/nar/gkab1007
[BibTeX] [Abstract]
Ensembl Genomes (https://www.ensemblgenomes.org) provides access to non-vertebrate genomes and analysis complementing vertebrate resources developed by the Ensembl project (https://www.ensembl.org). The two resources collectively present genome annotation through a consistent set of interfaces spanning the tree of life presenting genome sequence, annotation, variation, transcriptomic data and comparative analysis. Here, we present our largest increase in plant, metazoan and fungal genomes since the project’s inception creating one of the world’s most comprehensive genomic resources and describe our efforts to reduce genome redundancy in our Bacteria portal. We detail our new efforts in gene annotation, our emerging support for pangenome analysis, our efforts to accelerate data dissemination through the Ensembl Rapid Release resource and our new AlphaFold visualization. Finally, we present details of our future plans including updates on our integration with Ensembl, and how we plan to improve our support for the microbial research community. Software and data are made available without restriction via our website, online tools platform and programmatic interfaces (available under an Apache 2.0 license). Data updates are synchronised with Ensembl’s release cycle.
@Article{34791415, author = {Yates AD and Allen J and Amode RM and Azov AG and Barba M and Becerra A and Bhai J and Campbell LI and Carbajo Martinez M and Chakiachvili M and Chougule K and Christensen M and Contreras-Moreira B and Cuzick A and Da Rin Fioretto L and Davis P and De Silva NH and Diamantakis S and Dyer S and Elser J and Filippi CV and Gall A and Grigoriadis D and Guijarro-Clarke C and Gupta P and Hammond-Kosack KE and Howe KL and Jaiswal P and Kaikala V and Kumar V and Kumari S and Langridge N and Le T and Luypaert M and Maslen GL and Maurel T and Moore B and Muffato M and Mushtaq A and Naamati G and Naithani S and Olson A and Parker A and Paulini M and Pedro H and Perry E and Preece J and Quinton-Tulloch M and Rodgers F and Rosello M and Ruffier M and Seager J and Sitnik V and Szpak M and Tate J and Tello-Ruiz MK and Trevanion SJ and Urban M and Ware D and Wei S and Williams G and Winterbottom A and Zarowiecki M and Finn RD and Flicek P}, title = {Ensembl Genomes 2022: an expanding genome resource for non-vertebrates}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D996--D1003}, year = {2022}, doi = {10.1093/nar/gkab1007}, howpublished = {Advanced online publication: 13 November 2021}, abstract = {Ensembl Genomes (https://www.ensemblgenomes.org) provides access to non-vertebrate genomes and analysis complementing vertebrate resources developed by the Ensembl project (https://www.ensembl.org). The two resources collectively present genome annotation through a consistent set of interfaces spanning the tree of life presenting genome sequence, annotation, variation, transcriptomic data and comparative analysis. Here, we present our largest increase in plant, metazoan and fungal genomes since the project's inception creating one of the world's most comprehensive genomic resources and describe our efforts to reduce genome redundancy in our Bacteria portal. We detail our new efforts in gene annotation, our emerging support for pangenome analysis, our efforts to accelerate data dissemination through the Ensembl Rapid Release resource and our new AlphaFold visualization. Finally, we present details of our future plans including updates on our integration with Ensembl, and how we plan to improve our support for the microbial research community. Software and data are made available without restriction via our website, online tools platform and programmatic interfaces (available under an Apache 2.0 license). Data updates are synchronised with Ensembl's release cycle.},} - B Contreras-Moreira, G Naamati, M Rosello, JE Allen, SE Hunt, M Muffato, A Gall, P Flicek. Scripting Analyses of Genomes in Ensembl Plants. Methods Mol Biol 2022;2443:27–55. doi:10.1007/978-1-0716-2067-0_2
[BibTeX] [Abstract]
Ensembl Plants ( http://plants.ensembl.org ) offers genome-scale information for plants, with four releases per year. As of release 47 (April 2020) it features 79 species and includes genome sequence, gene models, and functional annotation. Comparative analyses help reconstruct the evolutionary history of gene families, genomes, and components of polyploid genomes. Some species have gene expression baseline reports or variation across genotypes. While the data can be accessed through the Ensembl genome browser, here we review specifically how our plant genomes can be interrogated programmatically and the data downloaded in bulk. These access routes are generally consistent across Ensembl for other non-plant species, including plant pathogens, pests, and pollinators.
@Article{35037199, author = {Contreras-Moreira B and Naamati G and Rosello M and Allen JE and Hunt SE and Muffato M and Gall A and Flicek P}, title = {Scripting Analyses of Genomes in Ensembl Plants}, journal = {Methods Mol Biol}, volume = {2443}, pages = {27--55}, year = {2022}, doi = {10.1007/978-1-0716-2067-0_2}, abstract = {Ensembl Plants ( http://plants.ensembl.org ) offers genome-scale information for plants, with four releases per year. As of release 47 (April 2020) it features 79 species and includes genome sequence, gene models, and functional annotation. Comparative analyses help reconstruct the evolutionary history of gene families, genomes, and components of polyploid genomes. Some species have gene expression baseline reports or variation across genotypes. While the data can be accessed through the Ensembl genome browser, here we review specifically how our plant genomes can be interrogated programmatically and the data downloaded in bulk. These access routes are generally consistent across Ensembl for other non-plant species, including plant pathogens, pests, and pollinators.},} - Darwin Tree of Life Project Consortium. Sequence locally, think globally: The Darwin Tree of Life Project. Proc Natl Acad Sci U S A 2022;119(4):e2115642118. doi:10.1073/pnas.2115642118
[BibTeX] [Abstract]
The goals of the Earth Biogenome Project-to sequence the genomes of all eukaryotic life on earth-are as daunting as they are ambitious. The Darwin Tree of Life Project was founded to demonstrate the credibility of these goals and to deliver at-scale genome sequences of unprecedented quality for a biogeographic region: the archipelago of islands that constitute Britain and Ireland. The Darwin Tree of Life Project is a collaboration between biodiversity organizations (museums, botanical gardens, and biodiversity institutes) and genomics institutes. Together, we have built a workflow that collects specimens from the field, robustly identifies them, performs sequencing, generates high-quality, curated assemblies, and releases these openly for the global community to use to build future science and conservation efforts.
@Article{35042805, author = {{Darwin Tree of Life Project Consortium}}, title = {Sequence locally, think globally: The Darwin Tree of Life Project}, journal = {Proc Natl Acad Sci U S A}, volume = {119}, number = {4}, pages = {e2115642118}, year = {2022}, doi = {10.1073/pnas.2115642118}, abstract = {The goals of the Earth Biogenome Project-to sequence the genomes of all eukaryotic life on earth-are as daunting as they are ambitious. The Darwin Tree of Life Project was founded to demonstrate the credibility of these goals and to deliver at-scale genome sequences of unprecedented quality for a biogeographic region: the archipelago of islands that constitute Britain and Ireland. The Darwin Tree of Life Project is a collaboration between biodiversity organizations (museums, botanical gardens, and biodiversity institutes) and genomics institutes. Together, we have built a workflow that collects specimens from the field, robustly identifies them, performs sequencing, generates high-quality, curated assemblies, and releases these openly for the global community to use to build future science and conservation efforts.},} - JM Mudge, J Ruiz-Orera, JR Prensner, MA Brunet, F Calvet, I Jungreis, JM Gonzalez, M Magrane, TF Martinez, JF Schulz, YT Yang, MM Albà, JL Aspden, PV Baranov, AA Bazzini, E Bruford, MJ Martin, L Calviello, AR Carvunis, J Chen, JP Couso, EW Deutsch, P Flicek, A Frankish, M Gerstein, N Hubner, NT Ingolia, M Kellis, G Menschaert, RL Moritz, U Ohler, X Roucou, A Saghatelian, JS Weissman, van S Heesch. Standardized annotation of translated open reading frames. Nat Biotechnol 2022;40(7):994–999. doi:10.1038/s41587-022-01369-0
[BibTeX]@Article{35831657, author = {Mudge JM and Ruiz-Orera J and Prensner JR and Brunet MA and Calvet F and Jungreis I and Gonzalez JM and Magrane M and Martinez TF and Schulz JF and Yang YT and Albà MM and Aspden JL and Baranov PV and Bazzini AA and Bruford E and Martin MJ and Calviello L and Carvunis AR and Chen J and Couso JP and Deutsch EW and Flicek P and Frankish A and Gerstein M and Hubner N and Ingolia NT and Kellis M and Menschaert G and Moritz RL and Ohler U and Roucou X and Saghatelian A and Weissman JS and van Heesch S}, title = {Standardized annotation of translated open reading frames}, journal = {Nat Biotechnol}, volume = {40}, number = {7}, pages = {994--999}, year = {2022}, doi = {10.1038/s41587-022-01369-0}, } - MKN Lawniczak, R Durbin, P Flicek, K Lindblad-Toh, X Wei, JM Archibald, WJ Baker, K Belov, ML Blaxter, T Marques Bonet, AK Childers, JA Coddington, KA Crandall, AJ Crawford, RP Davey, F Di Palma, Q Fang, W Haerty, N Hall, KJ Hoff, K Howe, ED Jarvis, WE Johnson, RN Johnson, PJ Kersey, X Liu, JV Lopez, EW Myers, OV Pettersson, AM Phillippy, MF Poelchau, KD Pruitt, A Rhie, JC Castilla-Rubio, SK Sahu, NA Salmon, PS Soltis, D Swarbreck, F Thibaud-Nissen, S Wang, JL Wegrzyn, G Zhang, H Zhang, HA Lewin, S Richards. Standards recommendations for the Earth BioGenome Project. Proc Natl Acad Sci U S A 2022;119(4):e2115639118. doi:10.1073/pnas.2115639118
[BibTeX] [Abstract]
A global international initiative, such as the Earth BioGenome Project (EBP), requires both agreement and coordination on standards to ensure that the collective effort generates rapid progress toward its goals. To this end, the EBP initiated five technical standards committees comprising volunteer members from the global genomics scientific community: Sample Collection and Processing, Sequencing and Assembly, Annotation, Analysis, and IT and Informatics. The current versions of the resulting standards documents are available on the EBP website, with the recognition that opportunities, technologies, and challenges may improve or change in the future, requiring flexibility for the EBP to meet its goals. Here, we describe some highlights from the proposed standards, and areas where additional challenges will need to be met.
@Article{35042802, author = {Lawniczak MKN and Durbin R and Flicek P and Lindblad-Toh K and Wei X and Archibald JM and Baker WJ and Belov K and Blaxter ML and Marques Bonet T and Childers AK and Coddington JA and Crandall KA and Crawford AJ and Davey RP and Di Palma F and Fang Q and Haerty W and Hall N and Hoff KJ and Howe K and Jarvis ED and Johnson WE and Johnson RN and Kersey PJ and Liu X and Lopez JV and Myers EW and Pettersson OV and Phillippy AM and Poelchau MF and Pruitt KD and Rhie A and Castilla-Rubio JC and Sahu SK and Salmon NA and Soltis PS and Swarbreck D and Thibaud-Nissen F and Wang S and Wegrzyn JL and Zhang G and Zhang H and Lewin HA and Richards S}, title = {Standards recommendations for the Earth BioGenome Project}, journal = {Proc Natl Acad Sci U S A}, volume = {119}, number = {4}, pages = {e2115639118}, year = {2022}, doi = {10.1073/pnas.2115639118}, abstract = {A global international initiative, such as the Earth BioGenome Project (EBP), requires both agreement and coordination on standards to ensure that the collective effort generates rapid progress toward its goals. To this end, the EBP initiated five technical standards committees comprising volunteer members from the global genomics scientific community: Sample Collection and Processing, Sequencing and Assembly, Annotation, Analysis, and IT and Informatics. The current versions of the resulting standards documents are available on the EBP website, with the recognition that opportunities, technologies, and challenges may improve or change in the future, requiring flexibility for the EBP to meet its goals. Here, we describe some highlights from the proposed standards, and areas where additional challenges will need to be met.},} - HA Lewin, S Richards, E Lieberman Aiden, ML Allende, JM Archibald, M Bálint, KB Barker, B Baumgartner, K Belov, G Bertorelle, ML Blaxter, J Cai, ND Caperello, K Carlson, JC Castilla-Rubio, SM Chaw, L Chen, AK Childers, JA Coddington, DA Conde, M Corominas, KA Crandall, AJ Crawford, F DiPalma, R Durbin, TE Ebenezer, SV Edwards, O Fedrigo, P Flicek, G Formenti, RA Gibbs, MTP Gilbert, MM Goldstein, JM Graves, HT Greely, IV Grigoriev, KJ Hackett, N Hall, D Haussler, KM Helgen, CJ Hogg, S Isobe, KS Jakobsen, A Janke, ED Jarvis, WE Johnson, SJM Jones, EK Karlsson, PJ Kersey, JH Kim, WJ Kress, S Kuraku, MKN Lawniczak, JH Leebens-Mack, X Li, K Lindblad-Toh, X Liu, JV Lopez, T Marques-Bonet, S Mazard, JAK Mazet, CJ Mazzoni, EW Myers, RJ O’Neill, S Paez, H Park, GE Robinson, C Roquet, OA Ryder, JSM Sabir, HB Shaffer, TM Shank, JS Sherkow, PS Soltis, B Tang, L Tedersoo, M Uliano-Silva, K Wang, X Wei, R Wetzer, JL Wilson, X Xu, H Yang, AD Yoder, G Zhang. The Earth BioGenome Project 2020: Starting the clock. Proc Natl Acad Sci U S A 2022;119(4):e2115635118. doi:10.1073/pnas.2115635118
[BibTeX]@Article{35042800, author = {Lewin HA and Richards S and Lieberman Aiden E and Allende ML and Archibald JM and Bálint M and Barker KB and Baumgartner B and Belov K and Bertorelle G and Blaxter ML and Cai J and Caperello ND and Carlson K and Castilla-Rubio JC and Chaw SM and Chen L and Childers AK and Coddington JA and Conde DA and Corominas M and Crandall KA and Crawford AJ and DiPalma F and Durbin R and Ebenezer TE and Edwards SV and Fedrigo O and Flicek P and Formenti G and Gibbs RA and Gilbert MTP and Goldstein MM and Graves JM and Greely HT and Grigoriev IV and Hackett KJ and Hall N and Haussler D and Helgen KM and Hogg CJ and Isobe S and Jakobsen KS and Janke A and Jarvis ED and Johnson WE and Jones SJM and Karlsson EK and Kersey PJ and Kim JH and Kress WJ and Kuraku S and Lawniczak MKN and Leebens-Mack JH and Li X and Lindblad-Toh K and Liu X and Lopez JV and Marques-Bonet T and Mazard S and Mazet JAK and Mazzoni CJ and Myers EW and O'Neill RJ and Paez S and Park H and Robinson GE and Roquet C and Ryder OA and Sabir JSM and Shaffer HB and Shank TM and Sherkow JS and Soltis PS and Tang B and Tedersoo L and Uliano-Silva M and Wang K and Wei X and Wetzer R and Wilson JL and Xu X and Yang H and Yoder AD and Zhang G}, title = {The Earth BioGenome Project 2020: Starting the clock}, journal = {Proc Natl Acad Sci U S A}, volume = {119}, number = {4}, pages = {e2115635118}, year = {2022}, doi = {10.1073/pnas.2115635118}, } - NH De Silva, J Bhai, M Chakiachvili, B Contreras-Moreira, C Cummins, A Frankish, A Gall, T Genez, KL Howe, SE Hunt, FJ Martin, B Moore, D Ogeh, A Parker, A Parton, M Ruffier, MP Sakthivel, D Sheppard, J Tate, A Thormann, D Thybert, SJ Trevanion, A Winterbottom, DR Zerbino, RD Finn, P Flicek, AD Yates. The Ensembl COVID-19 resource: ongoing integration of public SARS-CoV-2 data. Nucleic Acids Res 2022;50(D1):D765–D770. doi:10.1093/nar/gkab889
[BibTeX] [Abstract]
The COVID-19 pandemic has seen unprecedented use of SARS-CoV-2 genome sequencing for epidemiological tracking and identification of emerging variants. Understanding the potential impact of these variants on the infectivity of the virus and the efficacy of emerging therapeutics and vaccines has become a cornerstone of the fight against the disease. To support the maximal use of genomic information for SARS-CoV-2 research, we launched the Ensembl COVID-19 browser; the first virus to be encompassed within the Ensembl platform. This resource incorporates a new Ensembl gene set, multiple variant sets, and annotation from several relevant resources aligned to the reference SARS-CoV-2 assembly. Since the first release in May 2020, the content has been regularly updated using our new rapid release workflow, and tools such as the Ensembl Variant Effect Predictor have been integrated. The Ensembl COVID-19 browser is freely available at https://covid-19.ensembl.org.
@Article{34634797, author = {De Silva NH and Bhai J and Chakiachvili M and Contreras-Moreira B and Cummins C and Frankish A and Gall A and Genez T and Howe KL and Hunt SE and Martin FJ and Moore B and Ogeh D and Parker A and Parton A and Ruffier M and Sakthivel MP and Sheppard D and Tate J and Thormann A and Thybert D and Trevanion SJ and Winterbottom A and Zerbino DR and Finn RD and Flicek P and Yates AD}, title = {The Ensembl COVID-19 resource: ongoing integration of public SARS-CoV-2 data}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D765--D770}, year = {2022}, doi = {10.1093/nar/gkab889}, howpublished = {Advanced online publication: 11 October 2021}, note = {First posted as a preprint: 22 December 2020}, abstract = {The COVID-19 pandemic has seen unprecedented use of SARS-CoV-2 genome sequencing for epidemiological tracking and identification of emerging variants. Understanding the potential impact of these variants on the infectivity of the virus and the efficacy of emerging therapeutics and vaccines has become a cornerstone of the fight against the disease. To support the maximal use of genomic information for SARS-CoV-2 research, we launched the Ensembl COVID-19 browser; the first virus to be encompassed within the Ensembl platform. This resource incorporates a new Ensembl gene set, multiple variant sets, and annotation from several relevant resources aligned to the reference SARS-CoV-2 assembly. Since the first release in May 2020, the content has been regularly updated using our new rapid release workflow, and tools such as the Ensembl Variant Effect Predictor have been integrated. The Ensembl COVID-19 browser is freely available at https://covid-19.ensembl.org.},} - G Cantelli, A Bateman, C Brooksbank, AI Petrov, RS Malik-Sheriff, M Ide-Smith, H Hermjakob, P Flicek, R Apweiler, E Birney, J McEntyre. The European Bioinformatics Institute (EMBL-EBI) in 2021. Nucleic Acids Res 2022;50(D1):D11–D19. doi:10.1093/nar/gkab1127
[BibTeX] [Abstract]
The European Bioinformatics Institute (EMBL-EBI) maintains a comprehensive range of freely available and up-to-date molecular data resources, which includes over 40 resources covering every major data type in the life sciences. This year’s service update for EMBL-EBI includes new resources, PGS Catalog and AlphaFold DB, and updates on existing resources, including the COVID-19 Data Platform, trRosetta and RoseTTAfold models introduced in Pfam and InterPro, and the launch of Genome Integrations with Function and Sequence by UniProt and Ensembl. Furthermore, we highlight projects through which EMBL-EBI has contributed to the development of community-driven data standards and guidelines, including the Recommended Metadata for Biological Images (REMBI), and the BioModels Reproducibility Scorecard. Training is one of EMBL-EBI’s core missions and a key component of the provision of bioinformatics services to users: this year’s update includes many of the improvements that have been developed to EMBL-EBI’s online training offering.
@Article{34850134, author = {Cantelli G and Bateman A and Brooksbank C and Petrov AI and Malik-Sheriff RS and Ide-Smith M and Hermjakob H and Flicek P and Apweiler R and Birney E and McEntyre J}, title = {The European Bioinformatics Institute (EMBL-EBI) in 2021}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D11--D19}, year = {2022}, doi = {10.1093/nar/gkab1127}, howpublished = {Advanced online publication: 25 November 2021}, abstract = {The European Bioinformatics Institute (EMBL-EBI) maintains a comprehensive range of freely available and up-to-date molecular data resources, which includes over 40 resources covering every major data type in the life sciences. This year's service update for EMBL-EBI includes new resources, PGS Catalog and AlphaFold DB, and updates on existing resources, including the COVID-19 Data Platform, trRosetta and RoseTTAfold models introduced in Pfam and InterPro, and the launch of Genome Integrations with Function and Sequence by UniProt and Ensembl. Furthermore, we highlight projects through which EMBL-EBI has contributed to the development of community-driven data standards and guidelines, including the Recommended Metadata for Biological Images (REMBI), and the BioModels Reproducibility Scorecard. Training is one of EMBL-EBI's core missions and a key component of the provision of bioinformatics services to users: this year's update includes many of the improvements that have been developed to EMBL-EBI's online training offering.},} - MA Freeberg, LA Fromont, T D’Altri, AF Romero, JI Ciges, A Jene, G Kerry, M Moldes, R Ariosa, S Bahena, D Barrowdale, MC Barbero, D Fernandez-Orth, C Garcia-Linares, E Garcia-Rios, F Haziza, B Juhasz, OM Llobet, G Milla, A Mohan, M Rueda, A Sankar, D Shaju, A Shimpi, B Singh, C Thomas, de la S Torre, U Uyan, C Vasallo, P Flicek, R Guigo, A Navarro, H Parkinson, T Keane, J Rambla. The European Genome-phenome Archive in 2021. Nucleic Acids Res 2022;50(D1):D980–D987. doi:10.1093/nar/gkab1059
[BibTeX] [Abstract]
The European Genome-phenome Archive (EGA – https://ega-archive.org/) is a resource for long term secure archiving of all types of potentially identifiable genetic, phenotypic, and clinical data resulting from biomedical research projects. Its mission is to foster hosted data reuse, enable reproducibility, and accelerate biomedical and translational research in line with the FAIR principles. Launched in 2008, the EGA has grown quickly, currently archiving over 4,500 studies from nearly one thousand institutions. The EGA operates a distributed data access model in which requests are made to the data controller, not to the EGA, therefore, the submitter keeps control on who has access to the data and under which conditions. Given the size and value of data hosted, the EGA is constantly improving its value chain, that is, how the EGA can contribute to enhancing the value of human health data by facilitating its submission, discovery, access, and distribution, as well as leading the design and implementation of standards and methods necessary to deliver the value chain. The EGA has become a key GA4GH Driver Project, leading multiple development efforts and implementing new standards and tools, and has been appointed as an ELIXIR Core Data Resource.
@Article{34791407, author = {Freeberg MA and Fromont LA and D'Altri T and Romero AF and Ciges JI and Jene A and Kerry G and Moldes M and Ariosa R and Bahena S and Barrowdale D and Barbero MC and Fernandez-Orth D and Garcia-Linares C and Garcia-Rios E and Haziza F and Juhasz B and Llobet OM and Milla G and Mohan A and Rueda M and Sankar A and Shaju D and Shimpi A and Singh B and Thomas C and de la Torre S and Uyan U and Vasallo C and Flicek P and Guigo R and Navarro A and Parkinson H and Keane T and Rambla J}, title = {The European Genome-phenome Archive in 2021}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D980--D987}, year = {2022}, doi = {10.1093/nar/gkab1059}, howpublished = {Advanced online publication: 17 November 2021}, abstract = {The European Genome-phenome Archive (EGA - https://ega-archive.org/) is a resource for long term secure archiving of all types of potentially identifiable genetic, phenotypic, and clinical data resulting from biomedical research projects. Its mission is to foster hosted data reuse, enable reproducibility, and accelerate biomedical and translational research in line with the FAIR principles. Launched in 2008, the EGA has grown quickly, currently archiving over 4,500 studies from nearly one thousand institutions. The EGA operates a distributed data access model in which requests are made to the data controller, not to the EGA, therefore, the submitter keeps control on who has access to the data and under which conditions. Given the size and value of data hosted, the EGA is constantly improving its value chain, that is, how the EGA can contribute to enhancing the value of human health data by facilitating its submission, discovery, access, and distribution, as well as leading the design and implementation of standards and methods necessary to deliver the value chain. The EGA has become a key GA4GH Driver Project, leading multiple development efforts and implementing new standards and tools, and has been appointed as an ELIXIR Core Data Resource.},} - T Cezard, F Cunningham, SE Hunt, B Koylass, N Kumar, G Saunders, A Shen, AF Silva, K Tsukanov, S Venkataraman, P Flicek, H Parkinson, TM Keane. The European Variation Archive: a FAIR resource of genomic variation for all species. Nucleic Acids Res 2022;50(D1):D1216–D1220. doi:10.1093/nar/gkab960
[BibTeX] [Abstract]
The European Variation Archive (EVA; https://www.ebi.ac.uk/eva/) is a resource for sharing all types of genetic variation data (SNPs, indels, and structural variants) for all species. The EVA was created in 2014 to provide FAIR access to genetic variation data and has since grown to be a primary resource for genomic variants hosting >3 billion records. The EVA and dbSNP have established a compatible global system to assign unique identifiers to all submitted genetic variants. The EVA is active within the Global Alliance of Genomics and Health (GA4GH), maintaining, contributing and implementing standards such as VCF, Refget and Variant Representation Specification (VRS). In this article, we describe the submission and permanent accessioning services along with the different ways the data can be retrieved by the scientific community.
@Article{34718739, author = {Cezard T and Cunningham F and Hunt SE and Koylass B and Kumar N and Saunders G and Shen A and Silva AF and Tsukanov K and Venkataraman S and Flicek P and Parkinson H and Keane TM}, title = {The European Variation Archive: a FAIR resource of genomic variation for all species}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D1216--D1220}, year = {2022}, doi = {10.1093/nar/gkab960}, howpublished = {Advanced online publication: 28 October 2021}, abstract = {The European Variation Archive (EVA; https://www.ebi.ac.uk/eva/) is a resource for sharing all types of genetic variation data (SNPs, indels, and structural variants) for all species. The EVA was created in 2014 to provide FAIR access to genetic variation data and has since grown to be a primary resource for genomic variants hosting >3 billion records. The EVA and dbSNP have established a compatible global system to assign unique identifiers to all submitted genetic variants. The EVA is active within the Global Alliance of Genomics and Health (GA4GH), maintaining, contributing and implementing standards such as VCF, Refget and Variant Representation Specification (VRS). In this article, we describe the submission and permanent accessioning services along with the different ways the data can be retrieved by the scientific community.},} - T Wang, L Antonacci-Fulton, K Howe, HA Lawson, JK Lucas, AM Phillippy, AB Popejoy, M Asri, C Carson, MJP Chaisson, X Chang, R Cook-Deegan, AL Felsenfeld, RS Fulton, EP Garrison, NA Garrison, TA Graves-Lindsay, H Ji, EE Kenny, BA Koenig, D Li, T Marschall, JF McMichael, AM Novak, D Purushotham, VA Schneider, BI Schultz, MW Smith, HJ Sofia, T Weissman, P Flicek, H Li, KH Miga, B Paten, ED Jarvis, IM Hall, EE Eichler, D Haussler, PRC Human. The Human Pangenome Project: a global resource to map genomic diversity. Nature 2022;604(7906):437–446. doi:10.1038/s41586-022-04601-8
[BibTeX] [Abstract]
The human reference genome is the most widely used resource in human genetics and is due for a major update. Its current structure is a linear composite of merged haplotypes from more than 20 people, with a single individual comprising most of the sequence. It contains biases and errors within a framework that does not represent global human genomic variation. A high-quality reference with global representation of common variants, including single-nucleotide variants, structural variants and functional elements, is needed. The Human Pangenome Reference Consortium aims to create a more sophisticated and complete human reference genome with a graph-based, telomere-to-telomere representation of global genomic diversity. Here we leverage innovations in technology, study design and global partnerships with the goal of constructing the highest-possible quality human pangenome reference. Our goal is to improve data representation and streamline analyses to enable routine assembly of complete diploid genomes. With attention to ethical frameworks, the human pangenome reference will contain a more accurate and diverse representation of global genomic variation, improve gene-disease association studies across populations, expand the scope of genomics research to the most repetitive and polymorphic regions of the genome, and serve as the ultimate genetic resource for future biomedical research and precision medicine.
@Article{35444317, author = {Wang T and Antonacci-Fulton L and Howe K and Lawson HA and Lucas JK and Phillippy AM and Popejoy AB and Asri M and Carson C and Chaisson MJP and Chang X and Cook-Deegan R and Felsenfeld AL and Fulton RS and Garrison EP and Garrison NA and Graves-Lindsay TA and Ji H and Kenny EE and Koenig BA and Li D and Marschall T and McMichael JF and Novak AM and Purushotham D and Schneider VA and Schultz BI and Smith MW and Sofia HJ and Weissman T and Flicek P and Li H and Miga KH and Paten B and Jarvis ED and Hall IM and Eichler EE and Haussler D and Human PRC}, title = {The Human Pangenome Project: a global resource to map genomic diversity}, journal = {Nature}, volume = {604}, number = {7906}, pages = {437--446}, year = {2022}, doi = {10.1038/s41586-022-04601-8}, howpublished = {Advanced online publication: 20 April 2022}, abstract = {The human reference genome is the most widely used resource in human genetics and is due for a major update. Its current structure is a linear composite of merged haplotypes from more than 20 people, with a single individual comprising most of the sequence. It contains biases and errors within a framework that does not represent global human genomic variation. A high-quality reference with global representation of common variants, including single-nucleotide variants, structural variants and functional elements, is needed. The Human Pangenome Reference Consortium aims to create a more sophisticated and complete human reference genome with a graph-based, telomere-to-telomere representation of global genomic diversity. Here we leverage innovations in technology, study design and global partnerships with the goal of constructing the highest-possible quality human pangenome reference. Our goal is to improve data representation and streamline analyses to enable routine assembly of complete diploid genomes. With attention to ethical frameworks, the human pangenome reference will contain a more accurate and diverse representation of global genomic variation, improve gene-disease association studies across populations, expand the scope of genomics research to the most repetitive and polymorphic regions of the genome, and serve as the ultimate genetic resource for future biomedical research and precision medicine.},} - B Amos, C Aurrecoechea, M Barba, A Barreto, EY Basenko, W Bażant, R Belnap, AS Blevins, U Böhme, J Brestelli, BP Brunk, M Caddick, D Callan, L Campbell, MB Christensen, GK Christophides, K Crouch, K Davis, J DeBarry, R Doherty, Y Duan, M Dunn, D Falke, S Fisher, P Flicek, B Fox, B Gajria, GI Giraldo-Calderón, OS Harb, E Harper, C Hertz-Fowler, MJ Hickman, C Howington, S Hu, J Humphrey, J Iodice, A Jones, J Judkins, SA Kelly, JC Kissinger, DK Kwon, K Lamoureux, D Lawson, W Li, K Lies, D Lodha, J Long, RM MacCallum, G Maslen, MA McDowell, J Nabrzyski, DS Roos, SSC Rund, SW Schulman, A Shanmugasundram, V Sitnik, D Spruill, D Starns, CJ Stoeckert, SS Tomko, H Wang, S Warrenfeltz, R Wieck, PA Wilkinson, L Xu, J Zheng. VEuPathDB: the eukaryotic pathogen, vector and host bioinformatics resource center. Nucleic Acids Res 2022;50(D1):D898–D911. doi:10.1093/nar/gkab929
[BibTeX] [Abstract]
The Eukaryotic Pathogen, Vector and Host Informatics Resource (VEuPathDB, https://veupathdb.org) represents the 2019 merger of VectorBase with the EuPathDB projects. As a Bioinformatics Resource Center funded by the National Institutes of Health, with additional support from the Welllcome Trust, VEuPathDB supports >500 organisms comprising invertebrate vectors, eukaryotic pathogens (protists and fungi) and relevant free-living or non-pathogenic species or hosts. Designed to empower researchers with access to Omics data and bioinformatic analyses, VEuPathDB projects integrate >1700 pre-analysed datasets (and associated metadata) with advanced search capabilities, visualizations, and analysis tools in a graphic interface. Diverse data types are analysed with standardized workflows including an in-house OrthoMCL algorithm for predicting orthology. Comparisons are easily made across datasets, data types and organisms in this unique data mining platform. A new site-wide search facilitates access for both experienced and novice users. Upgraded infrastructure and workflows support numerous updates to the web interface, tools, searches and strategies, and Galaxy workspace where users can privately analyse their own data. Forthcoming upgrades include cloud-ready application architecture, expanded support for the Galaxy workspace, tools for interrogating host-pathogen interactions, and improved interactions with affiliated databases (ClinEpiDB, MicrobiomeDB) and other scientific resources, and increased interoperability with the Bacterial & Viral BRC.
@Article{34718728, author = {Amos B and Aurrecoechea C and Barba M and Barreto A and Basenko EY and Bażant W and Belnap R and Blevins AS and Böhme U and Brestelli J and Brunk BP and Caddick M and Callan D and Campbell L and Christensen MB and Christophides GK and Crouch K and Davis K and DeBarry J and Doherty R and Duan Y and Dunn M and Falke D and Fisher S and Flicek P and Fox B and Gajria B and Giraldo-Calderón GI and Harb OS and Harper E and Hertz-Fowler C and Hickman MJ and Howington C and Hu S and Humphrey J and Iodice J and Jones A and Judkins J and Kelly SA and Kissinger JC and Kwon DK and Lamoureux K and Lawson D and Li W and Lies K and Lodha D and Long J and MacCallum RM and Maslen G and McDowell MA and Nabrzyski J and Roos DS and Rund SSC and Schulman SW and Shanmugasundram A and Sitnik V and Spruill D and Starns D and Stoeckert CJ and Tomko SS and Wang H and Warrenfeltz S and Wieck R and Wilkinson PA and Xu L and Zheng J}, title = {VEuPathDB: the eukaryotic pathogen, vector and host bioinformatics resource center}, journal = {Nucleic Acids Res}, volume = {50}, number = {D1}, pages = {D898--D911}, year = {2022}, doi = {10.1093/nar/gkab929}, howpublished = {Advanced online publication: 28 October 2021}, abstract = {The Eukaryotic Pathogen, Vector and Host Informatics Resource (VEuPathDB, https://veupathdb.org) represents the 2019 merger of VectorBase with the EuPathDB projects. As a Bioinformatics Resource Center funded by the National Institutes of Health, with additional support from the Welllcome Trust, VEuPathDB supports >500 organisms comprising invertebrate vectors, eukaryotic pathogens (protists and fungi) and relevant free-living or non-pathogenic species or hosts. Designed to empower researchers with access to Omics data and bioinformatic analyses, VEuPathDB projects integrate >1700 pre-analysed datasets (and associated metadata) with advanced search capabilities, visualizations, and analysis tools in a graphic interface. Diverse data types are analysed with standardized workflows including an in-house OrthoMCL algorithm for predicting orthology. Comparisons are easily made across datasets, data types and organisms in this unique data mining platform. A new site-wide search facilitates access for both experienced and novice users. Upgraded infrastructure and workflows support numerous updates to the web interface, tools, searches and strategies, and Galaxy workspace where users can privately analyse their own data. Forthcoming upgrades include cloud-ready application architecture, expanded support for the Galaxy workspace, tools for interrogating host-pathogen interactions, and improved interactions with affiliated databases (ClinEpiDB, MicrobiomeDB) and other scientific resources, and increased interoperability with the Bacterial \& Viral BRC.},}
2021
- P Korlević, E McAlister, M Mayho, A Makunin, P Flicek, MKN Lawniczak. A minimally morphologically destructive approach for DNA retrieval and whole genome shotgun sequencing of pinned historic Dipteran vector species. Genome Biol Evol 2021;13(10):evab226. doi:10.1093/gbe/evab226
[BibTeX] [Abstract]
Museum collections contain enormous quantities of insect specimens collected over the past century, covering a period of increased and varied insecticide usage. These historic collections are therefore incredibly valuable as genomic snapshots of organisms before, during, and after exposure to novel selective pressures. However, these samples come with their own challenges compared to present-day collections, as they are fragile and retrievable DNA is low yield and fragmented. In this paper we tested several DNA extraction procedures across pinned historic Diptera specimens from four disease vector genera: Anopheles, Aedes, Culex and Glossina. We identify an approach that minimizes morphological damage while maximizing DNA retrieval for Illumina library preparation and sequencing that can accommodate the fragmented and low yield nature of historic DNA. We identify several key points in retrieving sufficient DNA while keeping morphological damage to a minimum: an initial rehydration step, a short incubation without agitation in a modified low salt Proteinase K buffer (referred to as “lysis buffer C” throughout), and critical point drying of samples post-extraction to prevent tissue collapse caused by air drying. The suggested method presented here provides a solid foundation for exploring the genomes and morphology of historic Diptera collections.
@Article{34599327, author = {Korlević P and McAlister E and Mayho M and Makunin A and Flicek P and Lawniczak MKN}, title = {A minimally morphologically destructive approach for DNA retrieval and whole genome shotgun sequencing of pinned historic Dipteran vector species}, journal = {Genome Biol Evol}, volume = {13}, number = {10}, pages = {evab226}, year = {2021}, doi = {10.1093/gbe/evab226}, note = {First posted as a preprint: 29 June 2021}, abstract = {Museum collections contain enormous quantities of insect specimens collected over the past century, covering a period of increased and varied insecticide usage. These historic collections are therefore incredibly valuable as genomic snapshots of organisms before, during, and after exposure to novel selective pressures. However, these samples come with their own challenges compared to present-day collections, as they are fragile and retrievable DNA is low yield and fragmented. In this paper we tested several DNA extraction procedures across pinned historic Diptera specimens from four disease vector genera: Anopheles, Aedes, Culex and Glossina. We identify an approach that minimizes morphological damage while maximizing DNA retrieval for Illumina library preparation and sequencing that can accommodate the fragmented and low yield nature of historic DNA. We identify several key points in retrieving sufficient DNA while keeping morphological damage to a minimum: an initial rehydration step, a short incubation without agitation in a modified low salt Proteinase K buffer (referred to as ``lysis buffer C'' throughout), and critical point drying of samples post-extraction to prevent tissue collapse caused by air drying. The suggested method presented here provides a solid foundation for exploring the genomes and morphology of historic Diptera collections.},} - A Joglekar, A Prjibelski, A Mahfouz, P Collier, S Lin, AK Schlusche, J Marrocco, SR Williams, B Haase, A Hayes, JG Chew, NI Weisenfeld, MY Wong, AN Stein, SA Hardwick, T Hunt, Q Wang, C Dieterich, Z Bent, O Fedrigo, SA Sloan, D Risso, ED Jarvis, P Flicek, W Luo, GS Pitt, A Frankish, AB Smit, ME Ross, HU Tilgner. A spatially resolved brain region- and cell type-specific isoform atlas of the postnatal mouse brain. Nat Commun 2021;12(1):463. doi:10.1038/s41467-020-20343-5 10.1016/j.neuron.2015.05.004 10.1038/nature07509 10.1038/nature08909 10.1016/j.molcel.2017.06.003 10.1038/nature13182 10.1016/j.tcb.2015.10.012 10.1126/science.1155390 10.1038/s41467-018-04559-0 10.1186/gb-2007-8-6-r108 10.1038/nbt.3242 10.1101/gr.230516.117 10.1186/s13059-018-1418-0 10.1186/s13059-015-0777-z 10.7554/eLife.03700 10.1073/pnas.1403244111 10.1016/j.neuron.2014.09.011 10.1186/gb-2004-5-10-r74 10.1073/pnas.95.22.13254 10.1186/1471-213X-5-14 10.1016/j.ajhg.2020.06.002 10.1146/annurev-neuro-062912-114322 10.1038/nrn.2016.27 10.3389/fnins.2012.00122 10.7554/eLife.11752 10.1038/nbt.4259 10.1016/j.cell.2019.05.031 10.1038/s41593-017-0056-2 10.1016/j.cell.2018.06.021 10.1523/JNEUROSCI.4515-09.2010 10.1038/s41586-020-2781-z 10.1016/j.cell.2018.06.021 10.1038/nn.4216 10.1214/aos/1013699998 10.1038/ncomms11020 10.1371/journal.pgen.1005907 10.1371/journal.pone.0017276 10.1016/j.cell.2014.11.035 10.1073/pnas.91.21.9975 10.1038/nrm.2017.103 10.1186/s12864-018-4772-0 10.1186/1471-2105-10-48 10.1371/journal.pone.0021800 10.1523/JNEUROSCI.5105-10.2011 10.1523/JNEUROSCI.15-10-06797.1995 10.1021/cn500337u 10.1016/j.bpj.2018.11.006 10.1038/s41398-018-0327-z 10.1002/pmic.201300196 10.1111/jnc.12816 10.1016/j.cell.2012.04.046 10.1523/JNEUROSCI.3470-14.2015 10.1096/fj.201901178R 10.1177/1073858414562217 10.1161/CIRCRESAHA.111.247957 10.1073/pnas.1521194113 10.1016/j.neuron.2017.01.009 10.1073/pnas.1614876114 10.1080/19336950.2016.1190055 10.1074/jbc.275.4.2589 10.1073/pnas.92.5.1510 10.1523/JNEUROSCI.1940-04.2004 10.1016/0378-1119(94)90773-0 10.1016/S0092-8674(03)00477-X 10.1242/jcs.216465 10.1016/0968-0004(91)90087-C 10.1016/j.celrep.2019.03.072 10.1038/s41593-019-0465-5 10.1073/pnas.1521194113 10.1186/s13059-019-1662-y 10.1038/nbt.4096 10.1016/j.cell.2019.05.031 10.1038/s41586-018-0414-6 10.1038/s41587-020-0591-3 10.1186/s13059-014-0560-6 10.1214/aoms/1177729380 10.2307/3001616 10.1214/aos/1013699998 10.1038/nbt.2705 10.1073/pnas.1400447111 10.1093/nar/gky955 10.1101/gr.135350.111 10.1093/
[BibTeX] [Abstract]
Splicing varies across brain regions, but the single-cell resolution of regional variation is unclear. We present a single-cell investigation of differential isoform expression (DIE) between brain regions using single-cell long-read sequencing in mouse hippocampus and prefrontal cortex in 45 cell types at postnatal day 7 ( www.isoformAtlas.com ). Isoform tests for DIE show better performance than exon tests. We detect hundreds of DIE events traceable to cell types, often corresponding to functionally distinct protein isoforms. Mostly, one cell type is responsible for brain-region specific DIE. However, for fewer genes, multiple cell types influence DIE. Thus, regional identity can, although rarely, override cell-type specificity. Cell types indigenous to one anatomic structure display distinctive DIE, e.g. the choroid plexus epithelium manifests distinct transcription-start-site usage. Spatial transcriptomics and long-read sequencing yield a spatially resolved splicing map. Our methods quantify isoform expression with cell-type and spatial resolution and it contributes to further our understanding of how the brain integrates molecular and cellular complexity.
@Article{33469025, author = {Joglekar A and Prjibelski A and Mahfouz A and Collier P and Lin S and Schlusche AK and Marrocco J and Williams SR and Haase B and Hayes A and Chew JG and Weisenfeld NI and Wong MY and Stein AN and Hardwick SA and Hunt T and Wang Q and Dieterich C and Bent Z and Fedrigo O and Sloan SA and Risso D and Jarvis ED and Flicek P and Luo W and Pitt GS and Frankish A and Smit AB and Ross ME and Tilgner HU}, title = {A spatially resolved brain region- and cell type-specific isoform atlas of the postnatal mouse brain}, journal = {Nat Commun}, volume = {12}, number = {1}, pages = {463}, year = {2021}, doi = {10.1038/s41467-020-20343-5 10.1016/j.neuron.2015.05.004 10.1038/nature07509 10.1038/nature08909 10.1016/j.molcel.2017.06.003 10.1038/nature13182 10.1016/j.tcb.2015.10.012 10.1126/science.1155390 10.1038/s41467-018-04559-0 10.1186/gb-2007-8-6-r108 10.1038/nbt.3242 10.1101/gr.230516.117 10.1186/s13059-018-1418-0 10.1186/s13059-015-0777-z 10.7554/eLife.03700 10.1073/pnas.1403244111 10.1016/j.neuron.2014.09.011 10.1186/gb-2004-5-10-r74 10.1073/pnas.95.22.13254 10.1186/1471-213X-5-14 10.1016/j.ajhg.2020.06.002 10.1146/annurev-neuro-062912-114322 10.1038/nrn.2016.27 10.3389/fnins.2012.00122 10.7554/eLife.11752 10.1038/nbt.4259 10.1016/j.cell.2019.05.031 10.1038/s41593-017-0056-2 10.1016/j.cell.2018.06.021 10.1523/JNEUROSCI.4515-09.2010 10.1038/s41586-020-2781-z 10.1016/j.cell.2018.06.021 10.1038/nn.4216 10.1214/aos/1013699998 10.1038/ncomms11020 10.1371/journal.pgen.1005907 10.1371/journal.pone.0017276 10.1016/j.cell.2014.11.035 10.1073/pnas.91.21.9975 10.1038/nrm.2017.103 10.1186/s12864-018-4772-0 10.1186/1471-2105-10-48 10.1371/journal.pone.0021800 10.1523/JNEUROSCI.5105-10.2011 10.1523/JNEUROSCI.15-10-06797.1995 10.1021/cn500337u 10.1016/j.bpj.2018.11.006 10.1038/s41398-018-0327-z 10.1002/pmic.201300196 10.1111/jnc.12816 10.1016/j.cell.2012.04.046 10.1523/JNEUROSCI.3470-14.2015 10.1096/fj.201901178R 10.1177/1073858414562217 10.1161/CIRCRESAHA.111.247957 10.1073/pnas.1521194113 10.1016/j.neuron.2017.01.009 10.1073/pnas.1614876114 10.1080/19336950.2016.1190055 10.1074/jbc.275.4.2589 10.1073/pnas.92.5.1510 10.1523/JNEUROSCI.1940-04.2004 10.1016/0378-1119(94)90773-0 10.1016/S0092-8674(03)00477-X 10.1242/jcs.216465 10.1016/0968-0004(91)90087-C 10.1016/j.celrep.2019.03.072 10.1038/s41593-019-0465-5 10.1073/pnas.1521194113 10.1186/s13059-019-1662-y 10.1038/nbt.4096 10.1016/j.cell.2019.05.031 10.1038/s41586-018-0414-6 10.1038/s41587-020-0591-3 10.1186/s13059-014-0560-6 10.1214/aoms/1177729380 10.2307/3001616 10.1214/aos/1013699998 10.1038/nbt.2705 10.1073/pnas.1400447111 10.1093/nar/gky955 10.1101/gr.135350.111 10.1093/}, note = {First posted as a preprint: 27 August 2020}, abstract = {Splicing varies across brain regions, but the single-cell resolution of regional variation is unclear. We present a single-cell investigation of differential isoform expression (DIE) between brain regions using single-cell long-read sequencing in mouse hippocampus and prefrontal cortex in 45 cell types at postnatal day 7 ( www.isoformAtlas.com ). Isoform tests for DIE show better performance than exon tests. We detect hundreds of DIE events traceable to cell types, often corresponding to functionally distinct protein isoforms. Mostly, one cell type is responsible for brain-region specific DIE. However, for fewer genes, multiple cell types influence DIE. Thus, regional identity can, although rarely, override cell-type specificity. Cell types indigenous to one anatomic structure display distinctive DIE, e.g. the choroid plexus epithelium manifests distinct transcription-start-site usage. Spatial transcriptomics and long-read sequencing yield a spatially resolved splicing map. Our methods quantify isoform expression with cell-type and spatial resolution and it contributes to further our understanding of how the brain integrates molecular and cellular complexity.},} - FJ Martin, A Gall, M Szpak, P Flicek. Accessing Livestock Resources in Ensembl. Front Genet 2021;12:650228. doi:10.3389/fgene.2021.650228
[BibTeX] [Abstract]
Genome assembly is cheaper, more accurate and more automated than it has ever been. This is due to a combination of more cost-efficient chemistries, new sequencing technologies and better algorithms. The livestock community has been at the forefront of this new wave of genome assembly, generating some of the highest quality vertebrate genome sequences. Ensembl’s goal is to add functional and comparative annotation to these genomes, through our gene annotation, genomic alignments, gene trees, regulatory, and variation data. We run computationally complex analyses in a high throughput and consistent manner to help accelerate downstream science. Our livestock resources are continuously growing in both breadth and depth. We annotate reference genome assemblies for newly sequenced species and regularly update annotation for existing genomes. We are the only major resource to support the annotation of breeds and other non-reference assemblies. We currently provide resources for 13 pig breeds, maternal and paternal haplotypes for hybrid cattle and various other non-reference or wild type assemblies for livestock species. Here, we describe the livestock data present in Ensembl and provide protocols for how to view data in our genome browser, download via it our FTP site, manipulate it via our tools and interact with it programmatically via our REST API.
@Article{33995484, author = {Martin FJ and Gall A and Szpak M and Flicek P}, title = {Accessing Livestock Resources in Ensembl}, journal = {Front Genet}, volume = {12}, pages = {650228}, year = {2021}, doi = {10.3389/fgene.2021.650228}, abstract = {Genome assembly is cheaper, more accurate and more automated than it has ever been. This is due to a combination of more cost-efficient chemistries, new sequencing technologies and better algorithms. The livestock community has been at the forefront of this new wave of genome assembly, generating some of the highest quality vertebrate genome sequences. Ensembl's goal is to add functional and comparative annotation to these genomes, through our gene annotation, genomic alignments, gene trees, regulatory, and variation data. We run computationally complex analyses in a high throughput and consistent manner to help accelerate downstream science. Our livestock resources are continuously growing in both breadth and depth. We annotate reference genome assemblies for newly sequenced species and regularly update annotation for existing genomes. We are the only major resource to support the annotation of breeds and other non-reference assemblies. We currently provide resources for 13 pig breeds, maternal and paternal haplotypes for hybrid cattle and various other non-reference or wild type assemblies for livestock species. Here, we describe the livestock data present in Ensembl and provide protocols for how to view data in our genome browser, download via it our FTP site, manipulate it via our tools and interact with it programmatically via our REST API.},} - L Grassi, OG Izuogu, NAN Jorge, D Seyres, M Bustamante, F Burden, S Farrow, N Farahi, FJ Martin, A Frankish, JM Mudge, M Kostadima, R Petersen, JJ Lambourne, S Rowlston, E Martin-Rendon, L Clarke, K Downes, X Estivill, P Flicek, JHA Martens, ML Yaspo, HG Stunnenberg, WH Ouwehand, F Passetti, E Turro, M Frontini. Cell type specific novel lncRNAs and circRNAs in the BLUEPRINT haematopoietic transcriptomes atlas. Haematologica 2021;106(10):2613–2623. doi:10.3324/haematol.2019.238147
[BibTeX] [Abstract]
Transcriptional profiling of hematopoietic cell subpopulations has helped characterize the developmental stages of the hematopoietic system and the molecular bases of malignant and non-malignant blood diseases for the past three decades. Previously, only the genes targeted by expression microarrays could be profiled genome wide. High-throughput RNA sequencing (RNA-seq), however, encompasses a broader repertoire of RNA molecules, without restriction to previously annotated genes. We analysed the BLUEPRINT consortium RNA- seq data for mature hematopoietic cell types. The data comprised 90 total RNA-seq samples, each composed of one of 27 cell types, and 32 small RNA-seq samples, each composed of one of 11 cell types. We estimated gene and isoform expression levels for each cell type using existing annotations from Ensembl. We then used guided transcriptome assembly to discover unannotated transcripts. We identified hundreds of novel non-coding RNA genes and showed that the majority have cell type dependent expression. We also characterized the expression of circular RNAs and found that these are also cell type specific. These analyses refine the active transcriptional landscape of mature hematopoietic cells, highlight abundant genes and transcriptional isoforms for each blood cell type, and provide a valuable resource for researchers of hematological development and diseases. Finally, we made the data accessible via a web-based interface: https://blueprint.haem.cam.ac.uk/bloodatlas/.
@Article{32703790, author = {Grassi L and Izuogu OG and Jorge NAN and Seyres D and Bustamante M and Burden F and Farrow S and Farahi N and Martin FJ and Frankish A and Mudge JM and Kostadima M and Petersen R and Lambourne JJ and Rowlston S and Martin-Rendon E and Clarke L and Downes K and Estivill X and Flicek P and Martens JHA and Yaspo ML and Stunnenberg HG and Ouwehand WH and Passetti F and Turro E and Frontini M}, title = {Cell type specific novel lncRNAs and circRNAs in the BLUEPRINT haematopoietic transcriptomes atlas}, journal = {Haematologica}, volume = {106}, number = {10}, pages = {2613--2623}, year = {2021}, doi = {10.3324/haematol.2019.238147}, howpublished = {Advanced online publication: 23 July 2020}, note = {First posted as a preprint: 10 September 2019}, abstract = {Transcriptional profiling of hematopoietic cell subpopulations has helped characterize the developmental stages of the hematopoietic system and the molecular bases of malignant and non-malignant blood diseases for the past three decades. Previously, only the genes targeted by expression microarrays could be profiled genome wide. High-throughput RNA sequencing (RNA-seq), however, encompasses a broader repertoire of RNA molecules, without restriction to previously annotated genes. We analysed the BLUEPRINT consortium RNA- seq data for mature hematopoietic cell types. The data comprised 90 total RNA-seq samples, each composed of one of 27 cell types, and 32 small RNA-seq samples, each composed of one of 11 cell types. We estimated gene and isoform expression levels for each cell type using existing annotations from Ensembl. We then used guided transcriptome assembly to discover unannotated transcripts. We identified hundreds of novel non-coding RNA genes and showed that the majority have cell type dependent expression. We also characterized the expression of circular RNAs and found that these are also cell type specific. These analyses refine the active transcriptional landscape of mature hematopoietic cells, highlight abundant genes and transcriptional isoforms for each blood cell type, and provide a valuable resource for researchers of hematological development and diseases. Finally, we made the data accessible via a web-based interface: https://blueprint.haem.cam.ac.uk/bloodatlas/.},} - KL Howe, P Achuthan, J Allen, J Allen, J Alvarez-Jarreta, MR Amode, IM Armean, AG Azov, R Bennett, J Bhai, K Billis, S Boddu, M Charkhchi, C Cummins, L Da Rin Fioretto, C Davidson, K Dodiya, B El Houdaigui, R Fatima, A Gall, C Garcia Giron, T Grego, C Guijarro-Clarke, L Haggerty, A Hemrom, T Hourlier, OG Izuogu, T Juettemann, V Kaikala, M Kay, I Lavidas, T Le, D Lemos, J Gonzalez Martinez, JC Marugán, T Maurel, AC McMahon, S Mohanan, B Moore, M Muffato, DN Oheh, D Paraschas, A Parker, A Parton, I Prosovetskaia, MP Sakthivel, AIA Salam, BM Schmitt, H Schuilenburg, D Sheppard, E Steed, M Szpak, M Szuba, K Taylor, A Thormann, G Threadgold, B Walts, A Winterbottom, M Chakiachvili, A Chaubal, N De Silva, B Flint, A Frankish, SE Hunt, GR IIsley, N Langridge, JE Loveland, FJ Martin, JM Mudge, J Morales, E Perry, M Ruffier, J Tate, D Thybert, SJ Trevanion, F Cunningham, AD Yates, DR Zerbino, P Flicek. Ensembl 2021. Nucleic Acids Res 2021;49(D1):D884–D891. doi:10.1093/nar/gkaa942
[BibTeX] [Abstract]
The Ensembl project (https://www.ensembl.org) annotates genomes and disseminates genomic data for vertebrate species. We create detailed and comprehensive annotation of gene structures, regulatory elements and variants, and enable comparative genomics by inferring the evolutionary history of genes and genomes. Our integrated genomic data are made available in a variety of ways, including genome browsers, search interfaces, specialist tools such as the Ensembl Variant Effect Predictor, download files and programmatic interfaces. Here, we present recent Ensembl developments including two new website portals. Ensembl Rapid Release (http://rapid.ensembl.org) is designed to provide core tools and services for genomes as soon as possible and has been deployed to support large biodiversity sequencing projects. Our SARS-CoV-2 genome browser (https://covid-19.ensembl.org) integrates our own annotation with publicly available genomic data from numerous sources to facilitate the use of genomics in the international scientific response to the COVID-19 pandemic. We also report on other updates to our annotation resources, tools and services. All Ensembl data and software are freely available without restriction.
@Article{33137190, author = {Howe KL and Achuthan P and Allen J and Allen J and Alvarez-Jarreta J and Amode MR and Armean IM and Azov AG and Bennett R and Bhai J and Billis K and Boddu S and Charkhchi M and Cummins C and Da Rin Fioretto L and Davidson C and Dodiya K and El Houdaigui B and Fatima R and Gall A and Garcia Giron C and Grego T and Guijarro-Clarke C and Haggerty L and Hemrom A and Hourlier T and Izuogu OG and Juettemann T and Kaikala V and Kay M and Lavidas I and Le T and Lemos D and Gonzalez Martinez J and Marugán JC and Maurel T and McMahon AC and Mohanan S and Moore B and Muffato M and Oheh DN and Paraschas D and Parker A and Parton A and Prosovetskaia I and Sakthivel MP and Salam AIA and Schmitt BM and Schuilenburg H and Sheppard D and Steed E and Szpak M and Szuba M and Taylor K and Thormann A and Threadgold G and Walts B and Winterbottom A and Chakiachvili M and Chaubal A and De Silva N and Flint B and Frankish A and Hunt SE and IIsley GR and Langridge N and Loveland JE and Martin FJ and Mudge JM and Morales J and Perry E and Ruffier M and Tate J and Thybert D and Trevanion SJ and Cunningham F and Yates AD and Zerbino DR and Flicek P}, title = {Ensembl 2021}, journal = {Nucleic Acids Res}, volume = {49}, number = {D1}, pages = {D884--D891}, year = {2021}, doi = {10.1093/nar/gkaa942}, howpublished = {Advanced online publication: 2 November 2020}, abstract = {The Ensembl project (https://www.ensembl.org) annotates genomes and disseminates genomic data for vertebrate species. We create detailed and comprehensive annotation of gene structures, regulatory elements and variants, and enable comparative genomics by inferring the evolutionary history of genes and genomes. Our integrated genomic data are made available in a variety of ways, including genome browsers, search interfaces, specialist tools such as the Ensembl Variant Effect Predictor, download files and programmatic interfaces. Here, we present recent Ensembl developments including two new website portals. Ensembl Rapid Release (http://rapid.ensembl.org) is designed to provide core tools and services for genomes as soon as possible and has been deployed to support large biodiversity sequencing projects. Our SARS-CoV-2 genome browser (https://covid-19.ensembl.org) integrates our own annotation with publicly available genomic data from numerous sources to facilitate the use of genomics in the international scientific response to the COVID-19 pandemic. We also report on other updates to our annotation resources, tools and services. All Ensembl data and software are freely available without restriction.},} - C Kern, Y Wang, X Xu, Z Pan, M Halstead, G Chanthavixay, P Saelao, S Waters, R Xiang, A Chamberlain, I Korf, ME Delany, HH Cheng, JF Medrano, AL Van Eenennaam, CK Tuggle, C Ernst, P Flicek, G Quon, P Ross, H Zhou. Functional annotations of three domestic animal genomes provide vital resources for comparative and agricultural research. Nat Commun 2021;12(1):1821. doi:10.1038/s41467-021-22100-8 10.1016/j.gfs.2019.100325 10.1038/nature03030 10.1073/pnas.0903103106 10.1126/science.1105136 10.1186/gb-2012-13-1-r1 10.1038/nature11247 10.1126/science.1222794 10.1038/nature14248 10.1038/s41586-020-2449-8 10.1038/s41586-020-2093-3 10.1093/gbe/evq087 10.1126/science.aat7244 10.1093/molbev/msu309 10.1038/ncomms14229 10.1007/s00427-019-00629-5 10.1186/s12915-019-0726-5 10.1093/molbev/msx156 10.1186/s13059-015-0622-4 10.1111/age.12466 10.1111/age.12717 10.1146/annurev-animal-020518-114913 10.1186/s12864-020-07078-9 10.1186/s13059-020-02197-8 10.1038/nature13972 10.1038/nature13985 10.1126/science.1141319 10.1016/j.cell.2007.05.009 10.1101/gr.4074106 10.1038/nmeth.2688 10.1101/gr.136184.111 10.1038/nmeth.1906 10.1093/nar/gks1284 10.1016/j.cell.2007.05.042 10.1038/nature06008 10.1038/nature09990 10.1073/pnas.1016071107 10.1038/nature09906 10.1038/nature13992 10.1038/ng.808 10.1093/nar/28.1.27 10.1038/nature11212 10.1016/j.molcel.2010.05.004 10.1038/ng.2713 10.7150/ijbs.9442 10.1038/nature11082 10.1016/j.cell.2014.11.021 10.1038/nature13395 10.1038/nature12716 10.1093/hmg/ddg180 10.1073/pnas.0909344107 10.1186/1471-2105-12-155 10.1186/s12864-018-4902-8 10.1073/pnas.1904159116 10.1186/s12864-018-5037-7 10.1038/s41598-020-61678-9 10.1038/ng.759 10.1093/bioinformatics/bts635 10.1093/bioinformatics/btp352 10.1093/bioinformatics/btu638 10.1093/bioinformatics/btp616 10.1038/nbt.1508 10.1093/nar/gku365 10.1186/gb-2008-9-9-r137 10.1093/bioinformatics/btq033 10.1093/molbev/msx116 10.1038/nprot.2008.211 10.1038/nmeth.3772 10.1186/s13059-019-1642-2 10.1016/j.tig.2013.05.010 10.1016/j.febslet.2015.04.024 10.1186/s12915-018-0556-x 10.1186/s12864-018-4800-0 10.1186/s12864-016-2516-6 10.1093/bioinformatics/btr064 10.1101/gr.229102
[BibTeX] [Abstract]
Gene regulatory elements are central drivers of phenotypic variation and thus of critical importance towards understanding the genetics of complex traits. The Functional Annotation of Animal Genomes consortium was formed to collaboratively annotate the functional elements in animal genomes, starting with domesticated animals. Here we present an expansive collection of datasets from eight diverse tissues in three important agricultural species: chicken (Gallus gallus), pig (Sus scrofa), and cattle (Bos taurus). Comparative analysis of these datasets and those from the human and mouse Encyclopedia of DNA Elements projects reveal that a core set of regulatory elements are functionally conserved independent of divergence between species, and that tissue-specific transcription factor occupancy at regulatory elements and their predicted target genes are also conserved. These datasets represent a unique opportunity for the emerging field of comparative epigenomics, as well as the agricultural research community, including species that are globally important food resources.
@Article{33758196, author = {Kern C and Wang Y and Xu X and Pan Z and Halstead M and Chanthavixay G and Saelao P and Waters S and Xiang R and Chamberlain A and Korf I and Delany ME and Cheng HH and Medrano JF and Van Eenennaam AL and Tuggle CK and Ernst C and Flicek P and Quon G and Ross P and Zhou H}, title = {Functional annotations of three domestic animal genomes provide vital resources for comparative and agricultural research}, journal = {Nat Commun}, volume = {12}, number = {1}, pages = {1821}, year = {2021}, doi = {10.1038/s41467-021-22100-8 10.1016/j.gfs.2019.100325 10.1038/nature03030 10.1073/pnas.0903103106 10.1126/science.1105136 10.1186/gb-2012-13-1-r1 10.1038/nature11247 10.1126/science.1222794 10.1038/nature14248 10.1038/s41586-020-2449-8 10.1038/s41586-020-2093-3 10.1093/gbe/evq087 10.1126/science.aat7244 10.1093/molbev/msu309 10.1038/ncomms14229 10.1007/s00427-019-00629-5 10.1186/s12915-019-0726-5 10.1093/molbev/msx156 10.1186/s13059-015-0622-4 10.1111/age.12466 10.1111/age.12717 10.1146/annurev-animal-020518-114913 10.1186/s12864-020-07078-9 10.1186/s13059-020-02197-8 10.1038/nature13972 10.1038/nature13985 10.1126/science.1141319 10.1016/j.cell.2007.05.009 10.1101/gr.4074106 10.1038/nmeth.2688 10.1101/gr.136184.111 10.1038/nmeth.1906 10.1093/nar/gks1284 10.1016/j.cell.2007.05.042 10.1038/nature06008 10.1038/nature09990 10.1073/pnas.1016071107 10.1038/nature09906 10.1038/nature13992 10.1038/ng.808 10.1093/nar/28.1.27 10.1038/nature11212 10.1016/j.molcel.2010.05.004 10.1038/ng.2713 10.7150/ijbs.9442 10.1038/nature11082 10.1016/j.cell.2014.11.021 10.1038/nature13395 10.1038/nature12716 10.1093/hmg/ddg180 10.1073/pnas.0909344107 10.1186/1471-2105-12-155 10.1186/s12864-018-4902-8 10.1073/pnas.1904159116 10.1186/s12864-018-5037-7 10.1038/s41598-020-61678-9 10.1038/ng.759 10.1093/bioinformatics/bts635 10.1093/bioinformatics/btp352 10.1093/bioinformatics/btu638 10.1093/bioinformatics/btp616 10.1038/nbt.1508 10.1093/nar/gku365 10.1186/gb-2008-9-9-r137 10.1093/bioinformatics/btq033 10.1093/molbev/msx116 10.1038/nprot.2008.211 10.1038/nmeth.3772 10.1186/s13059-019-1642-2 10.1016/j.tig.2013.05.010 10.1016/j.febslet.2015.04.024 10.1186/s12915-018-0556-x 10.1186/s12864-018-4800-0 10.1186/s12864-016-2516-6 10.1093/bioinformatics/btr064 10.1101/gr.229102}, abstract = {Gene regulatory elements are central drivers of phenotypic variation and thus of critical importance towards understanding the genetics of complex traits. The Functional Annotation of Animal Genomes consortium was formed to collaboratively annotate the functional elements in animal genomes, starting with domesticated animals. Here we present an expansive collection of datasets from eight diverse tissues in three important agricultural species: chicken (Gallus gallus), pig (Sus scrofa), and cattle (Bos taurus). Comparative analysis of these datasets and those from the human and mouse Encyclopedia of DNA Elements projects reveal that a core set of regulatory elements are functionally conserved independent of divergence between species, and that tissue-specific transcription factor occupancy at regulatory elements and their predicted target genes are also conserved. These datasets represent a unique opportunity for the emerging field of comparative epigenomics, as well as the agricultural research community, including species that are globally important food resources.},} - HL Rehm, AJH Page, L Smith, JB Adams, G Alterovitz, LJ Babb, MP Barkley, M Baudis, MJS Beauvais, T Beck, JS Beckmann, S Beltran, D Bernick, A Bernier, JK Bonfield, TF Boughtwood, G Bourque, SR Bowers, AJ Brookes, M Brudno, MH Brush, D Bujold, T Burdett, OJ Buske, MN Cabili, DL Cameron, RJ Carroll, E Casas-Silva, D Chakravarty, BP Chaudhari, SH Chen, JM Cherry, J Chung, M Cline, HL Clissold, RM Cook-Deegan, M Courtot, F Cunningham, M Cupak, RM Davies, D Denisko, MJ Doerr, LI Dolman, ES Dove, LJ Dursi, SOM Dyke, JA Eddy, K Eilbeck, KP Ellrott, S Fairley, KA Fakhro, HV Firth, MS Fitzsimons, M Fiume, P Flicek, IM Fore, MA Freeberg, RR Freimuth, LA Fromont, J Fuerth, CL Gaff, W Gan, EM Ghanaim, D Glazer, RC Green, M Griffith, OL Griffith, RL Grossman, T Groza, JMG Auvil, R Guigó, D Gupta, MA Haendel, A Hamosh, DP Hansen, RK Hart, DM Hartley, D Haussler, RM Hendricks-Sturrup, CWL Ho, AE Hobb, MM Hoffman, OM Hofmann, P Holub, JS Hsu, JP Hubaux, SE Hunt, A Husami, JO Jacobsen, SS Jamuar, EL Janes, F Jeanson, A Jené, AL Johns, Y Joly, SJM Jones, A Kanitz, K Kato, TM Keane, K Kekesi-Lafrance, J Kelleher, G Kerry, SS Khor, BM Knoppers, MA Konopko, K Kosaki, M Kuba, J Lawson, R Leinonen, S Li, MF Lin, M Linden, X Liu, I Udara Liyanage, J Lopez, AM Lucassen, M Lukowski, AL Mann, J Marshall, M Mattioni, A Metke-Jimenez, A Middleton, RJ Milne, F Molnár-Gábor, N Mulder, MC Munoz-Torres, R Nag, H Nakagawa, J Nasir, A Navarro, TH Nelson, A Niewielska, A Nisselle, J Niu, TH Nyrönen, BD O’Connor, S Oesterle, S Ogishima, VO Wang, LAD Paglione, E Palumbo, HE Parkinson, AA Philippakis, AD Pizarro, A Prlic, J Rambla, A Rendon, RA Rider, PN Robinson, KW Rodarmer, LL Rodriguez, AF Rubin, M Rueda, GA Rushton, RS Ryan, GI Saunders, H Schuilenburg, T Schwede, S Scollen, A Senf, NC Sheffield, N Skantharajah, AV Smith, HJ Sofia, D Spalding, AB Spurdle, Z Stark, LD Stein, M Suematsu, P Tan, JA Tedds, AA Thomson, A Thorogood, TL Tickle, K Tokunaga, J Törnroos, D Torrents, S Upchurch, A Valencia, RV Guimera, J Vamathevan, S Varma, DF Vears, C Viner, C Voisin, AH Wagner, SE Wallace, BP Walsh, MS Williams, EC Winkler, BJ Wold, GM Wood, JP Woolley, C Yamasaki, AD Yates, CK Yung, LJ Zass, K Zaytseva, J Zhang, P Goodhand, K North, E Birney. GA4GH: International policies and standards for data sharing across genomic research and healthcare. Cell Genom 2021;1(2):100029. doi:10.1016/j.xgen.2021.100029
[BibTeX] [Abstract]
The Global Alliance for Genomics and Health (GA4GH) aims to accelerate biomedical advances by enabling the responsible sharing of clinical and genomic data through both harmonized data aggregation and federated approaches. The decreasing cost of genomic sequencing (along with other genome-wide molecular assays) and increasing evidence of its clinical utility will soon drive the generation of sequence data from tens of millions of humans, with increasing levels of diversity. In this perspective, we present the GA4GH strategies for addressing the major challenges of this data revolution. We describe the GA4GH organization, which is fueled by the development efforts of eight Work Streams and informed by the needs of 24 Driver Projects and other key stakeholders. We present the GA4GH suite of secure, interoperable technical standards and policy frameworks and review the current status of standards, their relevance to key domains of research and clinical care, and future plans of GA4GH. Broad international participation in building, adopting, and deploying GA4GH standards and frameworks will catalyze an unprecedented effort in data sharing that will be critical to advancing genomic medicine and ensuring that all populations can access its benefits.
@Article{35072136, author = {Rehm HL and Page AJH and Smith L and Adams JB and Alterovitz G and Babb LJ and Barkley MP and Baudis M and Beauvais MJS and Beck T and Beckmann JS and Beltran S and Bernick D and Bernier A and Bonfield JK and Boughtwood TF and Bourque G and Bowers SR and Brookes AJ and Brudno M and Brush MH and Bujold D and Burdett T and Buske OJ and Cabili MN and Cameron DL and Carroll RJ and Casas-Silva E and Chakravarty D and Chaudhari BP and Chen SH and Cherry JM and Chung J and Cline M and Clissold HL and Cook-Deegan RM and Courtot M and Cunningham F and Cupak M and Davies RM and Denisko D and Doerr MJ and Dolman LI and Dove ES and Dursi LJ and Dyke SOM and Eddy JA and Eilbeck K and Ellrott KP and Fairley S and Fakhro KA and Firth HV and Fitzsimons MS and Fiume M and Flicek P and Fore IM and Freeberg MA and Freimuth RR and Fromont LA and Fuerth J and Gaff CL and Gan W and Ghanaim EM and Glazer D and Green RC and Griffith M and Griffith OL and Grossman RL and Groza T and Auvil JMG and Guigó R and Gupta D and Haendel MA and Hamosh A and Hansen DP and Hart RK and Hartley DM and Haussler D and Hendricks-Sturrup RM and Ho CWL and Hobb AE and Hoffman MM and Hofmann OM and Holub P and Hsu JS and Hubaux JP and Hunt SE and Husami A and Jacobsen JO and Jamuar SS and Janes EL and Jeanson F and Jené A and Johns AL and Joly Y and Jones SJM and Kanitz A and Kato K and Keane TM and Kekesi-Lafrance K and Kelleher J and Kerry G and Khor SS and Knoppers BM and Konopko MA and Kosaki K and Kuba M and Lawson J and Leinonen R and Li S and Lin MF and Linden M and Liu X and Udara Liyanage I and Lopez J and Lucassen AM and Lukowski M and Mann AL and Marshall J and Mattioni M and Metke-Jimenez A and Middleton A and Milne RJ and Molnár-Gábor F and Mulder N and Munoz-Torres MC and Nag R and Nakagawa H and Nasir J and Navarro A and Nelson TH and Niewielska A and Nisselle A and Niu J and Nyrönen TH and O'Connor BD and Oesterle S and Ogishima S and Wang VO and Paglione LAD and Palumbo E and Parkinson HE and Philippakis AA and Pizarro AD and Prlic A and Rambla J and Rendon A and Rider RA and Robinson PN and Rodarmer KW and Rodriguez LL and Rubin AF and Rueda M and Rushton GA and Ryan RS and Saunders GI and Schuilenburg H and Schwede T and Scollen S and Senf A and Sheffield NC and Skantharajah N and Smith AV and Sofia HJ and Spalding D and Spurdle AB and Stark Z and Stein LD and Suematsu M and Tan P and Tedds JA and Thomson AA and Thorogood A and Tickle TL and Tokunaga K and Törnroos J and Torrents D and Upchurch S and Valencia A and Guimera RV and Vamathevan J and Varma S and Vears DF and Viner C and Voisin C and Wagner AH and Wallace SE and Walsh BP and Williams MS and Winkler EC and Wold BJ and Wood GM and Woolley JP and Yamasaki C and Yates AD and Yung CK and Zass LJ and Zaytseva K and Zhang J and Goodhand P and North K and Birney E}, title = {GA4GH: International policies and standards for data sharing across genomic research and healthcare}, journal = {Cell Genom}, volume = {1}, number = {2}, pages = {100029}, year = {2021}, doi = {10.1016/j.xgen.2021.100029}, abstract = {The Global Alliance for Genomics and Health (GA4GH) aims to accelerate biomedical advances by enabling the responsible sharing of clinical and genomic data through both harmonized data aggregation and federated approaches. The decreasing cost of genomic sequencing (along with other genome-wide molecular assays) and increasing evidence of its clinical utility will soon drive the generation of sequence data from tens of millions of humans, with increasing levels of diversity. In this perspective, we present the GA4GH strategies for addressing the major challenges of this data revolution. We describe the GA4GH organization, which is fueled by the development efforts of eight Work Streams and informed by the needs of 24 Driver Projects and other key stakeholders. We present the GA4GH suite of secure, interoperable technical standards and policy frameworks and review the current status of standards, their relevance to key domains of research and clinical care, and future plans of GA4GH. Broad international participation in building, adopting, and deploying GA4GH standards and frameworks will catalyze an unprecedented effort in data sharing that will be critical to advancing genomic medicine and ensuring that all populations can access its benefits.},} - A Frankish, M Diekhans, I Jungreis, J Lagarde, JE Loveland, JM Mudge, C Sisu, JC Wright, J Armstrong, I Barnes, A Berry, A Bignell, C Boix, S Carbonell Sala, F Cunningham, T Di Domenico, S Donaldson, IT Fiddes, C García Girón, JM Gonzalez, T Grego, M Hardy, T Hourlier, KL Howe, T Hunt, OG Izuogu, R Johnson, FJ Martin, L Martínez, S Mohanan, P Muir, FCP Navarro, A Parker, B Pei, F Pozo, FC Riera, M Ruffier, BM Schmitt, E Stapleton, MM Suner, I Sycheva, B Uszczynska-Ratajczak, MY Wolf, J Xu, YT Yang, A Yates, D Zerbino, Y Zhang, JS Choudhary, M Gerstein, R Guigó, TJP Hubbard, M Kellis, B Paten, ML Tress, P Flicek. GENCODE 2021. Nucleic Acids Res 2021;49(D1):D916–D923. doi:10.1093/nar/gkaa1087
[BibTeX] [Abstract]
The GENCODE project annotates human and mouse genes and transcripts supported by experimental data with high accuracy, providing a foundational resource that supports genome biology and clinical genomics. GENCODE annotation processes make use of primary data and bioinformatic tools and analysis generated both within the consortium and externally to support the creation of transcript structures and the determination of their function. Here, we present improvements to our annotation infrastructure, bioinformatics tools, and analysis, and the advances they support in the annotation of the human and mouse genomes including: the completion of first pass manual annotation for the mouse reference genome; targeted improvements to the annotation of genes associated with SARS-CoV-2 infection; collaborative projects to achieve convergence across reference annotation databases for the annotation of human and mouse protein-coding genes; and the first GENCODE manually supervised automated annotation of lncRNAs. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.
@Article{33270111, author = {Frankish A and Diekhans M and Jungreis I and Lagarde J and Loveland JE and Mudge JM and Sisu C and Wright JC and Armstrong J and Barnes I and Berry A and Bignell A and Boix C and Carbonell Sala S and Cunningham F and Di Domenico T and Donaldson S and Fiddes IT and García Girón C and Gonzalez JM and Grego T and Hardy M and Hourlier T and Howe KL and Hunt T and Izuogu OG and Johnson R and Martin FJ and Martínez L and Mohanan S and Muir P and Navarro FCP and Parker A and Pei B and Pozo F and Riera FC and Ruffier M and Schmitt BM and Stapleton E and Suner MM and Sycheva I and Uszczynska-Ratajczak B and Wolf MY and Xu J and Yang YT and Yates A and Zerbino D and Zhang Y and Choudhary JS and Gerstein M and Guigó R and Hubbard TJP and Kellis M and Paten B and Tress ML and Flicek P}, title = {GENCODE 2021}, journal = {Nucleic Acids Res}, volume = {49}, number = {D1}, pages = {D916--D923}, year = {2021}, doi = {10.1093/nar/gkaa1087}, howpublished = {Advanced online publication: 3 December 2020}, abstract = {The GENCODE project annotates human and mouse genes and transcripts supported by experimental data with high accuracy, providing a foundational resource that supports genome biology and clinical genomics. GENCODE annotation processes make use of primary data and bioinformatic tools and analysis generated both within the consortium and externally to support the creation of transcript structures and the determination of their function. Here, we present improvements to our annotation infrastructure, bioinformatics tools, and analysis, and the advances they support in the annotation of the human and mouse genomes including: the completion of first pass manual annotation for the mouse reference genome; targeted improvements to the annotation of genes associated with SARS-CoV-2 infection; collaborative projects to achieve convergence across reference annotation databases for the annotation of human and mouse protein-coding genes; and the first GENCODE manually supervised automated annotation of lncRNAs. Our annotation is accessible via Ensembl, the UCSC Genome Browser and https://www.gencodegenes.org.},} - S Watt, L Vasquez, K Walter, AL Mann, K Kundu, L Chen, Y Sims, S Ecker, F Burden, S Farrow, B Farr, V Iotchkova, H Elding, D Mead, M Tardaguila, H Ponstingl, D Richardson, A Datta, P Flicek, L Clarke, K Downes, T Pastinen, P Fraser, M Frontini, BM Javierre, M Spivakov, N Soranzo. Genetic perturbation of PU.1 binding and chromatin looping at neutrophil enhancers associates with autoimmune disease. Nat Commun 2021;12(1):2298. doi:10.1038/s41467-021-22548-8 10.1016/j.chom.2014.04.011 10.1038/ni.2921 10.3389/fcimb.2017.00373 10.4049/jimmunol.1201719 10.1056/NEJM198902093200606 10.1378/chest.121.5_suppl.151S 10.1111/eci.12983 10.1038/nri3024 10.1126/science.1222794 10.1038/nature13835 10.1016/j.cell.2015.08.001 10.1016/j.cell.2015.07.048 10.1016/j.cell.2016.03.041 10.1016/j.cell.2016.10.026 10.1038/nature24277 10.1016/j.cell.2016.09.037 10.1038/ng.3963 10.1002/j.1460-2075.1996.tb00949.x 10.1182/blood-2016-07-730135 10.1016/j.molcel.2010.05.004 10.1101/gad.176826.111 10.1182/blood-2002-03-0835 10.1038/s41590-019-0343-z 10.1186/s13024-018-0277-1 10.1186/s13059-017-1156-8 10.1038/ng.3467 10.1038/s41467-019-13960-2 10.1016/j.immuni.2018.04.024 10.1016/j.cell.2013.02.029 10.1126/science.aat8266 10.1073/pnas.1530509100 10.1084/jem.20041535 10.1186/s13059-016-0992-2 10.7554/eLife.21926 10.1053/j.gastro.2017.04.002 10.1038/ng.3286 10.1101/gad.293910.116 10.1016/j.cell.2016.10.042 10.1038/s41588-018-0322-6 10.1038/ng.543 10.1038/ng.3359 10.1038/ng.2770 10.1038/ng.3245 10.1038/srep22223 10.1038/ng.2383 10.1038/ng.3570 10.1038/ng.3432 10.1016/j.tibs.2007.04.004 10.1038/ncomms13507 10.1183/13993003.00970-2017 10.1038/ng.717 10.1038/ncomms16021 10.1038/ng.2686 10.1074/jbc.M513471200 10.1038/ng.2504 10.1126/science.1242463 10.1186/s12864-018-4957-6 10.1038/s41588-020-00745-3 10.1038/s41467-019-08940-5 10.1186/s13059-019-1855-4 10.1101/gad.12.15.2403 10.1002/art.37885 10.1016/j.cell.2013.07.007 10.1016/j.cell.2018.04.018 10.1038/nrrheum.2014.80 10.1038/nri2779 10.1186/gb-2013-14-11-r124 10.1186/gb-2008-9-9-r137 10.1101/gr.136184.111 10.1093/nar/gkw257 10.1101/gr.155192.113 10.1038/nprot.2015.127 10.1093/bioinformatics/btp352 10.1038/s41586-018-0321-x 10.1038/nmeth.3582 10.1086/519795 10.1371/journal.pgen.1004383
[BibTeX] [Abstract]
Neutrophils play fundamental roles in innate immune response, shape adaptive immunity, and are a potentially causal cell type underpinning genetic associations with immune system traits and diseases. Here, we profile the binding of myeloid master regulator PU.1 in primary neutrophils across nearly a hundred volunteers. We show that variants associated with differential PU.1 binding underlie genetically-driven differences in cell count and susceptibility to autoimmune and inflammatory diseases. We integrate these results with other multi-individual genomic readouts, revealing coordinated effects of PU.1 binding variants on the local chromatin state, enhancer-promoter contacts and downstream gene expression, and providing a functional interpretation for 27 genes underlying immune traits. Collectively, these results demonstrate the functional role of PU.1 and its target enhancers in neutrophil transcriptional control and immune disease susceptibility.
@Article{33863903, author = {Watt S and Vasquez L and Walter K and Mann AL and Kundu K and Chen L and Sims Y and Ecker S and Burden F and Farrow S and Farr B and Iotchkova V and Elding H and Mead D and Tardaguila M and Ponstingl H and Richardson D and Datta A and Flicek P and Clarke L and Downes K and Pastinen T and Fraser P and Frontini M and Javierre BM and Spivakov M and Soranzo N}, title = {Genetic perturbation of PU.1 binding and chromatin looping at neutrophil enhancers associates with autoimmune disease}, journal = {Nat Commun}, volume = {12}, number = {1}, pages = {2298}, year = {2021}, doi = {10.1038/s41467-021-22548-8 10.1016/j.chom.2014.04.011 10.1038/ni.2921 10.3389/fcimb.2017.00373 10.4049/jimmunol.1201719 10.1056/NEJM198902093200606 10.1378/chest.121.5_suppl.151S 10.1111/eci.12983 10.1038/nri3024 10.1126/science.1222794 10.1038/nature13835 10.1016/j.cell.2015.08.001 10.1016/j.cell.2015.07.048 10.1016/j.cell.2016.03.041 10.1016/j.cell.2016.10.026 10.1038/nature24277 10.1016/j.cell.2016.09.037 10.1038/ng.3963 10.1002/j.1460-2075.1996.tb00949.x 10.1182/blood-2016-07-730135 10.1016/j.molcel.2010.05.004 10.1101/gad.176826.111 10.1182/blood-2002-03-0835 10.1038/s41590-019-0343-z 10.1186/s13024-018-0277-1 10.1186/s13059-017-1156-8 10.1038/ng.3467 10.1038/s41467-019-13960-2 10.1016/j.immuni.2018.04.024 10.1016/j.cell.2013.02.029 10.1126/science.aat8266 10.1073/pnas.1530509100 10.1084/jem.20041535 10.1186/s13059-016-0992-2 10.7554/eLife.21926 10.1053/j.gastro.2017.04.002 10.1038/ng.3286 10.1101/gad.293910.116 10.1016/j.cell.2016.10.042 10.1038/s41588-018-0322-6 10.1038/ng.543 10.1038/ng.3359 10.1038/ng.2770 10.1038/ng.3245 10.1038/srep22223 10.1038/ng.2383 10.1038/ng.3570 10.1038/ng.3432 10.1016/j.tibs.2007.04.004 10.1038/ncomms13507 10.1183/13993003.00970-2017 10.1038/ng.717 10.1038/ncomms16021 10.1038/ng.2686 10.1074/jbc.M513471200 10.1038/ng.2504 10.1126/science.1242463 10.1186/s12864-018-4957-6 10.1038/s41588-020-00745-3 10.1038/s41467-019-08940-5 10.1186/s13059-019-1855-4 10.1101/gad.12.15.2403 10.1002/art.37885 10.1016/j.cell.2013.07.007 10.1016/j.cell.2018.04.018 10.1038/nrrheum.2014.80 10.1038/nri2779 10.1186/gb-2013-14-11-r124 10.1186/gb-2008-9-9-r137 10.1101/gr.136184.111 10.1093/nar/gkw257 10.1101/gr.155192.113 10.1038/nprot.2015.127 10.1093/bioinformatics/btp352 10.1038/s41586-018-0321-x 10.1038/nmeth.3582 10.1086/519795 10.1371/journal.pgen.1004383}, note = {First posted as a preprint: 29 April 2019}, abstract = {Neutrophils play fundamental roles in innate immune response, shape adaptive immunity, and are a potentially causal cell type underpinning genetic associations with immune system traits and diseases. Here, we profile the binding of myeloid master regulator PU.1 in primary neutrophils across nearly a hundred volunteers. We show that variants associated with differential PU.1 binding underlie genetically-driven differences in cell count and susceptibility to autoimmune and inflammatory diseases. We integrate these results with other multi-individual genomic readouts, revealing coordinated effects of PU.1 binding variants on the local chromatin state, enhancer-promoter contacts and downstream gene expression, and providing a functional interpretation for 27 genes underlying immune traits. Collectively, these results demonstrate the functional role of PU.1 and its target enhancers in neutrophil transcriptional control and immune disease susceptibility.},} - MK Tello-Ruiz, S Naithani, P Gupta, A Olson, S Wei, J Preece, Y Jiao, B Wang, K Chougule, P Garg, J Elser, S Kumari, V Kumar, B Contreras-Moreira, G Naamati, N George, J Cook, D Bolser, P D’Eustachio, LD Stein, A Gupta, W Xu, J Regala, I Papatheodorou, PJ Kersey, P Flicek, C Taylor, P Jaiswal, D Ware. Gramene 2021: harnessing the power of comparative genomics and pathways for plant research. Nucleic Acids Res 2021;49(D1):D1452–D1463. doi:10.1093/nar/gkaa979
[BibTeX] [Abstract]
Gramene (http://www.gramene.org), a knowledgebase founded on comparative functional analyses of genomic and pathway data for model plants and major crops, supports agricultural researchers worldwide. The resource is committed to open access and reproducible science based on the FAIR data principles. Since the last NAR update, we made nine releases; doubled the genome portal’s content; expanded curated genes, pathways and expression sets; and implemented the Domain Informational Vocabulary Extraction (DIVE) algorithm for extracting gene function information from publications. The current release, \#63 (October 2020), hosts 93 reference genomes-over 3.9 million genes in 122 947 families with orthologous and paralogous classifications. Plant Reactome portrays pathway networks using a combination of manual biocuration in rice (320 reference pathways) and orthology-based projections to 106 species. The Reactome platform facilitates comparison between reference and projected pathways, gene expression analyses and overlays of gene-gene interactions. Gramene integrates ontology-based protein structure-function annotation; information on genetic, epigenetic, expression, and phenotypic diversity; and gene functional annotations extracted from plant-focused journals using DIVE. We train plant researchers in biocuration of genes and pathways; host curated maize gene structures as tracks in the maize genome browser; and integrate curated rice genes and pathways in the Plant Reactome.
@Article{33170273, author = {Tello-Ruiz MK and Naithani S and Gupta P and Olson A and Wei S and Preece J and Jiao Y and Wang B and Chougule K and Garg P and Elser J and Kumari S and Kumar V and Contreras-Moreira B and Naamati G and George N and Cook J and Bolser D and D'Eustachio P and Stein LD and Gupta A and Xu W and Regala J and Papatheodorou I and Kersey PJ and Flicek P and Taylor C and Jaiswal P and Ware D}, title = {Gramene 2021: harnessing the power of comparative genomics and pathways for plant research}, journal = {Nucleic Acids Res}, volume = {49}, number = {D1}, pages = {D1452--D1463}, year = {2021}, doi = {10.1093/nar/gkaa979}, howpublished = {Advanced online publication: 10 November 2020}, abstract = {Gramene (http://www.gramene.org), a knowledgebase founded on comparative functional analyses of genomic and pathway data for model plants and major crops, supports agricultural researchers worldwide. The resource is committed to open access and reproducible science based on the FAIR data principles. Since the last NAR update, we made nine releases; doubled the genome portal's content; expanded curated genes, pathways and expression sets; and implemented the Domain Informational Vocabulary Extraction (DIVE) algorithm for extracting gene function information from publications. The current release, \#63 (October 2020), hosts 93 reference genomes-over 3.9 million genes in 122 947 families with orthologous and paralogous classifications. Plant Reactome portrays pathway networks using a combination of manual biocuration in rice (320 reference pathways) and orthology-based projections to 106 species. The Reactome platform facilitates comparison between reference and projected pathways, gene expression analyses and overlays of gene-gene interactions. Gramene integrates ontology-based protein structure-function annotation; information on genetic, epigenetic, expression, and phenotypic diversity; and gene functional annotations extracted from plant-focused journals using DIVE. We train plant researchers in biocuration of genes and pathways; host curated maize gene structures as tracks in the maize genome browser; and integrate curated rice genes and pathways in the Plant Reactome.},} - P Ebert, PA Audano, Q Zhu, B Rodriguez-Martin, D Porubsky, MJ Bonder, A Sulovari, J Ebler, W Zhou, R Serra Mari, F Yilmaz, X Zhao, P Hsieh, J Lee, S Kumar, J Lin, T Rausch, Y Chen, J Ren, M Santamarina, W Höps, H Ashraf, NT Chuang, X Yang, KM Munson, AP Lewis, S Fairley, LJ Tallon, WE Clarke, AO Basile, M Byrska-Bishop, A Corvelo, US Evani, TY Lu, MJP Chaisson, J Chen, C Li, H Brand, AM Wenger, M Ghareghani, WT Harvey, B Raeder, P Hasenfeld, AA Regier, HJ Abel, IM Hall, P Flicek, O Stegle, MB Gerstein, JMC Tubio, Z Mu, YI Li, X Shi, AR Hastie, K Ye, Z Chong, AD Sanders, MC Zody, ME Talkowski, RE Mills, SE Devine, C Lee, JO Korbel, T Marschall, EE Eichler. Haplotype-resolved diverse human genomes and integrated analysis of structural variation. Science 2021;372(6537). doi:10.1126/science.abf7117
[BibTeX] [Abstract]
Long-read and strand-specific sequencing technologies together facilitate the de novo assembly of high-quality haplotype-resolved human genomes without parent-child trio data. We present 64 assembled haplotypes from 32 diverse human genomes. These highly contiguous haplotype assemblies (average contig N50: 26 Mbp) integrate all forms of genetic variation even across complex loci. We identify 107,590 structural variants (SVs), of which 68\% are not discovered by short-read sequencing, and 278 SV hotspots (spanning megabases of gene-rich sequence). We characterize 130 of the most active mobile element source elements and find that 63\% of all SVs arise by homology-mediated mechanisms. This resource enables reliable graph-based genotyping from short reads of up to 50,340 SVs, resulting in the identification of 1,526 expression quantitative trait loci as well as SV candidates for adaptive selection within the human population.
@Article{33632895, author = {Ebert P and Audano PA and Zhu Q and Rodriguez-Martin B and Porubsky D and Bonder MJ and Sulovari A and Ebler J and Zhou W and Serra Mari R and Yilmaz F and Zhao X and Hsieh P and Lee J and Kumar S and Lin J and Rausch T and Chen Y and Ren J and Santamarina M and Höps W and Ashraf H and Chuang NT and Yang X and Munson KM and Lewis AP and Fairley S and Tallon LJ and Clarke WE and Basile AO and Byrska-Bishop M and Corvelo A and Evani US and Lu TY and Chaisson MJP and Chen J and Li C and Brand H and Wenger AM and Ghareghani M and Harvey WT and Raeder B and Hasenfeld P and Regier AA and Abel HJ and Hall IM and Flicek P and Stegle O and Gerstein MB and Tubio JMC and Mu Z and Li YI and Shi X and Hastie AR and Ye K and Chong Z and Sanders AD and Zody MC and Talkowski ME and Mills RE and Devine SE and Lee C and Korbel JO and Marschall T and Eichler EE}, title = {Haplotype-resolved diverse human genomes and integrated analysis of structural variation}, journal = {Science}, volume = {372}, number = {6537}, year = {2021}, doi = {10.1126/science.abf7117}, howpublished = {Advanced online publication: 25 February 2021}, note = {First posted as a preprint: 16 December 2020}, abstract = {Long-read and strand-specific sequencing technologies together facilitate the de novo assembly of high-quality haplotype-resolved human genomes without parent-child trio data. We present 64 assembled haplotypes from 32 diverse human genomes. These highly contiguous haplotype assemblies (average contig N50: 26 Mbp) integrate all forms of genetic variation even across complex loci. We identify 107,590 structural variants (SVs), of which 68\% are not discovered by short-read sequencing, and 278 SV hotspots (spanning megabases of gene-rich sequence). We characterize 130 of the most active mobile element source elements and find that 63\% of all SVs arise by homology-mediated mechanisms. This resource enables reliable graph-based genotyping from short reads of up to 50,340 SVs, resulting in the identification of 1,526 expression quantitative trait loci as well as SV candidates for adaptive selection within the human population.},} - B Contreras-Moreira, CV Filippi, G Naamati, C García Girón, JE Allen, P Flicek. K-mer counting and curated libraries drive efficient annotation of repeats in plant genomes. Plant Genome 2021;14(3):e20143. doi:10.1002/tpg2.20143
[BibTeX] [Abstract]
The annotation of repetitive sequences within plant genomes can help in the interpretation of observed phenotypes. Moreover, repeat masking is required for tasks such as whole-genome alignment, promoter analysis, or pangenome exploration. Although homology-based annotation methods are computationally expensive, k-mer strategies for masking are orders of magnitude faster. Here, we benchmarked a two-step approach, where repeats were first called by k-mer counting and then annotated by comparison to curated libraries. This hybrid protocol was tested on 20 plant genomes from Ensembl, with the k-mer-based Repeat Detector (Red) and two repeat libraries (REdat, last updated in 2013, and nrTEplants, curated for this work). Custom libraries produced by RepeatModeler were also tested. We obtained repeated genome fractions that matched those reported in the literature but with shorter repeated elements than those produced directly by sequence homology. Inspection of the masked regions that overlapped genes revealed no preference for specific protein domains. Most Red-masked sequences could be successfully classified by sequence similarity, with the complete protocol taking less than 2 h on a desktop Linux box. A guide to curating your own repeat libraries and the scripts for masking and annotating plant genomes can be obtained at https://github.com/Ensembl/plant-scripts.
@Article{34562304, author = {Contreras-Moreira B and Filippi CV and Naamati G and García Girón C and Allen JE and Flicek P}, title = {K-mer counting and curated libraries drive efficient annotation of repeats in plant genomes}, journal = {Plant Genome}, volume = {14}, number = {3}, pages = {e20143}, year = {2021}, doi = {10.1002/tpg2.20143}, howpublished = {Advanced online publication: 25 September 2021}, note = {First posted as a preprint: 24 March 2021}, abstract = {The annotation of repetitive sequences within plant genomes can help in the interpretation of observed phenotypes. Moreover, repeat masking is required for tasks such as whole-genome alignment, promoter analysis, or pangenome exploration. Although homology-based annotation methods are computationally expensive, k-mer strategies for masking are orders of magnitude faster. Here, we benchmarked a two-step approach, where repeats were first called by k-mer counting and then annotated by comparison to curated libraries. This hybrid protocol was tested on 20 plant genomes from Ensembl, with the k-mer-based Repeat Detector (Red) and two repeat libraries (REdat, last updated in 2013, and nrTEplants, curated for this work). Custom libraries produced by RepeatModeler were also tested. We obtained repeated genome fractions that matched those reported in the literature but with shorter repeated elements than those produced directly by sequence homology. Inspection of the masked regions that overlapped genes revealed no preference for specific protein domains. Most Red-masked sequences could be successfully classified by sequence similarity, with the complete protocol taking less than 2 h on a desktop Linux box. A guide to curating your own repeat libraries and the scripts for masking and annotating plant genomes can be obtained at https://github.com/Ensembl/plant-scripts.},} - M Roller, E Stamper, D Villar, O Izuogu, F Martin, AM Redmond, R Ramachanderan, L Harewood, DT Odom, P Flicek. LINE retrotransposons characterize mammalian tissue-specific and evolutionarily dynamic regulatory regions. Genome Biol 2021;22(1):62. doi:10.1186/s13059-021-02260-y 10.1038/s41576-019-0173-8 10.1038/nature14217 10.1038/ng.3884 10.1038/s41588-019-0494-8 10.1038/nature12787 10.1371/journal.pbio.1000384 10.1038/nature09033 10.1016/j.molcel.2011.12.021 10.1038/s41467-018-06544-z 10.1038/nature10532 10.1038/s41586-019-1338-5 10.1126/science.1230612 10.1016/j.cell.2015.01.006 10.1093/gbe/evz134 10.1016/j.scr.2019.101456 10.1038/s41559-017-0447-5 10.1038/nature13985 10.1126/science.1246426 10.1016/j.cels.2018.01.002 10.1038/nature13182 10.1101/gr.190546.115 10.1186/s13059-018-1577-z 10.1371/journal.pgen.1003504 10.1101/gr.216150.116 10.1186/s12864-018-4850-3 10.1101/gr.218149.116 10.1101/gr.235747.118 10.1007/s10577-017-9570-z 10.1093/gbe/evv005 10.1126/science.aac7247 10.1093/nar/gkq132 10.1093/molbev/msy143 10.7554/eLife.13926 10.1093/nar/gkw1067 10.1074/jbc.273.2.891 10.1371/journal.pgen.1008036 10.1371/journal.pone.0027513 10.1093/nar/gkz1138 10.1038/nature11243 10.1016/j.cell.2005.01.001 10.1073/pnas.1016071107 10.1016/j.molcel.2013.01.038 10.1038/s41576-019-0128-0 10.1016/j.stem.2015.02.013 10.1073/pnas.1507125112 10.1038/nn.4229 10.1016/j.celrep.2013.05.031 10.1016/j.cell.2019.12.015 10.1038/nature07829 10.1038/ng.3286 10.1186/s13059-018-1432-2 10.1093/gbe/evx194 10.1101/gr.234096.117 10.1093/oxfordjournals.molbev.a003768 10.1038/nrg3802 10.1093/molbev/msx219 10.1016/j.ymeth.2009.03.001 10.1186/gb-2013-14-11-r124 10.1093/bioinformatics/btp324 10.1093/bioinformatics/btp352 10.1101/gr.136184.111 10.1186/gb-2008-9-9-r137 10.1038/ng1966 10.1038/nature09692 10.1038/nature01080 10.1038/ncb1076 10.1038/nature03877 10.1038/ng.154 10.1093/nar/gky955 10.1093/bioinformatics/bty191 10.1093/nar/gkg006 10.1093/nar/gkj112 10.1016/S0022-2836(05)80360-2 10.1186/1748-7188-6-26 10.1093/bioinformatics/btt509 10.1093/bioinformatics/btu170 10.1093/bioinformatics/bts635 10.1038/nbt.1621 10.1093/nar/gkv1189 10.1093/nar/gkw257 10.1093/bioinformatics/btt737 10.1093/bib/bbs017 10.1093/bioinformatics/btx364 10.1186/s13059-014-0550-8 10.1007/978-3-319-24277-4 10.1101/gr.09
[BibTeX] [Abstract]
BACKGROUND: To investigate the mechanisms driving regulatory evolution across tissues, we experimentally mapped promoters, enhancers, and gene expression in the liver, brain, muscle, and testis from ten diverse mammals. RESULTS: The regulatory landscape around genes included both tissue-shared and tissue-specific regulatory regions, where tissue-specific promoters and enhancers evolved most rapidly. Genomic regions switching between promoters and enhancers were more common across species, and less common across tissues within a single species. Long Interspersed Nuclear Elements (LINEs) played recurrent evolutionary roles: LINE L1s were associated with tissue-specific regulatory regions, whereas more ancient LINE L2s were associated with tissue-shared regulatory regions and with those switching between promoter and enhancer signatures across species. CONCLUSIONS: Our analyses of the tissue-specificity and evolutionary stability among promoters and enhancers reveal how specific LINE families have helped shape the dynamic mammalian regulome.
@Article{33602314, author = {Roller M and Stamper E and Villar D and Izuogu O and Martin F and Redmond AM and Ramachanderan R and Harewood L and Odom DT and Flicek P}, title = {LINE retrotransposons characterize mammalian tissue-specific and evolutionarily dynamic regulatory regions}, journal = {Genome Biol}, volume = {22}, number = {1}, pages = {62}, year = {2021}, doi = {10.1186/s13059-021-02260-y 10.1038/s41576-019-0173-8 10.1038/nature14217 10.1038/ng.3884 10.1038/s41588-019-0494-8 10.1038/nature12787 10.1371/journal.pbio.1000384 10.1038/nature09033 10.1016/j.molcel.2011.12.021 10.1038/s41467-018-06544-z 10.1038/nature10532 10.1038/s41586-019-1338-5 10.1126/science.1230612 10.1016/j.cell.2015.01.006 10.1093/gbe/evz134 10.1016/j.scr.2019.101456 10.1038/s41559-017-0447-5 10.1038/nature13985 10.1126/science.1246426 10.1016/j.cels.2018.01.002 10.1038/nature13182 10.1101/gr.190546.115 10.1186/s13059-018-1577-z 10.1371/journal.pgen.1003504 10.1101/gr.216150.116 10.1186/s12864-018-4850-3 10.1101/gr.218149.116 10.1101/gr.235747.118 10.1007/s10577-017-9570-z 10.1093/gbe/evv005 10.1126/science.aac7247 10.1093/nar/gkq132 10.1093/molbev/msy143 10.7554/eLife.13926 10.1093/nar/gkw1067 10.1074/jbc.273.2.891 10.1371/journal.pgen.1008036 10.1371/journal.pone.0027513 10.1093/nar/gkz1138 10.1038/nature11243 10.1016/j.cell.2005.01.001 10.1073/pnas.1016071107 10.1016/j.molcel.2013.01.038 10.1038/s41576-019-0128-0 10.1016/j.stem.2015.02.013 10.1073/pnas.1507125112 10.1038/nn.4229 10.1016/j.celrep.2013.05.031 10.1016/j.cell.2019.12.015 10.1038/nature07829 10.1038/ng.3286 10.1186/s13059-018-1432-2 10.1093/gbe/evx194 10.1101/gr.234096.117 10.1093/oxfordjournals.molbev.a003768 10.1038/nrg3802 10.1093/molbev/msx219 10.1016/j.ymeth.2009.03.001 10.1186/gb-2013-14-11-r124 10.1093/bioinformatics/btp324 10.1093/bioinformatics/btp352 10.1101/gr.136184.111 10.1186/gb-2008-9-9-r137 10.1038/ng1966 10.1038/nature09692 10.1038/nature01080 10.1038/ncb1076 10.1038/nature03877 10.1038/ng.154 10.1093/nar/gky955 10.1093/bioinformatics/bty191 10.1093/nar/gkg006 10.1093/nar/gkj112 10.1016/S0022-2836(05)80360-2 10.1186/1748-7188-6-26 10.1093/bioinformatics/btt509 10.1093/bioinformatics/btu170 10.1093/bioinformatics/bts635 10.1038/nbt.1621 10.1093/nar/gkv1189 10.1093/nar/gkw257 10.1093/bioinformatics/btt737 10.1093/bib/bbs017 10.1093/bioinformatics/btx364 10.1186/s13059-014-0550-8 10.1007/978-3-319-24277-4 10.1101/gr.09}, note = {First posted as a preprint: 31 May 2020}, abstract = {BACKGROUND: To investigate the mechanisms driving regulatory evolution across tissues, we experimentally mapped promoters, enhancers, and gene expression in the liver, brain, muscle, and testis from ten diverse mammals. RESULTS: The regulatory landscape around genes included both tissue-shared and tissue-specific regulatory regions, where tissue-specific promoters and enhancers evolved most rapidly. Genomic regions switching between promoters and enhancers were more common across species, and less common across tissues within a single species. Long Interspersed Nuclear Elements (LINEs) played recurrent evolutionary roles: LINE L1s were associated with tissue-specific regulatory regions, whereas more ancient LINE L2s were associated with tissue-shared regulatory regions and with those switching between promoter and enhancer signatures across species. CONCLUSIONS: Our analyses of the tissue-specificity and evolutionary stability among promoters and enhancers reveal how specific LINE families have helped shape the dynamic mammalian regulome.},} - G Cantelli, G Cochrane, C Brooksbank, E McDonagh, P Flicek, J McEntyre, E Birney, R Apweiler. The European Bioinformatics Institute: empowering cooperation in response to a global health crisis. Nucleic Acids Res 2021;49(D1):D29–D37. doi:10.1093/nar/gkaa1077
[BibTeX] [Abstract]
The European Bioinformatics Institute (EMBL-EBI; https://www.ebi.ac.uk/) provides freely available data and bioinformatics services to the scientific community, alongside its research activity and training provision. The 2020 COVID-19 pandemic has brought to the forefront a need for the scientific community to work even more cooperatively to effectively tackle a global health crisis. EMBL-EBI has been able to build on its position to contribute to the fight against COVID-19 in a number of ways. Firstly, EMBL-EBI has used its infrastructure, expertise and network of international collaborations to help build the European COVID-19 Data Platform (https://www.covid19dataportal.org/), which brings together COVID-19 biomolecular data and connects it to researchers, clinicians and public health professionals. By September 2020, the COVID-19 Data Platform has integrated in excess of 170 000 COVID-19 biomolecular data and literature records, collected through a number of EMBL-EBI resources. Secondly, EMBL-EBI has strived to continue its support of the life science communities through the crisis, with updated Training provision and improved service provision throughout its resources. The COVID-19 pandemic has highlighted the importance of EMBL-EBI’s core principles, including international cooperation, resource sharing and central data brokering, and has further empowered scientific cooperation.
@Article{33245775, author = {Cantelli G and Cochrane G and Brooksbank C and McDonagh E and Flicek P and McEntyre J and Birney E and Apweiler R}, title = {The European Bioinformatics Institute: empowering cooperation in response to a global health crisis}, journal = {Nucleic Acids Res}, volume = {49}, number = {D1}, pages = {D29--D37}, year = {2021}, doi = {10.1093/nar/gkaa1077}, howpublished = {Advanced online publication: 27 November 2020}, abstract = {The European Bioinformatics Institute (EMBL-EBI; https://www.ebi.ac.uk/) provides freely available data and bioinformatics services to the scientific community, alongside its research activity and training provision. The 2020 COVID-19 pandemic has brought to the forefront a need for the scientific community to work even more cooperatively to effectively tackle a global health crisis. EMBL-EBI has been able to build on its position to contribute to the fight against COVID-19 in a number of ways. Firstly, EMBL-EBI has used its infrastructure, expertise and network of international collaborations to help build the European COVID-19 Data Platform (https://www.covid19dataportal.org/), which brings together COVID-19 biomolecular data and connects it to researchers, clinicians and public health professionals. By September 2020, the COVID-19 Data Platform has integrated in excess of 170 000 COVID-19 biomolecular data and literature records, collected through a number of EMBL-EBI resources. Secondly, EMBL-EBI has strived to continue its support of the life science communities through the crisis, with updated Training provision and improved service provision throughout its resources. The COVID-19 pandemic has highlighted the importance of EMBL-EBI's core principles, including international cooperation, resource sharing and central data brokering, and has further empowered scientific cooperation.},} - PW Harrison, A Sokolov, A Nayak, J Fan, D Zerbino, G Cochrane, P Flicek. The FAANG Data Portal: Global, Open-Access, “FAIR”, and Richly Validated Genotype to Phenotype Data for High-Quality Functional Annotation of Animal Genomes. Front Genet 2021;12:639238. doi:10.3389/fgene.2021.639238
[BibTeX] [Abstract]
The Functional Annotation of ANimal Genomes (FAANG) project is a worldwide coordinated action creating high-quality functional annotation of farmed and companion animal genomes. The generation of a rich genome-to-phenome resource and supporting informatic infrastructure advances the scope of comparative genomics and furthers the understanding of functional elements. The project also provides terrestrial and aquatic animal agriculture community powerful resources for supporting improvements to farmed animal production, disease resistance, and genetic diversity. The FAANG Data Portal (https://data.faang.org) ensures Findable, Accessible, Interoperable and Reusable (FAIR) open access to the wealth of sample, sequencing, and analysis data produced by an ever-growing number of FAANG consortia. It is developed and maintained by the FAANG Data Coordination Centre (DCC) at the European Molecular Biology Laboratory’s European Bioinformatics Institute (EMBL-EBI). FAANG projects produce a standardised set of multi-omic assays with resulting data placed into a range of specialised open data archives. To ensure this data is easily findable and accessible by the community, the portal automatically identifies and collates all submitted FAANG data into a single easily searchable resource. The Data Portal supports direct download from the multiple underlying archives to enable seamless access to all FAANG data from within the portal itself. The portal provides a range of predefined filters, powerful predictive search, and a catalogue of sampling and analysis protocols and automatically identifies publications associated with any dataset. To ensure all FAANG data submissions are high-quality, the portal includes powerful contextual metadata validation and data submissions brokering to the underlying EMBL-EBI archives. The portal will incorporate extensive new technical infrastructure to effectively deliver and standardise FAANG’s shift to single-cellomics, cell atlases, pangenomes, and novel phenotypic prediction models. The Data Portal plays a key role for FAANG by supporting high-quality functional annotation of animal genomes, through open FAIR sharing of data, complete with standardised rich metadata. Future Data Portal features developed by the DCC will support new technological developments for continued improvement for FAANG projects.
@Article{34220930, author = {Harrison PW and Sokolov A and Nayak A and Fan J and Zerbino D and Cochrane G and Flicek P}, title = {The FAANG Data Portal: Global, Open-Access, ``FAIR'', and Richly Validated Genotype to Phenotype Data for High-Quality Functional Annotation of Animal Genomes}, journal = {Front Genet}, volume = {12}, pages = {639238}, year = {2021}, doi = {10.3389/fgene.2021.639238}, abstract = {The Functional Annotation of ANimal Genomes (FAANG) project is a worldwide coordinated action creating high-quality functional annotation of farmed and companion animal genomes. The generation of a rich genome-to-phenome resource and supporting informatic infrastructure advances the scope of comparative genomics and furthers the understanding of functional elements. The project also provides terrestrial and aquatic animal agriculture community powerful resources for supporting improvements to farmed animal production, disease resistance, and genetic diversity. The FAANG Data Portal (https://data.faang.org) ensures Findable, Accessible, Interoperable and Reusable (FAIR) open access to the wealth of sample, sequencing, and analysis data produced by an ever-growing number of FAANG consortia. It is developed and maintained by the FAANG Data Coordination Centre (DCC) at the European Molecular Biology Laboratory's European Bioinformatics Institute (EMBL-EBI). FAANG projects produce a standardised set of multi-omic assays with resulting data placed into a range of specialised open data archives. To ensure this data is easily findable and accessible by the community, the portal automatically identifies and collates all submitted FAANG data into a single easily searchable resource. The Data Portal supports direct download from the multiple underlying archives to enable seamless access to all FAANG data from within the portal itself. The portal provides a range of predefined filters, powerful predictive search, and a catalogue of sampling and analysis protocols and automatically identifies publications associated with any dataset. To ensure all FAANG data submissions are high-quality, the portal includes powerful contextual metadata validation and data submissions brokering to the underlying EMBL-EBI archives. The portal will incorporate extensive new technical infrastructure to effectively deliver and standardise FAANG's shift to single-cellomics, cell atlases, pangenomes, and novel phenotypic prediction models. The Data Portal plays a key role for FAANG by supporting high-quality functional annotation of animal genomes, through open FAIR sharing of data, complete with standardised rich metadata. Future Data Portal features developed by the DCC will support new technological developments for continued improvement for FAANG projects.},} - J Morales, AC McMahon, J Loveland, E Perry, A Frankish, S Hunt, IM Armean, P Flicek, F Cunningham. The value of primary transcripts to the clinical and non-clinical genomics community: Survey results and roadmap for improvements. Mol Genet Genomic Med 2021;9(12):e1786. doi:10.1002/mgg3.1786
[BibTeX] [Abstract]
BACKGROUND: Variant interpretation is dependent on transcript annotation and remains time consuming and challenging. There are major obstacles for historical data reuse and for interpretation of new variants. First, both RefSeq and Ensembl/GENCODE produce transcript sets in common use, but there is currently no easy way to translate between the two. Second, the resources often used for variant interpretation (e.g. ClinVar, gnomAD, UniProt) do not use the same transcript set, nor default transcript or protein sequence. METHOD: Ensembl ran a survey in 2018 to sample attitudes to choosing one default transcript per locus, and to gather data on reference sequences used by the scientific community. This was publicised on the Ensembl and UCSC genome browsers, by email and on social media. RESULTS: The survey had 788 responses from 32 different countries, the results of which we report here. CONCLUSIONS: We present our roadmap to create an effective default set of transcripts for resources, and for reporting interpretation of clinical variants.
@Article{34435752, author = {Morales J and McMahon AC and Loveland J and Perry E and Frankish A and Hunt S and Armean IM and Flicek P and Cunningham F}, title = {The value of primary transcripts to the clinical and non-clinical genomics community: Survey results and roadmap for improvements}, journal = {Mol Genet Genomic Med}, volume = {9}, number = {12}, pages = {e1786}, year = {2021}, doi = {10.1002/mgg3.1786}, howpublished = {Advanced online publication: 26 August 2021}, note = {First posted as a preprint: 26 March 2021}, abstract = {BACKGROUND: Variant interpretation is dependent on transcript annotation and remains time consuming and challenging. There are major obstacles for historical data reuse and for interpretation of new variants. First, both RefSeq and Ensembl/GENCODE produce transcript sets in common use, but there is currently no easy way to translate between the two. Second, the resources often used for variant interpretation (e.g. ClinVar, gnomAD, UniProt) do not use the same transcript set, nor default transcript or protein sequence. METHOD: Ensembl ran a survey in 2018 to sample attitudes to choosing one default transcript per locus, and to gather data on reference sequences used by the scientific community. This was publicised on the Ensembl and UCSC genome browsers, by email and on social media. RESULTS: The survey had 788 responses from 32 different countries, the results of which we report here. CONCLUSIONS: We present our roadmap to create an effective default set of transcripts for resources, and for reporting interpretation of clinical variants.},} - A Rhie, SA McCarthy, O Fedrigo, J Damas, G Formenti, S Koren, M Uliano-Silva, W Chow, A Fungtammasan, J Kim, C Lee, BJ Ko, M Chaisson, GL Gedman, LJ Cantin, F Thibaud-Nissen, L Haggerty, I Bista, M Smith, B Haase, J Mountcastle, S Winkler, S Paez, J Howard, SC Vernes, TM Lama, F Grutzner, WC Warren, CN Balakrishnan, D Burt, JM George, MT Biegler, D Iorns, A Digby, D Eason, B Robertson, T Edwards, M Wilkinson, G Turner, A Meyer, AF Kautt, P Franchini, HW Detrich, H Svardal, M Wagner, GJP Naylor, M Pippel, M Malinsky, M Mooney, M Simbirsky, BT Hannigan, T Pesout, M Houck, A Misuraca, SB Kingan, R Hall, Z Kronenberg, I Sović, C Dunn, Z Ning, A Hastie, J Lee, S Selvaraj, RE Green, NH Putnam, I Gut, J Ghurye, E Garrison, Y Sims, J Collins, S Pelan, J Torrance, A Tracey, J Wood, RE Dagnew, D Guan, SE London, DF Clayton, CV Mello, SR Friedrich, PV Lovell, E Osipova, FO Al-Ajli, S Secomandi, H Kim, C Theofanopoulou, M Hiller, Y Zhou, RS Harris, KD Makova, P Medvedev, J Hoffman, P Masterson, K Clark, F Martin, K Howe, P Flicek, BP Walenz, W Kwak, H Clawson, M Diekhans, L Nassar, B Paten, RHS Kraus, AJ Crawford, MTP Gilbert, G Zhang, B Venkatesh, RW Murphy, KP Koepfli, B Shapiro, WE Johnson, F Di Palma, T Marques-Bonet, EC Teeling, T Warnow, JM Graves, OA Ryder, D Haussler, SJ O’Brien, J Korlach, HA Lewin, K Howe, EW Myers, R Durbin, AM Phillippy, ED Jarvis. Towards complete and error-free genome assemblies of all vertebrate species. Nature 2021;592(7856):737–746. doi:10.1038/s41586-021-03451-0
[BibTeX] [Abstract]
High-quality and complete reference genome assemblies are fundamental for the application of genomics to biology, disease, and biodiversity conservation. However, such assemblies are available for only a few non-microbial species\textsuperscript{1-4}. To address this issue, the international Genome 10K (G10K) consortium\textsuperscript{5,6} has worked over a five-year period to evaluate and develop cost-effective methods for assembling highly accurate and nearly complete reference genomes. Here we present lessons learned from generating assemblies for 16 species that represent six major vertebrate lineages. We confirm that long-read sequencing technologies are essential for maximizing genome quality, and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Our assemblies correct substantial errors, add missing sequence in some of the best historical reference genomes, and reveal biological discoveries. These include the identification of many false gene duplications, increases in gene sizes, chromosome rearrangements that are specific to lineages, a repeated independent chromosome breakpoint in bat genomes, and a canonical GC-rich pattern in protein-coding genes and their regulatory regions. Adopting these lessons, we have embarked on the Vertebrate Genomes Project (VGP), an international effort to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species and to help to enable a new era of discovery across the life sciences.
@Article{33911273, author = {Rhie A and McCarthy SA and Fedrigo O and Damas J and Formenti G and Koren S and Uliano-Silva M and Chow W and Fungtammasan A and Kim J and Lee C and Ko BJ and Chaisson M and Gedman GL and Cantin LJ and Thibaud-Nissen F and Haggerty L and Bista I and Smith M and Haase B and Mountcastle J and Winkler S and Paez S and Howard J and Vernes SC and Lama TM and Grutzner F and Warren WC and Balakrishnan CN and Burt D and George JM and Biegler MT and Iorns D and Digby A and Eason D and Robertson B and Edwards T and Wilkinson M and Turner G and Meyer A and Kautt AF and Franchini P and Detrich HW and Svardal H and Wagner M and Naylor GJP and Pippel M and Malinsky M and Mooney M and Simbirsky M and Hannigan BT and Pesout T and Houck M and Misuraca A and Kingan SB and Hall R and Kronenberg Z and Sović I and Dunn C and Ning Z and Hastie A and Lee J and Selvaraj S and Green RE and Putnam NH and Gut I and Ghurye J and Garrison E and Sims Y and Collins J and Pelan S and Torrance J and Tracey A and Wood J and Dagnew RE and Guan D and London SE and Clayton DF and Mello CV and Friedrich SR and Lovell PV and Osipova E and Al-Ajli FO and Secomandi S and Kim H and Theofanopoulou C and Hiller M and Zhou Y and Harris RS and Makova KD and Medvedev P and Hoffman J and Masterson P and Clark K and Martin F and Howe K and Flicek P and Walenz BP and Kwak W and Clawson H and Diekhans M and Nassar L and Paten B and Kraus RHS and Crawford AJ and Gilbert MTP and Zhang G and Venkatesh B and Murphy RW and Koepfli KP and Shapiro B and Johnson WE and Di Palma F and Marques-Bonet T and Teeling EC and Warnow T and Graves JM and Ryder OA and Haussler D and O'Brien SJ and Korlach J and Lewin HA and Howe K and Myers EW and Durbin R and Phillippy AM and Jarvis ED}, title = {Towards complete and error-free genome assemblies of all vertebrate species}, journal = {Nature}, volume = {592}, number = {7856}, pages = {737--746}, year = {2021}, doi = {10.1038/s41586-021-03451-0}, note = {First posted as a preprint: 23 May 2020}, abstract = {High-quality and complete reference genome assemblies are fundamental for the application of genomics to biology, disease, and biodiversity conservation. However, such assemblies are available for only a few non-microbial species\textsuperscript{1-4}. To address this issue, the international Genome 10K (G10K) consortium\textsuperscript{5,6} has worked over a five-year period to evaluate and develop cost-effective methods for assembling highly accurate and nearly complete reference genomes. Here we present lessons learned from generating assemblies for 16 species that represent six major vertebrate lineages. We confirm that long-read sequencing technologies are essential for maximizing genome quality, and that unresolved complex repeats and haplotype heterozygosity are major sources of assembly error when not handled correctly. Our assemblies correct substantial errors, add missing sequence in some of the best historical reference genomes, and reveal biological discoveries. These include the identification of many false gene duplications, increases in gene sizes, chromosome rearrangements that are specific to lineages, a repeated independent chromosome breakpoint in bat genomes, and a canonical GC-rich pattern in protein-coding genes and their regulatory regions. Adopting these lessons, we have embarked on the Vertebrate Genomes Project (VGP), an international effort to generate high-quality, complete reference genomes for all of the roughly 70,000 extant vertebrate species and to help to enable a new era of discovery across the life sciences.},}
2020
- A Warr, N Affara, B Aken, H Beiki, DM Bickhart, K Billis, W Chow, L Eory, HA Finlayson, P Flicek, CG Girón, DK Griffin, R Hall, G Hannum, T Hourlier, K Howe, DA Hume, O Izuogu, K Kim, S Koren, H Liu, N Manchanda, FJ Martin, DJ Nonneman, RE O’Connor, AM Phillippy, GA Rohrer, BD Rosen, LA Rund, CA Sargent, LB Schook, SG Schroeder, AS Schwartz, BM Skinner, R Talbot, E Tseng, CK Tuggle, M Watson, TPL Smith, AL Archibald. An improved pig reference genome sequence to enable pig genetics and genomics research. Gigascience 2020;9(6):giaa051. doi:10.1093/gigascience/giaa051
[BibTeX] [Abstract]
BACKGROUND: The domestic pig (Sus scrofa) is important both as a food source and as a biomedical model given its similarity in size, anatomy, physiology, metabolism, pathology, and pharmacology to humans. The draft reference genome (Sscrofa10.2) of a purebred Duroc female pig established using older clone-based sequencing methods was incomplete, and unresolved redundancies, short-range order and orientation errors, and associated misassembled genes limited its utility. RESULTS: We present 2 annotated highly contiguous chromosome-level genome assemblies created with more recent long-read technologies and a whole-genome shotgun strategy, 1 for the same Duroc female (Sscrofa11.1) and 1 for an outbred, composite-breed male (USMARCv1.0). Both assemblies are of substantially higher (>90-fold) continuity and accuracy than Sscrofa10.2. CONCLUSIONS: These highly contiguous assemblies plus annotation of a further 11 short-read assemblies provide an unprecedented view of the genetic make-up of this important agricultural and biomedical model species. We propose that the improved Duroc assembly (Sscrofa11.1) become the reference genome for genomic research in pigs.
@Article{32543654, author = {Warr A and Affara N and Aken B and Beiki H and Bickhart DM and Billis K and Chow W and Eory L and Finlayson HA and Flicek P and Girón CG and Griffin DK and Hall R and Hannum G and Hourlier T and Howe K and Hume DA and Izuogu O and Kim K and Koren S and Liu H and Manchanda N and Martin FJ and Nonneman DJ and O'Connor RE and Phillippy AM and Rohrer GA and Rosen BD and Rund LA and Sargent CA and Schook LB and Schroeder SG and Schwartz AS and Skinner BM and Talbot R and Tseng E and Tuggle CK and Watson M and Smith TPL and Archibald AL}, title = {An improved pig reference genome sequence to enable pig genetics and genomics research}, journal = {Gigascience}, volume = {9}, number = {6}, pages = {giaa051}, year = {2020}, doi = {10.1093/gigascience/giaa051}, note = {First posted as a preprint: 13 June 2019}, abstract = {BACKGROUND: The domestic pig (Sus scrofa) is important both as a food source and as a biomedical model given its similarity in size, anatomy, physiology, metabolism, pathology, and pharmacology to humans. The draft reference genome (Sscrofa10.2) of a purebred Duroc female pig established using older clone-based sequencing methods was incomplete, and unresolved redundancies, short-range order and orientation errors, and associated misassembled genes limited its utility. RESULTS: We present 2 annotated highly contiguous chromosome-level genome assemblies created with more recent long-read technologies and a whole-genome shotgun strategy, 1 for the same Duroc female (Sscrofa11.1) and 1 for an outbred, composite-breed male (USMARCv1.0). Both assemblies are of substantially higher (>90-fold) continuity and accuracy than Sscrofa10.2. CONCLUSIONS: These highly contiguous assemblies plus annotation of a further 11 short-read assemblies provide an unprecedented view of the genetic make-up of this important agricultural and biomedical model species. We propose that the improved Duroc assembly (Sscrofa11.1) become the reference genome for genomic research in pigs.},} - R Ordonez, M Kulis, N Russinol, V Chapaprieta, A Carrasco-Leon, B Garcia-Torre, S Charampopoulou, G Clot, R Beekman, C Meydan, M Duran-Ferrer, N Verdaguer-Dot, R Vilarrasa-Blasi, P Soler-Vila, L Garate, E Miranda, E San Jose-Eneriz, JR Rodriguez-Madoz, T Ezponda, R Martinez-Turrilas, A Vilas-Zornoza, D Lara-Astiaso, D Dupere-Richer, JH Martens, H El-Omri, RY Taha, MJ Calasanz, B Paiva, J San Miguel, PH Flicek, IG Gut, A Melnick, CS Mitsiades, JD Licht, E Campo, HG Stunnenberg, X Agirre, F Prosper, I Martin-Subero. Chromatin activation as a unifying principle underlying pathogenic mechanisms in multiple myeloma. Genome Res 2020;30(9):1217–1227. doi:10.1101/gr.265520.120
[BibTeX] [Abstract]
Multiple myeloma (MM) is a plasma cell neoplasm associated with a broad variety of genetic lesions. In spite of this genetic heterogeneity, MMs share a characteristic malignant phenotype whose underlying molecular basis remains poorly characterized. In the present study, we examined plasma cells from MM using a multi-epigenomics approach and demonstrated that when compared to normal B cells, malignant plasma cells showed an extensive activation of regulatory elements, in part affecting coregulated adjacent genes. Among target genes upregulated by this process, we found members of the NOTCH, NF-kB, MTOR signaling and TP53 signaling pathways. Other activated genes included sets involved in osteoblast differentiation and response to oxidative stress, all of which have been shown to be associated with the MM phenotype and clinical behavior. We functionally characterized MM specific active distant enhancers controlling the expression of thioredoxin (TXN), a major regulator of cellular redox status, and in addition identified PRDM5 as a novel essential gene for MM. Collectively our data indicates that aberrant chromatin activation is a unifying feature underlying the malignant plasma cell phenotype.
@Article{32820006, author = {Ordonez R and Kulis M and Russinol N and Chapaprieta V and Carrasco-Leon A and Garcia-Torre B and Charampopoulou S and Clot G and Beekman R and Meydan C and Duran-Ferrer M and Verdaguer-Dot N and Vilarrasa-Blasi R and Soler-Vila P and Garate L and Miranda E and San Jose-Eneriz E and Rodriguez-Madoz JR and Ezponda T and Martinez-Turrilas R and Vilas-Zornoza A and Lara-Astiaso D and Dupere-Richer D and Martens JH and El-Omri H and Taha RY and Calasanz MJ and Paiva B and San Miguel J and Flicek PH and Gut IG and Melnick A and Mitsiades CS and Licht JD and Campo E and Stunnenberg HG and Agirre X and Prosper F and Martin-Subero I}, title = {Chromatin activation as a unifying principle underlying pathogenic mechanisms in multiple myeloma}, journal = {Genome Res}, volume = {30}, number = {9}, pages = {1217--1227}, year = {2020}, doi = {10.1101/gr.265520.120}, howpublished = {Advanced online publication: 20 August 2020}, note = {First posted as a preprint: 20 August 2019}, abstract = {Multiple myeloma (MM) is a plasma cell neoplasm associated with a broad variety of genetic lesions. In spite of this genetic heterogeneity, MMs share a characteristic malignant phenotype whose underlying molecular basis remains poorly characterized. In the present study, we examined plasma cells from MM using a multi-epigenomics approach and demonstrated that when compared to normal B cells, malignant plasma cells showed an extensive activation of regulatory elements, in part affecting coregulated adjacent genes. Among target genes upregulated by this process, we found members of the NOTCH, NF-kB, MTOR signaling and TP53 signaling pathways. Other activated genes included sets involved in osteoblast differentiation and response to oxidative stress, all of which have been shown to be associated with the MM phenotype and clinical behavior. We functionally characterized MM specific active distant enhancers controlling the expression of thioredoxin (TXN), a major regulator of cellular redox status, and in addition identified PRDM5 as a novel essential gene for MM. Collectively our data indicates that aberrant chromatin activation is a unifying feature underlying the malignant plasma cell phenotype.},} - E Kentepozidou, SJ Aitken, C Feig, K Stefflova, X Ibarra-Soria, DT Odom, M Roller, P Flicek. Clustered CTCF binding is an evolutionary mechanism to maintain topologically associating domains. Genome Biol 2020;21(1):5. doi:10.1186/s13059-019-1894-x
[BibTeX] [Abstract]
BACKGROUND: CTCF binding contributes to the establishment of a higher-order genome structure by demarcating the boundaries of large-scale topologically associating domains (TADs). However, despite the importance and conservation of TADs, the role of CTCF binding in their evolution and stability remains elusive. RESULTS: We carry out an experimental and computational study that exploits the natural genetic variation across five closely related species to assess how CTCF binding patterns stably fixed by evolution in each species contribute to the establishment and evolutionary dynamics of TAD boundaries. We perform CTCF ChIP-seq in multiple mouse species to create genome-wide binding profiles and associate them with TAD boundaries. Our analyses reveal that CTCF binding is maintained at TAD boundaries by a balance of selective constraints and dynamic evolutionary processes. Regardless of their conservation across species, CTCF binding sites at TAD boundaries are subject to stronger sequence and functional constraints compared to other CTCF sites. TAD boundaries frequently harbor dynamically evolving clusters containing both evolutionarily old and young CTCF sites as a result of the repeated acquisition of new species-specific sites close to conserved ones. The overwhelming majority of clustered CTCF sites colocalize with cohesin and are significantly closer to gene transcription start sites than nonclustered CTCF sites, suggesting that CTCF clusters particularly contribute to cohesin stabilization and transcriptional regulation. CONCLUSIONS: Dynamic conservation of CTCF site clusters is an apparently important feature of CTCF binding evolution that is critical to the functional stability of a higher-order chromatin structure.
@Article{31910870, author = {Kentepozidou E and Aitken SJ and Feig C and Stefflova K and Ibarra-Soria X and Odom DT and Roller M and Flicek P}, title = {Clustered CTCF binding is an evolutionary mechanism to maintain topologically associating domains}, journal = {Genome Biol}, volume = {21}, number = {1}, pages = {5}, year = {2020}, doi = {10.1186/s13059-019-1894-x}, note = {First posted as a preprint: 12 June 2019}, abstract = {BACKGROUND: CTCF binding contributes to the establishment of a higher-order genome structure by demarcating the boundaries of large-scale topologically associating domains (TADs). However, despite the importance and conservation of TADs, the role of CTCF binding in their evolution and stability remains elusive. RESULTS: We carry out an experimental and computational study that exploits the natural genetic variation across five closely related species to assess how CTCF binding patterns stably fixed by evolution in each species contribute to the establishment and evolutionary dynamics of TAD boundaries. We perform CTCF ChIP-seq in multiple mouse species to create genome-wide binding profiles and associate them with TAD boundaries. Our analyses reveal that CTCF binding is maintained at TAD boundaries by a balance of selective constraints and dynamic evolutionary processes. Regardless of their conservation across species, CTCF binding sites at TAD boundaries are subject to stronger sequence and functional constraints compared to other CTCF sites. TAD boundaries frequently harbor dynamically evolving clusters containing both evolutionarily old and young CTCF sites as a result of the repeated acquisition of new species-specific sites close to conserved ones. The overwhelming majority of clustered CTCF sites colocalize with cohesin and are significantly closer to gene transcription start sites than nonclustered CTCF sites, suggesting that CTCF clusters particularly contribute to cohesin stabilization and transcriptional regulation. CONCLUSIONS: Dynamic conservation of CTCF site clusters is an apparently important feature of CTCF binding evolution that is critical to the functional stability of a higher-order chromatin structure.},} - AD Yates, P Achuthan, W Akanni, J Allen, J Allen, J Alvarez-Jarreta, MR Amode, IM Armean, AG Azov, R Bennett, J Bhai, K Billis, S Boddu, JC Marugán, C Cummins, C Davidson, K Dodiya, R Fatima, A Gall, CG Giron, L Gil, T Grego, L Haggerty, E Haskell, T Hourlier, OG Izuogu, SH Janacek, T Juettemann, M Kay, I Lavidas, T Le, D Lemos, JG Martinez, T Maurel, M McDowall, A McMahon, S Mohanan, B Moore, M Nuhn, DN Oheh, A Parker, A Parton, M Patricio, MP Sakthivel, AI Abdul Salam, BM Schmitt, H Schuilenburg, D Sheppard, M Sycheva, M Szuba, K Taylor, A Thormann, G Threadgold, A Vullo, B Walts, A Winterbottom, A Zadissa, M Chakiachvili, B Flint, A Frankish, SE Hunt, G IIsley, M Kostadima, N Langridge, JE Loveland, FJ Martin, J Morales, JM Mudge, M Muffato, E Perry, M Ruffier, SJ Trevanion, F Cunningham, KL Howe, DR Zerbino, P Flicek. Ensembl 2020. Nucleic Acids Res 2020;48(D1):D682–D688. doi:10.1093/nar/gkz966
[BibTeX] [Abstract]
The Ensembl (https://www.ensembl.org) is a system for generating and distributing genome annotation such as genes, variation, regulation and comparative genomics across the vertebrate subphylum and key model organisms. The Ensembl annotation pipeline is capable of integrating experimental and reference data from multiple providers into a single integrated resource. Here, we present 94 newly annotated and re-annotated genomes, bringing the total number of genomes offered by Ensembl to 227. This represents the single largest expansion of the resource since its inception. We also detail our continued efforts to improve human annotation, developments in our epigenome analysis and display, a new tool for imputing causal genes from genome-wide association studies and visualisation of variation within a 3D protein model. Finally, we present information on our new website. Both software and data are made available without restriction via our website, online tools platform and programmatic interfaces (available under an Apache 2.0 license) and data updates made available four times a year.
@Article{31691826, author = {Yates AD and Achuthan P and Akanni W and Allen J and Allen J and Alvarez-Jarreta J and Amode MR and Armean IM and Azov AG and Bennett R and Bhai J and Billis K and Boddu S and Marugán JC and Cummins C and Davidson C and Dodiya K and Fatima R and Gall A and Giron CG and Gil L and Grego T and Haggerty L and Haskell E and Hourlier T and Izuogu OG and Janacek SH and Juettemann T and Kay M and Lavidas I and Le T and Lemos D and Martinez JG and Maurel T and McDowall M and McMahon A and Mohanan S and Moore B and Nuhn M and Oheh DN and Parker A and Parton A and Patricio M and Sakthivel MP and Abdul Salam AI and Schmitt BM and Schuilenburg H and Sheppard D and Sycheva M and Szuba M and Taylor K and Thormann A and Threadgold G and Vullo A and Walts B and Winterbottom A and Zadissa A and Chakiachvili M and Flint B and Frankish A and Hunt SE and IIsley G and Kostadima M and Langridge N and Loveland JE and Martin FJ and Morales J and Mudge JM and Muffato M and Perry E and Ruffier M and Trevanion SJ and Cunningham F and Howe KL and Zerbino DR and Flicek P}, title = {Ensembl 2020}, journal = {Nucleic Acids Res}, volume = {48}, number = {D1}, pages = {D682--D688}, year = {2020}, doi = {10.1093/nar/gkz966}, howpublished = {Advanced online publication: 6 November 2019}, abstract = {The Ensembl (https://www.ensembl.org) is a system for generating and distributing genome annotation such as genes, variation, regulation and comparative genomics across the vertebrate subphylum and key model organisms. The Ensembl annotation pipeline is capable of integrating experimental and reference data from multiple providers into a single integrated resource. Here, we present 94 newly annotated and re-annotated genomes, bringing the total number of genomes offered by Ensembl to 227. This represents the single largest expansion of the resource since its inception. We also detail our continued efforts to improve human annotation, developments in our epigenome analysis and display, a new tool for imputing causal genes from genome-wide association studies and visualisation of variation within a 3D protein model. Finally, we present information on our new website. Both software and data are made available without restriction via our website, online tools platform and programmatic interfaces (available under an Apache 2.0 license) and data updates made available four times a year.},} - KL Howe, B Contreras-Moreira, N De Silva, G Maslen, W Akanni, J Allen, J Alvarez-Jarreta, M Barba, DM Bolser, L Cambell, M Carbajo, M Chakiachvili, M Christensen, C Cummins, A Cuzick, P Davis, S Fexova, A Gall, N George, L Gil, P Gupta, KE Hammond-Kosack, E Haskell, SE Hunt, P Jaiswal, SH Janacek, PJ Kersey, N Langridge, U Maheswari, T Maurel, MD McDowall, B Moore, M Muffato, G Naamati, S Naithani, A Olson, I Papatheodorou, M Patricio, M Paulini, H Pedro, E Perry, J Preece, M Rosello, M Russell, V Sitnik, DM Staines, J Stein, MK Tello-Ruiz, SJ Trevanion, M Urban, S Wei, D Ware, G Williams, AD Yates, P Flicek. Ensembl Genomes 2020-enabling non-vertebrate genomic research. Nucleic Acids Res 2020;48(D1):D689–D695. doi:10.1093/nar/gkz890
[BibTeX] [Abstract]
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the context of the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of interfaces to genomic data across the tree of life, including reference genome sequence, gene models, transcriptional data, genetic variation and comparative analysis. Data may be accessed via our website, online tools platform and programmatic interfaces, with updates made four times per year (in synchrony with Ensembl). Here, we provide an overview of Ensembl Genomes, with a focus on recent developments. These include the continued growth, more robust and reproducible sets of orthologues and paralogues, and enriched views of gene expression and gene function in plants. Finally, we report on our continued deeper integration with the Ensembl project, which forms a key part of our future strategy for dealing with the increasing quantity of available genome-scale data across the tree of life.
@Article{31598706, author = {Howe KL and Contreras-Moreira B and De Silva N and Maslen G and Akanni W and Allen J and Alvarez-Jarreta J and Barba M and Bolser DM and Cambell L and Carbajo M and Chakiachvili M and Christensen M and Cummins C and Cuzick A and Davis P and Fexova S and Gall A and George N and Gil L and Gupta P and Hammond-Kosack KE and Haskell E and Hunt SE and Jaiswal P and Janacek SH and Kersey PJ and Langridge N and Maheswari U and Maurel T and McDowall MD and Moore B and Muffato M and Naamati G and Naithani S and Olson A and Papatheodorou I and Patricio M and Paulini M and Pedro H and Perry E and Preece J and Rosello M and Russell M and Sitnik V and Staines DM and Stein J and Tello-Ruiz MK and Trevanion SJ and Urban M and Wei S and Ware D and Williams G and Yates AD and Flicek P}, title = {Ensembl Genomes 2020-enabling non-vertebrate genomic research}, journal = {Nucleic Acids Res}, volume = {48}, number = {D1}, pages = {D689--D695}, year = {2020}, doi = {10.1093/nar/gkz890}, howpublished = {Advanced online publication: 10 October 2019}, abstract = {Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the context of the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of interfaces to genomic data across the tree of life, including reference genome sequence, gene models, transcriptional data, genetic variation and comparative analysis. Data may be accessed via our website, online tools platform and programmatic interfaces, with updates made four times per year (in synchrony with Ensembl). Here, we provide an overview of Ensembl Genomes, with a focus on recent developments. These include the continued growth, more robust and reproducible sets of orthologues and paralogues, and enriched views of gene expression and gene function in plants. Finally, we report on our continued deeper integration with the Ensembl project, which forms a key part of our future strategy for dealing with the increasing quantity of available genome-scale data across the tree of life.},} - ENCODE Project Consortium, JE Moore, MJ Purcaro, HE Pratt, CB Epstein, N Shoresh, J Adrian, T Kawli, CA Davis, A Dobin, R Kaul, J Halow, EL Van Nostrand, P Freese, DU Gorkin, Y Shen, Y He, M Mackiewicz, F Pauli-Behn, BA Williams, A Mortazavi, CA Keller, XO Zhang, SI Elhajjajy, J Huey, DE Dickel, V Snetkova, X Wei, X Wang, JC Rivera-Mulia, J Rozowsky, J Zhang, SB Chhetri, J Zhang, A Victorsen, KP White, A Visel, GW Yeo, CB Burge, E Lécuyer, DM Gilbert, J Dekker, J Rinn, EM Mendenhall, JR Ecker, M Kellis, RJ Klein, WS Noble, A Kundaje, R Guigó, PJ Farnham, JM Cherry, RM Myers, B Ren, BR Graveley, MB Gerstein, LA Pennacchio, MP Snyder, BE Bernstein, B Wold, RC Hardison, TR Gingeras, JA Stamatoyannopoulos, Z Weng. Expanded encyclopaedias of DNA elements in the human and mouse genomes. Nature 2020;583(7818):699–710. doi:10.1038/s41586-020-2493-4
[BibTeX] [Abstract]
The human and mouse genomes contain instructions that specify RNAs and proteins and govern the timing, magnitude, and cellular context of their production. To better delineate these elements, phase III of the Encyclopedia of DNA Elements (ENCODE) Project has expanded analysis of the cell and tissue repertoires of RNA transcription, chromatin structure and modification, DNA methylation, chromatin looping, and occupancy by transcription factors and RNA-binding proteins. Here we summarize these efforts, which have produced 5,992 new experimental datasets, including systematic determinations across mouse fetal development. All data are available through the ENCODE data portal (https://www.encodeproject.org), including phase II ENCODE\textsuperscript{1} and Roadmap Epigenomics\textsuperscript{2} data. We have developed a registry of 926,535 human and 339,815 mouse candidate cis-regulatory elements, covering 7.9 and 3.4\% of their respective genomes, by integrating selected datatypes associated with gene regulation, and constructed a web-based server (SCREEN; http://screen.encodeproject.org) to provide flexible, user-defined access to this resource. Collectively, the ENCODE data and registry provide an expansive resource for the scientific community to build a better understanding of the organization and function of the human and mouse genomes.
@Article{32728249, author = {{ENCODE Project Consortium} and Moore JE and Purcaro MJ and Pratt HE and Epstein CB and Shoresh N and Adrian J and Kawli T and Davis CA and Dobin A and Kaul R and Halow J and Van Nostrand EL and Freese P and Gorkin DU and Shen Y and He Y and Mackiewicz M and Pauli-Behn F and Williams BA and Mortazavi A and Keller CA and Zhang XO and Elhajjajy SI and Huey J and Dickel DE and Snetkova V and Wei X and Wang X and Rivera-Mulia JC and Rozowsky J and Zhang J and Chhetri SB and Zhang J and Victorsen A and White KP and Visel A and Yeo GW and Burge CB and Lécuyer E and Gilbert DM and Dekker J and Rinn J and Mendenhall EM and Ecker JR and Kellis M and Klein RJ and Noble WS and Kundaje A and Guigó R and Farnham PJ and Cherry JM and Myers RM and Ren B and Graveley BR and Gerstein MB and Pennacchio LA and Snyder MP and Bernstein BE and Wold B and Hardison RC and Gingeras TR and Stamatoyannopoulos JA and Weng Z}, title = {Expanded encyclopaedias of DNA elements in the human and mouse genomes}, journal = {Nature}, volume = {583}, number = {7818}, pages = {699--710}, year = {2020}, doi = {10.1038/s41586-020-2493-4}, abstract = {The human and mouse genomes contain instructions that specify RNAs and proteins and govern the timing, magnitude, and cellular context of their production. To better delineate these elements, phase III of the Encyclopedia of DNA Elements (ENCODE) Project has expanded analysis of the cell and tissue repertoires of RNA transcription, chromatin structure and modification, DNA methylation, chromatin looping, and occupancy by transcription factors and RNA-binding proteins. Here we summarize these efforts, which have produced 5,992 new experimental datasets, including systematic determinations across mouse fetal development. All data are available through the ENCODE data portal (https://www.encodeproject.org), including phase II ENCODE\textsuperscript{1} and Roadmap Epigenomics\textsuperscript{2} data. We have developed a registry of 926,535 human and 339,815 mouse candidate cis-regulatory elements, covering 7.9 and 3.4\% of their respective genomes, by integrating selected datatypes associated with gene regulation, and constructed a web-based server (SCREEN; http://screen.encodeproject.org) to provide flexible, user-defined access to this resource. Collectively, the ENCODE data and registry provide an expansive resource for the scientific community to build a better understanding of the organization and function of the human and mouse genomes.},} - D Azazi, JM Mudge, DT Odom, P Flicek. Functional signatures of evolutionarily young CTCF binding sites. BMC Biol 2020;18(1):132. doi:10.1186/s12915-020-00863-8
[BibTeX] [Abstract]
BACKGROUND: The introduction of novel CTCF binding sites in gene regulatory regions in the rodent lineage is partly the effect of transposable element expansion, particularly in the murine lineage. The exact mechanism and functional impact of evolutionarily novel CTCF binding sites are not yet fully understood. We investigated the impact of novel subspecies-specific CTCF binding sites in two Mus genus subspecies, Mus musculus domesticus and Mus musculus castaneus, that diverged 0.5 million years ago. RESULTS: CTCF binding site evolution is influenced by the action of the B2-B4 family of transposable elements independently in both lineages, leading to the proliferation of novel CTCF binding sites. A subset of evolutionarily young sites may harbour transcriptional functionality as evidenced by the stability of their binding across multiple tissues in M. musculus domesticus (BL6), while overall the distance of subspecies-specific CTCF binding to the nearest transcription start sites and/or topologically associated domains (TADs) is largely similar to musculus-common CTCF sites. Remarkably, we discovered a recurrent regulatory architecture consisting of a CTCF binding site and an interferon gene that appears to have been tandemly duplicated to create a 15-gene cluster on chromosome 4, thus forming a novel BL6 specific immune locus in which CTCF may play a regulatory role. CONCLUSIONS: Our results demonstrate that thousands of CTCF binding sites show multiple functional signatures rapidly after incorporation into the genome.
@Article{32988407, author = {Azazi D and Mudge JM and Odom DT and Flicek P}, title = {Functional signatures of evolutionarily young CTCF binding sites}, journal = {BMC Biol}, volume = {18}, number = {1}, pages = {132}, year = {2020}, doi = {10.1186/s12915-020-00863-8}, note = {First posted as a preprint: 31 January 2020}, abstract = {BACKGROUND: The introduction of novel CTCF binding sites in gene regulatory regions in the rodent lineage is partly the effect of transposable element expansion, particularly in the murine lineage. The exact mechanism and functional impact of evolutionarily novel CTCF binding sites are not yet fully understood. We investigated the impact of novel subspecies-specific CTCF binding sites in two Mus genus subspecies, Mus musculus domesticus and Mus musculus castaneus, that diverged 0.5 million years ago. RESULTS: CTCF binding site evolution is influenced by the action of the B2-B4 family of transposable elements independently in both lineages, leading to the proliferation of novel CTCF binding sites. A subset of evolutionarily young sites may harbour transcriptional functionality as evidenced by the stability of their binding across multiple tissues in M. musculus domesticus (BL6), while overall the distance of subspecies-specific CTCF binding to the nearest transcription start sites and/or topologically associated domains (TADs) is largely similar to musculus-common CTCF sites. Remarkably, we discovered a recurrent regulatory architecture consisting of a CTCF binding site and an interferon gene that appears to have been tandemly duplicated to create a 15-gene cluster on chromosome 4, thus forming a novel BL6 specific immune locus in which CTCF may play a regulatory role. CONCLUSIONS: Our results demonstrate that thousands of CTCF binding sites show multiple functional signatures rapidly after incorporation into the genome.},} - S Kongsstovu Í, HA Dahl, H Gislason, E Homrum Í, JA Jacobsen, P Flicek, SO Mikalsen. Identification of male heterogametic sex determining regions on the Atlantic herring Clupea harengus genome. J Fish Biol 2020;97(1):190–201. doi:10.1111/jfb.14349
[BibTeX] [Abstract]
The sex determination system of Atlantic herring Clupea harengus L., a commercially important fish, was investigated. Low coverage whole genome sequencing of 48 females and 55 males and a genome-wide association study revealed two regions on chromosomes 8 and 21 associated with sex. The genotyping data of the single nucleotide polymorphisms associated with sex showed that 99.4\% of the available female genotypes were homozygous, whereas 68.6\% of the available male genotypes were heterozygous. This is close to the theoretical expectation of homo/heterozygous distribution at low sequencing coverage when the males are factually heterozygous. This suggested a male heterogametic sex determination system in C. harengus, consistent with other species within the Clupeiformes group. There were 76 protein coding genes on the sex regions but none of these genes were previously reported master sex regulation genes, or obviously related to sex determination. However, many of these genes are expressed in testis or ovary in other species, but the exact genes controlling sex determination in C. harengus could not be identified. This article is protected by copyright. All rights reserved.
@Article{32293027, author = {Í Kongsstovu S and Dahl HA and Gislason H and Í Homrum E and Jacobsen JA and Flicek P and Mikalsen SO}, title = {Identification of male heterogametic sex determining regions on the Atlantic herring Clupea harengus genome}, journal = {J Fish Biol}, volume = {97}, number = {1}, pages = {190--201}, year = {2020}, doi = {10.1111/jfb.14349}, howpublished = {Advanced online publication: 15 April 2020}, abstract = {The sex determination system of Atlantic herring Clupea harengus L., a commercially important fish, was investigated. Low coverage whole genome sequencing of 48 females and 55 males and a genome-wide association study revealed two regions on chromosomes 8 and 21 associated with sex. The genotyping data of the single nucleotide polymorphisms associated with sex showed that 99.4\% of the available female genotypes were homozygous, whereas 68.6\% of the available male genotypes were heterozygous. This is close to the theoretical expectation of homo/heterozygous distribution at low sequencing coverage when the males are factually heterozygous. This suggested a male heterogametic sex determination system in C. harengus, consistent with other species within the Clupeiformes group. There were 76 protein coding genes on the sex regions but none of these genes were previously reported master sex regulation genes, or obviously related to sex determination. However, many of these genes are expressed in testis or ovary in other species, but the exact genes controlling sex determination in C. harengus could not be identified. This article is protected by copyright. All rights reserved.},} - J Robinson, DJ Barker, X Georgiou, MA Cooper, P Flicek, SGE Marsh. IPD-IMGT/HLA Database. Nucleic Acids Res 2020;48(D1):D948–D955. doi:10.1093/nar/gkz950
[BibTeX] [Abstract]
The IPD-IMGT/HLA Database, http://www.ebi.ac.uk/ipd/imgt/hla/, currently contains over 25 000 allele sequence for 45 genes, which are located within the Major Histocompatibility Complex (MHC) of the human genome. This region is the most polymorphic region of the human genome, and the levels of polymorphism seen exceed most other genes. Some of the genes have several thousand variants and are now termed hyperpolymorphic, rather than just simply polymorphic. The IPD-IMGT/HLA Database has provided a stable, highly accessible, user-friendly repository for this information, providing the scientific and medical community access to the many variant sequences of this gene system, that are critical for the successful outcome of transplantation. The number of currently known variants, and dramatic increase in the number of new variants being identified has necessitated a dedicated resource with custom tools for curation and publication. The challenge for the database is to continue to provide a highly curated database of sequence variants, while supporting the increased number of submissions and complexity of sequences. In order to do this, traditional methods of accessing and presenting data will be challenged, and new methods will need to be utilized to keep pace with new discoveries.
@Article{31667505, author = {Robinson J and Barker DJ and Georgiou X and Cooper MA and Flicek P and Marsh SGE}, title = {IPD-IMGT/HLA Database}, journal = {Nucleic Acids Res}, volume = {48}, number = {D1}, pages = {D948--D955}, year = {2020}, doi = {10.1093/nar/gkz950}, howpublished = {Advanced online publication: 31 October 2019}, abstract = {The IPD-IMGT/HLA Database, http://www.ebi.ac.uk/ipd/imgt/hla/, currently contains over 25 000 allele sequence for 45 genes, which are located within the Major Histocompatibility Complex (MHC) of the human genome. This region is the most polymorphic region of the human genome, and the levels of polymorphism seen exceed most other genes. Some of the genes have several thousand variants and are now termed hyperpolymorphic, rather than just simply polymorphic. The IPD-IMGT/HLA Database has provided a stable, highly accessible, user-friendly repository for this information, providing the scientific and medical community access to the many variant sequences of this gene system, that are critical for the successful outcome of transplantation. The number of currently known variants, and dramatic increase in the number of new variants being identified has necessitated a dedicated resource with custom tools for curation and publication. The challenge for the database is to continue to provide a highly curated database of sequence variants, while supporting the increased number of submissions and complexity of sequences. In order to do this, traditional methods of accessing and presenting data will be challenged, and new methods will need to be utilized to keep pace with new discoveries.},} - AL Swan, C Schütt, J Rozman, M Del Mar Muñiz Moreno, S Brandmaier, M Simon, S Leuchtenberger, M Griffiths, R Brommage, P Keskivali-Bond, H Grallert, T Werner, R Teperino, L Becker, G Miller, A Moshiri, JR Seavitt, DD Cissell, TF Meehan, EF Acar, CJ Lelliott, AM Flenniken, MF Champy, T Sorg, A Ayadi, RE Braun, H Cater, ME Dickinson, P Flicek, J Gallegos, EJ Ghirardello, JD Heaney, S Jacquot, C Lally, JG Logan, L Teboul, J Mason, N Spielmann, C McKerlie, SA Murray, LMJ Nutter, KF Odfalk, H Parkinson, J Prochazka, CL Reynolds, M Selloum, F Spoutil, KL Svenson, TS Vales, SE Wells, JK White, R Sedlacek, W Wurst, KKC Lloyd, PI Croucher, H Fuchs, GR Williams, D Bassett, V Gailus-Durner, Y Herault, AM Mallon, SDM Brown, P Mayer-Kuckuk, de M Hrabe Angelis, C IMPC. Mouse mutant phenotyping at scale reveals novel genes controlling bone mineral density. PLoS Genet 2020;16(12):e1009190. doi:10.1371/journal.pgen.1009190
[BibTeX] [Abstract]
The genetic landscape of diseases associated with changes in bone mineral density (BMD), such as osteoporosis, is only partially understood. Here, we explored data from 3,823 mutant mouse strains for BMD, a measure that is frequently altered in a range of bone pathologies, including osteoporosis. A total of 200 genes were found to significantly affect BMD. This pool of BMD genes comprised 141 genes with previously unknown functions in bone biology and was complementary to pools derived from recent human studies. Nineteen of the 141 genes also caused skeletal abnormalities. Examination of the BMD genes in osteoclasts and osteoblasts underscored BMD pathways, including vesicle transport, in these cells and together with in silico bone turnover studies resulted in the prioritization of candidate genes for further investigation. Overall, the results add novel pathophysiological and molecular insight into bone health and disease.
@Article{33370286, author = {Swan AL and Schütt C and Rozman J and Del Mar Muñiz Moreno M and Brandmaier S and Simon M and Leuchtenberger S and Griffiths M and Brommage R and Keskivali-Bond P and Grallert H and Werner T and Teperino R and Becker L and Miller G and Moshiri A and Seavitt JR and Cissell DD and Meehan TF and Acar EF and Lelliott CJ and Flenniken AM and Champy MF and Sorg T and Ayadi A and Braun RE and Cater H and Dickinson ME and Flicek P and Gallegos J and Ghirardello EJ and Heaney JD and Jacquot S and Lally C and Logan JG and Teboul L and Mason J and Spielmann N and McKerlie C and Murray SA and Nutter LMJ and Odfalk KF and Parkinson H and Prochazka J and Reynolds CL and Selloum M and Spoutil F and Svenson KL and Vales TS and Wells SE and White JK and Sedlacek R and Wurst W and Lloyd KKC and Croucher PI and Fuchs H and Williams GR and Bassett D and Gailus-Durner V and Herault Y and Mallon AM and Brown SDM and Mayer-Kuckuk P and Hrabe de Angelis M and IMPC C}, title = {Mouse mutant phenotyping at scale reveals novel genes controlling bone mineral density}, journal = {PLoS Genet}, volume = {16}, number = {12}, pages = {e1009190}, year = {2020}, doi = {10.1371/journal.pgen.1009190}, abstract = {The genetic landscape of diseases associated with changes in bone mineral density (BMD), such as osteoporosis, is only partially understood. Here, we explored data from 3,823 mutant mouse strains for BMD, a measure that is frequently altered in a range of bone pathologies, including osteoporosis. A total of 200 genes were found to significantly affect BMD. This pool of BMD genes comprised 141 genes with previously unknown functions in bone biology and was complementary to pools derived from recent human studies. Nineteen of the 141 genes also caused skeletal abnormalities. Examination of the BMD genes in osteoclasts and osteoblasts underscored BMD pathways, including vesicle transport, in these cells and together with in silico bone turnover studies resulted in the prioritization of candidate genes for further investigation. Overall, the results add novel pathophysiological and molecular insight into bone health and disease.},} - of {ICGC/TCGA Pan-Cancer Analysis Whole Genomes Consortium}. Pan-cancer analysis of whole genomes. Nature 2020;578(7793):82–93. doi:10.1038/s41586-020-1969-6
[BibTeX] [Abstract]
Cancer is driven by genetic change, and the advent of massively parallel sequencing has enabled systematic documentation of this variation at the whole-genome scale\textsuperscript{1-3}. Here we report the integrative analysis of 2,658 whole-cancer genomes and their matching normal tissues across 38 tumour types from the Pan-Cancer Analysis of Whole Genomes (PCAWG) Consortium of the International Cancer Genome Consortium (ICGC) and The Cancer Genome Atlas (TCGA). We describe the generation of the PCAWG resource, facilitated by international data sharing using compute clouds. On average, cancer genomes contained 4-5 driver mutations when combining coding and non-coding genomic elements; however, in around 5\% of cases no drivers were identified, suggesting that cancer driver discovery is not yet complete. Chromothripsis, in which many clustered structural variants arise in a single catastrophic event, is frequently an early event in tumour evolution; in acral melanoma, for example, these events precede most somatic point mutations and affect several cancer-associated genes simultaneously. Cancers with abnormal telomere maintenance often originate from tissues with low replicative activity and show several mechanisms of preventing telomere attrition to critical levels. Common and rare germline variants affect patterns of somatic mutation, including point mutations, structural variants and somatic retrotransposition. A collection of papers from the PCAWG Consortium describes non-coding mutations that drive cancer beyond those in the TERT promoter\textsuperscript{4}; identifies new signatures of mutational processes that cause base substitutions, small insertions and deletions and structural variation\textsuperscript{5,6}; analyses timings and patterns of tumour evolution\textsuperscript{7}; describes the diverse transcriptional consequences of somatic mutation on splicing, expression levels, fusion genes and promoter activity\textsuperscript{8,9}; and evaluates a range of more-specialized features of cancer genomes\textsuperscript{8,10-18}.
@Article{32025007, author = {{ICGC/TCGA, Pan-Cancer Analysis of Whole Genomes Consortium}}, title = {Pan-cancer analysis of whole genomes}, journal = {Nature}, volume = {578}, number = {7793}, pages = {82--93}, year = {2020}, doi = {10.1038/s41586-020-1969-6}, abstract = {Cancer is driven by genetic change, and the advent of massively parallel sequencing has enabled systematic documentation of this variation at the whole-genome scale\textsuperscript{1-3}. Here we report the integrative analysis of 2,658 whole-cancer genomes and their matching normal tissues across 38 tumour types from the Pan-Cancer Analysis of Whole Genomes (PCAWG) Consortium of the International Cancer Genome Consortium (ICGC) and The Cancer Genome Atlas (TCGA). We describe the generation of the PCAWG resource, facilitated by international data sharing using compute clouds. On average, cancer genomes contained 4-5 driver mutations when combining coding and non-coding genomic elements; however, in around 5\% of cases no drivers were identified, suggesting that cancer driver discovery is not yet complete. Chromothripsis, in which many clustered structural variants arise in a single catastrophic event, is frequently an early event in tumour evolution; in acral melanoma, for example, these events precede most somatic point mutations and affect several cancer-associated genes simultaneously. Cancers with abnormal telomere maintenance often originate from tissues with low replicative activity and show several mechanisms of preventing telomere attrition to critical levels. Common and rare germline variants affect patterns of somatic mutation, including point mutations, structural variants and somatic retrotransposition. A collection of papers from the PCAWG Consortium describes non-coding mutations that drive cancer beyond those in the TERT promoter\textsuperscript{4}; identifies new signatures of mutational processes that cause base substitutions, small insertions and deletions and structural variation\textsuperscript{5,6}; analyses timings and patterns of tumour evolution\textsuperscript{7}; describes the diverse transcriptional consequences of somatic mutation on splicing, expression levels, fusion genes and promoter activity\textsuperscript{8,9}; and evaluates a range of more-specialized features of cancer genomes\textsuperscript{8,10-18}.},} - ENCODE Project Consortium, MP Snyder, TR Gingeras, JE Moore, Z Weng, MB Gerstein, B Ren, RC Hardison, JA Stamatoyannopoulos, BR Graveley, EA Feingold, MJ Pazin, M Pagan, DA Gilchrist, BC Hitz, JM Cherry, BE Bernstein, EM Mendenhall, DR Zerbino, A Frankish, P Flicek, RM Myers. Perspectives on ENCODE. Nature 2020;583(7818):693–698. doi:10.1038/s41586-020-2449-8
[BibTeX] [Abstract]
The Encylopedia of DNA Elements (ENCODE) Project launched in 2003 with the long-term goal of developing a comprehensive map of functional elements in the human genome. These included genes, biochemical regions associated with gene regulation (for example, transcription factor binding sites, open chromatin, and histone marks) and transcript isoforms. The marks serve as sites for candidate cis-regulatory elements (cCREs) that may serve functional roles in regulating gene expression\textsuperscript{1}. The project has been extended to model organisms, particularly the mouse. In the third phase of ENCODE, nearly a million and more than 300,000 cCRE annotations have been generated for human and mouse, respectively, and these have provided a valuable resource for the scientific community.
@Article{32728248, author = {{ENCODE Project Consortium} and Snyder MP and Gingeras TR and Moore JE and Weng Z and Gerstein MB and Ren B and Hardison RC and Stamatoyannopoulos JA and Graveley BR and Feingold EA and Pazin MJ and Pagan M and Gilchrist DA and Hitz BC and Cherry JM and Bernstein BE and Mendenhall EM and Zerbino DR and Frankish A and Flicek P and Myers RM}, title = {Perspectives on ENCODE}, journal = {Nature}, volume = {583}, number = {7818}, pages = {693--698}, year = {2020}, doi = {10.1038/s41586-020-2449-8}, abstract = {The Encylopedia of DNA Elements (ENCODE) Project launched in 2003 with the long-term goal of developing a comprehensive map of functional elements in the human genome. These included genes, biochemical regions associated with gene regulation (for example, transcription factor binding sites, open chromatin, and histone marks) and transcript isoforms. The marks serve as sites for candidate cis-regulatory elements (cCREs) that may serve functional roles in regulating gene expression\textsuperscript{1}. The project has been extended to model organisms, particularly the mouse. In the third phase of ENCODE, nearly a million and more than 300,000 cCRE annotations have been generated for human and mouse, respectively, and these have provided a valuable resource for the scientific community.},} - SJ Aitken, CJ Anderson, F Connor, O Pich, V Sundaram, C Feig, TF Rayner, M Lukk, S Aitken, J Luft, E Kentepozidou, C Arnedo-Pac, SV Beentjes, SE Davies, RM Drews, A Ewing, VB Kaiser, A Khamseh, E López-Arribillaga, AM Redmond, J Santoyo-Lopez, I Sentís, L Talmane, AD Yates, CEC Liver, CA Semple, N López-Bigas, P Flicek, DT Odom, MS Taylor. Pervasive lesion segregation shapes cancer genome evolution. Nature 2020;583(7815):265–270. doi:10.1038/s41586-020-2435-1
[BibTeX] [Abstract]
Cancers arise through the acquisition of oncogenic mutations and grow by clonal expansion\textsuperscript{1,2}. Here we reveal that most mutagenic DNA lesions are not resolved into a mutated DNA base pair within a single cell cycle. Instead, DNA lesions segregate, unrepaired, into daughter cells for multiple cell generations, resulting in the chromosome-scale phasing of subsequent mutations. We characterize this process in mutagen-induced mouse liver tumours and show that DNA replication across persisting lesions can produce multiple alternative alleles in successive cell divisions, thereby generating both multiallelic and combinatorial genetic diversity. The phasing of lesions enables accurate measurement of strand-biased repair processes, quantification of oncogenic selection and fine mapping of sister-chromatid-exchange events. Finally, we demonstrate that lesion segregation is a unifying property of exogenous mutagens, including UV light and chemotherapy agents in human cells and tumours, which has profound implications for the evolution and adaptation of cancer genomes.
@Article{32581361, author = {Aitken SJ and Anderson CJ and Connor F and Pich O and Sundaram V and Feig C and Rayner TF and Lukk M and Aitken S and Luft J and Kentepozidou E and Arnedo-Pac C and Beentjes SV and Davies SE and Drews RM and Ewing A and Kaiser VB and Khamseh A and López-Arribillaga E and Redmond AM and Santoyo-Lopez J and Sentís I and Talmane L and Yates AD and Liver CEC and Semple CA and López-Bigas N and Flicek P and Odom DT and Taylor MS}, title = {Pervasive lesion segregation shapes cancer genome evolution}, journal = {Nature}, volume = {583}, number = {7815}, pages = {265--270}, year = {2020}, doi = {10.1038/s41586-020-2435-1}, howpublished = {Advanced online publication: 24 June 2020}, note = {First posted as a preprint: 8 December 2019}, abstract = {Cancers arise through the acquisition of oncogenic mutations and grow by clonal expansion\textsuperscript{1,2}. Here we reveal that most mutagenic DNA lesions are not resolved into a mutated DNA base pair within a single cell cycle. Instead, DNA lesions segregate, unrepaired, into daughter cells for multiple cell generations, resulting in the chromosome-scale phasing of subsequent mutations. We characterize this process in mutagen-induced mouse liver tumours and show that DNA replication across persisting lesions can produce multiple alternative alleles in successive cell divisions, thereby generating both multiallelic and combinatorial genetic diversity. The phasing of lesions enables accurate measurement of strand-biased repair processes, quantification of oncogenic selection and fine mapping of sister-chromatid-exchange events. Finally, we demonstrate that lesion segregation is a unifying property of exogenous mutagens, including UV light and chemotherapy agents in human cells and tumours, which has profound implications for the evolution and adaptation of cancer genomes.},} - DR Zerbino, A Frankish, P Flicek. Progress, Challenges, and Surprises in Annotating the Human Genome. Annu Rev Genomics Hum Genet 2020;21:55–79. doi:10.1146/annurev-genom-121119-083418
[BibTeX] [Abstract]
Our understanding of the human genome has continuously expanded since its draft publication in 2001. Over the years, novel assays have allowed us to progressively overlay layers of knowledge above the raw sequence of A’s, T’s, G’s, and C’s. The reference human genome sequence is now a complex knowledge base maintained under the shared stewardship of multiple specialist communities. Its complexity stems from the fact that it is simultaneously a template for transcription, a record of evolution, a vehicle for genetics, and a functional molecule. In short, the human genome serves as a frame of reference at the intersection of a diversity of scientific fields. In recent years, the progressive fall in sequencing costs has given increasing importance to the quality of the human reference genome, as hundreds of thousands of individuals are being sequenced yearly, often for clinical applications. Also, novel sequencing-based assays shed light on novel functions of the genome, especially with respect to gene expression regulation. Keeping the human genome annotation up to date and accurate is therefore an ongoing partnership between reference annotation projects and the greater community worldwide. Expected final online publication date for the Annual Review of Genomics and Human Genetics, Volume 21 is August 31, 2020. Please see http://www.annualreviews.org/page/journal/pubdates for revised estimates.
@Article{32421357, author = {Zerbino DR and Frankish A and Flicek P}, title = {Progress, Challenges, and Surprises in Annotating the Human Genome}, journal = {Annu Rev Genomics Hum Genet}, volume = {21}, pages = {55--79}, year = {2020}, doi = {10.1146/annurev-genom-121119-083418}, howpublished = {Advanced online publication: 18 May 2020}, abstract = {Our understanding of the human genome has continuously expanded since its draft publication in 2001. Over the years, novel assays have allowed us to progressively overlay layers of knowledge above the raw sequence of A's, T's, G's, and C's. The reference human genome sequence is now a complex knowledge base maintained under the shared stewardship of multiple specialist communities. Its complexity stems from the fact that it is simultaneously a template for transcription, a record of evolution, a vehicle for genetics, and a functional molecule. In short, the human genome serves as a frame of reference at the intersection of a diversity of scientific fields. In recent years, the progressive fall in sequencing costs has given increasing importance to the quality of the human reference genome, as hundreds of thousands of individuals are being sequenced yearly, often for clinical applications. Also, novel sequencing-based assays shed light on novel functions of the genome, especially with respect to gene expression regulation. Keeping the human genome annotation up to date and accurate is therefore an ongoing partnership between reference annotation projects and the greater community worldwide. Expected final online publication date for the Annual Review of Genomics and Human Genetics, Volume 21 is August 31, 2020. Please see http://www.annualreviews.org/page/journal/pubdates for revised estimates.},} - H Haselimashhadi, JC Mason, V Munoz-Fuentes, F López-Gómez, K Babalola, EF Acar, V Kumar, J White, AM Flenniken, R King, E Straiton, JR Seavitt, A Gaspero, A Garza, AE Christianson, CW Hsu, CL Reynolds, DG Lanza, I Lorenzo, JR Green, JJ Gallegos, R Bohat, RC Samaco, S Veeraragavan, JK Kim, G Miller, H Fuchs, L Garrett, L Becker, YK Kang, D Clary, SY Cho, M Tamura, N Tanaka, KD Soo, A Bezginov, GB About, MF Champy, L Vasseur, S Leblanc, H Meziane, M Selloum, PT Reilly, N Spielmann, H Maier, V Gailus-Durner, T Sorg, M Hiroshi, O Yuichi, JD Heaney, ME Dickinson, W Wolfgang, GP Tocchini-Valentini, KCK Lloyd, C McKerlie, JK Seong, H Yann, de MH Angelis, SDM Brown, D Smedley, P Flicek, AM Mallon, H Parkinson, TF Meehan. Soft Windowing Application to Improve Analysis of High-throughput Phenotyping Data. Bioinformatics 2020;36(5):1492–1500. doi:10.1093/bioinformatics/btz744
[BibTeX] [Abstract]
MOTIVATION: High-throughput phenomic projects generate complex data from small treatment and large control groups that increase the power of the analyses but introduce variation over time. A method is needed to utlize a set of temporally local controls that maximises analytic power while minimising noise from unspecified environmental factors. RESULTS: Here we introduce “soft windowing”, a methodological approach that selects a window of time that includes the most appropriate controls for analysis. Using phenotype data from the International Mouse Phenotyping Consortium (IMPC), adaptive windows were applied such that control data collected proximally to mutants were assigned the maximal weight, while data collected earlier or later had less weight. We applied this method to IMPC data and compared the results with those obtained from a standard non-windowed approach. Validation was performed using a resampling approach in which we demonstrate a 10\% reduction of false positives from 2.5 million analyses. We applied the method to our production analysis pipeline that establishes genotype-phenotype associations by comparing mutant versus control data. We report an increase of 30\% in significant p-values, as well as linkage to 106 versus 99 disease models via phenotype overlap with the soft-windowed and non-windowed approaches, respectively, from a set of 2,082 mutant mouse lines. Our method is generalisable and can benefit large-scale human phenomic projects such as the UK Biobank and the All of Us resources. AVAILABILITY AND IMPLEMENTATION: The method is freely available in the R package SmoothWin, available on CRAN http://CRAN.R-project.org/package=SmoothWin. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
@Article{31591642, author = {Haselimashhadi H and Mason JC and Munoz-Fuentes V and López-Gómez F and Babalola K and Acar EF and Kumar V and White J and Flenniken AM and King R and Straiton E and Seavitt JR and Gaspero A and Garza A and Christianson AE and Hsu CW and Reynolds CL and Lanza DG and Lorenzo I and Green JR and Gallegos JJ and Bohat R and Samaco RC and Veeraragavan S and Kim JK and Miller G and Fuchs H and Garrett L and Becker L and Kang YK and Clary D and Cho SY and Tamura M and Tanaka N and Soo KD and Bezginov A and About GB and Champy MF and Vasseur L and Leblanc S and Meziane H and Selloum M and Reilly PT and Spielmann N and Maier H and Gailus-Durner V and Sorg T and Hiroshi M and Yuichi O and Heaney JD and Dickinson ME and Wolfgang W and Tocchini-Valentini GP and Lloyd KCK and McKerlie C and Seong JK and Yann H and de Angelis MH and Brown SDM and Smedley D and Flicek P and Mallon AM and Parkinson H and Meehan TF}, title = {Soft Windowing Application to Improve Analysis of High-throughput Phenotyping Data}, journal = {Bioinformatics}, volume = {36}, number = {5}, pages = {1492--1500}, year = {2020}, doi = {10.1093/bioinformatics/btz744}, howpublished = {Advanced online publication: 8 October 2019}, note = {First posted as a preprint: 13 June 2019}, abstract = {MOTIVATION: High-throughput phenomic projects generate complex data from small treatment and large control groups that increase the power of the analyses but introduce variation over time. A method is needed to utlize a set of temporally local controls that maximises analytic power while minimising noise from unspecified environmental factors. RESULTS: Here we introduce ``soft windowing'', a methodological approach that selects a window of time that includes the most appropriate controls for analysis. Using phenotype data from the International Mouse Phenotyping Consortium (IMPC), adaptive windows were applied such that control data collected proximally to mutants were assigned the maximal weight, while data collected earlier or later had less weight. We applied this method to IMPC data and compared the results with those obtained from a standard non-windowed approach. Validation was performed using a resampling approach in which we demonstrate a 10\% reduction of false positives from 2.5 million analyses. We applied the method to our production analysis pipeline that establishes genotype-phenotype associations by comparing mutant versus control data. We report an increase of 30\% in significant p-values, as well as linkage to 106 versus 99 disease models via phenotype overlap with the soft-windowed and non-windowed approaches, respectively, from a set of 2,082 mutant mouse lines. Our method is generalisable and can benefit large-scale human phenomic projects such as the UK Biobank and the All of Us resources. AVAILABILITY AND IMPLEMENTATION: The method is freely available in the R package SmoothWin, available on CRAN http://CRAN.R-project.org/package=SmoothWin. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.},} - KCK Lloyd, DJ Adams, G Baynam, AL Beaudet, F Bosch, KM Boycott, RE Braun, M Caulfield, R Cohn, ME Dickinson, MS Dobbie, AM Flenniken, P Flicek, S Galande, X Gao, A Grobler, JD Heaney, Y Herault, de MH Angelis, JR Lupski, S Lyonnet, AM Mallon, F Mammano, CA MacRae, R McInnes, C McKerlie, TF Meehan, SA Murray, LMJ Nutter, Y Obata, H Parkinson, MS Pepper, R Sedlacek, JK Seong, T Shiroishi, D Smedley, G Tocchini-Valentini, D Valle, CL Wang, S Wells, J White, W Wurst, Y Xu, SDM Brown. The Deep Genome Project. Genome Biol 2020;21(1):18. doi:10.1186/s13059-020-1931-9
[BibTeX]@Article{32008577, author = {Lloyd KCK and Adams DJ and Baynam G and Beaudet AL and Bosch F and Boycott KM and Braun RE and Caulfield M and Cohn R and Dickinson ME and Dobbie MS and Flenniken AM and Flicek P and Galande S and Gao X and Grobler A and Heaney JD and Herault Y and de Angelis MH and Lupski JR and Lyonnet S and Mallon AM and Mammano F and MacRae CA and McInnes R and McKerlie C and Meehan TF and Murray SA and Nutter LMJ and Obata Y and Parkinson H and Pepper MS and Sedlacek R and Seong JK and Shiroishi T and Smedley D and Tocchini-Valentini G and Valle D and Wang CL and Wells S and White J and Wurst W and Xu Y and Brown SDM}, title = {The Deep Genome Project}, journal = {Genome Biol}, volume = {21}, number = {1}, pages = {18}, year = {2020}, doi = {10.1186/s13059-020-1931-9}, } - GTEx Consortium. The GTEx Consortium atlas of genetic regulatory effects across human tissues. Science 2020;369(6509):1318–1330. doi:10.1126/science.aaz1776
[BibTeX] [Abstract]
The Genotype-Tissue Expression (GTEx) project was established to characterize genetic effects on the transcriptome across human tissues and to link these regulatory mechanisms to trait and disease associations. Here, we present analyses of the version 8 data, examining 15,201 RNA-sequencing samples from 49 tissues of 838 postmortem donors. We comprehensively characterize genetic associations for gene expression and splicing in cis and trans, showing that regulatory associations are found for almost all genes, and describe the underlying molecular mechanisms and their contribution to allelic heterogeneity and pleiotropy of complex traits. Leveraging the large diversity of tissues, we provide insights into the tissue specificity of genetic effects and show that cell type composition is a key factor in understanding gene regulatory mechanisms in human tissues.
@Article{32913098, author = {{GTEx Consortium}}, title = {The GTEx Consortium atlas of genetic regulatory effects across human tissues}, journal = {Science}, volume = {369}, number = {6509}, pages = {1318--1330}, year = {2020}, doi = {10.1126/science.aaz1776}, note = {First posted as a preprint: 3 October 2019}, abstract = {The Genotype-Tissue Expression (GTEx) project was established to characterize genetic effects on the transcriptome across human tissues and to link these regulatory mechanisms to trait and disease associations. Here, we present analyses of the version 8 data, examining 15,201 RNA-sequencing samples from 49 tissues of 838 postmortem donors. We comprehensively characterize genetic associations for gene expression and splicing in cis and trans, showing that regulatory associations are found for almost all genes, and describe the underlying molecular mechanisms and their contribution to allelic heterogeneity and pleiotropy of complex traits. Leveraging the large diversity of tissues, we provide insights into the tissue specificity of genetic effects and show that cell type composition is a key factor in understanding gene regulatory mechanisms in human tissues.},} - S Fairley, E Lowy-Gallego, E Perry, P Flicek. The International Genome Sample Resource (IGSR) collection of open human genomic variation resources. Nucleic Acids Res 2020;48(D1):D941–D947. doi:10.1093/nar/gkz836
[BibTeX] [Abstract]
To sustain and develop the largest fully open human genomic resources the International Genome Sample Resource (IGSR) (https://www.internationalgenome.org) was established. It is built on the foundation of the 1000 Genomes Project, which created the largest openly accessible catalogue of human genomic variation developed from samples spanning five continents. IGSR (i) maintains access to 1000 Genomes Project resources, (ii) updates 1000 Genomes Project resources to the GRCh38 human reference assembly, (iii) adds new data generated on 1000 Genomes Project cell lines, (iv) shares data from samples with a similarly open consent to increase the number of samples and populations represented in the resources and (v) provides support to users of these resources. Among recent updates are the release of variation calls from 1000 Genomes Project data calculated directly on GRCh38 and the addition of high coverage sequence data for the 2504 samples in the 1000 Genomes Project phase three panel. The data portal, which facilitates web-based exploration of the IGSR resources, has been updated to include samples which were not part of the 1000 Genomes Project and now presents a unified view of data and samples across almost 5000 samples from multiple studies. All data is fully open and publicly accessible.
@Article{31584097, author = {Fairley S and Lowy-Gallego E and Perry E and Flicek P}, title = {The International Genome Sample Resource (IGSR) collection of open human genomic variation resources}, journal = {Nucleic Acids Res}, volume = {48}, number = {D1}, pages = {D941--D947}, year = {2020}, doi = {10.1093/nar/gkz836}, howpublished = {Advanced online publication: 4 October 2019}, abstract = {To sustain and develop the largest fully open human genomic resources the International Genome Sample Resource (IGSR) (https://www.internationalgenome.org) was established. It is built on the foundation of the 1000 Genomes Project, which created the largest openly accessible catalogue of human genomic variation developed from samples spanning five continents. IGSR (i) maintains access to 1000 Genomes Project resources, (ii) updates 1000 Genomes Project resources to the GRCh38 human reference assembly, (iii) adds new data generated on 1000 Genomes Project cell lines, (iv) shares data from samples with a similarly open consent to increase the number of samples and populations represented in the resources and (v) provides support to users of these resources. Among recent updates are the release of variation calls from 1000 Genomes Project data calculated directly on GRCh38 and the addition of high coverage sequence data for the 2504 samples in the 1000 Genomes Project phase three panel. The data portal, which facilitates web-based exploration of the IGSR resources, has been updated to include samples which were not part of the 1000 Genomes Project and now presents a unified view of data and samples across almost 5000 samples from multiple studies. All data is fully open and publicly accessible.},} - NJ Gemmell, K Rutherford, S Prost, M Tollis, D Winter, JR Macey, DL Adelson, A Suh, T Bertozzi, JH Grau, C Organ, PP Gardner, M Muffato, M Patricio, K Billis, FJ Martin, P Flicek, B Petersen, L Kang, P Michalak, TR Buckley, M Wilson, Y Cheng, H Miller, RK Schott, MD Jordan, RD Newcomb, JI Arroyo, N Valenzuela, TA Hore, J Renart, V Peona, CR Peart, VM Warmuth, L Zeng, RD Kortschak, JM Raison, VV Zapata, Z Wu, D Santesmasses, M Mariotti, R Guigó, SM Rupp, VG Twort, N Dussex, H Taylor, H Abe, DM Bond, JM Paterson, DG Mulcahy, VL Gonzalez, CG Barbieri, DP DeMeo, S Pabinger, T Van Stijn, S Clarke, O Ryder, SV Edwards, SL Salzberg, L Anderson, N Nelson, C Stone, TB Ngatiwai. The tuatara genome reveals ancient features of amniote evolution. Nature 2020;584(7821):403–409. doi:10.1038/s41586-020-2561-9
[BibTeX] [Abstract]
The tuatara (Sphenodon punctatus)-the only living member of the reptilian order Rhynchocephalia (Sphenodontia), once widespread across Gondwana\textsuperscript{1,2}-is an iconic species that is endemic to New Zealand\textsuperscript{2,3}. A key link to the now-extinct stem reptiles (from which dinosaurs, modern reptiles, birds and mammals evolved), the tuatara provides key insights into the ancestral amniotes\textsuperscript{2,4}. Here we analyse the genome of the tuatara, which-at approximately 5 Gb-is among the largest of the vertebrate genomes yet assembled. Our analyses of this genome, along with comparisons with other vertebrate genomes, reinforce the uniqueness of the tuatara. Phylogenetic analyses indicate that the tuatara lineage diverged from that of snakes and lizards around 250 million years ago. This lineage also shows moderate rates of molecular evolution, with instances of punctuated evolution. Our genome sequence analysis identifies expansions of proteins, non-protein-coding RNA families and repeat elements, the latter of which show an amalgam of reptilian and mammalian features. The sequencing of the tuatara genome provides a valuable resource for deep comparative analyses of tetrapods, as well as for tuatara biology and conservation. Our study also provides important insights into both the technical challenges and the cultural obligations that are associated with genome sequencing.
@Article{32760000, author = {Gemmell NJ and Rutherford K and Prost S and Tollis M and Winter D and Macey JR and Adelson DL and Suh A and Bertozzi T and Grau JH and Organ C and Gardner PP and Muffato M and Patricio M and Billis K and Martin FJ and Flicek P and Petersen B and Kang L and Michalak P and Buckley TR and Wilson M and Cheng Y and Miller H and Schott RK and Jordan MD and Newcomb RD and Arroyo JI and Valenzuela N and Hore TA and Renart J and Peona V and Peart CR and Warmuth VM and Zeng L and Kortschak RD and Raison JM and Zapata VV and Wu Z and Santesmasses D and Mariotti M and Guigó R and Rupp SM and Twort VG and Dussex N and Taylor H and Abe H and Bond DM and Paterson JM and Mulcahy DG and Gonzalez VL and Barbieri CG and DeMeo DP and Pabinger S and Van Stijn T and Clarke S and Ryder O and Edwards SV and Salzberg SL and Anderson L and Nelson N and Stone C and Ngatiwai TB}, title = {The tuatara genome reveals ancient features of amniote evolution}, journal = {Nature}, volume = {584}, number = {7821}, pages = {403--409}, year = {2020}, doi = {10.1038/s41586-020-2561-9}, howpublished = {Advanced online publication: 5 August 2020}, note = {First posted as a preprint: 8 December 2019}, abstract = {The tuatara (Sphenodon punctatus)-the only living member of the reptilian order Rhynchocephalia (Sphenodontia), once widespread across Gondwana\textsuperscript{1,2}-is an iconic species that is endemic to New Zealand\textsuperscript{2,3}. A key link to the now-extinct stem reptiles (from which dinosaurs, modern reptiles, birds and mammals evolved), the tuatara provides key insights into the ancestral amniotes\textsuperscript{2,4}. Here we analyse the genome of the tuatara, which-at approximately 5 Gb-is among the largest of the vertebrate genomes yet assembled. Our analyses of this genome, along with comparisons with other vertebrate genomes, reinforce the uniqueness of the tuatara. Phylogenetic analyses indicate that the tuatara lineage diverged from that of snakes and lizards around 250 million years ago. This lineage also shows moderate rates of molecular evolution, with instances of punctuated evolution. Our genome sequence analysis identifies expansions of proteins, non-protein-coding RNA families and repeat elements, the latter of which show an amalgam of reptilian and mammalian features. The sequencing of the tuatara genome provides a valuable resource for deep comparative analyses of tetrapods, as well as for tuatara biology and conservation. Our study also provides important insights into both the technical challenges and the cultural obligations that are associated with genome sequencing.},} - C Sisu, P Muir, A Frankish, I Fiddes, M Diekhans, D Thybert, DT Odom, P Flicek, TM Keane, T Hubbard, J Harrow, M Gerstein. Transcriptional activity and strain-specific history of mouse pseudogenes. Nat Commun 2020;11(1):3695. doi:10.1038/s41467-020-17157-w
[BibTeX] [Abstract]
Pseudogenes are ideal markers of genome remodelling. In turn, the mouse is an ideal platform for studying them, particularly with the recent availability of strain-sequencing and transcriptional data. Here, combining both manual curation and automatic pipelines, we present a genome-wide annotation of the pseudogenes in the mouse reference genome and 18 inbred mouse strains (available via the mouse.pseudogene.org resource). We also annotate 165 unitary pseudogenes in mouse, and 303, in human. The overall pseudogene repertoire in mouse is similar to that in human in terms of size, biotype distribution, and family composition (e.g. with GAPDH and ribosomal proteins being the largest families). Notable differences arise in the pseudogene age distribution, with multiple retro-transpositional bursts in mouse evolutionary history and only one in human. Furthermore, in each strain about a fifth of all pseudogenes are unique, reflecting strain-specific evolution. Finally, we find that ~15\% of the mouse pseudogenes are transcribed, and that highly transcribed parent genes tend to give rise to many processed pseudogenes.
@Article{32728065, author = {Sisu C and Muir P and Frankish A and Fiddes I and Diekhans M and Thybert D and Odom DT and Flicek P and Keane TM and Hubbard T and Harrow J and Gerstein M}, title = {Transcriptional activity and strain-specific history of mouse pseudogenes}, journal = {Nat Commun}, volume = {11}, number = {1}, pages = {3695}, year = {2020}, doi = {10.1038/s41467-020-17157-w}, note = {First posted as a preprint: 7 August 2018}, abstract = {Pseudogenes are ideal markers of genome remodelling. In turn, the mouse is an ideal platform for studying them, particularly with the recent availability of strain-sequencing and transcriptional data. Here, combining both manual curation and automatic pipelines, we present a genome-wide annotation of the pseudogenes in the mouse reference genome and 18 inbred mouse strains (available via the mouse.pseudogene.org resource). We also annotate 165 unitary pseudogenes in mouse, and 303, in human. The overall pseudogene repertoire in mouse is similar to that in human in terms of size, biotype distribution, and family composition (e.g. with GAPDH and ribosomal proteins being the largest families). Notable differences arise in the pseudogene age distribution, with multiple retro-transpositional bursts in mouse evolutionary history and only one in human. Furthermore, in each strain about a fifth of all pseudogenes are unique, reflecting strain-specific evolution. Finally, we find that ~15\% of the mouse pseudogenes are transcribed, and that highly transcribed parent genes tend to give rise to many processed pseudogenes.},}
2019
- ME Pettersson, CM Rochus, F Han, J Chen, J Hill, O Wallerman, G Fan, X Hong, Q Xu, H Zhang, S Liu, X Liu, L Haggerty, T Hunt, FJ Martin, P Flicek, I Bunikis, A Folkvord, L Andersson. A chromosome-level assembly of the Atlantic herring genome-detection of a supergene and other signals of selection. Genome Res 2019;29(11):1919–1928. doi:10.1101/gr.253435.119
[BibTeX] [Abstract]
The Atlantic herring is a model species for exploring the genetic basis for ecological adaptation, due to its huge population size and extremely low genetic differentiation at selectively neutral loci. However, such studies have so far been hampered because of a highly fragmented genome assembly. Here, we deliver a chromosome-level genome assembly based on a hybrid approach combining a de novo Pacific Biosciences (PacBio) assembly with Hi-C-supported scaffolding. The assembly comprises 26 autosomes with sizes ranging from 12.4 to 33.1 Mb and a total size, in chromosomes, of 726 Mb, which has been corroborated by a high-resolution linkage map. A comparison between the herring genome assembly with other high-quality assemblies from bony fishes revealed few inter-chromosomal but frequent intra-chromosomal rearrangements. The improved assembly facilitates analysis of previously intractable large-scale structural variation, allowing, for example, the detection of a 7.8-Mb inversion on Chromosome 12 underlying ecological adaptation. This supergene shows strong genetic differentiation between populations. The chromosome-based assembly also markedly improves the interpretation of previously detected signals of selection, allowing us to reveal hundreds of independent loci associated with ecological adaptation.
@Article{31649060, author = {Pettersson ME and Rochus CM and Han F and Chen J and Hill J and Wallerman O and Fan G and Hong X and Xu Q and Zhang H and Liu S and Liu X and Haggerty L and Hunt T and Martin FJ and Flicek P and Bunikis I and Folkvord A and Andersson L}, title = {A chromosome-level assembly of the Atlantic herring genome-detection of a supergene and other signals of selection}, journal = {Genome Res}, volume = {29}, number = {11}, pages = {1919--1928}, year = {2019}, doi = {10.1101/gr.253435.119}, howpublished = {Advanced online publication: 24 October 2019}, note = {First posted as a preprint: 11 June 2019}, abstract = {The Atlantic herring is a model species for exploring the genetic basis for ecological adaptation, due to its huge population size and extremely low genetic differentiation at selectively neutral loci. However, such studies have so far been hampered because of a highly fragmented genome assembly. Here, we deliver a chromosome-level genome assembly based on a hybrid approach combining a de novo Pacific Biosciences (PacBio) assembly with Hi-C-supported scaffolding. The assembly comprises 26 autosomes with sizes ranging from 12.4 to 33.1 Mb and a total size, in chromosomes, of 726 Mb, which has been corroborated by a high-resolution linkage map. A comparison between the herring genome assembly with other high-quality assemblies from bony fishes revealed few inter-chromosomal but frequent intra-chromosomal rearrangements. The improved assembly facilitates analysis of previously intractable large-scale structural variation, allowing, for example, the detection of a 7.8-Mb inversion on Chromosome 12 underlying ecological adaptation. This supergene shows strong genetic differentiation between populations. The chromosome-based assembly also markedly improves the interpretation of previously detected signals of selection, allowing us to reveal hundreds of independent loci associated with ecological adaptation.},} - C Berthelot, J Clarke, T Desvignes, H William Detrich, P Flicek, LS Peck, M Peters, JH Postlethwait, MS Clark. Adaptation of Proteins to the Cold in Antarctic Fish: A Role for Methionine. Genome Biol Evol 2019;11(1):220–231. doi:10.1093/gbe/evy262
[BibTeX] [Abstract]
The evolution of antifreeze glycoproteins has enabled notothenioid fish to flourish in the freezing waters of the Southern Ocean. Whereas successful at the biodiversity level to life in the cold, paradoxically at the cellular level these stenothermal animals have problems producing, folding, and degrading proteins at their ambient temperatures of -1.86 °C. In this first multi-species transcriptome comparison of the amino acid composition of notothenioid proteins with temperate teleost proteins, we show that, unlike psychrophilic bacteria, Antarctic fish provide little evidence for the mass alteration of protein amino acid composition to enhance protein folding and reduce protein denaturation in the cold. The exception was the significant overrepresentation of positions where leucine in temperate fish proteins was replaced by methionine in the notothenioid orthologues. We hypothesize that these extra methionines have been preferentially assimilated into the genome to act as redox sensors in the highly oxygenated waters of the Southern Ocean. This redox hypothesis is supported by analyses of notothenioids showing enrichment of genes associated with responses to environmental stress, particularly reactive oxygen species. So overall, although notothenioid fish show cold-associated problems with protein homeostasis, they may have modified only a selected number of biochemical pathways to work efficiently below 0 °C. Even a slight warming of the Southern Ocean might disrupt the critical functions of this handful of key pathways with considerable impacts for the functioning of this ecosystem in the future.
@Article{30496401, author = {Berthelot C and Clarke J and Desvignes T and William Detrich H and Flicek P and Peck LS and Peters M and Postlethwait JH and Clark MS}, title = {Adaptation of Proteins to the Cold in Antarctic Fish: A Role for Methionine}, journal = {Genome Biol Evol}, volume = {11}, number = {1}, pages = {220--231}, year = {2019}, doi = {10.1093/gbe/evy262}, howpublished = {Advanced online publication: 29 November 2018}, note = {First posted as a preprint: 9 August 2018}, abstract = {The evolution of antifreeze glycoproteins has enabled notothenioid fish to flourish in the freezing waters of the Southern Ocean. Whereas successful at the biodiversity level to life in the cold, paradoxically at the cellular level these stenothermal animals have problems producing, folding, and degrading proteins at their ambient temperatures of -1.86 °C. In this first multi-species transcriptome comparison of the amino acid composition of notothenioid proteins with temperate teleost proteins, we show that, unlike psychrophilic bacteria, Antarctic fish provide little evidence for the mass alteration of protein amino acid composition to enhance protein folding and reduce protein denaturation in the cold. The exception was the significant overrepresentation of positions where leucine in temperate fish proteins was replaced by methionine in the notothenioid orthologues. We hypothesize that these extra methionines have been preferentially assimilated into the genome to act as redox sensors in the highly oxygenated waters of the Southern Ocean. This redox hypothesis is supported by analyses of notothenioids showing enrichment of genes associated with responses to environmental stress, particularly reactive oxygen species. So overall, although notothenioid fish show cold-associated problems with protein homeostasis, they may have modified only a selected number of biochemical pathways to work efficiently below 0 °C. Even a slight warming of the Southern Ocean might disrupt the critical functions of this handful of key pathways with considerable impacts for the functioning of this ecosystem in the future.},} - B Benjelloun, F Boyer, I Streeter, W Zamani, S Engelen, A Alberti, FJ Alberto, M BenBati, M Ibnelbachyr, M Chentouf, A Bechchari, HR Rezaei, S Naderi, A Stella, A Chikhi, L Clarke, J Kijas, P Flicek, P Taberlet, F Pompanon. An evaluation of sequencing coverage and genotyping strategies to assess neutral and adaptive diversity. Mol Ecol Resour 2019;19(6):1497–1515. doi:10.1111/1755-0998.13070
[BibTeX] [Abstract]
Whole genome sequences (WGS) greatly increase our ability to precisely infer population genetic parameters, demographic processes, and selection signatures. However WGS can still be not affordable for a representative number of individuals/populations. In this context, our goal was to assess the efficiency of several SNP genotyping strategies by testing their ability to accurately estimate parameters describing neutral diversity and to detect signatures of selection. We analysed 110 WGS at 12X coverage for four different species, i.e. sheep, goats and their wild counterparts. From these data we generated 946 datasets corresponding to random panels of 1K to 5M variants, commercial SNP chips and exome capture, for sample sizes of 5 to 48 individuals. We also extracted low-coverage genome re-sequencing of 1X, 2X and 5X by randomly sub-sampling reads from the 12X re-sequencing data. Globally, 5K to 10K random variants were enough for an accurate estimation of genome diversity. Conversely, commercial panels and exome capture displayed strong ascertainment biases. Besides the characterization of neutral diversity, the detection of the signature of selection and the accurate estimation of linkage disequilibrium required high-density panels of at least 1M variants. Finally, genotype likelihoods increased the quality of variant calling from low coverage re-sequencing but proportions of incorrect genotypes remained substantial, especially for heterozygote sites. Whole genome re-sequencing coverage of at least 5X appeared to be necessary for accurate assessment of genomic variations. These results have implications for studies seeking to deploy low-density SNP collections or genome scans across genetically diverse populations/species showing similar genetic characteristics and patterns of LD decay for a wide variety of purposes. This article is protected by copyright. All rights reserved.
@Article{31359622, author = {Benjelloun B and Boyer F and Streeter I and Zamani W and Engelen S and Alberti A and Alberto FJ and BenBati M and Ibnelbachyr M and Chentouf M and Bechchari A and Rezaei HR and Naderi S and Stella A and Chikhi A and Clarke L and Kijas J and Flicek P and Taberlet P and Pompanon F}, title = {An evaluation of sequencing coverage and genotyping strategies to assess neutral and adaptive diversity}, journal = {Mol Ecol Resour}, volume = {19}, number = {6}, pages = {1497--1515}, year = {2019}, doi = {10.1111/1755-0998.13070}, howpublished = {Advanced online publication: 29 July 2019}, abstract = {Whole genome sequences (WGS) greatly increase our ability to precisely infer population genetic parameters, demographic processes, and selection signatures. However WGS can still be not affordable for a representative number of individuals/populations. In this context, our goal was to assess the efficiency of several SNP genotyping strategies by testing their ability to accurately estimate parameters describing neutral diversity and to detect signatures of selection. We analysed 110 WGS at 12X coverage for four different species, i.e. sheep, goats and their wild counterparts. From these data we generated 946 datasets corresponding to random panels of 1K to 5M variants, commercial SNP chips and exome capture, for sample sizes of 5 to 48 individuals. We also extracted low-coverage genome re-sequencing of 1X, 2X and 5X by randomly sub-sampling reads from the 12X re-sequencing data. Globally, 5K to 10K random variants were enough for an accurate estimation of genome diversity. Conversely, commercial panels and exome capture displayed strong ascertainment biases. Besides the characterization of neutral diversity, the detection of the signature of selection and the accurate estimation of linkage disequilibrium required high-density panels of at least 1M variants. Finally, genotype likelihoods increased the quality of variant calling from low coverage re-sequencing but proportions of incorrect genotypes remained substantial, especially for heterozygote sites. Whole genome re-sequencing coverage of at least 5X appeared to be necessary for accurate assessment of genomic variations. These results have implications for studies seeking to deploy low-density SNP collections or genome scans across genetically diverse populations/species showing similar genetic characteristics and patterns of LD decay for a wide variety of purposes. This article is protected by copyright. All rights reserved.},} - G Yi, ATJ Wierenga, F Petraglia, P Narang, EM Janssen-Megens, A Mandoli, A Merkel, K Berentsen, B Kim, F Matarese, AA Singh, E Habibi, KHM Prange, AB Mulder, JH Jansen, L Clarke, S Heath, van der BA Reijden, P Flicek, ML Yaspo, I Gut, C Bock, JJ Schuringa, L Altucci, E Vellenga, HG Stunnenberg, JHA Martens. Chromatin-Based Classification of Genetically Heterogeneous AMLs into Two Distinct Subtypes with Diverse Stemness Phenotypes. Cell Rep 2019;26(4):1059–1069.e6. doi:10.1016/j.celrep.2018.12.098
[BibTeX] [Abstract]
Global investigation of histone marks in acute myeloid leukemia (AML) remains limited. Analyses of 38 AML samples through integrated transcriptional and chromatin mark analysis exposes 2 major subtypes. One subtype is dominated by patients with NPM1 mutations or MLL-fusion genes, shows activation of the regulatory pathways involving HOX-family genes as targets, and displays high self-renewal capacity and stemness. The second subtype is enriched for RUNX1 or spliceosome mutations, suggesting potential interplay between the 2 aberrations, and mainly depends on IRF family regulators. Cellular consequences in prognosis predict a relatively worse outcome for the first subtype. Our integrated profiling establishes a rich resource to probe AML subtypes on the basis of expression and chromatin data.
@Article{30673601, author = {Yi G and Wierenga ATJ and Petraglia F and Narang P and Janssen-Megens EM and Mandoli A and Merkel A and Berentsen K and Kim B and Matarese F and Singh AA and Habibi E and Prange KHM and Mulder AB and Jansen JH and Clarke L and Heath S and van der Reijden BA and Flicek P and Yaspo ML and Gut I and Bock C and Schuringa JJ and Altucci L and Vellenga E and Stunnenberg HG and Martens JHA}, title = {Chromatin-Based Classification of Genetically Heterogeneous AMLs into Two Distinct Subtypes with Diverse Stemness Phenotypes}, journal = {Cell Rep}, volume = {26}, number = {4}, pages = {1059--1069.e6}, year = {2019}, doi = {10.1016/j.celrep.2018.12.098}, abstract = {Global investigation of histone marks in acute myeloid leukemia (AML) remains limited. Analyses of 38 AML samples through integrated transcriptional and chromatin mark analysis exposes 2 major subtypes. One subtype is dominated by patients with NPM1 mutations or MLL-fusion genes, shows activation of the regulatory pathways involving HOX-family genes as targets, and displays high self-renewal capacity and stemness. The second subtype is enriched for RUNX1 or spliceosome mutations, suggesting potential interplay between the 2 aberrations, and mainly depends on IRF family regulators. Cellular consequences in prognosis predict a relatively worse outcome for the first subtype. Our integrated profiling establishes a rich resource to probe AML subtypes on the basis of expression and chromatin data.},} - F Cunningham, P Achuthan, W Akanni, J Allen, MR Amode, IM Armean, R Bennett, J Bhai, K Billis, S Boddu, C Cummins, C Davidson, KJ Dodiya, A Gall, CG Girón, L Gil, T Grego, L Haggerty, E Haskell, T Hourlier, OG Izuogu, SH Janacek, T Juettemann, M Kay, MR Laird, I Lavidas, Z Liu, JE Loveland, JC Marugán, T Maurel, AC McMahon, B Moore, J Morales, JM Mudge, M Nuhn, D Ogeh, A Parker, A Parton, M Patricio, AI Abdul Salam, BM Schmitt, H Schuilenburg, D Sheppard, H Sparrow, E Stapleton, M Szuba, K Taylor, G Threadgold, A Thormann, A Vullo, B Walts, A Winterbottom, A Zadissa, M Chakiachvili, A Frankish, SE Hunt, M Kostadima, N Langridge, FJ Martin, M Muffato, E Perry, M Ruffier, DM Staines, SJ Trevanion, BL Aken, AD Yates, DR Zerbino, P Flicek. Ensembl 2019. Nucleic Acids Res 2019;47(D1):D745–D751. doi:10.1093/nar/gky1113
[BibTeX] [Abstract]
The Ensembl project (https://www.ensembl.org) makes key genomic data sets available to the entire scientific community without restrictions. Ensembl seeks to be a fundamental resource driving scientific progress by creating, maintaining and updating reference genome annotation and comparative genomics resources. This year we describe our new and expanded gene, variant and comparative annotation capabilities, which led to a 50\% increase in the number of vertebrate genomes we support. We have also doubled the number of available human variants and added regulatory regions for many mouse cell types and developmental stages. Our data sets and tools are available via the Ensembl website as well as a through a RESTful webservice, Perl application programming interface and as data files for download.
@Article{30407521, author = {Cunningham F and Achuthan P and Akanni W and Allen J and Amode MR and Armean IM and Bennett R and Bhai J and Billis K and Boddu S and Cummins C and Davidson C and Dodiya KJ and Gall A and Girón CG and Gil L and Grego T and Haggerty L and Haskell E and Hourlier T and Izuogu OG and Janacek SH and Juettemann T and Kay M and Laird MR and Lavidas I and Liu Z and Loveland JE and Marugán JC and Maurel T and McMahon AC and Moore B and Morales J and Mudge JM and Nuhn M and Ogeh D and Parker A and Parton A and Patricio M and Abdul Salam AI and Schmitt BM and Schuilenburg H and Sheppard D and Sparrow H and Stapleton E and Szuba M and Taylor K and Threadgold G and Thormann A and Vullo A and Walts B and Winterbottom A and Zadissa A and Chakiachvili M and Frankish A and Hunt SE and Kostadima M and Langridge N and Martin FJ and Muffato M and Perry E and Ruffier M and Staines DM and Trevanion SJ and Aken BL and Yates AD and Zerbino DR and Flicek P}, title = {Ensembl 2019}, journal = {Nucleic Acids Res}, volume = {47}, number = {D1}, pages = {D745--D751}, year = {2019}, doi = {10.1093/nar/gky1113}, howpublished = {Advanced online publication: 8 November 2018}, abstract = {The Ensembl project (https://www.ensembl.org) makes key genomic data sets available to the entire scientific community without restrictions. Ensembl seeks to be a fundamental resource driving scientific progress by creating, maintaining and updating reference genome annotation and comparative genomics resources. This year we describe our new and expanded gene, variant and comparative annotation capabilities, which led to a 50\% increase in the number of vertebrate genomes we support. We have also doubled the number of available human variants and added regulatory regions for many mouse cell types and developmental stages. Our data sets and tools are available via the Ensembl website as well as a through a RESTful webservice, Perl application programming interface and as data files for download.},} - M Fiume, M Cupak, S Keenan, J Rambla, de la S Torre, SOM Dyke, AJ Brookes, K Carey, D Lloyd, P Goodhand, M Haeussler, M Baudis, H Stockinger, L Dolman, I Lappalainen, J Törnroos, M Linden, JD Spalding, S Ur-Rehman, A Page, P Flicek, S Sherry, D Haussler, S Varma, G Saunders, S Scollen. Federated discovery and sharing of genomic data using Beacons. Nat Biotechnol 2019;37(3):220–224. doi:10.1038/s41587-019-0046-x
[BibTeX]@Article{30833764, author = {Fiume M and Cupak M and Keenan S and Rambla J and de la Torre S and Dyke SOM and Brookes AJ and Carey K and Lloyd D and Goodhand P and Haeussler M and Baudis M and Stockinger H and Dolman L and Lappalainen I and Törnroos J and Linden M and Spalding JD and Ur-Rehman S and Page A and Flicek P and Sherry S and Haussler D and Varma S and Saunders G and Scollen S}, title = {Federated discovery and sharing of genomic data using Beacons}, journal = {Nat Biotechnol}, volume = {37}, number = {3}, pages = {220--224}, year = {2019}, doi = {10.1038/s41587-019-0046-x}, } - A Frankish, M Diekhans, AM Ferreira, R Johnson, I Jungreis, J Loveland, JM Mudge, C Sisu, J Wright, J Armstrong, I Barnes, A Berry, A Bignell, S Carbonell Sala, J Chrast, F Cunningham, T Di Domenico, S Donaldson, IT Fiddes, C García Girón, JM Gonzalez, T Grego, M Hardy, T Hourlier, T Hunt, OG Izuogu, J Lagarde, FJ Martin, L Martínez, S Mohanan, P Muir, FCP Navarro, A Parker, B Pei, F Pozo, M Ruffier, BM Schmitt, E Stapleton, MM Suner, I Sycheva, B Uszczynska-Ratajczak, J Xu, A Yates, D Zerbino, Y Zhang, B Aken, JS Choudhary, M Gerstein, R Guigó, TJP Hubbard, M Kellis, B Paten, A Reymond, ML Tress, P Flicek. GENCODE reference annotation for the human and mouse genomes. Nucleic Acids Res 2019;47(D1):D766–D773. doi:10.1093/nar/gky955
[BibTeX] [Abstract]
The accurate identification and description of the genes in the human and mouse genomes is a fundamental requirement for high quality analysis of data informing both genome biology and clinical genomics. Over the last 15 years, the GENCODE consortium has been producing reference quality gene annotations to provide this foundational resource. The GENCODE consortium includes both experimental and computational biology groups who work together to improve and extend the GENCODE gene annotation. Specifically, we generate primary data, create bioinformatics tools and provide analysis to support the work of expert manual gene annotators and automated gene annotation pipelines. In addition, manual and computational annotation workflows use any and all publicly available data and analysis, along with the research literature to identify and characterise gene loci to the highest standard. GENCODE gene annotations are accessible via the Ensembl and UCSC Genome Browsers, the Ensembl FTP site, Ensembl Biomart, Ensembl Perl and REST APIs as well as https://www.gencodegenes.org.
@Article{30357393, author = {Frankish A and Diekhans M and Ferreira AM and Johnson R and Jungreis I and Loveland J and Mudge JM and Sisu C and Wright J and Armstrong J and Barnes I and Berry A and Bignell A and Carbonell Sala S and Chrast J and Cunningham F and Di Domenico T and Donaldson S and Fiddes IT and García Girón C and Gonzalez JM and Grego T and Hardy M and Hourlier T and Hunt T and Izuogu OG and Lagarde J and Martin FJ and Martínez L and Mohanan S and Muir P and Navarro FCP and Parker A and Pei B and Pozo F and Ruffier M and Schmitt BM and Stapleton E and Suner MM and Sycheva I and Uszczynska-Ratajczak B and Xu J and Yates A and Zerbino D and Zhang Y and Aken B and Choudhary JS and Gerstein M and Guigó R and Hubbard TJP and Kellis M and Paten B and Reymond A and Tress ML and Flicek P}, title = {GENCODE reference annotation for the human and mouse genomes}, journal = {Nucleic Acids Res}, volume = {47}, number = {D1}, pages = {D766--D773}, year = {2019}, doi = {10.1093/nar/gky955}, howpublished = {Advanced online publication: 24 October 2018}, abstract = {The accurate identification and description of the genes in the human and mouse genomes is a fundamental requirement for high quality analysis of data informing both genome biology and clinical genomics. Over the last 15 years, the GENCODE consortium has been producing reference quality gene annotations to provide this foundational resource. The GENCODE consortium includes both experimental and computational biology groups who work together to improve and extend the GENCODE gene annotation. Specifically, we generate primary data, create bioinformatics tools and provide analysis to support the work of expert manual gene annotators and automated gene annotation pipelines. In addition, manual and computational annotation workflows use any and all publicly available data and analysis, along with the research literature to identify and characterise gene loci to the highest standard. GENCODE gene annotations are accessible via the Ensembl and UCSC Genome Browsers, the Ensembl FTP site, Ensembl Biomart, Ensembl Perl and REST APIs as well as https://www.gencodegenes.org.},} - SA Hardwick, A Joglekar, P Flicek, A Frankish, HU Tilgner. Getting the Entire Message: Progress in Isoform Sequencing. Front Genet 2019;10:709. doi:10.3389/fgene.2019.00709
[BibTeX] [Abstract]
The advent of second-generation sequencing and its application to RNA sequencing have revolutionized the field of genomics by allowing quantification of gene expression, as well as the definition of transcription start/end sites, exons, splice sites and RNA editing sites. However, due to the sequencing of fragments of cDNAs, these methods have not given a reliable picture of complete RNA isoforms. Third-generation sequencing has filled this gap and allows end-to-end sequencing of entire RNA/cDNA molecules. This approach to transcriptomics has been a “niche” technology for a couple of years but now is becoming mainstream with many different applications. Here, we review the background and progress made to date in this rapidly growing field. We start by reviewing the progressive realization that alternative splicing is omnipresent. We then focus on long-noncoding RNA isoforms and the distinct combination patterns of exons in noncoding and coding genes. We consider the implications of the recent technologies of direct RNA sequencing and single-cell isoform RNA sequencing. Finally, we discuss the parameters that define the success of long-read RNA sequencing experiments and strategies commonly used to make the most of such data.
@Article{31475029, author = {Hardwick SA and Joglekar A and Flicek P and Frankish A and Tilgner HU}, title = {Getting the Entire Message: Progress in Isoform Sequencing}, journal = {Front Genet}, volume = {10}, pages = {709}, year = {2019}, doi = {10.3389/fgene.2019.00709}, abstract = {The advent of second-generation sequencing and its application to RNA sequencing have revolutionized the field of genomics by allowing quantification of gene expression, as well as the definition of transcription start/end sites, exons, splice sites and RNA editing sites. However, due to the sequencing of fragments of cDNAs, these methods have not given a reliable picture of complete RNA isoforms. Third-generation sequencing has filled this gap and allows end-to-end sequencing of entire RNA/cDNA molecules. This approach to transcriptomics has been a ``niche'' technology for a couple of years but now is becoming mainstream with many different applications. Here, we review the background and progress made to date in this rapidly growing field. We start by reviewing the progressive realization that alternative splicing is omnipresent. We then focus on long-noncoding RNA isoforms and the distinct combination patterns of exons in noncoding and coding genes. We consider the implications of the recent technologies of direct RNA sequencing and single-cell isoform RNA sequencing. Finally, we discuss the parameters that define the success of long-read RNA sequencing experiments and strategies commonly used to make the most of such data.},} - G Saunders, M Baudis, R Becker, S Beltran, C Béroud, E Birney, C Brooksbank, S Brunak, den M Van Bulcke, R Drysdale, S Capella-Gutierrez, P Flicek, F Florindi, P Goodhand, I Gut, J Heringa, P Holub, J Hooyberghs, N Juty, TM Keane, JO Korbel, I Lappalainen, B Leskosek, G Matthijs, MT Mayrhofer, A Metspalu, A Navarro, S Newhouse, T Nyrönen, A Page, B Persson, A Palotie, H Parkinson, J Rambla, D Salgado, E Steinfelder, MA Swertz, A Valencia, S Varma, N Blomberg, S Scollen. Leveraging European infrastructures to access 1 million human genomes by 2022. Nat Rev Genet 2019;20(11):693–701. doi:10.1038/s41576-019-0156-9
[BibTeX] [Abstract]
Human genomics is undergoing a step change from being a predominantly research-driven activity to one driven through health care as many countries in Europe now have nascent precision medicine programmes. To maximize the value of the genomic data generated, these data will need to be shared between institutions and across countries. In recognition of this challenge, 21 European countries recently signed a declaration to transnationally share data on at least 1 million human genomes by 2022. In this Roadmap, we identify the challenges of data sharing across borders and demonstrate that European research infrastructures are well-positioned to support the rapid implementation of widespread genomic data access.
@Article{31455890, author = {Saunders G and Baudis M and Becker R and Beltran S and Béroud C and Birney E and Brooksbank C and Brunak S and Van den Bulcke M and Drysdale R and Capella-Gutierrez S and Flicek P and Florindi F and Goodhand P and Gut I and Heringa J and Holub P and Hooyberghs J and Juty N and Keane TM and Korbel JO and Lappalainen I and Leskosek B and Matthijs G and Mayrhofer MT and Metspalu A and Navarro A and Newhouse S and Nyrönen T and Page A and Persson B and Palotie A and Parkinson H and Rambla J and Salgado D and Steinfelder E and Swertz MA and Valencia A and Varma S and Blomberg N and Scollen S}, title = {Leveraging European infrastructures to access 1 million human genomes by 2022}, journal = {Nat Rev Genet}, volume = {20}, number = {11}, pages = {693--701}, year = {2019}, doi = {10.1038/s41576-019-0156-9}, howpublished = {Advanced online publication: 27 August 2019}, abstract = {Human genomics is undergoing a step change from being a predominantly research-driven activity to one driven through health care as many countries in Europe now have nascent precision medicine programmes. To maximize the value of the genomic data generated, these data will need to be shared between institutions and across countries. In recognition of this challenge, 21 European countries recently signed a declaration to transnationally share data on at least 1 million human genomes by 2022. In this Roadmap, we identify the challenges of data sharing across borders and demonstrate that European research infrastructures are well-positioned to support the rapid implementation of widespread genomic data access.},} - MJP Chaisson, AD Sanders, X Zhao, A Malhotra, D Porubsky, T Rausch, EJ Gardner, OL Rodriguez, L Guo, RL Collins, X Fan, J Wen, RE Handsaker, S Fairley, ZN Kronenberg, X Kong, F Hormozdiari, D Lee, AM Wenger, AR Hastie, D Antaki, T Anantharaman, PA Audano, H Brand, S Cantsilieris, H Cao, E Cerveira, C Chen, X Chen, CS Chin, Z Chong, NT Chuang, CC Lambert, DM Church, L Clarke, A Farrell, J Flores, T Galeev, DU Gorkin, M Gujral, V Guryev, WH Heaton, J Korlach, S Kumar, JY Kwon, ET Lam, JE Lee, J Lee, WP Lee, SP Lee, S Li, P Marks, K Viaud-Martinez, S Meiers, KM Munson, FCP Navarro, BJ Nelson, C Nodzak, A Noor, S Kyriazopoulou-Panagiotopoulou, AWC Pang, Y Qiu, G Rosanio, M Ryan, A Stütz, DCJ Spierings, A Ward, AE Welch, M Xiao, W Xu, C Zhang, Q Zhu, X Zheng-Bradley, E Lowy, S Yakneen, S McCarroll, G Jun, L Ding, CL Koh, B Ren, P Flicek, K Chen, MB Gerstein, PY Kwok, PM Lansdorp, GT Marth, J Sebat, X Shi, A Bashir, K Ye, SE Devine, ME Talkowski, RE Mills, T Marschall, JO Korbel, EE Eichler, C Lee. Multi-platform discovery of haplotype-resolved structural variation in human genomes. Nat Commun 2019;10(1):1784. doi:10.1038/s41467-018-08148-z
[BibTeX] [Abstract]
The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs (≥50 bp) per genome. We also discover 156 inversions per genome and 58 of the inversions intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a three to sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The methods and the dataset presented serve as a gold standard for the scientific community allowing us to make recommendations for maximizing structural variation sensitivity for future genome sequencing studies.
@Article{30992455, author = {Chaisson MJP and Sanders AD and Zhao X and Malhotra A and Porubsky D and Rausch T and Gardner EJ and Rodriguez OL and Guo L and Collins RL and Fan X and Wen J and Handsaker RE and Fairley S and Kronenberg ZN and Kong X and Hormozdiari F and Lee D and Wenger AM and Hastie AR and Antaki D and Anantharaman T and Audano PA and Brand H and Cantsilieris S and Cao H and Cerveira E and Chen C and Chen X and Chin CS and Chong Z and Chuang NT and Lambert CC and Church DM and Clarke L and Farrell A and Flores J and Galeev T and Gorkin DU and Gujral M and Guryev V and Heaton WH and Korlach J and Kumar S and Kwon JY and Lam ET and Lee JE and Lee J and Lee WP and Lee SP and Li S and Marks P and Viaud-Martinez K and Meiers S and Munson KM and Navarro FCP and Nelson BJ and Nodzak C and Noor A and Kyriazopoulou-Panagiotopoulou S and Pang AWC and Qiu Y and Rosanio G and Ryan M and Stütz A and Spierings DCJ and Ward A and Welch AE and Xiao M and Xu W and Zhang C and Zhu Q and Zheng-Bradley X and Lowy E and Yakneen S and McCarroll S and Jun G and Ding L and Koh CL and Ren B and Flicek P and Chen K and Gerstein MB and Kwok PY and Lansdorp PM and Marth GT and Sebat J and Shi X and Bashir A and Ye K and Devine SE and Talkowski ME and Mills RE and Marschall T and Korbel JO and Eichler EE and Lee C}, title = {Multi-platform discovery of haplotype-resolved structural variation in human genomes}, journal = {Nat Commun}, volume = {10}, number = {1}, pages = {1784}, year = {2019}, doi = {10.1038/s41467-018-08148-z}, note = {First posted as a preprint: 23 September 2017}, abstract = {The incomplete identification of structural variants (SVs) from whole-genome sequencing data limits studies of human genetic diversity and disease association. Here, we apply a suite of long-read, short-read, strand-specific sequencing technologies, optical mapping, and variant discovery algorithms to comprehensively analyze three trios to define the full spectrum of human genetic variation in a haplotype-resolved manner. We identify 818,054 indel variants (<50 bp) and 27,622 SVs (≥50 bp) per genome. We also discover 156 inversions per genome and 58 of the inversions intersect with the critical regions of recurrent microdeletion and microduplication syndromes. Taken together, our SV callsets represent a three to sevenfold increase in SV detection compared to most standard high-throughput sequencing studies, including those from the 1000 Genomes Project. The methods and the dataset presented serve as a gold standard for the scientific community allowing us to make recommendations for maximizing structural variation sensitivity for future genome sequencing studies.},} - CA Steward, J Roovers, MM Suner, JM Gonzalez, B Uszczynska-Ratajczak, D Pervouchine, S Fitzgerald, M Viola, H Stamberger, FF Hamdan, B Ceulemans, P Leroy, C Nava, A Lepine, E Tapanari, D Keiller, S Abbs, A Sanchis-Juan, D Grozeva, AS Rogers, M Diekhans, R Guigó, R Petryszak, BA Minassian, G Cavalleri, D Vitsios, S Petrovski, J Harrow, P Flicek, F Lucy Raymond, NJ Lench, P Jonghe, JM Mudge, S Weckhuysen, SM Sisodiya, A Frankish. Re-annotation of 191 developmental and epileptic encephalopathy-associated genes unmasks de novo variants in SCN1A. NPJ Genom Med 2019;4:31. doi:10.1038/s41525-019-0106-7
[BibTeX] [Abstract]
The developmental and epileptic encephalopathies (DEE) are a group of rare, severe neurodevelopmental disorders, where even the most thorough sequencing studies leave 60-65\% of patients without a molecular diagnosis. Here, we explore the incompleteness of transcript models used for exome and genome analysis as one potential explanation for a lack of current diagnoses. Therefore, we have updated the GENCODE gene annotation for 191 epilepsy-associated genes, using human brain-derived transcriptomic libraries and other data to build 3,550 putative transcript models. Our annotations increase the transcriptional `footprint' of these genes by over 674 kb. Using SCN1A as a case study, due to its close phenotype/genotype correlation with Dravet syndrome, we screened 122 people with Dravet syndrome or a similar phenotype with a panel of exon sequences representing eight established genes and identified two de novo SCN1A variants that now - through improved gene annotation - are ascribed to residing among our exons. These two (from 122 screened people, 1.6\%) molecular diagnoses carry significant clinical implications. Furthermore, we identified a previously classified SCN1A intronic Dravet syndrome-associated variant that now lies within a deeply conserved exon. Our findings illustrate the potential gains of thorough gene annotation in improving diagnostic yields for genetic disorders.
@Article{31814998, author = {Steward CA and Roovers J and Suner MM and Gonzalez JM and Uszczynska-Ratajczak B and Pervouchine D and Fitzgerald S and Viola M and Stamberger H and Hamdan FF and Ceulemans B and Leroy P and Nava C and Lepine A and Tapanari E and Keiller D and Abbs S and Sanchis-Juan A and Grozeva D and Rogers AS and Diekhans M and Guigó R and Petryszak R and Minassian BA and Cavalleri G and Vitsios D and Petrovski S and Harrow J and Flicek P and Lucy Raymond F and Lench NJ and Jonghe P and Mudge JM and Weckhuysen S and Sisodiya SM and Frankish A}, title = {Re-annotation of 191 developmental and epileptic encephalopathy-associated genes unmasks de novo variants in SCN1A}, journal = {NPJ Genom Med}, volume = {4}, pages = {31}, year = {2019}, doi = {10.1038/s41525-019-0106-7}, note = {First posted as a preprint: 30 May 2019}, abstract = {The developmental and epileptic encephalopathies (DEE) are a group of rare, severe neurodevelopmental disorders, where even the most thorough sequencing studies leave 60-65\% of patients without a molecular diagnosis. Here, we explore the incompleteness of transcript models used for exome and genome analysis as one potential explanation for a lack of current diagnoses. Therefore, we have updated the GENCODE gene annotation for 191 epilepsy-associated genes, using human brain-derived transcriptomic libraries and other data to build 3,550 putative transcript models. Our annotations increase the transcriptional `footprint' of these genes by over 674 kb. Using SCN1A as a case study, due to its close phenotype/genotype correlation with Dravet syndrome, we screened 122 people with Dravet syndrome or a similar phenotype with a panel of exon sequences representing eight established genes and identified two de novo SCN1A variants that now - through improved gene annotation - are ascribed to residing among our exons. These two (from 122 screened people, 1.6\%) molecular diagnoses carry significant clinical implications. Furthermore, we identified a previously classified SCN1A intronic Dravet syndrome-associated variant that now lies within a deeply conserved exon. Our findings illustrate the potential gains of thorough gene annotation in improving diagnostic yields for genetic disorders.},} - A Buniello, JAL MacArthur, M Cerezo, LW Harris, J Hayhurst, C Malangone, A McMahon, J Morales, E Mountjoy, E Sollis, D Suveges, O Vrousgou, PL Whetzel, R Amode, JA Guillen, HS Riat, SJ Trevanion, P Hall, H Junkins, P Flicek, T Burdett, LA Hindorff, F Cunningham, H Parkinson. The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Res 2019;47(D1):D1005–D1012. doi:10.1093/nar/gky1120
[BibTeX] [Abstract]
The GWAS Catalog delivers a high-quality curated collection of all published genome-wide association studies enabling investigations to identify causal variants, understand disease mechanisms, and establish targets for novel therapies. The scope of the Catalog has also expanded to targeted and exome arrays with 1000 new associations added for these technologies. As of September 2018, the Catalog contains 5687 GWAS comprising 71673 variant-trait associations from 3567 publications. New content includes 284 full P-value summary statistics datasets for genome-wide and new targeted array studies, representing 6 × 109 individual variant-trait statistics. In the last 12 months, the Catalog's user interface was accessed by ∼90000 unique users who viewed >1 million pages. We have improved data access with the release of a new RESTful API to support high-throughput programmatic access, an improved web interface and a new summary statistics database. Summary statistics provision is supported by a new format proposed as a community standard for summary statistics data representation. This format was derived from our experience in standardizing heterogeneous submissions, mapping formats and in harmonizing content. Availability: https://www.ebi.ac.uk/gwas/.
@Article{30445434, author = {Buniello A and MacArthur JAL and Cerezo M and Harris LW and Hayhurst J and Malangone C and McMahon A and Morales J and Mountjoy E and Sollis E and Suveges D and Vrousgou O and Whetzel PL and Amode R and Guillen JA and Riat HS and Trevanion SJ and Hall P and Junkins H and Flicek P and Burdett T and Hindorff LA and Cunningham F and Parkinson H}, title = {The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019}, journal = {Nucleic Acids Res}, volume = {47}, number = {D1}, pages = {D1005--D1012}, year = {2019}, doi = {10.1093/nar/gky1120}, howpublished = {Advanced online publication: 16 November 2018}, abstract = {The GWAS Catalog delivers a high-quality curated collection of all published genome-wide association studies enabling investigations to identify causal variants, understand disease mechanisms, and establish targets for novel therapies. The scope of the Catalog has also expanded to targeted and exome arrays with 1000 new associations added for these technologies. As of September 2018, the Catalog contains 5687 GWAS comprising 71673 variant-trait associations from 3567 publications. New content includes 284 full P-value summary statistics datasets for genome-wide and new targeted array studies, representing 6 × 109 individual variant-trait statistics. In the last 12 months, the Catalog's user interface was accessed by ∼90000 unique users who viewed >1 million pages. We have improved data access with the release of a new RESTful API to support high-throughput programmatic access, an improved web interface and a new summary statistics database. Summary statistics provision is supported by a new format proposed as a community standard for summary statistics data representation. This format was derived from our experience in standardizing heterogeneous submissions, mapping formats and in harmonizing content. Availability: https://www.ebi.ac.uk/gwas/.},} - S Kongsstovu Í, SO Mikalsen, EÍ Homrum, JA Jacobsen, P Flicek, HA Dahl. Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly. Sci Rep 2019;9(1):17716. doi:10.1038/s41598-019-54151-9
[BibTeX] [Abstract]
Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.
@Article{31776409, author = {Í Kongsstovu S and Mikalsen SO and Homrum EÍ and Jacobsen JA and Flicek P and Dahl HA}, title = {Using long and linked reads to improve an Atlantic herring (Clupea harengus) genome assembly}, journal = {Sci Rep}, volume = {9}, number = {1}, pages = {17716}, year = {2019}, doi = {10.1038/s41598-019-54151-9}, abstract = {Atlantic herring (Clupea harengus) is one of the most abundant fish species in the world. It is an important economical and nutritional resource, as well as a crucial part of the North Atlantic ecosystem. In 2016, a draft herring genome assembly was published. Being a species of such importance, we sought to independently verify and potentially improve the herring genome assembly. We sequenced the herring genome generating paired-end, mate-pair, linked and long reads. Three assembly versions of the herring genome were generated based on a de novo assembly (A1), which was scaffolded using linked and long reads (A2) and then merged with the previously published assembly (A3). The resulting assemblies were compared using parameters describing the size, fragmentation, correctness, and completeness of the assemblies. Results showed that the A2 assembly was less fragmented, more complete and more correct than A1. A3 showed improvement in fragmentation and correctness compared with A2 and the published assembly but was slightly less complete than the published assembly. Thus, we here confirmed the previously published herring assembly, and made improvements by further scaffolding the assembly and removing low-quality sequences using linked and long reads and merging of assemblies.},} - E Lowy-Gallego, S Fairley, X Zheng-Bradley, M Ruffier, L Clarke, P Flicek, GPC 1000. Variant calling on the GRCh38 assembly with the data from phase three of the 1000 Genomes Project. Wellcome Open Res 2019;4:50. doi:10.12688/wellcomeopenres.15126.2
[BibTeX] [Abstract]
We present a set of biallelic SNVs and INDELs, from 2,548 samples spanning 26 populations from the 1000 Genomes Project, called de novo on GRCh38. We believe this will be a useful reference resource for those using GRCh38. It represents an improvement over the ``lift-overs'' of the 1000 Genomes Project data that have been available to date by encompassing all of the GRCh38 primary assembly autosomes and pseudo-autosomal regions, including novel, medically relevant loci. Here, we describe how the data set was created and benchmark our call set against that produced by the final phase of the 1000 Genomes Project on GRCh37 and the lift-over of that data to GRCh38.
@Article{32175479, author = {Lowy-Gallego E and Fairley S and Zheng-Bradley X and Ruffier M and Clarke L and Flicek P and 1000 GPC}, title = {Variant calling on the GRCh38 assembly with the data from phase three of the 1000 Genomes Project}, journal = {Wellcome Open Res}, volume = {4}, pages = {50}, year = {2019}, doi = {10.12688/wellcomeopenres.15126.2}, note = {First posted as a preprint: 11 March 2019}, abstract = {We present a set of biallelic SNVs and INDELs, from 2,548 samples spanning 26 populations from the 1000 Genomes Project, called de novo on GRCh38. We believe this will be a useful reference resource for those using GRCh38. It represents an improvement over the ``lift-overs'' of the 1000 Genomes Project data that have been available to date by encompassing all of the GRCh38 primary assembly autosomes and pseudo-autosomal regions, including novel, medically relevant loci. Here, we describe how the data set was created and benchmark our call set against that produced by the final phase of the 1000 Genomes Project on GRCh37 and the lift-over of that data to GRCh38.},}
2018
- J Morales, D Welter, EH Bowler, M Cerezo, LW Harris, AC McMahon, P Hall, HA Junkins, A Milano, E Hastings, C Malangone, A Buniello, T Burdett, P Flicek, H Parkinson, F Cunningham, LA Hindorff, JAL MacArthur. A standardized framework for representation of ancestry data in genomics studies, with application to the NHGRI-EBI GWAS Catalog. Genome Biol 2018;19(1):489. doi:10.1186/s13059-018-1396-2
[BibTeX] [Abstract]
The accurate description of ancestry is essential to interpret, access, and integrate human genomics data, and to ensure that these benefit individuals from all ancestral backgrounds. However, there are no established guidelines for the representation of ancestry information. Here we describe a framework for the accurate and standardized description of sample ancestry, and validate it by application to the NHGRI-EBI GWAS Catalog. We confirm known biases and gaps in diversity, and find that African and Hispanic or Latin American ancestry populations contribute a disproportionately high number of associations. It is our hope that widespread adoption of this framework will lead to improved analysis, interpretation, and integration of human genomics data.
@Article{29448949, author = {Morales J and Welter D and Bowler EH and Cerezo M and Harris LW and McMahon AC and Hall P and Junkins HA and Milano A and Hastings E and Malangone C and Buniello A and Burdett T and Flicek P and Parkinson H and Cunningham F and Hindorff LA and MacArthur JAL}, title = {A standardized framework for representation of ancestry data in genomics studies, with application to the NHGRI-EBI GWAS Catalog}, journal = {Genome Biol}, volume = {19}, number = {1}, pages = {489}, year = {2018}, doi = {10.1186/s13059-018-1396-2}, note = {First posted as a preprint: 21 April 2017}, abstract = {The accurate description of ancestry is essential to interpret, access, and integrate human genomics data, and to ensure that these benefit individuals from all ancestral backgrounds. However, there are no established guidelines for the representation of ancestry information. Here we describe a framework for the accurate and standardized description of sample ancestry, and validate it by application to the NHGRI-EBI GWAS Catalog. We confirm known biases and gaps in diversity, and find that African and Hispanic or Latin American ancestry populations contribute a disproportionately high number of associations. It is our hope that widespread adoption of this framework will lead to improved analysis, interpretation, and integration of human genomics data.},} - M Kolmogorov, J Armstrong, BJ Raney, I Streeter, M Dunn, F Yang, D Odom, P Flicek, TM Keane, D Thybert, B Paten, S Pham. Chromosome assembly of large and complex genomes using multiple references. Genome Res 2018;28(11):1720–1732. doi:10.1101/gr.236273.118
[BibTeX] [Abstract]
Despite the rapid development of sequencing technologies, the assembly of mammalian-scale genomes into complete chromosomes remains one of the most challenging problems in bioinformatics. To help address this difficulty, we developed Ragout 2, a reference-assisted assembly tool that works for large and complex genomes. By taking one or more target assemblies (generated from an NGS assembler) and one or multiple related reference genomes, Ragout 2 infers the evolutionary relationships between the genomes and builds the final assemblies using a genome rearrangement approach. By using Ragout 2, we transformed NGS assemblies of 16 laboratory mouse strains into sets of complete chromosomes, leaving <5\% of sequence unlocalized per set. Various benchmarks, including PCR testing and realigning of long Pacific Biosciences (PacBio) reads, suggest only a small number of structural errors in the final assemblies, comparable with direct assembly approaches. We applied Ragout 2 to the Mus caroli and Mus pahari genomes, which exhibit karyotype-scale variations compared with other genomes from the Muridae family. Chromosome painting maps confirmed most large-scale rearrangements that Ragout 2 detected. We applied Ragout 2 to improve draft sequences of three ape genomes that have recently been published. Ragout 2 transformed three sets of contigs (generated using PacBio reads only) into chromosome-scale assemblies with accuracy comparable to chromosome assemblies generated in the original study using BioNano maps, Hi-C, BAC clones, and FISH.
@Article{30341161, author = {Kolmogorov M and Armstrong J and Raney BJ and Streeter I and Dunn M and Yang F and Odom D and Flicek P and Keane TM and Thybert D and Paten B and Pham S}, title = {Chromosome assembly of large and complex genomes using multiple references}, journal = {Genome Res}, volume = {28}, number = {11}, pages = {1720--1732}, year = {2018}, doi = {10.1101/gr.236273.118}, howpublished = {Advanced online publication: 19 October 2018}, note = {First posted as a preprint: 19 November 2016}, abstract = {Despite the rapid development of sequencing technologies, the assembly of mammalian-scale genomes into complete chromosomes remains one of the most challenging problems in bioinformatics. To help address this difficulty, we developed Ragout 2, a reference-assisted assembly tool that works for large and complex genomes. By taking one or more target assemblies (generated from an NGS assembler) and one or multiple related reference genomes, Ragout 2 infers the evolutionary relationships between the genomes and builds the final assemblies using a genome rearrangement approach. By using Ragout 2, we transformed NGS assemblies of 16 laboratory mouse strains into sets of complete chromosomes, leaving <5\% of sequence unlocalized per set. Various benchmarks, including PCR testing and realigning of long Pacific Biosciences (PacBio) reads, suggest only a small number of structural errors in the final assemblies, comparable with direct assembly approaches. We applied Ragout 2 to the Mus caroli and Mus pahari genomes, which exhibit karyotype-scale variations compared with other genomes from the Muridae family. Chromosome painting maps confirmed most large-scale rearrangements that Ragout 2 detected. We applied Ragout 2 to improve draft sequences of three ape genomes that have recently been published. Ragout 2 transformed three sets of contigs (generated using PacBio reads only) into chromosome-scale assemblies with accuracy comparable to chromosome assemblies generated in the original study using BioNano maps, Hi-C, BAC clones, and FISH.},} - F Petraglia, AA Singh, V Carafa, A Nebbioso, M Conte, L Scisciola, S Valente, A Baldi, A Mandoli, VB Petrizzi, C Ingenito, S De Falco, V Cicatiello, I Apicella, EM Janssen-Megens, B Kim, G Yi, C Logie, S Heath, M Ruvo, ATJ Wierenga, P Flicek, ML Yaspo, V Della Valle, O Bernard, S Tomassi, E Novellino, A Feoli, G Sbardella, I Gut, E Vellenga, HG Stunnenberg, A Mai, JHA Martens, L Altucci. Combined HAT/EZH2 modulation leads to cancer-selective cell death. Oncotarget 2018;9(39):25630–25646. doi:10.18632/oncotarget.25428
[BibTeX] [Abstract]
Epigenetic alterations have been associated with both pathogenesis and progression of cancer. By screening of library compounds, we identified a novel hybrid epi-drug MC2884, a HAT/EZH2 inhibitor, able to induce
@Article{29876013, author = {Petraglia F and Singh AA and Carafa V and Nebbioso A and Conte M and Scisciola L and Valente S and Baldi A and Mandoli A and Petrizzi VB and Ingenito C and De Falco S and Cicatiello V and Apicella I and Janssen-Megens EM and Kim B and Yi G and Logie C and Heath S and Ruvo M and Wierenga ATJ and Flicek P and Yaspo ML and Della Valle V and Bernard O and Tomassi S and Novellino E and Feoli A and Sbardella G and Gut I and Vellenga E and Stunnenberg HG and Mai A and Martens JHA and Altucci L}, title = {Combined HAT/EZH2 modulation leads to cancer-selective cell death}, journal = {Oncotarget}, volume = {9}, number = {39}, pages = {25630--25646}, year = {2018}, doi = {10.18632/oncotarget.25428}, abstract = {Epigenetic alterations have been associated with both pathogenesis and progression of cancer. By screening of library compounds, we identified a novel hybrid epi-drug MC2884, a HAT/EZH2 inhibitor, able to induce},} - C Berthelot, D Villar, JE Horvath, DT Odom, P Flicek. Complexity and conservation of regulatory landscapes underlie evolutionary resilience of mammalian gene expression. Nat Ecol Evol 2018;2(1):152–163. doi:10.1038/s41559-017-0377-2
[BibTeX] [Abstract]
To gain insight into how mammalian gene expression is controlled by rapidly evolving regulatory elements, we jointly analysed promoter and enhancer activity with downstream transcription levels in liver samples from 15 species. Genes associated with complex regulatory landscapes generally exhibit high expression levels that remain evolutionarily stable. While the number of regulatory elements is the key driver of transcriptional output and resilience, regulatory conservation matters: elements active across mammals most effectively stabilize gene expression. In contrast, recently evolved enhancers typically contribute weakly, consistent with their high evolutionary plasticity. These effects are observed across the entire mammalian clade and are robust to potential confounders, such as the gene expression level. Using liver as a representative somatic tissue, our results illuminate how the evolutionary stability of gene expression is profoundly entwined with both the number and conservation of surrounding promoters and enhancers.
@Article{29180706, author = {Berthelot C and Villar D and Horvath JE and Odom DT and Flicek P}, title = {Complexity and conservation of regulatory landscapes underlie evolutionary resilience of mammalian gene expression}, journal = {Nat Ecol Evol}, volume = {2}, number = {1}, pages = {152--163}, year = {2018}, doi = {10.1038/s41559-017-0377-2}, howpublished = {Advanced online publication: 27 November 2017}, note = {First posted as a preprint: 7 April 2017}, abstract = {To gain insight into how mammalian gene expression is controlled by rapidly evolving regulatory elements, we jointly analysed promoter and enhancer activity with downstream transcription levels in liver samples from 15 species. Genes associated with complex regulatory landscapes generally exhibit high expression levels that remain evolutionarily stable. While the number of regulatory elements is the key driver of transcriptional output and resilience, regulatory conservation matters: elements active across mammals most effectively stabilize gene expression. In contrast, recently evolved enhancers typically contribute weakly, consistent with their high evolutionary plasticity. These effects are observed across the entire mammalian clade and are robust to potential confounders, such as the gene expression level. Using liver as a representative somatic tissue, our results illuminate how the evolutionary stability of gene expression is profoundly entwined with both the number and conservation of surrounding promoters and enhancers.},} - FJ Alberto, F Boyer, P Orozco-terWengel, I Streeter, B Servin, de P Villemereuil, B Benjelloun, P Librado, F Biscarini, L Colli, M Barbato, W Zamani, A Alberti, S Engelen, A Stella, S Joost, P Ajmone-Marsan, R Negrini, L Orlando, HR Rezaei, S Naderi, L Clarke, P Flicek, P Wincker, E Coissac, J Kijas, G Tosser-Klopp, A Chikhi, MW Bruford, P Taberlet, F Pompanon. Convergent genomic signatures of domestication in sheep and goats. Nat Commun 2018;9(1):813. doi:10.1038/s41467-018-03206-y
[BibTeX] [Abstract]
The evolutionary basis of domestication has been a longstanding question and its genetic architecture is becoming more tractable as more domestic species become genome-enabled. Before becoming established worldwide, sheep and goats were domesticated in the fertile crescent 10,500 years before present (YBP) where their wild relatives remain. Here we sequence the genomes of wild Asiatic mouflon and Bezoar ibex in the sheep and goat domestication center and compare their genomes with that of domestics from local, traditional, and improved breeds. Among the genomic regions carrying selective sweeps differentiating domestic breeds from wild populations, which are associated among others to genes involved in nervous system, immunity and productivity traits, 20 are common to Capra and Ovis. The patterns of selection vary between species, suggesting that while common targets of selection related to domestication and improvement exist, different solutions have arisen to achieve similar phenotypic end-points within these closely related livestock species.
@Article{29511174, author = {Alberto FJ and Boyer F and Orozco-terWengel P and Streeter I and Servin B and de Villemereuil P and Benjelloun B and Librado P and Biscarini F and Colli L and Barbato M and Zamani W and Alberti A and Engelen S and Stella A and Joost S and Ajmone-Marsan P and Negrini R and Orlando L and Rezaei HR and Naderi S and Clarke L and Flicek P and Wincker P and Coissac E and Kijas J and Tosser-Klopp G and Chikhi A and Bruford MW and Taberlet P and Pompanon F}, title = {Convergent genomic signatures of domestication in sheep and goats}, journal = {Nat Commun}, volume = {9}, number = {1}, pages = {813}, year = {2018}, doi = {10.1038/s41467-018-03206-y}, abstract = {The evolutionary basis of domestication has been a longstanding question and its genetic architecture is becoming more tractable as more domestic species become genome-enabled. Before becoming established worldwide, sheep and goats were domesticated in the fertile crescent 10,500 years before present (YBP) where their wild relatives remain. Here we sequence the genomes of wild Asiatic mouflon and Bezoar ibex in the sheep and goat domestication center and compare their genomes with that of domestics from local, traditional, and improved breeds. Among the genomic regions carrying selective sweeps differentiating domestic breeds from wild populations, which are associated among others to genes involved in nervous system, immunity and productivity traits, 20 are common to Capra and Ovis. The patterns of selection vary between species, suggesting that while common targets of selection related to domestication and improvement exist, different solutions have arisen to achieve similar phenotypic end-points within these closely related livestock species.},} - SJ Aitken, X Ibarra-Soria, E Kentepozidou, P Flicek, C Feig, JC Marioni, DT Odom. CTCF maintains regulatory homeostasis of cancer pathways. Genome Biol 2018;19(1):106. doi:10.1186/s13059-018-1484-3
[BibTeX] [Abstract]
BACKGROUND: CTCF binding to DNA helps partition the mammalian genome into discrete structural and regulatory domains. Complete removal of CTCF from mammalian cells causes catastrophic genome dysregulation, likely due to widespread collapse of 3D chromatin looping and alterations to inter- and intra-TAD interactions within the nucleus. In contrast, Ctcf hemizygous mice with lifelong reduction of CTCF expression are viable, albeit with increased cancer incidence. Here, we exploit chronic Ctcf hemizygosity to reveal its homeostatic roles in maintaining genome function and integrity. RESULTS: We find that Ctcf hemizygous cells show modest but robust changes in almost a thousand sites of genomic CTCF occupancy; these are enriched for lower affinity binding events with weaker evolutionary conservation across the mouse lineage. Furthermore, we observe dysregulation of the expression of several hundred genes, which are concentrated in cancer-related pathways, and are caused by changes in transcriptional regulation. Chromatin structure is preserved but some loop interactions are destabilized; these are often found around differentially expressed genes and their enhancers. Importantly, the transcriptional alterations identified in vitro are recapitulated in mouse tumors and also in human cancers. CONCLUSIONS: This multi-dimensional genomic and epigenomic profiling of a Ctcf hemizygous mouse model system shows that chronic depletion of CTCF dysregulates steady-state gene expression by subtly altering transcriptional regulation, changes which can also be observed in primary tumors.
@Article{30086769, author = {Aitken SJ and Ibarra-Soria X and Kentepozidou E and Flicek P and Feig C and Marioni JC and Odom DT}, title = {CTCF maintains regulatory homeostasis of cancer pathways}, journal = {Genome Biol}, volume = {19}, number = {1}, pages = {106}, year = {2018}, doi = {10.1186/s13059-018-1484-3}, abstract = {BACKGROUND: CTCF binding to DNA helps partition the mammalian genome into discrete structural and regulatory domains. Complete removal of CTCF from mammalian cells causes catastrophic genome dysregulation, likely due to widespread collapse of 3D chromatin looping and alterations to inter- and intra-TAD interactions within the nucleus. In contrast, Ctcf hemizygous mice with lifelong reduction of CTCF expression are viable, albeit with increased cancer incidence. Here, we exploit chronic Ctcf hemizygosity to reveal its homeostatic roles in maintaining genome function and integrity. RESULTS: We find that Ctcf hemizygous cells show modest but robust changes in almost a thousand sites of genomic CTCF occupancy; these are enriched for lower affinity binding events with weaker evolutionary conservation across the mouse lineage. Furthermore, we observe dysregulation of the expression of several hundred genes, which are concentrated in cancer-related pathways, and are caused by changes in transcriptional regulation. Chromatin structure is preserved but some loop interactions are destabilized; these are often found around differentially expressed genes and their enhancers. Importantly, the transcriptional alterations identified in vitro are recapitulated in mouse tumors and also in human cancers. CONCLUSIONS: This multi-dimensional genomic and epigenomic profiling of a Ctcf hemizygous mouse model system shows that chronic depletion of CTCF dysregulates steady-state gene expression by subtly altering transcriptional regulation, changes which can also be observed in primary tumors.},} - L Grassi, F Pourfarzad, S Ullrich, A Merkel, F Were, E Carrillo-de-Santa-Pau, G Yi, IH Hiemstra, ATJ Tool, E Mul, J Perner, E Janssen-Megens, K Berentsen, H Kerstens, E Habibi, M Gut, ML Yaspo, M Linser, E Lowy, A Datta, L Clarke, P Flicek, M Vingron, D Roos, van den TK Berg, S Heath, D Rico, M Frontini, M Kostadima, I Gut, A Valencia, WH Ouwehand, HG Stunnenberg, JHA Martens, TW Kuijpers. Dynamics of Transcription Regulation in Human Bone Marrow Myeloid Differentiation to Mature Blood Neutrophils. Cell Rep 2018;24(10):2784–2794. doi:10.1016/j.celrep.2018.08.018
[BibTeX] [Abstract]
Neutrophils are short-lived blood cells that play a critical role in host defense against infections. To better comprehend neutrophil functions and their regulation, we provide a complete epigenetic overview, assessing important functional features of their differentiation stages from bone marrow-residing progenitors to mature circulating cells. Integration of chromatin modifications, methylation, and transcriptome dynamics reveals an enforced regulation of differentiation, for cellular functions such as release of proteases, respiratory burst, cell cycle regulation, and apoptosis. We observe an early establishment of the cytotoxic capability, while the signaling components that activate these antimicrobial mechanisms are transcribed at later stages, outside the bone marrow, thus preventing toxic effects in the bone marrow niche. Altogether, these data reveal how the developmental dynamics of the chromatin landscape orchestrate the daily production of a large number of neutrophils required for innate host defense and provide a comprehensive overview of differentiating human neutrophils.
@Article{30184510, author = {Grassi L and Pourfarzad F and Ullrich S and Merkel A and Were F and Carrillo-de-Santa-Pau E and Yi G and Hiemstra IH and Tool ATJ and Mul E and Perner J and Janssen-Megens E and Berentsen K and Kerstens H and Habibi E and Gut M and Yaspo ML and Linser M and Lowy E and Datta A and Clarke L and Flicek P and Vingron M and Roos D and van den Berg TK and Heath S and Rico D and Frontini M and Kostadima M and Gut I and Valencia A and Ouwehand WH and Stunnenberg HG and Martens JHA and Kuijpers TW}, title = {Dynamics of Transcription Regulation in Human Bone Marrow Myeloid Differentiation to Mature Blood Neutrophils}, journal = {Cell Rep}, volume = {24}, number = {10}, pages = {2784--2794}, year = {2018}, doi = {10.1016/j.celrep.2018.08.018}, note = {First posted as a preprint: 5 April 2018}, abstract = {Neutrophils are short-lived blood cells that play a critical role in host defense against infections. To better comprehend neutrophil functions and their regulation, we provide a complete epigenetic overview, assessing important functional features of their differentiation stages from bone marrow-residing progenitors to mature circulating cells. Integration of chromatin modifications, methylation, and transcriptome dynamics reveals an enforced regulation of differentiation, for cellular functions such as release of proteases, respiratory burst, cell cycle regulation, and apoptosis. We observe an early establishment of the cytotoxic capability, while the signaling components that activate these antimicrobial mechanisms are transcribed at later stages, outside the bone marrow, thus preventing toxic effects in the bone marrow niche. Altogether, these data reveal how the developmental dynamics of the chromatin landscape orchestrate the daily production of a large number of neutrophils required for innate host defense and provide a comprehensive overview of differentiating human neutrophils.},} - DR Zerbino, P Achuthan, W Akanni, MR Amode, D Barrell, J Bhai, K Billis, C Cummins, A Gall, CG Girón, L Gil, L Gordon, L Haggerty, E Haskell, T Hourlier, OG Izuogu, SH Janacek, T Juettemann, JK To, MR Laird, I Lavidas, Z Liu, JE Loveland, T Maurel, W McLaren, B Moore, J Mudge, DN Murphy, V Newman, M Nuhn, D Ogeh, CK Ong, A Parker, M Patricio, HS Riat, H Schuilenburg, D Sheppard, H Sparrow, K Taylor, A Thormann, A Vullo, B Walts, A Zadissa, A Frankish, SE Hunt, M Kostadima, N Langridge, FJ Martin, M Muffato, E Perry, M Ruffier, DM Staines, SJ Trevanion, BL Aken, F Cunningham, A Yates, P Flicek. Ensembl 2018. Nucleic Acids Res 2018;46(D1):D754–D761. doi:10.1093/nar/gkx1098
[BibTeX] [Abstract]
The Ensembl project has been aggregating, processing, integrating and redistributing genomic datasets since the initial releases of the draft human genome, with the aim of accelerating genomics research through rapid open distribution of public data. Large amounts of raw data are thus transformed into knowledge, which is made available via a multitude of channels, in particular our browser (http://www.ensembl.org). Over time, we have expanded in multiple directions. First, our resources describe multiple fields of genomics, in particular gene annotation, comparative genomics, genetics and epigenomics. Second, we cover a growing number of genome assemblies; Ensembl Release 90 contains exactly 100. Third, our databases feed simultaneously into an array of services designed around different use cases, ranging from quick browsing to genome-wide bioinformatic analysis. We present here the latest developments of the Ensembl project, with a focus on managing an increasing number of assemblies, supporting efforts in genome interpretation and improving our browser.
@Article{29155950, author = {Zerbino DR and Achuthan P and Akanni W and Amode MR and Barrell D and Bhai J and Billis K and Cummins C and Gall A and Girón CG and Gil L and Gordon L and Haggerty L and Haskell E and Hourlier T and Izuogu OG and Janacek SH and Juettemann T and To JK and Laird MR and Lavidas I and Liu Z and Loveland JE and Maurel T and McLaren W and Moore B and Mudge J and Murphy DN and Newman V and Nuhn M and Ogeh D and Ong CK and Parker A and Patricio M and Riat HS and Schuilenburg H and Sheppard D and Sparrow H and Taylor K and Thormann A and Vullo A and Walts B and Zadissa A and Frankish A and Hunt SE and Kostadima M and Langridge N and Martin FJ and Muffato M and Perry E and Ruffier M and Staines DM and Trevanion SJ and Aken BL and Cunningham F and Yates A and Flicek P}, title = {Ensembl 2018}, journal = {Nucleic Acids Res}, volume = {46}, number = {D1}, pages = {D754--D761}, year = {2018}, doi = {10.1093/nar/gkx1098}, howpublished = {Advanced online publication: 16 November 2017}, abstract = {The Ensembl project has been aggregating, processing, integrating and redistributing genomic datasets since the initial releases of the draft human genome, with the aim of accelerating genomics research through rapid open distribution of public data. Large amounts of raw data are thus transformed into knowledge, which is made available via a multitude of channels, in particular our browser (http://www.ensembl.org). Over time, we have expanded in multiple directions. First, our resources describe multiple fields of genomics, in particular gene annotation, comparative genomics, genetics and epigenomics. Second, we cover a growing number of genome assemblies; Ensembl Release 90 contains exactly 100. Third, our databases feed simultaneously into an array of services designed around different use cases, ranging from quick browsing to genome-wide bioinformatic analysis. We present here the latest developments of the Ensembl project, with a focus on managing an increasing number of assemblies, supporting efforts in genome interpretation and improving our browser.},} - SE Hunt, W McLaren, L Gil, A Thormann, H Schuilenburg, D Sheppard, A Parton, IM Armean, SJ Trevanion, P Flicek, F Cunningham. Ensembl variation resources. Database (Oxford) 2018;2018:bay119. doi:10.1093/database/bay119
[BibTeX] [Abstract]
The major goal of sequencing humans and many other species is to understand the link between genomic variation, phenotype and disease. There are numerous valuable and well-established variation resources, but collating and making sense of non-homogeneous, often large-scale data sets from disparate sources remains a challenge. Without a systematic catalogue of these data and appropriate query and annotation tools, understanding the genome sequence of an individual and assessing their disease risk is impossible. In Ensembl, we substantially solve this problem: we develop methods to facilitate data integration and broad access; aggregate information in a consistent manner and make it available a variety of standard formats, both visually and programmatically; build analysis pipelines to compare variants to comprehensive genomic annotation sets; and make all tools and data publicly available.
@Article{30576484, author = {Hunt SE and McLaren W and Gil L and Thormann A and Schuilenburg H and Sheppard D and Parton A and Armean IM and Trevanion SJ and Flicek P and Cunningham F}, title = {Ensembl variation resources}, journal = {Database (Oxford)}, volume = {2018}, pages = {bay119}, year = {2018}, doi = {10.1093/database/bay119}, howpublished = {Advanced online publication: 6 November 2018}, abstract = {The major goal of sequencing humans and many other species is to understand the link between genomic variation, phenotype and disease. There are numerous valuable and well-established variation resources, but collating and making sense of non-homogeneous, often large-scale data sets from disparate sources remains a challenge. Without a systematic catalogue of these data and appropriate query and annotation tools, understanding the genome sequence of an individual and assessing their disease risk is impossible. In Ensembl, we substantially solve this problem: we develop methods to facilitate data integration and broad access; aggregate information in a consistent manner and make it available a variety of standard formats, both visually and programmatically; build analysis pipelines to compare variants to comprehensive genomic annotation sets; and make all tools and data publicly available.},} - PW Harrison, J Fan, D Richardson, L Clarke, D Zerbino, G Cochrane, AL Archibald, CJ Schmidt, P Flicek. FAANG, establishing metadata standards, validation and best practices for the farmed and companion animal community. Anim Genet 2018;49(6):520–526. doi:10.1111/age.12736
[BibTeX] [Abstract]
The Functional Annotation of ANimal Genomes (FAANG) project aims, through a coordinated international effort, to provide high quality functional annotation of animal genomes with an initial focus on farmed and companion animals. A key goal of the initiative is to ensure high quality and rich supporting metadata to describe the project's animals, specimens, cell cultures and experimental assays. By defining rich sample and experimental metadata standards and promoting best practices in data descriptions, deposition and openness, FAANG champions higher quality and reusability of published datasets. FAANG has established a Data Coordination Centre, which sits at the heart of the Metadata and Data Sharing Committee. It continues to evolve the metadata standards, support submissions and, crucially, create powerful and accessible tools to support deposition and validation of metadata. FAANG conforms to the findable, accessible, interoperable, and reusable (FAIR) data principles, with high quality, open access and functionally interlinked data. In addition to data generated by FAANG members and specific FAANG projects, existing datasets that meet the main-or more permissive legacy-standards are incorporated into a central, focused, functional data resource portal for the entire farmed and companion animal community. Through clear and effective metadata standards, validation and conversion software, combined with promotion of best practices in metadata implementation, FAANG aims to maximise effectiveness and inter-comparability of assay data. This supports the community to create a rich genome-to-phenotype resource and promotes continuing improvements in animal data standards as a whole.
@Article{30311252, author = {Harrison PW and Fan J and Richardson D and Clarke L and Zerbino D and Cochrane G and Archibald AL and Schmidt CJ and Flicek P}, title = {FAANG, establishing metadata standards, validation and best practices for the farmed and companion animal community}, journal = {Anim Genet}, volume = {49}, number = {6}, pages = {520--526}, year = {2018}, doi = {10.1111/age.12736}, howpublished = {Advanced online publication: 12 October 2018}, abstract = {The Functional Annotation of ANimal Genomes (FAANG) project aims, through a coordinated international effort, to provide high quality functional annotation of animal genomes with an initial focus on farmed and companion animals. A key goal of the initiative is to ensure high quality and rich supporting metadata to describe the project's animals, specimens, cell cultures and experimental assays. By defining rich sample and experimental metadata standards and promoting best practices in data descriptions, deposition and openness, FAANG champions higher quality and reusability of published datasets. FAANG has established a Data Coordination Centre, which sits at the heart of the Metadata and Data Sharing Committee. It continues to evolve the metadata standards, support submissions and, crucially, create powerful and accessible tools to support deposition and validation of metadata. FAANG conforms to the findable, accessible, interoperable, and reusable (FAIR) data principles, with high quality, open access and functionally interlinked data. In addition to data generated by FAANG members and specific FAANG projects, existing datasets that meet the main-or more permissive legacy-standards are incorporated into a central, focused, functional data resource portal for the entire farmed and companion animal community. Through clear and effective metadata standards, validation and conversion software, combined with promotion of best practices in metadata implementation, FAANG aims to maximise effectiveness and inter-comparability of assay data. This supports the community to create a rich genome-to-phenotype resource and promotes continuing improvements in animal data standards as a whole.},} - W Spooner, W McLaren, T Slidel, DK Finch, R Butler, J Campbell, L Eghobamien, D Rider, CM Kiefer, MJ Robinson, C Hardman, F Cunningham, T Vaughan, P Flicek, CC Huntington. Haplosaurus computes protein haplotypes for use in precision drug design. Nat Commun 2018;9(1):4128. doi:10.1038/s41467-018-06542-1
[BibTeX] [Abstract]
Selecting the most appropriate protein sequences is critical for precision drug design. Here we describe Haplosaurus, a bioinformatic tool for computation of protein haplotypes. Haplosaurus computes protein haplotypes from pre-existing chromosomally-phased genomic variation data. Integration into the Ensembl resource provides rapid and detailed protein haplotypes retrieval. Using Haplosaurus, we build a database of unique protein haplotypes from the 1000 Genomes dataset reflecting real-world protein sequence variability and their prevalence. For one in seven genes, their most common protein haplotype differs from the reference sequence and a similar number differs on their most common haplotype between human populations. Three case studies show how knowledge of the range of commonly encountered protein forms predicted in populations leads to insights into therapeutic efficacy. Haplosaurus and its associated database is expected to find broad applications in many disciplines using protein sequences and particularly impactful for therapeutics design.
@Article{30297836, author = {Spooner W and McLaren W and Slidel T and Finch DK and Butler R and Campbell J and Eghobamien L and Rider D and Kiefer CM and Robinson MJ and Hardman C and Cunningham F and Vaughan T and Flicek P and Huntington CC}, title = {Haplosaurus computes protein haplotypes for use in precision drug design}, journal = {Nat Commun}, volume = {9}, number = {1}, pages = {4128}, year = {2018}, doi = {10.1038/s41467-018-06542-1}, abstract = {Selecting the most appropriate protein sequences is critical for precision drug design. Here we describe Haplosaurus, a bioinformatic tool for computation of protein haplotypes. Haplosaurus computes protein haplotypes from pre-existing chromosomally-phased genomic variation data. Integration into the Ensembl resource provides rapid and detailed protein haplotypes retrieval. Using Haplosaurus, we build a database of unique protein haplotypes from the 1000 Genomes dataset reflecting real-world protein sequence variability and their prevalence. For one in seven genes, their most common protein haplotype differs from the reference sequence and a similar number differs on their most common haplotype between human populations. Three case studies show how knowledge of the range of commonly encountered protein forms predicted in populations leads to insights into therapeutic efficacy. Haplosaurus and its associated database is expected to find broad applications in many disciplines using protein sequences and particularly impactful for therapeutics design.},} - BA Moore, BC Leonard, L Sebbag, SG Edwards, A Cooper, DM Imai, E Straiton, L Santos, C Reilly, SM Griffey, L Bower, D Clary, J Mason, MJ Roux, H Meziane, Y Herault, International Mouse Phenotyping Consortium, C McKerlie, AM Flenniken, LMJ Nutter, Z Berberovic, C Owen, S Newbigging, H Adissu, M Eskandarian, CW Hsu, S Kalaga, U Udensi, C Asomugha, R Bohat, JJ Gallegos, JR Seavitt, JD Heaney, AL Beaudet, ME Dickinson, MJ Justice, V Philip, V Kumar, KL Svenson, RE Braun, S Wells, H Cater, M Stewart, S Clementson-Mobbs, R Joynson, X Gao, T Suzuki, S Wakana, D Smedley, JK Seong, G Tocchini-Valentini, M Moore, C Fletcher, N Karp, R Ramirez-Solis, JK White, de MH Angelis, W Wurst, SM Thomasy, P Flicek, H Parkinson, SDM Brown, TF Meehan, PM Nishina, SA Murray, MP Krebs, AM Mallon, KCK Lloyd, CJ Murphy, A Moshiri. Identification of genes required for eye development by high-throughput screening of mouse knockouts. Commun Biol 2018;1:236. doi:10.1038/s42003-018-0226-0
[BibTeX] [Abstract]
Despite advances in next generation sequencing technologies, determining the genetic basis of ocular disease remains a major challenge due to the limited access and prohibitive cost of human forward genetics. Thus, less than 4,000 genes currently have available phenotype information for any organ system. Here we report the ophthalmic findings from the International Mouse Phenotyping Consortium, a large-scale functional genetic screen with the goal of generating and phenotyping a null mutant for every mouse gene. Of 4364 genes evaluated, 347 were identified to influence ocular phenotypes, 75\% of which are entirely novel in ocular pathology. This discovery greatly increases the current number of genes known to contribute to ophthalmic disease, and it is likely that many of the genes will subsequently prove to be important in human ocular development and disease.
@Article{30588515, author = {Moore BA and Leonard BC and Sebbag L and Edwards SG and Cooper A and Imai DM and Straiton E and Santos L and Reilly C and Griffey SM and Bower L and Clary D and Mason J and Roux MJ and Meziane H and Herault Y and {International Mouse Phenotyping Consortium} and McKerlie C and Flenniken AM and Nutter LMJ and Berberovic Z and Owen C and Newbigging S and Adissu H and Eskandarian M and Hsu CW and Kalaga S and Udensi U and Asomugha C and Bohat R and Gallegos JJ and Seavitt JR and Heaney JD and Beaudet AL and Dickinson ME and Justice MJ and Philip V and Kumar V and Svenson KL and Braun RE and Wells S and Cater H and Stewart M and Clementson-Mobbs S and Joynson R and Gao X and Suzuki T and Wakana S and Smedley D and Seong JK and Tocchini-Valentini G and Moore M and Fletcher C and Karp N and Ramirez-Solis R and White JK and de Angelis MH and Wurst W and Thomasy SM and Flicek P and Parkinson H and Brown SDM and Meehan TF and Nishina PM and Murray SA and Krebs MP and Mallon AM and Lloyd KCK and Murphy CJ and Moshiri A}, title = {Identification of genes required for eye development by high-throughput screening of mouse knockouts}, journal = {Commun Biol}, volume = {1}, pages = {236}, year = {2018}, doi = {10.1038/s42003-018-0226-0}, abstract = {Despite advances in next generation sequencing technologies, determining the genetic basis of ocular disease remains a major challenge due to the limited access and prohibitive cost of human forward genetics. Thus, less than 4,000 genes currently have available phenotype information for any organ system. Here we report the ophthalmic findings from the International Mouse Phenotyping Consortium, a large-scale functional genetic screen with the goal of generating and phenotyping a null mutant for every mouse gene. Of 4364 genes evaluated, 347 were identified to influence ocular phenotypes, 75\% of which are entirely novel in ocular pathology. This discovery greatly increases the current number of genes known to contribute to ophthalmic disease, and it is likely that many of the genes will subsequently prove to be important in human ocular development and disease.},} - J Rozman, B Rathkolb, MA Oestereicher, C Schütt, AC Ravindranath, S Leuchtenberger, S Sharma, M Kistler, M Willershäuser, R Brommage, TF Meehan, J Mason, H Haselimashhadi, C IMPC, T Hough, AM Mallon, S Wells, L Santos, CJ Lelliott, JK White, T Sorg, MF Champy, LR Bower, CL Reynolds, AM Flenniken, SA Murray, LMJ Nutter, KL Svenson, D West, GP Tocchini-Valentini, AL Beaudet, F Bosch, RB Braun, MS Dobbie, X Gao, Y Herault, A Moshiri, BA Moore, KC Kent Lloyd, C McKerlie, H Masuya, N Tanaka, P Flicek, HE Parkinson, R Sedlacek, JK Seong, CL Wang, M Moore, SD Brown, MH Tschöp, W Wurst, M Klingenspor, E Wolf, J Beckers, F Machicao, A Peter, H Staiger, HU Häring, H Grallert, M Campillos, H Maier, H Fuchs, V Gailus-Durner, T Werner, de M Hrabe Angelis. Identification of genetic elements in metabolism by high-throughput mouse phenotyping. Nat Commun 2018;9(1):288. doi:10.1038/s41467-017-01995-2
[BibTeX] [Abstract]
Metabolic diseases are a worldwide problem but the underlying genetic factors and their relevance to metabolic disease remain incompletely understood. Genome-wide research is needed to characterize so-far unannotated mammalian metabolic genes. Here, we generate and analyze metabolic phenotypic data of 2016 knockout mouse strains under the aegis of the International Mouse Phenotyping Consortium (IMPC) and find 974 gene knockouts with strong metabolic phenotypes. 429 of those had no previous link to metabolism and 51 genes remain functionally completely unannotated. We compared human orthologues of these uncharacterized genes in five GWAS consortia and indeed 23 candidate genes are associated with metabolic disease. We further identify common regulatory elements in promoters of candidate genes. As each regulatory element is composed of several transcription factor binding sites, our data reveal an extensive metabolic phenotype-associated network of co-regulated genes. Our systematic mouse phenotype analysis thus paves the way for full functional annotation of the genome.
@Article{29348434, author = {Rozman J and Rathkolb B and Oestereicher MA and Schütt C and Ravindranath AC and Leuchtenberger S and Sharma S and Kistler M and Willershäuser M and Brommage R and Meehan TF and Mason J and Haselimashhadi H and IMPC C and Hough T and Mallon AM and Wells S and Santos L and Lelliott CJ and White JK and Sorg T and Champy MF and Bower LR and Reynolds CL and Flenniken AM and Murray SA and Nutter LMJ and Svenson KL and West D and Tocchini-Valentini GP and Beaudet AL and Bosch F and Braun RB and Dobbie MS and Gao X and Herault Y and Moshiri A and Moore BA and Kent Lloyd KC and McKerlie C and Masuya H and Tanaka N and Flicek P and Parkinson HE and Sedlacek R and Seong JK and Wang CL and Moore M and Brown SD and Tschöp MH and Wurst W and Klingenspor M and Wolf E and Beckers J and Machicao F and Peter A and Staiger H and Häring HU and Grallert H and Campillos M and Maier H and Fuchs H and Gailus-Durner V and Werner T and Hrabe de Angelis M}, title = {Identification of genetic elements in metabolism by high-throughput mouse phenotyping}, journal = {Nat Commun}, volume = {9}, number = {1}, pages = {288}, year = {2018}, doi = {10.1038/s41467-017-01995-2}, abstract = {Metabolic diseases are a worldwide problem but the underlying genetic factors and their relevance to metabolic disease remain incompletely understood. Genome-wide research is needed to characterize so-far unannotated mammalian metabolic genes. Here, we generate and analyze metabolic phenotypic data of 2016 knockout mouse strains under the aegis of the International Mouse Phenotyping Consortium (IMPC) and find 974 gene knockouts with strong metabolic phenotypes. 429 of those had no previous link to metabolism and 51 genes remain functionally completely unannotated. We compared human orthologues of these uncharacterized genes in five GWAS consortia and indeed 23 candidate genes are associated with metabolic disease. We further identify common regulatory elements in promoters of candidate genes. As each regulatory element is composed of several transcription factor binding sites, our data reveal an extensive metabolic phenotype-associated network of co-regulated genes. Our systematic mouse phenotype analysis thus paves the way for full functional annotation of the genome.},} - AA Singh, F Petraglia, A Nebbioso, G Yi, M Conte, S Valente, A Mandoli, L Scisciola, R Lindeboom, H Kerstens, EM Janssen-Megens, F Pourfarzad, E Habibi, K Berentsen, B Kim, C Logie, S Heath, ATJ Wierenga, L Clarke, P Flicek, JH Jansen, T Kuijpers, ML Yaspo, VD Valle, O Bernard, I Gut, E Vellenga, HG Stunnenberg, A Mai, L Altucci, JHA Martens. Multi-omics profiling reveals a distinctive epigenome signature for high-risk acute promyelocytic leukemia. Oncotarget 2018;9(39):25647–25660. doi:10.18632/oncotarget.25429
[BibTeX] [Abstract]
Epigenomic alterations have been associated with both pathogenesis and progression of cancer. Here, we analyzed the epigenome of two high-risk APL (hrAPL) patients and compared it to non-high-risk APL cases. Despite the lack of common genetic signatures, we found that human hrAPL blasts from patients with extremely poor prognosis display specific patterns of histone H3 acetylation, specifically hyperacetylation at a common set of enhancer regions. In addition, unique profiles of the repressive marks H3K27me3 and DNA methylation were exposed in high-risk APLs. Epigenetic comparison with low/intermediate-risk APLs and AMLs revealed hrAPL-specific patterns of histone acetylation and DNA methylation, suggesting these could be further developed into markers for clinical identification. The epigenetic drug MC2884, a newly generated general HAT/EZH2 inhibitor, induces apoptosis of high-risk APL blasts and reshapes their epigenomes by targeting both active and repressive marks. Together, our analysis uncovers distinctive epigenome signatures of hrAPL patients, and provides proof of concept for use of epigenome profiling coupled to epigenetic drugs to `personalize' precision medicine.
@Article{29876014, author = {Singh AA and Petraglia F and Nebbioso A and Yi G and Conte M and Valente S and Mandoli A and Scisciola L and Lindeboom R and Kerstens H and Janssen-Megens EM and Pourfarzad F and Habibi E and Berentsen K and Kim B and Logie C and Heath S and Wierenga ATJ and Clarke L and Flicek P and Jansen JH and Kuijpers T and Yaspo ML and Valle VD and Bernard O and Gut I and Vellenga E and Stunnenberg HG and Mai A and Altucci L and Martens JHA}, title = {Multi-omics profiling reveals a distinctive epigenome signature for high-risk acute promyelocytic leukemia}, journal = {Oncotarget}, volume = {9}, number = {39}, pages = {25647--25660}, year = {2018}, doi = {10.18632/oncotarget.25429}, abstract = {Epigenomic alterations have been associated with both pathogenesis and progression of cancer. Here, we analyzed the epigenome of two high-risk APL (hrAPL) patients and compared it to non-high-risk APL cases. Despite the lack of common genetic signatures, we found that human hrAPL blasts from patients with extremely poor prognosis display specific patterns of histone H3 acetylation, specifically hyperacetylation at a common set of enhancer regions. In addition, unique profiles of the repressive marks H3K27me3 and DNA methylation were exposed in high-risk APLs. Epigenetic comparison with low/intermediate-risk APLs and AMLs revealed hrAPL-specific patterns of histone acetylation and DNA methylation, suggesting these could be further developed into markers for clinical identification. The epigenetic drug MC2884, a newly generated general HAT/EZH2 inhibitor, induces apoptosis of high-risk APL blasts and reshapes their epigenomes by targeting both active and repressive marks. Together, our analysis uncovers distinctive epigenome signatures of hrAPL patients, and provides proof of concept for use of epigenome profiling coupled to epigenetic drugs to `personalize' precision medicine.},} - H Lochmüller, DM Badowska, R Thompson, NV Knoers, A Aartsma-Rus, I Gut, L Wood, T Harmuth, A Durudas, H Graessner, F Schaefer, O Riess, C RD-Connect, C NeurOmics, C EURenOmics. RD-Connect, NeurOmics and EURenOmics: collaborative European initiative for rare diseases. Eur J Hum Genet 2018;26(6):778–785. doi:10.1038/s41431-018-0115-5
[BibTeX] [Abstract]
Although individually uncommon, rare diseases (RDs) collectively affect 6-8\% of the population. The unmet need of the rare disease community was recognized by the European Commission which in 2012 funded three flagship projects, RD-Connect, NeurOmics, and EURenOmics, to help move the field forward with the ambition of advancing -omics research and data sharing at their core in line with the goals of IRDiRC (International Rare Disease Research Consortium). NeurOmics and EURenOmics generate -omics data and improve diagnosis and therapy in rare renal and neurological diseases, with RD-Connect developing an infrastructure to facilitate the sharing, systematic integration and analysis of these data. Here, we summarize the achievements of these three projects, their impact on the RD community and their vision for the future. We also report from the Joint Outreach Day organized by the three projects on the 3rd of May 2017 in Berlin. The workshop stimulated an open, multi-stakeholder discussion on the challenges of the rare diseases, and highlighted the cross-project cooperation and the common goal: the use of innovative genomic technologies in rare disease research.
@Article{29487416, author = {Lochmüller H and Badowska DM and Thompson R and Knoers NV and Aartsma-Rus A and Gut I and Wood L and Harmuth T and Durudas A and Graessner H and Schaefer F and Riess O and RD-Connect C and NeurOmics C and EURenOmics C}, title = {RD-Connect, NeurOmics and EURenOmics: collaborative European initiative for rare diseases}, journal = {Eur J Hum Genet}, volume = {26}, number = {6}, pages = {778--785}, year = {2018}, doi = {10.1038/s41431-018-0115-5}, howpublished = {Advanced online publication: 27 February 2018}, abstract = {Although individually uncommon, rare diseases (RDs) collectively affect 6-8\% of the population. The unmet need of the rare disease community was recognized by the European Commission which in 2012 funded three flagship projects, RD-Connect, NeurOmics, and EURenOmics, to help move the field forward with the ambition of advancing -omics research and data sharing at their core in line with the goals of IRDiRC (International Rare Disease Research Consortium). NeurOmics and EURenOmics generate -omics data and improve diagnosis and therapy in rare renal and neurological diseases, with RD-Connect developing an infrastructure to facilitate the sharing, systematic integration and analysis of these data. Here, we summarize the achievements of these three projects, their impact on the RD community and their vision for the future. We also report from the Joint Outreach Day organized by the three projects on the 3rd of May 2017 in Berlin. The workshop stimulated an open, multi-stakeholder discussion on the challenges of the rare diseases, and highlighted the cross-project cooperation and the common goal: the use of innovative genomic technologies in rare disease research.},} - SOM Dyke, M Linden, I Lappalainen, JR De Argila, K Carey, D Lloyd, JD Spalding, MN Cabili, G Kerry, J Foreman, T Cutts, M Shabani, LL Rodriguez, M Haeussler, B Walsh, X Jiang, S Wang, D Perrett, T Boughtwood, A Matern, AJ Brookes, M Cupak, M Fiume, R Pandya, I Tulchinsky, S Scollen, J Törnroos, S Das, AC Evans, BA Malin, S Beck, SE Brenner, T Nyrönen, N Blomberg, HV Firth, M Hurles, AA Philippakis, G Rätsch, M Brudno, KM Boycott, HL Rehm, M Baudis, ST Sherry, K Kato, BM Knoppers, D Baker, P Flicek. Registered access: authorizing data access. Eur J Hum Genet 2018;26(12):1721–1731. doi:10.1038/s41431-018-0219-y
[BibTeX] [Abstract]
The Global Alliance for Genomics and Health (GA4GH) proposes a data access policy model-''registered access''-to increase and improve access to data requiring an agreement to basic terms and conditions, such as the use of DNA sequence and health data in research. A registered access policy would enable a range of categories of users to gain access, starting with researchers and clinical care professionals. It would also facilitate general use and reuse of data but within the bounds of consent restrictions and other ethical obligations. In piloting registered access with the Scientific Demonstration data sharing projects of GA4GH, we provide additional ethics, policy and technical guidance to facilitate the implementation of this access model in an international setting.
@Article{30069064, author = {Dyke SOM and Linden M and Lappalainen I and De Argila JR and Carey K and Lloyd D and Spalding JD and Cabili MN and Kerry G and Foreman J and Cutts T and Shabani M and Rodriguez LL and Haeussler M and Walsh B and Jiang X and Wang S and Perrett D and Boughtwood T and Matern A and Brookes AJ and Cupak M and Fiume M and Pandya R and Tulchinsky I and Scollen S and Törnroos J and Das S and Evans AC and Malin BA and Beck S and Brenner SE and Nyrönen T and Blomberg N and Firth HV and Hurles M and Philippakis AA and Rätsch G and Brudno M and Boycott KM and Rehm HL and Baudis M and Sherry ST and Kato K and Knoppers BM and Baker D and Flicek P}, title = {Registered access: authorizing data access}, journal = {Eur J Hum Genet}, volume = {26}, number = {12}, pages = {1721--1731}, year = {2018}, doi = {10.1038/s41431-018-0219-y}, howpublished = {Advanced online publication: 2 August 2018}, abstract = {The Global Alliance for Genomics and Health (GA4GH) proposes a data access policy model-''registered access''-to increase and improve access to data requiring an agreement to basic terms and conditions, such as the use of DNA sequence and health data in research. A registered access policy would enable a range of categories of users to gain access, starting with researchers and clinical care professionals. It would also facilitate general use and reuse of data but within the bounds of consent restrictions and other ethical obligations. In piloting registered access with the Scientific Demonstration data sharing projects of GA4GH, we provide additional ethics, policy and technical guidance to facilitate the implementation of this access model in an international setting.},} - D Thybert, M Roller, FCP Navarro, I Fiddes, I Streeter, C Feig, D Martin-Galvez, M Kolmogorov, V Janoušek, W Akanni, B Aken, S Aldridge, V Chakrapani, W Chow, L Clarke, C Cummins, A Doran, M Dunn, L Goodstadt, K Howe, M Howell, AA Josselin, RC Karn, CM Laukaitis, L Jingtao, F Martin, M Muffato, S Nachtweide, MA Quail, C Sisu, M Stanke, K Stefflova, C Van Oosterhout, F Veyrunes, B Ward, F Yang, G Yazdanifar, A Zadissa, DJ Adams, A Brazma, M Gerstein, B Paten, S Pham, TM Keane, DT Odom, P Flicek. Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes. Genome Res 2018;28(4):448–459. doi:10.1101/gr.234096.117
[BibTeX] [Abstract]
Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 and 6 million yr ago, but that are absent in the Hominidae. Hominidae show between four- and sevenfold lower rates of nucleotide change and feature turnover in both neutral and functional sequences, suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. Recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli, which resulted in thousands of novel, species-specific CTCF binding sites. Our results show that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.
@Article{29563166, author = {Thybert D and Roller M and Navarro FCP and Fiddes I and Streeter I and Feig C and Martin-Galvez D and Kolmogorov M and Janoušek V and Akanni W and Aken B and Aldridge S and Chakrapani V and Chow W and Clarke L and Cummins C and Doran A and Dunn M and Goodstadt L and Howe K and Howell M and Josselin AA and Karn RC and Laukaitis CM and Jingtao L and Martin F and Muffato M and Nachtweide S and Quail MA and Sisu C and Stanke M and Stefflova K and Van Oosterhout C and Veyrunes F and Ward B and Yang F and Yazdanifar G and Zadissa A and Adams DJ and Brazma A and Gerstein M and Paten B and Pham S and Keane TM and Odom DT and Flicek P}, title = {Repeat associated mechanisms of genome evolution and function revealed by the Mus caroli and Mus pahari genomes}, journal = {Genome Res}, volume = {28}, number = {4}, pages = {448--459}, year = {2018}, doi = {10.1101/gr.234096.117}, howpublished = {Advanced online publication: 21 March 2018}, note = {First posted as a preprint: 2 July 2017}, abstract = {Understanding the mechanisms driving lineage-specific evolution in both primates and rodents has been hindered by the lack of sister clades with a similar phylogenetic structure having high-quality genome assemblies. Here, we have created chromosome-level assemblies of the Mus caroli and Mus pahari genomes. Together with the Mus musculus and Rattus norvegicus genomes, this set of rodent genomes is similar in divergence times to the Hominidae (human-chimpanzee-gorilla-orangutan). By comparing the evolutionary dynamics between the Muridae and Hominidae, we identified punctate events of chromosome reshuffling that shaped the ancestral karyotype of Mus musculus and Mus caroli between 3 and 6 million yr ago, but that are absent in the Hominidae. Hominidae show between four- and sevenfold lower rates of nucleotide change and feature turnover in both neutral and functional sequences, suggesting an underlying coherence to the Muridae acceleration. Our system of matched, high-quality genome assemblies revealed how specific classes of repeats can play lineage-specific roles in related species. Recent LINE activity has remodeled protein-coding loci to a greater extent across the Muridae than the Hominidae, with functional consequences at the species level such as reproductive isolation. Furthermore, we charted a Muridae-specific retrotransposon expansion at unprecedented resolution, revealing how a single nucleotide mutation transformed a specific SINE element into an active CTCF binding site carrier specifically in Mus caroli, which resulted in thousands of novel, species-specific CTCF binding sites. Our results show that the comparison of matched phylogenetic sets of genomes will be an increasingly powerful strategy for understanding mammalian biology.},} - J Lilue, AG Doran, IT Fiddes, M Abrudan, J Armstrong, R Bennett, W Chow, J Collins, S Collins, A Czechanski, P Danecek, M Diekhans, DD Dolle, M Dunn, R Durbin, D Earl, A Ferguson-Smith, P Flicek, J Flint, A Frankish, B Fu, M Gerstein, J Gilbert, L Goodstadt, J Harrow, K Howe, X Ibarra-Soria, M Kolmogorov, CJ Lelliott, DW Logan, J Loveland, CE Mathews, R Mott, P Muir, S Nachtweide, FCP Navarro, DT Odom, N Park, S Pelan, SK Pham, M Quail, L Reinholdt, L Romoth, L Shirley, C Sisu, M Sjoberg-Herrera, M Stanke, C Steward, M Thomas, G Threadgold, D Thybert, J Torrance, K Wong, J Wood, B Yalcin, F Yang, DJ Adams, B Paten, TM Keane. Sixteen diverse laboratory mouse reference genomes define strain-specific haplotypes and novel functional loci. Nat Genet 2018;50(11):1574–1583. doi:10.1038/s41588-018-0223-8
[BibTeX] [Abstract]
We report full-length draft de novo genome assemblies for 16 widely used inbred mouse strains and find extensive strain-specific haplotype variation. We identify and characterize 2,567 regions on the current mouse reference genome exhibiting the greatest sequence diversity. These regions are enriched for genes involved in pathogen defence and immunity and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. We used these genomes to improve the mouse reference genome, resulting in the completion of 10 new gene structures. Also, 62 new coding loci were added to the reference genome annotation. These genomes identified a large, previously unannotated, gene (Efcab3-like) encoding 5,874 amino acids. Mutant Efcab3-like mice display anomalies in multiple brain regions, suggesting a possible role for this gene in the regulation of brain development.
@Article{30275530, author = {Lilue J and Doran AG and Fiddes IT and Abrudan M and Armstrong J and Bennett R and Chow W and Collins J and Collins S and Czechanski A and Danecek P and Diekhans M and Dolle DD and Dunn M and Durbin R and Earl D and Ferguson-Smith A and Flicek P and Flint J and Frankish A and Fu B and Gerstein M and Gilbert J and Goodstadt L and Harrow J and Howe K and Ibarra-Soria X and Kolmogorov M and Lelliott CJ and Logan DW and Loveland J and Mathews CE and Mott R and Muir P and Nachtweide S and Navarro FCP and Odom DT and Park N and Pelan S and Pham SK and Quail M and Reinholdt L and Romoth L and Shirley L and Sisu C and Sjoberg-Herrera M and Stanke M and Steward C and Thomas M and Threadgold G and Thybert D and Torrance J and Wong K and Wood J and Yalcin B and Yang F and Adams DJ and Paten B and Keane TM}, title = {Sixteen diverse laboratory mouse reference genomes define strain-specific haplotypes and novel functional loci}, journal = {Nat Genet}, volume = {50}, number = {11}, pages = {1574--1583}, year = {2018}, doi = {10.1038/s41588-018-0223-8}, howpublished = {Advanced online publication: 1 October 2018}, note = {First posted as a preprint: 12 February 2018}, abstract = {We report full-length draft de novo genome assemblies for 16 widely used inbred mouse strains and find extensive strain-specific haplotype variation. We identify and characterize 2,567 regions on the current mouse reference genome exhibiting the greatest sequence diversity. These regions are enriched for genes involved in pathogen defence and immunity and exhibit enrichment of transposable elements and signatures of recent retrotransposition events. Combinations of alleles and genes unique to an individual strain are commonly observed at these loci, reflecting distinct strain phenotypes. We used these genomes to improve the mouse reference genome, resulting in the completion of 10 new gene structures. Also, 62 new coding loci were added to the reference genome annotation. These genomes identified a large, previously unannotated, gene (Efcab3-like) encoding 5,874 amino acids. Mutant Efcab3-like mice display anomalies in multiple brain regions, suggesting a possible role for this gene in the regulation of brain development.},} - V Muñoz-Fuentes, P Cacheiro, TF Meehan, JA Aguilar-Pimentel, SDM Brown, AM Flenniken, P Flicek, A Galli, HH Mashhadi, de M Hrabě Angelis, JK Kim, KCK Lloyd, C McKerlie, H Morgan, SA Murray, LMJ Nutter, PT Reilly, JR Seavitt, JK Seong, M Simon, H Wardle-Jones, A-M Mallon, D Smedley, HE Parkinson. The International Mouse Phenotyping Consortium (IMPC): a functional catalogue of the mammalian genome that informs conservation. Conserv Genet 2018;19(4):995–1005. doi:10.1007/s10592-018-1072-9
[BibTeX]@Article{, author = {Muñoz-Fuentes V and Cacheiro P and Meehan TF and Aguilar-Pimentel JA and Brown SDM and Flenniken AM and Flicek P and Galli A and Mashhadi HH and Hrabě de Angelis M and Kim JK and Lloyd KCK and McKerlie C and Morgan H and Murray SA and Nutter LMJ and Reilly PT and Seavitt JR and Seong JK and Simon M and Wardle-Jones H and Mallon A-M and Smedley D and Parkinson HE}, title = {The International Mouse Phenotyping Consortium (IMPC): a functional catalogue of the mammalian genome that informs conservation}, journal = {Conserv Genet}, volume = {19}, number = {4}, pages = {995--1005}, year = {2018}, doi = {10.1007/s10592-018-1072-9}, howpublished = {Advanced online publication: 19 May 2018}, } - R Beekman, V Chapaprieta, N Russiñol, R Vilarrasa-Blasi, N Verdaguer-Dot, JHA Martens, M Duran-Ferrer, M Kulis, F Serra, BM Javierre, SW Wingett, G Clot, AC Queirós, G Castellano, J Blanc, M Gut, A Merkel, S Heath, A Vlasova, S Ullrich, E Palumbo, A Enjuanes, D Martín-García, S Beà, M Pinyol, M Aymerich, R Royo, M Puiggros, D Torrents, A Datta, E Lowy, M Kostadima, M Roller, L Clarke, P Flicek, X Agirre, F Prosper, T Baumann, J Delgado, A López-Guillermo, P Fraser, ML Yaspo, R Guigó, R Siebert, MA Martí-Renom, XS Puente, C López-Otín, I Gut, HG Stunnenberg, E Campo, JI Martin-Subero. The reference epigenome and regulatory chromatin landscape of chronic lymphocytic leukemia. Nat Med 2018;24(6):868–880. doi:10.1038/s41591-018-0028-4
[BibTeX] [Abstract]
Chronic lymphocytic leukemia (CLL) is a frequent hematological neoplasm in which underlying epigenetic alterations are only partially understood. Here, we analyze the reference epigenome of seven primary CLLs and the regulatory chromatin landscape of 107 primary cases in the context of normal B cell differentiation. We identify that the CLL chromatin landscape is largely influenced by distinct dynamics during normal B cell maturation. Beyond this, we define extensive catalogues of regulatory elements de novo reprogrammed in CLL as a whole and in its major clinico-biological subtypes classified by IGHV somatic hypermutation levels. We uncover that IGHV-unmutated CLLs harbor more active and open chromatin than IGHV-mutated cases. Furthermore, we show that de novo active regions in CLL are enriched for NFAT, FOX and TCF/LEF transcription factor family binding sites. Although most genetic alterations are not associated with consistent epigenetic profiles, CLLs with MYD88 mutations and trisomy 12 show distinct chromatin configurations. Furthermore, we observe that non-coding mutations in IGHV-mutated CLLs are enriched in H3K27ac-associated regulatory elements outside accessible chromatin. Overall, this study provides an integrative portrait of the CLL epigenome, identifies extensive networks of altered regulatory elements and sheds light on the relationship between the genetic and epigenetic architecture of the disease.
@Article{29785028, author = {Beekman R and Chapaprieta V and Russiñol N and Vilarrasa-Blasi R and Verdaguer-Dot N and Martens JHA and Duran-Ferrer M and Kulis M and Serra F and Javierre BM and Wingett SW and Clot G and Queirós AC and Castellano G and Blanc J and Gut M and Merkel A and Heath S and Vlasova A and Ullrich S and Palumbo E and Enjuanes A and Martín-García D and Beà S and Pinyol M and Aymerich M and Royo R and Puiggros M and Torrents D and Datta A and Lowy E and Kostadima M and Roller M and Clarke L and Flicek P and Agirre X and Prosper F and Baumann T and Delgado J and López-Guillermo A and Fraser P and Yaspo ML and Guigó R and Siebert R and Martí-Renom MA and Puente XS and López-Otín C and Gut I and Stunnenberg HG and Campo E and Martin-Subero JI}, title = {The reference epigenome and regulatory chromatin landscape of chronic lymphocytic leukemia}, journal = {Nat Med}, volume = {24}, number = {6}, pages = {868--880}, year = {2018}, doi = {10.1038/s41591-018-0028-4}, howpublished = {Advanced online publication: 21 May 2018}, abstract = {Chronic lymphocytic leukemia (CLL) is a frequent hematological neoplasm in which underlying epigenetic alterations are only partially understood. Here, we analyze the reference epigenome of seven primary CLLs and the regulatory chromatin landscape of 107 primary cases in the context of normal B cell differentiation. We identify that the CLL chromatin landscape is largely influenced by distinct dynamics during normal B cell maturation. Beyond this, we define extensive catalogues of regulatory elements de novo reprogrammed in CLL as a whole and in its major clinico-biological subtypes classified by IGHV somatic hypermutation levels. We uncover that IGHV-unmutated CLLs harbor more active and open chromatin than IGHV-mutated cases. Furthermore, we show that de novo active regions in CLL are enriched for NFAT, FOX and TCF/LEF transcription factor family binding sites. Although most genetic alterations are not associated with consistent epigenetic profiles, CLLs with MYD88 mutations and trisomy 12 show distinct chromatin configurations. Furthermore, we observe that non-coding mutations in IGHV-mutated CLLs are enriched in H3K27ac-associated regulatory elements outside accessible chromatin. Overall, this study provides an integrative portrait of the CLL epigenome, identifies extensive networks of altered regulatory elements and sheds light on the relationship between the genetic and epigenetic architecture of the disease.},}
2017
- MR Bowl, MM Simon, NJ Ingham, S Greenaway, L Santos, H Cater, S Taylor, J Mason, N Kurbatova, S Pearson, LR Bower, DA Clary, H Meziane, P Reilly, O Minowa, L Kelsey, MPC International, GP Tocchini-Valentini, X Gao, A Bradley, WC Skarnes, M Moore, AL Beaudet, MJ Justice, J Seavitt, ME Dickinson, W Wurst, de MH Angelis, Y Herault, S Wakana, LMJ Nutter, AM Flenniken, C McKerlie, SA Murray, KL Svenson, RE Braun, DB West, KCK Lloyd, DJ Adams, J White, N Karp, P Flicek, D Smedley, TF Meehan, HE Parkinson, LM Teboul, S Wells, KP Steel, AM Mallon, SDM Brown. A large scale hearing loss screen reveals an extensive unexplored genetic landscape for auditory dysfunction. Nat Commun 2017;8(1):886. doi:10.1038/s41467-017-00595-4
[BibTeX] [Abstract]
The developmental and physiological complexity of the auditory system is likely reflected in the underlying set of genes involved in auditory function. In humans, over 150 non-syndromic loci have been identified, and there are more than 400 human genetic syndromes with a hearing loss component. Over 100 non-syndromic hearing loss genes have been identified in mouse and human, but we remain ignorant of the full extent of the genetic landscape involved in auditory dysfunction. As part of the International Mouse Phenotyping Consortium, we undertook a hearing loss screen in a cohort of 3006 mouse knockout strains. In total, we identify 67 candidate hearing loss genes. We detect known hearing loss genes, but the vast majority, 52, of the candidate genes were novel. Our analysis reveals a large and unexplored genetic landscape involved with auditory function.The full extent of the genetic basis for hearing impairment is unknown. Here, as part of the International Mouse Phenotyping Consortium, the authors perform a hearing loss screen in 3006 mouse knockout strains and identify 52 new candidate genes for genetic hearing loss.
@Article{29026089, author = {Bowl MR and Simon MM and Ingham NJ and Greenaway S and Santos L and Cater H and Taylor S and Mason J and Kurbatova N and Pearson S and Bower LR and Clary DA and Meziane H and Reilly P and Minowa O and Kelsey L and International MPC and Tocchini-Valentini GP and Gao X and Bradley A and Skarnes WC and Moore M and Beaudet AL and Justice MJ and Seavitt J and Dickinson ME and Wurst W and de Angelis MH and Herault Y and Wakana S and Nutter LMJ and Flenniken AM and McKerlie C and Murray SA and Svenson KL and Braun RE and West DB and Lloyd KCK and Adams DJ and White J and Karp N and Flicek P and Smedley D and Meehan TF and Parkinson HE and Teboul LM and Wells S and Steel KP and Mallon AM and Brown SDM}, title = {A large scale hearing loss screen reveals an extensive unexplored genetic landscape for auditory dysfunction}, journal = {Nat Commun}, volume = {8}, number = {1}, pages = {886}, year = {2017}, doi = {10.1038/s41467-017-00595-4}, abstract = {The developmental and physiological complexity of the auditory system is likely reflected in the underlying set of genes involved in auditory function. In humans, over 150 non-syndromic loci have been identified, and there are more than 400 human genetic syndromes with a hearing loss component. Over 100 non-syndromic hearing loss genes have been identified in mouse and human, but we remain ignorant of the full extent of the genetic landscape involved in auditory dysfunction. As part of the International Mouse Phenotyping Consortium, we undertook a hearing loss screen in a cohort of 3006 mouse knockout strains. In total, we identify 67 candidate hearing loss genes. We detect known hearing loss genes, but the vast majority, 52, of the candidate genes were novel. Our analysis reveals a large and unexplored genetic landscape involved with auditory function.The full extent of the genetic basis for hearing impairment is unknown. Here, as part of the International Mouse Phenotyping Consortium, the authors perform a hearing loss screen in 3006 mouse knockout strains and identify 52 new candidate genes for genetic hearing loss.},} - JL Raisaro, F Tramèr, Z Ji, D Bu, Y Zhao, K Carey, D Lloyd, H Sofia, D Baker, P Flicek, S Shringarpure, C Bustamante, S Wang, X Jiang, L Ohno-Machado, H Tang, X Wang, JP Hubaux. Addressing Beacon re-identification attacks: quantification and mitigation of privacy risks. J Am Med Inform Assoc 2017;24(4):799–805. doi:10.1093/jamia/ocw167
[BibTeX] [Abstract]
The Global Alliance for Genomics and Health (GA4GH) created the Beacon Project as a means of testing the willingness of data holders to share genetic data in the simplest technical context-a query for the presence of a specified nucleotide at a given position within a chromosome. Each participating site (or ``beacon'') is responsible for assuring that genomic data are exposed through the Beacon service only with the permission of the individual to whom the data pertains and in accordance with the GA4GH policy and standards.While recognizing the inference risks associated with large-scale data aggregation, and the fact that some beacons contain sensitive phenotypic associations that increase privacy risk, the GA4GH adjudged the risk of re-identification based on the binary yes/no allele-presence query responses as acceptable. However, recent work demonstrated that, given a beacon with specific characteristics (including relatively small sample size and an adversary who possesses an individual's whole genome sequence), the individual's membership in a beacon can be inferred through repeated queries for variants present in the individual's genome.In this paper, we propose three practical strategies for reducing re-identification risks in beacons. The first two strategies manipulate the beacon such that the presence of rare alleles is obscured; the third strategy budgets the number of accesses per user for each individual genome. Using a beacon containing data from the 1000 Genomes Project, we demonstrate that the proposed strategies can effectively reduce re-identification risk in beacon-like datasets.
@Article{28339683, author = {Raisaro JL and Tramèr F and Ji Z and Bu D and Zhao Y and Carey K and Lloyd D and Sofia H and Baker D and Flicek P and Shringarpure S and Bustamante C and Wang S and Jiang X and Ohno-Machado L and Tang H and Wang X and Hubaux JP}, title = {Addressing Beacon re-identification attacks: quantification and mitigation of privacy risks}, journal = {J Am Med Inform Assoc}, volume = {24}, number = {4}, pages = {799--805}, year = {2017}, doi = {10.1093/jamia/ocw167}, howpublished = {Advanced online publication: 20 February 2017}, abstract = {The Global Alliance for Genomics and Health (GA4GH) created the Beacon Project as a means of testing the willingness of data holders to share genetic data in the simplest technical context-a query for the presence of a specified nucleotide at a given position within a chromosome. Each participating site (or ``beacon'') is responsible for assuring that genomic data are exposed through the Beacon service only with the permission of the individual to whom the data pertains and in accordance with the GA4GH policy and standards.While recognizing the inference risks associated with large-scale data aggregation, and the fact that some beacons contain sensitive phenotypic associations that increase privacy risk, the GA4GH adjudged the risk of re-identification based on the binary yes/no allele-presence query responses as acceptable. However, recent work demonstrated that, given a beacon with specific characteristics (including relatively small sample size and an adversary who possesses an individual's whole genome sequence), the individual's membership in a beacon can be inferred through repeated queries for variants present in the individual's genome.In this paper, we propose three practical strategies for reducing re-identification risks in beacons. The first two strategies manipulate the beacon such that the presence of rare alleles is obscured; the third strategy budgets the number of accesses per user for each individual genome. Using a beacon containing data from the 1000 Genomes Project, we demonstrate that the proposed strategies can effectively reduce re-identification risk in beacon-like datasets.},} - X Zheng-Bradley, I Streeter, S Fairley, D Richardson, L Clarke, P Flicek, 1000 Genomes Project Consortium. Alignment of 1000 Genomes Project Reads to Reference Assembly GRCh38. Gigascience 2017;6(7):1–8. doi:10.1093/gigascience/gix038
[BibTeX] [Abstract]
BACKGROUND: The 1000 Genomes Project produced more than 100 trillion basepairs of short read sequence from more than 2600 samples in 26 populations over a period of five years. In its final phase, the project released over 85 million genotyped and phased variants on human reference genome assembly GRCh37. An updated reference assembly, GRCh38, was released in late 2013, but there was insufficient time for the final phase of the project analysis to change to the new assembly. Although it is possible to lift the coordinates of the 1000 Genomes project variants to the new assembly, this is a potentially error prone process as coordinate remapping is most appropriate only for non-repetitive regions of the genome and those that did not see significant change between the two assemblies. It will also miss variants in any region that was that was newly added to GRCh38. Thus, to produce the highest quality variants and genotypes on GRCh38, the best strategy is to realign the reads and recall the variants based on the new alignment. FINDINGS: As the first step of variant calling for the 1000 Genomes Project data, we have finished remapping all of the 1000 Genomes sequence reads to GRCh38 with ALT-aware BWA-MEM. The resulting alignments are available as CRAM, a reference-based sequence compression format. CONCLUSIONS: The data have been released on our FTP site and are also available from European Nucleotide Archive (ENA) to facilitate researchers discovering variants on the primary sequences and alternative contigs of GRCh38.
@Article{28531267, author = {Zheng-Bradley X and Streeter I and Fairley S and Richardson D and Clarke L and Flicek P and {1000 Genomes Project Consortium}}, title = {Alignment of 1000 Genomes Project Reads to Reference Assembly GRCh38}, journal = {Gigascience}, volume = {6}, number = {7}, pages = {1--8}, year = {2017}, doi = {10.1093/gigascience/gix038}, howpublished = {Advanced online publication: 20 May 2017}, abstract = {BACKGROUND: The 1000 Genomes Project produced more than 100 trillion basepairs of short read sequence from more than 2600 samples in 26 populations over a period of five years. In its final phase, the project released over 85 million genotyped and phased variants on human reference genome assembly GRCh37. An updated reference assembly, GRCh38, was released in late 2013, but there was insufficient time for the final phase of the project analysis to change to the new assembly. Although it is possible to lift the coordinates of the 1000 Genomes project variants to the new assembly, this is a potentially error prone process as coordinate remapping is most appropriate only for non-repetitive regions of the genome and those that did not see significant change between the two assemblies. It will also miss variants in any region that was that was newly added to GRCh38. Thus, to produce the highest quality variants and genotypes on GRCh38, the best strategy is to realign the reads and recall the variants based on the new alignment. FINDINGS: As the first step of variant calling for the 1000 Genomes Project data, we have finished remapping all of the 1000 Genomes sequence reads to GRCh38 with ALT-aware BWA-MEM. The resulting alignments are available as CRAM, a reference-based sequence compression format. CONCLUSIONS: The data have been released on our FTP site and are also available from European Nucleotide Archive (ENA) to facilitate researchers discovering variants on the primary sequences and alternative contigs of GRCh38.},} - X Zheng-Bradley, P Flicek. Applications of the 1000 Genomes Project resources. Brief Funct Genomics 2017;16(3):163–170. doi:10.1093/bfgp/elw027
[BibTeX] [Abstract]
The 1000 Genomes Project created a valuable, worldwide reference for human genetic variation. Common uses of the 1000 Genomes dataset include genotype imputation supporting Genome-wide Association Studies, mapping expression Quantitative Trait Loci, filtering non-pathogenic variants from exome, whole genome and cancer genome sequencing projects, and genetic analysis of population structure and molecular evolution. In this article, we will highlight some of the multiple ways that the 1000 Genomes data can be and has been utilized for genetic studies.
@Article{27436001, author = {Zheng-Bradley X and Flicek P}, title = {Applications of the 1000 Genomes Project resources}, journal = {Brief Funct Genomics}, volume = {16}, number = {3}, pages = {163--170}, year = {2017}, doi = {10.1093/bfgp/elw027}, howpublished = {Advanced online publication: 19 July 2016}, abstract = {The 1000 Genomes Project created a valuable, worldwide reference for human genetic variation. Common uses of the 1000 Genomes dataset include genotype imputation supporting Genome-wide Association Studies, mapping expression Quantitative Trait Loci, filtering non-pathogenic variants from exome, whole genome and cancer genome sequencing projects, and genetic analysis of population structure and molecular evolution. In this article, we will highlight some of the multiple ways that the 1000 Genomes data can be and has been utilized for genetic studies.},} - TF Meehan, N Conte, DB West, JO Jacobsen, J Mason, J Warren, CK Chen, I Tudose, M Relac, P Matthews, N Karp, L Santos, T Fiegel, N Ring, H Westerberg, S Greenaway, D Sneddon, H Morgan, GF Codner, ME Stewart, J Brown, N Horner, MPC International, M Haendel, N Washington, CJ Mungall, CL Reynolds, J Gallegos, V Gailus-Durner, T Sorg, G Pavlovic, LR Bower, M Moore, I Morse, X Gao, GP Tocchini-Valentini, Y Obata, SY Cho, JK Seong, J Seavitt, AL Beaudet, ME Dickinson, Y Herault, W Wurst, de MH Angelis, KCK Lloyd, AM Flenniken, LMJ Nutter, S Newbigging, C McKerlie, MJ Justice, SA Murray, KL Svenson, RE Braun, JK White, A Bradley, P Flicek, S Wells, WC Skarnes, DJ Adams, H Parkinson, AM Mallon, SDM Brown, D Smedley. Disease model discovery from 3,328 gene knockouts by The International Mouse Phenotyping Consortium. Nat Genet 2017;49(8):1231–1238. doi:10.1038/ng.3901
[BibTeX] [Abstract]
Although next-generation sequencing has revolutionized the ability to associate variants with human diseases, diagnostic rates and development of new therapies are still limited by a lack of knowledge of the functions and pathobiological mechanisms of most genes. To address this challenge, the International Mouse Phenotyping Consortium is creating a genome- and phenome-wide catalog of gene function by characterizing new knockout-mouse strains across diverse biological systems through a broad set of standardized phenotyping tests. All mice will be readily available to the biomedical community. Analyzing the first 3,328 genes identified models for 360 diseases, including the first models, to our knowledge, for type C Bernard-Soulier, Bardet-Biedl-5 and Gordon Holmes syndromes. 90\% of our phenotype annotations were novel, providing functional evidence for 1,092 genes and candidates in genetically uncharacterized diseases including arrhythmogenic right ventricular dysplasia 3. Finally, we describe our role in variant functional validation with The 100,000 Genomes Project and others.
@Article{28650483, author = {Meehan TF and Conte N and West DB and Jacobsen JO and Mason J and Warren J and Chen CK and Tudose I and Relac M and Matthews P and Karp N and Santos L and Fiegel T and Ring N and Westerberg H and Greenaway S and Sneddon D and Morgan H and Codner GF and Stewart ME and Brown J and Horner N and International MPC and Haendel M and Washington N and Mungall CJ and Reynolds CL and Gallegos J and Gailus-Durner V and Sorg T and Pavlovic G and Bower LR and Moore M and Morse I and Gao X and Tocchini-Valentini GP and Obata Y and Cho SY and Seong JK and Seavitt J and Beaudet AL and Dickinson ME and Herault Y and Wurst W and de Angelis MH and Lloyd KCK and Flenniken AM and Nutter LMJ and Newbigging S and McKerlie C and Justice MJ and Murray SA and Svenson KL and Braun RE and White JK and Bradley A and Flicek P and Wells S and Skarnes WC and Adams DJ and Parkinson H and Mallon AM and Brown SDM and Smedley D}, title = {Disease model discovery from 3,328 gene knockouts by The International Mouse Phenotyping Consortium}, journal = {Nat Genet}, volume = {49}, number = {8}, pages = {1231--1238}, year = {2017}, doi = {10.1038/ng.3901}, howpublished = {Advanced online publication: 26 June 2017}, abstract = {Although next-generation sequencing has revolutionized the ability to associate variants with human diseases, diagnostic rates and development of new therapies are still limited by a lack of knowledge of the functions and pathobiological mechanisms of most genes. To address this challenge, the International Mouse Phenotyping Consortium is creating a genome- and phenome-wide catalog of gene function by characterizing new knockout-mouse strains across diverse biological systems through a broad set of standardized phenotyping tests. All mice will be readily available to the biomedical community. Analyzing the first 3,328 genes identified models for 360 diseases, including the first models, to our knowledge, for type C Bernard-Soulier, Bardet-Biedl-5 and Gordon Holmes syndromes. 90\% of our phenotype annotations were novel, providing functional evidence for 1,092 genes and candidates in genetically uncharacterized diseases including arrhythmogenic right ventricular dysplasia 3. Finally, we describe our role in variant functional validation with The 100,000 Genomes Project and others.},} - BL Aken, P Achuthan, W Akanni, MR Amode, F Bernsdorff, J Bhai, K Billis, D Carvalho-Silva, C Cummins, P Clapham, L Gil, CG Girón, L Gordon, T Hourlier, SE Hunt, SH Janacek, T Juettemann, S Keenan, MR Laird, I Lavidas, T Maurel, W McLaren, B Moore, DN Murphy, R Nag, V Newman, M Nuhn, CK Ong, A Parker, M Patricio, HS Riat, D Sheppard, H Sparrow, K Taylor, A Thormann, A Vullo, B Walts, SP Wilder, A Zadissa, M Kostadima, FJ Martin, M Muffato, E Perry, M Ruffier, DM Staines, SJ Trevanion, F Cunningham, A Yates, DR Zerbino, P Flicek. Ensembl 2017. Nucleic Acids Res 2017;45(D1):D635–D642. doi:10.1093/nar/gkw1104
[BibTeX] [Abstract]
Ensembl (www.ensembl.org) is a database and genome browser for enabling research on vertebrate genomes. We import, analyse, curate and integrate a diverse collection of large-scale reference data to create a more comprehensive view of genome biology than would be possible from any individual dataset. Our extensive data resources include evidence-based gene and regulatory region annotation, genome variation and gene trees. An accompanying suite of tools, infrastructure and programmatic access methods ensure uniform data analysis and distribution for all supported species. Together, these provide a comprehensive solution for large-scale and targeted genomics applications alike. Among many other developments over the past year, we have improved our resources for gene regulation and comparative genomics, and added CRISPR/Cas9 target sites. We released new browser functionality and tools, including improved filtering and prioritization of genome variation, Manhattan plot visualization for linkage disequilibrium and eQTL data, and an ontology search for phenotypes, traits and disease. We have also enhanced data discovery and access with a track hub registry and a selection of new REST end points. All Ensembl data are freely released to the scientific community and our source code is available via the open source Apache 2.0 license.
@Article{27899575, author = {Aken BL and Achuthan P and Akanni W and Amode MR and Bernsdorff F and Bhai J and Billis K and Carvalho-Silva D and Cummins C and Clapham P and Gil L and Girón CG and Gordon L and Hourlier T and Hunt SE and Janacek SH and Juettemann T and Keenan S and Laird MR and Lavidas I and Maurel T and McLaren W and Moore B and Murphy DN and Nag R and Newman V and Nuhn M and Ong CK and Parker A and Patricio M and Riat HS and Sheppard D and Sparrow H and Taylor K and Thormann A and Vullo A and Walts B and Wilder SP and Zadissa A and Kostadima M and Martin FJ and Muffato M and Perry E and Ruffier M and Staines DM and Trevanion SJ and Cunningham F and Yates A and Zerbino DR and Flicek P}, title = {Ensembl 2017}, journal = {Nucleic Acids Res}, volume = {45}, number = {D1}, pages = {D635--D642}, year = {2017}, doi = {10.1093/nar/gkw1104}, howpublished = {Advanced online publication: 28 November 2016}, abstract = {Ensembl (www.ensembl.org) is a database and genome browser for enabling research on vertebrate genomes. We import, analyse, curate and integrate a diverse collection of large-scale reference data to create a more comprehensive view of genome biology than would be possible from any individual dataset. Our extensive data resources include evidence-based gene and regulatory region annotation, genome variation and gene trees. An accompanying suite of tools, infrastructure and programmatic access methods ensure uniform data analysis and distribution for all supported species. Together, these provide a comprehensive solution for large-scale and targeted genomics applications alike. Among many other developments over the past year, we have improved our resources for gene regulation and comparative genomics, and added CRISPR/Cas9 target sites. We released new browser functionality and tools, including improved filtering and prioritization of genome variation, Manhattan plot visualization for linkage disequilibrium and eQTL data, and an ontology search for phenotypes, traits and disease. We have also enhanced data discovery and access with a track hub registry and a selection of new REST end points. All Ensembl data are freely released to the scientific community and our source code is available via the open source Apache 2.0 license.},} - M Ruffier, A Kähäri, M Komorowska, S Keenan, MR Laird, I Longden, G Proctor, S Searle, D Staines, K Taylor, A Vullo, A Yates, D Zerbino, P Flicek. Ensembl Core Software Resources: storage and programmatic access for DNA sequence and genome annotation. Database (Oxford) 2017;2017:bax20. doi:https://doi.org/10.1093/database/bax020
[BibTeX]@Article{28365736, author = {Ruffier M and Kähäri A and Komorowska M and Keenan S and Laird MR and Longden I and Proctor G and Searle S and Staines D and Taylor K and Vullo A and Yates A and Zerbino D and Flicek P}, title = {Ensembl Core Software Resources: storage and programmatic access for DNA sequence and genome annotation}, journal = {Database (Oxford)}, volume = {2017}, pages = {bax20}, year = {2017}, doi = {https://doi.org/10.1093/database/bax020}, howpublished = {Advanced online publication: 18 March 2017}, note = {First posted as a preprint: 11 November 2016}, } - VA Schneider, T Graves-Lindsay, K Howe, N Bouk, HC Chen, PA Kitts, TD Murphy, KD Pruitt, F Thibaud-Nissen, D Albracht, RS Fulton, M Kremitzki, V Magrini, C Markovic, S McGrath, KM Steinberg, K Auger, W Chow, J Collins, G Harden, T Hubbard, S Pelan, JT Simpson, G Threadgold, J Torrance, JM Wood, L Clarke, S Koren, M Boitano, P Peluso, H Li, CS Chin, AM Phillippy, R Durbin, RK Wilson, P Flicek, EE Eichler, DM Church. Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly. Genome Res 2017;27(5):849–864. doi:10.1101/gr.213611.116
[BibTeX] [Abstract]
The human reference genome assembly plays a central role in nearly all aspects of today's basic and clinical research. GRCh38 is the first coordinate-changing assembly update since 2009; it reflects the resolution of roughly 1000 issues and encompasses modifications ranging from thousands of single base changes to megabase-scale path reorganizations, gap closures, and localization of previously orphaned sequences. We developed a new approach to sequence generation for targeted base updates and used data from new genome mapping technologies and single haplotype resources to identify and resolve larger assembly issues. For the first time, the reference assembly contains sequence-based representations for the centromeres. We also expanded the number of alternate loci to create a reference that provides a more robust representation of human population variation. We demonstrate that the updates render the reference an improved annotation substrate, alter read alignments in unchanged regions, and impact variant interpretation at clinically relevant loci. We additionally evaluated a collection of new de novo long-read haploid assemblies and conclude that although the new assemblies compare favorably to the reference with respect to continuity, error rate, and gene completeness, the reference still provides the best representation for complex genomic regions and coding sequences. We assert that the collected updates in GRCh38 make the newer assembly a more robust substrate for comprehensive analyses that will promote our understanding of human biology and advance our efforts to improve health.
@Article{28396521, author = {Schneider VA and Graves-Lindsay T and Howe K and Bouk N and Chen HC and Kitts PA and Murphy TD and Pruitt KD and Thibaud-Nissen F and Albracht D and Fulton RS and Kremitzki M and Magrini V and Markovic C and McGrath S and Steinberg KM and Auger K and Chow W and Collins J and Harden G and Hubbard T and Pelan S and Simpson JT and Threadgold G and Torrance J and Wood JM and Clarke L and Koren S and Boitano M and Peluso P and Li H and Chin CS and Phillippy AM and Durbin R and Wilson RK and Flicek P and Eichler EE and Church DM}, title = {Evaluation of GRCh38 and de novo haploid genome assemblies demonstrates the enduring quality of the reference assembly}, journal = {Genome Res}, volume = {27}, number = {5}, pages = {849--864}, year = {2017}, doi = {10.1101/gr.213611.116}, howpublished = {Advanced online publication: 10 April 2017}, note = {First posted as a preprint: 30 August 2016}, abstract = {The human reference genome assembly plays a central role in nearly all aspects of today's basic and clinical research. GRCh38 is the first coordinate-changing assembly update since 2009; it reflects the resolution of roughly 1000 issues and encompasses modifications ranging from thousands of single base changes to megabase-scale path reorganizations, gap closures, and localization of previously orphaned sequences. We developed a new approach to sequence generation for targeted base updates and used data from new genome mapping technologies and single haplotype resources to identify and resolve larger assembly issues. For the first time, the reference assembly contains sequence-based representations for the centromeres. We also expanded the number of alternate loci to create a reference that provides a more robust representation of human population variation. We demonstrate that the updates render the reference an improved annotation substrate, alter read alignments in unchanged regions, and impact variant interpretation at clinically relevant loci. We additionally evaluated a collection of new de novo long-read haploid assemblies and conclude that although the new assemblies compare favorably to the reference with respect to continuity, error rate, and gene completeness, the reference still provides the best representation for complex genomic regions and coding sequences. We assert that the collected updates in GRCh38 make the newer assembly a more robust substrate for comprehensive analyses that will promote our understanding of human biology and advance our efforts to improve health.},} - WA Cheung, X Shao, A Morin, V Siroux, T Kwan, B Ge, D Aïssi, L Chen, L Vasquez, F Allum, F Guénard, E Bouzigon, MM Simon, E Boulier, A Redensek, S Watt, A Datta, L Clarke, P Flicek, D Mead, DS Paul, S Beck, G Bourque, M Lathrop, A Tchernof, MC Vohl, F Demenais, I Pin, K Downes, HG Stunnenberg, N Soranzo, T Pastinen, E Grundberg. Functional variation in allelic methylomes underscores a strong genetic contribution and reveals novel epigenetic alterations in the human epigenome. Genome Biol 2017;18(1):50. doi:10.1186/s13059-017-1173-7
[BibTeX] [Abstract]
BACKGROUND: The functional impact of genetic variation has been extensively surveyed, revealing that genetic changes correlated to phenotypes lie mostly in non-coding genomic regions. Studies have linked allele-specific genetic changes to gene expression, DNA methylation, and histone marks but these investigations have only been carried out in a limited set of samples. RESULTS: We describe a large-scale coordinated study of allelic and non-allelic effects on DNA methylation, histone mark deposition, and gene expression, detecting the interrelations between epigenetic and functional features at unprecedented resolution. We use information from whole genome and targeted bisulfite sequencing from 910 samples to perform genotype-dependent analyses of allele-specific methylation (ASM) and non-allelic methylation (mQTL). In addition, we introduce a novel genotype-independent test to detect methylation imbalance between chromosomes. Of the ~2.2 million CpGs tested for ASM, mQTL, and genotype-independent effects, we identify ~32\% as being genetically regulated (ASM or mQTL) and ~14\% as being putatively epigenetically regulated. We also show that epigenetically driven effects are strongly enriched in repressed regions and near transcription start sites, whereas the genetically regulated CpGs are enriched in enhancers. Known imprinted regions are enriched among epigenetically regulated loci, but we also observe several novel genomic regions (e.g., HOX genes) as being epigenetically regulated. Finally, we use our ASM datasets for functional interpretation of disease-associated loci and show the advantage of utilizing naïve T cells for understanding autoimmune diseases. CONCLUSIONS: Our rich catalogue of haploid methylomes across multiple tissues will allow validation of epigenome association studies and exploration of new biological models for allelic exclusion in the human genome.
@Article{28283040, author = {Cheung WA and Shao X and Morin A and Siroux V and Kwan T and Ge B and Aïssi D and Chen L and Vasquez L and Allum F and Guénard F and Bouzigon E and Simon MM and Boulier E and Redensek A and Watt S and Datta A and Clarke L and Flicek P and Mead D and Paul DS and Beck S and Bourque G and Lathrop M and Tchernof A and Vohl MC and Demenais F and Pin I and Downes K and Stunnenberg HG and Soranzo N and Pastinen T and Grundberg E}, title = {Functional variation in allelic methylomes underscores a strong genetic contribution and reveals novel epigenetic alterations in the human epigenome}, journal = {Genome Biol}, volume = {18}, number = {1}, pages = {50}, year = {2017}, doi = {10.1186/s13059-017-1173-7}, abstract = {BACKGROUND: The functional impact of genetic variation has been extensively surveyed, revealing that genetic changes correlated to phenotypes lie mostly in non-coding genomic regions. Studies have linked allele-specific genetic changes to gene expression, DNA methylation, and histone marks but these investigations have only been carried out in a limited set of samples. RESULTS: We describe a large-scale coordinated study of allelic and non-allelic effects on DNA methylation, histone mark deposition, and gene expression, detecting the interrelations between epigenetic and functional features at unprecedented resolution. We use information from whole genome and targeted bisulfite sequencing from 910 samples to perform genotype-dependent analyses of allele-specific methylation (ASM) and non-allelic methylation (mQTL). In addition, we introduce a novel genotype-independent test to detect methylation imbalance between chromosomes. Of the ~2.2 million CpGs tested for ASM, mQTL, and genotype-independent effects, we identify ~32\% as being genetically regulated (ASM or mQTL) and ~14\% as being putatively epigenetically regulated. We also show that epigenetically driven effects are strongly enriched in repressed regions and near transcription start sites, whereas the genetically regulated CpGs are enriched in enhancers. Known imprinted regions are enriched among epigenetically regulated loci, but we also observe several novel genomic regions (e.g., HOX genes) as being epigenetically regulated. Finally, we use our ASM datasets for functional interpretation of disease-associated loci and show the advantage of utilizing naïve T cells for understanding autoimmune diseases. CONCLUSIONS: Our rich catalogue of haploid methylomes across multiple tissues will allow validation of epigenome association studies and exploration of new biological models for allelic exclusion in the human genome.},} - GTEx Consortium.. Genetic effects on gene expression across human tissues. Nature 2017;550(7675):204–213. doi:10.1038/nature24277
[BibTeX] [Abstract]
Characterization of the molecular function of the human genome and its variation across individuals is essential for identifying the cellular mechanisms that underlie human genetic traits and diseases. The Genotype-Tissue Expression (GTEx) project aims to characterize variation in gene expression levels across individuals and diverse tissues of the human body, many of which are not easily accessible. Here we describe genetic effects on gene expression levels across 44 human tissues. We find that local genetic variation affects gene expression levels for the majority of genes, and we further identify inter-chromosomal genetic effects for 93 genes and 112 loci. On the basis of the identified genetic effects, we characterize patterns of tissue specificity, compare local and distal effects, and evaluate the functional properties of the genetic effects. We also demonstrate that multi-tissue, multi-individual data can be used to identify genes and pathways affected by human disease-associated variation, enabling a mechanistic interpretation of gene regulation and the genetic basis of disease.
@Article{29022597, author = {{GTEx Consortium.}}, title = {Genetic effects on gene expression across human tissues}, journal = {Nature}, volume = {550}, number = {7675}, pages = {204--213}, year = {2017}, doi = {10.1038/nature24277}, abstract = {Characterization of the molecular function of the human genome and its variation across individuals is essential for identifying the cellular mechanisms that underlie human genetic traits and diseases. The Genotype-Tissue Expression (GTEx) project aims to characterize variation in gene expression levels across individuals and diverse tissues of the human body, many of which are not easily accessible. Here we describe genetic effects on gene expression levels across 44 human tissues. We find that local genetic variation affects gene expression levels for the majority of genes, and we further identify inter-chromosomal genetic effects for 93 genes and 112 loci. On the basis of the identified genetic effects, we characterize patterns of tissue specificity, compare local and distal effects, and evaluate the functional properties of the genetic effects. We also demonstrate that multi-tissue, multi-individual data can be used to identify genes and pathways affected by human disease-associated variation, enabling a mechanistic interpretation of gene regulation and the genetic basis of disease.},} - AJ Jasinska, I Zelaya, SK Service, CB Peterson, RM Cantor, OW Choi, J DeYoung, E Eskin, LA Fairbanks, S Fears, AE Furterer, YS Huang, V Ramensky, CA Schmitt, H Svardal, MJ Jorgensen, JR Kaplan, D Villar, BL Aken, P Flicek, R Nag, ES Wong, J Blangero, TD Dyer, M Bogomolov, Y Benjamini, GM Weinstock, K Dewar, C Sabatti, RK Wilson, JD Jentsch, W Warren, G Coppola, RP Woods, NB Freimer. Genetic variation and gene expression across multiple tissues and developmental stages in a nonhuman primate. Nat Genet 2017;49(12):1714–1721. doi:10.1038/ng.3959
[BibTeX] [Abstract]
By analyzing multitissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalog of expression quantitative trait loci (eQTLs) in a nonhuman primate model. This catalog contains more genome-wide significant eQTLs per sample than comparable human resources and identifies sex- and age-related expression patterns. Findings include a master regulatory locus that likely has a role in immune function and a locus regulating hippocampal long noncoding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.
@Article{29083405, author = {Jasinska AJ and Zelaya I and Service SK and Peterson CB and Cantor RM and Choi OW and DeYoung J and Eskin E and Fairbanks LA and Fears S and Furterer AE and Huang YS and Ramensky V and Schmitt CA and Svardal H and Jorgensen MJ and Kaplan JR and Villar D and Aken BL and Flicek P and Nag R and Wong ES and Blangero J and Dyer TD and Bogomolov M and Benjamini Y and Weinstock GM and Dewar K and Sabatti C and Wilson RK and Jentsch JD and Warren W and Coppola G and Woods RP and Freimer NB}, title = {Genetic variation and gene expression across multiple tissues and developmental stages in a nonhuman primate}, journal = {Nat Genet}, volume = {49}, number = {12}, pages = {1714--1721}, year = {2017}, doi = {10.1038/ng.3959}, howpublished = {Advanced online publication: 30 October 2017}, note = {First posted as a preprint: 9 December 2016}, abstract = {By analyzing multitissue gene expression and genome-wide genetic variation data in samples from a vervet monkey pedigree, we generated a transcriptome resource and produced the first catalog of expression quantitative trait loci (eQTLs) in a nonhuman primate model. This catalog contains more genome-wide significant eQTLs per sample than comparable human resources and identifies sex- and age-related expression patterns. Findings include a master regulatory locus that likely has a role in immune function and a locus regulating hippocampal long noncoding RNAs (lncRNAs), whose expression correlates with hippocampal volume. This resource will facilitate genetic investigation of quantitative traits, including brain and behavioral phenotypes relevant to neuropsychiatric disorders.},} - D Martín-Gálvez, de D Dunoyer Segonzac, MCJ Ma, AE Kwitek, D Thybert, P Flicek. Genome variation and conserved regulation identify genomic regions responsible for strain specific phenotypes in rat. BMC Genomics 2017;18(1):986. doi:10.1186/s12864-017-4351-9
[BibTeX] [Abstract]
BACKGROUND: The genomes of laboratory rat strains are characterised by a mosaic haplotype structure caused by their unique breeding history. These mosaic haplotypes have been recently mapped by extensive sequencing of key strains. Comparison of genomic variation between two closely related rat strains with different phenotypes has been proposed as an effective strategy for the discovery of candidate strain-specific regions involved in phenotypic differences. We developed a method to prioritise strain-specific haplotypes by integrating genomic variation and genomic regulatory data predicted to be involved in specific phenotypes. Specifically, we aimed to identify genomic regions associated with Metabolic Syndrome (MetS), a disorder of energy utilization and storage affecting several organ systems. RESULTS: We compared two Lyon rat strains, Lyon Hypertensive (LH) which is susceptible to MetS, and Lyon Low pressure (LL), which is susceptible to obesity as an intermediate MetS phenotype, with a third strain (Lyon Normotensive, LN) that is resistant to both MetS and obesity. Applying a novel metric, we ranked the identified strain-specific haplotypes using evolutionary conservation of the occupancy three liver-specific transcription factors (HNF4A, CEBPA, and FOXA1) in five rodents including rat. Consideration of regulatory information effectively identified regions with liver-associated genes and rat orthologues of human GWAS variants related to obesity and metabolic traits. We attempted to find possible causative variants and compared them with the candidate genes proposed by previous studies. In strain-specific regions with conserved regulation, we found a significant enrichment for published evidence to obesity-one of the metabolic symptoms shown by the Lyon strains-amongst the genes assigned to promoters with strain-specific variation. CONCLUSIONS: Our results show that the use of functional regulatory conservation is a potentially effective approach to select strain-specific genomic regions associated with phenotypic differences among Lyon rats and could be extended to other systems.
@Article{29272997, author = {Martín-Gálvez D and Dunoyer de Segonzac D and Ma MCJ and Kwitek AE and Thybert D and Flicek P}, title = {Genome variation and conserved regulation identify genomic regions responsible for strain specific phenotypes in rat}, journal = {BMC Genomics}, volume = {18}, number = {1}, pages = {986}, year = {2017}, doi = {10.1186/s12864-017-4351-9}, note = {First posted as a preprint: 28 July 2017}, abstract = {BACKGROUND: The genomes of laboratory rat strains are characterised by a mosaic haplotype structure caused by their unique breeding history. These mosaic haplotypes have been recently mapped by extensive sequencing of key strains. Comparison of genomic variation between two closely related rat strains with different phenotypes has been proposed as an effective strategy for the discovery of candidate strain-specific regions involved in phenotypic differences. We developed a method to prioritise strain-specific haplotypes by integrating genomic variation and genomic regulatory data predicted to be involved in specific phenotypes. Specifically, we aimed to identify genomic regions associated with Metabolic Syndrome (MetS), a disorder of energy utilization and storage affecting several organ systems. RESULTS: We compared two Lyon rat strains, Lyon Hypertensive (LH) which is susceptible to MetS, and Lyon Low pressure (LL), which is susceptible to obesity as an intermediate MetS phenotype, with a third strain (Lyon Normotensive, LN) that is resistant to both MetS and obesity. Applying a novel metric, we ranked the identified strain-specific haplotypes using evolutionary conservation of the occupancy three liver-specific transcription factors (HNF4A, CEBPA, and FOXA1) in five rodents including rat. Consideration of regulatory information effectively identified regions with liver-associated genes and rat orthologues of human GWAS variants related to obesity and metabolic traits. We attempted to find possible causative variants and compared them with the candidate genes proposed by previous studies. In strain-specific regions with conserved regulation, we found a significant enrichment for published evidence to obesity-one of the metabolic symptoms shown by the Lyon strains-amongst the genes assigned to promoters with strain-specific variation. CONCLUSIONS: Our results show that the use of functional regulatory conservation is a potentially effective approach to select strain-specific genomic regions associated with phenotypic differences among Lyon rats and could be extended to other systems.},} - ES Wong, BM Schmitt, A Kazachenka, D Thybert, A Redmond, F Connor, TF Rayner, C Feig, AC Ferguson-Smith, JC Marioni, DT Odom, P Flicek. Interplay of cis and trans mechanisms driving transcription factor binding and gene expression evolution. Nat Commun 2017;8(1):1092. doi:10.1038/s41467-017-01037-x
[BibTeX] [Abstract]
Noncoding regulatory variants play a central role in the genetics of human diseases and in evolution. Here we measure allele-specific transcription factor binding occupancy of three liver-specific transcription factors between crosses of two inbred mouse strains to elucidate the regulatory mechanisms underlying transcription factor binding variations in mammals. Our results highlight the pre-eminence of cis-acting variants on transcription factor occupancy divergence. Transcription factor binding differences linked to cis-acting variants generally exhibit additive inheritance, while those linked to trans-acting variants are most often dominantly inherited. Cis-acting variants lead to local coordination of transcription factor occupancies that decay with distance; distal coordination is also observed and may be modulated by long-range chromatin contacts. Our results reveal the regulatory mechanisms that interplay to drive transcription factor occupancy, chromatin state, and gene expression in complex mammalian cell states.
@Article{29061983, author = {Wong ES and Schmitt BM and Kazachenka A and Thybert D and Redmond A and Connor F and Rayner TF and Feig C and Ferguson-Smith AC and Marioni JC and Odom DT and Flicek P}, title = {Interplay of cis and trans mechanisms driving transcription factor binding and gene expression evolution}, journal = {Nat Commun}, volume = {8}, number = {1}, pages = {1092}, year = {2017}, doi = {10.1038/s41467-017-01037-x}, note = {First posted as a preprint: 19 June 2016}, abstract = {Noncoding regulatory variants play a central role in the genetics of human diseases and in evolution. Here we measure allele-specific transcription factor binding occupancy of three liver-specific transcription factors between crosses of two inbred mouse strains to elucidate the regulatory mechanisms underlying transcription factor binding variations in mammals. Our results highlight the pre-eminence of cis-acting variants on transcription factor occupancy divergence. Transcription factor binding differences linked to cis-acting variants generally exhibit additive inheritance, while those linked to trans-acting variants are most often dominantly inherited. Cis-acting variants lead to local coordination of transcription factor occupancies that decay with distance; distal coordination is also observed and may be modulated by long-range chromatin contacts. Our results reveal the regulatory mechanisms that interplay to drive transcription factor occupancy, chromatin state, and gene expression in complex mammalian cell states.},} - G Maccari, J Robinson, K Ballingall, LA Guethlein, U Grimholt, J Kaufman, CS Ho, de NG Groot, P Flicek, RE Bontrop, JA Hammond, SG Marsh. IPD-MHC 2.0: an improved inter-species database for the study of the major histocompatibility complex. Nucleic Acids Res 2017;45(D1):D860–D864. doi:10.1093/nar/gkw1050
[BibTeX] [Abstract]
The IPD-MHC Database project (http://www.ebi.ac.uk/ipd/mhc/) collects and expertly curates sequences of the major histocompatibility complex from non-human species and provides the infrastructure and tools to enable accurate analysis. Since the first release of the database in 2003, IPD-MHC has grown and currently hosts a number of specific sections, with more than 7000 alleles from 70 species, including non-human primates, canines, felines, equids, ovids, suids, bovins, salmonids and murids. These sequences are expertly curated and made publicly available through an open access website. The IPD-MHC Database is a key resource in its field, and this has led to an average of 1500 unique visitors and more than 5000 viewed pages per month. As the database has grown in size and complexity, it has created a number of challenges in maintaining and organizing information, particularly the need to standardize nomenclature and taxonomic classification, while incorporating new allele submissions. Here, we describe the latest database release, the IPD-MHC 2.0 and discuss planned developments. This release incorporates sequence updates and new tools that enhance database queries and improve the submission procedure by utilizing common tools that are able to handle the varied requirements of each MHC-group.
@Article{27899604, author = {Maccari G and Robinson J and Ballingall K and Guethlein LA and Grimholt U and Kaufman J and Ho CS and de Groot NG and Flicek P and Bontrop RE and Hammond JA and Marsh SG}, title = {IPD-MHC 2.0: an improved inter-species database for the study of the major histocompatibility complex}, journal = {Nucleic Acids Res}, volume = {45}, number = {D1}, pages = {D860--D864}, year = {2017}, doi = {10.1093/nar/gkw1050}, howpublished = {Advanced online publication: 28 November 2016}, abstract = {The IPD-MHC Database project (http://www.ebi.ac.uk/ipd/mhc/) collects and expertly curates sequences of the major histocompatibility complex from non-human species and provides the infrastructure and tools to enable accurate analysis. Since the first release of the database in 2003, IPD-MHC has grown and currently hosts a number of specific sections, with more than 7000 alleles from 70 species, including non-human primates, canines, felines, equids, ovids, suids, bovins, salmonids and murids. These sequences are expertly curated and made publicly available through an open access website. The IPD-MHC Database is a key resource in its field, and this has led to an average of 1500 unique visitors and more than 5000 viewed pages per month. As the database has grown in size and complexity, it has created a number of challenges in maintaining and organizing information, particularly the need to standardize nomenclature and taxonomic classification, while incorporating new allele submissions. Here, we describe the latest database release, the IPD-MHC 2.0 and discuss planned developments. This release incorporates sequence updates and new tools that enhance database queries and improve the submission procedure by utilizing common tools that are able to handle the varied requirements of each MHC-group.},} - R Petersen, JJ Lambourne, BM Javierre, L Grassi, R Kreuzhuber, D Ruklisa, IM Rosa, AR Tomé, H Elding, van JP Geffen, T Jiang, S Farrow, J Cairns, AM Al-Subaie, S Ashford, A Attwood, J Batista, H Bouman, F Burden, FA Choudry, L Clarke, P Flicek, SF Garner, M Haimel, C Kempster, V Ladopoulos, AS Lenaerts, PM Materek, H McKinney, S Meacham, D Mead, M Nagy, CJ Penkett, A Rendon, D Seyres, B Sun, S Tuna, van der ME Weide, SW Wingett, JH Martens, O Stegle, S Richardson, L Vallier, DJ Roberts, K Freson, L Wernisch, HG Stunnenberg, J Danesh, P Fraser, N Soranzo, AS Butterworth, JW Heemskerk, E Turro, M Spivakov, WH Ouwehand, WJ Astle, K Downes, M Kostadima, M Frontini. Platelet function is modified by common sequence variation in megakaryocyte super enhancers. Nat Commun 2017;8:16058. doi:10.1038/ncomms16058
[BibTeX] [Abstract]
Linking non-coding genetic variants associated with the risk of diseases or disease-relevant traits to target genes is a crucial step to realize GWAS potential in the introduction of precision medicine. Here we set out to determine the mechanisms underpinning variant association with platelet quantitative traits using cell type-matched epigenomic data and promoter long-range interactions. We identify potential regulatory functions for 423 of 565 (75\%) non-coding variants associated with platelet traits and we demonstrate, through ex vivo and proof of principle genome editing validation, that variants in super enhancers play an important role in controlling archetypical platelet functions.
@Article{28703137, author = {Petersen R and Lambourne JJ and Javierre BM and Grassi L and Kreuzhuber R and Ruklisa D and Rosa IM and Tomé AR and Elding H and van Geffen JP and Jiang T and Farrow S and Cairns J and Al-Subaie AM and Ashford S and Attwood A and Batista J and Bouman H and Burden F and Choudry FA and Clarke L and Flicek P and Garner SF and Haimel M and Kempster C and Ladopoulos V and Lenaerts AS and Materek PM and McKinney H and Meacham S and Mead D and Nagy M and Penkett CJ and Rendon A and Seyres D and Sun B and Tuna S and van der Weide ME and Wingett SW and Martens JH and Stegle O and Richardson S and Vallier L and Roberts DJ and Freson K and Wernisch L and Stunnenberg HG and Danesh J and Fraser P and Soranzo N and Butterworth AS and Heemskerk JW and Turro E and Spivakov M and Ouwehand WH and Astle WJ and Downes K and Kostadima M and Frontini M}, title = {Platelet function is modified by common sequence variation in megakaryocyte super enhancers}, journal = {Nat Commun}, volume = {8}, pages = {16058}, year = {2017}, doi = {10.1038/ncomms16058}, abstract = {Linking non-coding genetic variants associated with the risk of diseases or disease-relevant traits to target genes is a crucial step to realize GWAS potential in the introduction of precision medicine. Here we set out to determine the mechanisms underpinning variant association with platelet quantitative traits using cell type-matched epigenomic data and promoter long-range interactions. We identify potential regulatory functions for 423 of 565 (75\%) non-coding variants associated with platelet traits and we demonstrate, through ex vivo and proof of principle genome editing validation, that variants in super enhancers play an important role in controlling archetypical platelet functions.},} - I Streeter, PW Harrison, A Faulconbridge, The HipSci Consortium, P Flicek, H Parkinson, L Clarke. The human-induced pluripotent stem cell initiative-data resources for cellular genetics. Nucleic Acids Res 2017;45(D1):D691–D697. doi:10.1093/nar/gkw928
[BibTeX] [Abstract]
The Human Induced Pluripotent Stem Cell Initiative (HipSci) isf establishing a large catalogue of human iPSC lines, arguably the most well characterized collection to date. The HipSci portal enables researchers to choose the right cell line for their experiment, and makes HipSci's rich catalogue of assay data easy to discover and reuse. Each cell line has genomic, transcriptomic, proteomic and cellular phenotyping data. Data are deposited in the appropriate EMBL-EBI archives, including the European Nucleotide Archive (ENA), European Genome-phenome Archive (EGA), ArrayExpress and PRoteomics IDEntifications (PRIDE) databases. The project will make 500 cell lines from healthy individuals, and from 150 patients with rare genetic diseases; these will be available through the European Collection of Authenticated Cell Cultures (ECACC). As of August 2016, 238 cell lines are available for purchase. Project data is presented through the HipSci data portal (http://www.hipsci.org/lines) and is downloadable from the associated FTP site (ftp://ftp.hipsci.ebi.ac.uk/vol1/ftp). The data portal presents a summary matrix of the HipSci cell lines, showing available data types. Each line has its own page containing descriptive metadata, quality information, and links to archived assay data. Analysis results are also available in a Track Hub, allowing visualization in the context of public genomic annotations (http://www.hipsci.org/data/trackhubs).
@Article{27733501, author = {Streeter I and Harrison PW and Faulconbridge A and The HipSci Consortium and Flicek P and Parkinson H and Clarke L}, title = {The human-induced pluripotent stem cell initiative-data resources for cellular genetics}, journal = {Nucleic Acids Res}, volume = {45}, number = {D1}, pages = {D691--D697}, year = {2017}, doi = {10.1093/nar/gkw928}, howpublished = {Advanced online publication: 12 October 2016}, abstract = {The Human Induced Pluripotent Stem Cell Initiative (HipSci) isf establishing a large catalogue of human iPSC lines, arguably the most well characterized collection to date. The HipSci portal enables researchers to choose the right cell line for their experiment, and makes HipSci's rich catalogue of assay data easy to discover and reuse. Each cell line has genomic, transcriptomic, proteomic and cellular phenotyping data. Data are deposited in the appropriate EMBL-EBI archives, including the European Nucleotide Archive (ENA), European Genome-phenome Archive (EGA), ArrayExpress and PRoteomics IDEntifications (PRIDE) databases. The project will make 500 cell lines from healthy individuals, and from 150 patients with rare genetic diseases; these will be available through the European Collection of Authenticated Cell Cultures (ECACC). As of August 2016, 238 cell lines are available for purchase. Project data is presented through the HipSci data portal (http://www.hipsci.org/lines) and is downloadable from the associated FTP site (ftp://ftp.hipsci.ebi.ac.uk/vol1/ftp). The data portal presents a summary matrix of the HipSci cell lines, showing available data types. Each line has its own page containing descriptive metadata, quality information, and links to archived assay data. Analysis results are also available in a Track Hub, allowing visualization in the context of public genomic annotations (http://www.hipsci.org/data/trackhubs).},} - L Clarke, S Fairley, X Zheng-Bradley, I Streeter, E Perry, E Lowy, AM Tassé, P Flicek. The international Genome sample resource (IGSR): A worldwide collection of genome variation incorporating the 1000 Genomes Project data. Nucleic Acids Res 2017;45(D1):D854–D859. doi:10.1093/nar/gkw829
[BibTeX] [Abstract]
The International Genome Sample Resource (IGSR; http://www.internationalgenome.org) expands in data type and population diversity the resources from the 1000 Genomes Project. IGSR represents the largest open collection of human variation data and provides easy access to these resources. IGSR was established in 2015 to maintain and extend the 1000 Genomes Project data, which has been widely used as a reference set of human variation and by researchers developing analysis methods. IGSR has mapped all of the 1000 Genomes sequence to the newest human reference (GRCh38), and will release updated variant calls to ensure maximal usefulness of the existing data. IGSR is collecting new structural variation data on the 1000 Genomes samples from long read sequencing and other technologies, and will collect relevant functional data into a single comprehensive resource. IGSR is extending coverage with new populations sequenced by collaborating groups. Here, we present the new data and analysis that IGSR has made available. We have also introduced a new data portal that increases discoverability of our data-previously only browseable through our FTP site-by focusing on particular samples, populations or data sets of interest.
@Article{27638885, author = {Clarke L and Fairley S and Zheng-Bradley X and Streeter I and Perry E and Lowy E and Tassé AM and Flicek P}, title = {The international Genome sample resource (IGSR): A worldwide collection of genome variation incorporating the 1000 Genomes Project data}, journal = {Nucleic Acids Res}, volume = {45}, number = {D1}, pages = {D854--D859}, year = {2017}, doi = {10.1093/nar/gkw829}, howpublished = {Advanced online publication: 15 September 2016}, abstract = {The International Genome Sample Resource (IGSR; http://www.internationalgenome.org) expands in data type and population diversity the resources from the 1000 Genomes Project. IGSR represents the largest open collection of human variation data and provides easy access to these resources. IGSR was established in 2015 to maintain and extend the 1000 Genomes Project data, which has been widely used as a reference set of human variation and by researchers developing analysis methods. IGSR has mapped all of the 1000 Genomes sequence to the newest human reference (GRCh38), and will release updated variant calls to ensure maximal usefulness of the existing data. IGSR is collecting new structural variation data on the 1000 Genomes samples from long read sequencing and other technologies, and will collect relevant functional data into a single comprehensive resource. IGSR is extending coverage with new populations sequenced by collaborating groups. Here, we present the new data and analysis that IGSR has made available. We have also introduced a new data portal that increases discoverability of our data-previously only browseable through our FTP site-by focusing on particular samples, populations or data sets of interest.},} - J MacArthur, E Bowler, M Cerezo, L Gil, P Hall, E Hastings, H Junkins, A McMahon, A Milano, J Morales, ZM Pendlington, D Welter, T Burdett, L Hindorff, P Flicek, F Cunningham, H Parkinson. The new NHGRI-EBI Catalog of published genome-wide association studies (GWAS Catalog). Nucleic Acids Res 2017;45(D1):D896–D901. doi:10.1093/nar/gkw1133
[BibTeX] [Abstract]
The NHGRI-EBI GWAS Catalog has provided data from published genome-wide association studies since 2008. In 2015, the database was redesigned and relocated to EMBL-EBI. The new infrastructure includes a new graphical user interface (www.ebi.ac.uk/gwas/), ontology supported search functionality and an improved curation interface. These developments have improved the data release frequency by increasing automation of curation and providing scaling improvements. The range of available Catalog data has also been extended with structured ancestry and recruitment information added for all studies. The infrastructure improvements also support scaling for larger arrays, exome and sequencing studies, allowing the Catalog to adapt to the needs of evolving study design, genotyping technologies and user needs in the future.
@Article{27899670, author = {MacArthur J and Bowler E and Cerezo M and Gil L and Hall P and Hastings E and Junkins H and McMahon A and Milano A and Morales J and Pendlington ZM and Welter D and Burdett T and Hindorff L and Flicek P and Cunningham F and Parkinson H}, title = {The new NHGRI-EBI Catalog of published genome-wide association studies (GWAS Catalog)}, journal = {Nucleic Acids Res}, volume = {45}, number = {D1}, pages = {D896--D901}, year = {2017}, doi = {10.1093/nar/gkw1133}, howpublished = {Advanced online publication: 29 November 2016}, abstract = {The NHGRI-EBI GWAS Catalog has provided data from published genome-wide association studies since 2008. In 2015, the database was redesigned and relocated to EMBL-EBI. The new infrastructure includes a new graphical user interface (www.ebi.ac.uk/gwas/), ontology supported search functionality and an improved curation interface. These developments have improved the data release frequency by increasing automation of curation and providing scaling improvements. The range of available Catalog data has also been extended with structured ancestry and recruitment information added for all studies. The infrastructure improvements also support scaling for larger arrays, exome and sequencing studies, allowing the Catalog to adapt to the needs of evolving study design, genotyping technologies and user needs in the future.},}
2016
- for Genomics {Global Alliance, Health.}. A federated ecosystem for sharing genomic, clinical data. Science 2016;352(6291):1278–1280. doi:10.1126/science.aaf6162
[BibTeX]@Article{27284183, author = {{Global Alliance for Genomics and Health.}}, title = {A federated ecosystem for sharing genomic, clinical data}, journal = {Science}, volume = {352}, number = {6291}, pages = {1278--1280}, year = {2016}, doi = {10.1126/science.aaf6162}, } - AC Queirós, R Beekman, R Vilarrasa-Blasi, M Duran-Ferrer, G Clot, A Merkel, E Raineri, N Russiñol, G Castellano, S Beà, A Navarro, M Kulis, N Verdaguer-Dot, P Jares, A Enjuanes, MJ Calasanz, A Bergmann, I Vater, I Salaverría, van de HJ Werken, WH Wilson, A Datta, P Flicek, R Royo, J Martens, E Giné, A Lopez-Guillermo, HG Stunnenberg, W Klapper, C Pott, S Heath, IG Gut, R Siebert, E Campo, JI Martín-Subero. Decoding the DNA Methylome of Mantle Cell Lymphoma in the Light of the Entire B Cell Lineage. Cancer Cell 2016;30(5):806–821. doi:10.1016/j.ccell.2016.09.014
[BibTeX] [Abstract]
We analyzed the in silico purified DNA methylation signatures of 82 mantle cell lymphomas (MCL) in comparison with cell subpopulations spanning the entire B cell lineage. We identified two MCL subgroups, respectively carrying epigenetic imprints of germinal-center-inexperienced and germinal-center-experienced B cells, and we found that DNA methylation profiles during lymphomagenesis are largely influenced by the methylation dynamics in normal B cells. An integrative epigenomic approach revealed 10,504 differentially methylated regions in regulatory elements marked by H3K27ac in MCL primary cases, including a distant enhancer showing de novo looping to the MCL oncogene SOX11. Finally, we observed that the magnitude of DNA methylation changes per case is highly variable and serves as an independent prognostic factor for MCL outcome.
@Article{27846393, author = {Queirós AC and Beekman R and Vilarrasa-Blasi R and Duran-Ferrer M and Clot G and Merkel A and Raineri E and Russiñol N and Castellano G and Beà S and Navarro A and Kulis M and Verdaguer-Dot N and Jares P and Enjuanes A and Calasanz MJ and Bergmann A and Vater I and Salaverría I and van de Werken HJ and Wilson WH and Datta A and Flicek P and Royo R and Martens J and Giné E and Lopez-Guillermo A and Stunnenberg HG and Klapper W and Pott C and Heath S and Gut IG and Siebert R and Campo E and Martín-Subero JI}, title = {Decoding the DNA Methylome of Mantle Cell Lymphoma in the Light of the Entire B Cell Lineage}, journal = {Cancer Cell}, volume = {30}, number = {5}, pages = {806--821}, year = {2016}, doi = {10.1016/j.ccell.2016.09.014}, abstract = {We analyzed the in silico purified DNA methylation signatures of 82 mantle cell lymphomas (MCL) in comparison with cell subpopulations spanning the entire B cell lineage. We identified two MCL subgroups, respectively carrying epigenetic imprints of germinal-center-inexperienced and germinal-center-experienced B cells, and we found that DNA methylation profiles during lymphomagenesis are largely influenced by the methylation dynamics in normal B cells. An integrative epigenomic approach revealed 10,504 differentially methylated regions in regulatory elements marked by H3K27ac in MCL primary cases, including a distant enhancer showing de novo looping to the MCL oncogene SOX11. Finally, we observed that the magnitude of DNA methylation changes per case is highly variable and serves as an independent prognostic factor for MCL outcome.},} - RP Schuyler, A Merkel, E Raineri, L Altucci, E Vellenga, JH Martens, F Pourfarzad, TW Kuijpers, F Burden, S Farrow, K Downes, WH Ouwehand, L Clarke, A Datta, E Lowy, P Flicek, M Frontini, HG Stunnenberg, JI Martín-Subero, I Gut, S Heath. Distinct Trends of DNA Methylation Patterning in the Innate and Adaptive Immune Systems. Cell Rep 2016;17(8):2101–2111. doi:10.1016/j.celrep.2016.10.054
[BibTeX] [Abstract]
DNA methylation and the localization and post-translational modification of nucleosomes are interdependent factors that contribute to the generation of distinct phenotypes from genetically identical cells. With 112 whole-genome bisulfite sequencing datasets from the BLUEPRINT Epigenome Project, we analyzed the global development of DNA methylation patterns during lineage commitment and maturation of a range of immune system effector cells and the cancers that arise from them. We show clear trends in methylation patterns that are distinct in the innate and adaptive arms of the human immune system, both globally and in relation to consistently positioned nucleosomes. Most notable are a progressive loss of methylation in developing lymphocytes and the consistent occurrence of non-CG methylation in specific cell types. Cancer samples from the two lineages are further polarized, suggesting the involvement of distinct lineage-specific epigenetic mechanisms. We anticipate broad utility for this resource as a basis for further comparative epigenetic analyses.
@Article{27851971, author = {Schuyler RP and Merkel A and Raineri E and Altucci L and Vellenga E and Martens JH and Pourfarzad F and Kuijpers TW and Burden F and Farrow S and Downes K and Ouwehand WH and Clarke L and Datta A and Lowy E and Flicek P and Frontini M and Stunnenberg HG and Martín-Subero JI and Gut I and Heath S}, title = {Distinct Trends of DNA Methylation Patterning in the Innate and Adaptive Immune Systems}, journal = {Cell Rep}, volume = {17}, number = {8}, pages = {2101--2111}, year = {2016}, doi = {10.1016/j.celrep.2016.10.054}, abstract = {DNA methylation and the localization and post-translational modification of nucleosomes are interdependent factors that contribute to the generation of distinct phenotypes from genetically identical cells. With 112 whole-genome bisulfite sequencing datasets from the BLUEPRINT Epigenome Project, we analyzed the global development of DNA methylation patterns during lineage commitment and maturation of a range of immune system effector cells and the cancers that arise from them. We show clear trends in methylation patterns that are distinct in the innate and adaptive arms of the human immune system, both globally and in relation to consistently positioned nucleosomes. Most notable are a progressive loss of methylation in developing lymphocytes and the consistent occurrence of non-CG methylation in specific cell types. Cancer samples from the two lineages are further polarized, suggesting the involvement of distinct lineage-specific epigenetic mechanisms. We anticipate broad utility for this resource as a basis for further comparative epigenetic analyses.},} - S Uebbing, A Künstner, H Mäkinen, N Backström, P Bolivar, R Burri, L Dutoit, CF Mugal, A Nater, B Aken, P Flicek, FJ Martin, SMJ Searle, H Ellegren. Divergence in gene expression within and between two closely related flycatcher species. Mol Ecol 2016;25(9):2015–2028. doi:10.1111/mec.13596
[BibTeX] [Abstract]
Relatively little is known about the character of gene expression evolution as species diverge. It is for instance unclear if gene expression generally evolves in a clock-like manner (by stabilizing selection or neutral evolution) or if there are frequent episodes of directional selection. To gain insights into the evolutionary divergence of gene expression, we sequenced and compared the transcriptomes of multiple organs from population samples of collared (Ficedula albicollis) and pied flycatchers (F. hypoleuca), two species which diverged less than one million years ago. Ordination analysis separated samples by organ rather than by species. Organs differed in their degrees of expression variance within species and expression divergence between species. Variance was negatively correlated with expression breadth and protein interactivity, suggesting that pleiotropic constraints reduce gene expression variance within species. Variance was correlated with between-species divergence, consistent with a pattern expected from stabilizing selection and neutral evolution. Using an expression PST approach, we identified genes differentially expressed between species and found 16 genes uniquely expressed in one of the species. For one of these, DPP7, uniquely expressed in collared flycatcher, the absence of expression in pied flycatcher could be associated with a ≈ 20 kb deletion including 11 out of 13 exons. This study of a young vertebrate speciation model system expands our knowledge of how gene expression evolves as natural populations become reproductively isolated. This article is protected by copyright. All rights reserved
@Article{26928872, author = {Uebbing S and Künstner A and Mäkinen H and Backström N and Bolivar P and Burri R and Dutoit L and Mugal CF and Nater A and Aken B and Flicek P and Martin FJ and Searle SMJ and Ellegren H}, title = {Divergence in gene expression within and between two closely related flycatcher species}, journal = {Mol Ecol}, volume = {25}, number = {9}, pages = {2015--2028}, year = {2016}, doi = {10.1111/mec.13596}, howpublished = {Advanced online publication: 2 April 2016}, abstract = {Relatively little is known about the character of gene expression evolution as species diverge. It is for instance unclear if gene expression generally evolves in a clock-like manner (by stabilizing selection or neutral evolution) or if there are frequent episodes of directional selection. To gain insights into the evolutionary divergence of gene expression, we sequenced and compared the transcriptomes of multiple organs from population samples of collared (Ficedula albicollis) and pied flycatchers (F. hypoleuca), two species which diverged less than one million years ago. Ordination analysis separated samples by organ rather than by species. Organs differed in their degrees of expression variance within species and expression divergence between species. Variance was negatively correlated with expression breadth and protein interactivity, suggesting that pleiotropic constraints reduce gene expression variance within species. Variance was correlated with between-species divergence, consistent with a pattern expected from stabilizing selection and neutral evolution. Using an expression PST approach, we identified genes differentially expressed between species and found 16 genes uniquely expressed in one of the species. For one of these, DPP7, uniquely expressed in collared flycatcher, the absence of expression in pied flycatcher could be associated with a ≈ 20 kb deletion including 11 out of 13 exons. This study of a young vertebrate speciation model system expands our knowledge of how gene expression evolves as natural populations become reproductively isolated. This article is protected by copyright. All rights reserved},} - A Yates, W Akanni, MR Amode, D Barrell, K Billis, D Carvalho-Silva, C Cummins, P Clapham, S Fitzgerald, L Gil, CG Girón, L Gordon, T Hourlier, SE Hunt, SH Janacek, N Johnson, T Juettemann, S Keenan, I Lavidas, FJ Martin, T Maurel, W McLaren, DN Murphy, R Nag, M Nuhn, A Parker, M Patricio, M Pignatelli, M Rahtz, HS Riat, D Sheppard, K Taylor, A Thormann, A Vullo, SP Wilder, A Zadissa, E Birney, J Harrow, M Muffato, E Perry, M Ruffier, G Spudich, SJ Trevanion, F Cunningham, BL Aken, DR Zerbino, P Flicek. Ensembl 2016. Nucleic Acids Res 2016;44(D1):D710–6. doi:10.1093/nar/gkv1157
[BibTeX] [Abstract]
The Ensembl project (http://www.ensembl.org) is a system for genome annotation, analysis, storage and dissemination designed to facilitate the access of genomic annotation from chordates and key model organisms. It provides access to data from 87 species across our main and early access Pre! websites. This year we introduced three newly annotated species and released numerous updates across our supported species with a concentration on data for the latest genome assemblies of human, mouse, zebrafish and rat. We also provided two data updates for the previous human assembly, GRCh37, through a dedicated website (http://grch37.ensembl.org). Our tools, in particular the VEP, have been improved significantly through integration of additional third party data. REST is now capable of larger-scale analysis and our regulatory data BioMart can deliver faster results. The website is now capable of displaying long-range interactions such as those found in cis-regulated datasets. Finally we have launched a website optimized for mobile devices providing views of genes, variants and phenotypes. Our data is made available without restriction and all code is available from our GitHub organization site (http://github.com/Ensembl) under an Apache 2.0 license
@Article{26687719, author = {Yates A and Akanni W and Amode MR and Barrell D and Billis K and Carvalho-Silva D and Cummins C and Clapham P and Fitzgerald S and Gil L and Girón CG and Gordon L and Hourlier T and Hunt SE and Janacek SH and Johnson N and Juettemann T and Keenan S and Lavidas I and Martin FJ and Maurel T and McLaren W and Murphy DN and Nag R and Nuhn M and Parker A and Patricio M and Pignatelli M and Rahtz M and Riat HS and Sheppard D and Taylor K and Thormann A and Vullo A and Wilder SP and Zadissa A and Birney E and Harrow J and Muffato M and Perry E and Ruffier M and Spudich G and Trevanion SJ and Cunningham F and Aken BL and Zerbino DR and Flicek P}, title = {Ensembl 2016}, journal = {Nucleic Acids Res}, volume = {44}, number = {D1}, pages = {D710--6}, year = {2016}, doi = {10.1093/nar/gkv1157}, howpublished = {Advanced online publication: 19 December 2015}, abstract = {The Ensembl project (http://www.ensembl.org) is a system for genome annotation, analysis, storage and dissemination designed to facilitate the access of genomic annotation from chordates and key model organisms. It provides access to data from 87 species across our main and early access Pre! websites. This year we introduced three newly annotated species and released numerous updates across our supported species with a concentration on data for the latest genome assemblies of human, mouse, zebrafish and rat. We also provided two data updates for the previous human assembly, GRCh37, through a dedicated website (http://grch37.ensembl.org). Our tools, in particular the VEP, have been improved significantly through integration of additional third party data. REST is now capable of larger-scale analysis and our regulatory data BioMart can deliver faster results. The website is now capable of displaying long-range interactions such as those found in cis-regulated datasets. Finally we have launched a website optimized for mobile devices providing views of genes, variants and phenotypes. Our data is made available without restriction and all code is available from our GitHub organization site (http://github.com/Ensembl) under an Apache 2.0 license},} - J Herrero, M Muffato, K Beal, S Fitzgerald, L Gordon, M Pignatelli, AJ Vilella, SMJ Searle, R Amode, S Brent, W Spooner, E Kulesha, A Yates, P Flicek. Ensembl comparative genomics resources. Database (Oxford) 2016;2016:bav096. doi:10.1093/database/bav096
[BibTeX] [Abstract]
Evolution provides the unifying framework with which to understand biology. The coherent investigation of genic and genomic data often requires comparative genomics analyses based on whole-genome alignments, sets of homologous genes and other relevant datasets in order to evaluate and answer evolutionary-related questions. However, the complexity and computational requirements of producing such data are substantial: this has led to only a small number of reference resources that are used for most comparative analyses. The Ensembl comparative genomics resources are one such reference set that facilitates comprehensive and reproducible analysis of chordate genome data. Ensembl computes pairwise and multiple whole-genome alignments from which large-scale synteny, per-base conservation scores and constrained elements are obtained. Gene alignments are used to define Ensembl Protein Families, GeneTrees and homologies for both protein-coding and non-coding RNA genes. These resources are updated frequently and have a consistent informatics infrastructure and data presentation across all supported species. Specialized web-based visualizations are also available including synteny displays, collapsible gene tree plots, a gene family locator and different alignment views. The Ensembl comparative genomics infrastructure is extensively reused for the analysis of non-vertebrate species by other projects including Ensembl Genomes and Gramene and much of the information here is relevant to these projects. The consistency of the annotation across species and the focus on vertebrates makes Ensembl an ideal system to perform and support vertebrate comparative genomic analyses. We use robust software and pipelines to produce reference comparative data and make it freely available.Database URL: http://www.ensembl.org
@Article{26896847, author = {Herrero J and Muffato M and Beal K and Fitzgerald S and Gordon L and Pignatelli M and Vilella AJ and Searle SMJ and Amode R and Brent S and Spooner W and Kulesha E and Yates A and Flicek P}, title = {Ensembl comparative genomics resources}, journal = {Database (Oxford)}, volume = {2016}, pages = {bav096}, year = {2016}, doi = {10.1093/database/bav096}, abstract = {Evolution provides the unifying framework with which to understand biology. The coherent investigation of genic and genomic data often requires comparative genomics analyses based on whole-genome alignments, sets of homologous genes and other relevant datasets in order to evaluate and answer evolutionary-related questions. However, the complexity and computational requirements of producing such data are substantial: this has led to only a small number of reference resources that are used for most comparative analyses. The Ensembl comparative genomics resources are one such reference set that facilitates comprehensive and reproducible analysis of chordate genome data. Ensembl computes pairwise and multiple whole-genome alignments from which large-scale synteny, per-base conservation scores and constrained elements are obtained. Gene alignments are used to define Ensembl Protein Families, GeneTrees and homologies for both protein-coding and non-coding RNA genes. These resources are updated frequently and have a consistent informatics infrastructure and data presentation across all supported species. Specialized web-based visualizations are also available including synteny displays, collapsible gene tree plots, a gene family locator and different alignment views. The Ensembl comparative genomics infrastructure is extensively reused for the analysis of non-vertebrate species by other projects including Ensembl Genomes and Gramene and much of the information here is relevant to these projects. The consistency of the annotation across species and the focus on vertebrates makes Ensembl an ideal system to perform and support vertebrate comparative genomic analyses. We use robust software and pipelines to produce reference comparative data and make it freely available.Database URL: http://www.ensembl.org},} - DR Zerbino, N Johnson, T Juetteman, D Sheppard, SP Wilder, I Lavidas, M Nuhn, E Perry, Q Raffaillac-Desfosses, D Sobral, D Keefe, S Gräf, I Ahmed, R Kinsella, B Pritchard, S Brent, R Amode, A Parker, S Trevanion, E Birney, I Dunham, P Flicek. Ensembl regulation resources. Database (Oxford) 2016;2016:bav119. doi:10.1093/database/bav119
[BibTeX] [Abstract]
New experimental techniques in epigenomics allow researchers to assay a diversity of highly dynamic features such as histone marks, DNA modifications or chromatin structure. The study of their fluctuations should provide insights into gene expression regulation, cell differentiation and disease. The Ensembl project collects and maintains the Ensembl regulation data resources on epigenetic marks, transcription factor binding and DNA methylation for human and mouse, as well as microarray probe mappings and annotations for a variety of chordate genomes. From this data, we produce a functional annotation of the regulatory elements along the human and mouse genomes with plans to expand to other species as data becomes available. Starting from well-studied cell lines, we will progressively expand our library of measurements to a greater variety of samples. Ensembl's regulation resources provide a central and easy-to-query repository for reference epigenomes. As with all Ensembl data, it is freely available at http://www.ensembl.org, from the Perl and REST APIs and from the public Ensembl MySQL database server at ensembldb.ensembl.org.Database URL: http://www.ensembl.org
@Article{26888907, author = {Zerbino DR and Johnson N and Juetteman T and Sheppard D and Wilder SP and Lavidas I and Nuhn M and Perry E and Raffaillac-Desfosses Q and Sobral D and Keefe D and Gräf S and Ahmed I and Kinsella R and Pritchard B and Brent S and Amode R and Parker A and Trevanion S and Birney E and Dunham I and Flicek P}, title = {Ensembl regulation resources}, journal = {Database (Oxford)}, volume = {2016}, pages = {bav119}, year = {2016}, doi = {10.1093/database/bav119}, abstract = {New experimental techniques in epigenomics allow researchers to assay a diversity of highly dynamic features such as histone marks, DNA modifications or chromatin structure. The study of their fluctuations should provide insights into gene expression regulation, cell differentiation and disease. The Ensembl project collects and maintains the Ensembl regulation data resources on epigenetic marks, transcription factor binding and DNA methylation for human and mouse, as well as microarray probe mappings and annotations for a variety of chordate genomes. From this data, we produce a functional annotation of the regulatory elements along the human and mouse genomes with plans to expand to other species as data becomes available. Starting from well-studied cell lines, we will progressively expand our library of measurements to a greater variety of samples. Ensembl's regulation resources provide a central and easy-to-query repository for reference epigenomes. As with all Ensembl data, it is freely available at http://www.ensembl.org, from the Perl and REST APIs and from the public Ensembl MySQL database server at ensembldb.ensembl.org.Database URL: http://www.ensembl.org},} - L Chen, B Ge, FP Casale, L Vasquez, T Kwan, D Garrido-Martín, S Watt, Y Yan, K Kundu, S Ecker, A Datta, D Richardson, F Burden, D Mead, AL Mann, JM Fernandez, S Rowlston, SP Wilder, S Farrow, X Shao, JJ Lambourne, A Redensek, CA Albers, V Amstislavskiy, S Ashford, K Berentsen, L Bomba, G Bourque, D Bujold, S Busche, M Caron, SH Chen, W Cheung, O Delaneau, ET Dermitzakis, H Elding, I Colgiu, FO Bagger, P Flicek, E Habibi, V Iotchkova, E Janssen-Megens, B Kim, H Lehrach, E Lowy, A Mandoli, F Matarese, MT Maurano, JA Morris, V Pancaldi, F Pourfarzad, K Rehnstrom, A Rendon, T Risch, N Sharifi, MM Simon, M Sultan, A Valencia, K Walter, SY Wang, M Frontini, SE Antonarakis, L Clarke, ML Yaspo, S Beck, R Guigo, D Rico, JH Martens, WH Ouwehand, TW Kuijpers, DS Paul, HG Stunnenberg, O Stegle, K Downes, T Pastinen, N Soranzo. Genetic Drivers of Epigenetic and Transcriptional Variation in Human Immune Cells. Cell 2016;167(5):1398–1414.e24. doi:10.1016/j.cell.2016.10.026
[BibTeX] [Abstract]
Characterizing the multifaceted contribution of genetic and epigenetic factors to disease phenotypes is a major challenge in human genetics and medicine. We carried out high-resolution genetic, epigenetic, and transcriptomic profiling in three major human immune cell types (CD14(+) monocytes, CD16(+) neutrophils, and naive CD4(+) T cells) from up to 197 individuals. We assess, quantitatively, the relative contribution of cis-genetic and epigenetic factors to transcription and evaluate their impact as potential sources of confounding in epigenome-wide association studies. Further, we characterize highly coordinated genetic effects on gene expression, methylation, and histone variation through quantitative trait locus (QTL) mapping and allele-specific (AS) analyses. Finally, we demonstrate colocalization of molecular trait QTLs at 345 unique immune disease loci. This expansive, high-resolution atlas of multi-omics changes yields insights into cell-type-specific correlation between diverse genomic inputs, more generalizable correlations between these inputs, and defines molecular events that may underpin complex disease risk.
@Article{27863251, author = {Chen L and Ge B and Casale FP and Vasquez L and Kwan T and Garrido-Martín D and Watt S and Yan Y and Kundu K and Ecker S and Datta A and Richardson D and Burden F and Mead D and Mann AL and Fernandez JM and Rowlston S and Wilder SP and Farrow S and Shao X and Lambourne JJ and Redensek A and Albers CA and Amstislavskiy V and Ashford S and Berentsen K and Bomba L and Bourque G and Bujold D and Busche S and Caron M and Chen SH and Cheung W and Delaneau O and Dermitzakis ET and Elding H and Colgiu I and Bagger FO and Flicek P and Habibi E and Iotchkova V and Janssen-Megens E and Kim B and Lehrach H and Lowy E and Mandoli A and Matarese F and Maurano MT and Morris JA and Pancaldi V and Pourfarzad F and Rehnstrom K and Rendon A and Risch T and Sharifi N and Simon MM and Sultan M and Valencia A and Walter K and Wang SY and Frontini M and Antonarakis SE and Clarke L and Yaspo ML and Beck S and Guigo R and Rico D and Martens JH and Ouwehand WH and Kuijpers TW and Paul DS and Stunnenberg HG and Stegle O and Downes K and Pastinen T and Soranzo N}, title = {Genetic Drivers of Epigenetic and Transcriptional Variation in Human Immune Cells}, journal = {Cell}, volume = {167}, number = {5}, pages = {1398--1414.e24}, year = {2016}, doi = {10.1016/j.cell.2016.10.026}, abstract = {Characterizing the multifaceted contribution of genetic and epigenetic factors to disease phenotypes is a major challenge in human genetics and medicine. We carried out high-resolution genetic, epigenetic, and transcriptomic profiling in three major human immune cell types (CD14(+) monocytes, CD16(+) neutrophils, and naive CD4(+) T cells) from up to 197 individuals. We assess, quantitatively, the relative contribution of cis-genetic and epigenetic factors to transcription and evaluate their impact as potential sources of confounding in epigenome-wide association studies. Further, we characterize highly coordinated genetic effects on gene expression, methylation, and histone variation through quantitative trait locus (QTL) mapping and allele-specific (AS) analyses. Finally, we demonstrate colocalization of molecular trait QTLs at 345 unique immune disease loci. This expansive, high-resolution atlas of multi-omics changes yields insights into cell-type-specific correlation between diverse genomic inputs, more generalizable correlations between these inputs, and defines molecular events that may underpin complex disease risk.},} - C Kaloff, K Anastassiadis, A Ayadi, R Baldock, J Beig, M-C Birling, A Bradley, SDM Brown, A Bürger, W Bushell, F Chiani, FS Collins, B Doe, JT Eppig, RH Finnell, C Fletcher, P Flicek, M Fray, RH Friedel, A Gambadoro, H Gates, J Hansen, Y Herault, GG Hicks, A Hörlein, de M Hrabé Angelis, V Iyer, de PJ Jong, G Koscielny, R Kühn, P Liu, KCK Lloyd, RG Lopez, S Marschall, S Martínez, C McKerlie, T Meehan, von H Melchner, M Moore, SA Murray, A Nagy, LMJ Nutter, G Pavlovic, A Pombero, H Prosser, R Ramirez-Solis, M Ringwald, B Rosen, N Rosenthal, J Rossant, P Ruiz Noppinger, E Ryder, WC Skarnes, J Schick, F Schnütgen, P Schofield, C Seisenberger, M Selloum, D Smedley, EM Simpson, AF Stewart, L Teboul, GP Tocchini Valentini, D Valenzuela, AP West, W Wurst. Genome wide conditional mouse knockout resources. Drug Discovery Today: Disease Models 2016;20:3–12. doi:10.1016/j.ddmod.2017.08.002
[BibTeX] [Abstract]
The International Knockout Mouse Consortium (IKMC) developed high throughput gene trapping and gene targeting pipelines that produced mostly conditional mutations of more than 18,500 genes in C57BL/6N mouse embryonic stem (ES) cells which have been archived and are freely available to the research community as a frozen resource. From this unprecedented resource more than 6000 mutant mouse strains have been generated by the IKMC in collaboration with the International Mouse Phenotyping Consortium (IMPC). In addition, a cre-driver resource was established including 250 C57BL/6 cre-inducible mouse strains. Complementing the cre-driver resource, a collection comprising 27 rAAVs expressing cre in a tissue-specific manner has also been produced. All resources are easily accessible from the IKMC/IMPC web portal (www.mousephenotype.org). The IKMC/IMPC resource is a standardized reference library of mouse models with defined genetic backgrounds enabling the analysis of gene-disease associations in mice of different genetic makeup and should therefore have a major impact on biomedical research.
@Article{39132094, author = {Kaloff C and Anastassiadis K and Ayadi A and Baldock R and Beig J and Birling M-C and Bradley A and Brown SDM and Bürger A and Bushell W and Chiani F and Collins FS and Doe B and Eppig JT and Finnell RH and Fletcher C and Flicek P and Fray M and Friedel RH and Gambadoro A and Gates H and Hansen J and Herault Y and Hicks GG and Hörlein A and Hrabé de Angelis M and Iyer V and de Jong PJ and Koscielny G and Kühn R and Liu P and Lloyd KCK and Lopez RG and Marschall S and Martínez S and McKerlie C and Meehan T and von Melchner H and Moore M and Murray SA and Nagy A and Nutter LMJ and Pavlovic G and Pombero A and Prosser H and Ramirez-Solis R and Ringwald M and Rosen B and Rosenthal N and Rossant J and Ruiz Noppinger P and Ryder E and Skarnes WC and Schick J and Schnütgen F and Schofield P and Seisenberger C and Selloum M and Smedley D and Simpson EM and Stewart AF and Teboul L and Tocchini Valentini GP and Valenzuela D and West AP and Wurst W}, title = {Genome wide conditional mouse knockout resources}, journal = {Drug Discovery Today: Disease Models}, volume = {20}, pages = {3--12}, year = {2016}, doi = {10.1016/j.ddmod.2017.08.002}, abstract = {The International Knockout Mouse Consortium (IKMC) developed high throughput gene trapping and gene targeting pipelines that produced mostly conditional mutations of more than 18,500 genes in C57BL/6N mouse embryonic stem (ES) cells which have been archived and are freely available to the research community as a frozen resource. From this unprecedented resource more than 6000 mutant mouse strains have been generated by the IKMC in collaboration with the International Mouse Phenotyping Consortium (IMPC). In addition, a cre-driver resource was established including 250 C57BL/6 cre-inducible mouse strains. Complementing the cre-driver resource, a collection comprising 27 rAAVs expressing cre in a tissue-specific manner has also been produced. All resources are easily accessible from the IKMC/IMPC web portal (www.mousephenotype.org). The IKMC/IMPC resource is a standardized reference library of mouse models with defined genetic backgrounds enabling the analysis of gene-disease associations in mice of different genetic makeup and should therefore have a major impact on biomedical research.},} - ME Dickinson, AM Flenniken, X Ji, L Teboul, MD Wong, JK White, TF Meehan, WJ Weninger, H Westerberg, H Adissu, CN Baker, L Bower, JM Brown, LB Caddle, F Chiani, D Clary, J Cleak, MJ Daly, JM Denegre, B Doe, ME Dolan, SM Edie, H Fuchs, V Gailus-Durner, A Galli, A Gambadoro, J Gallegos, S Guo, NR Horner, CW Hsu, SJ Johnson, S Kalaga, LC Keith, L Lanoue, TN Lawson, M Lek, M Mark, S Marschall, J Mason, ML McElwee, S Newbigging, LM Nutter, KA Peterson, R Ramirez-Solis, DJ Rowland, E Ryder, KE Samocha, JR Seavitt, M Selloum, Z Szoke-Kovacs, M Tamura, AG Trainor, I Tudose, S Wakana, J Warren, O Wendling, DB West, L Wong, A Yoshiki, MPC International, L Jackson, ICDLSICS Infrastructure Nationale PHENOMIN, RL Charles, H MRC, CFP Toronto, TSI Wellcome, BC RIKEN, DG MacArthur, GP Tocchini-Valentini, X Gao, P Flicek, A Bradley, WC Skarnes, MJ Justice, HE Parkinson, M Moore, S Wells, RE Braun, KL Svenson, de MH Angelis, Y Herault, T Mohun, AM Mallon, RM Henkelman, SD Brown, DJ Adams, KC Lloyd, C McKerlie, AL Beaudet, M Bućan, SA Murray. High-throughput discovery of novel developmental phenotypes. Nature 2016;537(7621):508–514. doi:10.1038/nature19356
[BibTeX] [Abstract]
Approximately one-third of all mammalian genes are essential for life. Phenotypes resulting from knockouts of these genes in mice have provided tremendous insight into gene function and congenital disorders. As part of the International Mouse Phenotyping Consortium effort to generate and phenotypically characterize 5,000 knockout mouse lines, here we identify 410 lethal genes during the production of the first 1,751 unique gene knockouts. Using a standardized phenotyping platform that incorporates high-resolution 3D imaging, we identify phenotypes at multiple time points for previously uncharacterized genes and additional phenotypes for genes with previously reported mutant phenotypes. Unexpectedly, our analysis reveals that incomplete penetrance and variable expressivity are common even on a defined genetic background. In addition, we show that human disease genes are enriched for essential genes, thus providing a dataset that facilitates the prioritization and validation of mutations identified in clinical sequencing efforts.
@Article{27626380, author = {Dickinson ME and Flenniken AM and Ji X and Teboul L and Wong MD and White JK and Meehan TF and Weninger WJ and Westerberg H and Adissu H and Baker CN and Bower L and Brown JM and Caddle LB and Chiani F and Clary D and Cleak J and Daly MJ and Denegre JM and Doe B and Dolan ME and Edie SM and Fuchs H and Gailus-Durner V and Galli A and Gambadoro A and Gallegos J and Guo S and Horner NR and Hsu CW and Johnson SJ and Kalaga S and Keith LC and Lanoue L and Lawson TN and Lek M and Mark M and Marschall S and Mason J and McElwee ML and Newbigging S and Nutter LM and Peterson KA and Ramirez-Solis R and Rowland DJ and Ryder E and Samocha KE and Seavitt JR and Selloum M and Szoke-Kovacs Z and Tamura M and Trainor AG and Tudose I and Wakana S and Warren J and Wendling O and West DB and Wong L and Yoshiki A and International MPC and Jackson L and Infrastructure Nationale PHENOMIN ICDLSICS and Charles RL and MRC H and Toronto CFP and Wellcome TSI and RIKEN BC and MacArthur DG and Tocchini-Valentini GP and Gao X and Flicek P and Bradley A and Skarnes WC and Justice MJ and Parkinson HE and Moore M and Wells S and Braun RE and Svenson KL and de Angelis MH and Herault Y and Mohun T and Mallon AM and Henkelman RM and Brown SD and Adams DJ and Lloyd KC and McKerlie C and Beaudet AL and Bućan M and Murray SA}, title = {High-throughput discovery of novel developmental phenotypes}, journal = {Nature}, volume = {537}, number = {7621}, pages = {508--514}, year = {2016}, doi = {10.1038/nature19356}, howpublished = {Advanced online publication: 14 September 2016}, abstract = {Approximately one-third of all mammalian genes are essential for life. Phenotypes resulting from knockouts of these genes in mice have provided tremendous insight into gene function and congenital disorders. As part of the International Mouse Phenotyping Consortium effort to generate and phenotypically characterize 5,000 knockout mouse lines, here we identify 410 lethal genes during the production of the first 1,751 unique gene knockouts. Using a standardized phenotyping platform that incorporates high-resolution 3D imaging, we identify phenotypes at multiple time points for previously uncharacterized genes and additional phenotypes for genes with previously reported mutant phenotypes. Unexpectedly, our analysis reveals that incomplete penetrance and variable expressivity are common even on a defined genetic background. In addition, we show that human disease genes are enriched for essential genes, thus providing a dataset that facilitates the prioritization and validation of mutations identified in clinical sequencing efforts.},} - DS Paul, AE Teschendorff, MA Dang, R Lowe, MI Hawa, S Ecker, H Beyan, S Cunningham, AR Fouts, A Ramelius, F Burden, S Farrow, S Rowlston, K Rehnstrom, M Frontini, K Downes, S Busche, WA Cheung, B Ge, MM Simon, D Bujold, T Kwan, G Bourque, A Datta, E Lowy, L Clarke, P Flicek, E Libertini, S Heath, M Gut, IG Gut, WH Ouwehand, T Pastinen, N Soranzo, SE Hofer, B Karges, T Meissner, BO Boehm, C Cilio, H Elding Larsson, Å Lernmark, AK Steck, VK Rakyan, S Beck, RD Leslie. Increased DNA methylation variability in type 1 diabetes across three immune effector cell types. Nat Commun 2016;7:13555. doi:10.1038/ncomms13555
[BibTeX] [Abstract]
The incidence of type 1 diabetes (T1D) has substantially increased over the past decade, suggesting a role for non-genetic factors such as epigenetic mechanisms in disease development. Here we present an epigenome-wide association study across 406,365 CpGs in 52 monozygotic twin pairs discordant for T1D in three immune effector cell types. We observe a substantial enrichment of differentially variable CpG positions (DVPs) in T1D twins when compared with their healthy co-twins and when compared with healthy, unrelated individuals. These T1D-associated DVPs are found to be temporally stable and enriched at gene regulatory elements. Integration with cell type-specific gene regulatory circuits highlight pathways involved in immune cell metabolism and the cell cycle, including mTOR signalling. Evidence from cord blood of newborns who progress to overt T1D suggests that the DVPs likely emerge after birth. Our findings, based on 772 methylomes, implicate epigenetic changes that could contribute to disease pathogenesis in T1D.
@Article{27898055, author = {Paul DS and Teschendorff AE and Dang MA and Lowe R and Hawa MI and Ecker S and Beyan H and Cunningham S and Fouts AR and Ramelius A and Burden F and Farrow S and Rowlston S and Rehnstrom K and Frontini M and Downes K and Busche S and Cheung WA and Ge B and Simon MM and Bujold D and Kwan T and Bourque G and Datta A and Lowy E and Clarke L and Flicek P and Libertini E and Heath S and Gut M and Gut IG and Ouwehand WH and Pastinen T and Soranzo N and Hofer SE and Karges B and Meissner T and Boehm BO and Cilio C and Elding Larsson H and Lernmark Å and Steck AK and Rakyan VK and Beck S and Leslie RD}, title = {Increased DNA methylation variability in type 1 diabetes across three immune effector cell types}, journal = {Nat Commun}, volume = {7}, pages = {13555}, year = {2016}, doi = {10.1038/ncomms13555}, abstract = {The incidence of type 1 diabetes (T1D) has substantially increased over the past decade, suggesting a role for non-genetic factors such as epigenetic mechanisms in disease development. Here we present an epigenome-wide association study across 406,365 CpGs in 52 monozygotic twin pairs discordant for T1D in three immune effector cell types. We observe a substantial enrichment of differentially variable CpG positions (DVPs) in T1D twins when compared with their healthy co-twins and when compared with healthy, unrelated individuals. These T1D-associated DVPs are found to be temporally stable and enriched at gene regulatory elements. Integration with cell type-specific gene regulatory circuits highlight pathways involved in immune cell metabolism and the cell cycle, including mTOR signalling. Evidence from cord blood of newborns who progress to overt T1D suggests that the DVPs likely emerge after birth. Our findings, based on 772 methylomes, implicate epigenetic changes that could contribute to disease pathogenesis in T1D.},} - L Fontanesi, F Di Palma, P Flicek, AT Smith, C-G Thulin, PC Alves. LaGomiCs - Lagomorph Genomics Consortium: an international collaborative effort for sequencing the genomes of an entire mammalian order. J Hered 2016;107(4):295–308. doi:10.1093/jhered/esw010
[BibTeX] [Abstract]
The order Lagomorpha comprises about 90 living species, divided in two families: the pikas (Family Ochotonidae), and the rabbits, hare and jackrabbits (Family Leporidae). Lagomorphs are important economically and scientifically as a major human food resource, valued game species, pests of agricultural significance, model laboratory animals, and key elements in food webs. A quarter of the lagomorph species are listed as threatened. They are native to all continents except Antarctica, and occur up to 5,000 m above sea level, from the equator to the Arctic, spanning a wide range of environmental conditions. The order has notable taxonomic problems presenting significant difficulties for defining a species due to broad phenotypic variation, overlap of morphological characteristics, and relatively recent speciation events. At present, only the genomes of two species, the European rabbit (Oryctolagus cuniculus) and American pika (Ochotona princeps) have been sequenced and assembled. Starting from a paucity of genome information, the main scientific aim of the Lagomorph Genomics Consortium (LaGomiCs), born from a cooperative initiative of the European COST Action ``A Collaborative European Network on Rabbit Genome Biology - RGB-Net'' and the World Lagomorph Society (WLS), is to provide an international framework for the sequencing of the genome of all extant and selected extinct lagomorphs. Sequencing the genomes of an entire order will provide a large amount of information to address biological problems not only related to lagomorphs but also to all mammals. We present current and planned sequencing programs and outline the final objective of LaGomiCs possible through broad international collaboration
@Article{26921276, author = {Fontanesi L and Di Palma F and Flicek P and Smith AT and Thulin C-G and Alves PC}, title = {LaGomiCs - Lagomorph Genomics Consortium: an international collaborative effort for sequencing the genomes of an entire mammalian order}, journal = {J Hered}, volume = {107}, number = {4}, pages = {295--308}, year = {2016}, doi = {10.1093/jhered/esw010}, howpublished = {Advanced online publication: 26 February 2016}, abstract = {The order Lagomorpha comprises about 90 living species, divided in two families: the pikas (Family Ochotonidae), and the rabbits, hare and jackrabbits (Family Leporidae). Lagomorphs are important economically and scientifically as a major human food resource, valued game species, pests of agricultural significance, model laboratory animals, and key elements in food webs. A quarter of the lagomorph species are listed as threatened. They are native to all continents except Antarctica, and occur up to 5,000 m above sea level, from the equator to the Arctic, spanning a wide range of environmental conditions. The order has notable taxonomic problems presenting significant difficulties for defining a species due to broad phenotypic variation, overlap of morphological characteristics, and relatively recent speciation events. At present, only the genomes of two species, the European rabbit (Oryctolagus cuniculus) and American pika (Ochotona princeps) have been sequenced and assembled. Starting from a paucity of genome information, the main scientific aim of the Lagomorph Genomics Consortium (LaGomiCs), born from a cooperative initiative of the European COST Action ``A Collaborative European Network on Rabbit Genome Biology - RGB-Net'' and the World Lagomorph Society (WLS), is to provide an international framework for the sequencing of the genome of all extant and selected extinct lagomorphs. Sequencing the genomes of an entire order will provide a large amount of information to address biological problems not only related to lagomorphs but also to all mammals. We present current and planned sequencing programs and outline the final objective of LaGomiCs possible through broad international collaboration},} - C Auffray, R Balling, I Barroso, L Bencze, M Benson, J Bergeron, E Bernal-Delgado, N Blomberg, C Bock, A Conesa, S Del Signore, C Delogne, P Devilee, A Di Meglio, M Eijkemans, P Flicek, N Graf, V Grimm, H-J Guchelaar, Y-K Guo, IG Gut, A Hanbury, S Hanif, R-D Hilgers, Á Honrado, DR Hose, J Houwing-Duistermaat, T Hubbard, SH Janacek, H Karanikas, T Kievits, M Kohler, A Kremer, J Lanfear, T Lengauer, E Maes, T Meert, W Müller, D Nickel, P Oledzki, B Pedersen, M Petkovic, K Pliakos, M Rattray, JR IMàs, R Schneider, T Sengstag, X Serra-Picamal, W Spek, LAI Vaas, van O Batenburg, M Vandelaer, P Varnai, P Villoslada, JA Vizcaíno, JPM Wubbe, G Zanetti. Making sense of big data in health research: Towards an EU action plan. Genome Med 2016;8(1):71. doi:10.1186/s13073-016-0323-y
[BibTeX] [Abstract]
Medicine and healthcare are undergoing profound changes. Whole-genome sequencing and high-resolution imaging technologies are key drivers of this rapid and crucial transformation. Technological innovation combined with automation and miniaturization has triggered an explosion in data production that will soon reach exabyte proportions. How are we going to deal with this exponential increase in data production? The potential of ``big data'' for improving health is enormous but, at the same time, we face a wide range of challenges to overcome urgently. Europe is very proud of its cultural diversity; however, exploitation of the data made available through advances in genomic medicine, imaging, and a wide range of mobile health applications or connected devices is hampered by numerous historical, technical, legal, and political barriers. European health systems and databases are diverse and fragmented. There is a lack of harmonization of data formats, processing, analysis, and data transfer, which leads to incompatibilities and lost opportunities. Legal frameworks for data sharing are evolving. Clinicians, researchers, and citizens need improved methods, tools, and training to generate, analyze, and query data effectively. Addressing these barriers will contribute to creating the European Single Market for health, which will improve health and healthcare for all Europeans
@Article{27338147, author = {Auffray C and Balling R and Barroso I and Bencze L and Benson M and Bergeron J and Bernal-Delgado E and Blomberg N and Bock C and Conesa A and Del Signore S and Delogne C and Devilee P and Di Meglio A and Eijkemans M and Flicek P and Graf N and Grimm V and Guchelaar H-J and Guo Y-K and Gut IG and Hanbury A and Hanif S and Hilgers R-D and Honrado Á and Hose DR and Houwing-Duistermaat J and Hubbard T and Janacek SH and Karanikas H and Kievits T and Kohler M and Kremer A and Lanfear J and Lengauer T and Maes E and Meert T and Müller W and Nickel D and Oledzki P and Pedersen B and Petkovic M and Pliakos K and Rattray M and I Màs JR and Schneider R and Sengstag T and Serra-Picamal X and Spek W and Vaas LAI and van Batenburg O and Vandelaer M and Varnai P and Villoslada P and Vizcaíno JA and Wubbe JPM and Zanetti G}, title = {Making sense of big data in health research: Towards an EU action plan}, journal = {Genome Med}, volume = {8}, number = {1}, pages = {71}, year = {2016}, doi = {10.1186/s13073-016-0323-y}, abstract = {Medicine and healthcare are undergoing profound changes. Whole-genome sequencing and high-resolution imaging technologies are key drivers of this rapid and crucial transformation. Technological innovation combined with automation and miniaturization has triggered an explosion in data production that will soon reach exabyte proportions. How are we going to deal with this exponential increase in data production? The potential of ``big data'' for improving health is enormous but, at the same time, we face a wide range of challenges to overcome urgently. Europe is very proud of its cultural diversity; however, exploitation of the data made available through advances in genomic medicine, imaging, and a wide range of mobile health applications or connected devices is hampered by numerous historical, technical, legal, and political barriers. European health systems and databases are diverse and fragmented. There is a lack of harmonization of data formats, processing, analysis, and data transfer, which leads to incompatibilities and lost opportunities. Legal frameworks for data sharing are evolving. Clinicians, researchers, and citizens need improved methods, tools, and training to generate, analyze, and query data effectively. Addressing these barriers will contribute to creating the European Single Market for health, which will improve health and healthcare for all Europeans},} - T Rensch, D Villar, J Horvath, DT Odom, P Flicek. Mitochondrial heteroplasmy in vertebrates using ChIP-sequencing data. Genome Biol 2016;17(1):139. doi:10.1186/s13059-016-0996-y
[BibTeX] [Abstract]
\textbf{BACKGROUND:} Mitochondrial heteroplasmy, the presence of more than one mitochondrial DNA (mtDNA) variant in a cell or individual, is not as uncommon as previously thought. It is mostly due to the high mutation rate of the mtDNA and limited repair mechanisms present in the mitochondrion. Motivated by mitochondrial diseases, much focus has been placed into studying this phenomenon in human samples and in medical contexts. To place these results in an evolutionary context and to explore general principles of heteroplasmy, we describe an integrated cross-species evaluation of heteroplasmy in mammals that exploits previously reported NGS data. Focusing on ChIP-seq experiments, we developed a novel approach to detect heteroplasmy from the concomitant mitochondrial DNA fraction sequenced in these experiments.
\textbf{RESULTS:} We first demonstrate that the sequencing coverage of mtDNA in ChIP-seq experiments is sufficient for heteroplasmy detection. We then describe a novel detection method for accurate detection of heteroplasmies, which also accounts for the error rate of NGS technology. Applying this method to 79 individuals from 16 species resulted in 107 heteroplasmic positions present in a total of 45 individuals. Further analysis revealed that the majority of detected heteroplasmies occur in intergenic regions.
\textbf{CONCLUSION:} In addition to documenting the prevalence of mtDNA in ChIP-seq data, the results of our mitochondrial heteroplasmy detection method suggest that mitochondrial heteroplasmies identified across vertebrates share similar characteristics as found for human heteroplasmies. Although largely consistent with previous studies in individual vertebrates, our integrated cross-species analysis provides valuable insights into the evolutionary dynamics of mitochondrial heteroplasmy@Article{27349964, author = {Rensch T and Villar D and Horvath J and Odom DT and Flicek P}, title = {Mitochondrial heteroplasmy in vertebrates using ChIP-sequencing data}, journal = {Genome Biol}, volume = {17}, number = {1}, pages = {139}, year = {2016}, doi = {10.1186/s13059-016-0996-y}, abstract = {\textbf{BACKGROUND:} Mitochondrial heteroplasmy, the presence of more than one mitochondrial DNA (mtDNA) variant in a cell or individual, is not as uncommon as previously thought. It is mostly due to the high mutation rate of the mtDNA and limited repair mechanisms present in the mitochondrion. Motivated by mitochondrial diseases, much focus has been placed into studying this phenomenon in human samples and in medical contexts. To place these results in an evolutionary context and to explore general principles of heteroplasmy, we describe an integrated cross-species evaluation of heteroplasmy in mammals that exploits previously reported NGS data. Focusing on ChIP-seq experiments, we developed a novel approach to detect heteroplasmy from the concomitant mitochondrial DNA fraction sequenced in these experiments.
\textbf{RESULTS:} We first demonstrate that the sequencing coverage of mtDNA in ChIP-seq experiments is sufficient for heteroplasmy detection. We then describe a novel detection method for accurate detection of heteroplasmies, which also accounts for the error rate of NGS technology. Applying this method to 79 individuals from 16 species resulted in 107 heteroplasmic positions present in a total of 45 individuals. Further analysis revealed that the majority of detected heteroplasmies occur in intergenic regions.
\textbf{CONCLUSION:} In addition to documenting the prevalence of mtDNA in ChIP-seq data, the results of our mitochondrial heteroplasmy detection method suggest that mitochondrial heteroplasmies identified across vertebrates share similar characteristics as found for human heteroplasmies. Although largely consistent with previous studies in individual vertebrates, our integrated cross-species analysis provides valuable insights into the evolutionary dynamics of mitochondrial heteroplasmy},} - M Pignatelli, AJ Vilella, M Muffato, L Gordon, S White, P Flicek, J Herrero. ncRNA orthologies in the vertebrate lineage. Database (Oxford) 2016;2016:bav127. doi:10.1093/database/bav127
[BibTeX] [Abstract]
Annotation of orthologous and paralogous genes is necessary for many aspects of evolutionary analysis. Methods to infer these homology relationships have traditionally focused on protein-coding genes and evolutionary models used by these methods normally assume the positions in the protein evolve independently. However, as our appreciation for the roles of non-coding RNA genes has increased, consistently annotated sets of orthologous and paralogous ncRNA genes are increasingly needed. At the same time, methods such as PHASE or RAxML have implemented substitution models that consider pairs of sites to enable proper modelling of the loops and other features of RNA secondary structure. Here, we present a comprehensive analysis pipeline for the automatic detection of orthologues and paralogues for ncRNA genes. We focus on gene families represented in Rfam and for which a specific covariance model is provided. For each family ncRNA genes found in all Ensembl species are aligned using Infernal, and several trees are built using different substitution models. In parallel, a genomic alignment that includes the ncRNA genes and their flanking sequence regions is built with PRANK. This alignment is used to create two additional phylogenetic trees using the neighbour-joining (NJ) and maximum-likelihood (ML) methods. The trees arising from both the ncRNA and genomic alignments are merged using TreeBeST, which reconciles them with the species tree in order to identify speciation and duplication events. The final tree is used to infer the orthologues and paralogues following Fitch's definition. We also determine gene gain and loss events for each family using CAFE. All data are accessible through the Ensembl Comparative Genomics ('Compara') API, on our FTP site and are fully integrated in the Ensembl genome browser, where they can be accessed in a user-friendly manner.Database URL: http://www.ensembl.org
@Article{26980512, author = {Pignatelli M and Vilella AJ and Muffato M and Gordon L and White S and Flicek P and Herrero J}, title = {ncRNA orthologies in the vertebrate lineage}, journal = {Database (Oxford)}, volume = {2016}, pages = {bav127}, year = {2016}, doi = {10.1093/database/bav127}, abstract = {Annotation of orthologous and paralogous genes is necessary for many aspects of evolutionary analysis. Methods to infer these homology relationships have traditionally focused on protein-coding genes and evolutionary models used by these methods normally assume the positions in the protein evolve independently. However, as our appreciation for the roles of non-coding RNA genes has increased, consistently annotated sets of orthologous and paralogous ncRNA genes are increasingly needed. At the same time, methods such as PHASE or RAxML have implemented substitution models that consider pairs of sites to enable proper modelling of the loops and other features of RNA secondary structure. Here, we present a comprehensive analysis pipeline for the automatic detection of orthologues and paralogues for ncRNA genes. We focus on gene families represented in Rfam and for which a specific covariance model is provided. For each family ncRNA genes found in all Ensembl species are aligned using Infernal, and several trees are built using different substitution models. In parallel, a genomic alignment that includes the ncRNA genes and their flanking sequence regions is built with PRANK. This alignment is used to create two additional phylogenetic trees using the neighbour-joining (NJ) and maximum-likelihood (ML) methods. The trees arising from both the ncRNA and genomic alignments are merged using TreeBeST, which reconciles them with the species tree in order to identify speciation and duplication events. The final tree is used to infer the orthologues and paralogues following Fitch's definition. We also determine gene gain and loss events for each family using CAFE. All data are accessible through the Ensembl Comparative Genomics ('Compara') API, on our FTP site and are fully integrated in the Ensembl genome browser, where they can be accessed in a user-friendly manner.Database URL: http://www.ensembl.org},} - GD Poznik, Y Xue, FL Mendez, TF Willems, A Massaia, MA Wilson Sayres, Q Ayub, SA McCarthy, A Narechania, S Kashin, Y Chen, R Banerjee, JL Rodriguez-Flores, M Cerezo, H Shao, M Gymrek, A Malhotra, S Louzada, R Desalle, GRS Ritchie, E Cerveira, TW Fitzgerald, E Garrison, A Marcketta, D Mittelman, M Romanovitch, C Zhang, X Zheng-Bradley, GR Abecasis, SA McCarroll, P Flicek, PA Underhill, L Coin, DR Zerbino, F Yang, C Lee, L Clarke, A Auton, Y Erlich, RE Handsaker, CD Bustamante, C Tyler-Smith. Punctuated bursts in human male demography inferred from 1,244 worldwide Y-chromosome sequences. Nat Genet 2016;48(6):593–599. doi:10.1038/ng.3559
[BibTeX] [Abstract]
We report the sequences of 1,244 human Y chromosomes randomly ascertained from 26 worldwide populations by the 1000 Genomes Project. We discovered more than 65,000 variants, including single-nucleotide variants, multiple-nucleotide variants, insertions and deletions, short tandem repeats, and copy number variants. Of these, copy number variants contribute the greatest predicted functional impact. We constructed a calibrated phylogenetic tree on the basis of binary single-nucleotide variants and projected the more complex variants onto it, estimating the number of mutations for each class. Our phylogeny shows bursts of extreme expansion in male numbers that have occurred independently among each of the five continental superpopulations examined, at times of known migrations and technological innovations
@Article{27111036, author = {Poznik GD and Xue Y and Mendez FL and Willems TF and Massaia A and Wilson Sayres MA and Ayub Q and McCarthy SA and Narechania A and Kashin S and Chen Y and Banerjee R and Rodriguez-Flores JL and Cerezo M and Shao H and Gymrek M and Malhotra A and Louzada S and Desalle R and Ritchie GRS and Cerveira E and Fitzgerald TW and Garrison E and Marcketta A and Mittelman D and Romanovitch M and Zhang C and Zheng-Bradley X and Abecasis GR and McCarroll SA and Flicek P and Underhill PA and Coin L and Zerbino DR and Yang F and Lee C and Clarke L and Auton A and Erlich Y and Handsaker RE and Bustamante CD and Tyler-Smith C}, title = {Punctuated bursts in human male demography inferred from 1,244 worldwide Y-chromosome sequences}, journal = {Nat Genet}, volume = {48}, number = {6}, pages = {593--599}, year = {2016}, doi = {10.1038/ng.3559}, howpublished = {Advanced online publication: 25 April 2016}, abstract = {We report the sequences of 1,244 human Y chromosomes randomly ascertained from 26 worldwide populations by the 1000 Genomes Project. We discovered more than 65,000 variants, including single-nucleotide variants, multiple-nucleotide variants, insertions and deletions, short tandem repeats, and copy number variants. Of these, copy number variants contribute the greatest predicted functional impact. We constructed a calibrated phylogenetic tree on the basis of binary single-nucleotide variants and projected the more complex variants onto it, estimating the number of mutations for each class. Our phylogeny shows bursts of extreme expansion in male numbers that have occurred independently among each of the five continental superpopulations examined, at times of known migrations and technological innovations},} - JM Fernández, de la V Torre, D Richardson, R Royo, M Puiggròs, V Moncunill, S Fragkogianni, L Clarke, BLUEPRINT Consortium, P Flicek, D Rico, D Torrents, de E Carrillo Santa Pau, A Valencia. The BLUEPRINT Data Analysis Portal. Cell Syst 2016;3(5):491–495.e5. doi:10.1016/j.cels.2016.10.021
[BibTeX] [Abstract]
The impact of large and complex epigenomic datasets on biological insights or clinical applications is limited by the lack of accessibility by easy, intuitive, and fast tools. Here, we describe an epigenomics comparative cyber-infrastructure (EPICO), an open-access reference set of libraries to develop comparative epigenomic data portals. Using EPICO, large epigenome projects can make available their rich datasets to the community without requiring specific technical skills. As a first instance of EPICO, we implemented the BLUEPRINT Data Analysis Portal (BDAP). BDAP provides a desktop for the comparative analysis of epigenomes of hematopoietic cell types based on results, such as the position of epigenetic features, from basic analysis pipelines. The BDAP interface facilitates interactive exploration of genomic regions, genes, and pathways in the context of differentiation of hematopoietic lineages. This work represents initial steps toward broadly accessible integrative analysis of epigenomic data across international consortia. EPICO can be accessed at https://github.com/inab, and BDAP can be accessed at http://blueprint-data.bsc.es.
@Article{27863955, author = {Fernández JM and de la Torre V and Richardson D and Royo R and Puiggròs M and Moncunill V and Fragkogianni S and Clarke L and {BLUEPRINT Consortium} and Flicek P and Rico D and Torrents D and Carrillo de Santa Pau E and Valencia A}, title = {The BLUEPRINT Data Analysis Portal}, journal = {Cell Syst}, volume = {3}, number = {5}, pages = {491--495.e5}, year = {2016}, doi = {10.1016/j.cels.2016.10.021}, howpublished = {Advanced online publication: 15 November 2016}, abstract = {The impact of large and complex epigenomic datasets on biological insights or clinical applications is limited by the lack of accessibility by easy, intuitive, and fast tools. Here, we describe an epigenomics comparative cyber-infrastructure (EPICO), an open-access reference set of libraries to develop comparative epigenomic data portals. Using EPICO, large epigenome projects can make available their rich datasets to the community without requiring specific technical skills. As a first instance of EPICO, we implemented the BLUEPRINT Data Analysis Portal (BDAP). BDAP provides a desktop for the comparative analysis of epigenomes of hematopoietic cell types based on results, such as the position of epigenetic features, from basic analysis pipelines. The BDAP interface facilitates interactive exploration of genomic regions, genes, and pathways in the context of differentiation of hematopoietic lineages. This work represents initial steps toward broadly accessible integrative analysis of epigenomic data across international consortia. EPICO can be accessed at https://github.com/inab, and BDAP can be accessed at http://blueprint-data.bsc.es.},} - BL Aken, S Ayling, D Barrell, L Clarke, V Curwen, S Fairley, J Fernandez Banet, K Billis, C García Girón, T Hourlier, K Howe, A Kähäri, F Kokocinski, FJ Martin, DN Murphy, R Nag, M Ruffier, M Schuster, YA Tang, J-H Vogel, S White, A Zadissa, P Flicek, SMJ Searle. The Ensembl gene annotation system. Database (Oxford) 2016;2016:baw093. doi:10.1093/database/baw093
[BibTeX] [Abstract]
The Ensembl gene annotation system has been used to annotate over 70 different vertebrate species across a wide range of genome projects. Furthermore, it generates the automatic alignment-based annotation for the human and mouse GENCODE gene sets. The system is based on the alignment of biological sequences, including cDNAs, proteins and RNA-seq reads, to the target genome in order to construct candidate transcript models. Careful assessment and filtering of these candidate transcripts ultimately leads to the final gene set, which is made available on the Ensembl website. Here, we describe the annotation process in detail.Database URL: http://www.ensembl.org/index.html
@Article{27337980, author = {Aken BL and Ayling S and Barrell D and Clarke L and Curwen V and Fairley S and Fernandez Banet J and Billis K and García Girón C and Hourlier T and Howe K and Kähäri A and Kokocinski F and Martin FJ and Murphy DN and Nag R and Ruffier M and Schuster M and Tang YA and Vogel J-H and White S and Zadissa A and Flicek P and Searle SMJ}, title = {The Ensembl gene annotation system}, journal = {Database (Oxford)}, volume = {2016}, pages = {baw093}, year = {2016}, doi = {10.1093/database/baw093}, abstract = {The Ensembl gene annotation system has been used to annotate over 70 different vertebrate species across a wide range of genome projects. Furthermore, it generates the automatic alignment-based annotation for the human and mouse GENCODE gene sets. The system is based on the alignment of biological sequences, including cDNAs, proteins and RNA-seq reads, to the target genome in order to construct candidate transcript models. Careful assessment and filtering of these candidate transcripts ultimately leads to the final gene set, which is made available on the Ensembl website. Here, we describe the annotation process in detail.Database URL: http://www.ensembl.org/index.html},} - W McLaren, L Gil, SE Hunt, HS Riat, GRS Ritchie, A Thormann, P Flicek, F Cunningham. The Ensembl Variant Effect Predictor. Genome Biol 2016;17(1):122. doi:10.1186/s13059-016-0974-4
[BibTeX] [Abstract]
The Ensembl Variant Effect Predictor is a powerful toolset for the analysis, annotation, and prioritization of genomic variants in coding and non-coding regions. It provides access to an extensive collection of genomic annotation, with a variety of interfaces to suit different requirements, and simple options for configuring and extending analysis. It is open source, free to use, and supports full reproducibility of results. The Ensembl Variant Effect Predictor can simplify and accelerate variant interpretation in a wide range of study designs
@Article{27268795, author = {McLaren W and Gil L and Hunt SE and Riat HS and Ritchie GRS and Thormann A and Flicek P and Cunningham F}, title = {The Ensembl Variant Effect Predictor}, journal = {Genome Biol}, volume = {17}, number = {1}, pages = {122}, year = {2016}, doi = {10.1186/s13059-016-0974-4}, note = {First posted as a preprint: 4 March 2016}, abstract = {The Ensembl Variant Effect Predictor is a powerful toolset for the analysis, annotation, and prioritization of genomic variants in coding and non-coding regions. It provides access to an extensive collection of genomic annotation, with a variety of interfaces to suit different requirements, and simple options for configuring and extending analysis. It is open source, free to use, and supports full reproducibility of results. The Ensembl Variant Effect Predictor can simplify and accelerate variant interpretation in a wide range of study designs},} - AP Morgan, JM Holt, RC McMullan, TA Bell, AM-F Clayshulte, JP Didion, L Yadgary, D Thybert, DT Odom, P Flicek, L McMillan, de F Pardo-Manuel Villena. The Evolutionary Fates of a Large Segmental Duplication in Mouse. Genetics 2016;204(1):267–285. doi:10.1534/genetics.116.191007
[BibTeX] [Abstract]
Gene duplication and loss are major sources of genetic polymorphism in populations, and are important forces shaping the evolution of genome content and organization. We have reconstructed the origin and history of a 127 kbp segmental duplication, R2d, in the house mouse (Mus musculus). R2d contains a single protein-coding gene, Cwc22 De novo assembly of both the ancestral (R2d1) and the derived (R2d2) copies reveals that they have been subject to non-allelic gene conversion events spanning tens of kilobases. R2d2 is also a hotspot for structural variation: its diploid copy number ranges from zero in the mouse reference genome to more than 80 in wild mice sampled from around the globe. Hemizgyosity for high-copy-number alleles of R2d2 is associated in cis with meiotic drive, suppression of meiotic crossovers, and copy-number instability, with a mutation rate in excess of 1 per 100 transmissions in some laboratory populations. Our results provide a striking example of allelic diversity generated by duplication and demonstrate the value of de novo assembly in a phylogenetic context for understanding the mutational processes affecting duplicate genes
@Article{27371833, author = {Morgan AP and Holt JM and McMullan RC and Bell TA and Clayshulte AM-F and Didion JP and Yadgary L and Thybert D and Odom DT and Flicek P and McMillan L and Pardo-Manuel de Villena F}, title = {The Evolutionary Fates of a Large Segmental Duplication in Mouse}, journal = {Genetics}, volume = {204}, number = {1}, pages = {267--285}, year = {2016}, doi = {10.1534/genetics.116.191007}, howpublished = {Advanced online publication: 30 June 2016}, note = {First posted as a preprint: 29 April 2016}, abstract = {Gene duplication and loss are major sources of genetic polymorphism in populations, and are important forces shaping the evolution of genome content and organization. We have reconstructed the origin and history of a 127 kbp segmental duplication, R2d, in the house mouse (Mus musculus). R2d contains a single protein-coding gene, Cwc22 De novo assembly of both the ancestral (R2d1) and the derived (R2d2) copies reveals that they have been subject to non-allelic gene conversion events spanning tens of kilobases. R2d2 is also a hotspot for structural variation: its diploid copy number ranges from zero in the mouse reference genome to more than 80 in wild mice sampled from around the globe. Hemizgyosity for high-copy-number alleles of R2d2 is associated in cis with meiotic drive, suppression of meiotic crossovers, and copy-number instability, with a mutation rate in excess of 1 per 100 transmissions in some laboratory populations. Our results provide a striking example of allelic diversity generated by duplication and demonstrate the value of de novo assembly in a phylogenetic context for understanding the mutational processes affecting duplicate genes},} - HG Stunnenberg, International Human Epigenome Consortium, M Hirst. The International Human Epigenome Consortium: A Blueprint for Scientific Collaboration and Discovery. Cell 2016;167(5):1145–1149. doi:10.1016/j.cell.2016.11.007
[BibTeX] [Abstract]
The International Human Epigenome Consortium (IHEC) coordinates the generation of a catalog of high-resolution reference epigenomes of major primary human cell types. The studies now presented (see the Cell Press IHEC web portal at http://www.cell.com/consortium/IHEC) highlight the coordinated achievements of IHEC teams to gather and interpret comprehensive epigenomic datasets to gain insights in the epigenetic control of cell states relevant for human health and disease. PAPERCLIP.
@Article{27863232, author = {Stunnenberg HG and {International Human Epigenome Consortium} and Hirst M}, title = {The International Human Epigenome Consortium: A Blueprint for Scientific Collaboration and Discovery}, journal = {Cell}, volume = {167}, number = {5}, pages = {1145--1149}, year = {2016}, doi = {10.1016/j.cell.2016.11.007}, abstract = {The International Human Epigenome Consortium (IHEC) coordinates the generation of a catalog of high-resolution reference epigenomes of major primary human cell types. The studies now presented (see the Cell Press IHEC web portal at http://www.cell.com/consortium/IHEC) highlight the coordinated achievements of IHEC teams to gather and interpret comprehensive epigenomic datasets to gain insights in the epigenetic control of cell states relevant for human health and disease. PAPERCLIP.},} - B Novakovic, E Habibi, SY Wang, RJ Arts, R Davar, W Megchelenbrink, B Kim, T Kuznetsova, M Kox, J Zwaag, F Matarese, van SJ Heeringen, EM Janssen-Megens, N Sharifi, C Wang, F Keramati, V Schoonenberg, P Flicek, L Clarke, P Pickkers, S Heath, I Gut, MG Netea, JH Martens, C Logie, HG Stunnenberg. β-Glucan Reverses the Epigenetic State of LPS-Induced Immunological Tolerance. Cell 2016;167(5):1354–1368.e14. doi:10.1016/j.cell.2016.09.034
[BibTeX] [Abstract]
Innate immune memory is the phenomenon whereby innate immune cells such as monocytes or macrophages undergo functional reprogramming after exposure to microbial components such as lipopolysaccharide (LPS). We apply an integrated epigenomic approach to characterize the molecular events involved in LPS-induced tolerance in a time-dependent manner. Mechanistically, LPS-treated monocytes fail to accumulate active histone marks at promoter and enhancers of genes in the lipid metabolism and phagocytic pathways. Transcriptional inactivity in response to a second LPS exposure in tolerized macrophages is accompanied by failure to deposit active histone marks at promoters of tolerized genes. In contrast, β-glucan partially reverses the LPS-induced tolerance in vitro. Importantly, ex vivo β-glucan treatment of monocytes from volunteers with experimental endotoxemia re-instates their capacity for cytokine production. Tolerance is reversed at the level of distal element histone modification and transcriptional reactivation of otherwise unresponsive genes. VIDEO ABSTRACT.
@Article{27863248, author = {Novakovic B and Habibi E and Wang SY and Arts RJ and Davar R and Megchelenbrink W and Kim B and Kuznetsova T and Kox M and Zwaag J and Matarese F and van Heeringen SJ and Janssen-Megens EM and Sharifi N and Wang C and Keramati F and Schoonenberg V and Flicek P and Clarke L and Pickkers P and Heath S and Gut I and Netea MG and Martens JH and Logie C and Stunnenberg HG}, title = {β-Glucan Reverses the Epigenetic State of LPS-Induced Immunological Tolerance}, journal = {Cell}, volume = {167}, number = {5}, pages = {1354--1368.e14}, year = {2016}, doi = {10.1016/j.cell.2016.09.034}, abstract = {Innate immune memory is the phenomenon whereby innate immune cells such as monocytes or macrophages undergo functional reprogramming after exposure to microbial components such as lipopolysaccharide (LPS). We apply an integrated epigenomic approach to characterize the molecular events involved in LPS-induced tolerance in a time-dependent manner. Mechanistically, LPS-treated monocytes fail to accumulate active histone marks at promoter and enhancers of genes in the lipid metabolism and phagocytic pathways. Transcriptional inactivity in response to a second LPS exposure in tolerized macrophages is accompanied by failure to deposit active histone marks at promoters of tolerized genes. In contrast, β-glucan partially reverses the LPS-induced tolerance in vitro. Importantly, ex vivo β-glucan treatment of monocytes from volunteers with experimental endotoxemia re-instates their capacity for cytokine production. Tolerance is reversed at the level of distal element histone modification and transcriptional reactivation of otherwise unresponsive genes. VIDEO ABSTRACT.},}
2015
- The 1000 Genomes Project Consortium.. A global reference for human genetic variation. Nature 2015;526(7571):68–74. doi:10.1038/nature15393
[BibTeX] [Abstract]
The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations. Here we report completion of the project, having reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. We characterized a broad spectrum of genetic variation, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes >99\% of SNP variants with a frequency of >1\% for a variety of ancestries. We describe the distribution of genetic variation across the global sample, and discuss the implications for common disease studies
@Article{26432245, author = {{The 1000 Genomes Project Consortium.}}, title = {A global reference for human genetic variation}, journal = {Nature}, volume = {526}, number = {7571}, pages = {68--74}, year = {2015}, doi = {10.1038/nature15393}, abstract = {The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations. Here we report completion of the project, having reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-genome sequencing, deep exome sequencing, and dense microarray genotyping. We characterized a broad spectrum of genetic variation, in total over 88 million variants (84.7 million single nucleotide polymorphisms (SNPs), 3.6 million short insertions/deletions (indels), and 60,000 structural variants), all phased onto high-quality haplotypes. This resource includes \>99\% of SNP variants with a frequency of \>1\% for a variety of ancestries. We describe the distribution of genetic variation across the global sample, and discuss the implications for common disease studies},} - PH Sudmant, T Rausch, EJ Gardner, RE Handsaker, A Abyzov, J Huddleston, Y Zhang, K Ye, G Jun, M Hsi-Yang Fritz, MK Konkel, A Malhotra, AM Stütz, X Shi, F Paolo Casale, J Chen, F Hormozdiari, G Dayama, K Chen, M Malig, MJP Chaisson, K Walter, S Meiers, S Kashin, E Garrison, A Auton, HYK Lam, X Jasmine Mu, C Alkan, D Antaki, T Bae, E Cerveira, P Chines, Z Chong, L Clarke, E Dal, L Ding, S Emery, X Fan, M Gujral, F Kahveci, JM Kidd, Y Kong, E-W Lameijer, S McCarthy, P Flicek, RA Gibbs, G Marth, CE Mason, A Menelaou, DM Muzny, BJ Nelson, A Noor, NF Parrish, M Pendleton, A Quitadamo, B Raeder, EE Schadt, M Romanovitch, A Schlattl, R Sebra, AA Shabalin, A Untergasser, JA Walker, M Wang, F Yu, C Zhang, J Zhang, X Zheng-Bradley, W Zhou, T Zichner, J Sebat, MA Batzer, SA McCarroll, RE Mills, MB Gerstein, A Bashir, O Stegle, SE Devine, C Lee, EE Eichler, JO Korbel. An integrated map of structural variation in 2,504 human genomes. Nature 2015;526(7571):75–81. doi:10.1038/nature15394
[BibTeX] [Abstract]
Structural variants are implicated in numerous diseases and make up the majority of varying nucleotides among human genomes. Here we describe an integrated set of eight structural variant classes comprising both balanced and unbalanced variants, which we constructed using short-read DNA sequencing data and statistically phased onto haplotype blocks in 26 human populations. Analysing this set, we identify numerous gene-intersecting structural variants exhibiting population stratification and describe naturally occurring homozygous gene knockouts that suggest the dispensability of a variety of human genes. We demonstrate that structural variants are enriched on haplotypes identified by genome-wide association studies and exhibit enrichment for expression quantitative trait loci. Additionally, we uncover appreciable levels of structural variant complexity at different scales, including genic loci subject to clusters of repeated rearrangement and complex structural variants with multiple breakpoints likely to have formed through individual mutational events. Our catalogue will enhance future studies into structural variant demography, functional impact and disease association
@Article{26432246, author = {Sudmant PH and Rausch T and Gardner EJ and Handsaker RE and Abyzov A and Huddleston J and Zhang Y and Ye K and Jun G and Hsi-Yang Fritz M and Konkel MK and Malhotra A and Stütz AM and Shi X and Paolo Casale F and Chen J and Hormozdiari F and Dayama G and Chen K and Malig M and Chaisson MJP and Walter K and Meiers S and Kashin S and Garrison E and Auton A and Lam HYK and Jasmine Mu X and Alkan C and Antaki D and Bae T and Cerveira E and Chines P and Chong Z and Clarke L and Dal E and Ding L and Emery S and Fan X and Gujral M and Kahveci F and Kidd JM and Kong Y and Lameijer E-W and McCarthy S and Flicek P and Gibbs RA and Marth G and Mason CE and Menelaou A and Muzny DM and Nelson BJ and Noor A and Parrish NF and Pendleton M and Quitadamo A and Raeder B and Schadt EE and Romanovitch M and Schlattl A and Sebra R and Shabalin AA and Untergasser A and Walker JA and Wang M and Yu F and Zhang C and Zhang J and Zheng-Bradley X and Zhou W and Zichner T and Sebat J and Batzer MA and McCarroll SA and Mills RE and Gerstein MB and Bashir A and Stegle O and Devine SE and Lee C and Eichler EE and Korbel JO}, title = {An integrated map of structural variation in 2,504 human genomes}, journal = {Nature}, volume = {526}, number = {7571}, pages = {75--81}, year = {2015}, doi = {10.1038/nature15394}, abstract = {Structural variants are implicated in numerous diseases and make up the majority of varying nucleotides among human genomes. Here we describe an integrated set of eight structural variant classes comprising both balanced and unbalanced variants, which we constructed using short-read DNA sequencing data and statistically phased onto haplotype blocks in 26 human populations. Analysing this set, we identify numerous gene-intersecting structural variants exhibiting population stratification and describe naturally occurring homozygous gene knockouts that suggest the dispensability of a variety of human genes. We demonstrate that structural variants are enriched on haplotypes identified by genome-wide association studies and exhibit enrichment for expression quantitative trait loci. Additionally, we uncover appreciable levels of structural variant complexity at different scales, including genic loci subject to clusters of repeated rearrangement and complex structural variants with multiple breakpoints likely to have formed through individual mutational events. Our catalogue will enhance future studies into structural variant demography, functional impact and disease association},} - A Raposo, F Vasconcelos, D Drechsel, C Marie, C Johnston, D Dolle, A Bithell, S Gillotin, van den Berg D, L Ettwiller, P Flicek, G Crawford, C Parras, B Berninger, N Buckley, F Guillemot, D Castro. Ascl1 Coordinately Regulates Gene Expression and the Chromatin Landscape during Neurogenesis. Cell Reports 2015;10(9):1544–1556. doi:10.1016/j.celrep.2015.02.025
[BibTeX]@Article{25753420, author = {Raposo A and Vasconcelos F and Drechsel D and Marie C and Johnston C and Dolle D and Bithell A and Gillotin S and van den Berg D and Ettwiller L and Flicek P and Crawford G and Parras C and Berninger B and Buckley N and Guillemot F and Castro D}, title = {Ascl1 Coordinately Regulates Gene Expression and the Chromatin Landscape during Neurogenesis}, journal = {Cell Reports}, volume = {10}, number = {9}, pages = {1544--1556}, year = {2015}, doi = {10.1016/j.celrep.2015.02.025}, } - L Eöry, MTP Gilbert, C Li, B Li, A Archibald, BL Aken, G Zhang, E Jarvis, P Flicek, DW Burt. Avianbase: a community resource for bird genomics. Genome Biology 2015;16(1):21. doi:10.1186/s13059-015-0588-2
[BibTeX] [Abstract]
Giving access to sequence and annotation data for genome assemblies is important because, while facilitating research, it places both assembly and annotation quality under scrutiny, resulting in improvements to both. Therefore we announce Avianbase, a resource for bird genomics, which provides access to data released by the Avian Phylogenomics Consortium.
@Article{25723810, author = {Eöry L and Gilbert MTP and Li C and Li B and Archibald A and Aken BL and Zhang G and Jarvis E and Flicek P and Burt DW}, title = {Avianbase: a community resource for bird genomics}, journal = {Genome Biology}, volume = {16}, number = {1}, pages = {21}, year = {2015}, doi = {10.1186/s13059-015-0588-2}, abstract = {Giving access to sequence and annotation data for genome assemblies is important because, while facilitating research, it places both assembly and annotation quality under scrutiny, resulting in improvements to both. Therefore we announce Avianbase, a resource for bird genomics, which provides access to data released by the Avian Phylogenomics Consortium.},} - JL Mateo, van den DLC Berg, M Haeussler, D Drechsel, ZB Gaber, DS Castro, P Robson, GE Crawford, P Flicek, L Ettwiller, J Wittbrodt, F Guillemot, B Martynoga. Characterization of the neural stem cell gene regulatory network identifies OLIG2 as a multi-functional regulator of self-renewal. Genome Res 2015;25(1):41–56. doi:10.1101/gr.173435.114
[BibTeX] [Abstract]
The gene regulatory network (GRN) that supports neural stem cell (NS cell) self-renewal has so far been poorly characterised. Knowledge of the central transcription factors (TFs), the non-coding gene regulatory regions that they bind to and the genes whose expression they modulate will be crucial in unlocking the full therapeutic potential of these cells. Here, we use DNase-seq in combination with analysis of histone modifications to identify multiple classes of epigenetically and functionally distinct cis-regulatory elements (CREs). Through motif analysis and ChIP-seq we identify several of the crucial TF regulators of NS cells. At the core of the network are TFs of the basic helix-loop-helix (bHLH), nuclear factor I (NFI), SOX and FOX families, with CREs often densely bound by several of these different TFs. We use machine learning to highlight several crucial regulatory features of the network that underpin NS cell self-renewal and multipotency. We validate our predictions by functional analysis of the bHLH TF OLIG2. This TF makes an important contribution to NS cell self-renewal by concurrently activating pro-proliferation genes and preventing the untimely activation of genes promoting neuronal differentiation and stem cell quiescence
@Article{25294244, author = {Mateo JL and van den Berg DLC and Haeussler M and Drechsel D and Gaber ZB and Castro DS and Robson P and Crawford GE and Flicek P and Ettwiller L and Wittbrodt J and Guillemot F and Martynoga B}, title = {Characterization of the neural stem cell gene regulatory network identifies OLIG2 as a multi-functional regulator of self-renewal}, journal = {Genome Res}, volume = {25}, number = {1}, pages = {41--56}, year = {2015}, doi = {10.1101/gr.173435.114}, howpublished = {Advanced online publication: 7 October 2014}, abstract = {The gene regulatory network (GRN) that supports neural stem cell (NS cell) self-renewal has so far been poorly characterised. Knowledge of the central transcription factors (TFs), the non-coding gene regulatory regions that they bind to and the genes whose expression they modulate will be crucial in unlocking the full therapeutic potential of these cells. Here, we use DNase-seq in combination with analysis of histone modifications to identify multiple classes of epigenetically and functionally distinct cis-regulatory elements (CREs). Through motif analysis and ChIP-seq we identify several of the crucial TF regulators of NS cells. At the core of the network are TFs of the basic helix-loop-helix (bHLH), nuclear factor I (NFI), SOX and FOX families, with CREs often densely bound by several of these different TFs. We use machine learning to highlight several crucial regulatory features of the network that underpin NS cell self-renewal and multipotency. We validate our predictions by functional analysis of the bHLH TF OLIG2. This TF makes an important contribution to NS cell self-renewal by concurrently activating pro-proliferation genes and preventing the untimely activation of genes promoting neuronal differentiation and stem cell quiescence},} - B Benjelloun, FJ Alberto, I Streeter, F Boyer, E Coissac, S Stucki, M BenBati, M Ibnelbachyr, M Chentouf, A Bechchari, K Leempoel, A Alberti, S Engelen, A Chikhi, L Clarke, P Flicek, S Joost, P Taberlet, F Pompanon, N Consortium. Characterizing neutral genomic diversity and selection signatures in indigenous populations of Moroccan goats (Capra hircus) using WGS data. Front Genet 2015;6:107. doi:10.3389/fgene.2015.00107
[BibTeX] [Abstract]
Since the time of their domestication, goats (Capra hircus) have evolved in a large variety of locally adapted populations in response to different human and environmental pressures. In the present era, many indigenous populations are threatened with extinction due to their substitution by cosmopolitan breeds, while they might represent highly valuable genomic resources. It is thus crucial to characterize the neutral and adaptive genetic diversity of indigenous populations. A fine characterization of whole genome variation in farm animals is now possible by using new sequencing technologies. We sequenced the complete genome at 12× coverage of 44 goats geographically representative of the three phenotypically distinct indigenous populations in Morocco. The study of mitochondrial genomes showed a high diversity exclusively restricted to the haplogroup A. The 44 nuclear genomes showed a very high diversity (24 million variants) associated with low linkage disequilibrium. The overall genetic diversity was weakly structured according to geography and phenotypes. When looking for signals of positive selection in each population we identified many candidate genes, several of which gave insights into the metabolic pathways or biological processes involved in the adaptation to local conditions (e.g., panting in warm/desert conditions). This study highlights the interest of WGS data to characterize livestock genomic diversity. It illustrates the valuable genetic richness present in indigenous populations that have to be sustainably managed and may represent valuable genetic resources for the long-term preservation of the species
@Article{25904931, author = {Benjelloun B and Alberto FJ and Streeter I and Boyer F and Coissac E and Stucki S and BenBati M and Ibnelbachyr M and Chentouf M and Bechchari A and Leempoel K and Alberti A and Engelen S and Chikhi A and Clarke L and Flicek P and Joost S and Taberlet P and Pompanon F and Consortium N}, title = {Characterizing neutral genomic diversity and selection signatures in indigenous populations of Moroccan goats (Capra hircus) using WGS data}, journal = {Front Genet}, volume = {6}, pages = {107}, year = {2015}, doi = {10.3389/fgene.2015.00107}, abstract = {Since the time of their domestication, goats (Capra hircus) have evolved in a large variety of locally adapted populations in response to different human and environmental pressures. In the present era, many indigenous populations are threatened with extinction due to their substitution by cosmopolitan breeds, while they might represent highly valuable genomic resources. It is thus crucial to characterize the neutral and adaptive genetic diversity of indigenous populations. A fine characterization of whole genome variation in farm animals is now possible by using new sequencing technologies. We sequenced the complete genome at 12× coverage of 44 goats geographically representative of the three phenotypically distinct indigenous populations in Morocco. The study of mitochondrial genomes showed a high diversity exclusively restricted to the haplogroup A. The 44 nuclear genomes showed a very high diversity (24 million variants) associated with low linkage disequilibrium. The overall genetic diversity was weakly structured according to geography and phenotypes. When looking for signals of positive selection in each population we identified many candidate genes, several of which gave insights into the metabolic pathways or biological processes involved in the adaptation to local conditions (e.g., panting in warm/desert conditions). This study highlights the interest of WGS data to characterize livestock genomic diversity. It illustrates the valuable genetic richness present in indigenous populations that have to be sustainably managed and may represent valuable genetic resources for the long-term preservation of the species},} - ES Wong, D Thybert, BM Schmitt, K Stefflova, DT Odom, P Flicek. Decoupling of evolutionary changes in transcription factor binding and gene expression in mammals. Genome Res 2015;25(2):167–178. doi:10.1101/gr.177840.114
[BibTeX] [Abstract]
To understand the evolutionary dynamics between transcription factor (TF) binding and gene expression in mammals, we compared transcriptional output and the binding intensities for three tissue-specific TFs in livers from four closely related mouse species. For each transcription factor, TF dependent genes and the TF binding sites most likely to influence mRNA expression were identified by comparing mRNA expression levels between wildtype and TF knockout mice. Independent evolution was observed genome-wide between the rate of change in TF binding and the rate of change in mRNA expression across taxa, with the exception of a small number of TF dependent genes. We also found that binding intensities are preferentially conserved near genes whose expression is dependent on the TF, and the conservation is shared among binding peaks in close proximity to each other near the TSS. Expression of TF dependent genes typically showed an increased sensitivity to changes in binding levels, as measured by mRNA abundance. Taken together, these results highlight a significant tolerance to evolutionary changes in TF binding intensity in mammalian transcriptional networks, and suggest that some TF dependent genes may be largely regulated by a single TF across evolution
@Article{25394363, author = {Wong ES and Thybert D and Schmitt BM and Stefflova K and Odom DT and Flicek P}, title = {Decoupling of evolutionary changes in transcription factor binding and gene expression in mammals}, journal = {Genome Res}, volume = {25}, number = {2}, pages = {167--178}, year = {2015}, doi = {10.1101/gr.177840.114}, howpublished = {Advanced online publication: 13 November 2014}, abstract = {To understand the evolutionary dynamics between transcription factor (TF) binding and gene expression in mammals, we compared transcriptional output and the binding intensities for three tissue-specific TFs in livers from four closely related mouse species. For each transcription factor, TF dependent genes and the TF binding sites most likely to influence mRNA expression were identified by comparing mRNA expression levels between wildtype and TF knockout mice. Independent evolution was observed genome-wide between the rate of change in TF binding and the rate of change in mRNA expression across taxa, with the exception of a small number of TF dependent genes. We also found that binding intensities are preferentially conserved near genes whose expression is dependent on the TF, and the conservation is shared among binding peaks in close proximity to each other near the TSS. Expression of TF dependent genes typically showed an increased sensitivity to changes in binding levels, as measured by mRNA abundance. Taken together, these results highlight a significant tolerance to evolutionary changes in TF binding intensity in mammalian transcriptional networks, and suggest that some TF dependent genes may be largely regulated by a single TF across evolution},} - H Kretzmer, SH Bernhart, W Wang, A Haake, MA Weniger, AK Bergmann, MJ Betts, E Carrillo-de-Santa-Pau, G Doose, J Gutwein, J Richter, V Hovestadt, B Huang, D Rico, F Jühling, J Kolarova, Q Lu, C Otto, R Wagener, J Arnolds, B Burkhardt, A Claviez, HG Drexler, S Eberth, R Eils, P Flicek, S Haas, M Hummel, D Karsch, HHD Kerstens, W Klapper, M Kreuz, C Lawerenz, D Lenze, M Loeffler, C López, RAF MacLeod, JHA Martens, M Kulis, JI Martín-Subero, P Möller, I Nagel, S Picelli, I Vater, M Rohde, P Rosenstiel, M Rosolowski, RB Russell, M Schilhabel, M Schlesner, PF Stadler, M Szczepanowski, L Trümper, HG Stunnenberg, R Küppers, O Ammerpohl, P Lichter, R Siebert, S Hoffmann, B Radlwimmer. DNA methylome analysis in Burkitt and follicular lymphomas identifies differentially methylated regions linked to somatic mutation and transcriptional control. Nat Genet 2015;47(11):1316–1325. doi:10.1038/ng.3413
[BibTeX] [Abstract]
Although Burkitt lymphomas and follicular lymphomas both have features of germinal center B cells, they are biologically and clinically quite distinct. Here we performed whole-genome bisulfite, genome and transcriptome sequencing in 13 IG-MYC translocation-positive Burkitt lymphoma, nine BCL2 translocation-positive follicular lymphoma and four normal germinal center B cell samples. Comparison of Burkitt and follicular lymphoma samples showed differential methylation of intragenic regions that strongly correlated with expression of associated genes, for example, genes active in germinal center dark-zone and light-zone B cells. Integrative pathway analyses of regions differentially methylated in Burkitt and follicular lymphomas implicated DNA methylation as cooperating with somatic mutation of sphingosine phosphate signaling, as well as the TCF3-ID3 and SWI/SNF complexes, in a large fraction of Burkitt lymphomas. Taken together, our results demonstrate a tight connection between somatic mutation, DNA methylation and transcriptional control in key B cell pathways deregulated differentially in Burkitt lymphoma and other germinal center B cell lymphomas
@Article{26437030, author = {Kretzmer H and Bernhart SH and Wang W and Haake A and Weniger MA and Bergmann AK and Betts MJ and Carrillo-de-Santa-Pau E and Doose G and Gutwein J and Richter J and Hovestadt V and Huang B and Rico D and Jühling F and Kolarova J and Lu Q and Otto C and Wagener R and Arnolds J and Burkhardt B and Claviez A and Drexler HG and Eberth S and Eils R and Flicek P and Haas S and Hummel M and Karsch D and Kerstens HHD and Klapper W and Kreuz M and Lawerenz C and Lenze D and Loeffler M and López C and MacLeod RAF and Martens JHA and Kulis M and Martín-Subero JI and Möller P and Nagel I and Picelli S and Vater I and Rohde M and Rosenstiel P and Rosolowski M and Russell RB and Schilhabel M and Schlesner M and Stadler PF and Szczepanowski M and Trümper L and Stunnenberg HG and Küppers R and Ammerpohl O and Lichter P and Siebert R and Hoffmann S and Radlwimmer B}, title = {DNA methylome analysis in Burkitt and follicular lymphomas identifies differentially methylated regions linked to somatic mutation and transcriptional control}, journal = {Nat Genet}, volume = {47}, number = {11}, pages = {1316--1325}, year = {2015}, doi = {10.1038/ng.3413}, abstract = {Although Burkitt lymphomas and follicular lymphomas both have features of germinal center B cells, they are biologically and clinically quite distinct. Here we performed whole-genome bisulfite, genome and transcriptome sequencing in 13 IG-MYC translocation-positive Burkitt lymphoma, nine BCL2 translocation-positive follicular lymphoma and four normal germinal center B cell samples. Comparison of Burkitt and follicular lymphoma samples showed differential methylation of intragenic regions that strongly correlated with expression of associated genes, for example, genes active in germinal center dark-zone and light-zone B cells. Integrative pathway analyses of regions differentially methylated in Burkitt and follicular lymphomas implicated DNA methylation as cooperating with somatic mutation of sphingosine phosphate signaling, as well as the TCF3-ID3 and SWI/SNF complexes, in a large fraction of Burkitt lymphomas. Taken together, our results demonstrate a tight connection between somatic mutation, DNA methylation and transcriptional control in key B cell pathways deregulated differentially in Burkitt lymphoma and other germinal center B cell lymphomas},} - D Villar, C Berthelot, S Aldridge, TF Rayner, M Lukk, M Pignatelli, TJ Park, R Deaville, JT Erichsen, AJ Jasinska, JMA Turner, MF Bertelsen, EP Murchison, P Flicek, DT Odom. Enhancer Evolution across 20 Mammalian Species. Cell 2015;160(3):554–566. doi:10.1016/j.cell.2015.01.006
[BibTeX] [Abstract]
The mammalian radiation has corresponded with rapid changes in noncoding regions of the genome, but we lack a comprehensive understanding of regulatory evolution in mammals. Here, we track the evolution of promoters and enhancers active in liver across 20 mammalian species from six diverse orders by profiling genomic enrichment of H3K27 acetylation and H3K4 trimethylation. We report that rapid evolution of enhancers is a universal feature of mammalian genomes. Most of the recently evolved enhancers arise from ancestral DNA exaptation, rather than lineage-specific expansions of repeat elements. In contrast, almost all liver promoters are partially or fully conserved across these species. Our data further reveal that recently evolved enhancers can be associated with genes under positive selection, demonstrating the power of this approach for annotating regulatory adaptations in genomic sequences. These results provide important insight into the functional genetics underpinning mammalian regulatory evolution
@Article{25635462, author = {Villar D and Berthelot C and Aldridge S and Rayner TF and Lukk M and Pignatelli M and Park TJ and Deaville R and Erichsen JT and Jasinska AJ and Turner JMA and Bertelsen MF and Murchison EP and Flicek P and Odom DT}, title = {Enhancer Evolution across 20 Mammalian Species}, journal = {Cell}, volume = {160}, number = {3}, pages = {554--566}, year = {2015}, doi = {10.1016/j.cell.2015.01.006}, abstract = {The mammalian radiation has corresponded with rapid changes in noncoding regions of the genome, but we lack a comprehensive understanding of regulatory evolution in mammals. Here, we track the evolution of promoters and enhancers active in liver across 20 mammalian species from six diverse orders by profiling genomic enrichment of H3K27 acetylation and H3K4 trimethylation. We report that rapid evolution of enhancers is a universal feature of mammalian genomes. Most of the recently evolved enhancers arise from ancestral DNA exaptation, rather than lineage-specific expansions of repeat elements. In contrast, almost all liver promoters are partially or fully conserved across these species. Our data further reveal that recently evolved enhancers can be associated with genes under positive selection, demonstrating the power of this approach for annotating regulatory adaptations in genomic sequences. These results provide important insight into the functional genetics underpinning mammalian regulatory evolution},} - F Cunningham, MR Amode, D Barrell, K Beal, K Billis, S Brent, D Carvalho-Silva, P Clapham, G Coates, S Fitzgerald, L Gil, CG Girón, L Gordon, T Hourlier, SE Hunt, SH Janacek, N Johnson, T Juettemann, AK Kähäri, S Keenan, FJ Martin, T Maurel, W McLaren, DN Murphy, R Nag, B Overduin, A Parker, M Patricio, E Perry, M Pignatelli, HS Riat, D Sheppard, K Taylor, A Thormann, A Vullo, SP Wilder, A Zadissa, BL Aken, E Birney, J Harrow, R Kinsella, M Muffato, M Ruffier, SMJ Searle, G Spudich, SJ Trevanion, A Yates, DR Zerbino, P Flicek. Ensembl 2015. Nucleic Acids Res 2015;43(Database):D662–D669. doi:10.1093/nar/gku1010
[BibTeX] [Abstract]
Ensembl (http://www.ensembl.org) is a genomic interpretation system providing the most up-to-date annotations, querying tools and access methods for chordates and key model organisms. This year we released updated annotation (gene models, comparative genomics, regulatory regions and variation) on the new human assembly, GRCh38, although we continue to support researchers using the GRCh37.p13 assembly through a dedicated site (http://grch37.ensembl.org). Our Regulatory Build has been revamped to identify regulatory regions of interest and to efficiently highlight their activity across disparate epigenetic data sets. A number of new interfaces allow users to perform large-scale comparisons of their data against our annotations. The REST server (http://rest.ensembl.org), which allows programs written in any language to query our databases, has moved to a full service alongside our upgraded website tools. Our online Variant Effect Predictor tool has been updated to process more variants and calculate summary statistics. Lastly, the WiggleTools package enables users to summarize large collections of data sets and view them as single tracks in Ensembl. The Ensembl code base itself is more accessible: it is now hosted on our GitHub organization page (https://github.com/Ensembl) under an Apache 2.0 open source license
@Article{25352552, author = {Cunningham F and Amode MR and Barrell D and Beal K and Billis K and Brent S and Carvalho-Silva D and Clapham P and Coates G and Fitzgerald S and Gil L and Girón CG and Gordon L and Hourlier T and Hunt SE and Janacek SH and Johnson N and Juettemann T and Kähäri AK and Keenan S and Martin FJ and Maurel T and McLaren W and Murphy DN and Nag R and Overduin B and Parker A and Patricio M and Perry E and Pignatelli M and Riat HS and Sheppard D and Taylor K and Thormann A and Vullo A and Wilder SP and Zadissa A and Aken BL and Birney E and Harrow J and Kinsella R and Muffato M and Ruffier M and Searle SMJ and Spudich G and Trevanion SJ and Yates A and Zerbino DR and Flicek P}, title = {Ensembl 2015}, journal = {Nucleic Acids Res}, volume = {43}, number = {Database}, pages = {D662--D669}, year = {2015}, doi = {10.1093/nar/gku1010}, howpublished = {Advanced online publication: 28 October 2014}, abstract = {Ensembl (http://www.ensembl.org) is a genomic interpretation system providing the most up-to-date annotations, querying tools and access methods for chordates and key model organisms. This year we released updated annotation (gene models, comparative genomics, regulatory regions and variation) on the new human assembly, GRCh38, although we continue to support researchers using the GRCh37.p13 assembly through a dedicated site (http://grch37.ensembl.org). Our Regulatory Build has been revamped to identify regulatory regions of interest and to efficiently highlight their activity across disparate epigenetic data sets. A number of new interfaces allow users to perform large-scale comparisons of their data against our annotations. The REST server (http://rest.ensembl.org), which allows programs written in any language to query our databases, has moved to a full service alongside our upgraded website tools. Our online Variant Effect Predictor tool has been updated to process more variants and calculate summary statistics. Lastly, the WiggleTools package enables users to summarize large collections of data sets and view them as single tracks in Ensembl. The Ensembl code base itself is more accessible: it is now hosted on our GitHub organization page (https://github.com/Ensembl) under an Apache 2.0 open source license},} - SOM Dyke, WA Cheung, Y Joly, O Ammerpohl, P Lutsik, MA Rothstein, M Caron, S Busche, G Bourque, L Rönnblom, P Flicek, S Beck, M Hirst, H Stunnenberg, R Siebert, J Walter, T Pastinen. Epigenome data release: a participant-centered approach to privacy protection. Genome Biology 2015;16(1):142. doi:10.1186/s13059-015-0723-0
[BibTeX] [Abstract]
Large-scale epigenome mapping by the NIH Roadmap Epigenomics Project, the ENCODE Consortium and the International Human Epigenome Consortium (IHEC) produces genome-wide DNA methylation data at one base-pair resolution. We examine how such data can be made open-access while balancing appropriate interpretation and genomic privacy. We propose guidelines for data release that both reduce ambiguity in the interpretation of open-access data and limit immediate access to genetic variation data that are made available through controlled access.
@Article{26185018, author = {Dyke SOM and Cheung WA and Joly Y and Ammerpohl O and Lutsik P and Rothstein MA and Caron M and Busche S and Bourque G and Rönnblom L and Flicek P and Beck S and Hirst M and Stunnenberg H and Siebert R and Walter J and Pastinen T}, title = {Epigenome data release: a participant-centered approach to privacy protection}, journal = {Genome Biology}, volume = {16}, number = {1}, pages = {142}, year = {2015}, doi = {10.1186/s13059-015-0723-0}, abstract = {Large-scale epigenome mapping by the NIH Roadmap Epigenomics Project, the ENCODE Consortium and the International Human Epigenome Consortium (IHEC) produces genome-wide DNA methylation data at one base-pair resolution. We examine how such data can be made open-access while balancing appropriate interpretation and genomic privacy. We propose guidelines for data release that both reduce ambiguity in the interpretation of open-access data and limit immediate access to genetic variation data that are made available through controlled access.},} - DM Church, VA Schneider, KM Steinberg, MC Schatz, AR Quinlan, C-S Chin, PA Kitts, B Aken, GT Marth, MM Hoffman, J Herrero, MLZ Mendoza, R Durbin, P Flicek. Extending reference assembly models. Genome Biol 2015;16(1):13. doi:10.1186/s13059-015-0587-3
[BibTeX] [Abstract]
The human genome reference assembly is crucial for aligning and analyzing sequence data, and for genome annotation, among other roles. However, the models and analysis assumptions that underlie the current assembly need revising to fully represent human sequence diversity. Improved analysis tools and updated data reporting formats are also required.
@Article{25651527, author = {Church DM and Schneider VA and Steinberg KM and Schatz MC and Quinlan AR and Chin C-S and Kitts PA and Aken B and Marth GT and Hoffman MM and Herrero J and Mendoza MLZ and Durbin R and Flicek P}, title = {Extending reference assembly models}, journal = {Genome Biol}, volume = {16}, number = {1}, pages = {13}, year = {2015}, doi = {10.1186/s13059-015-0587-3}, abstract = {The human genome reference assembly is crucial for aligning and analyzing sequence data, and for genome annotation, among other roles. However, the models and analysis assumptions that underlie the current assembly need revising to fully represent human sequence diversity. Improved analysis tools and updated data reporting formats are also required.},} - RS Ritchie Graham, Paul Flicek. Functional Annotation of Rare Genetic Variants. In: E Zeggini, `editor`. In: Assessing Rare Variation in Complex Traits. New York: Springer, 2015, 57-70.
[BibTeX]@Incollection{Ritchie2015, author = {Ritchie Graham RS and Flicek Paul}, title = {Functional Annotation of Rare Genetic Variants}, booktitle = {Assessing Rare Variation in Complex Traits}, editor = {Zeggini E}, publisher = {Springer}, address = {New York}, pages = {57-70}, year = {2015}, } - INFRAFRONTIER Consortium.. INFRAFRONTIER-providing mutant mouse resources as research tools for the international scientific community. Nucleic Acids Res 2015;43:D1171–D1175. doi:10.1093/nar/gku1193
[BibTeX] [Abstract]
The laboratory mouse is a key model organism to investigate mechanism and therapeutics of human disease. The number of targeted genetic mouse models of disease is growing rapidly due to high-throughput production strategies employed by the International Mouse Phenotyping Consortium (IMPC) and the development of new, more efficient genome engineering techniques such as CRISPR based systems. We have previously described the European Mouse Mutant Archive (EMMA) resource and how this international infrastructure provides archiving and distribution worldwide for mutant mouse strains. EMMA has since evolved into INFRAFRONTIER (http://www.infrafrontier.eu), the pan-European research infrastructure for the systemic phenotyping, archiving and distribution of mouse disease models. Here we describe new features including improved search for mouse strains, support for new embryonic stem cell resources, access to training materials via a comprehensive knowledgebase and the promotion of innovative analytical and diagnostic techniques
@Article{25414328, author = {{INFRAFRONTIER Consortium.}}, title = {INFRAFRONTIER-providing mutant mouse resources as research tools for the international scientific community}, journal = {Nucleic Acids Res}, volume = {43}, pages = {D1171--D1175}, year = {2015}, doi = {10.1093/nar/gku1193}, howpublished = {Advanced online publication: 20 November 2014}, abstract = {The laboratory mouse is a key model organism to investigate mechanism and therapeutics of human disease. The number of targeted genetic mouse models of disease is growing rapidly due to high-throughput production strategies employed by the International Mouse Phenotyping Consortium (IMPC) and the development of new, more efficient genome engineering techniques such as CRISPR based systems. We have previously described the European Mouse Mutant Archive (EMMA) resource and how this international infrastructure provides archiving and distribution worldwide for mutant mouse strains. EMMA has since evolved into INFRAFRONTIER (http://www.infrafrontier.eu), the pan-European research infrastructure for the systemic phenotyping, archiving and distribution of mouse disease models. Here we describe new features including improved search for mouse strains, support for new embryonic stem cell resources, access to training materials via a comprehensive knowledgebase and the promotion of innovative analytical and diagnostic techniques},} - CJ Zepeda-Mendoza, S Mukhopadhyay, ES Wong, N Harder, E Splinter, de E Wit, MA Eckersley-Maslin, T Ried, R Eils, K Rohr, A Mills, de W Laat, P Flicek, AM Sengupta, DL Spector. Quantitative analysis of chromatin interaction changes upon a 4.3 Mb deletion at mouse 4E2. BMC Genomics 2015;16(1):982. doi:10.1186/s12864-015-2137-5
[BibTeX] [Abstract]
\textbf{BACKGROUND:} Circular chromosome conformation capture (4C) has provided important insights into three dimensional (3D) genome organization and its critical impact on the regulation of gene expression. We developed a new quantitative framework based on polymer physics for the analysis of paired-end sequencing 4C (PE-4Cseq) data. We applied this strategy to the study of chromatin interaction changes upon a 4.3 Mb DNA deletion in mouse region 4E2.
\textbf{RESULTS:} A significant number of differentially interacting regions (DIRs) and chromatin compaction changes were detected in the deletion chromosome compared to a wild-type (WT) control. Selected DIRs were validated by 3D DNA FISH experiments, demonstrating the robustness of our pipeline. Interestingly, significant overlaps of DIRs with CTCF/Smc1 binding sites and differentially expressed genes were observed.
\textbf{CONCLUSIONS:} Altogether, our PE-4Cseq analysis pipeline provides a comprehensive characterization of DNA deletion effects on chromatin structure and function@Article{26589460, author = {Zepeda-Mendoza CJ and Mukhopadhyay S and Wong ES and Harder N and Splinter E and de Wit E and Eckersley-Maslin MA and Ried T and Eils R and Rohr K and Mills A and de Laat W and Flicek P and Sengupta AM and Spector DL}, title = {Quantitative analysis of chromatin interaction changes upon a 4.3 Mb deletion at mouse 4E2}, journal = {BMC Genomics}, volume = {16}, number = {1}, pages = {982}, year = {2015}, doi = {10.1186/s12864-015-2137-5}, abstract = {\textbf{BACKGROUND:} Circular chromosome conformation capture (4C) has provided important insights into three dimensional (3D) genome organization and its critical impact on the regulation of gene expression. We developed a new quantitative framework based on polymer physics for the analysis of paired-end sequencing 4C (PE-4Cseq) data. We applied this strategy to the study of chromatin interaction changes upon a 4.3 Mb DNA deletion in mouse region 4E2.
\textbf{RESULTS:} A significant number of differentially interacting regions (DIRs) and chromatin compaction changes were detected in the deletion chromosome compared to a wild-type (WT) control. Selected DIRs were validated by 3D DNA FISH experiments, demonstrating the robustness of our pipeline. Interestingly, significant overlaps of DIRs with CTCF/Smc1 binding sites and differentially expressed genes were observed.
\textbf{CONCLUSIONS:} Altogether, our PE-4Cseq analysis pipeline provides a comprehensive characterization of DNA deletion effects on chromatin structure and function},} - S Leigh-Brown, A Goncalves, D Thybert, K Stefflova, S Watt, P Flicek, A Brazma, JC Marioni, DT Odom. Regulatory Divergence of Transcript Isoforms in a Mammalian Model System. PLoS One 2015;10(9):e0137367. doi:10.1371/journal.pone.0137367
[BibTeX] [Abstract]
Phenotypic differences between species are driven by changes in gene expression and, by extension, by modifications in the regulation of the transcriptome. Investigation of mammalian transcriptome divergence has been restricted to analysis of bulk gene expression levels and gene-internal splicing. Using allele-specific expression analysis in inter-strain hybrids of Mus musculus, we determined the contribution of multiple cellular regulatory systems to transcriptome divergence, including: alternative promoter usage, transcription start site selection, cassette exon usage, alternative last exon usage, and alternative polyadenylation site choice. Between mouse strains, a fifth of genes have variations in isoform usage that contribute to transcriptomic changes, half of which alter encoded amino acid sequence. Virtually all divergence in isoform usage altered the post-transcriptional regulatory instructions in gene UTRs. Furthermore, most genes with isoform differences between strains contain changes originating from multiple regulatory systems. This result indicates widespread cross-talk and coordination exists among different regulatory systems. Overall, isoform usage diverges in parallel with and independently to gene expression evolution, and the cis and trans regulatory contribution to each differs significantly
@Article{26339903, author = {Leigh-Brown S and Goncalves A and Thybert D and Stefflova K and Watt S and Flicek P and Brazma A and Marioni JC and Odom DT}, title = {Regulatory Divergence of Transcript Isoforms in a Mammalian Model System}, journal = {PLoS One}, volume = {10}, number = {9}, pages = {e0137367}, year = {2015}, doi = {10.1371/journal.pone.0137367}, abstract = {Phenotypic differences between species are driven by changes in gene expression and, by extension, by modifications in the regulation of the transcriptome. Investigation of mammalian transcriptome divergence has been restricted to analysis of bulk gene expression levels and gene-internal splicing. Using allele-specific expression analysis in inter-strain hybrids of Mus musculus, we determined the contribution of multiple cellular regulatory systems to transcriptome divergence, including: alternative promoter usage, transcription start site selection, cassette exon usage, alternative last exon usage, and alternative polyadenylation site choice. Between mouse strains, a fifth of genes have variations in isoform usage that contribute to transcriptomic changes, half of which alter encoded amino acid sequence. Virtually all divergence in isoform usage altered the post-transcriptional regulatory instructions in gene UTRs. Furthermore, most genes with isoform differences between strains contain changes originating from multiple regulatory systems. This result indicates widespread cross-talk and coordination exists among different regulatory systems. Overall, isoform usage diverges in parallel with and independently to gene expression evolution, and the cis and trans regulatory contribution to each differs significantly},} - E Ing-Simmons, V Seitan, A Faure, P Flicek, T Carroll, J Dekker, A Fisher, B Lenhard, M Merkenschlager. Spatial enhancer clustering and regulation of enhancer-proximal genes by cohesin. Genome Research 2015;25(4):504–513. doi:10.1101/gr.184986.114
[BibTeX] [Abstract]
In addition to mediating sister chromatid cohesion during the cell cycle, the cohesin complex associates with CTCF and with active gene regulatory elements to form long-range interactions between its binding sites. Genome-wide chromosome conformation capture had shown that cohesin's main role in interphase genome organization is in mediating interactions within architectural chromosome compartments, rather than specifying compartments per se. However, it remained unclear how cohesin-mediated interactions contribute to the regulation of gene expression. We have found that the binding of CTCF and cohesin is highly enriched at enhancers and in particular at enhancer arrays or `super-enhancers' in mouse thymocytes. Using local and global chromosome conformation capture we demonstrate that enhancer elements associate not just in linear sequence, but also in 3-D, and that spatial enhancer clustering is facilitated by cohesin. The conditional deletion of cohesin from non-cycling thymocytes preserved enhancer position, H3K27ac, H4K4me1 and enhancer transcription, but weakened interactions between enhancers. Interestingly, ~50\% of deregulated genes reside in the vicinity of enhancer elements, suggesting that cohesin regulates gene expression through spatial clustering of enhancer elements. We propose a model for cohesin-dependent gene regulation where spatial clustering of enhancer elements acts as a unified mechanism for both, enhancer-promoter `connections' and `insulation'.
@Article{25677180, author = {Ing-Simmons E and Seitan V and Faure A and Flicek P and Carroll T and Dekker J and Fisher A and Lenhard B and Merkenschlager M}, title = {Spatial enhancer clustering and regulation of enhancer-proximal genes by cohesin}, journal = {Genome Research}, volume = {25}, number = {4}, pages = {504--513}, year = {2015}, doi = {10.1101/gr.184986.114}, howpublished = {Advanced online publication: 12 February 2015}, abstract = {In addition to mediating sister chromatid cohesion during the cell cycle, the cohesin complex associates with CTCF and with active gene regulatory elements to form long-range interactions between its binding sites. Genome-wide chromosome conformation capture had shown that cohesin's main role in interphase genome organization is in mediating interactions within architectural chromosome compartments, rather than specifying compartments per se. However, it remained unclear how cohesin-mediated interactions contribute to the regulation of gene expression. We have found that the binding of CTCF and cohesin is highly enriched at enhancers and in particular at enhancer arrays or `super-enhancers' in mouse thymocytes. Using local and global chromosome conformation capture we demonstrate that enhancer elements associate not just in linear sequence, but also in 3-D, and that spatial enhancer clustering is facilitated by cohesin. The conditional deletion of cohesin from non-cycling thymocytes preserved enhancer position, H3K27ac, H4K4me1 and enhancer transcription, but weakened interactions between enhancers. Interestingly, ~50\% of deregulated genes reside in the vicinity of enhancer elements, suggesting that cohesin regulates gene expression through spatial clustering of enhancer elements. We propose a model for cohesin-dependent gene regulation where spatial clustering of enhancer elements acts as a unified mechanism for both, enhancer-promoter `connections' and `insulation'.},} - DR Zerbino, SP Wilder, N Johnson, T Juettemann, PR Flicek. The Ensembl Regulatory Build. Genome Biology 2015;16(1):56. doi:10.1186/s13059-015-0621-5
[BibTeX] [Abstract]
Most genomic variants associated with phenotypic traits or disease do not fall within gene coding regions, but in regulatory regions, rendering their interpretation difficult. We collected public data on epigenetic marks and transcription factor binding in human cell types and used it to construct an intuitive summary of regulatory regions in the human genome. We verified it against independent assays for sensitivity. The Ensembl Regulatory Build will be progressively enriched when more data is made available. It is freely available on the Ensembl browser, from the Ensembl Regulation MySQL database server and in a dedicated track hub.
@Article{25887522, author = {Zerbino DR and Wilder SP and Johnson N and Juettemann T and Flicek PR}, title = {The Ensembl Regulatory Build}, journal = {Genome Biology}, volume = {16}, number = {1}, pages = {56}, year = {2015}, doi = {10.1186/s13059-015-0621-5}, abstract = {Most genomic variants associated with phenotypic traits or disease do not fall within gene coding regions, but in regulatory regions, rendering their interpretation difficult. We collected public data on epigenetic marks and transcription factor binding in human cell types and used it to construct an intuitive summary of regulatory regions in the human genome. We verified it against independent assays for sensitivity. The Ensembl Regulatory Build will be progressively enriched when more data is made available. It is freely available on the Ensembl browser, from the Ensembl Regulation MySQL database server and in a dedicated track hub.},} - A Yates, K Beal, S Keenan, W McLaren, M Pignatelli, GRS Ritchie, M Ruffier, K Taylor, A Vullo, P Flicek. The Ensembl REST API: Ensembl Data for Any Language. Bioinformatics 2015;31(1):143–145. doi:10.1093/bioinformatics/btu613
[BibTeX] [Abstract]
\textbf{MOTIVATION:} We present a Web service to access Ensembl data using Representational State Transfer (REST). The Ensembl REST server enables the easy retrieval of a wide range of Ensembl data by most programming languages, using standard formats such as JSON and FASTA while minimizing client work. We also introduce bindings to the popular Ensembl Variant Effect Predictor tool permitting large-scale programmatic variant analysis independent of any specific programming language. Availability and implementation: The Ensembl REST API can be accessed at http://rest.ensembl.org and source code is freely available under an Apache 2.0 license from http://github.com/Ensembl/ensembl-rest.
\textbf{CONTACT:} ayates@ebi.ac.uk or flicek@ebi.ac.uk Supplementary information: Supplementary data are available at Bioinformatics online@Article{25236461, author = {Yates A and Beal K and Keenan S and McLaren W and Pignatelli M and Ritchie GRS and Ruffier M and Taylor K and Vullo A and Flicek P}, title = {The Ensembl REST API: Ensembl Data for Any Language}, journal = {Bioinformatics}, volume = {31}, number = {1}, pages = {143--145}, year = {2015}, doi = {10.1093/bioinformatics/btu613}, howpublished = {Advanced online publication: 17 September 2014}, abstract = {\textbf{MOTIVATION:} We present a Web service to access Ensembl data using Representational State Transfer (REST). The Ensembl REST server enables the easy retrieval of a wide range of Ensembl data by most programming languages, using standard formats such as JSON and FASTA while minimizing client work. We also introduce bindings to the popular Ensembl Variant Effect Predictor tool permitting large-scale programmatic variant analysis independent of any specific programming language. Availability and implementation: The Ensembl REST API can be accessed at http://rest.ensembl.org and source code is freely available under an Apache 2.0 license from http://github.com/Ensembl/ensembl-rest.
\textbf{CONTACT:} ayates@ebi.ac.uk or flicek@ebi.ac.uk Supplementary information: Supplementary data are available at Bioinformatics online},} - I Lappalainen, J Almeida-King, V Kumanduri, A Senf, JD Spalding, S Ur-Rehman, G Saunders, J Kandasamy, M Caccamo, R Leinonen, B Vaughan, T Laurent, F Rowland, P Marin-Garcia, J Barker, P Jokinen, AC Torres, de JR Argila, OM Llobet, I Medina, MS Puy, M Alberich, de la S Torre, A Navarro, J Paschall, P Flicek. The European Genome-phenome Archive of human data consented for biomedical research. Nat Genet 2015;47(7):692–695. doi:10.1038/ng.3312
[BibTeX]@Article{26111507, author = {Lappalainen I and Almeida-King J and Kumanduri V and Senf A and Spalding JD and Ur-Rehman S and Saunders G and Kandasamy J and Caccamo M and Leinonen R and Vaughan B and Laurent T and Rowland F and Marin-Garcia P and Barker J and Jokinen P and Torres AC and de Argila JR and Llobet OM and Medina I and Puy MS and Alberich M and de la Torre S and Navarro A and Paschall J and Flicek P}, title = {The European Genome-phenome Archive of human data consented for biomedical research}, journal = {Nat Genet}, volume = {47}, number = {7}, pages = {692--695}, year = {2015}, doi = {10.1038/ng.3312}, } - J Robinson, JA Halliwell, JD Hayhurst, P Flicek, P Parham, SGE Marsh. The IPD and IMGT/HLA database: allele variant databases. Nucleic Acids Res 2015;43:D423–D431. doi:10.1093/nar/gku1161
[BibTeX] [Abstract]
The Immuno Polymorphism Database (IPD) was developed to provide a centralized system for the study of polymorphism in genes of the immune system. Through the IPD project we have established a central platform for the curation and publication of locus-specific databases involved either directly or related to the function of the Major Histocompatibility Complex in a number of different species. We have collaborated with specialist groups or nomenclature committees that curate the individual sections before they are submitted to IPD for online publication. IPD consists of five core databases, with the IMGT/HLA Database as the primary database. Through the work of the various nomenclature committees, the HLA Informatics Group and in collaboration with the European Bioinformatics Institute we are able to provide public access to this data through the website http://www.ebi.ac.uk/ipd/. The IPD project continues to develop with new tools being added to address scientific developments, such as Next Generation Sequencing, and to address user feedback and requests. Regular updates to the website ensure that new and confirmatory sequences are dispersed to the immunogenetics community, and the wider research and clinical communities
@Article{25414341, author = {Robinson J and Halliwell JA and Hayhurst JD and Flicek P and Parham P and Marsh SGE}, title = {The IPD and IMGT/HLA database: allele variant databases}, journal = {Nucleic Acids Res}, volume = {43}, pages = {D423--D431}, year = {2015}, doi = {10.1093/nar/gku1161}, howpublished = {Advanced online publication: 20 November 2014}, abstract = {The Immuno Polymorphism Database (IPD) was developed to provide a centralized system for the study of polymorphism in genes of the immune system. Through the IPD project we have established a central platform for the curation and publication of locus-specific databases involved either directly or related to the function of the Major Histocompatibility Complex in a number of different species. We have collaborated with specialist groups or nomenclature committees that curate the individual sections before they are submitted to IPD for online publication. IPD consists of five core databases, with the IMGT/HLA Database as the primary database. Through the work of the various nomenclature committees, the HLA Informatics Group and in collaboration with the European Bioinformatics Institute we are able to provide public access to this data through the website http://www.ebi.ac.uk/ipd/. The IPD project continues to develop with new tools being added to address scientific developments, such as Next Generation Sequencing, and to address user feedback and requests. Regular updates to the website ensure that new and confirmatory sequences are dispersed to the immunogenetics community, and the wider research and clinical communities},} - UK10K Consortium.. The UK10K project identifies rare variants in health and disease. Nature 2015;526(7571):82–90. doi:10.1038/nature14962
[BibTeX] [Abstract]
The contribution of rare and low-frequency variants to human traits is largely unexplored. Here we describe insights from sequencing whole genomes (low read depth, 7×) or exomes (high read depth, 80×) of nearly 10,000 individuals from population-based and disease collections. In extensively phenotyped cohorts we characterize over 24 million novel sequence variants, generate a highly accurate imputation reference panel and identify novel alleles associated with levels of triglycerides (APOB), adiponectin (ADIPOQ) and low-density lipoprotein cholesterol (LDLR and RGAG1) from single-marker and rare variant aggregation tests. We describe population structure and functional annotation of rare and low-frequency variants, use the data to estimate the benefits of sequencing for association studies, and summarize lessons from disease-specific collections. Finally, we make available an extensive resource, including individual-level genetic and phenotypic data and web-based tools to facilitate the exploration of association results
@Article{26367797, author = {{UK10K Consortium.}}, title = {The UK10K project identifies rare variants in health and disease}, journal = {Nature}, volume = {526}, number = {7571}, pages = {82--90}, year = {2015}, doi = {10.1038/nature14962}, abstract = {The contribution of rare and low-frequency variants to human traits is largely unexplored. Here we describe insights from sequencing whole genomes (low read depth, 7×) or exomes (high read depth, 80×) of nearly 10,000 individuals from population-based and disease collections. In extensively phenotyped cohorts we characterize over 24 million novel sequence variants, generate a highly accurate imputation reference panel and identify novel alleles associated with levels of triglycerides (APOB), adiponectin (ADIPOQ) and low-density lipoprotein cholesterol (LDLR and RGAG1) from single-marker and rare variant aggregation tests. We describe population structure and functional annotation of rare and low-frequency variants, use the data to estimate the benefits of sequencing for association studies, and summarize lessons from disease-specific collections. Finally, we make available an extensive resource, including individual-level genetic and phenotypic data and web-based tools to facilitate the exploration of association results},} - M Schmid, J Smith, DW Burt, BL Aken, PB Antin, AL Archibald, C Ashwell, PJ Blackshear, C Boschiero, CT Brown, SC Burgess, HH Cheng, W Chow, DJ Coble, A Cooksey, RPMA Crooijmans, J Damas, RVN Davis, de DJ Koning, ME Delany, T Derrien, TT Desta, IC Dunn, M Dunn, H Ellegren, L Eöry, I Erb, M Farré, M Fasold, D Fleming, P Flicek, KE Fowler, L Frésard, DP Froman, V Garceau, PP Gardner, AA Gheyas, DK Griffin, MAM Groenen, T Haaf, O Hanotte, A Hart, J Häsler, SB Hedges, J Hertel, K Howe, A Hubbard, DA Hume, P Kaiser, D Kedra, SJ Kemp, C Klopp, KE Kniel, R Kuo, S Lagarrigue, SJ Lamont, DM Larkin, RA Lawal, SM Markland, F McCarthy, HA McCormack, MC McPherson, A Motegi, SA Muljo, A Münsterberg, R Nag, I Nanda, M Neuberger, A Nitsche, C Notredame, H Noyes, R O'Connor, EA O'Hare, AJ Oler, SC Ommeh, H Pais, M Persia, F Pitel, L Preeyanon, P Prieto Barja, EM Pritchett, DD Rhoads, CM Robinson, MN Romanov, M Rothschild, PF Roux, CJ Schmidt, AS Schneider, MG Schwartz, SM Searle, MA Skinner, CA Smith, PF Stadler, TE Steeves, C Steinlein, L Sun, M Takata, I Ulitsky, Q Wang, Y Wang, WC Warren, JMD Wood, D Wragg, H Zhou. Third Report on Chicken Genes and Chromosomes 2015. Cytogenetic and Genome Research 2015;145(2):78–179. doi:10.1159/000358072
[BibTeX]@Article{24480990, author = {Schmid M and Smith J and Burt DW and Aken BL and Antin PB and Archibald AL and Ashwell C and Blackshear PJ and Boschiero C and Brown CT and Burgess SC and Cheng HH and Chow W and Coble DJ and Cooksey A and Crooijmans RPMA and Damas J and Davis RVN and de Koning DJ and Delany ME and Derrien T and Desta TT and Dunn IC and Dunn M and Ellegren H and Eöry L and Erb I and Farré M and Fasold M and Fleming D and Flicek P and Fowler KE and Frésard L and Froman DP and Garceau V and Gardner PP and Gheyas AA and Griffin DK and Groenen MAM and Haaf T and Hanotte O and Hart A and Häsler J and Hedges SB and Hertel J and Howe K and Hubbard A and Hume DA and Kaiser P and Kedra D and Kemp SJ and Klopp C and Kniel KE and Kuo R and Lagarrigue S and Lamont SJ and Larkin DM and Lawal RA and Markland SM and McCarthy F and McCormack HA and McPherson MC and Motegi A and Muljo SA and Münsterberg A and Nag R and Nanda I and Neuberger M and Nitsche A and Notredame C and Noyes H and O'Connor R and O'Hare EA and Oler AJ and Ommeh SC and Pais H and Persia M and Pitel F and Preeyanon L and Prieto Barja P and Pritchett EM and Rhoads DD and Robinson CM and Romanov MN and Rothschild M and Roux PF and Schmidt CJ and Schneider AS and Schwartz MG and Searle SM and Skinner MA and Smith CA and Stadler PF and Steeves TE and Steinlein C and Sun L and Takata M and Ulitsky I and Wang Q and Wang Y and Warren WC and Wood JMD and Wragg D and Zhou H}, title = {Third Report on Chicken Genes and Chromosomes 2015}, journal = {Cytogenetic and Genome Research}, volume = {145}, number = {2}, pages = {78--179}, year = {2015}, doi = {10.1159/000358072}, howpublished = {Advanced online publication: 14 July 2015}, } - X Agirre, G Castellano, M Pascual, S Heath, M Kulis, V Segura, A Bergmann, A Esteve, A Merkel, E Raineri, L Agueda, J Blanc, D Richardson, L Clarke, A Datta, N Russiñol, AC Queiros, R Beekman, JR Rodríguez-Madoz, ES José-Enériz, F Fang, NC Gutiérrez, JM García-Verdugo, MI Robson, EC Schirmer, E Guruceaga, JHA Martens, M Gut, MJ Calasanz, P Flicek, R Siebert, E Campo, JFS Miguel, A Melnick, HG Stunnenberg, IG Gut, F Prosper, JI Martin-Subero. Whole-epigenome analysis in multiple myeloma reveals DNA hypermethylation of B cell-specific enhancers. Genome Research 2015;25(4):478–487. doi:10.1101/gr.180240.114
[BibTeX] [Abstract]
While analyzing the DNA methylome of multiple myeloma (MM), a plasma cell neoplasm, by whole-genome bisulfite sequencing and high-density arrays, we observed a highly heterogeneous pattern globally characterized by regional DNA hypermethylation embedded in extensive hypomethylation. In contrast to the widely reported DNA hypermethylation of promoter-associated CpG islands (CGIs) in cancer, hypermethylated sites in MM, as opposed to normal plasma cells, were located outside CpG islands and were unexpectedly associated with intronic enhancer regions defined in normal B cells and plasma cells. Both RNA-seq and in vitro reporter assays indicated that enhancer hypermethylation is globally associated with downregulation of its host genes. ChIP-seq and DNase-seq further revealed that DNA hypermethylation in these regions is related to enhancer decommissioning. Hypermethylated enhancer regions overlapped with binding sites of B cell-specific transcription factors (TFs) and the degree of enhancer methylation inversely correlated with expression levels of these TFs in MM. Furthermore, hypermethylated regions in MM were methylated in stem cells and gradually became demethylated during normal B-cell differentiation, suggesting that MM cells either reacquire epigenetic features of undifferentiated cells or maintain an epigenetic signature of a putative myeloma stem cell progenitor. Overall, we have identified DNA hypermethylation of developmentally-regulated enhancers as a new type of epigenetic modification associated with the pathogenesis of MM.
@Article{25644835, author = {Agirre X and Castellano G and Pascual M and Heath S and Kulis M and Segura V and Bergmann A and Esteve A and Merkel A and Raineri E and Agueda L and Blanc J and Richardson D and Clarke L and Datta A and Russiñol N and Queiros AC and Beekman R and Rodríguez-Madoz JR and José-Enériz ES and Fang F and Gutiérrez NC and García-Verdugo JM and Robson MI and Schirmer EC and Guruceaga E and Martens JHA and Gut M and Calasanz MJ and Flicek P and Siebert R and Campo E and Miguel JFS and Melnick A and Stunnenberg HG and Gut IG and Prosper F and Martin-Subero JI}, title = {Whole-epigenome analysis in multiple myeloma reveals DNA hypermethylation of B cell-specific enhancers}, journal = {Genome Research}, volume = {25}, number = {4}, pages = {478--487}, year = {2015}, doi = {10.1101/gr.180240.114}, howpublished = {Advanced online publication: 2 February 2015}, abstract = {While analyzing the DNA methylome of multiple myeloma (MM), a plasma cell neoplasm, by whole-genome bisulfite sequencing and high-density arrays, we observed a highly heterogeneous pattern globally characterized by regional DNA hypermethylation embedded in extensive hypomethylation. In contrast to the widely reported DNA hypermethylation of promoter-associated CpG islands (CGIs) in cancer, hypermethylated sites in MM, as opposed to normal plasma cells, were located outside CpG islands and were unexpectedly associated with intronic enhancer regions defined in normal B cells and plasma cells. Both RNA-seq and in vitro reporter assays indicated that enhancer hypermethylation is globally associated with downregulation of its host genes. ChIP-seq and DNase-seq further revealed that DNA hypermethylation in these regions is related to enhancer decommissioning. Hypermethylated enhancer regions overlapped with binding sites of B cell-specific transcription factors (TFs) and the degree of enhancer methylation inversely correlated with expression levels of these TFs in MM. Furthermore, hypermethylated regions in MM were methylated in stem cells and gradually became demethylated during normal B-cell differentiation, suggesting that MM cells either reacquire epigenetic features of undifferentiated cells or maintain an epigenetic signature of a putative myeloma stem cell progenitor. Overall, we have identified DNA hypermethylation of developmentally-regulated enhancers as a new type of epigenetic modification associated with the pathogenesis of MM.},} - M Kulis, A Merkel, S Heath, AC Queirós, RP Schuyler, G Castellano, R Beekman, E Raineri, A Esteve, G Clot, N Verdaguer-Dot, M Duran-Ferrer, N Russiñol, R Vilarrasa-Blasi, S Ecker, V Pancaldi, D Rico, L Agueda, J Blanc, D Richardson, L Clarke, A Datta, M Pascual, X Agirre, F Prosper, D Alignani, B Paiva, G Caron, T Fest, MO Muench, ME Fomin, S-T Lee, JL Wiemels, A Valencia, M Gut, P Flicek, HG Stunnenberg, R Siebert, R Küppers, IG Gut, E Campo, JI Martín-Subero. Whole-genome fingerprint of the DNA methylome during human B cell differentiation. Nat Genet 2015;47(7):746–756. doi:10.1038/ng.3291
[BibTeX] [Abstract]
We analyzed the DNA methylome of ten subpopulations spanning the entire B cell differentiation program by whole-genome bisulfite sequencing and high-density microarrays. We observed that non-CpG methylation disappeared upon B cell commitment, whereas CpG methylation changed extensively during B cell maturation, showing an accumulative pattern and affecting around 30\% of all measured CpG sites. Early differentiation stages mainly displayed enhancer demethylation, which was associated with upregulation of key B cell transcription factors and affected multiple genes involved in B cell biology. Late differentiation stages, in contrast, showed extensive demethylation of heterochromatin and methylation gain at Polycomb-repressed areas, and genes with apparent functional impact in B cells were not affected. This signature, which has previously been linked to aging and cancer, was particularly widespread in mature cells with an extended lifespan. Comparing B cell neoplasms with their normal counterparts, we determined that they frequently acquire methylation changes in regions already undergoing dynamic methylation during normal B cell differentiation
@Article{26053498, author = {Kulis M and Merkel A and Heath S and Queirós AC and Schuyler RP and Castellano G and Beekman R and Raineri E and Esteve A and Clot G and Verdaguer-Dot N and Duran-Ferrer M and Russiñol N and Vilarrasa-Blasi R and Ecker S and Pancaldi V and Rico D and Agueda L and Blanc J and Richardson D and Clarke L and Datta A and Pascual M and Agirre X and Prosper F and Alignani D and Paiva B and Caron G and Fest T and Muench MO and Fomin ME and Lee S-T and Wiemels JL and Valencia A and Gut M and Flicek P and Stunnenberg HG and Siebert R and Küppers R and Gut IG and Campo E and Martín-Subero JI}, title = {Whole-genome fingerprint of the DNA methylome during human B cell differentiation}, journal = {Nat Genet}, volume = {47}, number = {7}, pages = {746--756}, year = {2015}, doi = {10.1038/ng.3291}, abstract = {We analyzed the DNA methylome of ten subpopulations spanning the entire B cell differentiation program by whole-genome bisulfite sequencing and high-density microarrays. We observed that non-CpG methylation disappeared upon B cell commitment, whereas CpG methylation changed extensively during B cell maturation, showing an accumulative pattern and affecting around 30\% of all measured CpG sites. Early differentiation stages mainly displayed enhancer demethylation, which was associated with upregulation of key B cell transcription factors and affected multiple genes involved in B cell biology. Late differentiation stages, in contrast, showed extensive demethylation of heterochromatin and methylation gain at Polycomb-repressed areas, and genes with apparent functional impact in B cells were not affected. This signature, which has previously been linked to aging and cancer, was particularly widespread in mature cells with an extended lifespan. Comparing B cell neoplasms with their normal counterparts, we determined that they frequently acquire methylation changes in regions already undergoing dynamic methylation during normal B cell differentiation},}
2014
- F Yue, Y Cheng, A Breschi, J Vierstra, W Wu, T Ryba, R Sandstrom, Z Ma, C Davis, BD Pope, Y Shen, DD Pervouchine, S Djebali, RE Thurman, R Kaul, E Rynes, A Kirilusha, GK Marinov, BA Williams, D Trout, H Amrhein, K Fisher-Aylor, I Antoshechkin, G DeSalvo, L-H See, M Fastuca, J Drenkow, C Zaleski, A Dobin, P Prieto, J Lagarde, G Bussotti, A Tanzer, O Denas, K Li, MA Bender, M Zhang, R Byron, MT Groudine, D McCleary, L Pham, Z Ye, S Kuan, L Edsall, Y-C Wu, MD Rasmussen, MS Bansal, M Kellis, CA Keller, CS Morrissey, T Mishra, D Jain, N Dogan, RS Harris, P Cayting, T Kawli, AP Boyle, G Euskirchen, A Kundaje, S Lin, Y Lin, C Jansen, VS Malladi, MS Cline, DT Erickson, VM Kirkup, K Learned, CA Sloan, KR Rosenbloom, de B Lacerda Sousa, K Beal, M Pignatelli, P Flicek, J Lian, T Kahveci, D Lee, WJ Kent, M Ramalho Santos, J Herrero, C Notredame, A Johnson, S Vong, K Lee, D Bates, F Neri, M Diegel, T Canfield, PJ Sabo, MS Wilken, TA Reh, E Giste, A Shafer, T Kutyavin, E Haugen, D Dunn, AP Reynolds, S Neph, R Humbert, RS Hansen, M De Bruijn, L Selleri, A Rudensky, S Josefowicz, R Samstein, EE Eichler, SH Orkin, D Levasseur, T Papayannopoulou, K-H Chang, A Skoultchi, S Gosh, C Disteche, P Treuting, Y Wang, MJ Weiss, GA Blobel, X Cao, S Zhong, T Wang, PJ Good, RF Lowdon, LB Adams, X-Q Zhou, MJ Pazin, EA Feingold, B Wold, J Taylor, A Mortazavi, SM Weissman, JA Stamatoyannopoulos, MP Snyder, R Guigo, TR Gingeras, DM Gilbert, RC Hardison, MA Beer, B Ren, MENCODE Consortium. A comparative encyclopedia of DNA elements in the mouse genome. Nature 2014;515(7527):355–364. doi:10.1038/nature13992
[BibTeX] [Abstract]
The laboratory mouse shares the majority of its protein-coding genes with humans, making it the premier model organism in biomedical research, yet the two mammals differ in significant ways. To gain greater insights into both shared and species-specific transcriptional and cellular regulatory programs in the mouse, the Mouse ENCODE Consortium has mapped transcription, DNase I hypersensitivity, transcription factor binding, chromatin modifications and replication domains throughout the mouse genome in diverse cell and tissue types. By comparing with the human genome, we not only confirm substantial conservation in the newly annotated potential functional sequences, but also find a large degree of divergence of sequences involved in transcriptional regulation, chromatin state and higher order chromatin organization. Our results illuminate the wide range of evolutionary forces acting on genes and their regulatory regions, and provide a general resource for research into mammalian biology and mechanisms of human diseases
@Article{25409824, author = {Yue F and Cheng Y and Breschi A and Vierstra J and Wu W and Ryba T and Sandstrom R and Ma Z and Davis C and Pope BD and Shen Y and Pervouchine DD and Djebali S and Thurman RE and Kaul R and Rynes E and Kirilusha A and Marinov GK and Williams BA and Trout D and Amrhein H and Fisher-Aylor K and Antoshechkin I and DeSalvo G and See L-H and Fastuca M and Drenkow J and Zaleski C and Dobin A and Prieto P and Lagarde J and Bussotti G and Tanzer A and Denas O and Li K and Bender MA and Zhang M and Byron R and Groudine MT and McCleary D and Pham L and Ye Z and Kuan S and Edsall L and Wu Y-C and Rasmussen MD and Bansal MS and Kellis M and Keller CA and Morrissey CS and Mishra T and Jain D and Dogan N and Harris RS and Cayting P and Kawli T and Boyle AP and Euskirchen G and Kundaje A and Lin S and Lin Y and Jansen C and Malladi VS and Cline MS and Erickson DT and Kirkup VM and Learned K and Sloan CA and Rosenbloom KR and Lacerda de Sousa B and Beal K and Pignatelli M and Flicek P and Lian J and Kahveci T and Lee D and Kent WJ and Ramalho Santos M and Herrero J and Notredame C and Johnson A and Vong S and Lee K and Bates D and Neri F and Diegel M and Canfield T and Sabo PJ and Wilken MS and Reh TA and Giste E and Shafer A and Kutyavin T and Haugen E and Dunn D and Reynolds AP and Neph S and Humbert R and Hansen RS and De Bruijn M and Selleri L and Rudensky A and Josefowicz S and Samstein R and Eichler EE and Orkin SH and Levasseur D and Papayannopoulou T and Chang K-H and Skoultchi A and Gosh S and Disteche C and Treuting P and Wang Y and Weiss MJ and Blobel GA and Cao X and Zhong S and Wang T and Good PJ and Lowdon RF and Adams LB and Zhou X-Q and Pazin MJ and Feingold EA and Wold B and Taylor J and Mortazavi A and Weissman SM and Stamatoyannopoulos JA and Snyder MP and Guigo R and Gingeras TR and Gilbert DM and Hardison RC and Beer MA and Ren B and Consortium MENCODE}, title = {A comparative encyclopedia of DNA elements in the mouse genome}, journal = {Nature}, volume = {515}, number = {7527}, pages = {355--364}, year = {2014}, doi = {10.1038/nature13992}, abstract = {The laboratory mouse shares the majority of its protein-coding genes with humans, making it the premier model organism in biomedical research, yet the two mammals differ in significant ways. To gain greater insights into both shared and species-specific transcriptional and cellular regulatory programs in the mouse, the Mouse ENCODE Consortium has mapped transcription, DNase I hypersensitivity, transcription factor binding, chromatin modifications and replication domains throughout the mouse genome in diverse cell and tissue types. By comparing with the human genome, we not only confirm substantial conservation in the newly annotated potential functional sequences, but also find a large degree of divergence of sequences involved in transcriptional regulation, chromatin state and higher order chromatin organization. Our results illuminate the wide range of evolutionary forces acting on genes and their regulatory regions, and provide a general resource for research into mammalian biology and mechanisms of human diseases},} - EM Ramos, C Din-Lovinescu, JS Berg, LD Brooks, A Duncanson, M Dunn, P Good, TJP Hubbard, GP Jarvik, C O'Donnell, ST Sherry, N Aronson, LG Biesecker, B Blumberg, N Calonge, HM Colhoun, RS Epstein, P Flicek, ES Gordon, ED Green, RC Green, M Hurles, K Kawamoto, W Knaus, DH Ledbetter, HP Levy, E Lyon, D Maglott, HL McLeod, N Rahman, G Randhawa, C Wicklund, TA Manolio, RL Chisholm, MS Williams. Characterizing genetic variants for clinical action. Am J Med Genet C Semin Med Genet 2014;166(1):93–104. doi:10.1002/ajmg.c.31386
[BibTeX] [Abstract]
Genome-wide association studies, DNA sequencing studies, and other genomic studies are finding an increasing number of genetic variants associated with clinical phenotypes that may be useful in developing diagnostic, preventive, and treatment strategies for individual patients. However, few variants have been integrated into routine clinical practice. The reasons for this are several, but two of the most significant are limited evidence about the clinical implications of the variants and a lack of a comprehensive knowledge base that captures genetic variants, their phenotypic associations, and other pertinent phenotypic information that is openly accessible to clinical groups attempting to interpret sequencing data. As the field of medicine begins to incorporate genome-scale analysis into clinical care, approaches need to be developed for collecting and characterizing data on the clinical implications of variants, developing consensus on their actionability, and making this information available for clinical use. The National Human Genome Research Institute (NHGRI) and the Wellcome Trust thus convened a workshop to consider the processes and resources needed to: (1) identify clinically valid genetic variants; (2) decide whether they are actionable and what the action should be; and (3) provide this information for clinical use. This commentary outlines the key discussion points and recommendations from the workshop. {\copyright} 2014 Wiley Periodicals, Inc
@Article{24634402, author = {Ramos EM and Din-Lovinescu C and Berg JS and Brooks LD and Duncanson A and Dunn M and Good P and Hubbard TJP and Jarvik GP and O'Donnell C and Sherry ST and Aronson N and Biesecker LG and Blumberg B and Calonge N and Colhoun HM and Epstein RS and Flicek P and Gordon ES and Green ED and Green RC and Hurles M and Kawamoto K and Knaus W and Ledbetter DH and Levy HP and Lyon E and Maglott D and McLeod HL and Rahman N and Randhawa G and Wicklund C and Manolio TA and Chisholm RL and Williams MS}, title = {Characterizing genetic variants for clinical action}, journal = {Am J Med Genet C Semin Med Genet}, volume = {166}, number = {1}, pages = {93--104}, year = {2014}, doi = {10.1002/ajmg.c.31386}, abstract = {Genome-wide association studies, DNA sequencing studies, and other genomic studies are finding an increasing number of genetic variants associated with clinical phenotypes that may be useful in developing diagnostic, preventive, and treatment strategies for individual patients. However, few variants have been integrated into routine clinical practice. The reasons for this are several, but two of the most significant are limited evidence about the clinical implications of the variants and a lack of a comprehensive knowledge base that captures genetic variants, their phenotypic associations, and other pertinent phenotypic information that is openly accessible to clinical groups attempting to interpret sequencing data. As the field of medicine begins to incorporate genome-scale analysis into clinical care, approaches need to be developed for collecting and characterizing data on the clinical implications of variants, developing consensus on their actionability, and making this information available for clinical use. The National Human Genome Research Institute (NHGRI) and the Wellcome Trust thus convened a workshop to consider the processes and resources needed to: (1) identify clinically valid genetic variants; (2) decide whether they are actionable and what the action should be; and (3) provide this information for clinical use. This commentary outlines the key discussion points and recommendations from the workshop. {\copyright} 2014 Wiley Periodicals, Inc},} - GRS Ritchie, P Flicek. Computational approaches to interpreting genomic sequence variation. Genome Med 2014;6(10):87. doi:10.1186/s13073-014-0087-1
[BibTeX] [Abstract]
Identifying sequence variants that play a mechanistic role in human disease and other phenotypes is a fundamental goal in human genetics and will be important in translating the results of variation studies. Experimental validation to confirm that a variant causes the biochemical changes responsible for a given disease or phenotype is considered the gold standard, but this cannot currently be applied to the 3 million or so variants expected in an individual genome. This has prompted the development of a wide variety of computational approaches that use several different sources of information to identify functional variation. Here, we review and assess the limitations of computational techniques for categorizing variants according to functional classes, prioritizing variants for experimental follow-up and generating hypotheses about the possible molecular mechanisms to inform downstream experiments. We discuss the main current bioinformatics approaches to identifying functional variation, including widely used algorithms for coding variation such as SIFT and PolyPhen and also novel techniques for interpreting variation across the genome
@Article{25473426, author = {Ritchie GRS and Flicek P}, title = {Computational approaches to interpreting genomic sequence variation}, journal = {Genome Med}, volume = {6}, number = {10}, pages = {87}, year = {2014}, doi = {10.1186/s13073-014-0087-1}, abstract = {Identifying sequence variants that play a mechanistic role in human disease and other phenotypes is a fundamental goal in human genetics and will be important in translating the results of variation studies. Experimental validation to confirm that a variant causes the biochemical changes responsible for a given disease or phenotype is considered the gold standard, but this cannot currently be applied to the 3 million or so variants expected in an individual genome. This has prompted the development of a wide variety of computational approaches that use several different sources of information to identify functional variation. Here, we review and assess the limitations of computational techniques for categorizing variants according to functional classes, prioritizing variants for experimental follow-up and generating hypotheses about the possible molecular mechanisms to inform downstream experiments. We discuss the main current bioinformatics approaches to identifying functional variation, including widely used algorithms for coding variation such as SIFT and PolyPhen and also novel techniques for interpreting variation across the genome},} - P Flicek, MR Amode, D Barrell, K Beal, K Billis, S Brent, D Carvalho-Silva, P Clapham, G Coates, S Fitzgerald, L Gil, CG Girón, L Gordon, T Hourlier, S Hunt, N Johnson, T Juettemann, AK Kähäri, S Keenan, E Kulesha, FJ Martin, T Maurel, WM McLaren, DN Murphy, R Nag, B Overduin, M Pignatelli, B Pritchard, E Pritchard, HS Riat, M Ruffier, D Sheppard, K Taylor, A Thormann, SJ Trevanion, A Vullo, SP Wilder, M Wilson, A Zadissa, BL Aken, E Birney, F Cunningham, J Harrow, J Herrero, TJP Hubbard, R Kinsella, M Muffato, A Parker, G Spudich, A Yates, DR Zerbino, SMJ Searle. Ensembl 2014. Nucleic Acids Res 2014;42(Database issue):D749–755. doi:10.1093/nar/gkt1196
[BibTeX] [Abstract]
Ensembl (http://www.ensembl.org) creates tools and data resources to facilitate genomic analysis in chordate species with an emphasis on human, major vertebrate model organisms and farm animals. Over the past year we have increased the number of species that we support to 77 and expanded our genome browser with a new scrollable overview and improved variation and phenotype views. We also report updates to our core datasets and improvements to our gene homology relationships from the addition of new species. Our REST service has been extended with additional support for comparative genomics and ontology information. Finally, we provide updated information about our methods for data access and resources for user training
@Article{24316576, author = {Flicek P and Amode MR and Barrell D and Beal K and Billis K and Brent S and Carvalho-Silva D and Clapham P and Coates G and Fitzgerald S and Gil L and Girón CG and Gordon L and Hourlier T and Hunt S and Johnson N and Juettemann T and Kähäri AK and Keenan S and Kulesha E and Martin FJ and Maurel T and McLaren WM and Murphy DN and Nag R and Overduin B and Pignatelli M and Pritchard B and Pritchard E and Riat HS and Ruffier M and Sheppard D and Taylor K and Thormann A and Trevanion SJ and Vullo A and Wilder SP and Wilson M and Zadissa A and Aken BL and Birney E and Cunningham F and Harrow J and Herrero J and Hubbard TJP and Kinsella R and Muffato M and Parker A and Spudich G and Yates A and Zerbino DR and Searle SMJ}, title = {Ensembl 2014}, journal = {Nucleic Acids Res}, volume = {42}, number = {Database issue}, pages = {D749--755}, year = {2014}, doi = {10.1093/nar/gkt1196}, howpublished = {Advanced online publication: 6 December 2013}, abstract = {Ensembl (http://www.ensembl.org) creates tools and data resources to facilitate genomic analysis in chordate species with an emphasis on human, major vertebrate model organisms and farm animals. Over the past year we have increased the number of species that we support to 77 and expanded our genome browser with a new scrollable overview and improved variation and phenotype views. We also report updates to our core datasets and improvements to our gene homology relationships from the addition of new species. Our REST service has been extended with additional support for comparative genomics and ontology information. Finally, we provide updated information about our methods for data access and resources for user training},} - D Villar, P Flicek, DT Odom. Evolution of transcription factor binding in metazoans - mechanisms and functional implications. Nat Rev Genet 2014;15(4):221–233. doi:10.1038/nrg3481
[BibTeX] [Abstract]
Differences in transcription factor binding can contribute to organismal evolution by altering downstream gene expression programmes. Genome-wide studies in Drosophila melanogaster and mammals have revealed common quantitative and combinatorial properties of in vivo DNA binding, as well as marked differences in the rate and mechanisms of evolution of transcription factor binding in metazoans. Here, we review the recently discovered rapid `re-wiring' of in vivo transcription factor binding between related metazoan species and summarize general principles underlying the observed patterns of evolution. We then consider what might explain the differences in genome evolution between metazoan phyla and outline the conceptual and technological challenges facing this research field
@Article{24590227, author = {Villar D and Flicek P and Odom DT}, title = {Evolution of transcription factor binding in metazoans - mechanisms and functional implications}, journal = {Nat Rev Genet}, volume = {15}, number = {4}, pages = {221--233}, year = {2014}, doi = {10.1038/nrg3481}, abstract = {Differences in transcription factor binding can contribute to organismal evolution by altering downstream gene expression programmes. Genome-wide studies in Drosophila melanogaster and mammals have revealed common quantitative and combinatorial properties of in vivo DNA binding, as well as marked differences in the rate and mechanisms of evolution of transcription factor binding in metazoans. Here, we review the recently discovered rapid `re-wiring' of in vivo transcription factor binding between related metazoan species and summarize general principles underlying the observed patterns of evolution. We then consider what might explain the differences in genome evolution between metazoan phyla and outline the conceptual and technological challenges facing this research field},} - GRS Ritchie, I Dunham, E Zeggini, P Flicek. Functional annotation of noncoding sequence variants. Nat Methods 2014;11(3):294–296. doi:10.1038/nmeth.2832
[BibTeX] [Abstract]
Identifying functionally relevant variants against the background of ubiquitous genetic variation is a major challenge in human genetics. For variants in protein-coding regions, our understanding of the genetic code and splicing allows us to identify likely candidates, but interpreting variants outside genic regions is more difficult. Here we present genome-wide annotation of variants (GWAVA), a tool that supports prioritization of noncoding variants by integrating various genomic and epigenomic annotations
@Article{24487584, author = {Ritchie GRS and Dunham I and Zeggini E and Flicek P}, title = {Functional annotation of noncoding sequence variants}, journal = {Nat Methods}, volume = {11}, number = {3}, pages = {294--296}, year = {2014}, doi = {10.1038/nmeth.2832}, howpublished = {Advanced online publication: 2 February 2014}, abstract = {Identifying functionally relevant variants against the background of ubiquitous genetic variation is a major challenge in human genetics. For variants in protein-coding regions, our understanding of the genetic code and splicing allows us to identify likely candidates, but interpreting variants outside genic regions is more difficult. Here we present genome-wide annotation of variants (GWAVA), a tool that supports prioritization of noncoding variants by integrating various genomic and epigenomic annotations},} - L Carbone, RA Harris, S Gnerre, KR Veeramah, B Lorente-Galdos, J Huddleston, TJ Meyer, J Herrero, C Roos, B Aken, F Anaclerio, N Archidiacono, C Baker, D Barrell, MA Batzer, K Beal, A Blancher, CL Bohrson, M Brameier, MS Campbell, O Capozzi, C Casola, G Chiatante, A Cree, A Damert, de PJ Jong, L Dumas, M Fernandez-Callejo, P Flicek, NV Fuchs, I Gut, M Gut, MW Hahn, J Hernandez-Rodriguez, LW Hillier, R Hubley, B Ianc, Z Izsvák, NG Jablonski, LM Johnstone, A Karimpour-Fard, MK Konkel, D Kostka, NH Lazar, SL Lee, LR Lewis, Y Liu, DP Locke, S Mallick, FL Mendez, M Muffato, LV Nazareth, KA Nevonen, M O'Bleness, C Ochis, DT Odom, KS Pollard, J Quilez, D Reich, M Rocchi, GG Schumann, S Searle, JM Sikela, G Skollar, A Smit, K Sonmez, ten B Hallers, E Terhune, GWC Thomas, B Ullmer, M Ventura, JA Walker, JD Wall, L Walter, MC Ward, SJ Wheelan, CW Whelan, S White, LJ Wilhelm, AE Woerner, M Yandell, B Zhu, MF Hammer, T Marques-Bonet, EE Eichler, L Fulton, C Fronick, DM Muzny, WC Warren, KC Worley, J Rogers, RK Wilson, RA Gibbs. Gibbon genome and the fast karyotype evolution of small apes. Nature 2014;513(7517):195–201. doi:10.1038/nature13679
[BibTeX] [Abstract]
Gibbons are small arboreal apes that display an accelerated rate of evolutionary chromosomal rearrangement and occupy a key node in the primate phylogeny between Old World monkeys and great apes. Here we present the assembly and analysis of a northern white-cheeked gibbon (Nomascus leucogenys) genome. We describe the propensity for a gibbon-specific retrotransposon (LAVA) to insert into chromosome segregation genes and alter transcription by providing a premature termination site, suggesting a possible molecular mechanism for the genome plasticity of the gibbon lineage. We further show that the gibbon genera (Nomascus, Hylobates, Hoolock and Symphalangus) experienced a near-instantaneous radiation ∼5 million years ago, coincident with major geographical changes in southeast Asia that caused cycles of habitat compression and expansion. Finally, we identify signatures of positive selection in genes important for forelimb development (TBX5) and connective tissues (COL1A1) that may have been involved in the adaptation of gibbons to their arboreal habitat
@Article{25209798, author = {Carbone L and Harris RA and Gnerre S and Veeramah KR and Lorente-Galdos B and Huddleston J and Meyer TJ and Herrero J and Roos C and Aken B and Anaclerio F and Archidiacono N and Baker C and Barrell D and Batzer MA and Beal K and Blancher A and Bohrson CL and Brameier M and Campbell MS and Capozzi O and Casola C and Chiatante G and Cree A and Damert A and de Jong PJ and Dumas L and Fernandez-Callejo M and Flicek P and Fuchs NV and Gut I and Gut M and Hahn MW and Hernandez-Rodriguez J and Hillier LW and Hubley R and Ianc B and Izsvák Z and Jablonski NG and Johnstone LM and Karimpour-Fard A and Konkel MK and Kostka D and Lazar NH and Lee SL and Lewis LR and Liu Y and Locke DP and Mallick S and Mendez FL and Muffato M and Nazareth LV and Nevonen KA and O'Bleness M and Ochis C and Odom DT and Pollard KS and Quilez J and Reich D and Rocchi M and Schumann GG and Searle S and Sikela JM and Skollar G and Smit A and Sonmez K and ten Hallers B and Terhune E and Thomas GWC and Ullmer B and Ventura M and Walker JA and Wall JD and Walter L and Ward MC and Wheelan SJ and Whelan CW and White S and Wilhelm LJ and Woerner AE and Yandell M and Zhu B and Hammer MF and Marques-Bonet T and Eichler EE and Fulton L and Fronick C and Muzny DM and Warren WC and Worley KC and Rogers J and Wilson RK and Gibbs RA}, title = {Gibbon genome and the fast karyotype evolution of small apes}, journal = {Nature}, volume = {513}, number = {7517}, pages = {195--201}, year = {2014}, doi = {10.1038/nature13679}, abstract = {Gibbons are small arboreal apes that display an accelerated rate of evolutionary chromosomal rearrangement and occupy a key node in the primate phylogeny between Old World monkeys and great apes. Here we present the assembly and analysis of a northern white-cheeked gibbon (Nomascus leucogenys) genome. We describe the propensity for a gibbon-specific retrotransposon (LAVA) to insert into chromosome segregation genes and alter transcription by providing a premature termination site, suggesting a possible molecular mechanism for the genome plasticity of the gibbon lineage. We further show that the gibbon genera (Nomascus, Hylobates, Hoolock and Symphalangus) experienced a near-instantaneous radiation ∼5 million years ago, coincident with major geographical changes in southeast Asia that caused cycles of habitat compression and expansion. Finally, we identify signatures of positive selection in genes important for forelimb development (TBX5) and connective tissues (COL1A1) that may have been involved in the adaptation of gibbons to their arboreal habitat},} - AC Nelson, SJ Cutty, M Niini, DL Stemple, P Flicek, C Houart, A Bruce, FC Wardle. Global identification of Smad2 and Eomesodermin targets in zebrafish identifies a conserved transcriptional network in mesendoderm and a novel role for Eomesodermin in repression of ectodermal gene expression. BMC Biol 2014;12(1):81. doi:10.1186/s12915-014-0081-5
[BibTeX] [Abstract]
BackgroundNodal signalling is an absolute requirement for normal mesoderm and endoderm formation in vertebrate embryos, yet the transcriptional networks acting directly downstream of Nodal and the extent to which they are conserved is largely unexplored, particularly in vivo. Eomesodermin also plays a role in patterning mesoderm and endoderm in vertebrates, but its mechanisms of action, and how it interacts with the Nodal signalling pathway are still unclear.ResultsUsing a combination of ChIP-seq and expression analysis we identify direct targets of Smad2, the effector of Nodal signalling in blastula stage zebrafish embryos, including many novel target genes. Through comparison of these data with published ChIP-seq data in human, mouse and Xenopus we show that the transcriptional network driven by Smad2 in mesoderm and endoderm is conserved in these vertebrate species. We also show that Smad2 and zebrafish Eomesodermin a (Eomesa) bind common genomic regions proximal to genes involved in mesoderm and endoderm formation, suggesting Eomesa forms a general component of the Smad2 signalling complex in zebrafish. Combinatorial perturbation of Eomesa and Smad2-interacting factor Foxh1 results in loss of both mesoderm and endoderm markers, confirming the role of Eomesa in endoderm formation and its functional interaction with Foxh1 for correct Nodal signalling. Finally, we uncover a novel, role for Eomesa in repressing ectodermal genes in the early blastula.ConclusionOur data demonstrate that evolutionarily conserved developmental functions of Nodal signalling occur through maintenance of the transcriptional network directed by Smad2. This network is modulated by Eomesa in zebrafish which acts to promote mesoderm and endoderm formation in combination with Nodal signalling, whilst Eomesa also opposes ectoderm gene expression. Eomesa therefore regulates the formation of all three germ layers in the early zebrafish embryo
@Article{25277163, author = {Nelson AC and Cutty SJ and Niini M and Stemple DL and Flicek P and Houart C and Bruce A and Wardle FC}, title = {Global identification of Smad2 and Eomesodermin targets in zebrafish identifies a conserved transcriptional network in mesendoderm and a novel role for Eomesodermin in repression of ectodermal gene expression}, journal = {BMC Biol}, volume = {12}, number = {1}, pages = {81}, year = {2014}, doi = {10.1186/s12915-014-0081-5}, abstract = {BackgroundNodal signalling is an absolute requirement for normal mesoderm and endoderm formation in vertebrate embryos, yet the transcriptional networks acting directly downstream of Nodal and the extent to which they are conserved is largely unexplored, particularly in vivo. Eomesodermin also plays a role in patterning mesoderm and endoderm in vertebrates, but its mechanisms of action, and how it interacts with the Nodal signalling pathway are still unclear.ResultsUsing a combination of ChIP-seq and expression analysis we identify direct targets of Smad2, the effector of Nodal signalling in blastula stage zebrafish embryos, including many novel target genes. Through comparison of these data with published ChIP-seq data in human, mouse and Xenopus we show that the transcriptional network driven by Smad2 in mesoderm and endoderm is conserved in these vertebrate species. We also show that Smad2 and zebrafish Eomesodermin a (Eomesa) bind common genomic regions proximal to genes involved in mesoderm and endoderm formation, suggesting Eomesa forms a general component of the Smad2 signalling complex in zebrafish. Combinatorial perturbation of Eomesa and Smad2-interacting factor Foxh1 results in loss of both mesoderm and endoderm markers, confirming the role of Eomesa in endoderm formation and its functional interaction with Foxh1 for correct Nodal signalling. Finally, we uncover a novel, role for Eomesa in repressing ectodermal genes in the early blastula.ConclusionOur data demonstrate that evolutionarily conserved developmental functions of Nodal signalling occur through maintenance of the transcriptional network directed by Smad2. This network is modulated by Eomesa in zebrafish which acts to promote mesoderm and endoderm formation in combination with Nodal signalling, whilst Eomesa also opposes ectoderm gene expression. Eomesa therefore regulates the formation of all three germ layers in the early zebrafish embryo},} - JAL MacArthur, J Morales, RE Tully, A Astashyn, L Gil, EA Bruford, P Larsson, P Flicek, R Dalgleish, DR Maglott, F Cunningham. Locus Reference Genomic: reference sequences for the reporting of clinically relevant sequence variants. Nucleic Acids Res 2014;42(Database issue):D873–8. doi:10.1093/nar/gkt1198
[BibTeX] [Abstract]
Locus Reference Genomic (LRG; http://www.lrg-sequence.org/) records contain internationally recognized stable reference sequences designed specifically for reporting clinically relevant sequence variants. Each LRG is contained within a single file consisting of a stable `fixed' section and a regularly updated `updatable' section. The fixed section contains stable genomic DNA sequence for a genomic region, essential transcripts and proteins for variant reporting and an exon numbering system. The updatable section contains mapping information, annotation of all transcripts and overlapping genes in the region and legacy exon and amino acid numbering systems. LRGs provide a stable framework that is vital for reporting variants, according to Human Genome Variation Society (HGVS) conventions, in genomic DNA, transcript or protein coordinates. To enable translation of information between LRG and genomic coordinates, LRGs include mapping to the human genome assembly. LRGs are compiled and maintained by the National Center for Biotechnology Information (NCBI) and European Bioinformatics Institute (EBI). LRG reference sequences are selected in collaboration with the diagnostic and research communities, locus-specific database curators and mutation consortia. Currently >700 LRGs have been created, of which >400 are publicly available. The aim is to create an LRG for every locus with clinical implications
@Article{24285302, author = {MacArthur JAL and Morales J and Tully RE and Astashyn A and Gil L and Bruford EA and Larsson P and Flicek P and Dalgleish R and Maglott DR and Cunningham F}, title = {Locus Reference Genomic: reference sequences for the reporting of clinically relevant sequence variants}, journal = {Nucleic Acids Res}, volume = {42}, number = {Database issue}, pages = {D873--8}, year = {2014}, doi = {10.1093/nar/gkt1198}, howpublished = {Advanced online publication: 26 November 2013}, abstract = {Locus Reference Genomic (LRG; http://www.lrg-sequence.org/) records contain internationally recognized stable reference sequences designed specifically for reporting clinically relevant sequence variants. Each LRG is contained within a single file consisting of a stable `fixed' section and a regularly updated `updatable' section. The fixed section contains stable genomic DNA sequence for a genomic region, essential transcripts and proteins for variant reporting and an exon numbering system. The updatable section contains mapping information, annotation of all transcripts and overlapping genes in the region and legacy exon and amino acid numbering systems. LRGs provide a stable framework that is vital for reporting variants, according to Human Genome Variation Society (HGVS) conventions, in genomic DNA, transcript or protein coordinates. To enable translation of information between LRG and genomic coordinates, LRGs include mapping to the human genome assembly. LRGs are compiled and maintained by the National Center for Biotechnology Information (NCBI) and European Bioinformatics Institute (EBI). LRG reference sequences are selected in collaboration with the diagnostic and research communities, locus-specific database curators and mutation consortia. Currently \>700 LRGs have been created, of which \>400 are publicly available. The aim is to create an LRG for every locus with clinical implications},} - B Ballester, A Medina-Rivera, D Schmidt, M Gonzàlez-Porta, M Carlucci, X Chen, K Chessman, AJ Faure, AP Funnell, A Goncalves, C Kutter, M Lukk, S Menon, WM McLaren, K Stefflova, S Watt, MT Weirauch, M Crossley, JC Marioni, DT Odom, P Flicek, MD Wilson. Multi-species, multi-transcription factor binding highlights conserved control of tissue-specific biological pathways. Elife 2014;3:e02626. doi:10.7554/eLife.02626
[BibTeX] [Abstract]
As exome sequencing gives way to genome sequencing, the need to interpret the function of regulatory DNA becomes increasingly important. To test whether evolutionary conservation of cis-regulatory modules (CRMs) gives insight into human gene regulation, we determined transcription factor (TF) binding locations of four liver-essential TFs in liver tissue from human, macaque, mouse, rat, and dog. Approximately, two thirds of the TF-bound regions fell into CRMs. Less than half of the human CRMs were found as a CRM in the orthologous region of a second species. Shared CRMs were associated with liver pathways and disease loci identified by genome-wide association studies. Recurrent rare human disease causing mutations at the promoters of several blood coagulation and lipid metabolism genes were also identified within CRMs shared in multiple species. This suggests that multi-species analyses of experimentally determined combinatorial TF binding will help identify genomic regions critical for tissue-specific gene control
@Article{25279814, author = {Ballester B and Medina-Rivera A and Schmidt D and Gonzàlez-Porta M and Carlucci M and Chen X and Chessman K and Faure AJ and Funnell AP and Goncalves A and Kutter C and Lukk M and Menon S and McLaren WM and Stefflova K and Watt S and Weirauch MT and Crossley M and Marioni JC and Odom DT and Flicek P and Wilson MD}, title = {Multi-species, multi-transcription factor binding highlights conserved control of tissue-specific biological pathways}, journal = {Elife}, volume = {3}, pages = {e02626}, year = {2014}, doi = {10.7554/eLife.02626}, abstract = {As exome sequencing gives way to genome sequencing, the need to interpret the function of regulatory DNA becomes increasingly important. To test whether evolutionary conservation of cis-regulatory modules (CRMs) gives insight into human gene regulation, we determined transcription factor (TF) binding locations of four liver-essential TFs in liver tissue from human, macaque, mouse, rat, and dog. Approximately, two thirds of the TF-bound regions fell into CRMs. Less than half of the human CRMs were found as a CRM in the orthologous region of a second species. Shared CRMs were associated with liver pathways and disease loci identified by genome-wide association studies. Recurrent rare human disease causing mutations at the promoters of several blood coagulation and lipid metabolism genes were also identified within CRMs shared in multiple species. This suggests that multi-species analyses of experimentally determined combinatorial TF binding will help identify genomic regions critical for tissue-specific gene control},} - MA Eckersley-Maslin, D Thybert, JH Bergmann, JC Marioni, P Flicek, DL Spector. Random Monoallelic Gene Expression Increases upon Embryonic Stem Cell Differentiation. Dev Cell 2014;28(4):351–365. doi:10.1016/j.devcel.2014.01.017
[BibTeX] [Abstract]
Random autosomal monoallelic gene expression refers to the transcription of a gene from one of two homologous alleles. We assessed the dynamics of monoallelic expression during development through an allele-specific RNA-sequencing screen in clonal populations of hybrid mouse embryonic stem cells (ESCs) and neural progenitor cells (NPCs). We identified 67 and 376 inheritable autosomal random monoallelically expressed genes in ESCs and NPCs, respectively, a 5.6-fold increase upon differentiation. Although DNA methylation and nuclear positioning did not distinguish the active and inactive alleles, specific histone modifications were differentially enriched between the two alleles. Interestingly, expression levels of 8\% of the monoallelically expressed genes remained similar between monoallelic and biallelic clones. These results support a model in which random monoallelic expression occurs stochastically during differentiation and, for some genes, is compensated for by the cell to maintain the required transcriptional output of these genes
@Article{24576421, author = {Eckersley-Maslin MA and Thybert D and Bergmann JH and Marioni JC and Flicek P and Spector DL}, title = {Random Monoallelic Gene Expression Increases upon Embryonic Stem Cell Differentiation}, journal = {Dev Cell}, volume = {28}, number = {4}, pages = {351--365}, year = {2014}, doi = {10.1016/j.devcel.2014.01.017}, abstract = {Random autosomal monoallelic gene expression refers to the transcription of a gene from one of two homologous alleles. We assessed the dynamics of monoallelic expression during development through an allele-specific RNA-sequencing screen in clonal populations of hybrid mouse embryonic stem cells (ESCs) and neural progenitor cells (NPCs). We identified 67 and 376 inheritable autosomal random monoallelically expressed genes in ESCs and NPCs, respectively, a 5.6-fold increase upon differentiation. Although DNA methylation and nuclear positioning did not distinguish the active and inactive alleles, specific histone modifications were differentially enriched between the two alleles. Interestingly, expression levels of 8\% of the monoallelically expressed genes remained similar between monoallelic and biallelic clones. These results support a model in which random monoallelic expression occurs stochastically during differentiation and, for some genes, is compensated for by the cell to maintain the required transcriptional output of these genes},} - Sequencing {Marmoset Genome, Consortium.} Analysis. The common marmoset genome provides insight into primate biology and evolution. Nat Genet 2014;46(8):850–857. doi:10.1038/ng.3042
[BibTeX] [Abstract]
We report the whole-genome sequence of the common marmoset (Callithrix jacchus). The 2.26-Gb genome of a female marmoset was assembled using Sanger read data (6×) and a whole-genome shotgun strategy. A first analysis has permitted comparison with the genomes of apes and Old World monkeys and the identification of specific features that might contribute to the unique biology of this diminutive primate, including genetic changes that may influence body size, frequent twinning and chimerism. We observed positive selection in growth hormone/insulin-like growth factor genes (growth pathways), respiratory complex I genes (metabolic pathways), and genes encoding immunobiological factors and proteases (reproductive and immunity pathways). In addition, both protein-coding and microRNA genes related to reproduction exhibited evidence of rapid sequence evolution. This genome sequence for a New World monkey enables increased power for comparative analyses among available primate genomes and facilitates biomedical research application
@Article{25038751, author = {{Marmoset Genome Sequencing and Analysis Consortium.}}, title = {The common marmoset genome provides insight into primate biology and evolution}, journal = {Nat Genet}, volume = {46}, number = {8}, pages = {850--857}, year = {2014}, doi = {10.1038/ng.3042}, abstract = {We report the whole-genome sequence of the common marmoset (Callithrix jacchus). The 2.26-Gb genome of a female marmoset was assembled using Sanger read data (6×) and a whole-genome shotgun strategy. A first analysis has permitted comparison with the genomes of apes and Old World monkeys and the identification of specific features that might contribute to the unique biology of this diminutive primate, including genetic changes that may influence body size, frequent twinning and chimerism. We observed positive selection in growth hormone/insulin-like growth factor genes (growth pathways), respiratory complex I genes (metabolic pathways), and genes encoding immunobiological factors and proteases (reproductive and immunity pathways). In addition, both protein-coding and microRNA genes related to reproduction exhibited evidence of rapid sequence evolution. This genome sequence for a New World monkey enables increased power for comparative analyses among available primate genomes and facilitates biomedical research application},} - G Koscielny, G Yaikhom, V Iyer, TF Meehan, H Morgan, J Atienza-Herrero, A Blake, C-K Chen, R Easty, A Di Fenza, T Fiegel, M Grifiths, A Horne, NA Karp, N Kurbatova, JC Mason, P Matthews, DJ Oakley, A Qazi, J Regnart, A Retha, LA Santos, DJ Sneddon, J Warren, H Westerberg, RJ Wilson, DG Melvin, D Smedley, SDM Brown, P Flicek, WC Skarnes, A-M Mallon, H Parkinson. The International Mouse Phenotyping Consortium Web Portal, a unified point of access for knockout mice and related phenotyping data. Nucleic Acids Res 2014;42(Database issue):D802–9. doi:10.1093/nar/gkt977
[BibTeX] [Abstract]
The International Mouse Phenotyping Consortium (IMPC) web portal (http://www.mousephenotype.org) provides the biomedical community with a unified point of access to mutant mice and rich collection of related emerging and existing mouse phenotype data. IMPC mouse clinics worldwide follow rigorous highly structured and standardized protocols for the experimentation, collection and dissemination of data. Dedicated `data wranglers' work with each phenotyping center to collate data and perform quality control of data. An automated statistical analysis pipeline has been developed to identify knockout strains with a significant change in the phenotype parameters. Annotation with biomedical ontologies allows biologists and clinicians to easily find mouse strains with phenotypic traits relevant to their research. Data integration with other resources will provide insights into mammalian gene function and human disease. As phenotype data become available for every gene in the mouse, the IMPC web portal will become an invaluable tool for researchers studying the genetic contributions of genes to human diseases
@Article{24194600, author = {Koscielny G and Yaikhom G and Iyer V and Meehan TF and Morgan H and Atienza-Herrero J and Blake A and Chen C-K and Easty R and Di Fenza A and Fiegel T and Grifiths M and Horne A and Karp NA and Kurbatova N and Mason JC and Matthews P and Oakley DJ and Qazi A and Regnart J and Retha A and Santos LA and Sneddon DJ and Warren J and Westerberg H and Wilson RJ and Melvin DG and Smedley D and Brown SDM and Flicek P and Skarnes WC and Mallon A-M and Parkinson H}, title = {The International Mouse Phenotyping Consortium Web Portal, a unified point of access for knockout mice and related phenotyping data}, journal = {Nucleic Acids Res}, volume = {42}, number = {Database issue}, pages = {D802--9}, year = {2014}, doi = {10.1093/nar/gkt977}, howpublished = {Advanced online publication: 4 November 2013}, abstract = {The International Mouse Phenotyping Consortium (IMPC) web portal (http://www.mousephenotype.org) provides the biomedical community with a unified point of access to mutant mice and rich collection of related emerging and existing mouse phenotype data. IMPC mouse clinics worldwide follow rigorous highly structured and standardized protocols for the experimentation, collection and dissemination of data. Dedicated `data wranglers' work with each phenotyping center to collate data and perform quality control of data. An automated statistical analysis pipeline has been developed to identify knockout strains with a significant change in the phenotype parameters. Annotation with biomedical ontologies allows biologists and clinicians to easily find mouse strains with phenotypic traits relevant to their research. Data integration with other resources will provide insights into mammalian gene function and human disease. As phenotype data become available for every gene in the mouse, the IMPC web portal will become an invaluable tool for researchers studying the genetic contributions of genes to human diseases},} - D Welter, J MacArthur, J Morales, T Burdett, P Hall, H Junkins, A Klemm, P Flicek, T Manolio, L Hindorff, H Parkinson. The NHGRI GWAS Catalog, a curated resource of SNP-trait associations. Nucleic Acids Res 2014;42(Database issue):D1001–1006. doi:10.1093/nar/gkt1229
[BibTeX] [Abstract]
The National Human Genome Research Institute (NHGRI) Catalog of Published Genome-Wide Association Studies (GWAS) Catalog provides a publicly available manually curated collection of published GWAS assaying at least 100,000 single-nucleotide polymorphisms (SNPs) and all SNP-trait associations with P <1 × 10(-5). The Catalog includes 1751 curated publications of 11 912 SNPs. In addition to the SNP-trait association data, the Catalog also publishes a quarterly diagram of all SNP-trait associations mapped to the SNPs' chromosomal locations. The Catalog can be accessed via a tabular web interface, via a dynamic visualization on the human karyotype, as a downloadable tab-delimited file and as an OWL knowledge base. This article presents a number of recent improvements to the Catalog, including novel ways for users to interact with the Catalog and changes to the curation infrastructure
@Article{24316577, author = {Welter D and MacArthur J and Morales J and Burdett T and Hall P and Junkins H and Klemm A and Flicek P and Manolio T and Hindorff L and Parkinson H}, title = {The NHGRI GWAS Catalog, a curated resource of SNP-trait associations}, journal = {Nucleic Acids Res}, volume = {42}, number = {Database issue}, pages = {D1001--1006}, year = {2014}, doi = {10.1093/nar/gkt1229}, howpublished = {Advanced online publication: 6 December 2013}, abstract = {The National Human Genome Research Institute (NHGRI) Catalog of Published Genome-Wide Association Studies (GWAS) Catalog provides a publicly available manually curated collection of published GWAS assaying at least 100,000 single-nucleotide polymorphisms (SNPs) and all SNP-trait associations with P \<1 × 10(-5). The Catalog includes 1751 curated publications of 11 912 SNPs. In addition to the SNP-trait association data, the Catalog also publishes a quarterly diagram of all SNP-trait associations mapped to the SNPs' chromosomal locations. The Catalog can be accessed via a tabular web interface, via a dynamic visualization on the human karyotype, as a downloadable tab-delimited file and as an OWL knowledge base. This article presents a number of recent improvements to the Catalog, including novel ways for users to interact with the Catalog and changes to the curation infrastructure},} - Y Jiang, M Xie, W Chen, R Talbot, JF Maddox, T Faraut, C Wu, DM Muzny, Y Li, W Zhang, J-A Stanton, R Brauning, WC Barris, T Hourlier, BL Aken, SMJ Searle, DL Adelson, C Bian, GR Cam, Y Chen, S Cheng, U DeSilva, K Dixen, Y Dong, G Fan, IR Franklin, S Fu, P Fuentes-Utrilla, R Guan, MA Highland, ME Holder, G Huang, AB Ingham, SN Jhangiani, D Kalra, CL Kovar, SL Lee, W Liu, X Liu, C Lu, T Lv, T Mathew, S McWilliam, M Menzies, S Pan, D Robelin, B Servin, D Townley, W Wang, B Wei, SN White, X Yang, C Ye, Y Yue, P Zeng, Q Zhou, JB Hansen, K Kristiansen, RA Gibbs, P Flicek, CC Warkup, HE Jones, VH Oddy, FW Nicholas, JC McEwan, JW Kijas, J Wang, KC Worley, AL Archibald, N Cockett, X Xu, W Wang, BP Dalrymple. The sheep genome illuminates biology of the rumen and lipid metabolism. Science 2014;344(6188):1168–1173. doi:10.1126/science.1252806
[BibTeX] [Abstract]
Sheep (Ovis aries) are a major source of meat, milk, and fiber in the form of wool and represent a distinct class of animals that have a specialized digestive organ, the rumen, that carries out the initial digestion of plant material. We have developed and analyzed a high-quality reference sheep genome and transcriptomes from 40 different tissues. We identified highly expressed genes encoding keratin cross-linking proteins associated with rumen evolution. We also identified genes involved in lipid metabolism that had been amplified and/or had altered tissue expression patterns. This may be in response to changes in the barrier lipids of the skin, an interaction between lipid metabolism and wool synthesis, and an increased role of volatile fatty acids in ruminants compared with nonruminant animals
@Article{24904168, author = {Jiang Y and Xie M and Chen W and Talbot R and Maddox JF and Faraut T and Wu C and Muzny DM and Li Y and Zhang W and Stanton J-A and Brauning R and Barris WC and Hourlier T and Aken BL and Searle SMJ and Adelson DL and Bian C and Cam GR and Chen Y and Cheng S and DeSilva U and Dixen K and Dong Y and Fan G and Franklin IR and Fu S and Fuentes-Utrilla P and Guan R and Highland MA and Holder ME and Huang G and Ingham AB and Jhangiani SN and Kalra D and Kovar CL and Lee SL and Liu W and Liu X and Lu C and Lv T and Mathew T and McWilliam S and Menzies M and Pan S and Robelin D and Servin B and Townley D and Wang W and Wei B and White SN and Yang X and Ye C and Yue Y and Zeng P and Zhou Q and Hansen JB and Kristiansen K and Gibbs RA and Flicek P and Warkup CC and Jones HE and Oddy VH and Nicholas FW and McEwan JC and Kijas JW and Wang J and Worley KC and Archibald AL and Cockett N and Xu X and Wang W and Dalrymple BP}, title = {The sheep genome illuminates biology of the rumen and lipid metabolism}, journal = {Science}, volume = {344}, number = {6188}, pages = {1168--1173}, year = {2014}, doi = {10.1126/science.1252806}, abstract = {Sheep (Ovis aries) are a major source of meat, milk, and fiber in the form of wool and represent a distinct class of animals that have a specialized digestive organ, the rumen, that carries out the initial digestion of plant material. We have developed and analyzed a high-quality reference sheep genome and transcriptomes from 40 different tissues. We identified highly expressed genes encoding keratin cross-linking proteins associated with rumen evolution. We also identified genes involved in lipid metabolism that had been amplified and/or had altered tissue expression patterns. This may be in response to changes in the barrier lipids of the skin, an interaction between lipid metabolism and wool synthesis, and an increased role of volatile fatty acids in ruminants compared with nonruminant animals},} - L Chen, M Kostadima, JHA Martens, G Canu, SP Garcia, E Turro, K Downes, IC Macaulay, E Bielczyk-Maczynska, S Coe, S Farrow, P Poudel, F Burden, SBG Jansen, WJ Astle, A Attwood, T Bariana, de B Bono, A Breschi, JC Chambers, FA Choudry, L Clarke, P Coupland, van der M Ent, WN Erber, JH Jansen, R Favier, ME Fenech, N Foad, K Freson, van C Geet, K Gomez, R Guigo, D Hampshire, AM Kelly, HHD Kerstens, JS Kooner, M Laffan, C Lentaigne, C Labalette, T Martin, S Meacham, A Mumford, S Nürnberg, E Palumbo, van der BA Reijden, D Richardson, SJ Sammut, G Slodkowicz, AU Tamuri, L Vasquez, K Voss, S Watt, S Westbury, P Flicek, R Loos, N Goldman, P Bertone, RJ Read, S Richardson, A Cvejic, N Soranzo, WH Ouwehand, HG Stunnenberg, M Frontini, A Rendon. Transcriptional diversity during lineage commitment of human blood progenitors. Science 2014;345(6204):1251033. doi:10.1126/science.1251033
[BibTeX] [Abstract]
Blood cells derive from hematopoietic stem cells through stepwise fating events. To characterize gene expression programs driving lineage choice, we sequenced RNA from eight primary human hematopoietic progenitor populations representing the major myeloid commitment stages and the main lymphoid stage. We identified extensive cell type-specific expression changes: 6711 genes and 10,724 transcripts, enriched in non-protein-coding elements at early stages of differentiation. In addition, we found 7881 novel splice junctions and 2301 differentially used alternative splicing events, enriched in genes involved in regulatory processes. We demonstrated experimentally cell-specific isoform usage, identifying nuclear factor I/B (NFIB) as a regulator of megakaryocyte maturation-the platelet precursor. Our data highlight the complexity of fating events in closely related progenitor populations, the understanding of which is essential for the advancement of transplantation and regenerative medicine
@Article{25258084, author = {Chen L and Kostadima M and Martens JHA and Canu G and Garcia SP and Turro E and Downes K and Macaulay IC and Bielczyk-Maczynska E and Coe S and Farrow S and Poudel P and Burden F and Jansen SBG and Astle WJ and Attwood A and Bariana T and de Bono B and Breschi A and Chambers JC and Choudry FA and Clarke L and Coupland P and van der Ent M and Erber WN and Jansen JH and Favier R and Fenech ME and Foad N and Freson K and van Geet C and Gomez K and Guigo R and Hampshire D and Kelly AM and Kerstens HHD and Kooner JS and Laffan M and Lentaigne C and Labalette C and Martin T and Meacham S and Mumford A and Nürnberg S and Palumbo E and van der Reijden BA and Richardson D and Sammut SJ and Slodkowicz G and Tamuri AU and Vasquez L and Voss K and Watt S and Westbury S and Flicek P and Loos R and Goldman N and Bertone P and Read RJ and Richardson S and Cvejic A and Soranzo N and Ouwehand WH and Stunnenberg HG and Frontini M and Rendon A}, title = {Transcriptional diversity during lineage commitment of human blood progenitors}, journal = {Science}, volume = {345}, number = {6204}, pages = {1251033}, year = {2014}, doi = {10.1126/science.1251033}, abstract = {Blood cells derive from hematopoietic stem cells through stepwise fating events. To characterize gene expression programs driving lineage choice, we sequenced RNA from eight primary human hematopoietic progenitor populations representing the major myeloid commitment stages and the main lymphoid stage. We identified extensive cell type-specific expression changes: 6711 genes and 10,724 transcripts, enriched in non-protein-coding elements at early stages of differentiation. In addition, we found 7881 novel splice junctions and 2301 differentially used alternative splicing events, enriched in genes involved in regulatory processes. We demonstrated experimentally cell-specific isoform usage, identifying nuclear factor I/B (NFIB) as a regulator of megakaryocyte maturation-the platelet precursor. Our data highlight the complexity of fating events in closely related progenitor populations, the understanding of which is essential for the advancement of transplantation and regenerative medicine},} - DR Zerbino, N Johnson, T Juettemann, SP Wilder, P Flicek. WiggleTools: parallel processing of large collections of genome-wide datasets for visualization and statistical analysis. Bioinformatics 2014;30(7):1008–1009. doi:10.1093/bioinformatics/btt737
[BibTeX] [Abstract]
\textbf{MOTIVATION:} Using high-throughput sequencing, researchers are now generating hundreds of whole-genome assays to measure various features such as transcription factor binding, histone marks, DNA methylation or RNA transcription. Displaying so much data generally leads to a confusing accumulation of plots. We describe here a multithreaded library that computes statistics on large numbers of datasets (Wiggle, BigWig, Bed, BigBed and BAM), generating statistical summaries within minutes with limited memory requirements, whether on the whole genome or on selected regions.
\textbf{AVAILABILITY AND IMPLEMENTATION:} The code is freely available under Apache 2.0 license at www.github.com/Ensembl/Wiggletools
\textbf{CONTACT:} zerbino@ebi.ac.uk or flicek@ebi.ac.uk@Article{24363377, author = {Zerbino DR and Johnson N and Juettemann T and Wilder SP and Flicek P}, title = {WiggleTools: parallel processing of large collections of genome-wide datasets for visualization and statistical analysis}, journal = {Bioinformatics}, volume = {30}, number = {7}, pages = {1008--1009}, year = {2014}, doi = {10.1093/bioinformatics/btt737}, howpublished = {Advanced online publication: 19 December 2013}, abstract = {\textbf{MOTIVATION:} Using high-throughput sequencing, researchers are now generating hundreds of whole-genome assays to measure various features such as transcription factor binding, histone marks, DNA methylation or RNA transcription. Displaying so much data generally leads to a confusing accumulation of plots. We describe here a multithreaded library that computes statistics on large numbers of datasets (Wiggle, BigWig, Bed, BigBed and BAM), generating statistical summaries within minutes with limited memory requirements, whether on the whole genome or on selected regions.
\textbf{AVAILABILITY AND IMPLEMENTATION:} The code is freely available under Apache 2.0 license at www.github.com/Ensembl/Wiggletools
\textbf{CONTACT:} zerbino@ebi.ac.uk or flicek@ebi.ac.uk},}
2013
- APW Funnell, MD Wilson, B Ballester, KS Mak, J Burdach, N Magan, RCM Pearson, FP Lemaigre, KM Stowell, DT Odom, P Flicek, M Crossley. A CpG Mutational Hotspot in a ONECUT Binding Site Accounts for the Prevalent Variant of Hemophilia B Leyden. Am J Hum Genet 2013;92(3):460–467. doi:10.1016/j.ajhg.2013.02.003
[BibTeX] [Abstract]
Hemophilia B, or the ``royal disease,'' arises from mutations in coagulation factor IX (F9). Mutations within the F9 promoter are associated with a remarkable hemophilia B subtype, termed hemophilia B Leyden, in which symptoms ameliorate after puberty. Mutations at the -5/-6 site (nucleotides -5 and -6 relative to the transcription start site, designated +1) account for the majority of Leyden cases and have been postulated to disrupt the binding of a transcriptional activator, the identity of which has remained elusive for more than 20 years. Here, we show that ONECUT transcription factors (ONECUT1 and ONECUT2) bind to the -5/-6 site. The various hemophilia B Leyden mutations that have been reported in this site inhibit ONECUT binding to varying degrees, which correlate well with their associated clinical severities. In addition, expression of F9 is crucially dependent on ONECUT factors in vivo, and as such, mice deficient in ONECUT1, ONECUT2, or both exhibit depleted levels of F9. Taken together, our findings establish ONECUT transcription factors as the missing hemophilia B Leyden regulators that operate through the -5/-6 site
@Article{23472758, author = {Funnell APW and Wilson MD and Ballester B and Mak KS and Burdach J and Magan N and Pearson RCM and Lemaigre FP and Stowell KM and Odom DT and Flicek P and Crossley M}, title = {A CpG Mutational Hotspot in a ONECUT Binding Site Accounts for the Prevalent Variant of Hemophilia B Leyden}, journal = {Am J Hum Genet}, volume = {92}, number = {3}, pages = {460--467}, year = {2013}, doi = {10.1016/j.ajhg.2013.02.003}, abstract = {Hemophilia B, or the ``royal disease,'' arises from mutations in coagulation factor IX (F9). Mutations within the F9 promoter are associated with a remarkable hemophilia B subtype, termed hemophilia B Leyden, in which symptoms ameliorate after puberty. Mutations at the -5/-6 site (nucleotides -5 and -6 relative to the transcription start site, designated +1) account for the majority of Leyden cases and have been postulated to disrupt the binding of a transcriptional activator, the identity of which has remained elusive for more than 20 years. Here, we show that ONECUT transcription factors (ONECUT1 and ONECUT2) bind to the -5/-6 site. The various hemophilia B Leyden mutations that have been reported in this site inhibit ONECUT binding to varying degrees, which correlate well with their associated clinical severities. In addition, expression of F9 is crucially dependent on ONECUT factors in vivo, and as such, mice deficient in ONECUT1, ONECUT2, or both exhibit depleted levels of F9. Taken together, our findings establish ONECUT transcription factors as the missing hemophilia B Leyden regulators that operate through the -5/-6 site},} - T Schauer, PC Schwalie, A Handley, CE Margulies, P Flicek, AG Ladurner. CAST-ChIP Maps Cell-Type-Specific Chromatin States in the Drosophila Central Nervous System. Cell Rep 2013;5(1):271–282. doi:10.1016/j.celrep.2013.09.001
[BibTeX] [Abstract]
Chromatin organization and gene activity are responsive to developmental and environmental cues. Although many genes are transcribed throughout development and across cell types, much of gene regulation is highly cell-type specific. To readily track chromatin features at the resolution of cell types within complex tissues, we developed and validated chromatin affinity purification from specific cell types by chromatin immunoprecipitation (CAST-ChIP), a broadly applicable biochemical procedure. RNA polymerase II (Pol II) CAST-ChIP identifies ∼1,500 neuronal and glia-specific genes in differentiated cells within the adult Drosophila brain. In contrast, the histone H2A.Z is distributed similarly across cell types and throughout development, marking cell-type-invariant Pol II-bound regions. Our study identifies H2A.Z as an active chromatin signature that is refractory to changes across cell fates. Thus, CAST-ChIP powerfully identifies cell-type-specific as well as cell-type-invariant chromatin states, enabling the systematic dissection of chromatin structure and gene regulation within complex tissues such as the brain
@Article{24095734, author = {Schauer T and Schwalie PC and Handley A and Margulies CE and Flicek P and Ladurner AG}, title = {CAST-ChIP Maps Cell-Type-Specific Chromatin States in the Drosophila Central Nervous System}, journal = {Cell Rep}, volume = {5}, number = {1}, pages = {271--282}, year = {2013}, doi = {10.1016/j.celrep.2013.09.001}, howpublished = {Advanced online publication: 3 October 2013}, abstract = {Chromatin organization and gene activity are responsive to developmental and environmental cues. Although many genes are transcribed throughout development and across cell types, much of gene regulation is highly cell-type specific. To readily track chromatin features at the resolution of cell types within complex tissues, we developed and validated chromatin affinity purification from specific cell types by chromatin immunoprecipitation (CAST-ChIP), a broadly applicable biochemical procedure. RNA polymerase II (Pol II) CAST-ChIP identifies ∼1,500 neuronal and glia-specific genes in differentiated cells within the adult Drosophila brain. In contrast, the histone H2A.Z is distributed similarly across cell types and throughout development, marking cell-type-invariant Pol II-bound regions. Our study identifies H2A.Z as an active chromatin signature that is refractory to changes across cell fates. Thus, CAST-ChIP powerfully identifies cell-type-specific as well as cell-type-invariant chromatin states, enabling the systematic dissection of chromatin structure and gene regulation within complex tissues such as the brain},} - PC Schwalie, MC Ward, CE Cain, AJ Faure, Y Gilad, DT Odom, P Flicek. Co-binding by YY1 identifies the transcriptionally active, highly conserved set of CTCF-bound regions in primate genomes. Genome Biol 2013;14(12):R148. doi:10.1186/gb-2013-14-12-r148
[BibTeX] [Abstract]
\textbf{BACKGROUND:} The genomic binding of CTCF is highly conserved across mammals, but the mechanisms that underlie its stability are poorly understood. One transcription factor known to functionally interact with CTCF in the context of X-chromosome inactivation is the ubiquitously expressed YY1. Because combinatorial transcription factor binding can contribute to the evolutionary stabilization of regulatory regions, we tested whether YY1 and CTCF co-binding could in part account for conservation of CTCF binding.
\textbf{RESULTS:} Combined analysis of CTCF and YY1 binding in lymphoblastoid cell lines from seven primates, as well as in mouse and human livers, reveals extensive genome-wide co-localization specifically at evolutionarily stable CTCF-bound regions. CTCF-YY1 co-bound regions resemble regions bound by YY1 alone, as they enrich for co-bound transcription factors, RNA polymerase II and active histone marks. Although these highly conserved, transcriptionally active CTCF-YY1 co-bound regions are often promoter-proximal, gene-distal sites show similar molecular features.
\textbf{CONCLUSIONS:} Our results reveal that these two ubiquitously expressed, multi-functional zinc-finger proteins collaborate in functionally active regions to stabilize one another's genome-wide binding across primate evolution@Article{24380390, author = {Schwalie PC and Ward MC and Cain CE and Faure AJ and Gilad Y and Odom DT and Flicek P}, title = {Co-binding by YY1 identifies the transcriptionally active, highly conserved set of CTCF-bound regions in primate genomes}, journal = {Genome Biol}, volume = {14}, number = {12}, pages = {R148}, year = {2013}, doi = {10.1186/gb-2013-14-12-r148}, abstract = {\textbf{BACKGROUND:} The genomic binding of CTCF is highly conserved across mammals, but the mechanisms that underlie its stability are poorly understood. One transcription factor known to functionally interact with CTCF in the context of X-chromosome inactivation is the ubiquitously expressed YY1. Because combinatorial transcription factor binding can contribute to the evolutionary stabilization of regulatory regions, we tested whether YY1 and CTCF co-binding could in part account for conservation of CTCF binding.
\textbf{RESULTS:} Combined analysis of CTCF and YY1 binding in lymphoblastoid cell lines from seven primates, as well as in mouse and human livers, reveals extensive genome-wide co-localization specifically at evolutionarily stable CTCF-bound regions. CTCF-YY1 co-bound regions resemble regions bound by YY1 alone, as they enrich for co-bound transcription factors, RNA polymerase II and active histone marks. Although these highly conserved, transcriptionally active CTCF-YY1 co-bound regions are often promoter-proximal, gene-distal sites show similar molecular features.
\textbf{CONCLUSIONS:} Our results reveal that these two ubiquitously expressed, multi-functional zinc-finger proteins collaborate in functionally active regions to stabilize one another's genome-wide binding across primate evolution},} - VC Seitan, AJ Faure, Y Zhan, RP McCord, BR Lajoie, E Ing-Simmons, B Lenhard, L Giorgetti, E Heard, AG Fisher, P Flicek, J Dekker, M Merkenschlager. Cohesin-based chromatin interactions enable regulated gene expression within preexisting architectural compartments. Genome Res 2013;23(12):2066–2077. doi:10.1101/gr.161620.113
[BibTeX] [Abstract]
Chromosome conformation capture approaches have shown that interphase chromatin is partitioned into spatially segregated Mb-sized compartments and sub-Mb-sized topological domains. This compartmentalization is thought to facilitate the matching of genes and regulatory elements, but its precise function and mechanistic basis remain unknown. Cohesin controls chromosome topology to enable DNA repair and chromosome segregation in cycling cells. In addition, cohesin associates with active enhancers and promoters and with CTCF to form long-range interactions important for gene regulation. Although these findings suggest an important role for cohesin in genome organization, this role has not been assessed on a global scale. Unexpectedly, we find that architectural compartments are maintained in noncycling mouse thymocytes after genetic depletion of cohesin in vivo. Cohesin was, however, required for specific long-range interactions within compartments where cohesin-regulated genes reside. Cohesin depletion diminished interactions between cohesin-bound sites, whereas alternative interactions between chromatin features associated with transcriptional activation and repression became more prominent, with corresponding changes in gene expression. Our findings indicate that cohesin-mediated long-range interactions facilitate discrete gene expression states within preexisting chromosomal compartments
@Article{24002784, author = {Seitan VC and Faure AJ and Zhan Y and McCord RP and Lajoie BR and Ing-Simmons E and Lenhard B and Giorgetti L and Heard E and Fisher AG and Flicek P and Dekker J and Merkenschlager M}, title = {Cohesin-based chromatin interactions enable regulated gene expression within preexisting architectural compartments}, journal = {Genome Res}, volume = {23}, number = {12}, pages = {2066--2077}, year = {2013}, doi = {10.1101/gr.161620.113}, howpublished = {Advanced online publication: 3 September 2013}, abstract = {Chromosome conformation capture approaches have shown that interphase chromatin is partitioned into spatially segregated Mb-sized compartments and sub-Mb-sized topological domains. This compartmentalization is thought to facilitate the matching of genes and regulatory elements, but its precise function and mechanistic basis remain unknown. Cohesin controls chromosome topology to enable DNA repair and chromosome segregation in cycling cells. In addition, cohesin associates with active enhancers and promoters and with CTCF to form long-range interactions important for gene regulation. Although these findings suggest an important role for cohesin in genome organization, this role has not been assessed on a global scale. Unexpectedly, we find that architectural compartments are maintained in noncycling mouse thymocytes after genetic depletion of cohesin in vivo. Cohesin was, however, required for specific long-range interactions within compartments where cohesin-regulated genes reside. Cohesin depletion diminished interactions between cohesin-bound sites, whereas alternative interactions between chromatin features associated with transcriptional activation and repression became more prominent, with corresponding changes in gene expression. Our findings indicate that cohesin-mediated long-range interactions facilitate discrete gene expression states within preexisting chromosomal compartments},} - A Baud, R Hermsen, V Guryev, P Stridh, D Graham, MW McBride, T Foroud, S Calderari, M Diez, J Ockinger, AD Beyeen, A Gillett, N Abdelmagid, AO Guerreiro-Cacais, M Jagodic, J Tuncel, U Norin, E Beattie, N Huynh, WH Miller, DL Koller, I Alam, S Falak, M Osborne-Pellegrin, E Martinez-Membrives, T Canete, G Blazquez, E Vicens-Costa, C Mont-Cardona, S Diaz-Moran, A Tobena, O Hummel, D Zelenika, K Saar, G Patone, A Bauerfeind, M-T Bihoreau, M Heinig, Y-A Lee, C Rintisch, H Schulz, DA Wheeler, KC Worley, DM Muzny, RA Gibbs, M Lathrop, N Lansu, P Toonen, FP Ruzius, de E Bruijn, H Hauser, DJ Adams, T Keane, SS Atanur, TJ Aitman, P Flicek, T Malinauskas, EY Jones, D Ekman, R Lopez-Aumatell, AF Dominiczak, M Johannesson, R Holmdahl, T Olsson, D Gauguier, N Hubner, A Fernandez-Teruel, E Cuppen, R Mott, J Flint. Combined sequence-based and genetic mapping analysis of complex traits in outbred rats. Nat Genet 2013;45(7):767–775. doi:10.1038/ng.2644
[BibTeX] [Abstract]
Genetic mapping on fully sequenced individuals is transforming understanding of the relationship between molecular variation and variation in complex traits. Here we report a combined sequence and genetic mapping analysis in outbred rats that maps 355 quantitative trait loci for 122 phenotypes. We identify 35 causal genes involved in 31 phenotypes, implicating new genes in models of anxiety, heart disease and multiple sclerosis. The relationship between sequence and genetic variation is unexpectedly complex: at approximately 40\% of quantitative trait loci, a single sequence variant cannot account for the phenotypic effect. Using comparable sequence and mapping data from mice, we show that the extent and spatial pattern of variation in inbred rats differ substantially from those of inbred mice and that the genetic variants in orthologous genes rarely contribute to the same phenotype in both species
@Article{23708188, author = {Baud A and Hermsen R and Guryev V and Stridh P and Graham D and McBride MW and Foroud T and Calderari S and Diez M and Ockinger J and Beyeen AD and Gillett A and Abdelmagid N and Guerreiro-Cacais AO and Jagodic M and Tuncel J and Norin U and Beattie E and Huynh N and Miller WH and Koller DL and Alam I and Falak S and Osborne-Pellegrin M and Martinez-Membrives E and Canete T and Blazquez G and Vicens-Costa E and Mont-Cardona C and Diaz-Moran S and Tobena A and Hummel O and Zelenika D and Saar K and Patone G and Bauerfeind A and Bihoreau M-T and Heinig M and Lee Y-A and Rintisch C and Schulz H and Wheeler DA and Worley KC and Muzny DM and Gibbs RA and Lathrop M and Lansu N and Toonen P and Ruzius FP and de Bruijn E and Hauser H and Adams DJ and Keane T and Atanur SS and Aitman TJ and Flicek P and Malinauskas T and Jones EY and Ekman D and Lopez-Aumatell R and Dominiczak AF and Johannesson M and Holmdahl R and Olsson T and Gauguier D and Hubner N and Fernandez-Teruel A and Cuppen E and Mott R and Flint J}, title = {Combined sequence-based and genetic mapping analysis of complex traits in outbred rats}, journal = {Nat Genet}, volume = {45}, number = {7}, pages = {767--775}, year = {2013}, doi = {10.1038/ng.2644}, abstract = {Genetic mapping on fully sequenced individuals is transforming understanding of the relationship between molecular variation and variation in complex traits. Here we report a combined sequence and genetic mapping analysis in outbred rats that maps 355 quantitative trait loci for 122 phenotypes. We identify 35 causal genes involved in 31 phenotypes, implicating new genes in models of anxiety, heart disease and multiple sclerosis. The relationship between sequence and genetic variation is unexpectedly complex: at approximately 40\% of quantitative trait loci, a single sequence variant cannot account for the phenotypic effect. Using comparable sequence and mapping data from mice, we show that the extent and spatial pattern of variation in inbred rats differ substantially from those of inbred mice and that the genetic variants in orthologous genes rarely contribute to the same phenotype in both species},} - A Gonzalez-Perez, V Mustonen, B Reva, GRS Ritchie, P Creixell, R Karchin, M Vazquez, JL Fink, KS Kassahn, JV Pearson, GD Bader, PC Boutros, L Muthuswamy, BFF Ouellette, J Reimand, R Linding, T Shibata, A Valencia, A Butler, S Dronov, P Flicek, NB Shannon, H Carter, L Ding, C Sander, JM Stuart, LD Stein, N Lopez-Bigas. Computational approaches to identify functional genetic variants in cancer genomes. Nat Methods 2013;10(8):723–729. doi:10.1038/nmeth.2562
[BibTeX] [Abstract]
The International Cancer Genome Consortium (ICGC) aims to catalog genomic abnormalities in tumors from 50 different cancer types. Genome sequencing reveals hundreds to thousands of somatic mutations in each tumor but only a minority of these drive tumor progression. We present the result of discussions within the ICGC on how to address the challenge of identifying mutations that contribute to oncogenesis, tumor maintenance or response to therapy, and recommend computational techniques to annotate somatic variants and predict their impact on cancer phenotype
@Article{23900255, author = {Gonzalez-Perez A and Mustonen V and Reva B and Ritchie GRS and Creixell P and Karchin R and Vazquez M and Fink JL and Kassahn KS and Pearson JV and Bader GD and Boutros PC and Muthuswamy L and Ouellette BFF and Reimand J and Linding R and Shibata T and Valencia A and Butler A and Dronov S and Flicek P and Shannon NB and Carter H and Ding L and Sander C and Stuart JM and Stein LD and Lopez-Bigas N}, title = {Computational approaches to identify functional genetic variants in cancer genomes}, journal = {Nat Methods}, volume = {10}, number = {8}, pages = {723--729}, year = {2013}, doi = {10.1038/nmeth.2562}, abstract = {The International Cancer Genome Consortium (ICGC) aims to catalog genomic abnormalities in tumors from 50 different cancer types. Genome sequencing reveals hundreds to thousands of somatic mutations in each tumor but only a minority of these drive tumor progression. We present the result of discussions within the ICGC on how to address the challenge of identifying mutations that contribute to oncogenesis, tumor maintenance or response to therapy, and recommend computational techniques to annotate somatic variants and predict their impact on cancer phenotype},} - K Stefflova, D Thybert, MD Wilson, I Streeter, J Aleksic, P Karagianni, A Brazma, DJ Adams, I Talianidis, JC Marioni, P Flicek, DT Odom. Cooperativity and rapid evolution of cobound transcription factors in closely related mammals. Cell 2013;154(3):530–540. doi:10.1016/j.cell.2013.07.007
[BibTeX] [Abstract]
To mechanistically characterize the microevolutionary processes active in altering transcription factor (TF) binding among closely related mammals, we compared the genome-wide binding of three tissue-specific TFs that control liver gene expression in six rodents. Despite an overall fast turnover of TF binding locations between species, we identified thousands of TF regions of highly constrained TF binding intensity. Although individual mutations in bound sequence motifs can influence TF binding, most binding differences occur in the absence of nearby sequence variations. Instead, combinatorial binding was found to be significant for genetic and evolutionary stability; cobound TFs tend to disappear in concert and were sensitive to genetic knockout of partner TFs. The large, qualitative differences in genomic regions bound between closely related mammals, when contrasted with the smaller, quantitative TF binding differences among Drosophila species, illustrate how genome structure and population genetics together shape regulatory evolution
@Article{23911320, author = {Stefflova K and Thybert D and Wilson MD and Streeter I and Aleksic J and Karagianni P and Brazma A and Adams DJ and Talianidis I and Marioni JC and Flicek P and Odom DT}, title = {Cooperativity and rapid evolution of cobound transcription factors in closely related mammals}, journal = {Cell}, volume = {154}, number = {3}, pages = {530--540}, year = {2013}, doi = {10.1016/j.cell.2013.07.007}, abstract = {To mechanistically characterize the microevolutionary processes active in altering transcription factor (TF) binding among closely related mammals, we compared the genome-wide binding of three tissue-specific TFs that control liver gene expression in six rodents. Despite an overall fast turnover of TF binding locations between species, we identified thousands of TF regions of highly constrained TF binding intensity. Although individual mutations in bound sequence motifs can influence TF binding, most binding differences occur in the absence of nearby sequence variations. Instead, combinatorial binding was found to be significant for genetic and evolutionary stability; cobound TFs tend to disappear in concert and were sensitive to genetic knockout of partner TFs. The large, qualitative differences in genomic regions bound between closely related mammals, when contrasted with the smaller, quantitative TF binding differences among Drosophila species, illustrate how genome structure and population genetics together shape regulatory evolution},} - I Lappalainen, J Lopez, L Skipper, T Hefferon, JD Spalding, J Garner, C Chen, M Maguire, M Corbett, G Zhou, J Paschall, V Ananiev, P Flicek, DM Church. dbVar and DGVa: public archives for genomic structural variation. Nucleic Acids Res 2013;41(D1):D936–41. doi:10.1093/nar/gks1213
[BibTeX] [Abstract]
Much has changed in the last two years at DGVa (http://www.ebi.ac.uk/dgva) and dbVar (http://www.ncbi.nlm.nih.gov/dbvar). We are now processing direct submissions rather than only curating data from the literature and our joint study catalog includes data from over 100 studies in 11 organisms. Studies from human dominate with data from control and case populations, tumor samples as well as three large curated studies derived from multiple sources. During the processing of these data, we have made improvements to our data model, submission process and data representation. Additionally, we have made significant improvements in providing access to these data via web and FTP interfaces
@Article{23193291, author = {Lappalainen I and Lopez J and Skipper L and Hefferon T and Spalding JD and Garner J and Chen C and Maguire M and Corbett M and Zhou G and Paschall J and Ananiev V and Flicek P and Church DM}, title = {dbVar and DGVa: public archives for genomic structural variation}, journal = {Nucleic Acids Res}, volume = {41}, number = {D1}, pages = {D936--41}, year = {2013}, doi = {10.1093/nar/gks1213}, howpublished = {Advanced online publication: 27 November 2012}, abstract = {Much has changed in the last two years at DGVa (http://www.ebi.ac.uk/dgva) and dbVar (http://www.ncbi.nlm.nih.gov/dbvar). We are now processing direct submissions rather than only curating data from the literature and our joint study catalog includes data from over 100 studies in 11 organisms. Studies from human dominate with data from control and case populations, tumor samples as well as three large curated studies derived from multiple sources. During the processing of these data, we have made improvements to our data model, submission process and data representation. Additionally, we have made significant improvements in providing access to these data via web and FTP interfaces},} - P Flicek, I Ahmed, MR Amode, D Barrell, K Beal, S Brent, D Carvalho-Silva, P Clapham, G Coates, S Fairley, S Fitzgerald, L Gil, C García-Girón, L Gordon, T Hourlier, S Hunt, T Juettemann, AK Kähäri, S Keenan, M Komorowska, E Kulesha, I Longden, T Maurel, WM McLaren, M Muffato, R Nag, B Overduin, M Pignatelli, B Pritchard, E Pritchard, HS Riat, GRS Ritchie, M Ruffier, M Schuster, D Sheppard, D Sobral, K Taylor, A Thormann, S Trevanion, S White, SP Wilder, BL Aken, E Birney, F Cunningham, I Dunham, J Harrow, J Herrero, TJP Hubbard, N Johnson, R Kinsella, A Parker, G Spudich, A Yates, A Zadissa, SMJ Searle. Ensembl 2013. Nucleic Acids Res 2013;41(D1):D48–55. doi:10.1093/nar/gks1236
[BibTeX] [Abstract]
The Ensembl project (http://www.ensembl.org) provides genome information for sequenced chordate genomes with a particular focus on human, mouse, zebrafish and rat. Our resources include evidenced-based gene sets for all supported species; large-scale whole genome multiple species alignments across vertebrates and clade-specific alignments for eutherian mammals, primates, birds and fish; variation data resources for 17 species and regulation annotations based on ENCODE and other data sets. Ensembl data are accessible through the genome browser at http://www.ensembl.org and through other tools and programmatic interfaces
@Article{23203987, author = {Flicek P and Ahmed I and Amode MR and Barrell D and Beal K and Brent S and Carvalho-Silva D and Clapham P and Coates G and Fairley S and Fitzgerald S and Gil L and García-Girón C and Gordon L and Hourlier T and Hunt S and Juettemann T and Kähäri AK and Keenan S and Komorowska M and Kulesha E and Longden I and Maurel T and McLaren WM and Muffato M and Nag R and Overduin B and Pignatelli M and Pritchard B and Pritchard E and Riat HS and Ritchie GRS and Ruffier M and Schuster M and Sheppard D and Sobral D and Taylor K and Thormann A and Trevanion S and White S and Wilder SP and Aken BL and Birney E and Cunningham F and Dunham I and Harrow J and Herrero J and Hubbard TJP and Johnson N and Kinsella R and Parker A and Spudich G and Yates A and Zadissa A and Searle SMJ}, title = {Ensembl 2013}, journal = {Nucleic Acids Res}, volume = {41}, number = {D1}, pages = {D48--55}, year = {2013}, doi = {10.1093/nar/gks1236}, howpublished = {Advanced online publication: 30 November 2012}, abstract = {The Ensembl project (http://www.ensembl.org) provides genome information for sequenced chordate genomes with a particular focus on human, mouse, zebrafish and rat. Our resources include evidenced-based gene sets for all supported species; large-scale whole genome multiple species alignments across vertebrates and clade-specific alignments for eutherian mammals, primates, birds and fish; variation data resources for 17 species and regulation annotations based on ENCODE and other data sets. Ensembl data are accessible through the genome browser at http://www.ensembl.org and through other tools and programmatic interfaces},} - P Flicek. Evolutionary biology: The handiwork of tinkering. Nature 2013;500(7461):158–159. doi:10.1038/500158a
[BibTeX]@Article{23925235, author = {Flicek P}, title = {Evolutionary biology: The handiwork of tinkering}, journal = {Nature}, volume = {500}, number = {7461}, pages = {158--159}, year = {2013}, doi = {10.1038/500158a}, } - SS Atanur, AG Diaz, K Maratou, A Sarkis, M Rotival, L Game, MR Tschannen, PJ Kaisaki, GW Otto, MCJ Ma, TM Keane, O Hummel, K Saar, W Chen, V Guryev, K Gopalakrishnan, MR Garrett, B Joe, L Citterio, G Bianchi, M McBride, A Dominiczak, DJ Adams, T Serikawa, P Flicek, E Cuppen, N Hubner, E Petretto, D Gauguier, A Kwitek, H Jacob, TJ Aitman. Genome Sequencing Reveals Loci under Artificial Selection that Underlie Disease Phenotypes in the Laboratory Rat. Cell 2013;154(3):691–703. doi:10.1016/j.cell.2013.06.040
[BibTeX] [Abstract]
Large numbers of inbred laboratory rat strains have been developed for a range of complex disease phenotypes. To gain insights into the evolutionary pressures underlying selection for these phenotypes, we sequenced the genomes of 27 rat strains, including 11 models of hypertension, diabetes, and insulin resistance, along with their respective control strains. Altogether, we identified more than 13 million single-nucleotide variants, indels, and structural variants across these rat strains. Analysis of strain-specific selective sweeps and gene clusters implicated genes and pathways involved in cation transport, angiotensin production, and regulators of oxidative stress in the development of cardiovascular disease phenotypes in rats. Many of the rat loci that we identified overlap with previously mapped loci for related traits in humans, indicating the presence of shared pathways underlying these phenotypes in rats and humans. These data represent a step change in resources available for evolutionary analysis of complex traits in disease models. PAPERCLIP
@Article{23890820, author = {Atanur SS and Diaz AG and Maratou K and Sarkis A and Rotival M and Game L and Tschannen MR and Kaisaki PJ and Otto GW and Ma MCJ and Keane TM and Hummel O and Saar K and Chen W and Guryev V and Gopalakrishnan K and Garrett MR and Joe B and Citterio L and Bianchi G and McBride M and Dominiczak A and Adams DJ and Serikawa T and Flicek P and Cuppen E and Hubner N and Petretto E and Gauguier D and Kwitek A and Jacob H and Aitman TJ}, title = {Genome Sequencing Reveals Loci under Artificial Selection that Underlie Disease Phenotypes in the Laboratory Rat}, journal = {Cell}, volume = {154}, number = {3}, pages = {691--703}, year = {2013}, doi = {10.1016/j.cell.2013.06.040}, abstract = {Large numbers of inbred laboratory rat strains have been developed for a range of complex disease phenotypes. To gain insights into the evolutionary pressures underlying selection for these phenotypes, we sequenced the genomes of 27 rat strains, including 11 models of hypertension, diabetes, and insulin resistance, along with their respective control strains. Altogether, we identified more than 13 million single-nucleotide variants, indels, and structural variants across these rat strains. Analysis of strain-specific selective sweeps and gene clusters implicated genes and pathways involved in cation transport, angiotensin production, and regulators of oxidative stress in the development of cardiovascular disease phenotypes in rats. Many of the rat loci that we identified overlap with previously mapped loci for related traits in humans, indicating the presence of shared pathways underlying these phenotypes in rats and humans. These data represent a step change in resources available for evolutionary analysis of complex traits in disease models. PAPERCLIP},} - E Khurana, Y Fu, V Colonna, XJ Mu, HM Kang, T Lappalainen, A Sboner, L Lochovsky, J Chen, A Harmanci, J Das, A Abyzov, S Balasubramanian, K Beal, D Chakravarty, D Challis, Y Chen, D Clarke, L Clarke, F Cunningham, US Evani, P Flicek, R Fragoza, E Garrison, R Gibbs, ZH Gümüs, J Herrero, N Kitabayashi, Y Kong, K Lage, V Liluashvili, SM Lipkin, DG Macarthur, G Marth, D Muzny, TH Pers, GRS Ritchie, JA Rosenfeld, C Sisu, X Wei, M Wilson, Y Xue, F Yu, ET Dermitzakis, H Yu, MA Rubin, C Tyler-Smith, M Gerstein. Integrative Annotation of Variants from 1092 Humans: Application to Cancer Genomics. Science 2013;342(6154):1235587. doi:10.1126/science.1235587
[BibTeX] [Abstract]
Interpreting variants, especially noncoding ones, in the increasing number of personal genomes is challenging. We used patterns of polymorphisms in functionally annotated regions in 1092 humans to identify deleterious variants; then we experimentally validated candidates. We analyzed both coding and noncoding regions, with the former corroborating the latter. We found regions particularly sensitive to mutations (''ultrasensitive'') and variants that are disruptive because of mechanistic effects on transcription-factor binding (that is, ``motif-breakers''). We also found variants in regions with higher network centrality tend to be deleterious. Insertions and deletions followed a similar pattern to single-nucleotide variants, with some notable exceptions (e.g., certain deletions and enhancers). On the basis of these patterns, we developed a computational tool (FunSeq), whose application to ~90 cancer genomes reveals nearly a hundred candidate noncoding drivers
@Article{24092746, author = {Khurana E and Fu Y and Colonna V and Mu XJ and Kang HM and Lappalainen T and Sboner A and Lochovsky L and Chen J and Harmanci A and Das J and Abyzov A and Balasubramanian S and Beal K and Chakravarty D and Challis D and Chen Y and Clarke D and Clarke L and Cunningham F and Evani US and Flicek P and Fragoza R and Garrison E and Gibbs R and Gümüs ZH and Herrero J and Kitabayashi N and Kong Y and Lage K and Liluashvili V and Lipkin SM and Macarthur DG and Marth G and Muzny D and Pers TH and Ritchie GRS and Rosenfeld JA and Sisu C and Wei X and Wilson M and Xue Y and Yu F and Dermitzakis ET and Yu H and Rubin MA and Tyler-Smith C and Gerstein M}, title = {Integrative Annotation of Variants from 1092 Humans: Application to Cancer Genomics}, journal = {Science}, volume = {342}, number = {6154}, pages = {1235587}, year = {2013}, doi = {10.1126/science.1235587}, abstract = {Interpreting variants, especially noncoding ones, in the increasing number of personal genomes is challenging. We used patterns of polymorphisms in functionally annotated regions in 1092 humans to identify deleterious variants; then we experimentally validated candidates. We analyzed both coding and noncoding regions, with the former corroborating the latter. We found regions particularly sensitive to mutations (''ultrasensitive'') and variants that are disruptive because of mechanistic effects on transcription-factor binding (that is, ``motif-breakers''). We also found variants in regions with higher network centrality tend to be deleterious. Insertions and deletions followed a similar pattern to single-nucleotide variants, with some notable exceptions (e.g., certain deletions and enhancers). On the basis of these patterns, we developed a computational tool (FunSeq), whose application to ~90 cancer genomes reveals nearly a hundred candidate noncoding drivers},} - MC Ward, MD Wilson, NL Barbosa-Morais, D Schmidt, R Stark, Q Pan, PC Schwalie, S Menon, M Lukk, S Watt, D Thybert, C Kutter, K Kirschner, P Flicek, BJ Blencowe, DT Odom. Latent regulatory potential of human-specific repetitive elements. Mol Cell 2013;49(2):262–272. doi:10.1016/j.molcel.2012.11.013
[BibTeX] [Abstract]
At least half of the human genome is derived from repetitive elements, which are often lineage specific and silenced by a variety of genetic and epigenetic mechanisms. Using a transchromosomic mouse strain that transmits an almost complete single copy of human chromosome 21 via the female germline, we show that a heterologous regulatory environment can transcriptionally activate transposon-derived human regulatory regions. In the mouse nucleus, hundreds of locations on human chromosome 21 newly associate with activating histone modifications in both somatic and germline tissues, and influence the gene expression of nearby transcripts. These regions are enriched with primate and human lineage-specific transposable elements, and their activation corresponds to changes in DNA methylation at CpG dinucleotides. This study reveals the latent regulatory potential of the repetitive human genome and illustrates the species specificity of mechanisms that control it
@Article{23246434, author = {Ward MC and Wilson MD and Barbosa-Morais NL and Schmidt D and Stark R and Pan Q and Schwalie PC and Menon S and Lukk M and Watt S and Thybert D and Kutter C and Kirschner K and Flicek P and Blencowe BJ and Odom DT}, title = {Latent regulatory potential of human-specific repetitive elements}, journal = {Mol Cell}, volume = {49}, number = {2}, pages = {262--272}, year = {2013}, doi = {10.1016/j.molcel.2012.11.013}, abstract = {At least half of the human genome is derived from repetitive elements, which are often lineage specific and silenced by a variety of genetic and epigenetic mechanisms. Using a transchromosomic mouse strain that transmits an almost complete single copy of human chromosome 21 via the female germline, we show that a heterologous regulatory environment can transcriptionally activate transposon-derived human regulatory regions. In the mouse nucleus, hundreds of locations on human chromosome 21 newly associate with activating histone modifications in both somatic and germline tissues, and influence the gene expression of nearby transcripts. These regions are enriched with primate and human lineage-specific transposable elements, and their activation corresponds to changes in DNA methylation at CpG dinucleotides. This study reveals the latent regulatory potential of the repetitive human genome and illustrates the species specificity of mechanisms that control it},} - JJ Smith, S Kuraku, C Holt, T Sauka-Spengler, N Jiang, MS Campbell, MD Yandell, T Manousaki, A Meyer, OE Bloom, JR Morgan, JD Buxbaum, R Sachidanandam, C Sims, AS Garruss, M Cook, R Krumlauf, LM Wiedemann, SA Sower, WA Decatur, JA Hall, CT Amemiya, NR Saha, KM Buckley, JP Rast, S Das, M Hirano, N McCurley, P Guo, N Rohner, CJ Tabin, P Piccinelli, G Elgar, M Ruffier, BL Aken, SMJ Searle, M Muffato, M Pignatelli, J Herrero, M Jones, CT Brown, Y-W Chung-Davidson, KG Nanlohy, SV Libants, C-Y Yeh, DW McCauley, JA Langeland, Z Pancer, B Fritzsch, de PJ Jong, B Zhu, LL Fulton, B Theising, P Flicek, ME Bronner, WC Warren, SW Clifton, RK Wilson, W Li. Sequencing of the sea lamprey (Petromyzon marinus) genome provides insights into vertebrate evolution. Nat Genet 2013;45(4):415–421. doi:10.1038/ng.2568
[BibTeX] [Abstract]
Lampreys are representatives of an ancient vertebrate lineage that diverged from our own ∼500 million years ago. By virtue of this deeply shared ancestry, the sea lamprey (P. marinus) genome is uniquely poised to provide insight into the ancestry of vertebrate genomes and the underlying principles of vertebrate biology. Here, we present the first lamprey whole-genome sequence and assembly. We note challenges faced owing to its high content of repetitive elements and GC bases, as well as the absence of broad-scale sequence information from closely related species. Analyses of the assembly indicate that two whole-genome duplications likely occurred before the divergence of ancestral lamprey and gnathostome lineages. Moreover, the results help define key evolutionary events within vertebrate lineages, including the origin of myelin-associated proteins and the development of appendages. The lamprey genome provides an important resource for reconstructing vertebrate origins and the evolutionary events that have shaped the genomes of extant organisms
@Article{23435085, author = {Smith JJ and Kuraku S and Holt C and Sauka-Spengler T and Jiang N and Campbell MS and Yandell MD and Manousaki T and Meyer A and Bloom OE and Morgan JR and Buxbaum JD and Sachidanandam R and Sims C and Garruss AS and Cook M and Krumlauf R and Wiedemann LM and Sower SA and Decatur WA and Hall JA and Amemiya CT and Saha NR and Buckley KM and Rast JP and Das S and Hirano M and McCurley N and Guo P and Rohner N and Tabin CJ and Piccinelli P and Elgar G and Ruffier M and Aken BL and Searle SMJ and Muffato M and Pignatelli M and Herrero J and Jones M and Brown CT and Chung-Davidson Y-W and Nanlohy KG and Libants SV and Yeh C-Y and McCauley DW and Langeland JA and Pancer Z and Fritzsch B and de Jong PJ and Zhu B and Fulton LL and Theising B and Flicek P and Bronner ME and Warren WC and Clifton SW and Wilson RK and Li W}, title = {Sequencing of the sea lamprey (Petromyzon marinus) genome provides insights into vertebrate evolution}, journal = {Nat Genet}, volume = {45}, number = {4}, pages = {415--421}, year = {2013}, doi = {10.1038/ng.2568}, abstract = {Lampreys are representatives of an ancient vertebrate lineage that diverged from our own ∼500 million years ago. By virtue of this deeply shared ancestry, the sea lamprey (P. marinus) genome is uniquely poised to provide insight into the ancestry of vertebrate genomes and the underlying principles of vertebrate biology. Here, we present the first lamprey whole-genome sequence and assembly. We note challenges faced owing to its high content of repetitive elements and GC bases, as well as the absence of broad-scale sequence information from closely related species. Analyses of the assembly indicate that two whole-genome duplications likely occurred before the divergence of ancestral lamprey and gnathostome lineages. Moreover, the results help define key evolutionary events within vertebrate lineages, including the origin of myelin-associated proteins and the development of appendages. The lamprey genome provides an important resource for reconstructing vertebrate origins and the evolutionary events that have shaped the genomes of extant organisms},} - Z Wang, J Pascual-Anaya, A Zadissa, W Li, Y Niimura, Z Huang, C Li, S White, Z Xiong, D Fang, B Wang, Y Ming, Y Chen, Y Zheng, S Kuraku, M Pignatelli, J Herrero, K Beal, M Nozawa, Q Li, J Wang, H Zhang, L Yu, S Shigenobu, J Wang, J Liu, P Flicek, S Searle, J Wang, S Kuratani, Y Yin, B Aken, G Zhang, N Irie. The draft genomes of soft-shell turtle and green sea turtle yield insights into the development and evolution of the turtle-specific body plan. Nat Genet 2013;45(6):701–706. doi:10.1038/ng.2615
[BibTeX] [Abstract]
The unique anatomical features of turtles have raised unanswered questions about the origin of their unique body plan. We generated and analyzed draft genomes of the soft-shell turtle (Pelodiscus sinensis) and the green sea turtle (Chelonia mydas); our results indicated the close relationship of the turtles to the bird-crocodilian lineage, from which they split ∼267.9-248.3 million years ago (Upper Permian to Triassic). We also found extensive expansion of olfactory receptor genes in these turtles. Embryonic gene expression analysis identified an hourglass-like divergence of turtle and chicken embryogenesis, with maximal conservation around the vertebrate phylotypic period, rather than at later stages that show the amniote-common pattern. Wnt5a expression was found in the growth zone of the dorsal shell, supporting the possible co-option of limb-associated Wnt signaling in the acquisition of this turtle-specific novelty. Our results suggest that turtle evolution was accompanied by an unexpectedly conservative vertebrate phylotypic period, followed by turtle-specific repatterning of development to yield the novel structure of the shell
@Article{23624526, author = {Wang Z and Pascual-Anaya J and Zadissa A and Li W and Niimura Y and Huang Z and Li C and White S and Xiong Z and Fang D and Wang B and Ming Y and Chen Y and Zheng Y and Kuraku S and Pignatelli M and Herrero J and Beal K and Nozawa M and Li Q and Wang J and Zhang H and Yu L and Shigenobu S and Wang J and Liu J and Flicek P and Searle S and Wang J and Kuratani S and Yin Y and Aken B and Zhang G and Irie N}, title = {The draft genomes of soft-shell turtle and green sea turtle yield insights into the development and evolution of the turtle-specific body plan}, journal = {Nat Genet}, volume = {45}, number = {6}, pages = {701--706}, year = {2013}, doi = {10.1038/ng.2615}, abstract = {The unique anatomical features of turtles have raised unanswered questions about the origin of their unique body plan. We generated and analyzed draft genomes of the soft-shell turtle (Pelodiscus sinensis) and the green sea turtle (Chelonia mydas); our results indicated the close relationship of the turtles to the bird-crocodilian lineage, from which they split ∼267.9-248.3 million years ago (Upper Permian to Triassic). We also found extensive expansion of olfactory receptor genes in these turtles. Embryonic gene expression analysis identified an hourglass-like divergence of turtle and chicken embryogenesis, with maximal conservation around the vertebrate phylotypic period, rather than at later stages that show the amniote-common pattern. Wnt5a expression was found in the growth zone of the dorsal shell, supporting the possible co-option of limb-associated Wnt signaling in the acquisition of this turtle-specific novelty. Our results suggest that turtle evolution was accompanied by an unexpectedly conservative vertebrate phylotypic period, followed by turtle-specific repatterning of development to yield the novel structure of the shell},} - T Lappalainen, M Sammeth, MR Friedländer, PAC Hoen `t, J Monlong, MA Rivas, M Gonzàlez-Porta, N Kurbatova, T Griebel, PG Ferreira, M Barann, T Wieland, L Greger, van M Iterson, J Almlöf, P Ribeca, I Pulyakhina, D Esser, T Giger, A Tikhonov, M Sultan, G Bertier, DG MacArthur, M Lek, E Lizano, HPJ Buermans, I Padioleau, T Schwarzmayr, O Karlberg, H Ongen, H Kilpinen, S Beltran, M Gut, K Kahlem, V Amstislavskiy, O Stegle, M Pirinen, SB Montgomery, P Donnelly, MI McCarthy, P Flicek, TM Strom, TG Consortium, H Lehrach, S Schreiber, R Sudbrak, A Carracedo, SE Antonarakis, R Häsler, A-C Syvänen, van G-J Ommen, A Brazma, T Meitinger, P Rosenstiel, R Guigó, IG Gut, X Estivill, ET Dermitzakis. Transcriptome and genome sequencing uncovers functional variation in humans. Nature 2013;501(7468):506–511. doi:10.1038/nature12531
[BibTeX] [Abstract]
Genome sequencing projects are discovering millions of genetic variants in humans, and interpretation of their functional effects is essential for understanding the genetic basis of variation in human traits. Here we report sequencing and deep analysis of messenger RNA and microRNA from lymphoblastoid cell lines of 462 individuals from the 1000 Genomes Project–the first uniformly processed high-throughput RNA-sequencing data from multiple human populations with high-quality genome sequences. We discover extremely widespread genetic variation affecting the regulation of most genes, with transcript structure and expression level variation being equally common but genetically largely independent. Our characterization of causal regulatory variation sheds light on the cellular mechanisms of regulatory and loss-of-function variation, and allows us to infer putative causal variants for dozens of disease-associated loci. Altogether, this study provides a deep understanding of the cellular mechanisms of transcriptome variation and of the landscape of functional variants in the human genome
@Article{24037378, author = {Lappalainen T and Sammeth M and Friedländer MR and `t Hoen PAC and Monlong J and Rivas MA and Gonzàlez-Porta M and Kurbatova N and Griebel T and Ferreira PG and Barann M and Wieland T and Greger L and van Iterson M and Almlöf J and Ribeca P and Pulyakhina I and Esser D and Giger T and Tikhonov A and Sultan M and Bertier G and MacArthur DG and Lek M and Lizano E and Buermans HPJ and Padioleau I and Schwarzmayr T and Karlberg O and Ongen H and Kilpinen H and Beltran S and Gut M and Kahlem K and Amstislavskiy V and Stegle O and Pirinen M and Montgomery SB and Donnelly P and McCarthy MI and Flicek P and Strom TM and Consortium TG and Lehrach H and Schreiber S and Sudbrak R and Carracedo A and Antonarakis SE and Häsler R and Syvänen A-C and van Ommen G-J and Brazma A and Meitinger T and Rosenstiel P and Guigó R and Gut IG and Estivill X and Dermitzakis ET}, title = {Transcriptome and genome sequencing uncovers functional variation in humans}, journal = {Nature}, volume = {501}, number = {7468}, pages = {506--511}, year = {2013}, doi = {10.1038/nature12531}, abstract = {Genome sequencing projects are discovering millions of genetic variants in humans, and interpretation of their functional effects is essential for understanding the genetic basis of variation in human traits. Here we report sequencing and deep analysis of messenger RNA and microRNA from lymphoblastoid cell lines of 462 individuals from the 1000 Genomes Project--the first uniformly processed high-throughput RNA-sequencing data from multiple human populations with high-quality genome sequences. We discover extremely widespread genetic variation affecting the regulation of most genes, with transcript structure and expression level variation being equally common but genetically largely independent. Our characterization of causal regulatory variation sheds light on the cellular mechanisms of regulatory and loss-of-function variation, and allows us to infer putative causal variants for dozens of disease-associated loci. Altogether, this study provides a deep understanding of the cellular mechanisms of transcriptome variation and of the landscape of functional variants in the human genome},}
2012
- A-M Mallon, V Iyer, D Melvin, H Morgan, H Parkinson, SDM Brown, P Flicek, WC Skarnes. Accessing data from the International Mouse Phenotyping Consortium: state of the art and future plans. Mamm Genome 2012;23(9-10):641–652. doi:10.1007/s00335-012-9428-9
[BibTeX] [Abstract]
The International Mouse Phenotyping Consortium (IMPC) ( http://www.mousephenotype.org ) will reveal the pleiotropic functions of every gene in the mouse genome and uncover the wider role of genetic loci within diverse biological systems. Comprehensive informatics solutions are vital to ensuring that this vast array of data is captured in a standardised manner and made accessible to the scientific community for interrogation and analysis. Here we review the existing EuroPhenome and WTSI phenotype informatics systems and the IKMC portal, and present plans for extending these systems and lessons learned to the development of a robust IMPC informatics infrastructure
@Article{22991088, author = {Mallon A-M and Iyer V and Melvin D and Morgan H and Parkinson H and Brown SDM and Flicek P and Skarnes WC}, title = {Accessing data from the International Mouse Phenotyping Consortium: state of the art and future plans}, journal = {Mamm Genome}, volume = {23}, number = {9-10}, pages = {641--652}, year = {2012}, doi = {10.1007/s00335-012-9428-9}, abstract = {The International Mouse Phenotyping Consortium (IMPC) ( http://www.mousephenotype.org ) will reveal the pleiotropic functions of every gene in the mouse genome and uncover the wider role of genetic loci within diverse biological systems. Comprehensive informatics solutions are vital to ensuring that this vast array of data is captured in a standardised manner and made accessible to the scientific community for interrogation and analysis. Here we review the existing EuroPhenome and WTSI phenotype informatics systems and the IKMC portal, and present plans for extending these systems and lessons learned to the development of a robust IMPC informatics infrastructure},} - ENCODE Project Consortium.. An integrated encyclopedia of DNA elements in the human genome. Nature 2012;489(7414):57–74. doi:10.1038/nature11247
[BibTeX] [Abstract]
The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80\% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research
@Article{22955616, author = {{ENCODE Project Consortium.}}, title = {An integrated encyclopedia of DNA elements in the human genome}, journal = {Nature}, volume = {489}, number = {7414}, pages = {57--74}, year = {2012}, doi = {10.1038/nature11247}, abstract = {The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80\% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research},} - AC Nelson, N Pillay, S Henderson, N Presneau, R Tirabosco, D Halai, F Berisha, P Flicek, DL Stemple, C Stern, FC Wardle, AM Flanagan. An integrated functional genomics approach identifies the regulatory network directed by brachyury (T) in chordoma. J Pathol 2012;228(3):274–285. doi:10.1002/path.4082
[BibTeX] [Abstract]
Chordoma is a rare malignant tumour of bone, the molecular marker of which is the expression of the transcription factor, brachyury. Having recently demonstrated that silencing brachyury induces growth arrest in a chordoma cell line, we now seek to identify its downstream target genes. Here we use an integrated functional genomics approach involving shRNA-mediated brachyury knockdown, gene expression microarray, ChIP-seq experiments and bioinformatics analysis to achieve this goal. We confirm that the T-Box binding motif of human brachyury is identical to that found in mouse, Xenopus and zebrafish development, and that brachyury acts primarily as an activator of transcription. Using human chordoma samples for validation purposes we show that brachyury binds 99 direct targets and indirectly influences the expression of 64 other genes, thereby acting as a master regulator of an elaborate oncogenic transcriptional network encompassing diverse signalling pathways including components of the cell cycle, and extracellular matrix components. Given the wide repertoire of its active binding and the relative specific localisation of brachyury to the tumour cells, we propose that a RNA interference-based gene therapy approach is a plausible therapeutic avenue worthy of investigation. Copyright {\copyright} 2012 Pathological Society of Great Britain and Ireland. Published by John Wiley & Sons, Ltd
@Article{22847733, author = {Nelson AC and Pillay N and Henderson S and Presneau N and Tirabosco R and Halai D and Berisha F and Flicek P and Stemple DL and Stern C and Wardle FC and Flanagan AM}, title = {An integrated functional genomics approach identifies the regulatory network directed by brachyury (T) in chordoma}, journal = {J Pathol}, volume = {228}, number = {3}, pages = {274--285}, year = {2012}, doi = {10.1002/path.4082}, abstract = {Chordoma is a rare malignant tumour of bone, the molecular marker of which is the expression of the transcription factor, brachyury. Having recently demonstrated that silencing brachyury induces growth arrest in a chordoma cell line, we now seek to identify its downstream target genes. Here we use an integrated functional genomics approach involving shRNA-mediated brachyury knockdown, gene expression microarray, ChIP-seq experiments and bioinformatics analysis to achieve this goal. We confirm that the T-Box binding motif of human brachyury is identical to that found in mouse, Xenopus and zebrafish development, and that brachyury acts primarily as an activator of transcription. Using human chordoma samples for validation purposes we show that brachyury binds 99 direct targets and indirectly influences the expression of 64 other genes, thereby acting as a master regulator of an elaborate oncogenic transcriptional network encompassing diverse signalling pathways including components of the cell cycle, and extracellular matrix components. Given the wide repertoire of its active binding and the relative specific localisation of brachyury to the tumour cells, we propose that a RNA interference-based gene therapy approach is a plausible therapeutic avenue worthy of investigation. Copyright {\copyright} 2012 Pathological Society of Great Britain and Ireland. Published by John Wiley \& Sons, Ltd},} - The 1000 Genomes Project Consortium.. An integrated map of genetic variation from 1,092 human genomes. Nature 2012;491(7422):56–65. doi:10.1038/nature11632
[BibTeX] [Abstract]
By characterizing the geographic and functional spectrum of human genetic variation, the 1000 Genomes Project aims to build a resource to help to understand the genetic contribution to disease. Here we describe the genomes of 1,092 individuals from 14 populations, constructed using a combination of low-coverage whole-genome and exome sequencing. By developing methods to integrate information across several algorithms and diverse data sources, we provide a validated haplotype map of 38 million single nucleotide polymorphisms, 1.4 million short insertions and deletions, and more than 14,000 larger deletions. We show that individuals from different populations carry different profiles of rare and common variants, and that low-frequency variants show substantial geographic differentiation, which is further increased by the action of purifying selection. We show that evolutionary conservation and coding consequence are key determinants of the strength of purifying selection, that rare-variant load varies substantially across biological pathways, and that each individual contains hundreds of rare non-coding variants at conserved sites, such as motif-disrupting changes in transcription-factor-binding sites. This resource, which captures up to 98\% of accessible single nucleotide polymorphisms at a frequency of 1\% in related populations, enables analysis of common and low-frequency variants in individuals from diverse, including admixed, populations
@Article{23128226, author = {{The 1000 Genomes Project Consortium.}}, title = {An integrated map of genetic variation from 1,092 human genomes}, journal = {Nature}, volume = {491}, number = {7422}, pages = {56--65}, year = {2012}, doi = {10.1038/nature11632}, abstract = {By characterizing the geographic and functional spectrum of human genetic variation, the 1000 Genomes Project aims to build a resource to help to understand the genetic contribution to disease. Here we describe the genomes of 1,092 individuals from 14 populations, constructed using a combination of low-coverage whole-genome and exome sequencing. By developing methods to integrate information across several algorithms and diverse data sources, we provide a validated haplotype map of 38 million single nucleotide polymorphisms, 1.4 million short insertions and deletions, and more than 14,000 larger deletions. We show that individuals from different populations carry different profiles of rare and common variants, and that low-frequency variants show substantial geographic differentiation, which is further increased by the action of purifying selection. We show that evolutionary conservation and coding consequence are key determinants of the strength of purifying selection, that rare-variant load varies substantially across biological pathways, and that each individual contains hundreds of rare non-coding variants at conserved sites, such as motif-disrupting changes in transcription-factor-binding sites. This resource, which captures up to 98\% of accessible single nucleotide polymorphisms at a frequency of 1\% in related populations, enables analysis of common and low-frequency variants in individuals from diverse, including admixed, populations},} - D Adams, L Altucci, SE Antonarakis, J Ballesteros, S Beck, A Bird, C Bock, B Boehm, E Campo, A Caricasole, F Dahl, ET Dermitzakis, T Enver, M Esteller, X Estivill, A Ferguson-Smith, J Fitzgibbon, P Flicek, C Giehl, T Graf, F Grosveld, R Guigo, I Gut, K Helin, J Jarvius, R Küppers, H Lehrach, T Lengauer, A Lernmark, D Leslie, M Loeffler, E Macintyre, A Mai, JH Martens, S Minucci, WH Ouwehand, PG Pelicci, H Pendeville, B Porse, V Rakyan, W Reik, M Schrappe, D Schübeler, M Seifert, R Siebert, D Simmons, N Soranzo, S Spicuglia, M Stratton, HG Stunnenberg, A Tanay, D Torrents, A Valencia, E Vellenga, M Vingron, J Walter, S Willcocks. BLUEPRINT to decode the epigenetic signature written in blood. Nat Biotechnol 2012;30(3):224–226. doi:10.1038/nbt.2153
[BibTeX]@Article{22398613, author = {Adams D and Altucci L and Antonarakis SE and Ballesteros J and Beck S and Bird A and Bock C and Boehm B and Campo E and Caricasole A and Dahl F and Dermitzakis ET and Enver T and Esteller M and Estivill X and Ferguson-Smith A and Fitzgibbon J and Flicek P and Giehl C and Graf T and Grosveld F and Guigo R and Gut I and Helin K and Jarvius J and Küppers R and Lehrach H and Lengauer T and Lernmark A and Leslie D and Loeffler M and Macintyre E and Mai A and Martens JH and Minucci S and Ouwehand WH and Pelicci PG and Pendeville H and Porse B and Rakyan V and Reik W and Schrappe M and Schübeler D and Seifert M and Siebert R and Simmons D and Soranzo N and Spicuglia S and Stratton M and Stunnenberg HG and Tanay A and Torrents D and Valencia A and Vellenga E and Vingron M and Walter J and Willcocks S}, title = {BLUEPRINT to decode the epigenetic signature written in blood}, journal = {Nat Biotechnol}, volume = {30}, number = {3}, pages = {224--226}, year = {2012}, doi = {10.1038/nbt.2153}, } - AJ Faure, D Schmidt, S Watt, PC Schwalie, MD Wilson, H Xu, RG Ramsay, DT Odom, P Flicek. Cohesin regulates tissue-specific expression by stabilizing highly occupied cis-regulatory modules. Genome Res 2012;22(11):2163–2175. doi:10.1101/gr.136507.111
[BibTeX] [Abstract]
The cohesin protein complex contributes to transcriptional regulation in a CTCF-independent manner by colocalizing with master regulators at tissue-specific loci. The regulation of transcription involves the concerted action of multiple transcription factors (TFs) and cohesin's role in this context of combinatorial TF binding remains unexplored. To investigate cohesin-non-CTCF (CNC) binding events in vivo we mapped cohesin and CTCF, as well as a collection of tissue-specific and ubiquitous transcriptional regulators using ChIP-seq in primary mouse liver. We observe a positive correlation between the number of distinct TFs bound and the presence of CNC sites. In contrast to regions of the genome where cohesin and CTCF colocalize, CNC sites coincide with the binding of master regulators and enhancer-markers and are significantly associated with liver-specific expressed genes. We also show that cohesin presence partially explains the commonly observed discrepancy between TF motif score and ChIP signal. Evidence from these statistical analyses in wild-type cells, and comparisons to maps of TF binding in Rad21-cohesin haploinsufficient mouse liver, suggests that cohesin helps to stabilize large protein-DNA complexes. Finally, we observe that the presence of mirrored CTCF binding events at promoters and their nearby cohesin-bound enhancers is associated with elevated expression levels
@Article{22780989, author = {Faure AJ and Schmidt D and Watt S and Schwalie PC and Wilson MD and Xu H and Ramsay RG and Odom DT and Flicek P}, title = {Cohesin regulates tissue-specific expression by stabilizing highly occupied cis-regulatory modules}, journal = {Genome Res}, volume = {22}, number = {11}, pages = {2163--2175}, year = {2012}, doi = {10.1101/gr.136507.111}, howpublished = {Advanced online publication: 10 July 2012}, abstract = {The cohesin protein complex contributes to transcriptional regulation in a CTCF-independent manner by colocalizing with master regulators at tissue-specific loci. The regulation of transcription involves the concerted action of multiple transcription factors (TFs) and cohesin's role in this context of combinatorial TF binding remains unexplored. To investigate cohesin-non-CTCF (CNC) binding events in vivo we mapped cohesin and CTCF, as well as a collection of tissue-specific and ubiquitous transcriptional regulators using ChIP-seq in primary mouse liver. We observe a positive correlation between the number of distinct TFs bound and the presence of CNC sites. In contrast to regions of the genome where cohesin and CTCF colocalize, CNC sites coincide with the binding of master regulators and enhancer-markers and are significantly associated with liver-specific expressed genes. We also show that cohesin presence partially explains the commonly observed discrepancy between TF motif score and ChIP signal. Evidence from these statistical analyses in wild-type cells, and comparisons to maps of TF binding in Rad21-cohesin haploinsufficient mouse liver, suggests that cohesin helps to stabilize large protein-DNA complexes. Finally, we observe that the presence of mirrored CTCF binding events at promoters and their nearby cohesin-bound enhancers is associated with elevated expression levels},} - Z Iqbal, M Caccamo, I Turner, P Flicek, G McVean. De novo assembly and genotyping of variants using colored de Bruijn graphs. Nat Genet 2012;44(2):226–232. doi:10.1038/ng.1028
[BibTeX] [Abstract]
Detecting genetic variants that are highly divergent from a reference sequence remains a major challenge in genome sequencing. We introduce de novo assembly algorithms using colored de Bruijn graphs for detecting and genotyping simple and complex genetic variants in an individual or population. We provide an efficient software implementation, Cortex, the first de novo assembler capable of assembling multiple eukaryotic genomes simultaneously. Four applications of Cortex are presented. First, we detect and validate both simple and complex structural variations in a high-coverage human genome. Second, we identify more than 3 Mb of sequence absent from the human reference genome, in pooled low-coverage population sequence data from the 1000 Genomes Project. Third, we show how population information from ten chimpanzees enables accurate variant calls without a reference sequence. Last, we estimate classical human leukocyte antigen (HLA) genotypes at HLA-B, the most variable gene in the human genome
@Article{22231483, author = {Iqbal Z and Caccamo M and Turner I and Flicek P and McVean G}, title = {De novo assembly and genotyping of variants using colored de Bruijn graphs}, journal = {Nat Genet}, volume = {44}, number = {2}, pages = {226--232}, year = {2012}, doi = {10.1038/ng.1028}, abstract = {Detecting genetic variants that are highly divergent from a reference sequence remains a major challenge in genome sequencing. We introduce de novo assembly algorithms using colored de Bruijn graphs for detecting and genotyping simple and complex genetic variants in an individual or population. We provide an efficient software implementation, Cortex, the first de novo assembler capable of assembling multiple eukaryotic genomes simultaneously. Four applications of Cortex are presented. First, we detect and validate both simple and complex structural variations in a high-coverage human genome. Second, we identify more than 3 Mb of sequence absent from the human reference genome, in pooled low-coverage population sequence data from the 1000 Genomes Project. Third, we show how population information from ten chimpanzees enables accurate variant calls without a reference sequence. Last, we estimate classical human leukocyte antigen (HLA) genotypes at HLA-B, the most variable gene in the human genome},} - P Flicek, MR Amode, D Barrell, K Beal, S Brent, D Carvalho-Silva, P Clapham, G Coates, S Fairley, S Fitzgerald, L Gil, L Gordon, M Hendrix, T Hourlier, N Johnson, AK Kähäri, D Keefe, S Keenan, R Kinsella, M Komorowska, G Koscielny, E Kulesha, P Larsson, I Longden, W McLaren, M Muffato, B Overduin, M Pignatelli, B Pritchard, HS Riat, GRS Ritchie, M Ruffier, M Schuster, D Sobral, YA Tang, K Taylor, S Trevanion, J Vandrovcova, S White, M Wilson, SP Wilder, BL Aken, E Birney, F Cunningham, I Dunham, R Durbin, XM Fernández-Suarez, J Harrow, J Herrero, TJP Hubbard, A Parker, G Proctor, G Spudich, J Vogel, A Yates, A Zadissa, SMJ Searle. Ensembl 2012. Nucleic Acids Res 2012;40(1):D84–90. doi:10.1093/nar/gkr991
[BibTeX] [Abstract]
The Ensembl project (http://www.ensembl.org) provides genome resources for chordate genomes with a particular focus on human genome data as well as data for key model organisms such as mouse, rat and zebrafish. Five additional species were added in the last year including gibbon (Nomascus leucogenys) and Tasmanian devil (Sarcophilus harrisii) bringing the total number of supported species to 61 as of Ensembl release 64 (September 2011). Of these, 55 species appear on the main Ensembl website and six species are provided on the Ensembl preview site (Pre!Ensembl; http://pre.ensembl.org) with preliminary support. The past year has also seen improvements across the project.
@Article{22086963, author = {Flicek P and Amode MR and Barrell D and Beal K and Brent S and Carvalho-Silva D and Clapham P and Coates G and Fairley S and Fitzgerald S and Gil L and Gordon L and Hendrix M and Hourlier T and Johnson N and Kähäri AK and Keefe D and Keenan S and Kinsella R and Komorowska M and Koscielny G and Kulesha E and Larsson P and Longden I and McLaren W and Muffato M and Overduin B and Pignatelli M and Pritchard B and Riat HS and Ritchie GRS and Ruffier M and Schuster M and Sobral D and Tang YA and Taylor K and Trevanion S and Vandrovcova J and White S and Wilson M and Wilder SP and Aken BL and Birney E and Cunningham F and Dunham I and Durbin R and Fernández-Suarez XM and Harrow J and Herrero J and Hubbard TJP and Parker A and Proctor G and Spudich G and Vogel J and Yates A and Zadissa A and Searle SMJ}, title = {Ensembl 2012}, journal = {Nucleic Acids Res}, volume = {40}, number = {1}, pages = {D84--90}, year = {2012}, doi = {10.1093/nar/gkr991}, howpublished = {Advanced online publication: 15 November 2011}, abstract = {The Ensembl project (http://www.ensembl.org) provides genome resources for chordate genomes with a particular focus on human genome data as well as data for key model organisms such as mouse, rat and zebrafish. Five additional species were added in the last year including gibbon (Nomascus leucogenys) and Tasmanian devil (Sarcophilus harrisii) bringing the total number of supported species to 61 as of Ensembl release 64 (September 2011). Of these, 55 species appear on the main Ensembl website and six species are provided on the Ensembl preview site (Pre!Ensembl; http://pre.ensembl.org) with preliminary support. The past year has also seen improvements across the project.},} - A Goncalves, S Leigh-Brown, D Thybert, K Stefflova, E Turro, P Flicek, A Brazma, DT Odom, JC Marioni. Extensive compensatory cis-trans regulation in the evolution of mouse gene expression. Genome Res 2012;22(12):2376–2384. doi:10.1101/gr.142281.112
[BibTeX] [Abstract]
Gene expression levels are thought to diverge primarily via regulatory mutations in trans within species, and in cis between species. To test this hypothesis in mammals we used RNA-sequencing to measure gene expression divergence between C57BL/6J and CAST/EiJ mouse strains and allele-specific expression in their F1 progeny. We identified 535 genes with parent-of-origin specific expression patterns, although few of these showed full allelic silencing. This suggests that the number of imprinted genes in a typical mouse somatic tissue is relatively small. In the set of non-imprinted genes, 32\% showed evidence of divergent expression between the two strains. Of these, 2\% could be attributed purely to variants acting in trans, while 43\% were attributable only to variants acting in cis. The genes with expression divergence driven by changes in trans showed significantly higher sequence constraint than genes where the divergence was explained by variants acting in cis. The remaining genes with divergent patterns of expression (55\%) were regulated by a combination of variants acting in cis and variants acting in trans. Intriguingly, the changes in expression induced by the cis and trans variants were in opposite directions more frequently than expected by chance, implying that compensatory regulation to stabilize gene expression levels is widespread. We propose that expression levels of genes regulated by this mechanism are fine-tuned by cis variants that arise following regulatory changes in trans, suggesting that many cis variants are not the primary targets of natural selection
@Article{22919075, author = {Goncalves A and Leigh-Brown S and Thybert D and Stefflova K and Turro E and Flicek P and Brazma A and Odom DT and Marioni JC}, title = {Extensive compensatory cis-trans regulation in the evolution of mouse gene expression}, journal = {Genome Res}, volume = {22}, number = {12}, pages = {2376--2384}, year = {2012}, doi = {10.1101/gr.142281.112}, howpublished = {Advanced online publication: 23 August 2012}, abstract = {Gene expression levels are thought to diverge primarily via regulatory mutations in trans within species, and in cis between species. To test this hypothesis in mammals we used RNA-sequencing to measure gene expression divergence between C57BL/6J and CAST/EiJ mouse strains and allele-specific expression in their F1 progeny. We identified 535 genes with parent-of-origin specific expression patterns, although few of these showed full allelic silencing. This suggests that the number of imprinted genes in a typical mouse somatic tissue is relatively small. In the set of non-imprinted genes, 32\% showed evidence of divergent expression between the two strains. Of these, 2\% could be attributed purely to variants acting in trans, while 43\% were attributable only to variants acting in cis. The genes with expression divergence driven by changes in trans showed significantly higher sequence constraint than genes where the divergence was explained by variants acting in cis. The remaining genes with divergent patterns of expression (55\%) were regulated by a combination of variants acting in cis and variants acting in trans. Intriguingly, the changes in expression induced by the cis and trans variants were in opposite directions more frequently than expected by chance, implying that compensatory regulation to stabilize gene expression levels is widespread. We propose that expression levels of genes regulated by this mechanism are fine-tuned by cis variants that arise following regulatory changes in trans, suggesting that many cis variants are not the primary targets of natural selection},} - A Scally, JY Dutheil, LW Hillier, GE Jordan, I Goodhead, J Herrero, A Hobolth, T Lappalainen, T Mailund, T Marques-Bonet, S McCarthy, SH Montgomery, PC Schwalie, YA Tang, MC Ward, Y Xue, B Yngvadottir, C Alkan, LN Andersen, Q Ayub, EV Ball, K Beal, BJ Bradley, Y Chen, CM Clee, S Fitzgerald, TA Graves, Y Gu, P Heath, A Heger, E Karakoc, A Kolb-Kokocinski, GK Laird, G Lunter, S Meader, M Mort, JC Mullikin, K Munch, TD O'Connor, AD Phillips, J Prado-Martinez, AS Rogers, S Sajjadian, D Schmidt, K Shaw, JT Simpson, PD Stenson, DJ Turner, L Vigilant, AJ Vilella, W Whitener, B Zhu, DN Cooper, de P Jong, ET Dermitzakis, EE Eichler, P Flicek, N Goldman, NI Mundy, Z Ning, DT Odom, CP Ponting, MA Quail, OA Ryder, SM Searle, WC Warren, RK Wilson, MH Schierup, J Rogers, C Tyler-Smith, R Durbin. Insights into hominid evolution from the gorilla genome sequence. Nature 2012;483(7388):169–175. doi:10.1038/nature10842
[BibTeX] [Abstract]
Gorillas are humans' closest living relatives after chimpanzees, and are of comparable importance for the study of human origins and evolution. Here we present the assembly and analysis of a genome sequence for the western lowland gorilla, and compare the whole genomes of all extant great ape genera. We propose a synthesis of genetic and fossil evidence consistent with placing the human-chimpanzee and human-chimpanzee-gorilla speciation events at approximately 6 and 10 million years ago. In 30\% of the genome, gorilla is closer to human or chimpanzee than the latter are to each other; this is rarer around coding genes, indicating pervasive selection throughout great ape evolution, and has functional consequences in gene expression. A comparison of protein coding genes reveals approximately 500 genes showing accelerated evolution on each of the gorilla, human and chimpanzee lineages, and evidence for parallel acceleration, particularly of genes involved in hearing. We also compare the western and eastern gorilla species, estimating an average sequence divergence time 1.75 million years ago, but with evidence for more recent genetic exchange and a population bottleneck in the eastern species. The use of the genome sequence in these and future analyses will promote a deeper understanding of great ape biology and evolution
@Article{22398555, author = {Scally A and Dutheil JY and Hillier LW and Jordan GE and Goodhead I and Herrero J and Hobolth A and Lappalainen T and Mailund T and Marques-Bonet T and McCarthy S and Montgomery SH and Schwalie PC and Tang YA and Ward MC and Xue Y and Yngvadottir B and Alkan C and Andersen LN and Ayub Q and Ball EV and Beal K and Bradley BJ and Chen Y and Clee CM and Fitzgerald S and Graves TA and Gu Y and Heath P and Heger A and Karakoc E and Kolb-Kokocinski A and Laird GK and Lunter G and Meader S and Mort M and Mullikin JC and Munch K and O'Connor TD and Phillips AD and Prado-Martinez J and Rogers AS and Sajjadian S and Schmidt D and Shaw K and Simpson JT and Stenson PD and Turner DJ and Vigilant L and Vilella AJ and Whitener W and Zhu B and Cooper DN and de Jong P and Dermitzakis ET and Eichler EE and Flicek P and Goldman N and Mundy NI and Ning Z and Odom DT and Ponting CP and Quail MA and Ryder OA and Searle SM and Warren WC and Wilson RK and Schierup MH and Rogers J and Tyler-Smith C and Durbin R}, title = {Insights into hominid evolution from the gorilla genome sequence}, journal = {Nature}, volume = {483}, number = {7388}, pages = {169--175}, year = {2012}, doi = {10.1038/nature10842}, abstract = {Gorillas are humans' closest living relatives after chimpanzees, and are of comparable importance for the study of human origins and evolution. Here we present the assembly and analysis of a genome sequence for the western lowland gorilla, and compare the whole genomes of all extant great ape genera. We propose a synthesis of genetic and fossil evidence consistent with placing the human-chimpanzee and human-chimpanzee-gorilla speciation events at approximately 6 and 10 million years ago. In 30\% of the genome, gorilla is closer to human or chimpanzee than the latter are to each other; this is rarer around coding genes, indicating pervasive selection throughout great ape evolution, and has functional consequences in gene expression. A comparison of protein coding genes reveals approximately 500 genes showing accelerated evolution on each of the gorilla, human and chimpanzee lineages, and evidence for parallel acceleration, particularly of genes involved in hearing. We also compare the western and eastern gorilla species, estimating an average sequence divergence time 1.75 million years ago, but with evidence for more recent genetic exchange and a population bottleneck in the eastern species. The use of the genome sequence in these and future analyses will promote a deeper understanding of great ape biology and evolution},} - X Zheng-Bradley, P Flicek. Maps for the world of genomic medicine: The 2011 CSHL Personal Genomes Meeting. Hum Mutat 2012;33(6):1016–1019. doi:10.1002/humu.22024
[BibTeX] [Abstract]
The fourth Personal Genomes meeting was held at Cold Spring Harbor Laboratory, New York, from 30 September to 2 October and provided an exciting collection of science built on recent significant milestones in individual human genome sequencing, from the first personal genomes to thousands of human genomes sequenced. As ultra-high throughput sequencing platforms enable the production of more and more individual genomes, a growing number of scientists, physicians and clinical geneticists are actively exploring the promise and the implications of these new data. Personal Genomes brought many of these pioneers together with nearly 200 scientists, physicians, ethicists and others to discuss the progress and opportunities around the sequencing and medical interpretation of individual genome sequences
@Article{22253119, author = {Zheng-Bradley X and Flicek P}, title = {Maps for the world of genomic medicine: The 2011 CSHL Personal Genomes Meeting}, journal = {Hum Mutat}, volume = {33}, number = {6}, pages = {1016--1019}, year = {2012}, doi = {10.1002/humu.22024}, abstract = {The fourth Personal Genomes meeting was held at Cold Spring Harbor Laboratory, New York, from 30 September to 2 October and provided an exciting collection of science built on recent significant milestones in individual human genome sequencing, from the first personal genomes to thousands of human genomes sequenced. As ultra-high throughput sequencing platforms enable the production of more and more individual genomes, a growing number of scientists, physicians and clinical geneticists are actively exploring the promise and the implications of these new data. Personal Genomes brought many of these pioneers together with nearly 200 scientists, physicians, ethicists and others to discuss the progress and opportunities around the sequencing and medical interpretation of individual genome sequences},} - L Clarke, X Zheng-Bradley, R Smith, E Kulesha, C Xiao, I Toneva, B Vaughan, D Preuss, R Leinonen, M Shumway, S Sherry, P Flicek, The 1000 Genomes Project Consortium. The 1000 Genomes Project: data management and community access. Nat Methods 2012;9(5):459–462. doi:10.1038/nmeth.1974
[BibTeX] [Abstract]
The 1000 Genomes Project was launched as one of the largest distributed data collection and analysis projects ever undertaken in biology. In addition to the primary scientific goals of creating both a deep catalog of human genetic variation and extensive methods to accurately discover and characterize variation using new sequencing technologies, the project makes all of its data publicly available. Members of the project data coordination center have developed and deployed several tools to enable widespread data access
@Article{22543379, author = {Clarke L and Zheng-Bradley X and Smith R and Kulesha E and Xiao C and Toneva I and Vaughan B and Preuss D and Leinonen R and Shumway M and Sherry S and Flicek P and {The 1000 Genomes Project Consortium}}, title = {The 1000 Genomes Project: data management and community access}, journal = {Nat Methods}, volume = {9}, number = {5}, pages = {459--462}, year = {2012}, doi = {10.1038/nmeth.1974}, abstract = {The 1000 Genomes Project was launched as one of the largest distributed data collection and analysis projects ever undertaken in biology. In addition to the primary scientific goals of creating both a deep catalog of human genetic variation and extensive methods to accurately discover and characterize variation using new sequencing technologies, the project makes all of its data publicly available. Members of the project data coordination center have developed and deployed several tools to enable widespread data access},} - D Schmidt, PC Schwalie, MD Wilson, B Ballester, A Gonçalves, C Kutter, GD Brown, A Marshall, P Flicek, DT Odom. Waves of Retrotransposon Expansion Remodel Genome Organization and CTCF Binding in Multiple Mammalian Lineages. Cell 2012;148(1-2):335–348. doi:10.1016/j.cell.2011.11.058
[BibTeX] [Abstract]
CTCF-binding locations represent regulatory sequences that are highly constrained over the course of evolution. To gain insight into how these DNA elements are conserved and spread through the genome, we defined the full spectrum of CTCF-binding sites, including a 33/34-mer motif, and identified over five thousand highly conserved, robust, and tissue-independent CTCF-binding locations by comparing ChIP-seq data from six mammals. Our data indicate that activation of retroelements has produced species-specific expansions of CTCF binding in rodents, dogs, and opossum, which often functionally serve as chromatin and transcriptional insulators. We discovered fossilized repeat elements flanking deeply conserved CTCF-binding regions, indicating that similar retrotransposon expansions occurred hundreds of millions of years ago. Repeat-driven dispersal of CTCF binding is a fundamental, ancient, and still highly active mechanism of genome evolution in mammalian lineages. PAPERCLIP
@Article{22244452, author = {Schmidt D and Schwalie PC and Wilson MD and Ballester B and Gonçalves A and Kutter C and Brown GD and Marshall A and Flicek P and Odom DT}, title = {Waves of Retrotransposon Expansion Remodel Genome Organization and CTCF Binding in Multiple Mammalian Lineages}, journal = {Cell}, volume = {148}, number = {1-2}, pages = {335--348}, year = {2012}, doi = {10.1016/j.cell.2011.11.058}, abstract = {CTCF-binding locations represent regulatory sequences that are highly constrained over the course of evolution. To gain insight into how these DNA elements are conserved and spread through the genome, we defined the full spectrum of CTCF-binding sites, including a 33/34-mer motif, and identified over five thousand highly conserved, robust, and tissue-independent CTCF-binding locations by comparing ChIP-seq data from six mammals. Our data indicate that activation of retroelements has produced species-specific expansions of CTCF binding in rodents, dogs, and opossum, which often functionally serve as chromatin and transcriptional insulators. We discovered fossilized repeat elements flanking deeply conserved CTCF-binding regions, indicating that similar retrotransposon expansions occurred hundreds of millions of years ago. Repeat-driven dispersal of CTCF binding is a fundamental, ancient, and still highly active mechanism of genome evolution in mammalian lineages. PAPERCLIP},}
2011
- K Lindblad-Toh, M Garber, O Zuk, MF Lin, BJ Parker, S Washietl, P Kheradpour, J Ernst, G Jordan, E Mauceli, LD Ward, CB Lowe, AK Holloway, M Clamp, S Gnerre, J Alföldi, K Beal, J Chang, H Clawson, J Cuff, F Di Palma, S Fitzgerald, P Flicek, M Guttman, MJ Hubisz, DB Jaffe, I Jungreis, WJ Kent, D Kostka, M Lara, AL Martins, T Massingham, I Moltke, BJ Raney, MD Rasmussen, J Robinson, A Stark, AJ Vilella, J Wen, X Xie, MC Zody, J Baldwin, T Bloom, C Whye Chin, D Heiman, R Nicol, C Nusbaum, S Young, J Wilkinson, KC Worley, CL Kovar, DM Muzny, RA Gibbs, A Cree, HH Dihn, G Fowler, S Jhangiani, V Joshi, S Lee, LR Lewis, LV Nazareth, G Okwuonu, J Santibanez, WC Warren, ER Mardis, GM Weinstock, RK Wilson, K Delehaunty, D Dooling, C Fronik, L Fulton, B Fulton, T Graves, P Minx, E Sodergren, E Birney, EH Margulies, J Herrero, ED Green, D Haussler, A Siepel, N Goldman, KS Pollard, JS Pedersen, ES Lander, M Kellis. A high-resolution map of human evolutionary constraint using 29 mammals. Nature 2011;478(7370):476–482. doi:10.1038/nature10530
[BibTeX] [Abstract]
The comparison of related genomes has emerged as a powerful lens for genome interpretation. Here we report the sequencing and comparative analysis of 29 eutherian genomes. We confirm that at least 5.5\% of the human genome has undergone purifying selection, and locate constrained elements covering ∼4.2\% of the genome. We use evolutionary signatures and comparisons with experimental data sets to suggest candidate functions for ∼60\% of constrained bases. These elements reveal a small number of new coding exons, candidate stop codon readthrough events and over 10,000 regions of overlapping synonymous constraint within protein-coding exons. We find 220 candidate RNA structural families, and nearly a million elements overlapping potential promoter, enhancer and insulator regions. We report specific amino acid residues that have undergone positive selection, 280,000 non-coding elements exapted from mobile elements and more than 1,000 primate- and human-accelerated elements. Overlap with disease-associated variants indicates that our findings will be relevant for studies of human biology, health and disease.
@Article{21993624, author = {Lindblad-Toh K and Garber M and Zuk O and Lin MF and Parker BJ and Washietl S and Kheradpour P and Ernst J and Jordan G and Mauceli E and Ward LD and Lowe CB and Holloway AK and Clamp M and Gnerre S and Alföldi J and Beal K and Chang J and Clawson H and Cuff J and Di Palma F and Fitzgerald S and Flicek P and Guttman M and Hubisz MJ and Jaffe DB and Jungreis I and Kent WJ and Kostka D and Lara M and Martins AL and Massingham T and Moltke I and Raney BJ and Rasmussen MD and Robinson J and Stark A and Vilella AJ and Wen J and Xie X and Zody MC and Baldwin J and Bloom T and Whye Chin C and Heiman D and Nicol R and Nusbaum C and Young S and Wilkinson J and Worley KC and Kovar CL and Muzny DM and Gibbs RA and Cree A and Dihn HH and Fowler G and Jhangiani S and Joshi V and Lee S and Lewis LR and Nazareth LV and Okwuonu G and Santibanez J and Warren WC and Mardis ER and Weinstock GM and Wilson RK and Delehaunty K and Dooling D and Fronik C and Fulton L and Fulton B and Graves T and Minx P and Sodergren E and Birney E and Margulies EH and Herrero J and Green ED and Haussler D and Siepel A and Goldman N and Pollard KS and Pedersen JS and Lander ES and Kellis M}, title = {A high-resolution map of human evolutionary constraint using 29 mammals}, journal = {Nature}, volume = {478}, number = {7370}, pages = {476--482}, year = {2011}, doi = {10.1038/nature10530}, abstract = {The comparison of related genomes has emerged as a powerful lens for genome interpretation. Here we report the sequencing and comparative analysis of 29 eutherian genomes. We confirm that at least 5.5\% of the human genome has undergone purifying selection, and locate constrained elements covering ∼4.2\% of the genome. We use evolutionary signatures and comparisons with experimental data sets to suggest candidate functions for ∼60\% of constrained bases. These elements reveal a small number of new coding exons, candidate stop codon readthrough events and over 10,000 regions of overlapping synonymous constraint within protein-coding exons. We find 220 candidate RNA structural families, and nearly a million elements overlapping potential promoter, enhancer and insulator regions. We report specific amino acid residues that have undergone positive selection, 280,000 non-coding elements exapted from mobile elements and more than 1,000 primate- and human-accelerated elements. Overlap with disease-associated variants indicates that our findings will be relevant for studies of human biology, health and disease.},} - D Lipman, P Flicek, S Salzberg, M Gerstein, R Knight. Closure of the NCBI SRA and implications for the long-term future of genomics data storage. Genome Biol 2011;12(3):402. doi:10.1186/gb-2011-12-3-402
[BibTeX]@Article{21418618, author = {Lipman D and Flicek P and Salzberg S and Gerstein M and Knight R}, title = {Closure of the NCBI SRA and implications for the long-term future of genomics data storage}, journal = {Genome Biol}, volume = {12}, number = {3}, pages = {402}, year = {2011}, doi = {10.1186/gb-2011-12-3-402}, } - DP Locke, LW Hillier, WC Warren, KC Worley, LV Nazareth, DM Muzny, S-P Yang, Z Wang, AT Chinwalla, P Minx, M Mitreva, L Cook, KD Delehaunty, C Fronick, H Schmidt, LA Fulton, RS Fulton, JO Nelson, V Magrini, C Pohl, TA Graves, C Markovic, A Cree, HH Dinh, J Hume, CL Kovar, GR Fowler, G Lunter, S Meader, A Heger, CP Ponting, T Marques-Bonet, C Alkan, L Chen, Z Cheng, JM Kidd, EE Eichler, S White, S Searle, AJ Vilella, Y Chen, P Flicek, J Ma, B Raney, B Suh, R Burhans, J Herrero, D Haussler, R Faria, O Fernando, F Darré, D Farré, E Gazave, M Oliva, A Navarro, R Roberto, O Capozzi, N Archidiacono, GD Valle, S Purgato, M Rocchi, MK Konkel, JA Walker, B Ullmer, MA Batzer, AFA Smit, R Hubley, C Casola, DR Schrider, MW Hahn, V Quesada, XS Puente, GR Ordoñez, C López-Otín, T Vinar, B Brejova, A Ratan, RS Harris, W Miller, C Kosiol, HA Lawson, V Taliwal, AL Martins, A Siepel, A Roychoudhury, X Ma, J Degenhardt, CD Bustamante, RN Gutenkunst, T Mailund, JY Dutheil, A Hobolth, MH Schierup, OA Ryder, Y Yoshinaga, de PJ Jong, GM Weinstock, J Rogers, ER Mardis, RA Gibbs, RK Wilson. Comparative and demographic analysis of orang-utan genomes. Nature 2011;469(7331):529–533. doi:10.1038/nature09687
[BibTeX] [Abstract]
'Orang-utan' is derived from a Malay term meaning `man of the forest' and aptly describes the southeast Asian great apes native to Sumatra and Borneo. The orang-utan species, Pongo abelii (Sumatran) and Pongo pygmaeus (Bornean), are the most phylogenetically distant great apes from humans, thereby providing an informative perspective on hominid evolution. Here we present a Sumatran orang-utan draft genome assembly and short read sequence data from five Sumatran and five Bornean orang-utan genomes. Our analyses reveal that, compared to other primates, the orang-utan genome has many unique features. Structural evolution of the orang-utan genome has proceeded much more slowly than other great apes, evidenced by fewer rearrangements, less segmental duplication, a lower rate of gene family turnover and surprisingly quiescent Alu repeats, which have played a major role in restructuring other primate genomes. We also describe a primate polymorphic neocentromere, found in both Pongo species, emphasizing the gradual evolution of orang-utan genome structure. Orang-utans have extremely low energy usage for a eutherian mammal, far lower than their hominid relatives. Adding their genome to the repertoire of sequenced primates illuminates new signals of positive selection in several pathways including glycolipid metabolism. From the population perspective, both Pongo species are deeply diverse; however, Sumatran individuals possess greater diversity than their Bornean counterparts, and more species-specific variation. Our estimate of Bornean/Sumatran speciation time, 400,000 years ago, is more recent than most previous studies and underscores the complexity of the orang-utan speciation process. Despite a smaller modern census population size, the Sumatran effective population size (N(e)) expanded exponentially relative to the ancestral N(e) after the split, while Bornean N(e) declined over the same period. Overall, the resources and analyses presented here offer new opportunities in evolutionary genomics, insights into hominid biology, and an extensive database of variation for conservation efforts.
@Article{21270892, author = {Locke DP and Hillier LW and Warren WC and Worley KC and Nazareth LV and Muzny DM and Yang S-P and Wang Z and Chinwalla AT and Minx P and Mitreva M and Cook L and Delehaunty KD and Fronick C and Schmidt H and Fulton LA and Fulton RS and Nelson JO and Magrini V and Pohl C and Graves TA and Markovic C and Cree A and Dinh HH and Hume J and Kovar CL and Fowler GR and Lunter G and Meader S and Heger A and Ponting CP and Marques-Bonet T and Alkan C and Chen L and Cheng Z and Kidd JM and Eichler EE and White S and Searle S and Vilella AJ and Chen Y and Flicek P and Ma J and Raney B and Suh B and Burhans R and Herrero J and Haussler D and Faria R and Fernando O and Darré F and Farré D and Gazave E and Oliva M and Navarro A and Roberto R and Capozzi O and Archidiacono N and Valle GD and Purgato S and Rocchi M and Konkel MK and Walker JA and Ullmer B and Batzer MA and Smit AFA and Hubley R and Casola C and Schrider DR and Hahn MW and Quesada V and Puente XS and Ordoñez GR and López-Otín C and Vinar T and Brejova B and Ratan A and Harris RS and Miller W and Kosiol C and Lawson HA and Taliwal V and Martins AL and Siepel A and Roychoudhury A and Ma X and Degenhardt J and Bustamante CD and Gutenkunst RN and Mailund T and Dutheil JY and Hobolth A and Schierup MH and Ryder OA and Yoshinaga Y and de Jong PJ and Weinstock GM and Rogers J and Mardis ER and Gibbs RA and Wilson RK}, title = {Comparative and demographic analysis of orang-utan genomes}, journal = {Nature}, volume = {469}, number = {7331}, pages = {529--533}, year = {2011}, doi = {10.1038/nature09687}, abstract = {'Orang-utan' is derived from a Malay term meaning `man of the forest' and aptly describes the southeast Asian great apes native to Sumatra and Borneo. The orang-utan species, Pongo abelii (Sumatran) and Pongo pygmaeus (Bornean), are the most phylogenetically distant great apes from humans, thereby providing an informative perspective on hominid evolution. Here we present a Sumatran orang-utan draft genome assembly and short read sequence data from five Sumatran and five Bornean orang-utan genomes. Our analyses reveal that, compared to other primates, the orang-utan genome has many unique features. Structural evolution of the orang-utan genome has proceeded much more slowly than other great apes, evidenced by fewer rearrangements, less segmental duplication, a lower rate of gene family turnover and surprisingly quiescent Alu repeats, which have played a major role in restructuring other primate genomes. We also describe a primate polymorphic neocentromere, found in both Pongo species, emphasizing the gradual evolution of orang-utan genome structure. Orang-utans have extremely low energy usage for a eutherian mammal, far lower than their hominid relatives. Adding their genome to the repertoire of sequenced primates illuminates new signals of positive selection in several pathways including glycolipid metabolism. From the population perspective, both Pongo species are deeply diverse; however, Sumatran individuals possess greater diversity than their Bornean counterparts, and more species-specific variation. Our estimate of Bornean/Sumatran speciation time, 400,000 years ago, is more recent than most previous studies and underscores the complexity of the orang-utan speciation process. Despite a smaller modern census population size, the Sumatran effective population size (N(e)) expanded exponentially relative to the ancestral N(e) after the split, while Bornean N(e) declined over the same period. Overall, the resources and analyses presented here offer new opportunities in evolutionary genomics, insights into hominid biology, and an extensive database of variation for conservation efforts.},} - AJ Vilella, E Birney, P Flicek, J Herrero. Considerations for the inclusion of 2x mammalian genomes in phylogenetic analyses. Genome Biol 2011;12(2):401. doi:10.1186/gb-2011-12-2-401
[BibTeX] [Abstract]
: A response to 2x genomes - depth does matter by MC Milinkovitch, R Helaers, E Depiereux, AC Tzika and T Gabaldón. Genome Biol 2010, 11:R16.
@Article{21320298, author = {Vilella AJ and Birney E and Flicek P and Herrero J}, title = {Considerations for the inclusion of 2x mammalian genomes in phylogenetic analyses}, journal = {Genome Biol}, volume = {12}, number = {2}, pages = {401}, year = {2011}, doi = {10.1186/gb-2011-12-2-401}, abstract = {: A response to 2x genomes - depth does matter by MC Milinkovitch, R Helaers, E Depiereux, AC Tzika and T Gabaldón. Genome Biol 2010, 11:R16.},} - P Flicek, MR Amode, D Barrell, K Beal, S Brent, Y Chen, P Clapham, G Coates, S Fairley, S Fitzgerald, L Gordon, M Hendrix, T Hourlier, N Johnson, A Kähäri, D Keefe, S Keenan, R Kinsella, F Kokocinski, E Kulesha, P Larsson, I Longden, W McLaren, B Overduin, B Pritchard, HS Riat, D Rios, GRS Ritchie, M Ruffier, M Schuster, D Sobral, G Spudich, YA Tang, S Trevanion, J Vandrovcova, AJ Vilella, S White, SP Wilder, A Zadissa, J Zamora, BL Aken, E Birney, F Cunningham, I Dunham, R Durbin, XM Fernández-Suarez, J Herrero, TJP Hubbard, A Parker, G Proctor, J Vogel, SMJ Searle. Ensembl 2011. Nucleic Acids Res 2011;39(Database issue):D800–6. doi:10.1093/nar/gkq1064
[BibTeX] [Abstract]
The Ensembl project (http://www.ensembl.org) seeks to enable genomic science by providing high quality, integrated annotation on chordate and selected eukaryotic genomes within a consistent and accessible infrastructure. All supported species include comprehensive, evidence-based gene annotations and a selected set of genomes includes additional data focused on variation, comparative, evolutionary, functional and regulatory annotation. The most advanced resources are provided for key species including human, mouse, rat and zebrafish reflecting the popularity and importance of these species in biomedical research. As of Ensembl release 59 (August 2010), 56 species are supported of which 5 have been added in the past year. Since our previous report, we have substantially improved the presentation and integration of both data of disease relevance and the regulatory state of different cell types.
@Article{21045057, author = {Flicek P and Amode MR and Barrell D and Beal K and Brent S and Chen Y and Clapham P and Coates G and Fairley S and Fitzgerald S and Gordon L and Hendrix M and Hourlier T and Johnson N and Kähäri A and Keefe D and Keenan S and Kinsella R and Kokocinski F and Kulesha E and Larsson P and Longden I and McLaren W and Overduin B and Pritchard B and Riat HS and Rios D and Ritchie GRS and Ruffier M and Schuster M and Sobral D and Spudich G and Tang YA and Trevanion S and Vandrovcova J and Vilella AJ and White S and Wilder SP and Zadissa A and Zamora J and Aken BL and Birney E and Cunningham F and Dunham I and Durbin R and Fernández-Suarez XM and Herrero J and Hubbard TJP and Parker A and Proctor G and Vogel J and Searle SMJ}, title = {Ensembl 2011}, journal = {Nucleic Acids Res}, volume = {39}, number = {Database issue}, pages = {D800--6}, year = {2011}, doi = {10.1093/nar/gkq1064}, howpublished = {Advanced online publication: 2 November 2010}, abstract = {The Ensembl project (http://www.ensembl.org) seeks to enable genomic science by providing high quality, integrated annotation on chordate and selected eukaryotic genomes within a consistent and accessible infrastructure. All supported species include comprehensive, evidence-based gene annotations and a selected set of genomes includes additional data focused on variation, comparative, evolutionary, functional and regulatory annotation. The most advanced resources are provided for key species including human, mouse, rat and zebrafish reflecting the popularity and importance of these species in biomedical research. As of Ensembl release 59 (August 2010), 56 species are supported of which 5 have been added in the past year. Since our previous report, we have substantially improved the presentation and integration of both data of disease relevance and the regulatory state of different cell types.},} - RJ Kinsella, A Kähäri, S Haider, J Zamora, G Proctor, G Spudich, J Almeida-King, D Staines, P Derwent, A Kerhornou, P Kersey, P Flicek. Ensembl BioMarts: a hub for data retrieval across taxonomic space. Database (Oxford) 2011;2011:bar030. doi:10.1093/database/bar030
[BibTeX] [Abstract]
For a number of years the BioMart data warehousing system has proven to be a valuable resource for scientists seeking a fast and versatile means of accessing the growing volume of genomic data provided by the Ensembl project. The launch of the Ensembl Genomes project in 2009 complemented the Ensembl project by utilizing the same visualization, interactive and programming tools to provide users with a means for accessing genome data from a further five domains: protists, bacteria, metazoa, plants and fungi. The Ensembl and Ensembl Genomes BioMarts provide a point of access to the high-quality gene annotation, variation data, functional and regulatory annotation and evolutionary relationships from genomes spanning the taxonomic space. This article aims to give a comprehensive overview of the Ensembl and Ensembl Genomes BioMarts as well as some useful examples and a description of current data content and future objectives. Database URLs: http://www.ensembl.org/biomart/martview/; http://metazoa.ensembl.org/biomart/martview/; http://plants.ensembl.org/biomart/martview/; http://protists.ensembl.org/biomart/martview/; http://fungi.ensembl.org/biomart/martview/; http://bacteria.ensembl.org/biomart/martview/
@Article{21785142, author = {Kinsella RJ and Kähäri A and Haider S and Zamora J and Proctor G and Spudich G and Almeida-King J and Staines D and Derwent P and Kerhornou A and Kersey P and Flicek P}, title = {Ensembl BioMarts: a hub for data retrieval across taxonomic space}, journal = {Database (Oxford)}, volume = {2011}, pages = {bar030}, year = {2011}, doi = {10.1093/database/bar030}, abstract = {For a number of years the BioMart data warehousing system has proven to be a valuable resource for scientists seeking a fast and versatile means of accessing the growing volume of genomic data provided by the Ensembl project. The launch of the Ensembl Genomes project in 2009 complemented the Ensembl project by utilizing the same visualization, interactive and programming tools to provide users with a means for accessing genome data from a further five domains: protists, bacteria, metazoa, plants and fungi. The Ensembl and Ensembl Genomes BioMarts provide a point of access to the high-quality gene annotation, variation data, functional and regulatory annotation and evolutionary relationships from genomes spanning the taxonomic space. This article aims to give a comprehensive overview of the Ensembl and Ensembl Genomes BioMarts as well as some useful examples and a description of current data content and future objectives. Database URLs: http://www.ensembl.org/biomart/martview/; http://metazoa.ensembl.org/biomart/martview/; http://plants.ensembl.org/biomart/martview/; http://protists.ensembl.org/biomart/martview/; http://fungi.ensembl.org/biomart/martview/; http://bacteria.ensembl.org/biomart/martview/},} - MB Renfree, AT Papenfuss, JE Deakin, J Lindsay, T Heider, K Belov, W Rens, PD Waters, EA Pharo, G Shaw, ES Wong, CM Lefevre, KR Nicholas, Y Kuroki, MJ Wakefield, KR Zenger, C Wang, M Ferguson-Smith, FW Nicholas, D Hickford, H Yu, KR Short, HV Siddle, SR Frankenberg, KY Chew, BR Menzies, JM Stringer, S Suzuki, TA Hore, ML Delbridge, A Mohammadi, NY Schneider, Y Hu, W O'Hara, S Al Nadaf, C Wu, Z-P Feng, BG Cocks, J Wang, P Flicek, SM Searle, S Fairley, K Beal, J Herrero, DM Carone, Y Suzuki, S Sagano, A Toyoda, Y Sakaki, S Kondo, Y Nishida, S Tatsumoto, I Mandiou, A Hsu, KA McColl, B Landsell, G Weinstock, E Kuczek, A McGrath, P Wilson, A Men, M Hazar-Rethinam, A Hall, J Davies, D Wood, S Williams, Y Sundaravadanam, DM Muzny, SN Jhangiani, LR Lewis, MB Morgan, GO Okwuonu, SJ Ruiz, J Santibanez, L Nazareth, A Cree, G Fowler, CL Kovar, HH Dinh, V Joshi, C Jing, F Lara, R Thornton, L Chen, J Deng, Y Liu, JY Shen, X-Z Song, J Edson, C Troon, D Thomas, A Stephens, L Yapa, T Levchenko, RA Gibbs, DW Cooper, TP Speed, A Fujiyama, JA Graves, RJ O'Neill, AJ Pask, SM Forrest, KC Worley. Genome sequence of an Australian kangaroo, Macropus eugenii, provides insight into the evolution of mammalian reproduction and development. Genome Biol 2011;12(8):R81. doi:10.1186/gb-2011-12-8-r81
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: We present the genome sequence of the tammar wallaby, Macropus eugenii, which is a member of the kangaroo family and the first representative of the iconic hopping mammals that symbolize Australia to be sequenced. The tammar has many unusual biological characteristics, including the longest period of embryonic diapause of any mammal, extremely synchronized seasonal breeding and prolonged and sophisticated lactation within a well-defined pouch. Like other marsupials, it gives birth to highly altricial young, and has a small number of very large chromosomes, making it a valuable model for genomics, reproduction and development. RESULTS: The genome has been sequenced to 2x coverage using Sanger sequencing, enhanced with additional next generation sequencing and the integration of extensive physical and linkage maps to build the genome assembly. We also sequenced the tammar transcriptome across many tissues and developmental time points. Our analyses of these data shed light on mammalian reproduction, development and genome evolution: there is innovation in reproductive and lactational genes, rapid evolution of germ cell genes, and incomplete, locus-specific X inactivation. We also observe novel retrotransposons and a highly rearranged major histocompatibility complex, with many class I genes located outside the complex. Novel microRNAs in the tammar HOX clusters uncover new potential mammalian HOX regulatory elements. CONCLUSIONS: Analyses of these resources enhance our understanding of marsupial gene evolution, identify marsupial-specific conserved non-coding elements and critical genes across a range of biological systems, including reproduction, development and immunity, and provide new insight into marsupial and mammalian biology and genome evolution.
@Article{21854559, author = {Renfree MB and Papenfuss AT and Deakin JE and Lindsay J and Heider T and Belov K and Rens W and Waters PD and Pharo EA and Shaw G and Wong ES and Lefevre CM and Nicholas KR and Kuroki Y and Wakefield MJ and Zenger KR and Wang C and Ferguson-Smith M and Nicholas FW and Hickford D and Yu H and Short KR and Siddle HV and Frankenberg SR and Chew KY and Menzies BR and Stringer JM and Suzuki S and Hore TA and Delbridge ML and Mohammadi A and Schneider NY and Hu Y and O'Hara W and Al Nadaf S and Wu C and Feng Z-P and Cocks BG and Wang J and Flicek P and Searle SM and Fairley S and Beal K and Herrero J and Carone DM and Suzuki Y and Sagano S and Toyoda A and Sakaki Y and Kondo S and Nishida Y and Tatsumoto S and Mandiou I and Hsu A and McColl KA and Landsell B and Weinstock G and Kuczek E and McGrath A and Wilson P and Men A and Hazar-Rethinam M and Hall A and Davies J and Wood D and Williams S and Sundaravadanam Y and Muzny DM and Jhangiani SN and Lewis LR and Morgan MB and Okwuonu GO and Ruiz SJ and Santibanez J and Nazareth L and Cree A and Fowler G and Kovar CL and Dinh HH and Joshi V and Jing C and Lara F and Thornton R and Chen L and Deng J and Liu Y and Shen JY and Song X-Z and Edson J and Troon C and Thomas D and Stephens A and Yapa L and Levchenko T and Gibbs RA and Cooper DW and Speed TP and Fujiyama A and Graves JA and O'Neill RJ and Pask AJ and Forrest SM and Worley KC}, title = {Genome sequence of an Australian kangaroo, Macropus eugenii, provides insight into the evolution of mammalian reproduction and development}, journal = {Genome Biol}, volume = {12}, number = {8}, pages = {R81}, year = {2011}, doi = {10.1186/gb-2011-12-8-r81}, abstract = {ABSTRACT: BACKGROUND: We present the genome sequence of the tammar wallaby, Macropus eugenii, which is a member of the kangaroo family and the first representative of the iconic hopping mammals that symbolize Australia to be sequenced. The tammar has many unusual biological characteristics, including the longest period of embryonic diapause of any mammal, extremely synchronized seasonal breeding and prolonged and sophisticated lactation within a well-defined pouch. Like other marsupials, it gives birth to highly altricial young, and has a small number of very large chromosomes, making it a valuable model for genomics, reproduction and development. RESULTS: The genome has been sequenced to 2x coverage using Sanger sequencing, enhanced with additional next generation sequencing and the integration of extensive physical and linkage maps to build the genome assembly. We also sequenced the tammar transcriptome across many tissues and developmental time points. Our analyses of these data shed light on mammalian reproduction, development and genome evolution: there is innovation in reproductive and lactational genes, rapid evolution of germ cell genes, and incomplete, locus-specific X inactivation. We also observe novel retrotransposons and a highly rearranged major histocompatibility complex, with many class I genes located outside the complex. Novel microRNAs in the tammar HOX clusters uncover new potential mammalian HOX regulatory elements. CONCLUSIONS: Analyses of these resources enhance our understanding of marsupial gene evolution, identify marsupial-specific conserved non-coding elements and critical genes across a range of biological systems, including reproduction, development and immunity, and provide new insight into marsupial and mammalian biology and genome evolution.},} - DM Church, VA Schneider, T Graves, K Auger, F Cunningham, N Bouk, H-C Chen, R Agarwala, WM McLaren, GRS Ritchie, D Albracht, M Kremitzki, S Rock, H Kotkiewicz, C Kremitzki, A Wollam, L Trani, L Fulton, R Fulton, L Matthews, S Whitehead, W Chow, J Torrance, M Dunn, G Harden, G Threadgold, J Wood, J Collins, P Heath, G Griffiths, S Pelan, D Grafham, E Eichler E, G Weinstock, ER Mardis, RK Wilson, K Howe, P Flicek, T Hubbard. Modernizing reference genome assemblies. PLoS Biol 2011;9(7):e1001091. doi:10.1371/journal.pbio.1001091
[BibTeX]@Article{21750661, author = {Church DM and Schneider VA and Graves T and Auger K and Cunningham F and Bouk N and Chen H-C and Agarwala R and McLaren WM and Ritchie GRS and Albracht D and Kremitzki M and Rock S and Kotkiewicz H and Kremitzki C and Wollam A and Trani L and Fulton L and Fulton R and Matthews L and Whitehead S and Chow W and Torrance J and Dunn M and Harden G and Threadgold G and Wood J and Collins J and Heath P and Griffiths G and Pelan S and Grafham D and E Eichler E and Weinstock G and Mardis ER and Wilson RK and Howe K and Flicek P and Hubbard T}, title = {Modernizing reference genome assemblies}, journal = {PLoS Biol}, volume = {9}, number = {7}, pages = {e1001091}, year = {2011}, doi = {10.1371/journal.pbio.1001091}, } - GT Marth, F Yu, AR Indap, K Garimella, S Gravel, WF Leong, C Tyler-Smith, M Bainbridge, T Blackwell, X Zheng-Bradley, Y Chen, D Challis, L Clarke, EV Ball, K Cibulskis, DN Cooper, B Fulton, C Hartl, D Koboldt, D Muzny, R Smith, C Sougnez, C Stewart, A Ward, J Yu, Y Xue, D Altshuler, CD Bustamante, AG Clark, M Daly, M Depristo, P Flicek, S Gabriel, E Mardis, A Palotie, RA Gibbs, T Genomes Project 1000. The functional spectrum of low-frequency coding variation. Genome Biol 2011;12(9):R84. doi:10.1186/gb-2011-12-9-r84
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: Rare coding variants constitute an important class of human genetic variation, but are underrepresented in current databases that are based on small population samples. Recent studies show that variants altering amino acid sequence and protein function are enriched at low variant allele frequency, 2-5\%, but because of insufficient sample size it is not clear if the same trend holds for rare variants below 1\% allele frequency. RESULTS: The 1000 Genomes Exon Pilot Project has collected deep-coverage exon-capture data in roughly 1,000 human genes, for nearly 700 samples. Although medical whole-exome projects are currently afoot, this is still the deepest reported sampling of a large number of human genes with next-generation technologies. According to the goals of the 1000 Genomes Project, we created effective informatics pipelines to process and analyze the data, and discovered 12,758 exonic SNPs, 70\% of them novel, and 74\% below 1\% allele frequency in the seven population samples we examined. Our analysis confirms that coding variants below 1\% allele frequency show increased population-specificity and are enriched for functional variants. CONCLUSIONS: This study represents a large step toward detecting and interpreting low frequency coding variation: clearly lays out technical steps for effective analysis of DNA capture data, and articulates functional and population properties of this important class of genetic variation.
@Article{21917140, author = {Marth GT and Yu F and Indap AR and Garimella K and Gravel S and Leong WF and Tyler-Smith C and Bainbridge M and Blackwell T and Zheng-Bradley X and Chen Y and Challis D and Clarke L and Ball EV and Cibulskis K and Cooper DN and Fulton B and Hartl C and Koboldt D and Muzny D and Smith R and Sougnez C and Stewart C and Ward A and Yu J and Xue Y and Altshuler D and Bustamante CD and Clark AG and Daly M and Depristo M and Flicek P and Gabriel S and Mardis E and Palotie A and Gibbs RA and 1000 Genomes Project T}, title = {The functional spectrum of low-frequency coding variation}, journal = {Genome Biol}, volume = {12}, number = {9}, pages = {R84}, year = {2011}, doi = {10.1186/gb-2011-12-9-r84}, abstract = {ABSTRACT: BACKGROUND: Rare coding variants constitute an important class of human genetic variation, but are underrepresented in current databases that are based on small population samples. Recent studies show that variants altering amino acid sequence and protein function are enriched at low variant allele frequency, 2-5\%, but because of insufficient sample size it is not clear if the same trend holds for rare variants below 1\% allele frequency. RESULTS: The 1000 Genomes Exon Pilot Project has collected deep-coverage exon-capture data in roughly 1,000 human genes, for nearly 700 samples. Although medical whole-exome projects are currently afoot, this is still the deepest reported sampling of a large number of human genes with next-generation technologies. According to the goals of the 1000 Genomes Project, we created effective informatics pipelines to process and analyze the data, and discovered 12,758 exonic SNPs, 70\% of them novel, and 74\% below 1\% allele frequency in the seven population samples we examined. Our analysis confirms that coding variants below 1\% allele frequency show increased population-specificity and are enriched for functional variants. CONCLUSIONS: This study represents a large step toward detecting and interpreting low frequency coding variation: clearly lays out technical steps for effective analysis of DNA capture data, and articulates functional and population properties of this important class of genetic variation.},}
2010
- D Schmidt, PC Schwalie, CS Ross-Innes, A Hurtado, GD Brown, JS Carroll, P Flicek, DT Odom. A CTCF-independent role for cohesin in tissue-specific transcription. Genome Res 2010;20(5):578–588. doi:10.1101/gr.100479.109
[BibTeX] [Abstract]
The cohesin protein complex holds sister chromatids in dividing cells together and is essential for chromosome segregation. Recently, cohesin has been implicated in mediating transcriptional insulation, via its interactions with CTCF. Here, we show in different cell types that cohesin functionally behaves as a tissue-specific transcriptional regulator, independent of CTCF binding. By performing matched genome-wide binding assays (ChIP-seq) in human breast cancer cells (MCF-7), we discovered thousands of genomic sites that share cohesin and estrogen receptor alpha (ER) yet lack CTCF binding. By use of human hepatocellular carcinoma cells (HepG2), we found that liver-specific transcription factors colocalize with cohesin independently of CTCF at liver-specific targets that are distinct from those found in breast cancer cells. Furthermore, estrogen-regulated genes are preferentially bound by both ER and cohesin, and functionally, the silencing of cohesin caused aberrant re-entry of breast cancer cells into cell cycle after hormone treatment. We combined chromosomal interaction data in MCF-7 cells with our cohesin binding data to show that cohesin is highly enriched at ER-bound regions that capture inter-chromosomal loop anchors. Together, our data show that cohesin cobinds across the genome with transcription factors independently of CTCF, plays a functional role in estrogen-regulated transcription, and may help to mediate tissue-specific transcriptional responses via long-range chromosomal interactions.
@Article{20219941, author = {Schmidt D and Schwalie PC and Ross-Innes CS and Hurtado A and Brown GD and Carroll JS and Flicek P and Odom DT}, title = {A CTCF-independent role for cohesin in tissue-specific transcription}, journal = {Genome Res}, volume = {20}, number = {5}, pages = {578--588}, year = {2010}, doi = {10.1101/gr.100479.109}, howpublished = {Advanced online publication: 10 March 2010}, abstract = {The cohesin protein complex holds sister chromatids in dividing cells together and is essential for chromosome segregation. Recently, cohesin has been implicated in mediating transcriptional insulation, via its interactions with CTCF. Here, we show in different cell types that cohesin functionally behaves as a tissue-specific transcriptional regulator, independent of CTCF binding. By performing matched genome-wide binding assays (ChIP-seq) in human breast cancer cells (MCF-7), we discovered thousands of genomic sites that share cohesin and estrogen receptor alpha (ER) yet lack CTCF binding. By use of human hepatocellular carcinoma cells (HepG2), we found that liver-specific transcription factors colocalize with cohesin independently of CTCF at liver-specific targets that are distinct from those found in breast cancer cells. Furthermore, estrogen-regulated genes are preferentially bound by both ER and cohesin, and functionally, the silencing of cohesin caused aberrant re-entry of breast cancer cells into cell cycle after hormone treatment. We combined chromosomal interaction data in MCF-7 cells with our cohesin binding data to show that cohesin is highly enriched at ER-bound regions that capture inter-chromosomal loop anchors. Together, our data show that cohesin cobinds across the genome with transcription factors independently of CTCF, plays a functional role in estrogen-regulated transcription, and may help to mediate tissue-specific transcriptional responses via long-range chromosomal interactions.},} - D Rios, WM McLaren, Y Chen, E Birney, A Stabenau, P Flicek, F Cunningham. A database and API for variation, dense genotyping and resequencing data. BMC Bioinformatics 2010;11(1):238. doi:10.1186/1471-2105-11-238
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: Advances in sequencing and genotyping technologies are leading to the widespread availability of multi-species variation data, dense genotype data and large-scale resequencing projects. The 1000 Genomes Project and similar efforts in other species are challenging the methods previously used for storage and manipulation of such data necessitating the redesign of existing genome-wide bioinformatics resources. RESULTS: Ensembl has created a database and software library to support data storage, analysis and access to the existing and emerging variation data from large mammalian and vertebrate genomes. These tools scale to thousands of individual genome sequences and are integrated into the Ensembl infrastructure for genome annotation and visualisation. The database and software system is easily expanded to integrate both public and non-public data sources in the context of an Ensembl software installation and is already being used outside of the Ensembl project in a number of database and application environments. CONCLUSIONS: Ensembl's powerful, flexible and open source infrastructure for the management of variation, genotyping and resequencing data is freely available at http://www.ensembl.org.
@Article{20459810, author = {Rios D and McLaren WM and Chen Y and Birney E and Stabenau A and Flicek P and Cunningham F}, title = {A database and API for variation, dense genotyping and resequencing data}, journal = {BMC Bioinformatics}, volume = {11}, number = {1}, pages = {238}, year = {2010}, doi = {10.1186/1471-2105-11-238}, abstract = {ABSTRACT: BACKGROUND: Advances in sequencing and genotyping technologies are leading to the widespread availability of multi-species variation data, dense genotype data and large-scale resequencing projects. The 1000 Genomes Project and similar efforts in other species are challenging the methods previously used for storage and manipulation of such data necessitating the redesign of existing genome-wide bioinformatics resources. RESULTS: Ensembl has created a database and software library to support data storage, analysis and access to the existing and emerging variation data from large mammalian and vertebrate genomes. These tools scale to thousands of individual genome sequences and are integrated into the Ensembl infrastructure for genome annotation and visualisation. The database and software system is easily expanded to integrate both public and non-public data sources in the context of an Ensembl software installation and is already being used outside of the Ensembl project in a number of database and application environments. CONCLUSIONS: Ensembl's powerful, flexible and open source infrastructure for the management of variation, genotyping and resequencing data is freely available at http://www.ensembl.org.},} - The 1000 Genomes Project Consortium.. A map of human genome variation from population-scale sequencing. Nature 2010;467(7319):1061–1073. doi:10.1038/nature09534
[BibTeX] [Abstract]
The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother-father-child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95\% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately 10(-8) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.
@Article{20981092, author = {{The 1000 Genomes Project Consortium.}}, title = {A map of human genome variation from population-scale sequencing}, journal = {Nature}, volume = {467}, number = {7319}, pages = {1061--1073}, year = {2010}, doi = {10.1038/nature09534}, abstract = {The 1000 Genomes Project aims to provide a deep characterization of human genome sequence variation as a foundation for investigating the relationship between genotype and phenotype. Here we present results of the pilot phase of the project, designed to develop and compare different strategies for genome-wide sequencing with high-throughput platforms. We undertook three projects: low-coverage whole-genome sequencing of 179 individuals from four populations; high-coverage sequencing of two mother-father-child trios; and exon-targeted sequencing of 697 individuals from seven populations. We describe the location, allele frequency and local haplotype structure of approximately 15 million single nucleotide polymorphisms, 1 million short insertions and deletions, and 20,000 structural variants, most of which were previously undescribed. We show that, because we have catalogued the vast majority of common variation, over 95\% of the currently accessible variants found in any individual are present in this data set. On average, each person is found to carry approximately 250 to 300 loss-of-function variants in annotated genes and 50 to 100 variants previously implicated in inherited disorders. We demonstrate how these results can be used to inform association and functional studies. From the two trios, we directly estimate the rate of de novo germline base substitution mutations to be approximately 10(-8) per base pair per generation. We explore the data with regard to signatures of natural selection, and identify a marked reduction of genetic variation in the neighbourhood of genes, due to selection at linked sites. These methods and public data will support the next phase of human genetic research.},} - MG Reese, B Moore, C Batchelor, F Salas, F Cunningham, G Marth, L Stein, P Flicek, M Yandell, K Eilbeck. A standard variation file format for human genome sequences. Genome Biol 2010;11(8):R88. doi:10.1186/gb-2010-11-8-r88
[BibTeX] [Abstract]
ABSTRACT: Here we describe the Genome Variation Format (GVF) and the 10Gen dataset. GVF, an extension of Generic Feature Format version 3 (GFF3), is a simple tab-delimited format for DNA variant files, which uses Sequence Ontology to describe genome variation data. The 10Gen dataset, ten human genomes in GVF format, is freely available for community analysis from the Sequence Ontology website and from an Amazon elastic block storage (EBS) snapshot for use in Amazon's EC2 cloud computing environment.
@Article{20796305, author = {Reese MG and Moore B and Batchelor C and Salas F and Cunningham F and Marth G and Stein L and Flicek P and Yandell M and Eilbeck K}, title = {A standard variation file format for human genome sequences}, journal = {Genome Biol}, volume = {11}, number = {8}, pages = {R88}, year = {2010}, doi = {10.1186/gb-2010-11-8-r88}, abstract = {ABSTRACT: Here we describe the Genome Variation Format (GVF) and the 10Gen dataset. GVF, an extension of Generic Feature Format version 3 (GFF3), is a simple tab-delimited format for DNA variant files, which uses Sequence Ontology to describe genome variation data. The 10Gen dataset, ten human genomes in GVF format, is freely available for community analysis from the Sequence Ontology website and from an Amazon elastic block storage (EBS) snapshot for use in Amazon's EC2 cloud computing environment.},} - MP Schnetz, L Handoko, B Akhtar-Zaidi, CF Bartels, CF Pereira, AG Fisher, DJ Adams, P Flicek, GE Crawford, T Laframboise, P Tesar, C-L Wei, PC Scacheri. CHD7 targets active gene enhancer elements to modulate ES cell-specific gene expression. PLoS Genet 2010;6(7):e1001023. doi:10.1371/journal.pgen.1001023
[BibTeX] [Abstract]
CHD7 is one of nine members of the chromodomain helicase DNA-binding domain family of ATP-dependent chromatin remodeling enzymes found in mammalian cells. De novo mutation of CHD7 is a major cause of CHARGE syndrome, a genetic condition characterized by multiple congenital anomalies. To gain insights to the function of CHD7, we used the technique of chromatin immunoprecipitation followed by massively parallel DNA sequencing (ChIP-Seq) to map CHD7 sites in mouse ES cells. We identified 10,483 sites on chromatin bound by CHD7 at high confidence. Most of the CHD7 sites show features of gene enhancer elements. Specifically, CHD7 sites are predominantly located distal to transcription start sites, contain high levels of H3K4 mono-methylation, found within open chromatin that is hypersensitive to DNase I digestion, and correlate with ES cell-specific gene expression. Moreover, CHD7 co-localizes with P300, a known enhancer-binding protein and strong predictor of enhancer activity. Correlations with 18 other factors mapped by ChIP-seq in mouse ES cells indicate that CHD7 also co-localizes with ES cell master regulators OCT4, SOX2, and NANOG. Correlations between CHD7 sites and global gene expression profiles obtained from Chd7(+/+), Chd7(+/-), and Chd7(-/-) ES cells indicate that CHD7 functions at enhancers as a transcriptional rheostat to modulate, or fine-tune the expression levels of ES-specific genes. CHD7 can modulate genes in either the positive or negative direction, although negative regulation appears to be the more direct effect of CHD7 binding. These data indicate that enhancer-binding proteins can limit gene expression and are not necessarily co-activators. Although ES cells are not likely to be affected in CHARGE syndrome, we propose that enhancer-mediated gene dysregulation contributes to disease pathogenesis and that the critical CHD7 target genes may be subject to positive or negative regulation.
@Article{20657823, author = {Schnetz MP and Handoko L and Akhtar-Zaidi B and Bartels CF and Pereira CF and Fisher AG and Adams DJ and Flicek P and Crawford GE and Laframboise T and Tesar P and Wei C-L and Scacheri PC}, title = {CHD7 targets active gene enhancer elements to modulate ES cell-specific gene expression}, journal = {PLoS Genet}, volume = {6}, number = {7}, pages = {e1001023}, year = {2010}, doi = {10.1371/journal.pgen.1001023}, abstract = {CHD7 is one of nine members of the chromodomain helicase DNA-binding domain family of ATP-dependent chromatin remodeling enzymes found in mammalian cells. De novo mutation of CHD7 is a major cause of CHARGE syndrome, a genetic condition characterized by multiple congenital anomalies. To gain insights to the function of CHD7, we used the technique of chromatin immunoprecipitation followed by massively parallel DNA sequencing (ChIP-Seq) to map CHD7 sites in mouse ES cells. We identified 10,483 sites on chromatin bound by CHD7 at high confidence. Most of the CHD7 sites show features of gene enhancer elements. Specifically, CHD7 sites are predominantly located distal to transcription start sites, contain high levels of H3K4 mono-methylation, found within open chromatin that is hypersensitive to DNase I digestion, and correlate with ES cell-specific gene expression. Moreover, CHD7 co-localizes with P300, a known enhancer-binding protein and strong predictor of enhancer activity. Correlations with 18 other factors mapped by ChIP-seq in mouse ES cells indicate that CHD7 also co-localizes with ES cell master regulators OCT4, SOX2, and NANOG. Correlations between CHD7 sites and global gene expression profiles obtained from Chd7(+/+), Chd7(+/-), and Chd7(-/-) ES cells indicate that CHD7 functions at enhancers as a transcriptional rheostat to modulate, or fine-tune the expression levels of ES-specific genes. CHD7 can modulate genes in either the positive or negative direction, although negative regulation appears to be the more direct effect of CHD7 binding. These data indicate that enhancer-binding proteins can limit gene expression and are not necessarily co-activators. Although ES cells are not likely to be affected in CHARGE syndrome, we propose that enhancer-mediated gene dysregulation contributes to disease pathogenesis and that the critical CHD7 target genes may be subject to positive or negative regulation.},} - B Ballester, N Johnson, G Proctor, P Flicek. Consistent annotation of gene expression arrays. BMC Genomics 2010;11(1):294. doi:10.1186/1471-2164-11-294
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: Gene expression arrays are valuable and widely used tools for biomedical research. Today's commercial arrays attempt to measure the expression level of all of the genes in the genome. Effectively translating the results from the microarray into a biological interpretation requires an accurate mapping between the probesets on the array and the genes that they are targeting. Although major array manufacturers provide annotations of their gene expression arrays, the methods used by various manufacturers are different and the annotations are difficult to keep up to date in the rapidly changing world of biological sequence databases. RESULTS: We have created a consistent microarray annotation protocol applicable to all of the major array manufacturers. We constantly keep our annotations updated with the latest Ensembl Gene predictions, and thus cross-referenced with a large number of external biomedical sequence database identifiers. We show that these annotations are accurate and address in detail reasons for the minority of probesets that cannot be annotated. Annotations are publicly accessible through the Ensembl Genome Browser and programmatically through the Ensembl Application Programming Interface. They are also seamlessly integrated into the BioMart data-mining tool and the biomaRt package of BioConductor. CONCLUSIONS: Consistent, accurate and updated gene expression array annotations remain critical for biological research. Our annotations facilitate accurate biological interpretation of gene expression profiles.
@Article{20459806, author = {Ballester B and Johnson N and Proctor G and Flicek P}, title = {Consistent annotation of gene expression arrays}, journal = {BMC Genomics}, volume = {11}, number = {1}, pages = {294}, year = {2010}, doi = {10.1186/1471-2164-11-294}, abstract = {ABSTRACT: BACKGROUND: Gene expression arrays are valuable and widely used tools for biomedical research. Today's commercial arrays attempt to measure the expression level of all of the genes in the genome. Effectively translating the results from the microarray into a biological interpretation requires an accurate mapping between the probesets on the array and the genes that they are targeting. Although major array manufacturers provide annotations of their gene expression arrays, the methods used by various manufacturers are different and the annotations are difficult to keep up to date in the rapidly changing world of biological sequence databases. RESULTS: We have created a consistent microarray annotation protocol applicable to all of the major array manufacturers. We constantly keep our annotations updated with the latest Ensembl Gene predictions, and thus cross-referenced with a large number of external biomedical sequence database identifiers. We show that these annotations are accurate and address in detail reasons for the minority of probesets that cannot be annotated. Annotations are publicly accessible through the Ensembl Genome Browser and programmatically through the Ensembl Application Programming Interface. They are also seamlessly integrated into the BioMart data-mining tool and the biomaRt package of BioConductor. CONCLUSIONS: Consistent, accurate and updated gene expression array annotations remain critical for biological research. Our annotations facilitate accurate biological interpretation of gene expression profiles.},} - W McLaren, B Pritchard, D Rios, Y Chen, P Flicek, F Cunningham. Deriving the consequences of genomic variants with the Ensembl API and SNP Effect Predictor. Bioinformatics 2010;26(16):2069–2070. doi:10.1093/bioinformatics/btq330
[BibTeX] [Abstract]
SUMMARY: A tool to predict the effect newly discovered genomic variants have on known transcripts is indispensible in prioritizing and categorizing such variants. In Ensembl a web-based tool (the SNP Effect Predictor) and API interface can now functionally annotate variants in all Ensembl and Ensembl Genomes supported species. AVAILABILITY: The Ensembl SNP Effect Predictor can be accessed via the Ensembl website at http://www.ensembl.org/. The Ensembl API (see Ensembl A for installation instructions) is open source software. CONTACT: wm2@ebi.ac.uk, fiona@ebi.ac.uk.
@Article{20562413, author = {McLaren W and Pritchard B and Rios D and Chen Y and Flicek P and Cunningham F}, title = {Deriving the consequences of genomic variants with the Ensembl API and SNP Effect Predictor}, journal = {Bioinformatics}, volume = {26}, number = {16}, pages = {2069--2070}, year = {2010}, doi = {10.1093/bioinformatics/btq330}, abstract = {SUMMARY: A tool to predict the effect newly discovered genomic variants have on known transcripts is indispensible in prioritizing and categorizing such variants. In Ensembl a web-based tool (the SNP Effect Predictor) and API interface can now functionally annotate variants in all Ensembl and Ensembl Genomes supported species. AVAILABILITY: The Ensembl SNP Effect Predictor can be accessed via the Ensembl website at http://www.ensembl.org/. The Ensembl API (see Ensembl A for installation instructions) is open source software. CONTACT: wm2@ebi.ac.uk, fiona@ebi.ac.uk.},} - J Severin, K Beal, AJ Vilella, S Fitzgerald, M Schuster, L Gordon, A Ureta-Vidal, P Flicek, J Herrero. eHive: An Artificial Intelligence workflow system for genomic analysis. BMC Bioinformatics 2010;11(1):240. doi:10.1186/1471-2105-11-240
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: The Ensembl project produces updates to its comparative genomics resources with each of its several releases per year. During each release cycle approximately two weeks are allocated to generate all the genomic alignments and the protein homology predictions. The number of calculations required for this task grows approximately quadratically with the number of species. We currently support 51 species in Ensembl and we expect the number to continue to grow in the future. RESULTS: We present eHive, a new fault tolerant distributed processing system initially designed to support comparative genomic analysis, based on blackboard systems, network distributed autonomous agents, dataflow graphs and block-branch diagrams. In the eHive system a MySQL database serves as the central blackboard and the autonomous agent, a Perl script, queries the system and runs jobs as required. The system allows us to define dataflow and branching rules to suit all our production pipelines. We describe the implementation of three pipelines: (1) pairwise whole genome alignments, (2) multiple whole genome alignments and (3) gene trees with protein homology inference. Finally, we show the efficiency of the system in real case scenarios. CONCLUSIONS: eHive allows us to produce computationally demanding results in a reliable and efficient way with minimal supervision and high throughput. Further documentation is available at: http://www.ensembl.org/info/docs/eHive/
@Article{20459813, author = {Severin J and Beal K and Vilella AJ and Fitzgerald S and Schuster M and Gordon L and Ureta-Vidal A and Flicek P and Herrero J}, title = {eHive: An Artificial Intelligence workflow system for genomic analysis}, journal = {BMC Bioinformatics}, volume = {11}, number = {1}, pages = {240}, year = {2010}, doi = {10.1186/1471-2105-11-240}, abstract = {ABSTRACT: BACKGROUND: The Ensembl project produces updates to its comparative genomics resources with each of its several releases per year. During each release cycle approximately two weeks are allocated to generate all the genomic alignments and the protein homology predictions. The number of calculations required for this task grows approximately quadratically with the number of species. We currently support 51 species in Ensembl and we expect the number to continue to grow in the future. RESULTS: We present eHive, a new fault tolerant distributed processing system initially designed to support comparative genomic analysis, based on blackboard systems, network distributed autonomous agents, dataflow graphs and block-branch diagrams. In the eHive system a MySQL database serves as the central blackboard and the autonomous agent, a Perl script, queries the system and runs jobs as required. The system allows us to define dataflow and branching rules to suit all our production pipelines. We describe the implementation of three pipelines: (1) pairwise whole genome alignments, (2) multiple whole genome alignments and (3) gene trees with protein homology inference. Finally, we show the efficiency of the system in real case scenarios. CONCLUSIONS: eHive allows us to produce computationally demanding results in a reliable and efficient way with minimal supervision and high throughput. Further documentation is available at: http://www.ensembl.org/info/docs/eHive/},} - Y Chen, F Cunningham, D Rios, WM McLaren, J Smith, B Pritchard, GM Spudich, S Brent, E Kulesha, P Marin-Garcia, D Smedley, E Birney, P Flicek. Ensembl Variation Resources. BMC Genomics 2010;11(1):293. doi:10.1186/1471-2164-11-293
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: The maturing field of genomics is rapidly increasing the number of sequenced genomes and producing more information from those previously sequenced. Much of this additional information is variation data derived from sampling multiple individuals of a given species with the goal of discovering new variants and characterising the population frequencies of the variants that are already known. These data have immense value for many studies, including those designed to understand evolution and connect genotype to phenotype. Maximising the utility of the data requires that it be stored in an accessible manner that facilitates the integration of variation data with other genome resources such as gene annotation and comparative genomics. Description: The Ensembl project provides comprehensive and integrated variation resources for a wide variety of chordate genomes. This paper provides a detailed description of the sources of data and the methods for creating the Ensembl variation databases. It also explores the utility of the information by explaining the range of query options available, from using interactive web displays, to online data mining tools and connecting directly to the data servers programmatically. It gives a good overview of the variation resources and future plans for expanding the variation data within Ensembl. CONCLUSIONS: Variation data is an important key to understanding the functional and phenotypic differences between individuals. The development of new sequencing and genotyping technologies is greatly increasing the amount of variation data known for almost all genomes. The Ensembl variation resources are integrated into the Ensembl genome browser and provide a comprehensive way to access this data in the context of a widely used genome bioinformatics system. All Ensembl data is freely available at http://www.ensembl.org and from the public MySQL database server at ensembldb.ensembl.org.
@Article{20459805, author = {Chen Y and Cunningham F and Rios D and McLaren WM and Smith J and Pritchard B and Spudich GM and Brent S and Kulesha E and Marin-Garcia P and Smedley D and Birney E and Flicek P}, title = {Ensembl Variation Resources}, journal = {BMC Genomics}, volume = {11}, number = {1}, pages = {293}, year = {2010}, doi = {10.1186/1471-2164-11-293}, abstract = {ABSTRACT: BACKGROUND: The maturing field of genomics is rapidly increasing the number of sequenced genomes and producing more information from those previously sequenced. Much of this additional information is variation data derived from sampling multiple individuals of a given species with the goal of discovering new variants and characterising the population frequencies of the variants that are already known. These data have immense value for many studies, including those designed to understand evolution and connect genotype to phenotype. Maximising the utility of the data requires that it be stored in an accessible manner that facilitates the integration of variation data with other genome resources such as gene annotation and comparative genomics. Description: The Ensembl project provides comprehensive and integrated variation resources for a wide variety of chordate genomes. This paper provides a detailed description of the sources of data and the methods for creating the Ensembl variation databases. It also explores the utility of the information by explaining the range of query options available, from using interactive web displays, to online data mining tools and connecting directly to the data servers programmatically. It gives a good overview of the variation resources and future plans for expanding the variation data within Ensembl. CONCLUSIONS: Variation data is an important key to understanding the functional and phenotypic differences between individuals. The development of new sequencing and genotyping technologies is greatly increasing the amount of variation data known for almost all genomes. The Ensembl variation resources are integrated into the Ensembl genome browser and provide a comprehensive way to access this data in the context of a widely used genome bioinformatics system. All Ensembl data is freely available at http://www.ensembl.org and from the public MySQL database server at ensembldb.ensembl.org.},} - P Flicek, BL Aken, B Ballester, K Beal, E Bragin, S Brent, Y Chen, P Clapham, G Coates, S Fairley, S Fitzgerald, J Fernandez-Banet, L Gordon, S Gräf, S Haider, M Hammond, K Howe, A Jenkinson, N Johnson, A Kähäri, D Keefe, S Keenan, R Kinsella, F Kokocinski, G Koscielny, E Kulesha, D Lawson, I Longden, T Massingham, W McLaren, K Megy, B Overduin, B Pritchard, D Rios, M Ruffier, M Schuster, G Slater, D Smedley, G Spudich, YA Tang, S Trevanion, A Vilella, J Vogel, S White, SP Wilder, A Zadissa, E Birney, F Cunningham, I Dunham, R Durbin, XM Fernández-Suarez, J Herrero, TJP Hubbard, A Parker, G Proctor, J Smith, SMJ Searle. Ensembl's 10th year. Nucleic Acids Res 2010;38(Database issue):D557–62. doi:10.1093/nar/gkp972
[BibTeX] [Abstract]
Ensembl (http://www.ensembl.org) integrates genomic information for a comprehensive set of chordate genomes with a particular focus on resources for human, mouse, rat, zebrafish and other high-value sequenced genomes. We provide complete gene annotations for all supported species in addition to specific resources that target genome variation, function and evolution. Ensembl data is accessible in a variety of formats including via our genome browser, API and BioMart. This year marks the tenth anniversary of Ensembl and in that time the project has grown with advances in genome technology. As of release 56 (September 2009), Ensembl supports 51 species including marmoset, pig, zebra finch, lizard, gorilla and wallaby, which were added in the past year. Major additions and improvements to Ensembl since our previous report include the incorporation of the human GRCh37 assembly, enhanced visualisation and data-mining options for the Ensembl regulatory features and continued development of our software infrastructure.
@Article{19906699, author = {Flicek P and Aken BL and Ballester B and Beal K and Bragin E and Brent S and Chen Y and Clapham P and Coates G and Fairley S and Fitzgerald S and Fernandez-Banet J and Gordon L and Gräf S and Haider S and Hammond M and Howe K and Jenkinson A and Johnson N and Kähäri A and Keefe D and Keenan S and Kinsella R and Kokocinski F and Koscielny G and Kulesha E and Lawson D and Longden I and Massingham T and McLaren W and Megy K and Overduin B and Pritchard B and Rios D and Ruffier M and Schuster M and Slater G and Smedley D and Spudich G and Tang YA and Trevanion S and Vilella A and Vogel J and White S and Wilder SP and Zadissa A and Birney E and Cunningham F and Dunham I and Durbin R and Fernández-Suarez XM and Herrero J and Hubbard TJP and Parker A and Proctor G and Smith J and Searle SMJ}, title = {Ensembl's 10th year}, journal = {Nucleic Acids Res}, volume = {38}, number = {Database issue}, pages = {D557--62}, year = {2010}, doi = {10.1093/nar/gkp972}, howpublished = {Advanced online publication: 11 November 2009}, abstract = {Ensembl (http://www.ensembl.org) integrates genomic information for a comprehensive set of chordate genomes with a particular focus on resources for human, mouse, rat, zebrafish and other high-value sequenced genomes. We provide complete gene annotations for all supported species in addition to specific resources that target genome variation, function and evolution. Ensembl data is accessible in a variety of formats including via our genome browser, API and BioMart. This year marks the tenth anniversary of Ensembl and in that time the project has grown with advances in genome technology. As of release 56 (September 2009), Ensembl supports 51 species including marmoset, pig, zebra finch, lizard, gorilla and wallaby, which were added in the past year. Major additions and improvements to Ensembl since our previous report include the incorporation of the human GRCh37 assembly, enhanced visualisation and data-mining options for the Ensembl regulatory features and continued development of our software infrastructure.},} - D Smedley, P Schofield, C-K Chen, V Aidinis, C Ainali, J Bard, R Balling, E Birney, A Blake, E Bongcam-Rudloff, AJ Brookes, G Cesareni, C Chandras, J Eppig, P Flicek, G Gkoutos, S Greenaway, M Gruenberger, J-K Hériché, A Lyall, A-M Mallon, D Muddyman, F Reisinger, M Ringwald, N Rosenthal, K Schughart, M Swertz, GA Thorisson, M Zouberakis, JM Hancock. Finding and sharing: new approaches to registries of databases and services for the biomedical sciences. Database (Oxford) 2010;2010:baq014. doi:10.1093/database/baq014
[BibTeX] [Abstract]
The recent explosion of biological data and the concomitant proliferation of distributed databases make it challenging for biologists and bioinformaticians to discover the best data resources for their needs, and the most efficient way to access and use them. Despite a rapid acceleration in uptake of syntactic and semantic standards for interoperability, it is still difficult for users to find which databases support the standards and interfaces that they need. To solve these problems, several groups are developing registries of databases that capture key metadata describing the biological scope, utility, accessibility, ease-of-use and existence of web services allowing interoperability between resources. Here, we describe some of these initiatives including a novel formalism, the Database Description Framework, for describing database operations and functionality and encouraging good database practise. We expect such approaches will result in improved discovery, uptake and utilization of data resources. Database URL: http://www.casimir.org.uk/casimir_ddf.
@Article{20627863, author = {Smedley D and Schofield P and Chen C-K and Aidinis V and Ainali C and Bard J and Balling R and Birney E and Blake A and Bongcam-Rudloff E and Brookes AJ and Cesareni G and Chandras C and Eppig J and Flicek P and Gkoutos G and Greenaway S and Gruenberger M and Hériché J-K and Lyall A and Mallon A-M and Muddyman D and Reisinger F and Ringwald M and Rosenthal N and Schughart K and Swertz M and Thorisson GA and Zouberakis M and Hancock JM}, title = {Finding and sharing: new approaches to registries of databases and services for the biomedical sciences}, journal = {Database (Oxford)}, volume = {2010}, pages = {baq014}, year = {2010}, doi = {10.1093/database/baq014}, abstract = {The recent explosion of biological data and the concomitant proliferation of distributed databases make it challenging for biologists and bioinformaticians to discover the best data resources for their needs, and the most efficient way to access and use them. Despite a rapid acceleration in uptake of syntactic and semantic standards for interoperability, it is still difficult for users to find which databases support the standards and interfaces that they need. To solve these problems, several groups are developing registries of databases that capture key metadata describing the biological scope, utility, accessibility, ease-of-use and existence of web services allowing interoperability between resources. Here, we describe some of these initiatives including a novel formalism, the Database Description Framework, for describing database operations and functionality and encouraging good database practise. We expect such approaches will result in improved discovery, uptake and utilization of data resources. Database URL: http://www.casimir.org.uk/casimir_ddf.},} - D Schmidt, MD Wilson, B Ballester, PC Schwalie, GD Brown, A Marshall, C Kutter, S Watt, CP Martinez-Jimenez, S Mackay, I Talianidis, P Flicek, DT Odom. Five-Vertebrate ChIP-seq Reveals the Evolutionary Dynamics of Transcription Factor Binding. Science 2010;328(5981):1036–1040. doi:10.1126/science.1186176
[BibTeX] [Abstract]
Transcription factors (TFs) direct gene expression by binding to DNA regulatory regions. To explore the evolution of gene regulation, we experimentally determined the genome-wide occupancy of two TFs, CEBPA and HNF4A, in livers of five vertebrates. Although each TF displays highly conserved DNA binding preferences, most binding is species-specific, and aligned binding events present in all five species are rare. Regions near genes with expression levels dependent on a TF are often bound by the TF in multiple species, yet show no enhanced DNA sequence constraint. Binding divergence between species can be largely explained by sequence changes to the bound motifs. Among the binding events lost in one lineage, only half are recovered by another binding event within 10 kilobases. Our results reveal large interspecies differences in transcriptional regulation and provide insight into their evolution.
@Article{20378774, author = {Schmidt D and Wilson MD and Ballester B and Schwalie PC and Brown GD and Marshall A and Kutter C and Watt S and Martinez-Jimenez CP and Mackay S and Talianidis I and Flicek P and Odom DT}, title = {Five-Vertebrate ChIP-seq Reveals the Evolutionary Dynamics of Transcription Factor Binding}, journal = {Science}, volume = {328}, number = {5981}, pages = {1036--1040}, year = {2010}, doi = {10.1126/science.1186176}, abstract = {Transcription factors (TFs) direct gene expression by binding to DNA regulatory regions. To explore the evolution of gene regulation, we experimentally determined the genome-wide occupancy of two TFs, CEBPA and HNF4A, in livers of five vertebrates. Although each TF displays highly conserved DNA binding preferences, most binding is species-specific, and aligned binding events present in all five species are rare. Regions near genes with expression levels dependent on a TF are often bound by the TF in multiple species, yet show no enhanced DNA sequence constraint. Binding divergence between species can be largely explained by sequence changes to the bound motifs. Among the binding events lost in one lineage, only half are recovered by another binding event within 10 kilobases. Our results reveal large interspecies differences in transcriptional regulation and provide insight into their evolution.},} - International Cancer Genome Consortium.. International network of cancer genome projects. Nature 2010;464(7291):993–998. doi:10.1038/nature08987
[BibTeX] [Abstract]
The International Cancer Genome Consortium (ICGC) was launched to coordinate large-scale cancer genome studies in tumours from 50 different cancer types and/or subtypes that are of clinical and societal importance across the globe. Systematic studies of more than 25,000 cancer genomes at the genomic, epigenomic and transcriptomic levels will reveal the repertoire of oncogenic mutations, uncover traces of the mutagenic influences, define clinically relevant subtypes for prognosis and therapeutic management, and enable the development of new cancer therapies.
@Article{20393554, author = {{International Cancer Genome Consortium.}}, title = {International network of cancer genome projects}, journal = {Nature}, volume = {464}, number = {7291}, pages = {993--998}, year = {2010}, doi = {10.1038/nature08987}, abstract = {The International Cancer Genome Consortium (ICGC) was launched to coordinate large-scale cancer genome studies in tumours from 50 different cancer types and/or subtypes that are of clinical and societal importance across the globe. Systematic studies of more than 25,000 cancer genomes at the genomic, epigenomic and transcriptomic levels will reveal the repertoire of oncogenic mutations, uncover traces of the mutagenic influences, define clinically relevant subtypes for prognosis and therapeutic management, and enable the development of new cancer therapies.},} - P Flicek. Journal club. A computational geneticist looks at mechanisms of chromosomal evolution. Nature 2010;463(7282):713. doi:10.1038/463713e
[BibTeX]@Article{20147997, author = {Flicek P}, title = {Journal club. A computational geneticist looks at mechanisms of chromosomal evolution}, journal = {Nature}, volume = {463}, number = {7282}, pages = {713}, year = {2010}, doi = {10.1038/463713e}, } - R Dalgleish, P Flicek, F Cunningham, A Astashyn, RE Tully, G Proctor, Y Chen, WM McLaren, P Larsson, BW Vaughan, C Beroud, G Dobson, H Lehvaslaiho, PE Taschner, den JT Dunnen, A Devereau, E Birney, AJ Brookes, DR Maglott. Locus Reference Genomic sequences: an improved basis for describing human DNA variants. Genome Med 2010;2(4):24. doi:10.1186/gm145
[BibTeX] [Abstract]
ABSTRACT: As our knowledge of the complexity of gene architecture grows, and we increase our understanding of the subtleties of gene expression, the process of accurately describing disease-causing gene variants has become increasingly problematic. In part, this is due to current reference DNA sequence formats that do not fully meet present needs. Here we present the Locus Reference Genomic (LRG) sequence format which has been designed for the specific purpose of gene variant reporting. The format builds on the successful National Center for Biotechnology Information (NCBI) RefSeqGene project and provides a single-file record containing a uniquely stable reference DNA sequence along with all relevant transcript and protein sequences essential to the description of gene variants. In principle, LRGs can be created for any organism, not just human. In addition, we recognise the need to respect legacy numbering systems for exons and amino acids and the LRG format takes account of these. We hope that widespread adoption of LRGs – which will be created and maintained by the NCBI and the European Bioinformatics Institute (EBI) – along with consistent use of the Human Genome Variation Society (HGVS)-approved variant nomenclature will reduce errors in the reporting of variants in the literature and improve communication about variants affecting human health. Further information can be found on the LRG web site (http://www.lrg-sequence.org).
@Article{20398331, author = {Dalgleish R and Flicek P and Cunningham F and Astashyn A and Tully RE and Proctor G and Chen Y and McLaren WM and Larsson P and Vaughan BW and Beroud C and Dobson G and Lehvaslaiho H and Taschner PE and den Dunnen JT and Devereau A and Birney E and Brookes AJ and Maglott DR}, title = {Locus Reference Genomic sequences: an improved basis for describing human DNA variants}, journal = {Genome Med}, volume = {2}, number = {4}, pages = {24}, year = {2010}, doi = {10.1186/gm145}, abstract = {ABSTRACT: As our knowledge of the complexity of gene architecture grows, and we increase our understanding of the subtleties of gene expression, the process of accurately describing disease-causing gene variants has become increasingly problematic. In part, this is due to current reference DNA sequence formats that do not fully meet present needs. Here we present the Locus Reference Genomic (LRG) sequence format which has been designed for the specific purpose of gene variant reporting. The format builds on the successful National Center for Biotechnology Information (NCBI) RefSeqGene project and provides a single-file record containing a uniquely stable reference DNA sequence along with all relevant transcript and protein sequences essential to the description of gene variants. In principle, LRGs can be created for any organism, not just human. In addition, we recognise the need to respect legacy numbering systems for exons and amino acids and the LRG format takes account of these. We hope that widespread adoption of LRGs -- which will be created and maintained by the NCBI and the European Bioinformatics Institute (EBI) -- along with consistent use of the Human Genome Variation Society (HGVS)-approved variant nomenclature will reduce errors in the reporting of variants in the literature and improve communication about variants affecting human health. Further information can be found on the LRG web site (http://www.lrg-sequence.org).},} - K Ye, Z Jia, Y Wang, P Flicek, R Apweiler. Mining Unique-m Substrings from Genomes. J Proteomics Bioinform 2010;3:099–100. doi:10.4172/jpb.1000127
[BibTeX] [Abstract]
Unique substrings in genomes may indicate high level of specificity which is crucial and fundamental to many genetics studies, such as PCR, microarray hybridization, Southern and Northern blotting, RNA interference (RNAi), and genome (re)sequencing. However, being unique sequence in the genome alone is not adequate to guaranty high specificity. For example, nucleotides mismatches within a certain tolerance may impair specificity even if an interested substring occur only once in the genome. In this study we propose the concept of unique-m substrings of genomes for controlling specificity in genome-wide assays. A unique-m substring is defined if it only has a single perfect match on one strand of the entire genome while all other approximate matches must have more than m mismatches. We developed a pattern growth approach to systematically mine such unique-m substrings from a given genome. Our algorithm does not need a pre-processing step to extract sequential information which is required by most of other rival methods. The search for unique-m substrings from genomes is performed as a single task of regular data mining so that the similarities among queries are utilized to achieve tremendous speedup. The runtime of our algorithm is linear to the sizes of input genomes and the length of unique-m substrings. In addition, the unique-m mining algorithm has been parallelized to facilitate genome-wide computation on a cluster or a single machine of multiple CPUs with shared memory.
@Article{29657484, author = {Ye K and Jia Z and Wang Y and Flicek P and Apweiler R}, title = {Mining Unique-m Substrings from Genomes}, journal = {J Proteomics Bioinform}, volume = {3}, pages = {099--100}, year = {2010}, doi = {10.4172/jpb.1000127}, abstract = {Unique substrings in genomes may indicate high level of specificity which is crucial and fundamental to many genetics studies, such as PCR, microarray hybridization, Southern and Northern blotting, RNA interference (RNAi), and genome (re)sequencing. However, being unique sequence in the genome alone is not adequate to guaranty high specificity. For example, nucleotides mismatches within a certain tolerance may impair specificity even if an interested substring occur only once in the genome. In this study we propose the concept of unique-m substrings of genomes for controlling specificity in genome-wide assays. A unique-m substring is defined if it only has a single perfect match on one strand of the entire genome while all other approximate matches must have more than m mismatches. We developed a pattern growth approach to systematically mine such unique-m substrings from a given genome. Our algorithm does not need a pre-processing step to extract sequential information which is required by most of other rival methods. The search for unique-m substrings from genomes is performed as a single task of regular data mining so that the similarities among queries are utilized to achieve tremendous speedup. The runtime of our algorithm is linear to the sizes of input genomes and the length of unique-m substrings. In addition, the unique-m mining algorithm has been parallelized to facilitate genome-wide computation on a cluster or a single machine of multiple CPUs with shared memory.},} - D Peric-Hupkes, W Meuleman, L Pagie, SWM Bruggeman, I Solovei, W Brugman, S Gräf, P Flicek, RM Kerkhoven, van M Lohuizen, M Reinders, L Wessels, van B Steensel. Molecular maps of the reorganization of genome-nuclear lamina interactions during differentiation. Mol Cell 2010;38(4):603–613. doi:10.1016/j.molcel.2010.03.016
[BibTeX] [Abstract]
The three-dimensional organization of chromosomes within the nucleus and its dynamics during differentiation are largely unknown. To visualize this process in molecular detail, we generated high-resolution maps of genome-nuclear lamina interactions during subsequent differentiation of mouse embryonic stem cells via lineage-committed neural precursor cells into terminally differentiated astrocytes. This reveals that a basal chromosome architecture present in embryonic stem cells is cumulatively altered at hundreds of sites during lineage commitment and subsequent terminal differentiation. This remodeling involves both individual transcription units and multigene regions and affects many genes that determine cellular identity. Often, genes that move away from the lamina are concomitantly activated; many others, however, remain inactive yet become unlocked for activation in a next differentiation step. These results suggest that lamina-genome interactions are widely involved in the control of gene expression programs during lineage commitment and terminal differentiation.
@Article{20513434, author = {Peric-Hupkes D and Meuleman W and Pagie L and Bruggeman SWM and Solovei I and Brugman W and Gräf S and Flicek P and Kerkhoven RM and van Lohuizen M and Reinders M and Wessels L and van Steensel B}, title = {Molecular maps of the reorganization of genome-nuclear lamina interactions during differentiation}, journal = {Mol Cell}, volume = {38}, number = {4}, pages = {603--613}, year = {2010}, doi = {10.1016/j.molcel.2010.03.016}, abstract = {The three-dimensional organization of chromosomes within the nucleus and its dynamics during differentiation are largely unknown. To visualize this process in molecular detail, we generated high-resolution maps of genome-nuclear lamina interactions during subsequent differentiation of mouse embryonic stem cells via lineage-committed neural precursor cells into terminally differentiated astrocytes. This reveals that a basal chromosome architecture present in embryonic stem cells is cumulatively altered at hundreds of sites during lineage commitment and subsequent terminal differentiation. This remodeling involves both individual transcription units and multigene regions and affects many genes that determine cellular identity. Often, genes that move away from the lamina are concomitantly activated; many others, however, remain inactive yet become unlocked for activation in a next differentiation step. These results suggest that lamina-genome interactions are widely involved in the control of gene expression programs during lineage commitment and terminal differentiation.},} - RA Dalloul, JA Long, AV Zimin, L Aslam, K Beal, L Ann Blomberg, P Bouffard, DW Burt, O Crasta, RPMA Crooijmans, K Cooper, RA Coulombe, S De, ME Delany, JB Dodgson, JJ Dong, C Evans, KM Frederickson, P Flicek, L Florea, O Folkerts, MAM Groenen, TT Harkins, J Herrero, S Hoffmann, H-J Megens, A Jiang, de P Jong, P Kaiser, H Kim, K-W Kim, S Kim, D Langenberger, M-K Lee, T Lee, S Mane, G Marcais, M Marz, AP McElroy, T Modise, M Nefedov, C Notredame, IR Paton, WS Payne, G Pertea, D Prickett, D Puiu, D Qioa, E Raineri, M Ruffier, SL Salzberg, MC Schatz, C Scheuring, CJ Schmidt, S Schroeder, SMJ Searle, EJ Smith, J Smith, TS Sonstegard, PF Stadler, H Tafer, ZJ Tu, CP Van Tassell, AJ Vilella, KP Williams, JA Yorke, L Zhang, H-B Zhang, X Zhang, Y Zhang, KM Reed. Multi-Platform Next-Generation Sequencing of the Domestic Turkey (Meleagris gallopavo): Genome Assembly and Analysis. PLoS Biol 2010;8(9):e1000475. doi:10.1371/journal.pbio.1000475
[BibTeX] [Abstract]
A synergistic combination of two next-generation sequencing platforms with a detailed comparative BAC physical contig map provided a cost-effective assembly of the genome sequence of the domestic turkey (Meleagris gallopavo). Heterozygosity of the sequenced source genome allowed discovery of more than 600,000 high quality single nucleotide variants. Despite this heterozygosity, the current genome assembly (∼1.1 Gb) includes 917 Mb of sequence assigned to specific turkey chromosomes. Annotation identified nearly 16,000 genes, with 15,093 recognized as protein coding and 611 as non-coding RNA genes. Comparative analysis of the turkey, chicken, and zebra finch genomes, and comparing avian to mammalian species, supports the characteristic stability of avian genomes and identifies genes unique to the avian lineage. Clear differences are seen in number and variety of genes of the avian immune system where expansions and novel genes are less frequent than examples of gene loss. The turkey genome sequence provides resources to further understand the evolution of vertebrate genomes and genetic variation underlying economically important quantitative traits in poultry. This integrated approach may be a model for providing both gene and chromosome level assemblies of other species with agricultural, ecological, and evolutionary interest.
@Article{20838655, author = {Dalloul RA and Long JA and Zimin AV and Aslam L and Beal K and Ann Blomberg L and Bouffard P and Burt DW and Crasta O and Crooijmans RPMA and Cooper K and Coulombe RA and De S and Delany ME and Dodgson JB and Dong JJ and Evans C and Frederickson KM and Flicek P and Florea L and Folkerts O and Groenen MAM and Harkins TT and Herrero J and Hoffmann S and Megens H-J and Jiang A and de Jong P and Kaiser P and Kim H and Kim K-W and Kim S and Langenberger D and Lee M-K and Lee T and Mane S and Marcais G and Marz M and McElroy AP and Modise T and Nefedov M and Notredame C and Paton IR and Payne WS and Pertea G and Prickett D and Puiu D and Qioa D and Raineri E and Ruffier M and Salzberg SL and Schatz MC and Scheuring C and Schmidt CJ and Schroeder S and Searle SMJ and Smith EJ and Smith J and Sonstegard TS and Stadler PF and Tafer H and Tu ZJ and Van Tassell CP and Vilella AJ and Williams KP and Yorke JA and Zhang L and Zhang H-B and Zhang X and Zhang Y and Reed KM}, title = {Multi-Platform Next-Generation Sequencing of the Domestic Turkey (Meleagris gallopavo): Genome Assembly and Analysis}, journal = {PLoS Biol}, volume = {8}, number = {9}, pages = {e1000475}, year = {2010}, doi = {10.1371/journal.pbio.1000475}, abstract = {A synergistic combination of two next-generation sequencing platforms with a detailed comparative BAC physical contig map provided a cost-effective assembly of the genome sequence of the domestic turkey (Meleagris gallopavo). Heterozygosity of the sequenced source genome allowed discovery of more than 600,000 high quality single nucleotide variants. Despite this heterozygosity, the current genome assembly (∼1.1 Gb) includes 917 Mb of sequence assigned to specific turkey chromosomes. Annotation identified nearly 16,000 genes, with 15,093 recognized as protein coding and 611 as non-coding RNA genes. Comparative analysis of the turkey, chicken, and zebra finch genomes, and comparing avian to mammalian species, supports the characteristic stability of avian genomes and identifies genes unique to the avian lineage. Clear differences are seen in number and variety of genes of the avian immune system where expansions and novel genes are less frequent than examples of gene loss. The turkey genome sequence provides resources to further understand the evolution of vertebrate genomes and genetic variation underlying economically important quantitative traits in poultry. This integrated approach may be a model for providing both gene and chromosome level assemblies of other species with agricultural, ecological, and evolutionary interest.},} - DM Church, I Lappalainen, TP Sneddon, J Hinton, M Maguire, J Lopez, J Garner, J Paschall, M Dicuccio, E Yaschenko, SW Scherer, L Feuk, P Flicek. Public data archives for genomic structural variation. Nat Genet 2010;42(10):813–814. doi:10.1038/ng1010-813
[BibTeX]@Article{20877315, author = {Church DM and Lappalainen I and Sneddon TP and Hinton J and Maguire M and Lopez J and Garner J and Paschall J and Dicuccio M and Yaschenko E and Scherer SW and Feuk L and Flicek P}, title = {Public data archives for genomic structural variation}, journal = {Nat Genet}, volume = {42}, number = {10}, pages = {813--814}, year = {2010}, doi = {10.1038/ng1010-813}, } - WC Warren, DF Clayton, H Ellegren, AP Arnold, LW Hillier, A Künstner, S Searle, S White, AJ Vilella, S Fairley, A Heger, L Kong, CP Ponting, ED Jarvis, CV Mello, P Minx, P Lovell, TAF Velho, M Ferris, CN Balakrishnan, S Sinha, C Blatti, SE London, Y Li, Y-C Lin, J George, J Sweedler, B Southey, P Gunaratne, M Watson, K Nam, N Backström, L Smeds, B Nabholz, Y Itoh, O Whitney, AR Pfenning, J Howard, M Völker, BM Skinner, DK Griffin, L Ye, WM McLaren, P Flicek, V Quesada, G Velasco, C Lopez-Otin, XS Puente, T Olender, D Lancet, AFA Smit, R Hubley, MK Konkel, JA Walker, MA Batzer, W Gu, DD Pollock, L Chen, Z Cheng, EE Eichler, J Stapley, J Slate, R Ekblom, T Birkhead, T Burke, D Burt, C Scharff, I Adam, H Richard, M Sultan, A Soldatov, H Lehrach, SV Edwards, S-P Yang, X Li, T Graves, L Fulton, J Nelson, A Chinwalla, S Hou, ER Mardis, RK Wilson. The genome of a songbird. Nature 2010;464(7289):757–762. doi:10.1038/nature08819
[BibTeX] [Abstract]
The zebra finch is an important model organism in several fields with unique relevance to human neuroscience. Like other songbirds, the zebra finch communicates through learned vocalizations, an ability otherwise documented only in humans and a few other animals and lacking in the chicken-the only bird with a sequenced genome until now. Here we present a structural, functional and comparative analysis of the genome sequence of the zebra finch (Taeniopygia guttata), which is a songbird belonging to the large avian order Passeriformes. We find that the overall structures of the genomes are similar in zebra finch and chicken, but they differ in many intrachromosomal rearrangements, lineage-specific gene family expansions, the number of long-terminal-repeat-based retrotransposons, and mechanisms of sex chromosome dosage compensation. We show that song behaviour engages gene regulatory networks in the zebra finch brain, altering the expression of long non-coding RNAs, microRNAs, transcription factors and their targets. We also show evidence for rapid molecular evolution in the songbird lineage of genes that are regulated during song experience. These results indicate an active involvement of the genome in neural processes underlying vocal communication and identify potential genetic substrates for the evolution and regulation of this behaviour.
@Article{20360741, author = {Warren WC and Clayton DF and Ellegren H and Arnold AP and Hillier LW and Künstner A and Searle S and White S and Vilella AJ and Fairley S and Heger A and Kong L and Ponting CP and Jarvis ED and Mello CV and Minx P and Lovell P and Velho TAF and Ferris M and Balakrishnan CN and Sinha S and Blatti C and London SE and Li Y and Lin Y-C and George J and Sweedler J and Southey B and Gunaratne P and Watson M and Nam K and Backström N and Smeds L and Nabholz B and Itoh Y and Whitney O and Pfenning AR and Howard J and Völker M and Skinner BM and Griffin DK and Ye L and McLaren WM and Flicek P and Quesada V and Velasco G and Lopez-Otin C and Puente XS and Olender T and Lancet D and Smit AFA and Hubley R and Konkel MK and Walker JA and Batzer MA and Gu W and Pollock DD and Chen L and Cheng Z and Eichler EE and Stapley J and Slate J and Ekblom R and Birkhead T and Burke T and Burt D and Scharff C and Adam I and Richard H and Sultan M and Soldatov A and Lehrach H and Edwards SV and Yang S-P and Li X and Graves T and Fulton L and Nelson J and Chinwalla A and Hou S and Mardis ER and Wilson RK}, title = {The genome of a songbird}, journal = {Nature}, volume = {464}, number = {7289}, pages = {757--762}, year = {2010}, doi = {10.1038/nature08819}, abstract = {The zebra finch is an important model organism in several fields with unique relevance to human neuroscience. Like other songbirds, the zebra finch communicates through learned vocalizations, an ability otherwise documented only in humans and a few other animals and lacking in the chicken-the only bird with a sequenced genome until now. Here we present a structural, functional and comparative analysis of the genome sequence of the zebra finch (Taeniopygia guttata), which is a songbird belonging to the large avian order Passeriformes. We find that the overall structures of the genomes are similar in zebra finch and chicken, but they differ in many intrachromosomal rearrangements, lineage-specific gene family expansions, the number of long-terminal-repeat-based retrotransposons, and mechanisms of sex chromosome dosage compensation. We show that song behaviour engages gene regulatory networks in the zebra finch brain, altering the expression of long non-coding RNAs, microRNAs, transcription factors and their targets. We also show evidence for rapid molecular evolution in the songbird lineage of genes that are regulated during song experience. These results indicate an active involvement of the genome in neural processes underlying vocal communication and identify potential genetic substrates for the evolution and regulation of this behaviour.},} - SS Atanur, I Birol, V Guryev, M Hirst, O Hummel, C Morrissey, J Behmoaras, XM Fernandez-Suarez, MD Johnson, WM McLaren, G Patone, E Petretto, C Plessy, KS Rockland, C Rockland, K Saar, Y Zhao, P Carninci, P Flicek, T Kurtz, E Cuppen, M Pravenec, N Hubner, SJM Jones, E Birney, TJ Aitman. The genome sequence of the spontaneously hypertensive rat: Analysis and functional significance. Genome Res 2010;20(6):791–803. doi:10.1101/gr.103499.109
[BibTeX] [Abstract]
The spontaneously hypertensive rat (SHR) is the most widely studied animal model of hypertension. Scores of SHR quantitative loci (QTLs) have been mapped for hypertension and other phenotypes. We have sequenced the SHR/OlaIpcv genome at 10.7-fold coverage by paired-end sequencing on the Illumina platform. We identified 3.6 million high-quality single nucleotide polymorphisms (SNPs) between the SHR/OlaIpcv and Brown Norway (BN) reference genome, with a high rate of validation (sensitivity 96.3\%-98.0\% and specificity 99\%-100\%). We also identified 343,243 short indels between the SHR/OlaIpcv and reference genomes. These SNPs and indels resulted in 161 gain or loss of stop codons and 629 frameshifts compared with the BN reference sequence. We also identified 13,438 larger deletions that result in complete or partial absence of 107 genes in the SHR/OlaIpcv genome compared with the BN reference and 588 copy number variants (CNVs) that overlap with the gene regions of 688 genes. Genomic regions containing genes whose expression had been previously mapped as cis-regulated expression quantitative trait loci (eQTLs) were significantly enriched with SNPs, short indels, and larger deletions, suggesting that some of these variants have functional effects on gene expression. Genes that were affected by major alterations in their coding sequence were highly enriched for genes related to ion transport, transport, and plasma membrane localization, providing insights into the likely molecular and cellular basis of hypertension and other phenotypes specific to the SHR strain. This near complete catalog of genomic differences between two extensively studied rat strains provides the starting point for complete elucidation, at the molecular level, of the physiological and pathophysiological phenotypic differences between individuals from these strains.
@Article{20430781, author = {Atanur SS and Birol I and Guryev V and Hirst M and Hummel O and Morrissey C and Behmoaras J and Fernandez-Suarez XM and Johnson MD and McLaren WM and Patone G and Petretto E and Plessy C and Rockland KS and Rockland C and Saar K and Zhao Y and Carninci P and Flicek P and Kurtz T and Cuppen E and Pravenec M and Hubner N and Jones SJM and Birney E and Aitman TJ}, title = {The genome sequence of the spontaneously hypertensive rat: Analysis and functional significance}, journal = {Genome Res}, volume = {20}, number = {6}, pages = {791--803}, year = {2010}, doi = {10.1101/gr.103499.109}, howpublished = {Advanced online publication: 29 April 2010}, abstract = {The spontaneously hypertensive rat (SHR) is the most widely studied animal model of hypertension. Scores of SHR quantitative loci (QTLs) have been mapped for hypertension and other phenotypes. We have sequenced the SHR/OlaIpcv genome at 10.7-fold coverage by paired-end sequencing on the Illumina platform. We identified 3.6 million high-quality single nucleotide polymorphisms (SNPs) between the SHR/OlaIpcv and Brown Norway (BN) reference genome, with a high rate of validation (sensitivity 96.3\%-98.0\% and specificity 99\%-100\%). We also identified 343,243 short indels between the SHR/OlaIpcv and reference genomes. These SNPs and indels resulted in 161 gain or loss of stop codons and 629 frameshifts compared with the BN reference sequence. We also identified 13,438 larger deletions that result in complete or partial absence of 107 genes in the SHR/OlaIpcv genome compared with the BN reference and 588 copy number variants (CNVs) that overlap with the gene regions of 688 genes. Genomic regions containing genes whose expression had been previously mapped as cis-regulated expression quantitative trait loci (eQTLs) were significantly enriched with SNPs, short indels, and larger deletions, suggesting that some of these variants have functional effects on gene expression. Genes that were affected by major alterations in their coding sequence were highly enriched for genes related to ion transport, transport, and plasma membrane localization, providing insights into the likely molecular and cellular basis of hypertension and other phenotypes specific to the SHR strain. This near complete catalog of genomic differences between two extensively studied rat strains provides the starting point for complete elucidation, at the molecular level, of the physiological and pathophysiological phenotypic differences between individuals from these strains.},}
2009
- RH Morley, K Lachani, D Keefe, MJ Gilchrist, P Flicek, JC Smith, FC Wardle. A gene regulatory network directed by zebrafish No tail accounts for its roles in mesoderm formation. Proc Natl Acad Sci U S A 2009;106(10):3829–3834. doi:10.1073/pnas.0808382106
[BibTeX] [Abstract]
Using chromatin immunoprecipitation combined with genomic microarrays we have identified targets of No tail (Ntl), a zebrafish Brachyury ortholog that plays a central role in mesoderm formation. We show that Ntl regulates a downstream network of other transcription factors and identify an in vivo Ntl binding site that resembles the consensus T-box binding site (TBS) previously identified by in vitro studies. We show that the notochord-expressed gene floating head (flh) is a direct transcriptional target of Ntl and that a combination of TBSs in the flh upstream region are required for Ntl-directed expression. Using our genome-scale data we have assembled a preliminary gene regulatory network that begins to describe mesoderm formation and patterning in the early zebrafish embryo.
@Article{19225104, author = {Morley RH and Lachani K and Keefe D and Gilchrist MJ and Flicek P and Smith JC and Wardle FC}, title = {A gene regulatory network directed by zebrafish No tail accounts for its roles in mesoderm formation}, journal = {Proc Natl Acad Sci U S A}, volume = {106}, number = {10}, pages = {3829--3834}, year = {2009}, doi = {10.1073/pnas.0808382106}, abstract = {Using chromatin immunoprecipitation combined with genomic microarrays we have identified targets of No tail (Ntl), a zebrafish Brachyury ortholog that plays a central role in mesoderm formation. We show that Ntl regulates a downstream network of other transcription factors and identify an in vivo Ntl binding site that resembles the consensus T-box binding site (TBS) previously identified by in vitro studies. We show that the notochord-expressed gene floating head (flh) is a direct transcriptional target of Ntl and that a combination of TBSs in the flh upstream region are required for Ntl-directed expression. Using our genome-scale data we have assembled a preliminary gene regulatory network that begins to describe mesoderm formation and patterning in the early zebrafish embryo.},} - TJP Hubbard, BL Aken, S Ayling, B Ballester, K Beal, E Bragin, S Brent, Y Chen, P Clapham, L Clarke, G Coates, S Fairley, S Fitzgerald, J Fernandez-Banet, L Gordon, S Graf, S Haider, M Hammond, R Holland, K Howe, A Jenkinson, N Johnson, A Kahari, D Keefe, S Keenan, R Kinsella, F Kokocinski, E Kulesha, D Lawson, I Longden, K Megy, P Meidl, B Overduin, A Parker, B Pritchard, D Rios, M Schuster, G Slater, D Smedley, W Spooner, G Spudich, S Trevanion, A Vilella, J Vogel, S White, S Wilder, A Zadissa, E Birney, F Cunningham, V Curwen, R Durbin, XM Fernandez-Suarez, J Herrero, A Kasprzyk, G Proctor, J Smith, S Searle, P Flicek. Ensembl 2009. Nucleic Acids Res 2009;37(Database issue):D690–7. doi:10.1093/nar/gkn828
[BibTeX] [Abstract]
The Ensembl project (http://www.ensembl.org) is a comprehensive genome information system featuring an integrated set of genome annotation, databases, and other information for chordate, selected model organism and disease vector genomes. As of release 51 (November 2008), Ensembl fully supports 45 species, and three additional species have preliminary support. New species in the past year include orangutan and six additional low coverage mammalian genomes. Major additions and improvements to Ensembl since our previous report include a major redesign of our website; generation of multiple genome alignments and ancestral sequences using the new Enredo-Pecan-Ortheus pipeline and development of our software infrastructure, particularly to support the Ensembl Genomes project (http://www.ensemblgenomes.org/).
@Article{19033362, author = {Hubbard TJP and Aken BL and Ayling S and Ballester B and Beal K and Bragin E and Brent S and Chen Y and Clapham P and Clarke L and Coates G and Fairley S and Fitzgerald S and Fernandez-Banet J and Gordon L and Graf S and Haider S and Hammond M and Holland R and Howe K and Jenkinson A and Johnson N and Kahari A and Keefe D and Keenan S and Kinsella R and Kokocinski F and Kulesha E and Lawson D and Longden I and Megy K and Meidl P and Overduin B and Parker A and Pritchard B and Rios D and Schuster M and Slater G and Smedley D and Spooner W and Spudich G and Trevanion S and Vilella A and Vogel J and White S and Wilder S and Zadissa A and Birney E and Cunningham F and Curwen V and Durbin R and Fernandez-Suarez XM and Herrero J and Kasprzyk A and Proctor G and Smith J and Searle S and Flicek P}, title = {Ensembl 2009}, journal = {Nucleic Acids Res}, volume = {37}, number = {Database issue}, pages = {D690--7}, year = {2009}, doi = {10.1093/nar/gkn828}, howpublished = {Advanced online publication: 25 November 2008}, abstract = {The Ensembl project (http://www.ensembl.org) is a comprehensive genome information system featuring an integrated set of genome annotation, databases, and other information for chordate, selected model organism and disease vector genomes. As of release 51 (November 2008), Ensembl fully supports 45 species, and three additional species have preliminary support. New species in the past year include orangutan and six additional low coverage mammalian genomes. Major additions and improvements to Ensembl since our previous report include a major redesign of our website; generation of multiple genome alignments and ancestral sequences using the new Enredo-Pecan-Ortheus pipeline and development of our software infrastructure, particularly to support the Ensembl Genomes project (http://www.ensemblgenomes.org/).},} - AW Bruce, AJ López-Contreras, P Flicek, TA Down, P Dhami, SC Dillon, CM Koch, CF Langford, I Dunham, RM Andrews, D Vetrie. Functional diversity for REST (NRSF) is defined by in vivo binding affinity hierarchies at the DNA sequence level. Genome Res 2009;19(6):994–1005. doi:10.1101/gr.089086.108
[BibTeX] [Abstract]
The molecular events that contribute to, and result from, the in vivo binding of transcription factors to their cognate DNA sequence motifs in mammalian genomes are poorly understood. We demonstrate that variations within the DNA sequence motifs that bind the transcriptional repressor REST (NRSF) encode in vivo DNA binding affinity hierarchies that contribute to regulatory function during lineage-specific and developmental programs in fundamental ways. First, canonical sequence motifs for REST facilitate strong REST binding and control functional classes of REST targets that are common to all cell types, whilst atypical motifs participate in weak interactions and control those targets, which are cell- or tissue-specific. Second, variations in REST binding relate directly to variations in expression and chromatin configurations of REST's target genes. Third, REST clearance from its binding sites is also associated with variations in the RE1 motif. Finally, and most surprisingly, weak REST binding sites reside in DNA sequences that show the highest levels of constraint through evolution, thus facilitating their roles in maintaining tissue-specific functions. These relationships have never been reported in mammalian systems for any transcription factor.
@Article{19401398, author = {Bruce AW and López-Contreras AJ and Flicek P and Down TA and Dhami P and Dillon SC and Koch CM and Langford CF and Dunham I and Andrews RM and Vetrie D}, title = {Functional diversity for REST (NRSF) is defined by in vivo binding affinity hierarchies at the DNA sequence level}, journal = {Genome Res}, volume = {19}, number = {6}, pages = {994--1005}, year = {2009}, doi = {10.1101/gr.089086.108}, howpublished = {Advanced online publication: 28 April 2009}, abstract = {The molecular events that contribute to, and result from, the in vivo binding of transcription factors to their cognate DNA sequence motifs in mammalian genomes are poorly understood. We demonstrate that variations within the DNA sequence motifs that bind the transcriptional repressor REST (NRSF) encode in vivo DNA binding affinity hierarchies that contribute to regulatory function during lineage-specific and developmental programs in fundamental ways. First, canonical sequence motifs for REST facilitate strong REST binding and control functional classes of REST targets that are common to all cell types, whilst atypical motifs participate in weak interactions and control those targets, which are cell- or tissue-specific. Second, variations in REST binding relate directly to variations in expression and chromatin configurations of REST's target genes. Third, REST clearance from its binding sites is also associated with variations in the RE1 motif. Finally, and most surprisingly, weak REST binding sites reside in DNA sequences that show the highest levels of constraint through evolution, thus facilitating their roles in maintaining tissue-specific functions. These relationships have never been reported in mammalian systems for any transcription factor.},} - J Kaput, RGH Cotton, L Hardman, M Watson, AI Al Aqeel, JY Al-Aama, F Al-Mulla, S Alonso, S Aretz, AD Auerbach, B Bapat, IT Bernstein, J Bhak, SL Bleoo, H Blöcker, SE Brenner, J Burn, M Bustamante, R Calzone, A Cambon-Thomsen, M Cargill, P Carrera, L Cavedon, YS Cho, Y-J Chung, M Claustres, G Cutting, R Dalgleish, den JT Dunnen, C Díaz, S Dobrowolski, MRN Dos Santos, R Ekong, SB Flanagan, P Flicek, Y Furukawa, M Genuardi, H Ghang, MV Golubenko, MS Greenblatt, A Hamosh, JM Hancock, R Hardison, TM Harrison, R Hoffmann, R Horaitis, HJ Howard, CI Barash, N Izagirre, J Jung, T Kojima, S Laradi, Y-S Lee, J-Y Lee, VL Gil-da-Silva-Lopes, FA Macrae, D Maglott, MJ Marafie, SGE Marsh, Y Matsubara, LM Messiaen, G Möslein, MG Netea, ML Norton, PJ Oefner, WS Oetting, JC O'Leary, de AMO Ramirez, MH Paalman, J Parboosingh, GP Patrinos, G Perozzi, IR Phillips, S Povey, S Prasad, M Qi, DJ Quin, RS Ramesar, CS Richards, J Savige, DG Scheible, RJ Scott, D Seminara, EA Shephard, RH Sijmons, TD Smith, M-J Sobrido, T Tanaka, SV Tavtigian, GR Taylor, J Teague, T Töpel, M Ullman-Cullere, J Utsunomiya, van HJ Kranen, M Vihinen, E Webb, TK Weber, M Yeager, YI Yeom, S-H Yim, H-S Yoo, OBOCTTHVPP Meeting. Planning the Human Variome Project: The Spain Report. Human Mutation 2009;30(4):496–510. doi:10.1002/humu.20972
[BibTeX] [Abstract]
The remarkable progress in characterizing the human genome sequence, exemplified by the Human Genome Project and the HapMap Consortium, has led to the perception that knowledge and the tools (e.g., microarrays) are sufficient for many if not most biomedical research efforts. A large amount of data from diverse studies proves this perception inaccurate at best, and at worst, an impediment for further efforts to characterize the variation in the human genome. Because variation in genotype and environment are the fundamental basis to understand phenotypic variability and heritability at the population level, identifying the range of human genetic variation is crucial to the development of personalized nutrition and medicine. The Human Variome Project (HVP; http://www.humanvariomeproject.org/) was proposed initially to systematically collect mutations that cause human disease and create a cyber infrastructure to link locus specific databases (LSDB). We report here the discussions and recommendations from the 2008 HVP planning meeting held in San Feliu de Guixols, Spain, in May 2008. Hum Mutat 30, 496-510, 2009. (c) 2009 Wiley-Liss, Inc.
@Article{19306394, author = {Kaput J and Cotton RGH and Hardman L and Watson M and Al Aqeel AI and Al-Aama JY and Al-Mulla F and Alonso S and Aretz S and Auerbach AD and Bapat B and Bernstein IT and Bhak J and Bleoo SL and Blöcker H and Brenner SE and Burn J and Bustamante M and Calzone R and Cambon-Thomsen A and Cargill M and Carrera P and Cavedon L and Cho YS and Chung Y-J and Claustres M and Cutting G and Dalgleish R and den Dunnen JT and Díaz C and Dobrowolski S and Dos Santos MRN and Ekong R and Flanagan SB and Flicek P and Furukawa Y and Genuardi M and Ghang H and Golubenko MV and Greenblatt MS and Hamosh A and Hancock JM and Hardison R and Harrison TM and Hoffmann R and Horaitis R and Howard HJ and Barash CI and Izagirre N and Jung J and Kojima T and Laradi S and Lee Y-S and Lee J-Y and Gil-da-Silva-Lopes VL and Macrae FA and Maglott D and Marafie MJ and Marsh SGE and Matsubara Y and Messiaen LM and Möslein G and Netea MG and Norton ML and Oefner PJ and Oetting WS and O'Leary JC and de Ramirez AMO and Paalman MH and Parboosingh J and Patrinos GP and Perozzi G and Phillips IR and Povey S and Prasad S and Qi M and Quin DJ and Ramesar RS and Richards CS and Savige J and Scheible DG and Scott RJ and Seminara D and Shephard EA and Sijmons RH and Smith TD and Sobrido M-J and Tanaka T and Tavtigian SV and Taylor GR and Teague J and Töpel T and Ullman-Cullere M and Utsunomiya J and van Kranen HJ and Vihinen M and Webb E and Weber TK and Yeager M and Yeom YI and Yim S-H and Yoo H-S and Meeting OBOCTTHVPP}, title = {Planning the Human Variome Project: The Spain Report}, journal = {Human Mutation}, volume = {30}, number = {4}, pages = {496--510}, year = {2009}, doi = {10.1002/humu.20972}, abstract = {The remarkable progress in characterizing the human genome sequence, exemplified by the Human Genome Project and the HapMap Consortium, has led to the perception that knowledge and the tools (e.g., microarrays) are sufficient for many if not most biomedical research efforts. A large amount of data from diverse studies proves this perception inaccurate at best, and at worst, an impediment for further efforts to characterize the variation in the human genome. Because variation in genotype and environment are the fundamental basis to understand phenotypic variability and heritability at the population level, identifying the range of human genetic variation is crucial to the development of personalized nutrition and medicine. The Human Variome Project (HVP; http://www.humanvariomeproject.org/) was proposed initially to systematically collect mutations that cause human disease and create a cyber infrastructure to link locus specific databases (LSDB). We report here the discussions and recommendations from the 2008 HVP planning meeting held in San Feliu de Guixols, Spain, in May 2008. Hum Mutat 30, 496-510, 2009. (c) 2009 Wiley-Liss, Inc.},} - P Flicek, E Birney. Sense from sequence reads: methods for alignment and assembly. Nat Methods 2009;6(11 Suppl):S6–S12. doi:10.1038/nmeth.1376
[BibTeX] [Abstract]
The most important first step in understanding next-generation sequencing data is the initial alignment or assembly that determines whether an experiment has succeeded and provides a first glimpse into the results. In parallel with the growth of new sequencing technologies, several algorithms that align or assemble the large data output of today's sequencing machines have been developed. We discuss the current algorithmic approaches and future directions of these fundamental tools and provide specific examples for some commonly used tools.
@Article{19844229, author = {Flicek P and Birney E}, title = {Sense from sequence reads: methods for alignment and assembly}, journal = {Nat Methods}, volume = {6}, number = {11 Suppl}, pages = {S6--S12}, year = {2009}, doi = {10.1038/nmeth.1376}, abstract = {The most important first step in understanding next-generation sequencing data is the initial alignment or assembly that determines whether an experiment has succeeded and provides a first glimpse into the results. In parallel with the growth of new sequencing technologies, several algorithms that align or assemble the large data output of today's sequencing machines have been developed. We discuss the current algorithmic approaches and future directions of these fundamental tools and provide specific examples for some commonly used tools.},} - M Carlile, D Swan, K Jackson, K Preston-Fayers, B Ballester, P Flicek, A Werner. Strand selective generation of endo-siRNAs from the Na/phosphate transporter gene Slc34a1 in murine tissues. Nucleic Acids Res 2009;37(7):2274–2282. doi:10.1093/nar/gkp088
[BibTeX] [Abstract]
Natural antisense transcripts (NATs) are important regulators of gene expression. Recently, a link between antisense transcription and the formation of endo-siRNAs has emerged. We investigated the bi-directionally transcribed Na/phosphate cotransporter gene (Slc34a1) under the aspect of endo-siRNA processing. Mouse Slc34a1 produces an antisense transcript that represents an alternative splice product of the Pfn3 gene located downstream of Slc34a1. The antisense transcript is prominently found in testis and in kidney. Co-expression of in vitro synthesized sense/antisense transcripts in Xenopus oocytes indicated processing of the overlapping transcripts into endo-siRNAs in the nucleus. Truncation experiments revealed that an overlap of at least 29 base-pairs is required to induce processing. We detected endo-siRNAs in mouse tissues that co express Slc34a1 sense/antisense transcripts by northern blotting. The orientation of endo-siRNAs was tissue specific in mouse kidney and testis. In kidney where the Na/phosphate cotransporter fulfils its physiological function endo-siRNAs complementary to the NAT were detected, in testis both orientations were found. Considering the wide spread expression of NATs and the gene silencing potential of endo-siRNAs we hypothesized a genome-wide link between antisense transcription and monoallelic expression. Significant correlation between random imprinting and antisense transcription could indeed be established. Our findings suggest a novel, more general role for NATs in gene regulation.
@Article{19237395, author = {Carlile M and Swan D and Jackson K and Preston-Fayers K and Ballester B and Flicek P and Werner A}, title = {Strand selective generation of endo-siRNAs from the Na/phosphate transporter gene Slc34a1 in murine tissues}, journal = {Nucleic Acids Res}, volume = {37}, number = {7}, pages = {2274--2282}, year = {2009}, doi = {10.1093/nar/gkp088}, howpublished = {Advanced online publication: 23 February 2009}, abstract = {Natural antisense transcripts (NATs) are important regulators of gene expression. Recently, a link between antisense transcription and the formation of endo-siRNAs has emerged. We investigated the bi-directionally transcribed Na/phosphate cotransporter gene (Slc34a1) under the aspect of endo-siRNA processing. Mouse Slc34a1 produces an antisense transcript that represents an alternative splice product of the Pfn3 gene located downstream of Slc34a1. The antisense transcript is prominently found in testis and in kidney. Co-expression of in vitro synthesized sense/antisense transcripts in Xenopus oocytes indicated processing of the overlapping transcripts into endo-siRNAs in the nucleus. Truncation experiments revealed that an overlap of at least 29 base-pairs is required to induce processing. We detected endo-siRNAs in mouse tissues that co express Slc34a1 sense/antisense transcripts by northern blotting. The orientation of endo-siRNAs was tissue specific in mouse kidney and testis. In kidney where the Na/phosphate cotransporter fulfils its physiological function endo-siRNAs complementary to the NAT were detected, in testis both orientations were found. Considering the wide spread expression of NATs and the gene silencing potential of endo-siRNAs we hypothesized a genome-wide link between antisense transcription and monoallelic expression. Significant correlation between random imprinting and antisense transcription could indeed be established. Our findings suggest a novel, more general role for NATs in gene regulation.},} - P Flicek. The need for speed. Genome Biol 2009;10(3):212. doi:10.1186/gb-2009-10-3-212
[BibTeX] [Abstract]
ABSTRACT: DNA sequence data are being produced at an ever-increasing rate. The Bowtie sequence-alignment algorithm uses advanced data structures to help data analysis keep pace with data generation.
@Article{19344490, author = {Flicek P}, title = {The need for speed}, journal = {Genome Biol}, volume = {10}, number = {3}, pages = {212}, year = {2009}, doi = {10.1186/gb-2009-10-3-212}, abstract = {ABSTRACT: DNA sequence data are being produced at an ever-increasing rate. The Bowtie sequence-alignment algorithm uses advanced data structures to help data analysis keep pace with data generation.},} - Paul Flicek, Ewan Birney. Visualising the Epigenome. In: C Ferguson-Smith Anne, M Greally John, A Martienssen Robert, `editors`. In: Epigenomics. Springer Netherlands, 2009, 55-66.
[BibTeX] [Abstract]
The epigenome describes the collection of chromatin modifications that are stable through cell division and are apparently indicative of cell type. Epigenetic features include DNA methylation, histone modifications, and other aspects of chromatin organisation. Genome browsers were developed as part of the human genome project to visualise and organise genomic data and analysis using the common index of the genome sequence. These browsers and other tools are now being expanded to support the growing collection of epigenomic data. Ideally viewing data in genome browsers allows researchers to better understand their own data, draw new scientific conclusions, and present their data to others in a visually appealing and intuitive way.
@Incollection{Flicek2009b, author = {Flicek Paul and Birney Ewan}, title = {Visualising the Epigenome}, booktitle = {Epigenomics}, editor = {Ferguson-Smith Anne C and Greally John M and Martienssen Robert A}, publisher = {Springer Netherlands}, pages = {55-66}, year = {2009}, abstract = {The epigenome describes the collection of chromatin modifications that are stable through cell division and are apparently indicative of cell type. Epigenetic features include DNA methylation, histone modifications, and other aspects of chromatin organisation. Genome browsers were developed as part of the human genome project to visualise and organise genomic data and analysis using the common index of the genome sequence. These browsers and other tools are now being expanded to support the growing collection of epigenomic data. Ideally viewing data in genome browsers allows researchers to better understand their own data, draw new scientific conclusions, and present their data to others in a visually appealing and intuitive way.}, }
2008
- TA Down, VK Rakyan, DJ Turner, P Flicek, H Li, E Kulesha, S Gräf, N Johnson, J Herrero, EM Tomazou, NP Thorne, L Bäckdahl, M Herberth, KL Howe, DK Jackson, MM Miretti, JC Marioni, E Birney, TJP Hubbard, R Durbin, S Tavaré, S Beck. A Bayesian deconvolution strategy for immunoprecipitation-based DNA methylome analysis. Nat Biotechnol 2008;26(7):779–785. doi:10.1038/nbt1414
[BibTeX] [Abstract]
DNA methylation is an indispensible epigenetic modification required for regulating the expression of mammalian genomes. Immunoprecipitation-based methods for DNA methylome analysis are rapidly shifting the bottleneck in this field from data generation to data analysis, necessitating the development of better analytical tools. In particular, an inability to estimate absolute methylation levels remains a major analytical difficulty associated with immunoprecipitation-based DNA methylation profiling. To address this issue, we developed a cross-platform algorithm-Bayesian tool for methylation analysis (Batman)-for analyzing methylated DNA immunoprecipitation (MeDIP) profiles generated using oligonucleotide arrays (MeDIP-chip) or next-generation sequencing (MeDIP-seq). We developed the latter approach to provide a high-resolution whole-genome DNA methylation profile (DNA methylome) of a mammalian genome. Strong correlation of our data, obtained using mature human spermatozoa, with those obtained using bisulfite sequencing suggest that combining MeDIP-seq or MeDIP-chip with Batman provides a robust, quantitative and cost-effective functional genomic strategy for elucidating the function of DNA methylation.
@Article{18612301, author = {Down TA and Rakyan VK and Turner DJ and Flicek P and Li H and Kulesha E and Gräf S and Johnson N and Herrero J and Tomazou EM and Thorne NP and Bäckdahl L and Herberth M and Howe KL and Jackson DK and Miretti MM and Marioni JC and Birney E and Hubbard TJP and Durbin R and Tavaré S and Beck S}, title = {A Bayesian deconvolution strategy for immunoprecipitation-based DNA methylome analysis}, journal = {Nat Biotechnol}, volume = {26}, number = {7}, pages = {779--785}, year = {2008}, doi = {10.1038/nbt1414}, howpublished = {Advanced online publication: 8 July 2008}, abstract = {DNA methylation is an indispensible epigenetic modification required for regulating the expression of mammalian genomes. Immunoprecipitation-based methods for DNA methylome analysis are rapidly shifting the bottleneck in this field from data generation to data analysis, necessitating the development of better analytical tools. In particular, an inability to estimate absolute methylation levels remains a major analytical difficulty associated with immunoprecipitation-based DNA methylation profiling. To address this issue, we developed a cross-platform algorithm-Bayesian tool for methylation analysis (Batman)-for analyzing methylated DNA immunoprecipitation (MeDIP) profiles generated using oligonucleotide arrays (MeDIP-chip) or next-generation sequencing (MeDIP-seq). We developed the latter approach to provide a high-resolution whole-genome DNA methylation profile (DNA methylome) of a mammalian genome. Strong correlation of our data, obtained using mature human spermatozoa, with those obtained using bisulfite sequencing suggest that combining MeDIP-seq or MeDIP-chip with Batman provides a robust, quantitative and cost-effective functional genomic strategy for elucidating the function of DNA methylation.},} - VK Rakyan, TA Down, NP Thorne, P Flicek, E Kulesha, S Gräf, EM Tomazou, L Bäckdahl, N Johnson, M Herberth, KL Howe, DK Jackson, MM Miretti, H Fiegler, JC Marioni, E Birney, TJP Hubbard, NP Carter, S Tavaré, S Beck. An integrated resource for genome-wide identification and analysis of human tissue-specific differentially methylated regions (tDMRs). Genome Res 2008;18(9):1518–1529. doi:10.1101/gr.077479.108
[BibTeX] [Abstract]
We report a novel resource (methylation profiles of DNA, or mPod) for human genome-wide tissue-specific DNA methylation profiles. mPod consists of three fully integrated parts, genome-wide DNA methylation reference profiles of 13 normal somatic tissues, placenta, sperm, and an immortalized cell line, a visualization tool that has been integrated with the Ensembl genome browser and a new algorithm for the analysis of immunoprecipitation-based DNA methylation profiles. We demonstrate the utility of our resource by identifying the first comprehensive genome-wide set of tissue-specific differentially methylated regions (tDMRs) that may play a role in cellular identity and the regulation of tissue-specific genome function. We also discuss the implications of our findings with respect to the regulatory potential of regions with varied CpG density, gene expression, transcription factor motifs, gene ontology, and correlation with other epigenetic marks such as histone modifications.
@Article{18577705, author = {Rakyan VK and Down TA and Thorne NP and Flicek P and Kulesha E and Gräf S and Tomazou EM and Bäckdahl L and Johnson N and Herberth M and Howe KL and Jackson DK and Miretti MM and Fiegler H and Marioni JC and Birney E and Hubbard TJP and Carter NP and Tavaré S and Beck S}, title = {An integrated resource for genome-wide identification and analysis of human tissue-specific differentially methylated regions (tDMRs)}, journal = {Genome Res}, volume = {18}, number = {9}, pages = {1518--1529}, year = {2008}, doi = {10.1101/gr.077479.108}, howpublished = {Advanced online publication: 24 June 2008}, abstract = {We report a novel resource (methylation profiles of DNA, or mPod) for human genome-wide tissue-specific DNA methylation profiles. mPod consists of three fully integrated parts, genome-wide DNA methylation reference profiles of 13 normal somatic tissues, placenta, sperm, and an immortalized cell line, a visualization tool that has been integrated with the Ensembl genome browser and a new algorithm for the analysis of immunoprecipitation-based DNA methylation profiles. We demonstrate the utility of our resource by identifying the first comprehensive genome-wide set of tissue-specific differentially methylated regions (tDMRs) that may play a role in cellular identity and the regulation of tissue-specific genome function. We also discuss the implications of our findings with respect to the regulatory potential of regions with varied CpG density, gene expression, transcription factor motifs, gene ontology, and correlation with other epigenetic marks such as histone modifications.},} - P Flicek, BL Aken, K Beal, B Ballester, M Caccamo, Y Chen, L Clarke, G Coates, F Cunningham, T Cutts, T Down, SC Dyer, T Eyre, S Fitzgerald, J Fernandez-Banet, S Gräf, S Haider, M Hammond, R Holland, KL Howe, K Howe, N Johnson, A Jenkinson, A Kähäri, D Keefe, F Kokocinski, E Kulesha, D Lawson, I Longden, K Megy, P Meidl, B Overduin, A Parker, B Pritchard, A Prlic, S Rice, D Rios, M Schuster, I Sealy, G Slater, D Smedley, G Spudich, S Trevanion, AJ Vilella, J Vogel, S White, M Wood, E Birney, T Cox, V Curwen, R Durbin, XM Fernandez-Suarez, J Herrero, TJP Hubbard, A Kasprzyk, G Proctor, J Smith, A Ureta-Vidal, S Searle. Ensembl 2008. Nucleic Acids Res 2008;36(Database issue):D707–14. doi:10.1093/nar/gkm988
[BibTeX] [Abstract]
The Ensembl project (http://www.ensembl.org) is a comprehensive genome information system featuring an integrated set of genome annotation, databases and other information for chordate and selected model organism and disease vector genomes. As of release 47 (October 2007), Ensembl fully supports 35 species, with preliminary support for six additional species. New species in the past year include platypus and horse. Major additions and improvements to Ensembl since our previous report include extensive support for functional genomics data in the form of a specialized functional genomics database, genome-wide maps of protein-DNA interactions and the Ensembl regulatory build; support for customization of the Ensembl web interface through the addition of user accounts and user groups; and increased support for genome resequencing. We have also introduced new comparative genomics-based data mining options and report on the continued development of our software infrastructure.
@Article{18000006, author = {Flicek P and Aken BL and Beal K and Ballester B and Caccamo M and Chen Y and Clarke L and Coates G and Cunningham F and Cutts T and Down T and Dyer SC and Eyre T and Fitzgerald S and Fernandez-Banet J and Gräf S and Haider S and Hammond M and Holland R and Howe KL and Howe K and Johnson N and Jenkinson A and Kähäri A and Keefe D and Kokocinski F and Kulesha E and Lawson D and Longden I and Megy K and Meidl P and Overduin B and Parker A and Pritchard B and Prlic A and Rice S and Rios D and Schuster M and Sealy I and Slater G and Smedley D and Spudich G and Trevanion S and Vilella AJ and Vogel J and White S and Wood M and Birney E and Cox T and Curwen V and Durbin R and Fernandez-Suarez XM and Herrero J and Hubbard TJP and Kasprzyk A and Proctor G and Smith J and Ureta-Vidal A and Searle S}, title = {Ensembl 2008}, journal = {Nucleic Acids Res}, volume = {36}, number = {Database issue}, pages = {D707--14}, year = {2008}, doi = {10.1093/nar/gkm988}, howpublished = {Advanced online publication: 13 November 2007}, abstract = {The Ensembl project (http://www.ensembl.org) is a comprehensive genome information system featuring an integrated set of genome annotation, databases and other information for chordate and selected model organism and disease vector genomes. As of release 47 (October 2007), Ensembl fully supports 35 species, with preliminary support for six additional species. New species in the past year include platypus and horse. Major additions and improvements to Ensembl since our previous report include extensive support for functional genomics data in the form of a specialized functional genomics database, genome-wide maps of protein-DNA interactions and the Ensembl regulatory build; support for customization of the Ensembl web interface through the addition of user accounts and user groups; and increased support for genome resequencing. We have also introduced new comparative genomics-based data mining options and report on the continued development of our software infrastructure.},} - RGH Cotton, AD Auerbach, M Axton, CI Barash, SF Berkovic, AJ Brookes, J Burn, G Cutting, den JT Dunnen, P Flicek, N Freimer, MS Greenblatt, HJ Howard, M Katz, FA Macrae, D Maglott, G Möslein, S Povey, RS Ramesar, CS Richards, D Seminara, TD Smith, M-J Sobrido, JH Solbakk, RE Tanzi, SV Tavtigian, GR Taylor, J Utsunomiya, M Watson. GENETICS. The Human Variome Project. Science 2008;322(5903):861–862. doi:10.1126/science.1167363
[BibTeX] [Abstract]
An ambitious plan to collect, curate, and make accessible information on genetic variations affecting human health is beginning to be realized.
@Article{18988827, author = {Cotton RGH and Auerbach AD and Axton M and Barash CI and Berkovic SF and Brookes AJ and Burn J and Cutting G and den Dunnen JT and Flicek P and Freimer N and Greenblatt MS and Howard HJ and Katz M and Macrae FA and Maglott D and Möslein G and Povey S and Ramesar RS and Richards CS and Seminara D and Smith TD and Sobrido M-J and Solbakk JH and Tanzi RE and Tavtigian SV and Taylor GR and Utsunomiya J and Watson M}, title = {GENETICS. The Human Variome Project}, journal = {Science}, volume = {322}, number = {5903}, pages = {861--862}, year = {2008}, doi = {10.1126/science.1167363}, abstract = {An ambitious plan to collect, curate, and make accessible information on genetic variations affecting human health is beginning to be realized.},} - WC Warren, LW Hillier, JA Marshall Graves, E Birney, CP Ponting, F Grützner, K Belov, W Miller, L Clarke, AT Chinwalla, S-P Yang, A Heger, DP Locke, P Miethke, PD Waters, F Veyrunes, L Fulton, B Fulton, T Graves, J Wallis, XS Puente, C López-Otín, GR Ordóñez, EE Eichler, L Chen, Z Cheng, JE Deakin, A Alsop, K Thompson, P Kirby, AT Papenfuss, MJ Wakefield, T Olender, D Lancet, GA Huttley, AFA Smit, A Pask, P Temple-Smith, MA Batzer, JA Walker, MK Konkel, RS Harris, CM Whittington, ESW Wong, NJ Gemmell, E Buschiazzo, IM Vargas Jentzsch, A Merkel, J Schmitz, A Zemann, G Churakov, JO Kriegs, J Brosius, EP Murchison, R Sachidanandam, C Smith, GJ Hannon, E Tsend-Ayush, D McMillan, R Attenborough, W Rens, M Ferguson-Smith, CM Lefèvre, JA Sharp, KR Nicholas, DA Ray, M Kube, R Reinhardt, TH Pringle, J Taylor, RC Jones, B Nixon, J-L Dacheux, H Niwa, Y Sekita, X Huang, A Stark, P Kheradpour, M Kellis, P Flicek, Y Chen, C Webber, R Hardison, J Nelson, K Hallsworth-Pepin, K Delehaunty, C Markovic, P Minx, Y Feng, C Kremitzki, M Mitreva, J Glasscock, T Wylie, P Wohldmann, P Thiru, MN Nhan, CS Pohl, SM Smith, S Hou, MB Renfree, ER Mardis, RK Wilson. Genome analysis of the platypus reveals unique signatures of evolution. Nature 2008;453(7192):175–183. doi:10.1038/nature06936
[BibTeX] [Abstract]
We present a draft genome sequence of the platypus, Ornithorhynchus anatinus. This monotreme exhibits a fascinating combination of reptilian and mammalian characters. For example, platypuses have a coat of fur adapted to an aquatic lifestyle; platypus females lactate, yet lay eggs; and males are equipped with venom similar to that of reptiles. Analysis of the first monotreme genome aligned these features with genetic innovations. We find that reptile and platypus venom proteins have been co-opted independently from the same gene families; milk protein genes are conserved despite platypuses laying eggs; and immune gene family expansions are directly related to platypus biology. Expansions of protein, non-protein-coding RNA and microRNA families, as well as repeat elements, are identified. Sequencing of this genome now provides a valuable resource for deep mammalian comparative analyses, as well as for monotreme biology and conservation.
@Article{18464734, author = {Warren WC and Hillier LW and Marshall Graves JA and Birney E and Ponting CP and Grützner F and Belov K and Miller W and Clarke L and Chinwalla AT and Yang S-P and Heger A and Locke DP and Miethke P and Waters PD and Veyrunes F and Fulton L and Fulton B and Graves T and Wallis J and Puente XS and López-Otín C and Ordóñez GR and Eichler EE and Chen L and Cheng Z and Deakin JE and Alsop A and Thompson K and Kirby P and Papenfuss AT and Wakefield MJ and Olender T and Lancet D and Huttley GA and Smit AFA and Pask A and Temple-Smith P and Batzer MA and Walker JA and Konkel MK and Harris RS and Whittington CM and Wong ESW and Gemmell NJ and Buschiazzo E and Vargas Jentzsch IM and Merkel A and Schmitz J and Zemann A and Churakov G and Kriegs JO and Brosius J and Murchison EP and Sachidanandam R and Smith C and Hannon GJ and Tsend-Ayush E and McMillan D and Attenborough R and Rens W and Ferguson-Smith M and Lefèvre CM and Sharp JA and Nicholas KR and Ray DA and Kube M and Reinhardt R and Pringle TH and Taylor J and Jones RC and Nixon B and Dacheux J-L and Niwa H and Sekita Y and Huang X and Stark A and Kheradpour P and Kellis M and Flicek P and Chen Y and Webber C and Hardison R and Nelson J and Hallsworth-Pepin K and Delehaunty K and Markovic C and Minx P and Feng Y and Kremitzki C and Mitreva M and Glasscock J and Wylie T and Wohldmann P and Thiru P and Nhan MN and Pohl CS and Smith SM and Hou S and Renfree MB and Mardis ER and Wilson RK}, title = {Genome analysis of the platypus reveals unique signatures of evolution}, journal = {Nature}, volume = {453}, number = {7192}, pages = {175--183}, year = {2008}, doi = {10.1038/nature06936}, abstract = {We present a draft genome sequence of the platypus, Ornithorhynchus anatinus. This monotreme exhibits a fascinating combination of reptilian and mammalian characters. For example, platypuses have a coat of fur adapted to an aquatic lifestyle; platypus females lactate, yet lay eggs; and males are equipped with venom similar to that of reptiles. Analysis of the first monotreme genome aligned these features with genetic innovations. We find that reptile and platypus venom proteins have been co-opted independently from the same gene families; milk protein genes are conserved despite platypuses laying eggs; and immune gene family expansions are directly related to platypus biology. Expansions of protein, non-protein-coding RNA and microRNA families, as well as repeat elements, are identified. Sequencing of this genome now provides a valuable resource for deep mammalian comparative analyses, as well as for monotreme biology and conservation.},} - B Paten, J Herrero, S Fitzgerald, K Beal, P Flicek, I Holmes, E Birney. Genome-wide nucleotide-level mammalian ancestor reconstruction. Genome Res 2008;18(11):1829–1843. doi:10.1101/gr.076521.108
[BibTeX] [Abstract]
Recently attention has been turned to the problem of reconstructing complete ancestral sequences from large multiple alignments. Successful generation of these genome-wide reconstructions will facilitate a greater knowledge of the events that have driven evolution. We present a new evolutionary alignment modeler, called ``Ortheus,'' for inferring the evolutionary history of a multiple alignment, in terms of both substitutions and, importantly, insertions and deletions. Based on a multiple sequence probabilistic transducer model of the type proposed by Holmes, Ortheus uses efficient stochastic graph-based dynamic programming methods. Unlike other methods, Ortheus does not rely on a single fixed alignment from which to work. Ortheus is also more scaleable than previous methods while being fast, stable, and open source. Large-scale simulations show that Ortheus performs close to optimally on a deep mammalian phylogeny. Simulations also indicate that significant proportions of errors due to insertions and deletions can be avoided by not assuming a fixed alignment. We additionally use a challenging hold-out cross-validation procedure to test the method; using the reconstructions to predict extant sequence bases, we demonstrate significant improvements over using closest extant neighbor sequences. Accompanying this paper, a new, public, and genome-wide set of Ortheus ancestor alignments provide an intriguing new resource for evolutionary studies in mammals. As a first piece of analysis, we attempt to recover ``fossilized'' ancestral pseudogenes. We confidently find 31 cases in which the ancestral sequence had a more complete sequence than any of the extant sequences.
@Article{18849525, author = {Paten B and Herrero J and Fitzgerald S and Beal K and Flicek P and Holmes I and Birney E}, title = {Genome-wide nucleotide-level mammalian ancestor reconstruction}, journal = {Genome Res}, volume = {18}, number = {11}, pages = {1829--1843}, year = {2008}, doi = {10.1101/gr.076521.108}, howpublished = {Advanced online publication: 10 October 2008}, abstract = {Recently attention has been turned to the problem of reconstructing complete ancestral sequences from large multiple alignments. Successful generation of these genome-wide reconstructions will facilitate a greater knowledge of the events that have driven evolution. We present a new evolutionary alignment modeler, called ``Ortheus,'' for inferring the evolutionary history of a multiple alignment, in terms of both substitutions and, importantly, insertions and deletions. Based on a multiple sequence probabilistic transducer model of the type proposed by Holmes, Ortheus uses efficient stochastic graph-based dynamic programming methods. Unlike other methods, Ortheus does not rely on a single fixed alignment from which to work. Ortheus is also more scaleable than previous methods while being fast, stable, and open source. Large-scale simulations show that Ortheus performs close to optimally on a deep mammalian phylogeny. Simulations also indicate that significant proportions of errors due to insertions and deletions can be avoided by not assuming a fixed alignment. We additionally use a challenging hold-out cross-validation procedure to test the method; using the reconstructions to predict extant sequence bases, we demonstrate significant improvements over using closest extant neighbor sequences. Accompanying this paper, a new, public, and genome-wide set of Ortheus ancestor alignments provide an intriguing new resource for evolutionary studies in mammals. As a first piece of analysis, we attempt to recover ``fossilized'' ancestral pseudogenes. We confidently find 31 cases in which the ancestral sequence had a more complete sequence than any of the extant sequences.},} - A Coghlan, TJ Fiedler, SJ McKay, P Flicek, TW Harris, D Blasiar, TN Consortium, LD Stein. nGASP - the nematode genome annotation assessment project. BMC Bioinformatics 2008;9(1):549. doi:10.1186/1471-2105-9-549
[BibTeX] [Abstract]
ABSTRACT: BACKGROUND: ========== While the C. elegans genome is extensively annotated, relatively little information is available for other Caenorhabditis species. The nematode genome annotation assessment project (nGASP) was launched to objectively assess the accuracy of protein-coding gene prediction software in C. elegans, and to apply this knowledge to the annotation of the genomes of four additional Caenorhabditis species and other nematodes. Seventeen groups worldwide participated in nGASP, and submitted 47 prediction sets across 10 Mb of the C. elegans genome. Predictions were compared to reference gene sets consisting of confirmed or manually curated gene models from WormBase. RESULTS: ======= The most accurate gene-finders were `combiner' algorithms, which made use of transcript- and protein-alignments and multi-genome alignments, as well as gene predictions from other gene-finders. Gene-finders that used alignments of ESTs, mRNAs and proteins came in second. There was a tie for third place between gene-finders that used multi-genome alignments and ab initio gene-finders. The median gene level sensitivity of combiners was 78\% and their specificity was 42\%, which is nearly the same accuracy reported for combiners in the human genome. C. elegans genes with exons of unusual hexamer content, as well as those with unusually many exons, short exons, long introns, a weak translation start signal, weak splice sites, or poorly conserved orthologs posed the greatest difficulty for gene-finders. CONCLUSIONS: =========== This experiment establishes a baseline of gene prediction accuracy in Caenorhabditis genomes, and has guided the choice of gene-finders for the annotation of newly sequenced genomes of Caenorhabditis and other nematode species. We have created new gene sets for C. briggsae, C. remanei, C. brenneri, C. japonica, and Brugia malayi using some of the best-performing gene-finders.
@Article{19099578, author = {Coghlan A and Fiedler TJ and McKay SJ and Flicek P and Harris TW and Blasiar D and Consortium TN and Stein LD}, title = {nGASP - the nematode genome annotation assessment project}, journal = {BMC Bioinformatics}, volume = {9}, number = {1}, pages = {549}, year = {2008}, doi = {10.1186/1471-2105-9-549}, abstract = {ABSTRACT: BACKGROUND: ========== While the C. elegans genome is extensively annotated, relatively little information is available for other Caenorhabditis species. The nematode genome annotation assessment project (nGASP) was launched to objectively assess the accuracy of protein-coding gene prediction software in C. elegans, and to apply this knowledge to the annotation of the genomes of four additional Caenorhabditis species and other nematodes. Seventeen groups worldwide participated in nGASP, and submitted 47 prediction sets across 10 Mb of the C. elegans genome. Predictions were compared to reference gene sets consisting of confirmed or manually curated gene models from WormBase. RESULTS: ======= The most accurate gene-finders were `combiner' algorithms, which made use of transcript- and protein-alignments and multi-genome alignments, as well as gene predictions from other gene-finders. Gene-finders that used alignments of ESTs, mRNAs and proteins came in second. There was a tie for third place between gene-finders that used multi-genome alignments and ab initio gene-finders. The median gene level sensitivity of combiners was 78\% and their specificity was 42\%, which is nearly the same accuracy reported for combiners in the human genome. C. elegans genes with exons of unusual hexamer content, as well as those with unusually many exons, short exons, long introns, a weak translation start signal, weak splice sites, or poorly conserved orthologs posed the greatest difficulty for gene-finders. CONCLUSIONS: =========== This experiment establishes a baseline of gene prediction accuracy in Caenorhabditis genomes, and has guided the choice of gene-finders for the annotation of newly sequenced genomes of Caenorhabditis and other nematode species. We have created new gene sets for C. briggsae, C. remanei, C. brenneri, C. japonica, and Brugia malayi using some of the best-performing gene-finders.},} - STAR Consortium.. SNP and haplotype mapping for genetic analysis in the rat. Nat Genet 2008;40(5):560–566. doi:10.1038/ng.124
[BibTeX] [Abstract]
The laboratory rat is one of the most extensively studied model organisms. Inbred laboratory rat strains originated from limited Rattus norvegicus founder populations, and the inherited genetic variation provides an excellent resource for the correlation of genotype to phenotype. Here, we report a survey of genetic variation based on almost 3 million newly identified SNPs. We obtained accurate and complete genotypes for a subset of 20,238 SNPs across 167 distinct inbred rat strains, two rat recombinant inbred panels and an F2 intercross. Using 81\% of these SNPs, we constructed high-density genetic maps, creating a large dataset of fully characterized SNPs for disease gene mapping. Our data characterize the population structure and illustrate the degree of linkage disequilibrium. We provide a detailed SNP map and demonstrate its utility for mapping of quantitative trait loci. This community resource is openly available and augments the genetic tools for this workhorse of physiological studies.
@Article{18443594, author = {{STAR Consortium.}}, title = {SNP and haplotype mapping for genetic analysis in the rat}, journal = {Nat Genet}, volume = {40}, number = {5}, pages = {560--566}, year = {2008}, doi = {10.1038/ng.124}, abstract = {The laboratory rat is one of the most extensively studied model organisms. Inbred laboratory rat strains originated from limited Rattus norvegicus founder populations, and the inherited genetic variation provides an excellent resource for the correlation of genotype to phenotype. Here, we report a survey of genetic variation based on almost 3 million newly identified SNPs. We obtained accurate and complete genotypes for a subset of 20,238 SNPs across 167 distinct inbred rat strains, two rat recombinant inbred panels and an F2 intercross. Using 81\% of these SNPs, we constructed high-density genetic maps, creating a large dataset of fully characterized SNPs for disease gene mapping. Our data characterize the population structure and illustrate the degree of linkage disequilibrium. We provide a detailed SNP map and demonstrate its utility for mapping of quantitative trait loci. This community resource is openly available and augments the genetic tools for this workhorse of physiological studies.},} - DS Johnson, W Li, DB Gordon, A Bhattacharjee, B Curry, J Ghosh, L Brizuela, JS Carroll, M Brown, P Flicek, CM Koch, I Dunham, M Bieda, X Xu, PJ Farnham, P Kapranov, DA Nix, TR Gingeras, X Zhang, H Holster, N Jiang, RD Green, JS Song, SA McCuine, E Anton, L Nguyen, ND Trinklein, Z Ye, K Ching, D Hawkins, B Ren, PC Scacheri, J Rozowsky, A Karpikov, G Euskirchen, S Weissman, M Gerstein, M Snyder, A Yang, Z Moqtaderi, H Hirsch, HP Shulha, Y Fu, Z Weng, K Struhl, RM Myers, JD Lieb, XS Liu. Systematic evaluation of variability in ChIP-chip experiments using predefined DNA targets. Genome Res 2008;18(3):393–403. doi:10.1101/gr.7080508
[BibTeX] [Abstract]
The most widely used method for detecting genome-wide protein-DNA interactions is chromatin immunoprecipitation on tiling microarrays, commonly known as ChIP-chip. Here, we conducted the first objective analysis of tiling array platforms, amplification procedures, and signal detection algorithms in a simulated ChIP-chip experiment. Mixtures of human genomic DNA and ``spike-ins'' comprised of nearly 100 human sequences at various concentrations were hybridized to four tiling array platforms by eight independent groups. Blind to the number of spike-ins, their locations, and the range of concentrations, each group made predictions of the spike-in locations. We found that microarray platform choice is not the primary determinant of overall performance. In fact, variation in performance between labs, protocols, and algorithms within the same array platform was greater than the variation in performance between array platforms. However, each array platform had unique performance characteristics that varied with tiling resolution and the number of replicates, which have implications for cost versus detection power. Long oligonucleotide arrays were slightly more sensitive at detecting very low enrichment. On all platforms, simple sequence repeats and genome redundancy tended to result in false positives. LM-PCR and WGA, the most popular sample amplification techniques, reproduced relative enrichment levels with high fidelity. Performance among signal detection algorithms was heavily dependent on array platform. The spike-in DNA samples and the data presented here provide a stable benchmark against which future ChIP platforms, protocol improvements, and analysis methods can be evaluated.
@Article{18258921, author = {Johnson DS and Li W and Gordon DB and Bhattacharjee A and Curry B and Ghosh J and Brizuela L and Carroll JS and Brown M and Flicek P and Koch CM and Dunham I and Bieda M and Xu X and Farnham PJ and Kapranov P and Nix DA and Gingeras TR and Zhang X and Holster H and Jiang N and Green RD and Song JS and McCuine SA and Anton E and Nguyen L and Trinklein ND and Ye Z and Ching K and Hawkins D and Ren B and Scacheri PC and Rozowsky J and Karpikov A and Euskirchen G and Weissman S and Gerstein M and Snyder M and Yang A and Moqtaderi Z and Hirsch H and Shulha HP and Fu Y and Weng Z and Struhl K and Myers RM and Lieb JD and Liu XS}, title = {Systematic evaluation of variability in ChIP-chip experiments using predefined DNA targets}, journal = {Genome Res}, volume = {18}, number = {3}, pages = {393--403}, year = {2008}, doi = {10.1101/gr.7080508}, howpublished = {Advanced online publication: 7 February 2008}, abstract = {The most widely used method for detecting genome-wide protein-DNA interactions is chromatin immunoprecipitation on tiling microarrays, commonly known as ChIP-chip. Here, we conducted the first objective analysis of tiling array platforms, amplification procedures, and signal detection algorithms in a simulated ChIP-chip experiment. Mixtures of human genomic DNA and ``spike-ins'' comprised of nearly 100 human sequences at various concentrations were hybridized to four tiling array platforms by eight independent groups. Blind to the number of spike-ins, their locations, and the range of concentrations, each group made predictions of the spike-in locations. We found that microarray platform choice is not the primary determinant of overall performance. In fact, variation in performance between labs, protocols, and algorithms within the same array platform was greater than the variation in performance between array platforms. However, each array platform had unique performance characteristics that varied with tiling resolution and the number of replicates, which have implications for cost versus detection power. Long oligonucleotide arrays were slightly more sensitive at detecting very low enrichment. On all platforms, simple sequence repeats and genome redundancy tended to result in false positives. LM-PCR and WGA, the most popular sample amplification techniques, reproduced relative enrichment levels with high fidelity. Performance among signal detection algorithms was heavily dependent on array platform. The spike-in DNA samples and the data presented here provide a stable benchmark against which future ChIP platforms, protocol improvements, and analysis methods can be evaluated.},}
2007
- TJP Hubbard, BL Aken, K Beal, B Ballester, M Caccamo, Y Chen, L Clarke, G Coates, F Cunningham, T Cutts, T Down, SC Dyer, S Fitzgerald, J Fernandez-Banet, S Graf, S Haider, M Hammond, J Herrero, R Holland, K Howe, N Johnson, A Kahari, D Keefe, F Kokocinski, E Kulesha, D Lawson, I Longden, C Melsopp, K Megy, P Meidl, B Ouverdin, A Parker, A Prlic, S Rice, D Rios, M Schuster, I Sealy, J Severin, G Slater, D Smedley, G Spudich, S Trevanion, A Vilella, J Vogel, S White, M Wood, T Cox, V Curwen, R Durbin, XM Fernandez-Suarez, P Flicek, A Kasprzyk, G Proctor, S Searle, J Smith, A Ureta-Vidal, E Birney. Ensembl 2007. Nucleic Acids Res 2007;35(Database issue):D610–7. doi:10.1093/nar/gkl996
[BibTeX] [Abstract]
The Ensembl (http://www.ensembl.org/) project provides a comprehensive and integrated source of annotation of chordate genome sequences. Over the past year the number of genomes available from Ensembl has increased from 15 to 33, with the addition of sites for the mammalian genomes of elephant, rabbit, armadillo, tenrec, platypus, pig, cat, bush baby, common shrew, microbat and european hedgehog; the fish genomes of stickleback and medaka and the second example of the genomes of the sea squirt (Ciona savignyi) and the mosquito (Aedes aegypti). Some of the major features added during the year include the first complete gene sets for genomes with low-sequence coverage, the introduction of new strain variation data and the introduction of new orthology/paralog annotations based on gene trees.
@Article{17148474, author = {Hubbard TJP and Aken BL and Beal K and Ballester B and Caccamo M and Chen Y and Clarke L and Coates G and Cunningham F and Cutts T and Down T and Dyer SC and Fitzgerald S and Fernandez-Banet J and Graf S and Haider S and Hammond M and Herrero J and Holland R and Howe K and Johnson N and Kahari A and Keefe D and Kokocinski F and Kulesha E and Lawson D and Longden I and Melsopp C and Megy K and Meidl P and Ouverdin B and Parker A and Prlic A and Rice S and Rios D and Schuster M and Sealy I and Severin J and Slater G and Smedley D and Spudich G and Trevanion S and Vilella A and Vogel J and White S and Wood M and Cox T and Curwen V and Durbin R and Fernandez-Suarez XM and Flicek P and Kasprzyk A and Proctor G and Searle S and Smith J and Ureta-Vidal A and Birney E}, title = {Ensembl 2007}, journal = {Nucleic Acids Res}, volume = {35}, number = {Database issue}, pages = {D610--7}, year = {2007}, doi = {10.1093/nar/gkl996}, howpublished = {Advanced online publication: 5 December 2006}, abstract = {The Ensembl (http://www.ensembl.org/) project provides a comprehensive and integrated source of annotation of chordate genome sequences. Over the past year the number of genomes available from Ensembl has increased from 15 to 33, with the addition of sites for the mammalian genomes of elephant, rabbit, armadillo, tenrec, platypus, pig, cat, bush baby, common shrew, microbat and european hedgehog; the fish genomes of stickleback and medaka and the second example of the genomes of the sea squirt (Ciona savignyi) and the mosquito (Aedes aegypti). Some of the major features added during the year include the first complete gene sets for genomes with low-sequence coverage, the introduction of new strain variation data and the introduction of new orthology/paralog annotations based on gene trees.},} - P Flicek. Gene prediction: compare and CONTRAST. Genome Biol 2007;8(12):233. doi:10.1186/gb-2007-8-12-233
[BibTeX] [Abstract]
CONTRAST, a new gene-prediction algorithm that uses sophisticated machine-learning techniques, has pushed de novo prediction accuracy to new heights, and has significantly closed the gap between de novo and evidence-based methods for human genome annotation.
@Article{18096089, author = {Flicek P}, title = {Gene prediction: compare and CONTRAST}, journal = {Genome Biol}, volume = {8}, number = {12}, pages = {233}, year = {2007}, doi = {10.1186/gb-2007-8-12-233}, abstract = {CONTRAST, a new gene-prediction algorithm that uses sophisticated machine-learning techniques, has pushed de novo prediction accuracy to new heights, and has significantly closed the gap between de novo and evidence-based methods for human genome annotation.},} - ENCODE Project Consortium.. Identification and analysis of functional elements in 1\% of the human genome by the ENCODE pilot project. Nature 2007;447(7146):799–816. doi:10.1038/nature05874
[BibTeX] [Abstract]
We report the generation and analysis of functional data from multiple, diverse experiments performed on a targeted 1\% of the human genome as part of the pilot phase of the ENCODE Project. These data have been further integrated and augmented by a number of evolutionary and computational analyses. Together, our results advance the collective knowledge about human genome function in several major areas. First, our studies provide convincing evidence that the genome is pervasively transcribed, such that the majority of its bases can be found in primary transcripts, including non-protein-coding transcripts, and those that extensively overlap one another. Second, systematic examination of transcriptional regulation has yielded new understanding about transcription start sites, including their relationship to specific regulatory sequences and features of chromatin accessibility and histone modification. Third, a more sophisticated view of chromatin structure has emerged, including its inter-relationship with DNA replication and transcriptional regulation. Finally, integration of these new sources of information, in particular with respect to mammalian evolution based on inter- and intra-species sequence comparisons, has yielded new mechanistic and evolutionary insights concerning the functional landscape of the human genome. Together, these studies are defining a path for pursuit of a more comprehensive characterization of human genome function.
@Article{17571346, author = {{ENCODE Project Consortium.}}, title = {Identification and analysis of functional elements in 1\% of the human genome by the ENCODE pilot project}, journal = {Nature}, volume = {447}, number = {7146}, pages = {799--816}, year = {2007}, doi = {10.1038/nature05874}, abstract = {We report the generation and analysis of functional data from multiple, diverse experiments performed on a targeted 1\% of the human genome as part of the pilot phase of the ENCODE Project. These data have been further integrated and augmented by a number of evolutionary and computational analyses. Together, our results advance the collective knowledge about human genome function in several major areas. First, our studies provide convincing evidence that the genome is pervasively transcribed, such that the majority of its bases can be found in primary transcripts, including non-protein-coding transcripts, and those that extensively overlap one another. Second, systematic examination of transcriptional regulation has yielded new understanding about transcription start sites, including their relationship to specific regulatory sequences and features of chromatin accessibility and histone modification. Third, a more sophisticated view of chromatin structure has emerged, including its inter-relationship with DNA replication and transcriptional regulation. Finally, integration of these new sources of information, in particular with respect to mammalian evolution based on inter- and intra-species sequence comparisons, has yielded new mechanistic and evolutionary insights concerning the functional landscape of the human genome. Together, these studies are defining a path for pursuit of a more comprehensive characterization of human genome function.},} - S Gräf, FGG Nielsen, S Kurtz, MA Huynen, E Birney, H Stunnenberg, P Flicek. Optimized design and assessment of whole genome tiling arrays. Bioinformatics 2007;23(13):i195–204. doi:10.1093/bioinformatics/btm200
[BibTeX] [Abstract]
MOTIVATION: Recent advances in microarray technologies have made it feasible to interrogate whole genomes with tiling arrays and this technique is rapidly becoming one of the most important high-throughput functional genomics assays. For large mammalian genomes, analyzing oligonucleotide tiling array data is complicated by the presence of non-unique sequences on the array, which increases the overall noise in the data and may lead to false positive results due to cross-hybridization. The ability to create custom microarrays using maskless array synthesis has led us to consider ways to optimize array design characteristics for improving data quality and analysis. We have identified a number of design parameters to be optimized including uniqueness of the probe sequences within the whole genome, melting temperature and self-hybridization potential. RESULTS: We introduce the uniqueness score, U, a novel quality measure for oligonucleotide probes and present a method to quickly compute it. We show that U is equivalent to the number of shortest unique substrings in the probe and describe an efficient greedy algorithm to design mammalian whole genome tiling arrays using probes that maximize U. Using the mouse genome, we demonstrate how several optimizations influence the tiling array design characteristics. With a sensible set of parameters, our designs cover 78\% of the mouse genome including many regions previously considered `untilable' due to the presence of repetitive sequence. Finally, we compare our whole genome tiling array designs with commercially available designs. AVAILABILITY: Source code is available under an open source license from http://www.ebi.ac.uk/~graef/arraydesign/.
@Article{17646297, author = {Gräf S and Nielsen FGG and Kurtz S and Huynen MA and Birney E and Stunnenberg H and Flicek P}, title = {Optimized design and assessment of whole genome tiling arrays}, journal = {Bioinformatics}, volume = {23}, number = {13}, pages = {i195--204}, year = {2007}, doi = {10.1093/bioinformatics/btm200}, abstract = {MOTIVATION: Recent advances in microarray technologies have made it feasible to interrogate whole genomes with tiling arrays and this technique is rapidly becoming one of the most important high-throughput functional genomics assays. For large mammalian genomes, analyzing oligonucleotide tiling array data is complicated by the presence of non-unique sequences on the array, which increases the overall noise in the data and may lead to false positive results due to cross-hybridization. The ability to create custom microarrays using maskless array synthesis has led us to consider ways to optimize array design characteristics for improving data quality and analysis. We have identified a number of design parameters to be optimized including uniqueness of the probe sequences within the whole genome, melting temperature and self-hybridization potential. RESULTS: We introduce the uniqueness score, U, a novel quality measure for oligonucleotide probes and present a method to quickly compute it. We show that U is equivalent to the number of shortest unique substrings in the probe and describe an efficient greedy algorithm to design mammalian whole genome tiling arrays using probes that maximize U. Using the mouse genome, we demonstrate how several optimizations influence the tiling array design characteristics. With a sensible set of parameters, our designs cover 78\% of the mouse genome including many regions previously considered `untilable' due to the presence of repetitive sequence. Finally, we compare our whole genome tiling array designs with commercially available designs. AVAILABILITY: Source code is available under an open source license from http://www.ebi.ac.uk/~graef/arraydesign/.},} - BE Stranger, AC Nica, MS Forrest, A Dimas, CP Bird, C Beazley, CE Ingle, M Dunning, P Flicek, D Koller, S Montgomery, S Tavaré, P Deloukas, ET Dermitzakis. Population genomics of human gene expression. Nat Genet 2007;39(10):1217–1224. doi:10.1038/ng2142
[BibTeX] [Abstract]
Genetic variation influences gene expression, and this variation in gene expression can be efficiently mapped to specific genomic regions and variants. Here we have used gene expression profiling of Epstein-Barr virus-transformed lymphoblastoid cell lines of all 270 individuals genotyped in the HapMap Consortium to elucidate the detailed features of genetic variation underlying gene expression variation. We find that gene expression is heritable and that differentiation between populations is in agreement with earlier small-scale studies. A detailed association analysis of over 2.2 million common SNPs per population (5\% frequency in HapMap) with gene expression identified at least 1,348 genes with association signals in cis and at least 180 in trans. Replication in at least one independent population was achieved for 37\% of cis signals and 15\% of trans signals, respectively. Our results strongly support an abundance of cis-regulatory variation in the human genome. Detection of trans effects is limited but suggests that regulatory variation may be the key primary effect contributing to phenotypic variation in humans. We also explore several methodologies that improve the current state of analysis of gene expression variation.
@Article{17873874, author = {Stranger BE and Nica AC and Forrest MS and Dimas A and Bird CP and Beazley C and Ingle CE and Dunning M and Flicek P and Koller D and Montgomery S and Tavaré S and Deloukas P and Dermitzakis ET}, title = {Population genomics of human gene expression}, journal = {Nat Genet}, volume = {39}, number = {10}, pages = {1217--1224}, year = {2007}, doi = {10.1038/ng2142}, abstract = {Genetic variation influences gene expression, and this variation in gene expression can be efficiently mapped to specific genomic regions and variants. Here we have used gene expression profiling of Epstein-Barr virus-transformed lymphoblastoid cell lines of all 270 individuals genotyped in the HapMap Consortium to elucidate the detailed features of genetic variation underlying gene expression variation. We find that gene expression is heritable and that differentiation between populations is in agreement with earlier small-scale studies. A detailed association analysis of over 2.2 million common SNPs per population (5\% frequency in HapMap) with gene expression identified at least 1,348 genes with association signals in cis and at least 180 in trans. Replication in at least one independent population was achieved for 37\% of cis signals and 15\% of trans signals, respectively. Our results strongly support an abundance of cis-regulatory variation in the human genome. Detection of trans effects is limited but suggests that regulatory variation may be the key primary effect contributing to phenotypic variation in humans. We also explore several methodologies that improve the current state of analysis of gene expression variation.},} - RGH Cotton, HV Project, W Appelbe, AD Auerbach, K Becker, W Bodmer, DJ Boone, V Boulyjenkov, S Brahmachari, L Brody, A Brookes, AF Brown, P Byers, JM Cantu, J-J Cassiman, M Claustres, P Concannon, den JT Dunnen, P Flicek, R Gibbs, J Hall, J Hasler, M Katz, P-Y Kwok, S Laradi, A Lindblom, D Maglott, S Marsh, CM Masimirembwa, S Minoshima, de AMO Ramirez, R Pagon, R Ramesar, D Ravine, S Richards, D Rimoin, HZ Ring, CR Scriver, S Sherry, N Shimizu, L Stein, GO Tadmouri, G Taylor, M Watson. Recommendations of the 2006 Human Variome Project meeting. Nat Genet 2007;39(4):433–436. doi:10.1038/ng2024
[BibTeX] [Abstract]
Lists of variations in genomic DNA and their effects have been kept for some time and have been used in diagnostics and research. Although these lists have been carefully gathered and curated, there has been little standardization and coordination, complicating their use. Given the myriad possible variations in the estimated 24,000 genes in the human genome, it would be useful to have standard criteria for databases of variation. Incomplete collection and ascertainment of variants demonstrates a need for a universally accessible system. These and other problems led to the World Heath Organization-cosponsored meeting on June 20-23, 2006 in Melbourne, Australia, which launched the Human Variome Project. This meeting addressed all areas of human genetics relevant to collection of information on variation and its effects. Members of each of eight sessions (the clinic and phenotype, the diagnostic laboratory, the research laboratory, curation and collection, informatics, relevance to the emerging world, integration and federation and funding and sustainability) developed a number of recommendations that were then organized into a total of 96 recommendations to act as a foundation for future work worldwide. Here we summarize the background of the project, the meeting and its recommendations.
@Article{17392799, author = {Cotton RGH and Project HV and Appelbe W and Auerbach AD and Becker K and Bodmer W and Boone DJ and Boulyjenkov V and Brahmachari S and Brody L and Brookes A and Brown AF and Byers P and Cantu JM and Cassiman J-J and Claustres M and Concannon P and den Dunnen JT and Flicek P and Gibbs R and Hall J and Hasler J and Katz M and Kwok P-Y and Laradi S and Lindblom A and Maglott D and Marsh S and Masimirembwa CM and Minoshima S and de Ramirez AMO and Pagon R and Ramesar R and Ravine D and Richards S and Rimoin D and Ring HZ and Scriver CR and Sherry S and Shimizu N and Stein L and Tadmouri GO and Taylor G and Watson M}, title = {Recommendations of the 2006 Human Variome Project meeting}, journal = {Nat Genet}, volume = {39}, number = {4}, pages = {433--436}, year = {2007}, doi = {10.1038/ng2024}, abstract = {Lists of variations in genomic DNA and their effects have been kept for some time and have been used in diagnostics and research. Although these lists have been carefully gathered and curated, there has been little standardization and coordination, complicating their use. Given the myriad possible variations in the estimated 24,000 genes in the human genome, it would be useful to have standard criteria for databases of variation. Incomplete collection and ascertainment of variants demonstrates a need for a universally accessible system. These and other problems led to the World Heath Organization-cosponsored meeting on June 20-23, 2006 in Melbourne, Australia, which launched the Human Variome Project. This meeting addressed all areas of human genetics relevant to collection of information on variation and its effects. Members of each of eight sessions (the clinic and phenotype, the diagnostic laboratory, the research laboratory, curation and collection, informatics, relevance to the emerging world, integration and federation and funding and sustainability) developed a number of recommendations that were then organized into a total of 96 recommendations to act as a foundation for future work worldwide. Here we summarize the background of the project, the meeting and its recommendations.},} - CM Koch, RM Andrews, P Flicek, SC Dillon, U Karaöz, GK Clelland, S Wilcox, DM Beare, JC Fowler, P Couttet, KD James, GC Lefebvre, AW Bruce, OM Dovey, PD Ellis, P Dhami, CF Langford, Z Weng, E Birney, NP Carter, D Vetrie, I Dunham. The landscape of histone modifications across 1\% of the human genome in five human cell lines. Genome Res 2007;17(6):691–707. doi:10.1101/gr.5704207
[BibTeX] [Abstract]
We generated high-resolution maps of histone H3 lysine 9/14 acetylation (H3ac), histone H4 lysine 5/8/12/16 acetylation (H4ac), and histone H3 at lysine 4 mono-, di-, and trimethylation (H3K4me1, H3K4me2, H3K4me3, respectively) across the ENCODE regions. Studying each modification in five human cell lines including the ENCODE Consortium common cell lines GM06990 (lymphoblastoid) and HeLa-S3, as well as K562, HFL-1, and MOLT4, we identified clear patterns of histone modification profiles with respect to genomic features. H3K4me3, H3K4me2, and H3ac modifications are tightly associated with the transcriptional start sites (TSSs) of genes, while H3K4me1 and H4ac have more widespread distributions. TSSs reveal characteristic patterns of both types of modification present and the position relative to TSSs. These patterns differ between active and inactive genes and in particular the state of H3K4me3 and H3ac modifications is highly predictive of gene activity. Away from TSSs, modification sites are enriched in H3K4me1 and relatively depleted in H3K4me3 and H3ac. Comparison between cell lines identified differences in the histone modification profiles associated with transcriptional differences between the cell lines. These results provide an overview of the functional relationship among histone modifications and gene expression in human cells.
@Article{17567990, author = {Koch CM and Andrews RM and Flicek P and Dillon SC and Karaöz U and Clelland GK and Wilcox S and Beare DM and Fowler JC and Couttet P and James KD and Lefebvre GC and Bruce AW and Dovey OM and Ellis PD and Dhami P and Langford CF and Weng Z and Birney E and Carter NP and Vetrie D and Dunham I}, title = {The landscape of histone modifications across 1\% of the human genome in five human cell lines}, journal = {Genome Res}, volume = {17}, number = {6}, pages = {691--707}, year = {2007}, doi = {10.1101/gr.5704207}, abstract = {We generated high-resolution maps of histone H3 lysine 9/14 acetylation (H3ac), histone H4 lysine 5/8/12/16 acetylation (H4ac), and histone H3 at lysine 4 mono-, di-, and trimethylation (H3K4me1, H3K4me2, H3K4me3, respectively) across the ENCODE regions. Studying each modification in five human cell lines including the ENCODE Consortium common cell lines GM06990 (lymphoblastoid) and HeLa-S3, as well as K562, HFL-1, and MOLT4, we identified clear patterns of histone modification profiles with respect to genomic features. H3K4me3, H3K4me2, and H3ac modifications are tightly associated with the transcriptional start sites (TSSs) of genes, while H3K4me1 and H4ac have more widespread distributions. TSSs reveal characteristic patterns of both types of modification present and the position relative to TSSs. These patterns differ between active and inactive genes and in particular the state of H3K4me3 and H3ac modifications is highly predictive of gene activity. Away from TSSs, modification sites are enriched in H3K4me1 and relatively depleted in H3K4me3 and H3ac. Comparison between cell lines identified differences in the histone modification profiles associated with transcriptional differences between the cell lines. These results provide an overview of the functional relationship among histone modifications and gene expression in human cells.},} - S Oeder, J Mages, P Flicek, R Lang. Uncovering information on expression of natural antisense transcripts in Affymetrix MOE430 datasets. BMC Genomics 2007;8:200. doi:10.1186/1471-2164-8-200
[BibTeX] [Abstract]
BACKGROUND: The function and significance of the widespread expression of natural antisense transcripts (NATs) is largely unknown. The ability to quantitatively assess changes in NAT expression for many different transcripts in multiple samples would facilitate our understanding of this relatively new class of RNA molecules. RESULTS: Here, we demonstrate that standard expression analysis Affymetrix MOE430 and HG-U133 GeneChips contain hundreds of probe sets that detect NATs. Probe sets carrying a ``Negative Strand Matching Probes'' annotation in NetAffx were validated using Ensembl by manual and automated approaches. More than 50 \% of the 1,113 probe sets with ``Negative Strand Matching Probes'' on the MOE430 2.0 GeneChip were confirmed as detecting NATs. Expression of selected antisense transcripts as indicated by Affymetrix data was confirmed using strand-specific RT-PCR. Thus, Affymetrix datasets can be mined to reveal information about the regulated expression of a considerable number of NATs. In a correlation analysis of 179 sense-antisense (SAS) probe set pairs using publicly available data from 1637 MOE430 2.0 GeneChips a significant number of SAS transcript pairs were found to be positively correlated. CONCLUSION: Standard expression analysis Affymetrix GeneChips can be used to measure many different NATs. The large amount of samples deposited in microarray databases represents a valuable resource for a quantitative analysis of NAT expression and regulation in different cells, tissues and biological conditions.
@Article{17598913, author = {Oeder S and Mages J and Flicek P and Lang R}, title = {Uncovering information on expression of natural antisense transcripts in Affymetrix MOE430 datasets}, journal = {BMC Genomics}, volume = {8}, pages = {200}, year = {2007}, doi = {10.1186/1471-2164-8-200}, abstract = {BACKGROUND: The function and significance of the widespread expression of natural antisense transcripts (NATs) is largely unknown. The ability to quantitatively assess changes in NAT expression for many different transcripts in multiple samples would facilitate our understanding of this relatively new class of RNA molecules. RESULTS: Here, we demonstrate that standard expression analysis Affymetrix MOE430 and HG-U133 GeneChips contain hundreds of probe sets that detect NATs. Probe sets carrying a ``Negative Strand Matching Probes'' annotation in NetAffx were validated using Ensembl by manual and automated approaches. More than 50 \% of the 1,113 probe sets with ``Negative Strand Matching Probes'' on the MOE430 2.0 GeneChip were confirmed as detecting NATs. Expression of selected antisense transcripts as indicated by Affymetrix data was confirmed using strand-specific RT-PCR. Thus, Affymetrix datasets can be mined to reveal information about the regulated expression of a considerable number of NATs. In a correlation analysis of 179 sense-antisense (SAS) probe set pairs using publicly available data from 1637 MOE430 2.0 GeneChips a significant number of SAS transcript pairs were found to be positively correlated. CONCLUSION: Standard expression analysis Affymetrix GeneChips can be used to measure many different NATs. The large amount of samples deposited in microarray databases represents a valuable resource for a quantitative analysis of NAT expression and regulation in different cells, tissues and biological conditions.},}
2006
- R Guigó, P Flicek, JF Abril, A Reymond, J Lagarde, F Denoeud, S Antonarakis, M Ashburner, VB Bajic, E Birney, R Castelo, E Eyras, C Ucla, TR Gingeras, J Harrow, T Hubbard, SE Lewis, MG Reese. EGASP: the human ENCODE Genome Annotation Assessment Project. Genome Biol 2006;7 Suppl 1:S2.1–31. doi:10.1186/gb-2006-7-s1-s2
[BibTeX] [Abstract]
BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1\% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a `reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70\% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40\% to 50\% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90\% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2\%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.
@Article{16925836, author = {Guigó R and Flicek P and Abril JF and Reymond A and Lagarde J and Denoeud F and Antonarakis S and Ashburner M and Bajic VB and Birney E and Castelo R and Eyras E and Ucla C and Gingeras TR and Harrow J and Hubbard T and Lewis SE and Reese MG}, title = {EGASP: the human ENCODE Genome Annotation Assessment Project}, journal = {Genome Biol}, volume = {7 Suppl 1}, pages = {S2.1--31}, year = {2006}, doi = {10.1186/gb-2006-7-s1-s2}, abstract = {BACKGROUND: We present the results of EGASP, a community experiment to assess the state-of-the-art in genome annotation within the ENCODE regions, which span 1\% of the human genome sequence. The experiment had two major goals: the assessment of the accuracy of computational methods to predict protein coding genes; and the overall assessment of the completeness of the current human genome annotations as represented in the ENCODE regions. For the computational prediction assessment, eighteen groups contributed gene predictions. We evaluated these submissions against each other based on a `reference set' of annotations generated as part of the GENCODE project. These annotations were not available to the prediction groups prior to the submission deadline, so that their predictions were blind and an external advisory committee could perform a fair assessment. RESULTS: The best methods had at least one gene transcript correctly predicted for close to 70\% of the annotated genes. Nevertheless, the multiple transcript accuracy, taking into account alternative splicing, reached only approximately 40\% to 50\% accuracy. At the coding nucleotide level, the best programs reached an accuracy of 90\% in both sensitivity and specificity. Programs relying on mRNA and protein sequences were the most accurate in reproducing the manually curated annotations. Experimental validation shows that only a very small percentage (3.2\%) of the selected 221 computationally predicted exons outside of the existing annotation could be verified. CONCLUSION: This is the first such experiment in human DNA, and we have followed the standards established in a similar experiment, GASP1, in Drosophila melanogaster. We believe the results presented here contribute to the value of ongoing large-scale annotation projects and should guide further experimental methods when being scaled up to the entire human genome sequence.},} - E Birney, D Andrews, M Caccamo, Y Chen, L Clarke, G Coates, T Cox, F Cunningham, V Curwen, T Cutts, T Down, R Durbin, XM Fernandez-Suarez, P Flicek, S Gräf, M Hammond, J Herrero, K Howe, V Iyer, K Jekosch, A Kähäri, A Kasprzyk, D Keefe, F Kokocinski, E Kulesha, D London, I Longden, C Melsopp, P Meidl, B Overduin, A Parker, G Proctor, A Prlic, M Rae, D Rios, S Redmond, M Schuster, I Sealy, S Searle, J Severin, G Slater, D Smedley, J Smith, A Stabenau, J Stalker, S Trevanion, A Ureta-Vidal, J Vogel, S White, C Woodwark, TJP Hubbard. Ensembl 2006. Nucleic Acids Res 2006;34(Database issue):D556–61. doi:10.1093/nar/gkj133
[BibTeX] [Abstract]
The Ensembl (http://www.ensembl.org/) project provides a comprehensive and integrated source of annotation of large genome sequences. Over the last year the number of genomes available from the Ensembl site has increased from 4 to 19, with the addition of the mammalian genomes of Rhesus macaque and Opossum, the chordate genome of Ciona intestinalis and the import and integration of the yeast genome. The year has also seen extensive improvements to both data analysis and presentation, with the introduction of a redesigned website, the addition of RNA gene and regulatory annotation and substantial improvements to the integration of human genome variation data.
@Article{16381931, author = {Birney E and Andrews D and Caccamo M and Chen Y and Clarke L and Coates G and Cox T and Cunningham F and Curwen V and Cutts T and Down T and Durbin R and Fernandez-Suarez XM and Flicek P and Gräf S and Hammond M and Herrero J and Howe K and Iyer V and Jekosch K and Kähäri A and Kasprzyk A and Keefe D and Kokocinski F and Kulesha E and London D and Longden I and Melsopp C and Meidl P and Overduin B and Parker A and Proctor G and Prlic A and Rae M and Rios D and Redmond S and Schuster M and Sealy I and Searle S and Severin J and Slater G and Smedley D and Smith J and Stabenau A and Stalker J and Trevanion S and Ureta-Vidal A and Vogel J and White S and Woodwark C and Hubbard TJP}, title = {Ensembl 2006}, journal = {Nucleic Acids Res}, volume = {34}, number = {Database issue}, pages = {D556--61}, year = {2006}, doi = {10.1093/nar/gkj133}, abstract = {The Ensembl (http://www.ensembl.org/) project provides a comprehensive and integrated source of annotation of large genome sequences. Over the last year the number of genomes available from the Ensembl site has increased from 4 to 19, with the addition of the mammalian genomes of Rhesus macaque and Opossum, the chordate genome of Ciona intestinalis and the import and integration of the yeast genome. The year has also seen extensive improvements to both data analysis and presentation, with the introduction of a redesigned website, the addition of RNA gene and regulatory annotation and substantial improvements to the integration of human genome variation data.},} - F Cunningham, D Rios, M Griffiths, J Smith, Z Ning, T Cox, P Flicek, P Marin-Garcin, J Herrero, J Rogers, van der L Weyden, A Bradley, E Birney, DJ Adams. TranscriptSNPView: a genome-wide catalog of mouse coding variation. Nat Genet 2006;38(8):853. doi:10.1038/ng0806-853a
[BibTeX]@Article{16874317, author = {Cunningham F and Rios D and Griffiths M and Smith J and Ning Z and Cox T and Flicek P and Marin-Garcin P and Herrero J and Rogers J and van der Weyden L and Bradley A and Birney E and Adams DJ}, title = {TranscriptSNPView: a genome-wide catalog of mouse coding variation}, journal = {Nat Genet}, volume = {38}, number = {8}, pages = {853}, year = {2006}, doi = {10.1038/ng0806-853a}, } - P Flicek, MR Brent. Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts. Genome Biol 2006;7 Suppl 1:S8.1–9. doi:10.1186/gb-2006-7-s1-s8
[BibTeX] [Abstract]
BACKGROUND: As part of the ENCODE Genome Annotation Assessment Project (EGASP), we developed the MARS extension to the Twinscan algorithm. MARS is designed to find human alternatively spliced transcripts that are conserved in only one or a limited number of extant species. MARS is able to use an arbitrary number of informant sequences and predicts a number of alternative transcripts at each gene locus. RESULTS: MARS uses the mouse, rat, dog, opossum, chicken, and frog genome sequences as pairwise informant sources for Twinscan and combines the resulting transcript predictions into genes based on coding (CDS) region overlap. Based on the EGASP assessment, MARS is one of the more accurate dual-genome prediction programs. Compared to the GENCODE annotation, we find that predictive sensitivity increases, while specificity decreases, as more informant species are used. MARS correctly predicts alternatively spliced transcripts for 11 of the 236 multi-exon GENCODE genes that are alternatively spliced in the coding region of their transcripts. For these genes a total of 24 correct transcripts are predicted. CONCLUSION: The MARS algorithm is able to predict alternatively spliced transcripts without the use of expressed sequence information, although the number of loci in which multiple predicted transcripts match multiple alternatively spliced transcripts in the GENCODE annotation is relatively small.
@Article{16925842, author = {Flicek P and Brent MR}, title = {Using several pair-wise informant sequences for de novo prediction of alternatively spliced transcripts}, journal = {Genome Biol}, volume = {7 Suppl 1}, pages = {S8.1--9}, year = {2006}, doi = {10.1186/gb-2006-7-s1-s8}, abstract = {BACKGROUND: As part of the ENCODE Genome Annotation Assessment Project (EGASP), we developed the MARS extension to the Twinscan algorithm. MARS is designed to find human alternatively spliced transcripts that are conserved in only one or a limited number of extant species. MARS is able to use an arbitrary number of informant sequences and predicts a number of alternative transcripts at each gene locus. RESULTS: MARS uses the mouse, rat, dog, opossum, chicken, and frog genome sequences as pairwise informant sources for Twinscan and combines the resulting transcript predictions into genes based on coding (CDS) region overlap. Based on the EGASP assessment, MARS is one of the more accurate dual-genome prediction programs. Compared to the GENCODE annotation, we find that predictive sensitivity increases, while specificity decreases, as more informant species are used. MARS correctly predicts alternatively spliced transcripts for 11 of the 236 multi-exon GENCODE genes that are alternatively spliced in the coding region of their transcripts. For these genes a total of 24 correct transcripts are predicted. CONCLUSION: The MARS algorithm is able to predict alternatively spliced transcripts without the use of expressed sequence information, although the number of loci in which multiple predicted transcripts match multiple alternatively spliced transcripts in the GENCODE annotation is relatively small.},}
2005
- E Eyras, A Reymond, R Castelo, JM Bye, F Camara, P Flicek, EJ Huckle, G Parra, DD Shteynberg, C Wyss, J Rogers, SE Antonarakis, E Birney, R Guigo, MR Brent. Gene finding in the chicken genome. BMC Bioinformatics 2005;6:131. doi:10.1186/1471-2105-6-131
[BibTeX] [Abstract]
BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40\% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.
@Article{15924626, author = {Eyras E and Reymond A and Castelo R and Bye JM and Camara F and Flicek P and Huckle EJ and Parra G and Shteynberg DD and Wyss C and Rogers J and Antonarakis SE and Birney E and Guigo R and Brent MR}, title = {Gene finding in the chicken genome}, journal = {BMC Bioinformatics}, volume = {6}, pages = {131}, year = {2005}, doi = {10.1186/1471-2105-6-131}, abstract = {BACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40\% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods.},}
2004
- Paul Flicek. Methods for improving gene prediction with evolutionary conservation [ PhD Thesis], 2004.
[BibTeX]@phdthesis{Flicek2004, author = {Flicek Paul}, title = {Methods for improving gene prediction with evolutionary conservation}, school = {Washington University}, address = {St. Louis}, year = {2004}, } - International Chicken Genome Sequencing Consortium.. Sequence and comparative analysis of the chicken genome provide unique perspectives on vertebrate evolution. Nature 2004;432(7018):695–716. doi:10.1038/nature03154
[BibTeX] [Abstract]
We present here a draft genome sequence of the red jungle fowl, Gallus gallus. Because the chicken is a modern descendant of the dinosaurs and the first non-mammalian amniote to have its genome sequenced, the draft sequence of its genome–composed of approximately one billion base pairs of sequence and an estimated 20,000-23,000 genes–provides a new perspective on vertebrate genome evolution, while also improving the annotation of mammalian genomes. For example, the evolutionary distance between chicken and human provides high specificity in detecting functional elements, both non-coding and coding. Notably, many conserved non-coding sequences are far from genes and cannot be assigned to defined functional classes. In coding regions the evolutionary dynamics of protein domains and orthologous groups illustrate processes that distinguish the lineages leading to birds and mammals. The distinctive properties of avian microchromosomes, together with the inferred patterns of conserved synteny, provide additional insights into vertebrate chromosome architecture.
@Article{15592404, author = {{International Chicken Genome Sequencing Consortium.}}, title = {Sequence and comparative analysis of the chicken genome provide unique perspectives on vertebrate evolution}, journal = {Nature}, volume = {432}, number = {7018}, pages = {695--716}, year = {2004}, doi = {10.1038/nature03154}, abstract = {We present here a draft genome sequence of the red jungle fowl, Gallus gallus. Because the chicken is a modern descendant of the dinosaurs and the first non-mammalian amniote to have its genome sequenced, the draft sequence of its genome--composed of approximately one billion base pairs of sequence and an estimated 20,000-23,000 genes--provides a new perspective on vertebrate genome evolution, while also improving the annotation of mammalian genomes. For example, the evolutionary distance between chicken and human provides high specificity in detecting functional elements, both non-coding and coding. Notably, many conserved non-coding sequences are far from genes and cannot be assigned to defined functional classes. In coding regions the evolutionary dynamics of protein domains and orthologous groups illustrate processes that distinguish the lineages leading to birds and mammals. The distinctive properties of avian microchromosomes, together with the inferred patterns of conserved synteny, provide additional insights into vertebrate chromosome architecture.},}
2003
- P Flicek, E Keibler, P Hu, I Korf, MR Brent. Leveraging the mouse genome for gene prediction in human: from whole-genome shotgun reads to a global synteny map. Genome Res 2003;13(1):46–54. doi:10.1101/gr.830003
[BibTeX] [Abstract]
The availability of draft sequences for both the mouse and human genomes makes it possible, for the first time, to annotate whole mammalian genomes using comparative methods. TWINSCAN is a gene-prediction system that combines the methods of single-genome predictors like GENSCAN with information derived from genome comparison, thereby improving accuracy. Because TWINSCAN uses genomic sequence only, it is less biased toward highly and/or ubiquitously expressed genes than GENEWISE, GENOMESCAN, and other methods based on evidence derived from transcripts. We show that TWINSCAN improves gene prediction in human using intermediate products from various stages of the sequencing and analysis of the mouse genome, from low-redundancy, whole-genome shotgun reads to the draft assembly and the synteny map. TWINSCAN improves on the prior state of the art even when alignments from only 1X coverage of the mouse genome are available. Gene prediction accuracy improves steadily from 1X through 3X, more slowly from 3X to 4X, and relatively little thereafter. The assembly and the synteny map greatly speed the computations, however. Our human annotation using the mouse assembly is conservative, predicting only 25,622 genes, and appears to be one of the best de novo annotations of the human genome to date.
@Article{12529305, author = {Flicek P and Keibler E and Hu P and Korf I and Brent MR}, title = {Leveraging the mouse genome for gene prediction in human: from whole-genome shotgun reads to a global synteny map}, journal = {Genome Res}, volume = {13}, number = {1}, pages = {46--54}, year = {2003}, doi = {10.1101/gr.830003}, abstract = {The availability of draft sequences for both the mouse and human genomes makes it possible, for the first time, to annotate whole mammalian genomes using comparative methods. TWINSCAN is a gene-prediction system that combines the methods of single-genome predictors like GENSCAN with information derived from genome comparison, thereby improving accuracy. Because TWINSCAN uses genomic sequence only, it is less biased toward highly and/or ubiquitously expressed genes than GENEWISE, GENOMESCAN, and other methods based on evidence derived from transcripts. We show that TWINSCAN improves gene prediction in human using intermediate products from various stages of the sequencing and analysis of the mouse genome, from low-redundancy, whole-genome shotgun reads to the draft assembly and the synteny map. TWINSCAN improves on the prior state of the art even when alignments from only 1X coverage of the mouse genome are available. Gene prediction accuracy improves steadily from 1X through 3X, more slowly from 3X to 4X, and relatively little thereafter. The assembly and the synteny map greatly speed the computations, however. Our human annotation using the mouse assembly is conservative, predicting only 25,622 genes, and appears to be one of the best de novo annotations of the human genome to date.},} - LW Hillier, RS Fulton, LA Fulton, TA Graves, KH Pepin, C Wagner-McPherson, D Layman, J Maas, S Jaeger, R Walker, K Wylie, M Sekhon, MC Becker, MD O'Laughlin, ME Schaller, GA Fewell, KD Delehaunty, TL Miner, WE Nash, M Cordes, H Du, H Sun, J Edwards, H Bradshaw-Cordum, J Ali, S Andrews, A Isak, A Vanbrunt, C Nguyen, F Du, B Lamar, L Courtney, J Kalicki, P Ozersky, L Bielicki, K Scott, A Holmes, R Harkins, A Harris, CM Strong, S Hou, C Tomlinson, S Dauphin-Kohlberg, A Kozlowicz-Reilly, S Leonard, T Rohlfing, SM Rock, A-M Tin-Wollam, A Abbott, P Minx, R Maupin, C Strowmatt, P Latreille, N Miller, D Johnson, J Murray, JP Woessner, MC Wendl, S-P Yang, BR Schultz, JW Wallis, J Spieth, TA Bieri, JO Nelson, N Berkowicz, PE Wohldmann, LL Cook, MT Hickenbotham, J Eldred, D Williams, JA Bedell, ER Mardis, SW Clifton, SL Chissoe, MA Marra, C Raymond, E Haugen, W Gillett, Y Zhou, R James, K Phelps, S Iadanoto, K Bubb, E Simms, R Levy, J Clendenning, R Kaul, WJ Kent, TS Furey, RA Baertsch, MR Brent, E Keibler, P Flicek, P Bork, M Suyama, JA Bailey, ME Portnoy, D Torrents, AT Chinwalla, WR Gish, SR Eddy, JD McPherson, MV Olson, EE Eichler, ED Green, RH Waterston, RK Wilson. The DNA sequence of human chromosome 7. Nature 2003;424(6945):157–164. doi:10.1038/nature01782
[BibTeX] [Abstract]
Human chromosome 7 has historically received prominent attention in the human genetics community, primarily related to the search for the cystic fibrosis gene and the frequent cytogenetic changes associated with various forms of cancer. Here we present more than 153 million base pairs representing 99.4\% of the euchromatic sequence of chromosome 7, the first metacentric chromosome completed so far. The sequence has excellent concordance with previously established physical and genetic maps, and it exhibits an unusual amount of segmentally duplicated sequence (8.2\%), with marked differences between the two arms. Our initial analyses have identified 1,150 protein-coding genes, 605 of which have been confirmed by complementary DNA sequences, and an additional 941 pseudogenes. Of genes confirmed by transcript sequences, some are polymorphic for mutations that disrupt the reading frame.
@Article{12853948, author = {Hillier LW and Fulton RS and Fulton LA and Graves TA and Pepin KH and Wagner-McPherson C and Layman D and Maas J and Jaeger S and Walker R and Wylie K and Sekhon M and Becker MC and O'Laughlin MD and Schaller ME and Fewell GA and Delehaunty KD and Miner TL and Nash WE and Cordes M and Du H and Sun H and Edwards J and Bradshaw-Cordum H and Ali J and Andrews S and Isak A and Vanbrunt A and Nguyen C and Du F and Lamar B and Courtney L and Kalicki J and Ozersky P and Bielicki L and Scott K and Holmes A and Harkins R and Harris A and Strong CM and Hou S and Tomlinson C and Dauphin-Kohlberg S and Kozlowicz-Reilly A and Leonard S and Rohlfing T and Rock SM and Tin-Wollam A-M and Abbott A and Minx P and Maupin R and Strowmatt C and Latreille P and Miller N and Johnson D and Murray J and Woessner JP and Wendl MC and Yang S-P and Schultz BR and Wallis JW and Spieth J and Bieri TA and Nelson JO and Berkowicz N and Wohldmann PE and Cook LL and Hickenbotham MT and Eldred J and Williams D and Bedell JA and Mardis ER and Clifton SW and Chissoe SL and Marra MA and Raymond C and Haugen E and Gillett W and Zhou Y and James R and Phelps K and Iadanoto S and Bubb K and Simms E and Levy R and Clendenning J and Kaul R and Kent WJ and Furey TS and Baertsch RA and Brent MR and Keibler E and Flicek P and Bork P and Suyama M and Bailey JA and Portnoy ME and Torrents D and Chinwalla AT and Gish WR and Eddy SR and McPherson JD and Olson MV and Eichler EE and Green ED and Waterston RH and Wilson RK}, title = {The DNA sequence of human chromosome 7}, journal = {Nature}, volume = {424}, number = {6945}, pages = {157--164}, year = {2003}, doi = {10.1038/nature01782}, abstract = {Human chromosome 7 has historically received prominent attention in the human genetics community, primarily related to the search for the cystic fibrosis gene and the frequent cytogenetic changes associated with various forms of cancer. Here we present more than 153 million base pairs representing 99.4\% of the euchromatic sequence of chromosome 7, the first metacentric chromosome completed so far. The sequence has excellent concordance with previously established physical and genetic maps, and it exhibits an unusual amount of segmentally duplicated sequence (8.2\%), with marked differences between the two arms. Our initial analyses have identified 1,150 protein-coding genes, 605 of which have been confirmed by complementary DNA sequences, and an additional 941 pseudogenes. Of genes confirmed by transcript sequences, some are polymorphic for mutations that disrupt the reading frame.},}
2002
- Mouse Genome Sequencing Consortium.. Initial sequencing and comparative analysis of the mouse genome. Nature 2002;420(6915):520–562. doi:10.1038/nature01262
[BibTeX] [Abstract]
The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.
@Article{12466850, author = {{Mouse Genome Sequencing Consortium.}}, title = {Initial sequencing and comparative analysis of the mouse genome}, journal = {Nature}, volume = {420}, number = {6915}, pages = {520--562}, year = {2002}, doi = {10.1038/nature01262}, abstract = {The sequence of the mouse genome is a key informational tool for understanding the contents of the human genome and a key experimental tool for biomedical research. Here, we report the results of an international collaboration to produce a high-quality draft sequence of the mouse genome. We also present an initial comparative analysis of the mouse and human genomes, describing some of the insights that can be gleaned from the two sequences. We discuss topics including the analysis of the evolutionary forces shaping the size, structure and sequence of the genomes; the conservation of large-scale synteny across most of the genomes; the much lower extent of sequence orthology covering less than half of the genomes; the proportions of the genomes under selection; the number of protein-coding genes; the expansion of gene families related to reproduction and immunity; the evolution of proteins; and the identification of intraspecies polymorphism.},} - Paul Flicek. Twinscan: A Software Package for Homology-based Gene Prediction [ PhD Thesis], 2002.
[BibTeX]@phdthesis{Flicek2002, author = {Flicek Paul}, title = {Twinscan: A Software Package for Homology-based Gene Prediction}, school = {Washington University}, address = {St. Louis}, year = {2002}, }
2001
- I Korf, P Flicek, D Duan, MR Brent. Integrating genomic homology into gene structure prediction. Bioinformatics 2001;17 Suppl 1:S140–8. doi:10.1093/bioinformatics/17.suppl_1.s140
[BibTeX] [Abstract]
TWINSCAN is a new gene-structure prediction system that directly extends the probability model of GENSCAN, allowing it to exploit homology between two related genomes. Separate probability models are used for conservation in exons, introns, splice sites, and UTRs, reflecting the differences among their patterns of evolutionary conservation. TWINSCAN is specifically designed for the analysis of high-throughput genomic sequences containing an unknown number of genes. In experiments on high-throughput mouse sequences, using homologous sequences from the human genome, TWINSCAN shows notable improvement over GENSCAN in exon sensitivity and specificity and dramatic improvement in exact gene sensitivity and specificity. This improvement can be attributed entirely to modeling the patterns of evolutionary conservation in genomic sequence.
@Article{11473003, author = {Korf I and Flicek P and Duan D and Brent MR}, title = {Integrating genomic homology into gene structure prediction}, journal = {Bioinformatics}, volume = {17 Suppl 1}, pages = {S140--8}, year = {2001}, doi = {10.1093/bioinformatics/17.suppl_1.s140}, abstract = {TWINSCAN is a new gene-structure prediction system that directly extends the probability model of GENSCAN, allowing it to exploit homology between two related genomes. Separate probability models are used for conservation in exons, introns, splice sites, and UTRs, reflecting the differences among their patterns of evolutionary conservation. TWINSCAN is specifically designed for the analysis of high-throughput genomic sequences containing an unknown number of genes. In experiments on high-throughput mouse sequences, using homologous sequences from the human genome, TWINSCAN shows notable improvement over GENSCAN in exon sensitivity and specificity and dramatic improvement in exact gene sensitivity and specificity. This improvement can be attributed entirely to modeling the patterns of evolutionary conservation in genomic sequence.},}
1999
- TJ Miner, CD Shriver, PR Flicek, FC Miner, DP Jaques, ME Maniscalco-Theberge, DN Krag. Guidelines for the safe use of radioactive materials during localization and resection of the sentinel lymph node. Ann Surg Oncol 1999;6(1):75–82. doi:10.1007/s10434-999-0075-7
[BibTeX] [Abstract]
BACKGROUND: Several reports have demonstrated accurate prediction of nodal metastasis with radiolocalization and selective resection of the radiolocalized sentinel lymph node (SLN) in patients with breast cancer and melanoma. As reliance on this technique grows, its use by those without experience in radiation safety will increase. METHODS: Tissue obtained during radioguided SLN biopsies was examined for residual radioactivity. Specimens with a specific activity greater than the radiologic control level (RCL) of 0.002 microCi/g were considered radioactive. Radiation exposure to the surgical team was measured. RESULTS: A total of 24 primary tissue specimens and 318 lymph nodes were obtained during 57 operations (37 for breast cancer, 20 for melanoma). All 24 (100\%) of the specimens injected with radiopharmaceutical and 89 of 98 (91\%) of the localized nodes were radioactive after surgery. Activity fell below the RCL 71+/-3.6 hours in primary tissue specimens, 46+/-1.7 hours in nodes from melanoma patients, and 33+/-3.5 hours in nodes from breast cancer patients (P = .037). The hands of the surgical team (n = 22 cases) were exposed to 9.4+/-3.6 mrem/case. CONCLUSION: Although low levels of radiation exposure are associated with radiolocalization and resection of the SLN, the presented guidelines ensure conformity to existing regulations and allow timely pathologic analysis.
@Article{10030418, author = {Miner TJ and Shriver CD and Flicek PR and Miner FC and Jaques DP and Maniscalco-Theberge ME and Krag DN}, title = {Guidelines for the safe use of radioactive materials during localization and resection of the sentinel lymph node}, journal = {Ann Surg Oncol}, volume = {6}, number = {1}, pages = {75--82}, year = {1999}, doi = {10.1007/s10434-999-0075-7}, abstract = {BACKGROUND: Several reports have demonstrated accurate prediction of nodal metastasis with radiolocalization and selective resection of the radiolocalized sentinel lymph node (SLN) in patients with breast cancer and melanoma. As reliance on this technique grows, its use by those without experience in radiation safety will increase. METHODS: Tissue obtained during radioguided SLN biopsies was examined for residual radioactivity. Specimens with a specific activity greater than the radiologic control level (RCL) of 0.002 microCi/g were considered radioactive. Radiation exposure to the surgical team was measured. RESULTS: A total of 24 primary tissue specimens and 318 lymph nodes were obtained during 57 operations (37 for breast cancer, 20 for melanoma). All 24 (100\%) of the specimens injected with radiopharmaceutical and 89 of 98 (91\%) of the localized nodes were radioactive after surgery. Activity fell below the RCL 71+/-3.6 hours in primary tissue specimens, 46+/-1.7 hours in nodes from melanoma patients, and 33+/-3.5 hours in nodes from breast cancer patients (P = .037). The hands of the surgical team (n = 22 cases) were exposed to 9.4+/-3.6 mrem/case. CONCLUSION: Although low levels of radiation exposure are associated with radiolocalization and resection of the SLN, the presented guidelines ensure conformity to existing regulations and allow timely pathologic analysis.},}
1993
- BJ Albright, K Bartschat, PR Flicek. Core potentials for quasi-one electron systems. J Phys B: At Mol Opt Phys 1993;26:337–344. doi:10.1088/0953-4075/26/3/008
[BibTeX] [Abstract]
The authors have developed a method to calculate core potentials for quasi-one-electron systems and the corresponding single electron orbitals. It is shown that the approximate inclusion of exchange effects between the valence electrons and the core removes the unphysical structure in the potential function that is characteristic for potentials calculated by only including the effect of core polarization due to the valence electrons. Excellent agreement with experimental ionization potentials is achieved, and example results for various systems such as Na, Cs, Ba+, In and Tl are presented.
@Article{, author = {Albright BJ and Bartschat K and Flicek PR}, title = {Core potentials for quasi-one electron systems}, journal = {J Phys B: At Mol Opt Phys}, volume = {26}, pages = {337--344}, year = {1993}, doi = {10.1088/0953-4075/26/3/008}, abstract = {The authors have developed a method to calculate core potentials for quasi-one-electron systems and the corresponding single electron orbitals. It is shown that the approximate inclusion of exchange effects between the valence electrons and the core removes the unphysical structure in the potential function that is characteristic for potentials calculated by only including the effect of core polarization due to the valence electrons. Excellent agreement with experimental ionization potentials is achieved, and example results for various systems such as Na, Cs, Ba+, In and Tl are presented.},}