BEDTools: a flexible suite of utilities for comparing genomic features
Aaron R. Quinlan,Ira M. Hall +1 more
TLDR
A new software suite for the comparison, manipulation and annotation of genomic features in Browser Extensible Data (BED) and General Feature Format (GFF) format, which allows the user to compare large datasets (e.g. next-generation sequencing data) with both public and custom genome annotation tracks.Abstract:
Motivation: Testing for correlations between different sets of genomic features is a fundamental task in genomics research. However, searching for overlaps between features with existing webbased methods is complicated by the massive datasets that are routinely produced with current sequencing technologies. Fast and flexible tools are therefore required to ask complex questions of these data in an efficient manner. Results: This article introduces a new software suite for the comparison, manipulation and annotation of genomic features in Browser Extensible Data (BED) and General Feature Format (GFF) format. BEDTools also supports the comparison of sequence alignments in BAM format to both BED and GFF features. The tools are extremely efficient and allow the user to compare large datasets (e.g. next-generation sequencing data) with both public and custom genome annotation tracks. BEDTools can be combined with one another as well as with standard UNIX commands, thus facilitating routine genomics tasks as well as pipelines that can quickly answer intricate questions of large genomic datasets. Availability and implementation: BEDTools was written in C++. Source code and a comprehensive user manual are freely available at http://code.google.com/p/bedtoolsread more
Citations
More filters
Journal ArticleDOI
HTSeq—a Python framework to work with high-throughput sequencing data
TL;DR: This work presents HTSeq, a Python library to facilitate the rapid development of custom scripts for high-throughput sequencing data analysis, and presents htseq-count, a tool developed with HTSequ that preprocesses RNA-Seq data for differential expression analysis by counting the overlap of reads with genes.
Journal ArticleDOI
featureCounts: an efficient general-purpose program for assigning sequence reads to genomic features
TL;DR: FeatureCounts as discussed by the authors is a read summarization program suitable for counting reads generated from either RNA or genomic DNA sequencing experiments, which implements highly efficient chromosome hashing and feature blocking techniques.
Journal ArticleDOI
A global reference for human genetic variation.
Adam Auton,Gonçalo R. Abecasis,David Altshuler,Richard Durbin,David R. Bentley,Aravinda Chakravarti,Andrew G. Clark,Peter Donnelly,Evan E. Eichler,Paul Flicek,Stacey Gabriel,Richard A. Gibbs,Eric D. Green,Matthew E. Hurles,Bartha Maria Knoppers,Jan O. Korbel,Eric S. Lander,Charles Lee,Hans Lehrach,Elaine R. Mardis,Gabor T. Marth,Gil McVean,Deborah A. Nickerson,Jeanette Schmidt,Stephen T. Sherry,Jun Wang,Richard K. Wilson,Eric Boerwinkle,Harsha Doddapaneni,Yi Han,Viktoriya Korchina,Christie Kovar,Sandra L. Lee,Donna M. Muzny,Jeffrey G. Reid,Yiming Zhu,Yuqi Chang,Qiang Feng,Qiang Feng,Xiaodong Fang,Xiaodong Fang,Xiaosen Guo,Xiaosen Guo,Min Jian,Min Jian,Hui Jiang,Hui Jiang,Xin Jin,Tianming Lan,Guoqing Li,Jingxiang Li,Yingrui Li,Shengmao Liu,Xiao Liu,Xiao Liu,Yao Lu,Xuedi Ma,Meifang Tang,Bo Wang,Guangbiao Wang,Honglong Wu,Renhua Wu,Xun Xu,Ye Yin,Dandan Zhang,Wenwei Zhang,Jiao Zhao,Meiru Zhao,Xiaole Zheng,Namrata Gupta,Neda Gharani,Lorraine Toji,Norman P. Gerry,Alissa M. Resch,Jonathan Barker,Laura Clarke,Laurent Gil,Sarah E. Hunt,Gavin Kelman,Eugene Kulesha,Rasko Leinonen,William M. McLaren,Rajesh Radhakrishnan,Asier Roa,Dmitriy Smirnov,Richard Smith,Ian Streeter,Anja Thormann,Iliana Toneva,Brendan Vaughan,Xiangqun Zheng-Bradley,Russell J. Grocock,Sean Humphray,Terena James,Zoya Kingsbury,Ralf Sudbrak,M. Albrecht,Vyacheslav Amstislavskiy,Tatiana A. Borodina,Matthias Lienhard,Florian Mertes,Marc Sultan,Bernd Timmermann,Marie-Laure Yaspo,Lucinda Fulton,Victor Ananiev,Zinaida Belaia,Dimitriy Beloslyudtsev,Nathan Bouk,Chao Chen,Deanna M. Church,Robert M. Cohen,Charles Cook,John Garner,Timothy Hefferon,Mikhail Kimelman,Chunlei Liu,John Lopez,Peter Meric,Chris O’Sullivan,Yuri Ostapchuk,Lon Phan,Sergiy Ponomarov,Valerie A. Schneider,Eugene Shekhtman,Karl Sirotkin,Douglas J. Slotta,Hua Zhang,Senduran Balasubramaniam,John Burton,Petr Danecek,Thomas M. Keane,Anja Kolb-Kokocinski,Shane A. McCarthy,James Stalker,Michael A. Quail,Christopher Davies,Jeremy Gollub,Teresa Webster,Brant Wong,Yiping Zhan,Christopher L. Campbell,Yu Kong,Anthony Marcketta,Fuli Yu,Lilian Antunes,Matthew N. Bainbridge,Aniko Sabo,Zhuoyi Huang,Lachlan J. M. Coin,Lin Fang,Lin Fang,Qibin Li,Zhenyu Li,Haoxiang Lin,Binghang Liu,Ruibang Luo,Haojing Shao,Haojing Shao,Yinlong Xie,Chen Ye,Chang Yu,Fan Zhang,Hancheng Zheng,Zhu Hongmei,Can Alkan,Elif Dal,Fatma Kahveci,Erik Garrison,Deniz Kural,Wan-Ping Lee,Wen Fung Leong,Michael Strömberg,Alistair Ward,Jiantao Wu,Mengyao Zhang,Mark J. Daly,Mark A. DePristo,Robert E. Handsaker,Robert E. Handsaker,Eric Banks,Gaurav Bhatia,Guillermo del Angel,Giulio Genovese,Heng Li,Seva Kashin,Seva Kashin,Steven A. McCarroll,Steven A. McCarroll,James Nemesh,Ryan Poplin,Seungtai Yoon,Jayon Lihm,Vladimir Makarov,Srikanth Gottipati,Alon Keinan,Juan L. Rodriguez-Flores,Tobias Rausch,Markus Hsi-Yang Fritz,Adrian M. Stütz,Kathryn Beal,Avik Datta,Javier Herrero,Graham R. S. Ritchie,Daniel R. Zerbino,Pardis C. Sabeti,Pardis C. Sabeti,Ilya Shlyakhter,Ilya Shlyakhter,Stephen F. Schaffner,Stephen F. Schaffner,Joseph J. Vitti,Joseph J. Vitti,David Neil Cooper,Edward V. Ball,Peter D. Stenson,Bret Barnes,Markus J. Bauer,R. Keira Cheetham,Anthony J. Cox,Michael A. Eberle,Scott Kahn,Lisa Murray,John F. Peden,Richard Shaw,Eimear E. Kenny,Mark A. Batzer,Miriam K. Konkel,Jerilyn A. Walker,Daniel G. MacArthur,Monkol Lek,Ralf Herwig,Li Ding,Daniel C. Koboldt,David E. Larson,Kai Ye,Simon Gravel,Anand Swaroop,Emily Y. Chew,Tuuli Lappalainen,Yaniv Erlich,Melissa Gymrek,Melissa Gymrek,Thomas Willems,Jared T. Simpson,Mark D. Shriver,Jeffrey A. Rosenfeld,Carlos Bustamante,Stephen B. Montgomery,Francisco M. De La Vega,Jake K. Byrnes,Andrew Carroll,Marianne K. DeGorter,Phil Lacroute,Brian K. Maples,Alicia R. Martin,Andrés Moreno-Estrada,Andrés Moreno-Estrada,Suyash Shringarpure,Fouad Zakharia,Eran Halperin,Eran Halperin,Yael Baran,Eliza Cerveira,Jaeho Hwang,Ankit Malhotra,Dariusz Plewczynski,Kamen Radew,Mallory Romanovitch,Chengsheng Zhang,Fiona Hyland,David Craig,Alexis Christoforides,Nils Homer,Tyler Izatt,Ahmet Kurdoglu,Shripad Sinari,Kevin Squire,Chunlin Xiao,Jonathan Sebat,Danny Antaki,Madhusudan Gujral,Amina Noor,Kenny Ye,Esteban G. Burchard,Ryan D. Hernandez,Christopher R. Gignoux,David Haussler,David Haussler,Sol Katzman,W. James Kent,Bryan Howie,Andres Ruiz-Linares,Emmanouil T. Dermitzakis,Emmanouil T. Dermitzakis,Scott E. Devine,Hyun Min Kang,Jeffrey M. Kidd,Thomas W. Blackwell,Sean Caron,Wei Chen,S. Emery,Lars G. Fritsche,Christian Fuchsberger,Goo Jun,Goo Jun,Bingshan Li,Robert H. Lyons,Chris Scheller,Carlo Sidore,Carlo Sidore,Carlo Sidore,Shiya Song,Elzbieta Sliwerska,Daniel Taliun,Adrian Tan,Ryan P. Welch,Mary Kate Wing,Xiaowei Zhan,Philip Awadalla,Philip Awadalla,Alan Hodgkinson,Yun Li,Xinghua Shi,Andrew Quitadamo,Gerton Lunter,Jonathan Marchini,Simon Myers,Claire Churchhouse,Olivier Delaneau,Olivier Delaneau,Anjali Gupta-Hinch,Warren W. Kretzschmar,Zamin Iqbal,Iain Mathieson,Androniki Menelaou,Androniki Menelaou,Andy Rimmer,Dionysia Kiara Xifara,Taras K. Oleksyk,Yunxin Fu,Xiaoming Liu,Momiao Xiong,Lynn B. Jorde,David J. Witherspoon,Jinchuan Xing,Brian L. Browning,Sharon R. Browning,Fereydoun Hormozdiari,Peter H. Sudmant,Ekta Khurana,Chris Tyler-Smith,Cornelis A. Albers,Qasim Ayub,Yuan Chen,Vincenza Colonna,Vincenza Colonna,Luke Jostins,Klaudia Walter,Yali Xue,Mark Gerstein,Alexej Abyzov,Suganthi Balasubramanian,Jieming Chen,Declan Clarke,Yao Fu,Arif Harmanci,Mike Jin,Dong-Hoon Lee,Jeremy Liu,Xinmeng Jasmine Mu,Xinmeng Jasmine Mu,Jing Zhang,Yan Zhang,Christopher Hartl,Khalid Shakir,Jeremiah D. Degenhardt,Sascha Meiers,Benjamin Raeder,Francesco Paolo Casale,Oliver Stegle,Eric-Wubbo Lameijer,Ira M. Hall,Vineet Bafna,Jacob J. Michaelson,Eugene J. Gardner,Ryan E. Mills,Gargi Dayama,Ken Chen,Xian Fan,Zechen Chong,Tenghui Chen,Mark Chaisson,John Huddleston,Maika Malig,Bradley J. Nelson,Nicholas F. Parrish,Ben Blackburne,Sarah J. Lindsay,Zemin Ning,Yujun Zhang,Hugo Y. K. Lam,Cristina Sisu,Danny Challis,Uday S. Evani,James T. Lu,Uma Nagaswamy,Jin Yu,Wangshen Li,Lukas Habegger,Haiyuan Yu,Fiona Cunningham,Ian Dunham,Kasper Lage,Kasper Lage,Jakob Berg Jespersen,Jakob Berg Jespersen,Jakob Berg Jespersen,Heiko Horn,Heiko Horn,Donghoon Kim,Rob DeSalle,Apurva Narechania,Melissa A. Wilson Sayres,Fernando L. Mendez,G. David Poznik,Peter A. Underhill,David Mittelman,Ruby Banerjee,Maria Cerezo,Thomas W. Fitzgerald,Sandra Louzada,Andrea Massaia,Fengtang Yang,Divya Kalra,Walker Hale,Xu Dan,Kathleen C. Barnes,Christine Beiswanger,Hongyu Cai,Hongzhi Cao,Hongzhi Cao,Brenna M. Henn,Danielle Jones,Jane Kaye,Alastair Kent,Angeliki Kerasidou,Rasika A. Mathias,Pilar N. Ossorio,Michael Parker,Charles N. Rotimi,Charmaine D.M. Royal,Karla Sandoval,Yeyang Su,Zhongming Tian,Sarah A. Tishkoff,Marc Via,Yuhong Wang,Huanming Yang,Ling Yang,Jiayong Zhu,Walter F. Bodmer,Gabriel Bedoya,Zhiming Cai,Yang Gao,Jiayou Chu,Leena Peltonen,Andrés C. García-Montero,Alberto Orfao,Julie Dutil,Juan Carlos Martínez-Cruzado,R. Mathias,Anselm Hennis,Harold Watson,Colin A. McKenzie,Firdausi Qadri,Regina C. LaRocque,Xiaoyan Deng,Danny Asogun,Onikepe A. Folarin,Christian T. Happi,Omonwunmi Omoniwa,Matt Stremlau,Matt Stremlau,Ridhi Tariyal,Ridhi Tariyal,M Jallow,M Jallow,Fatoumatta Sisay Joof,Fatoumatta Sisay Joof,Tumani Corrah,Tumani Corrah,Kirk A. Rockett,Kirk A. Rockett,Dominic P. Kwiatkowski,Dominic P. Kwiatkowski,Jaspal S. Kooner,Tran Tinh Hien,Sarah J. Dunstan,Sarah J. Dunstan,Nguyen ThuyHang,Richard Fonnie,Robert F. Garry,Lansana Kanneh,Lina M. Moses,John S. Schieffelin,Donald S. Grant,Carla Gallo,Giovanni Poletti,Danish Saleheen,Asif Rasheed,Lisa D. Brooks,Adam Felsenfeld,Jean E. McEwen,Yekaterina Vaydylevich,Audrey Duncanson,Michael Dunn,Jeffery A. Schloss +517 more
TL;DR: The 1000 Genomes Project set out to provide a comprehensive description of common human genetic variation by applying whole-genome sequencing to a diverse set of individuals from multiple populations, and has reconstructed the genomes of 2,504 individuals from 26 populations using a combination of low-coverage whole-generation sequencing, deep exome sequencing, and dense microarray genotyping.
Journal ArticleDOI
RNA-Guided Human Genome Engineering via Cas9
Prashant Mali,Luhan Yang,Kevin M. Esvelt,John Aach,Marc Güell,James E. DiCarlo,Julie E. Norville,George M. Church,George M. Church +8 more
TL;DR: The type II bacterial CRISPR system is engineer to function with custom guide RNA (gRNA) in human cells to establish an RNA-guided editing tool for facile, robust, and multiplexable human genome engineering.
Journal ArticleDOI
Comprehensive Integration of Single-Cell Data.
Tim Stuart,Andrew Butler,Paul J. Hoffman,Christoph Hafemeister,Efthymia Papalexi,William M. Mauck,Yuhan Hao,Marlon Stoeckius,Peter Smibert,Rahul Satija +9 more
TL;DR: A strategy to "anchor" diverse datasets together, enabling us to integrate single-cell measurements not only across scRNA-seq technologies, but also across different modalities.
References
More filters
Journal ArticleDOI
The Sequence Alignment/Map format and SAMtools
Heng Li,Bob Handsaker,Alec Wysoker,T. J. Fennell,Jue Ruan,Nils Homer,Gabor T. Marth,Gonçalo R. Abecasis,Richard Durbin +8 more
TL;DR: SAMtools as discussed by the authors implements various utilities for post-processing alignments in the SAM format, such as indexing, variant caller and alignment viewer, and thus provides universal tools for processing read alignments.
Journal ArticleDOI
Improved tools for biological sequence comparison.
TL;DR: Three computer programs for comparisons of protein and DNA sequences can be used to search sequence data bases, evaluate similarity scores, and identify periodic structures based on local sequence similarity.
Journal ArticleDOI
The Human Genome Browser at UCSC
W. James Kent,Charles W. Sugnet,Terrence S. Furey,Krishna M. Roskin,Tom H. Pringle,Alan M. Zahler,and David Haussler +6 more
TL;DR: A mature web tool for rapid and reliable display of any requested portion of the genome at any scale, together with several dozen aligned annotation tracks, is provided at http://genome.ucsc.edu.
Journal ArticleDOI
Galaxy: A platform for interactive large-scale genome analysis
Belinda Giardine,Cathy Riemer,Ross C. Hardison,Richard Burhans,Laura Elnitski,Prachi Shah,Prachi Shah,Yi Zhang,Daniel Blankenberg,Istvan Albert,James Taylor,Webb Miller,W. James Kent,Anton Nekrutenko +13 more
TL;DR: An interactive system, Galaxy, that combines the power of existing genome annotation databases with a simple Web portal to enable users to search remote resources, combine data from independent queries, and visualize the results.