http://techcrunch.com/2014/08/15/one-codex-wants-to-be-the-google-for-genomic-data/
As hospitals and public health organizations switch to using genomic data for testing, searching through genomic data can still take some time. Y Combinator-backed startup, One Codex, wants to help researchers, clinicians and public health officials, who have sequenced more than 100,000 genomes and created petabytes of data, to search this data.
Just a collection of some random cool stuff. PS. Almost 99% of the contents here are not mine and I don't take credit for them, I reference and copy part of the interesting sections.
Showing posts with label genomics. Show all posts
Showing posts with label genomics. Show all posts
Monday, August 18, 2014
Monday, June 30, 2014
Data Analysis for Genomics
http://genomicsclass.github.io/book/
HarvardX: PH525x Data Analysis for Genomics
Data Analysis for Genomics
The repository of the R markdown files (.Rmd) for the labs shown here is:
Resources
Introduction (week 1)
- Introduction
- Exploratory Data Analysis
- Installing Bioconductor and finding help
- R refresher
- Robust summaries
Microarray and NGS basics (week 2)
- Installing packages from Github
- Reading microarray data
- Downloading data from GEO using GEOquery
- EDA plots for microarray
- Basic Bioconductor infrastructure
- EDA plots for next generation sequencing
Statistical inference and linear modeling (week 3)
- Inference
- Expressing design formula in R
- Linear models
- Basic inference for microarray
- Rank tests
- Monte Carlo methods
Background, modeling and normalization (week 4)
Distance and prediction (week 5)
- Distance lecture
- Distance and clustering lab
- Dimension reduction and heatmaps
- Prediction lecture
- Cross-validation
Batch effect (week 6)
Advanced differential expression (week 7)
- Hierarchical modeling and using limma
- Mapping features to genes
- Gene set analysis lecture
- Gene set testing in R
- Multiple testing
Advanced workflows (week 8)
- Visualizing NGS data
- Counting NGS reads in features
- Methylation
- Reading 450K idat files with the minfi package
- Interactive visualization of DNA methylation data analysis
- ChIP-seq
- RNA-seq
- Genome variation
Footnotes for all lectures
Friday, May 30, 2014
Tissue specific gene expression
http://www.cureffi.org/2013/07/11/tissue-specific-gene-expression-data-based-on-human-bodymap-2-0/
GTex
Allen Human Brain
Illumina Human Body Map 2.0
GTex
Allen Human Brain
Illumina Human Body Map 2.0
Wednesday, February 27, 2013
On the immortality of television sets: “function” in the human genome according to the evolution-free gospel of ENCODE
http://gbe.oxfordjournals.org/content/early/2013/02/20/gbe.evt028
A recent slew of ENCODE Consortium publications, specifically the article signed by all Consortium members, put forward the idea that more than 80% of the human genome is functional. This claim flies in the face of current estimates according to which the fraction of the genome that is evolutionarily conserved through purifying selection is under 10%. Thus, according to the ENCODE Consortium, a biological function can be maintained indefinitely without selection, which implies that at least 80 − 10 = 70% of the genome is perfectly invulnerable to deleterious mutations, either because no mutation can ever occur in these “functional” regions, or because no mutation in these regions can ever be deleterious. This absurd conclusion was reached through various means, chiefly (1) by employing the seldom used “causal role” definition of biological function and then applying it inconsistently to different biochemical properties, (2) by committing a logical fallacy known as “affirming the consequent,” (3) by failing to appreciate the crucial difference between “junk DNA” and “garbage DNA,” (4) by using analytical methods that yield biased errors and inflate estimates of functionality, (5) by favoring statistical sensitivity over specificity, and (6) by emphasizing statistical significance rather than the magnitude of the effect. Here, we detail the many logical and methodological transgressions involved in assigning functionality to almost every nucleotide in the human genome. The ENCODE results were predicted by one of its authors to necessitate the rewriting of textbooks. We agree, many textbooks dealing with marketing, mass-media hype, and public relations may well have to be rewritten.
A recent slew of ENCODE Consortium publications, specifically the article signed by all Consortium members, put forward the idea that more than 80% of the human genome is functional. This claim flies in the face of current estimates according to which the fraction of the genome that is evolutionarily conserved through purifying selection is under 10%. Thus, according to the ENCODE Consortium, a biological function can be maintained indefinitely without selection, which implies that at least 80 − 10 = 70% of the genome is perfectly invulnerable to deleterious mutations, either because no mutation can ever occur in these “functional” regions, or because no mutation in these regions can ever be deleterious. This absurd conclusion was reached through various means, chiefly (1) by employing the seldom used “causal role” definition of biological function and then applying it inconsistently to different biochemical properties, (2) by committing a logical fallacy known as “affirming the consequent,” (3) by failing to appreciate the crucial difference between “junk DNA” and “garbage DNA,” (4) by using analytical methods that yield biased errors and inflate estimates of functionality, (5) by favoring statistical sensitivity over specificity, and (6) by emphasizing statistical significance rather than the magnitude of the effect. Here, we detail the many logical and methodological transgressions involved in assigning functionality to almost every nucleotide in the human genome. The ENCODE results were predicted by one of its authors to necessitate the rewriting of textbooks. We agree, many textbooks dealing with marketing, mass-media hype, and public relations may well have to be rewritten.
Friday, December 14, 2012
Is Big Pharma's Drug Research Finally Speeding Up?
http://www.forbes.com/sites/matthewherper/2012/12/14/moores-law-backward-is-drug-research-finally-speeding-up/
Industry analyst Jack Scannell, writing in Nature Reviews Drug Discovery, gave this economic quicksand a name: Eroom’s Law. That’s Moore’s Law backward.
There’s even a good example, from Pfizer’s Dolsten, of a drug that has come directly out of these changes. Xalkori originally failed. But researchers discovered that it was incredibly potent in 5% of patients with non-small cell lung cancer, giving Pfizer one of its most important new drugs in years. There have been some other drugs that were clearly the result of genomics that have succeeded, including many other targeted cancer medicines and Vertex’s cystic fibrosis drug, Kalydeco, but it’s still early to say that these new tools have led to better medicines.
Industry analyst Jack Scannell, writing in Nature Reviews Drug Discovery, gave this economic quicksand a name: Eroom’s Law. That’s Moore’s Law backward.
There’s even a good example, from Pfizer’s Dolsten, of a drug that has come directly out of these changes. Xalkori originally failed. But researchers discovered that it was incredibly potent in 5% of patients with non-small cell lung cancer, giving Pfizer one of its most important new drugs in years. There have been some other drugs that were clearly the result of genomics that have succeeded, including many other targeted cancer medicines and Vertex’s cystic fibrosis drug, Kalydeco, but it’s still early to say that these new tools have led to better medicines.
Monday, November 5, 2012
Genotype-Tissue Expression (GTEx)
GTEx moves from pilot to scale up – The GTEx program is creating a comprehensive data and sample resource of genetic variation and gene expression profiles in multiple tissues from post-mortem donors. This resource will aid in the interpretation of Genome Wide Association Studies and help prioritize therapeutic targets by identifying those that affect gene expression.
http://commonfund.nih.gov/GTEx/
http://commonfund.nih.gov/GTEx/
Thursday, November 1, 2012
Charting a course for genomic medicine from base pairs to bedside
There has been much progress in genomics in the ten years since a draft sequence of the human genome was published.
Opportunities for understanding health and disease are now unprecedented, as advances in genomics are harnessed to
obtain robust foundational knowledge about the structure and function of the human genome and about the genetic
contributions to human health and disease. Here we articulate a 2011 vision for the future of genomics research and
describe the path towards an era of genomic medicine.
www.ncbi.nlm.nih.gov/pubmed/21307933
ED Green
Working Group on Data and Informatics
http://acd.od.nih.gov/diwg.htm
www.ncbi.nlm.nih.gov/pubmed/21307933
ED Green
Working Group on Data and Informatics
http://acd.od.nih.gov/diwg.htm
Sunday, October 14, 2012
Friday, September 14, 2012
ENCODE Animation
http://www.nature.com/nature/videoarchive/encode_animation/index.html
ENCODE: The story of you
Ever since a monk called Mendel started breeding pea plants we've been learning about our genomes. In 1953, Watson, Crick and Franklin described the structure of the molecule that makes up our genomes: the DNA double helix. Then, in 2000, scientists wrote down the entire 3-billion letter code contained in the average human genome. Now they're trying to interpret that code; to work out how it's used to make different types of cells and different people. The ENCODE project, as it's called, is the latest chapter in the story of you. This animation, narrated by Tim Minchin, shows how ENCODE is the culmination of two centuries of learning.
To read the research papers and more, visit nature.com/ENCODE
13 September 2011
Thursday, March 8, 2012
Genome Reference Consortium GRCm38, UCSC version mm10
We are pleased to announce the release of the latest Genome Browser for the December 2011 Mouse genome assembly. The Mus musculus genome assembly (Genome Reference Consortium GRCm38, UCSC version mm10) was produced by the Mouse Genome Reference Consortium.
GRCm38 includes approximately 2.6 Gb of sequence and is considered to be "essentially complete". The assembly includes chromosomes 1-19, X, Y, M (mitochondrial DNA) and chr*_random (unlocalized) and chrUn_* (unplaced clone contigs). For information about the process used to assemble this version, please see the GRC website.
Bulk downloads of the sequence and annotation data are available via the Genome Browser FTP server or Downloads page.
The Mouse browser annotation tracks were generated by UCSC and collaborators worldwide. See the Credits page for a detailed list of the organizations and individuals who contributed to the success of this release.
GRCm38 includes approximately 2.6 Gb of sequence and is considered to be "essentially complete". The assembly includes chromosomes 1-19, X, Y, M (mitochondrial DNA) and chr*_random (unlocalized) and chrUn_* (unplaced clone contigs). For information about the process used to assemble this version, please see the GRC website.
Bulk downloads of the sequence and annotation data are available via the Genome Browser FTP server or Downloads page.
The Mouse browser annotation tracks were generated by UCSC and collaborators worldwide. See the Credits page for a detailed list of the organizations and individuals who contributed to the success of this release.
Tuesday, January 3, 2012
Human genetics: Genomes on prescription
http://www.nature.com/news/2011/111005/full/478022a.html
As prices fall further, some say that prescribing a genome sequence or analysis will become akin to requesting a magnetic resonance imaging (MRI) scan. "It's just like any other test in medicine. There's nothing remotely special about it," says David Bick, a clinical geneticist at the Medical College of Wisconsin in Milwaukee. But, he adds, "people will cry and scream and yell about that statement". That's true: unlike the results of most medical tests, a genome sequence provides a vast amount of difficult-to-interpret data, not all of which will be necessary for diagnosing or treating the patient's condition and which could provide unwanted clues to future health risks.
they identified a mutation on the X chromosome in a gene called X-linked inhibitor of Apoptosis, or XIAP (ref. 3). A deficiency of the protein encoded by this gene is known to put patients at high risk for a deadly immune-cell disorder,
http://www.nature.com/news/2011/111005/full/478022a.html#B3
As prices fall further, some say that prescribing a genome sequence or analysis will become akin to requesting a magnetic resonance imaging (MRI) scan. "It's just like any other test in medicine. There's nothing remotely special about it," says David Bick, a clinical geneticist at the Medical College of Wisconsin in Milwaukee. But, he adds, "people will cry and scream and yell about that statement". That's true: unlike the results of most medical tests, a genome sequence provides a vast amount of difficult-to-interpret data, not all of which will be necessary for diagnosing or treating the patient's condition and which could provide unwanted clues to future health risks.
they identified a mutation on the X chromosome in a gene called X-linked inhibitor of Apoptosis, or XIAP (ref. 3). A deficiency of the protein encoded by this gene is known to put patients at high risk for a deadly immune-cell disorder,
http://www.nature.com/news/2011/111005/full/478022a.html#B3
Subscribe to:
Posts (Atom)