Thursday, August 4, 2011

Jian Pei - SFU Bioinformatics, data mining, gene expression

http://www.cs.sfu.ca/~jpei/publications.htm

Nature Articles -

Interdisciplinary studies: Seeking the right toolkit
* Bryn Nelson
http://www.nature.com/naturejobs/2011/110804/full/nj7358-115a.html?WT.ec_id=NATUREjobs-20110804

Barry Bozeman, a policy analyst at the University of Georgia in Athens who studies scientists' career trajectories, says that for now, an interdisciplinary background is “very rarely an advantage” when looking for a faculty position. Biotechnology and pharmaceutical firms might be more accommodating, as long as the applicant's unconventional research fits within the company's overall scientific aims. But formal interdisciplinary training may be less important than informal learning experiences in labs, institutes and universities that encourage the intermingling of a broad range of ideas.

“If you have ideas that the department likes and people think that what you're proposing to do is vigorous and interesting, then you will get a job,” says Anikeeva, who did her postdoctoral research at the Clark Center. “I don't think it really depends on if you have interdisciplinary training or not.”

it is much harder to get interdisciplinary faculty positions.

“gives them full citizenry in terms of access to financial and physical resources”, says the programme's website. And when they complete their graduate studies, students are “strongly advised to be strategic about their post-doctoral placement, since most must find a job in an existing more traditional field”.

“People who establish interdisciplinary degrees are also more likely to hire people with interdisciplinary degrees,” says Bozeman.

Scientists for sale:
http://www.nature.com/naturejobs/2011/110804/full/nj7358-117a.html?WT.ec_id=NATUREjobs-20110804

First, be realistic and make sure that your product fits the needs of your target audience.

Second, a sales meeting is a conversation. All the tips I found stressed that the salesperson must listen to potential buyers to understand their needs.

Finally, explain clearly what will happen after the sale. Buyers need to know how they will put you, the product, to use. Think of yourself as a new printer. Are you 'upgradable'? Be honest about what you need to get started. It's best to tell your department about the particle accelerator you'll need in your basement before the fleet of moving trucks arrives.

Of course, should everything else fail, you can always break out the car-salesman routine. Look the search-committee members squarely in the eye, give them your widest grin and ask, “Say, what will it take for me to get this job today?”


Graduate students: Aspirations and anxieties
* Gene Russo
http://www.nature.com/naturejobs/2011/110728/full/nj7357-533a.html?WT.ec_id=NATUREjobs-20110804

Across all disciplines, PhD students became less pleased with their experience as their degrees progressed. Of first-year students who responded to the survey, 76% were “satisfied” or “very satisfied”; that decreased to 66.8% for second-years and 61.3% for third-years, although the numbers varied with region (see 'Continental divide').

Hugh Kearns, a psychologist at Flinders University in Adelaide, Australia, who studies the graduate-student experience, says that the change could also be due to research results not turning out as expected. He notes that new students sometimes have unrealistically optimistic ideas about the feasibility of their research aims.

Also, getting a PhD typically takes three to four years in parts of Europe, whereas it can take five or more in the United States, which can cause dismay.

Adviser recognition is an “essential element” of quality supervision, says Marja Makarow

They have found that a lack of direction and clear advice from an adviser leads to significant declines in student satisfaction.

Thomas Skalak, vice-president for research at the University of Virginia in Charlottesville, emphasizes the need to impress upon students that they are, in the end, responsible for their own education. He likes to suggest that they act as 'intellectual entrepreneurs' by fastidiously minding their own education, graduate project, research focus and career prospects.

The survey implies that the longer students spend in graduate education, the less attractive an academic career becomes.

Intense competition for original results, publications and jobs seems to be a major factor in this change.

Among the 469 respondents, 42% of first-years wanted to be a “principal investigator at a research-intensive institution”; that dropped to 25% for third-year students. Of those who gave reasons, many cited the long work hours required, the challenge of getting funding, a distaste for daily tasks such as grant writing and the slow pace of research, and the intense competition for tenure. Some also had what Fuhrmann terms “positive” reasons for their change of preference — such as learning about an exciting new job opportunity.

Life and Years

And in the end, it's not the years in your life that count. It's the life in your years.

--Abraham Lincoln

Microarray Data Analysis Lectures and Tutorials

http://compdiag.molgen.mpg.de/ngfn/pma2005.shtml

Course materials of previous courses are online. Visit the pages for courses in 2004, 2003, 2002 and click on "Details" to get all slides, exercises and tutorials.
    
http://compdiag.molgen.mpg.de/ngfn/pma2005mar.shtml

Genomic data integration using guided clustering

http://bioinformatics.oxfordjournals.org/content/27/16/2231.abstract?etoc

Genomic data integration using guided clustering

1. Matthias Maneck1,
2. Alexandra Schrader2,
3. Dieter Kube2 and
4. Rainer Spang1,*

* Received March 16, 2011.
* Revision received June 8, 2011.
* Accepted June 13, 2011.

Motivation: In biomedical research transcriptomic, proteomic or metabolomic profiles of patient samples are often combined with genomic profiles from experiments in cell lines or animal models. Integrating experimental data with patient data is still a challenging task due to the lack of tailored statistical tools.

Results: Here we introduce guided clustering, a new data integration strategy that combines experimental and clinical high-throughput data. Guided clustering identifies sets of genes that stand out in experimental data while at the same time display coherent expression in clinical data. We report on two potential applications: The integration of clinical microarray data with (i) genome-wide chromatin immunoprecipitation assays and (ii) with cell perturbation assays. Unlike other analysis strategies, guided clustering does not analyze the two datasets sequentially but instead in a single joint analysis. In a simulation study and in several biological applications, guided clustering performs favorably when compared with sequential analysis approaches.

Availability: Guided clustering is available as a R-package from http://compdiag.uni-regensburg.de/software/guidedClustering.shtml. Documented R code of all our analysis is included in the Supplementary Materials. All newly generated data are available at the GEO database (GSE29700).

Contact: rainer.spang@klinik.uni-regensburg.de

Wednesday, August 3, 2011

Year of Science

http://yearofsciencebc.ca/calendar-of-events-and-news/

Linkage analysis HaploView

The LOD score (logarithm (base 10) of odds), developed by Newton E. Morton, is a statistical test often used for linkage analysis in human, animal, and plant populations. The LOD score compares the likelihood of obtaining the test data if the two loci are indeed linked, to the likelihood of observing the same data purely by chance. Positive LOD scores favor the presence of linkage, whereas negative LOD scores indicate that linkage is less likely.

http://en.wikipedia.org/wiki/Genetic_linkage

The deviation of the observed frequency of a haplotype from the expected is a quantity[2] called the linkage disequilibrium[3] and is commonly denoted by a capital D:
D = x11 − p1q1

In the genetic literature the phrase "two alleles are in LD" usually means that D ≠ 0. Contrariwise, "linkage equilibrium" means D = 0.

In summary, linkage disequilibrium reflects the difference between the expected haplotype frequencies under the assumption of independence, and observed haplotype frequencies. A value of 0 for D' indicates that the examined loci are in fact independent of one another, while a value of 1 demonstrates complete dependency.

http://en.wikipedia.org/wiki/Linkage_disequilibrium

Broad HaploView

http://www.broadinstitute.org/science/programs/medical-and-population-genetics/haploview/ld-display


http://www.sciencemag.org/content/296/5576/2225.long


Science. 2002 Jun 21;296(5576):2225-9. Epub 2002 May 23.

The structure of haplotype blocks in the human genome.

Gabriel SB, Schaffner SF, Nguyen H, Moore JM, Roy J, Blumenstiel B, Higgins J,
DeFelice M, Lochner A, Faggart M, Liu-Cordero SN, Rotimi C, Adeyemo A, Cooper R,
Ward R, Lander ES, Daly MJ, Altshuler D.

Whitehead/MIT Center for Genome Research, Cambridge, MA 02139, USA.

Haplotype-based methods offer a powerful approach to disease gene mapping, based
on the association between causal mutations and the ancestral haplotypes on which
they arose. As part of The SNP Consortium Allele Frequency Projects, we
characterized haplotype patterns across 51 autosomal regions (spanning 13
megabases of the human genome) in samples from Africa, Europe, and Asia. We show
that the human genome can be parsed objectively into haplotype blocks: sizable
regions over which there is little evidence for historical recombination and
within which only a few common haplotypes are observed. The boundaries of blocks
and specific haplotypes they contain are highly correlated across populations. We
demonstrate that such haplotype frameworks provide substantial statistical power
in association studies of common genetic variation across each region. Our
results provide a foundation for the construction of a haplotype map of the human
genome, facilitating comprehensive genetic association studies of human disease.


PMID: 12029063 [PubMed - indexed for MEDLINE]

Tuesday, August 2, 2011

Integrating large-scale functional genomic data to dissect the complexity of yeast regulatory networks.

Nat Genet. 2008 Jul;40(7):854-61. Epub 2008 Jun 15.

Integrating large-scale functional genomic data to dissect the complexity of
yeast regulatory networks.

Zhu J, Zhang B, Smith EN, Drees B, Brem RB, Kruglyak L, Bumgarner RE, Schadt EE.

Rosetta Inpharmatics, LLC, Seattle, Washington 98109, USA.

A key goal of biology is to construct networks that predict complex system
behavior. We combine multiple types of molecular data, including genotypic,
expression, transcription factor binding site (TFBS), and protein-protein
interaction (PPI) data previously generated from a number of yeast experiments,
in order to reconstruct causal gene networks. Networks based on different types
of data are compared using metrics devised to assess the predictive power of a
network. We show that a network reconstructed by integrating genotypic, TFBS and
PPI data is the most predictive. This network is used to predict causal
regulators responsible for hot spots of gene expression activity in a segregating
yeast population. We also show that the network can elucidate the mechanisms by
which causal regulators give rise to larger-scale changes in gene expression
activity. We then prospectively validate predictions, providing direct
experimental evidence that predictive networks can be constructed by integrating
multiple, appropriate data types.


PMCID: PMC2573859
PMID: 18552845 [PubMed - indexed for MEDLINE]

Inferring cellular networks using probabilistic graphical models.

http://www.ncbi.nlm.nih.gov/pubmed/14764868


1. Science. 2004 Feb 6;303(5659):799-805.

Inferring cellular networks using probabilistic graphical models.

Friedman N.

School of Computer Science and Engineering, Hebrew University, 91904 Jerusalem,
Israel. nir@cs.huji.ac.il

High-throughput genome-wide molecular assays, which probe cellular networks from
different perspectives, have become central to molecular biology. Probabilistic
graphical models are useful for extracting meaningful biological insights from
the resulting data sets. These models provide a concise representation of complex
cellular networks by composing simpler submodels. Procedures based on
well-understood principles for inferring such models from data facilitate a
model-based methodology for analysis and discovery. This methodology and its
capabilities are illustrated by several recent applications to gene expression
data.


PMID: 14764868 [PubMed - indexed for MEDLINE]

Boosting Signal-to-Noise in Complex Biology: Prior Knowledge Is Power

Trey Ideker, Janusz Dutkowski, and Leroy Hood, “Boosting Signal-to-Noise in Complex Biology: Prior Knowledge Is Power,” Cell 144, no. 6 (March 2011): 860-863.

http://www.cell.com/abstract/S0092-8674(11)00244-3

- filter
- integrators: combine many weak signals to increase statistical power

Examples
- Roach et al., 2010 - Familial SNV study found 3 candidate genes for Miller syndrome
- protein signaling network of kinases and phosphotases: McGary et al., 2010 - same network across species control diff. phenotypes and diff. network show similar phenotype (Erwin and Davidson 2009)

Applied bioinformatics for the identification of regulatory elements

http://www.nature.com/nrg/journal/v5/n4/full/nrg1315.html

Wyeth W Wasserman and Albin Sandelin, “Applied bioinformatics for the identification of regulatory elements,” Nature Reviews. Genetics 5, no. 4 (April 2004): 276-287.

The compilation of multiple metazoan genome sequences and the deluge of large-scale expression data have combined to motivate the maturation of bioinformatics methods for the analysis of sequences that regulate gene transcription. Historically, these bioinformatics methods have been plagued by poor predictive specificity, but new bioinformatics algorithms that accelerate the identification of regulatory regions are drawing disgruntled users back to their keyboards. However, these new approaches and software are not without problems. Here, we introduce the purpose and mechanisms of the leading algorithms, with a particular emphasis on metazoan sequence analysis. We identify key issues that users should take into consideration in interpreting the results and provide an online training example to help researchers who wish to test online tools before taking an independent foray into the bioinformatics of transcription regulation.

Monday, August 1, 2011

Adobe Edges Flash

http://www.pcmag.com/article2/0,2817,2389500,00.asp

Adobe Systems today released a preview version of an HTML5 development tool called Adobe Edge. The tool will allow Web developers to build those "little beautifully designed jewels on the Web featuring animations," Devin Fernandez, Adobe Group product manager, told PCMag last week.

http://labs.adobe.com/

Eisenhower

Pull the string, and it will follow wherever you wish. Push it, and it will go nowhere at all.