http://www.ted.com/talks/dan_gilbert_asks_why_are_we_happy.html#!
The lesson I want to leave you with from these data
is that our longings and our worries are both to some degree overblown,
because we have within us the capacity to manufacture the very commodity
we are constantly chasing when we choose experience.
Just a collection of some random cool stuff. PS. Almost 99% of the contents here are not mine and I don't take credit for them, I reference and copy part of the interesting sections.
Monday, December 30, 2013
An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics.
http://www.ncbi.nlm.nih.gov/pubmed/21210976?dopt=Abstract
BMC Bioinformatics. 2010 Dec 21;11 Suppl 12:S1. doi: 10.1186/1471-2105-11-S12-S1.
An overview of the Hadoop/MapReduce/HBase framework and its current applications in bioinformatics.
Abstract
BACKGROUND:
Bioinformatics researchers are now confronted with analysis of ultra large-scale data sets, a problem that will only increase at an alarming rate in coming years. Recent developments in open source software, that is, the Hadoop project and associated software, provide a foundation for scaling to petabyte scale data warehouses on Linux clusters, providing fault-tolerant parallelized analysis on such data using a programming style named MapReduce.
DESCRIPTION:
An overview is given of the current usage within the bioinformatics community of Hadoop, a top-level Apache Software Foundation project, and of associated open source software projects. The concepts behind Hadoop and the associated HBase project are defined, and current bioinformatics software that employ Hadoop is described. The focus is on next-generation sequencing, as the leading application area to date.
CONCLUSIONS:
Hadoop and the MapReduce programming paradigm already have a substantial base in the bioinformatics community, especially in the field of next-generation sequencing analysis, and such use is increasing. This is due to the cost-effectiveness of Hadoop-based analysis on commodity Linux clusters, and in the cloud via data upload to cloud vendors who have implemented Hadoop/HBase; and due to the effectiveness and ease-of-use of the MapReduce method in parallelization of many data analysis algorithms.
- PMID:
- 21210976
- [PubMed - indexed for MEDLINE]
- PMCID:
- PMC3040523
Friday, December 13, 2013
PDF to Text with OCR Ubuntu
$sudo apt-get install tesseract-ocr
$sudo apt-get install convert
$ convert -density 300 in.pdf out.png
$ tesseract out.png out
$ vi out.txt
For multiple files
for i in out-*.png ; do tesseract $i $i-txt; done
$sudo apt-get install convert
$ convert -density 300 in.pdf out.png
$ tesseract out.png out
$ vi out.txt
For multiple files
for i in out-*.png ; do tesseract $i $i-txt; done
Tuesday, December 10, 2013
Java GUI applications runs slowly over X11 in Ubuntu
Add the Java option
-Dsun.java2d.pmoffscreen=false
when running your remote application
e.g.
$ java -Dsun.java2d.pmoffscreen=false -cp . my.app
http://www.mathworks.com/matlabcentral/newsreader/view_thread/160387
-Dsun.java2d.pmoffscreen=false
when running your remote application
e.g.
$ java -Dsun.java2d.pmoffscreen=false -cp . my.app
http://www.mathworks.com/matlabcentral/newsreader/view_thread/160387
Monday, December 9, 2013
Aspect Oriented Programming (AOP) - AspectJ
http://www.voelter.de/data/articles/aop/aop.html
Object oriented programming has become mainstream over the last years, having almost completely replaced the procedural approach. One of the biggest advantages of object orientation is that a software system can be seen as being built of a collection of discrete classes. Each of these classes has a well defined task, its responsibilities are clearly defined. In an OO application, those classes collaborate to achieve the application's overall goal. However, there are parts of a system that cannot be viewed as being the responsibility of only one class, they cross-cut the complete system and affect parts of many classes. Examples might be locking in a distributed application, exception handling, or logging method calls. Of course, the code that handles these parts can be added to each class separately, but that would violate the principle that each class has well-defined responsibilities. This is where AOP comes into play: AOP defines a new program construct, called an aspect, which is used to capture cross-cutting aspects of a software system in separate program entities. The application classes keep their well-defined responsibilities. Additionally, each aspect captures cross-cutting behaviour.
Object oriented programming has become mainstream over the last years, having almost completely replaced the procedural approach. One of the biggest advantages of object orientation is that a software system can be seen as being built of a collection of discrete classes. Each of these classes has a well defined task, its responsibilities are clearly defined. In an OO application, those classes collaborate to achieve the application's overall goal. However, there are parts of a system that cannot be viewed as being the responsibility of only one class, they cross-cut the complete system and affect parts of many classes. Examples might be locking in a distributed application, exception handling, or logging method calls. Of course, the code that handles these parts can be added to each class separately, but that would violate the principle that each class has well-defined responsibilities. This is where AOP comes into play: AOP defines a new program construct, called an aspect, which is used to capture cross-cutting aspects of a software system in separate program entities. The application classes keep their well-defined responsibilities. Additionally, each aspect captures cross-cutting behaviour.
Friday, December 6, 2013
The Data Effect
http://cityage.tv/thedataeffect/program/
Data has been called our 21st-Century resource. For good reason. We live in a digital age, when data across all sectors of society is being created and stored at historic proportions. How we protect and use that data, as well as structure share and analyze it, stands to transform health care, research and other sectors in Canada.
The Data Effect, now in its fourth edition, has assembled the private, public and research leaders who are capitalizing on data to drive health research and innovation. They are part of the CityAge’s goal to make Canada a leader in the proper and innovative use of data for the public good.
The fourth edition of The Data Effect will build on the inaugural version held in Vancouver in June 2012, which explored how BC’s uniquely high quality population health data can be put to use to save lives and improve health research. Using focused discussions and use cases, the event will address the steps required to make British Columbia a centre for excellence on the use of data for 21st Century advanced health care and research.
Thursday, December 5, 2013
Combining Hibernate Cache and Ehcache for Better Java Scalability
http://www.developer.com/java/ent/article.php/3897536/Combining-Hibernate-Cache-and-Ehcache-for-Better-Java-Scalability.htm
Hence, it is critical to lower your database load. In particular, you can employ efficient caching strategies that reduce the number of SQL queries and the amount of data that needs to be transferred over the wire. One powerful caching strategy that meets these requirements is to combine Hibernate cache (second-level) and Ehcache.
Hence, it is critical to lower your database load. In particular, you can employ efficient caching strategies that reduce the number of SQL queries and the amount of data that needs to be transferred over the wire. One powerful caching strategy that meets these requirements is to combine Hibernate cache (second-level) and Ehcache.
Tuesday, December 3, 2013
Kurtosis - from the Greek word κυρτός, kyrtos or kurtos, meaning curved, arching
http://en.wikipedia.org/wiki/Kurtosis
In probability theory and statistics, kurtosis (from the Greek word κυρτός, kyrtos or kurtos, meaning curved, arching) is any measure of the "peakedness" of the probability distribution of a real-valued random variable.[1] In a similar way to the concept of skewness, kurtosis is a descriptor of the shape of a probability distribution and, just as for skewness, there are different ways of quantifying it for a theoretical distribution and corresponding ways of estimating it from a sample from a population.
In probability theory and statistics, kurtosis (from the Greek word κυρτός, kyrtos or kurtos, meaning curved, arching) is any measure of the "peakedness" of the probability distribution of a real-valued random variable.[1] In a similar way to the concept of skewness, kurtosis is a descriptor of the shape of a probability distribution and, just as for skewness, there are different ways of quantifying it for a theoretical distribution and corresponding ways of estimating it from a sample from a population.
Monday, December 2, 2013
Find out disk space usage - dispus
$ dispus
Get to the shell by typing "!" and "exit" to return to dispus
Get to the shell by typing "!" and "exit" to return to dispus
Wednesday, November 27, 2013
Biological Ontologies
http://www.berkeleybop.org/ontologies
The BBOP, located at the Lawrence Berkeley National Labs, is a diverse group of scientific researchers and software engineers dedicated to developing tools and applying computational technologies to solve biological problems. Members of the group contribute to a number of projects, including the Gene Ontology, OBO Foundry, the Phenotypic Quality Ontology, modENCODE, and the Generic Model Organism Database Project.
http://www.ontobee.org/
Ontobee: A linked data server designed for ontologies. Ontobee is aimed to facilitate ontology data sharing, visualization, query, integration, and analysis. Ontobee dynamically dereferences and presents individual ontology term URIs by generating RDF/XML source code for Semantic Web applications, and providing a well-structured HTML web page for user-friendly visualization in a web browser. Ontobee is the default linked data server for most OBO Foundry and OBO Library ontologies.
The BBOP, located at the Lawrence Berkeley National Labs, is a diverse group of scientific researchers and software engineers dedicated to developing tools and applying computational technologies to solve biological problems. Members of the group contribute to a number of projects, including the Gene Ontology, OBO Foundry, the Phenotypic Quality Ontology, modENCODE, and the Generic Model Organism Database Project.
http://www.ontobee.org/
Ontobee: A linked data server designed for ontologies. Ontobee is aimed to facilitate ontology data sharing, visualization, query, integration, and analysis. Ontobee dynamically dereferences and presents individual ontology term URIs by generating RDF/XML source code for Semantic Web applications, and providing a well-structured HTML web page for user-friendly visualization in a web browser. Ontobee is the default linked data server for most OBO Foundry and OBO Library ontologies.
Monday, November 25, 2013
Ehcache: Java's Most Widely-Used Cache
http://ehcache.org/
Ehcache: Java's Most Widely-Used Cache
Ehcache is an open source, standards-based cache for boosting performance, offloading your database, and simplifying scalability. It's the most widely-used Java-based cache because it's robust, proven, and full-featured. Ehcache scales from in-process, with one or more nodes, all the way to mixed in-process/out-of-process configurations with terabyte-sized caches.
Teaching the art of guessing
http://blogs.nature.com/thescepticalchymist/2013/08/teaching-the-art-of-guessing.html
Estimation of chemical quantities is, in my opinion, one of the most important skills a chemist can have. It is also one that my students are simultaneously skeptical of — why not just use a search engine to find the ‘real’ value or crunch numbers to 8 decimal places in your calculator — and intimidated by (I probably don’t help by telling this story about Fermi).
I emphasize the need to gather a tool kit of anchor points (key chemical values) and maps (relationships between quantities) and provide examples. Here is the handout (pdf)
Estimation of chemical quantities is, in my opinion, one of the most important skills a chemist can have. It is also one that my students are simultaneously skeptical of — why not just use a search engine to find the ‘real’ value or crunch numbers to 8 decimal places in your calculator — and intimidated by (I probably don’t help by telling this story about Fermi).
I emphasize the need to gather a tool kit of anchor points (key chemical values) and maps (relationships between quantities) and provide examples. Here is the handout (pdf)
Tuesday, November 5, 2013
Slides and Video: Metrics that Matter with Chris Geoghegan
http://www.net2van.com/2013/09/slides-and-video-metrics-that-matter-with-chris-geoghegan/
Chris has bad news. Most of the time you spend looking at your analytics is probably in vain. There is too much data that tells you too little information. You thought analytics was all about using information to make better decisions, but it isn’t working. How do you go beyond just giving yourself a nice pat on the back with how many visitors you got last month?
Chris has bad news. Most of the time you spend looking at your analytics is probably in vain. There is too much data that tells you too little information. You thought analytics was all about using information to make better decisions, but it isn’t working. How do you go beyond just giving yourself a nice pat on the back with how many visitors you got last month?
Subscribe to:
Posts (Atom)