Guest post from Rachid Ounit on CLARK: Fast and Accurate Classification of Metagenomic and Genomic Sequences

Recently I received and email from Rachid Ounit pointing me to a new open access paper he had on a metagenomics analysis tool called CLARK.  I asked him if he would be willing to write a guest post about it and, well, he did.  Here is it:

CLARK: Accurate metagenomic analysis of a million reads in 20 seconds or less…
At the University of California, Riverside, we have developed a new lightweight algorithm to classify accurately metagenomic samples while minimizing computational resources better than any other classifiers (e.g., Kraken).  While CLARK and Kraken have comparable accuracy, CLARK is significantly faster (cf. Fig. a) and uses less RAM and disk space (cf. Fig. b-c). In default mode and single-threaded, CLARK’s classification speed is higher than 3 million short reads per minute (cf. Fig. a), and it also scales better in multithreading (cf. Fig. d). Like Kraken, CLARK uses k-mers (short DNA words of length k) to solve the classification problem. However, while Kraken and other k-mers based classifiers consider the whole taxonomy tree and must resolve k-mers that match genomes from different taxa (by using the concept of “lowest common ancestor” from MEGAN), CLARK rather considers taxa defined for a unique taxonomy rank (e.g. species/genus), and, during the preprocessing, discards any k-mers that can be found in any pair of taxon. In other words, CLARK exploits specificities of each taxon (against all others) to populate its light and efficient data structure. It uses a customized dictionary of k-mers, in which each k-mer is associated to at most one taxon and results in fast k-mer queries. Then, the read is assigned to the taxon that has the highest amount of k-mers matches with it. Since these matches are discriminative, CLARK assignments are highly accurate. We also show that the choice of the value of k is critical for the optimal performance, and long k-mers (e.g., 31-mers) are not necessarily the best choice to perform accurate identification.  For example, high confidence assignments using 20-mers from real metagenomes show strong consistency with several published and independent results. 
Finally, CLARK can be used for detecting contamination in draft reference genome or, in genomics, chimera in sequenced BACs. We are currently investigating new techniques for improving the sensitivity and the speed of the tool, and we plan to release a new version later this year. We are also extending the tool for comparative genomics/metagenomics purposes. A “RAM-light” version of CLARK for your 4 GB RAM laptop is also available. CLARK is user-friendly (i.e., easy to use, it does not require strong background in programming/bioinformatics) and self-contained (i.e., does not need depend on any external software tool). The latest version of CLARK (v1.1.2) contains several features to analyze your results and is freely available under the GNU GPL license (for more details, please visit CLARK’s webpage). Experimental results and algorithm details can be found in the BMC genomics manuscript.
Performance of Kraken (v0.10.4-beta) and CLARK (v1.0) for the classification of a metagenome sample of 10,000 reads (average reads length 92bp).  a) The classification speed (in 103 reads per minute) in default mode. b) RAM usage (in GB) for the classification. c) Disk space (in GB) required for the database (bacterial genomes from NCBI/RefSeq). d) Classification speed (in 10^3 reads per minute) using 1, 2, 4 and 8 threads.

#UCDavis Summer Bioinformatics Workshops — Registration is Open!

Registration is Open for the 2015 Bioinformatics Summer Workshops!

Now in its 8th year, the UC Davis Bioinformatics Training Program will be holding two week-long workshops this summer:

June 15-19, 2015: Using Command Line for Analysis of High Throughput Sequence Data

Sept 14-18, 2015: Using Galaxy for Analysis of High Throughput Sequence Data

These workshops will be held on the UC Davis campus and will run from 9am to 5pm on the dates indicated.

Details

Both workshops will cover modern high throughput sequencing technologies, applications, and ancillary topics, including:

  • Illumina HiSeq / MiSeq, and PacBio RS technologies

  • Read Quality Assessment & Improvement

  • Genome assembly

  • SNP, indel, and CNV discovery

  • RNA-Seq differential expression analysis

  • Experimental design

  • Hardware and software considerations

  • Cloud Computing

Each workshop will include a rich collection of lectures and hands-on sessions, covering both theory and tools. We will explore the basics of several high throughput sequencing technologies, focusing on Illumina and PacBio data for hands-on exercises. Participants will explore software and protocols, create and modify workflows, and diagnose/treat problematic data. Both workshops will utilize computing power of the Amazon Cloud (http://aws.amazon.com/),

In June, exercises will be performed using the Linux command line. Therefore, for this workshop, it is strongly recommended that participants should also have basic familiarity with the Linux/Unix (or Mac) command line.

In September, workshop exercises will be performed using the popular Galaxy platform (http://usegalaxy.org) which allows for powerful web-based data analyses. There are no prerequisites other than basic familiarity with genomic concepts.

Who Should Attend

Prior course participants have included faculty, post docs, grad students, staff, and industry researchers. Anyone with an interest in sequence analysis is welcome!

Registration Info

Attendance is limited to 36 participants per workshop in order to foster an effective learning environment and ensure sufficient one-on-one attention. Course tuition is $1,500 for academic or non-profit participants and $1,800 for other participants. Amazon has kindly provided grants of $100 per participant for Amazon Web Services accounts. This will allow you to perform analysis during and after the course using Amazon’s resources, without purchasing your own high performance computing servers!

To register, click on the links above or go to training.bioinformatics.ucdavis.edu/. All registration is “first-come, first-served”. There is no application process. We accept credit cards, as well as UC recharge accounts, for payment. Registration fees include light breakfast, lunch, and snacks, but do not include dinner, lodging or parking fees.

Questions

If you have any questions, please don’t hesitate to contact us:

  • Core email: bioinformatics.core
  • Core main telephone line: 530-752-2698

See you this summer!

The UC Davis Bioinformatics Core Team

http://training.bioinformatics.ucdavis.edu

http://bioinformatics.ucdavis.edu/

Job: Director for our Biostatistics and Bioinformatics Core at Gladstone/UCSF.

From colleague at the Gladstone Inst.

***

The Gladstone Institutes is an independent, not-for-profit research organization affiliated with the University of California San Francisco (UCSF), contributing to the health and well-being of all people through medical research, education, and outreach in the areas of heart disease, HIV/AIDS, and neurological disease. Gladstone’s approximately 450 employees receive exceptional benefits. We are located in an award winning building adjacent to UCSF’s Mission Bay Campus.

We are seeking a Director for the Bioinformatics Core at Gladstone. The Bioinformatics Core is dedicated to providing outstanding professional services to our researchers in need of consulting, statistical analysis support, software development, and analysis of a wide range of ‘omics data types. The Bioinformatics Core Director is expected to lead our professional services team, following best practices for consulting that includes understanding and fitting solutions to needs, estimating efforts and skills, contributing high value to proposals and creating high quality deliverables for project reports and publications. The Director may apply to join UCSF as a faculty member in the adjunct series.

Gladstone recently launched an initiative to establish a “Convergence Zone”, which will be a vibrant group of labs and core facilities focused on emerging and highly innovative technologies for biomedical research (e.g., new approaches in imaging, single-cell analysis, nanotechnology, mass spectrometry, genome editing). Bioinformatics and biostatistics will play a central role in this team, including developing computational tools for analyzing and integrating the novel data types from Convergence Zone platforms. The Director will play a key role in ensuring that we have the technical expertise and collaborative culture for this effort to be successful.

Responsibilities

· Mentor and supervise the Bioinformatics Core team, which:

o Provides services to investigators across Gladstone and UCSF, including pre-experiment consulting, analysis, and visualization of results.

o Provides workshops in bioinformatics and statistical analysis.

o Develops pipelines for large-scale ‘omics data analysis.

o Assesses, tests, and implements best practices for ‘omics data analysis.

o Conducts statistical analysis of a variety of data types including longitudinal, hierarchical, and survival data.

o Develops tools to integrate commonly used open source bioinformatics software applications.

· Develop long-term vision for the Bioinformatics Core and work with Gladstone Investigators to prioritize core services to best meet user needs.

· Refine the funding model for expansion and sustainability of the core, including investments from research institutes within Gladstone, charge-back models, and collaborative grant applications.

· Develop staffing plans and oversee recruitments as needed.

· Work effectively with multiple stakeholders and collaborators, including experimental and computational faculty and researchers, as well as the Scientific Computing group within IT.

· Lead and collaborate on grant applications related to Bioinformatics Core projects.

Required Education and Experience

· PhD in statistics, computational biology, computer science, or related field.

· Previous experience with statistical analysis of biological data.

· Extensive experience with high-throughput sequence analysis, such as RNA-seq, ChIP-seq, and/or whole genome/exome variant analysis.

· Fluency in a scripting language (e.g., perl, python) and the R programming language.

· High performance computing experience in a Unix/Linux environment.

· Practical experience with web-based analytical platforms and cloud computing resources.

· Minimum of 3 years of experience supervising staff at PhD and/or MS level.

· Experience with budgeting and financial administration of grants.

· Strong record of publications and success obtaining extramural funding.

Rob Dunn seeking community participation in suveying & analyzing Duke Forest warming chambers

Just got this email from Rob Dunn from NC State.  He said it was OK to post it … so I am .. (I note – I just completely love this idea).

Hi folks,

As you might (or might not) know, we have for five years now been running a large-scale warming experiment in which we have warmed twelve 5 meter diameter open-top chambers in forest understory at Duke Forest. We have warmed these chambers in a regression design with the warmest chambers as warm as temperatures are predicted to be in the region in 2100 and the coolest chambers at ambient temperatures (We also have no-chamber controls). These are small worlds each of which mimics aspects of futures we might face. This entire set-up is replicated at Harvard Forest. In these chambers we have been studying the response of insects (with a focus on ants) and plants  over the last five years. When we built them these chambers were the biggest warming experiment in a forest understory in the world. I don’t know if it is still true, but it probably is, if only because chambers of this size are so hard to keep going (especially in the early we felt like Fitzcarraldo dragging a ship through the rainforest) that most people have decided against repeating them elsewhere. 
Some basics on the chambers… http://robdunnlab.com/projects/warming-chambers/
I’m writing because on May 25th we are taking the chambers down and doing a final inventory of the response of everything–all the life we can possibly evaluate–to this warming. To varying extents we have considered the phenology of plants in the chambers, many things about ants in the chambers, shifts in composition of invertebrates in the chambers and simple responses of bacterial and fungal assemblages in the chambers. But, we have done all of this delicately, always mindful to not overly disturb the future world we are simulating. Now though that the chambers are coming down we can and will consider roots, plant biomass, the abundance of insect pests, fungal pathogens and much, much, more. 
As we do this intensive survey, we are hoping to train as many different eyes, lenses and perspectives on the chambers as possible. If you are potentially interested in studying some aspect of the response of understory forest life to warming, let us know. If you are interested in studying something that can be extracted from soil or litter samples, we may be able to send you material you can work on. If you have something grander in mind (and we love grand things), then we may need more help from you. If interested, send an email to me, copied to MJ Epps (Mj Epps <mycota@gmail.com>).  This collaboration might be in the form of bringing a new method to the chambers (looking at microbial processes, for instance) or considering a group of organisms we’ve somewhat ignored (e.g., fly larvae) or it might be something totally off the wall. Feel free to share this email with likable folks that might be interested. 
I’m also delighted to hear creative ideas about visualizing the differences that have emerged over the years of this experiment (hence the inclusion of several artists of various sorts on this email list, if you were wondering why you were copied). 
I hope this email finds you well. 
Best,
Rob

Plant Sciences seminar – Jeff Ross-Ibarra – May 6

Plant Sciences Seminar Announcement…

Jeffrey Ross-Ibarra (Plant Sciences, UCD) will be presenting “Adaptation in maize: domestication and beyond”.

Wednesday, May 6 from 12-1pm in 3001 PES.

Jeff Ross-Ibarra May 6 seminar flyer.pdf

Seminar at #UCDavis today: Cross-kingdom molecular battles in the phyllosphere

MIC 291: Selected Topics in Microbiology

Work-in-Progress Seminars

Dr. Maeli Melotti
(Dept. of Plant Sciences)

"Cross-kingdom molecular battles in the phyllosphere"
Wednesday, April 29, 2015
4:10 pm

1022 Life Sciences

Melotto 4-22-15.doc

Some comments on Williams & Ceci (2015)

I wrote a blog post over at Nothing In Biology Makes Sense! on some of the methodological and interpretation issues I had with a recent PNAS paper titled “National hiring experiments reveal 2:1 faculty preference for women on STEM tenure track”. Long story short, I think there’s an underlying problem with the authors’ assumption that identically worded sentences are identically interpreted by human beings. I also struggled with their conclusion that women now have the advantage over men and with the general “Everything is great now!” feeling of the paper.

But please – go read the whole thing! And let me know what you think! I’d love discussion or comments on this – did you read the paper? How do you feel about their methods and conclusions? How do you feel about the climate for women in STEM in general?

My figure from “Isn’t that just…sexism?” over at Nothing in Biology Makes Sense.

Permanent Position at NSF in Population and Community Ecology

__________________Position Announcement_________________

Permanent position at NSF in Population and Community Ecology

The Division of Environmental Biology at NSF seeks applicants for a Program Director in its Population and Community Ecology cluster. This is a permanent position, with good benefits and a salary range of $107,325 – $167,252. Previous experience with NSF as a PI, panelist or rotator is helpful but not essential. The deadline is May 18, 2015 and the application process is relatively straightforward.

Applicants should have a PhD and expertise in at least one of the following: population ecology, species interactions, and community structure/dynamics in terrestrial, wetland or freshwater habitats.

For details, please see the USAjobs.gov posting (link below) or contact Doug Levey (dlevey@nsf.gov).

Members of underrepresented groups, including individuals with disabilities, are especially encouraged to apply.

https://www.usajobs.gov/GetJob/ViewDetails/399791500

C-DEBI Seeks an Education, Outreach, and Diversity Manager

Just got this announcement:


C_DEBIlogo3sq.jpgC-DEBI Seeks an Education, Outreach, and Diversity Manager

The Center for Dark Energy Biosphere Investigations (C-DEBI) is seeking a Program Manager to join its team. C-DEBI is a multi-institutional research and education center funded by the National Science Foundation with USC as its headquarters. In addition to the science of exploring microbial life beneath the seafloor, education and diversity are priorities to the Center’s efforts to strengthen the STEM pipeline by integrating research and educational programs for diverse future generations. The full-time Program Manager will serve as Education, Outreach, and Diversity Manager, helping to create, coordinate, and lead our education, outreach, and diversity efforts to serve our students, postdocs, faculty, and other participants at USC and across the nation. The Program Manager will also direct day-to-day project operations and administrative activities of C-DEBI at USC.

The ideal candidate for the position of C-DEBI Education, Outreach, and Diversity Manager has:

  • 5 years of experience developing and managing educational STEM programs with multiple institutions
  • Leadership and strong oral and written communication skills
  • Experience using social media for professional outreach
  • Research experience at Ph.D. level

See the USC jobs website for more information on this posting ID 1003195:
http://jobs.usc.edu/postings/43150

open.php?u=f36f3ea470b4b11934dd3374a&id=58bcf1dc5c&e=48f2265255

At #UCDavis 4/15 and 4/16: Dr. Tim Clutton-Brock

REMINDER: