Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Robert Olson Project Name: RAST Division: MCS Project title: Rapid Genome Annotation Technology Associated funding: NIH - National Institute of Allergy and Infectious Diseases Other Systems: MCS - bioinformatics group infrastructure Science: The objectives of this project are twofold. We will demonstrate that the genome annotation technology created for a small custom-configured compute cluster can be scaled up to run efficiently on a modern large cluster like the LCRC. We will also use the computational resources of the LCRC to compute phylogenetic trees and alignments corresponding to the protein families in the FIGfams library utilized by the genome annotation server. Project description: The RAST server is a genome annotation server developed by the bioinformatics group in MCS that was funded by the NIH Bioinformatics Resource Centers program. Throughout the history of the RAST server, we have made use only of the computational cluster that was funded by NIH for the purpose of hosting RAST. The configuration of this cluster was extensively hand-tuned for the purposes of running RAST. We believe the technology has reached a state of maturity such that we can develop the tools to enable the core computation to execute on a general-purpose scientific cluster like the LCRC. The real test of the success of this work will be to process real genomic data at scale. To this end we will work toward the integration of the RAST server (which provides facilities for the upload of new genomic data, the monitoring of the progress of the pipeline, and tools for online browsing and analysis of the completed annotation as well as downloads of the completed annotation in one of several export formats) with the LCRC as a computational backend. Initially this integration will be manually supervised, but we plan to investigate the technology for a more automated integration. Advances in the RAST technology have decreased the overall computational load on a per-genome basis, but the number of genomes submitted to RAST continues to rise. In 2010 we were successful at porting the core RAST computations to Fusion, but at the same time we brought the more highly optimized computation online on the core RAST server, highly reducing the requirement for the computational offload. In 2011, however, we envision the return for the need for this offload. We plan to further automate the integration of RAST and Fusion, as well as to leverage the Magellan cluster to engineer a family of computational offload facilities for RAST. The RAST project has also made effective use of the Fusion resource for other computations related to the curation of the protein families that underpin the RAST system. We perform detailed all-to-all analyses of the genomes within families of organisms (for instance, all sequenced genomes in the Brucella family of pathogenic bacteria) to determine the precise relationship between the genes in these organisms. Fusion is an excellent platform for such analyses, and we plan to continue such analyses on the other pathogen groups. Project URL: http://rast.nmpdr.org/ Current FY Hours Used: undetermined amount New FY Requested allocation: 50000 Justification: Thank You, The LCRC Accounts System