Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Robert Latham Project Name: radix Division: MCS Project title: Scalable Parallel System Software Associated funding: DoE core, SciDAC, NSF, FASTOS Other Systems: breadboard (unlimited) ALCF Intrepid/Surveyor (5,000,000 cpu hours) Science: Our mission is to develop the technologies required to dramatically increase the productivity of scientists developing applications for parallel supercomputers. The focus is fourfold: integration of parallel programming tools, reuse of parallel program components, development of scientific computing toolkits and portable libraries, and exploration of requirements of future parallel computers. Project description: This is a Computer Science request for 450,000 cpu-hours on Fusion for development and evaluation of a family of tools and libraries with a direct benefit to all users of high-performance computing resources. The Radix group in Argonnne's Mathematics and Computer Science division conducts research and development on facilities to better enable efficient Parallel Programming, I/O, and Scientific Understanding. The Fusion system represents an attractive milestone for our software development: projects start out on our breadboard cluster and can further explore scalability in the presence of a high speed network on Fusion before making the jump to BlueGene-scale parallelism. Radix experiments often are concerned with latency. The infiniband network on Fusion provides an excellent low-latency environment for message passing and storage management research while also providing ample bandwidth for I/O and message bandwidth testing. For FY2012 our Messaging middleware experiments include the usual MPICH2 development as well as application-oriented libraries like ADLB, a load-balancing library. Fusion hosts two quality parallel file systems, GPFS and PVFS. The wider HPC community often asks the Radix file system researchers to compare these file systems head-to-head. Fusion provides us the opportunity to do just that. For FY 2012 our file system research extends into data storage approaches suitable for exascale. Our research into object storage protocols, including high performance access and replication, fits well on Fusion, with the large amount of local storage on compute nodes. Hosting two parallel file systems allows us to do more than just drag-race (so to speak). The two file systems have distinct characteristics, and act differently in the broader role of the I/O software stack. On Fusion, we can utilize a single software stack (application, parallel-netcdf, ROMIO), and compare the impact of a change in the underlying file system on overall application behavior. We know from other research, for example, that the two-phase optimization in ROMIO needs to adapt to the underlying file system, and that aligning variables to file system boundaries can yield performance improvements for parallel-netcdf. In FY 2012 we are continuing research into new high level I/O libraries. We anticipate Fusion and Surveyor will be our two main test platforms. Fusion compute nodes actually look fairly attractive to active storage research. These nodes have enough disk space to store non-trivial datasets, and powerful enough processors to carry out complex computations without much impact on storage performance. Radix projects also focus on deriving insight both about application behavior as well as scientific insight. 8 core nodes make efforts like in-situ visualization, where for example an application renders a frame of a movie, ever more feasible. Even without fancy graphics accelerators, the processing power on Fusion nodes makes such analysis and visualization feasible, especially if the visualization can take advantage of multiple cores and multiple nodes. Naturally, the radix visualization tools have demonstrated scalability to Fusion sizes and beyond on BlueGene, so making full utilization of Fusion should not be a concern. As should be evident from the variety of planned experiments, the Radix group plans to make full use of Fusion. With Argonne employees and student collaborators, we expect around 20 members in FY2012. We plan on about 0% of our jobs being single core jobs. It's crazy that you even have to ask! While Fusion may contain a modest number of compute nodes, the machine still represents an attractive platform for testing: very fast, low latency interconnect; 8 cores per node; .25 TB of storage per compute node. The testing and experiments the Radix group can carry out on Fusion ensures the tools the group develops today will remain relevant even as supercomputers grow. Quite a few radix-developed projects are part of the Fusion software stack: CPU hours for this project yield improvements and benefits not just for the radix group, but for all users of Fusion and indeed users of high-end computational resources worldwide. Project URL: http://www.mcs.anl.gov/research/group_detail.php?id=1 Current FY Hours Used: undetermined amount New FY Requested allocation: 448000 Q1: 112000 Q2: 112000 Q3: 112000 Q4: 112000 Justification: Thank You, The LCRC Accounts System