[LCRC Accounts] Yearly Allocation Request from radix
Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Robert Latham Project Name: radix Division: MCS Project title: Scalable Parallel System Software Associated funding: DoE core, SciDAC, NSF, FASTOS Other Systems: breadboard (unlimited) ALCF Intrepid/Surveyor (5,000,000 cpu hours) Science: Our mission is to develop the technologies required to dramatically increase the productivity of scientists developing applications for parallel supercomputers. The focus is fourfold: integration of parallel programming tools, reuse of parallel program components, development of scientific computing toolkits and portable libraries, and exploration of requirements of future parallel computers. Project description: This is a Computer Science request for 49,500 cpu-hours on Fusion for development and evaluation of a family of tools and libraries with a direct benefit to all users of high-performance computing resources. The Radix group in Argonnne's Mathematics and Computer Science division conducts research and development on facilities to better enable efficient Parallel Programming, I/O, and Scientific Understanding. The Fusion system represents an attractive milestone for our software development: projects start out on our breadboard cluster and can further explore scalability in the presence of a high speed network on Fusion before making the jump to BlueGene-scale parallelism. Radix experiments often are concerned with latency. The infiniband network on Fusion provides an excellent low-latency environment for message passing and storage management research while also providing ample bandwidth for I/O and message bandwidth testing. For FY2013 our Messaging middleware experiments include the usual MPICH2 development. We will study the continued improvement of various MPI features including scalability and memory usage, performance tuning, high-level libraries on top of MPI, fault tolerance capabilities, one-sided communication, and collective operations. This allocation will also be used for studying MPI extensions including active messages, and compiler support for MPI refactoring. We also plan for developing application-oriented libraries like ADLB, a load-balancing library. We will port the ADLB load-balancing library to Fusion and test its ability to function with multi-core nodes at intermediate scale, taking advantage of the multiple cores to replace MPI communication with local shared memory. The NWChem library will continue to improve the performance of an important scientific application that is widely used for science by Argonne researchers. Specifically, we will develop inspect-executor load-balancing techniques that will reduce network contention and improve the overall performance and scalability of an important family of methods (coupled-cluster methods). Furthermore, we will explore hybrid programming models (that is, adding threading to the existing GA/ARMCI/MPI model), particularly in the TCE module, which has been used extensively by researchers in Argonne MSD and CSE. All the aforementioned developments will be deployed on Fusion in a production fashion and made available to all users. This allocation will enable Jeff Hammond to maintain the NWChem installations on Fusion, as he has done since the machine went live. We work closely with application scientists, and develop application tools under the 'radix' allocation. For one example, the KMI project defines a high-level standardized interface for computational biology applications relying on distributed search and match semantics for biological reads. We will study efficient data management and data movement techniques to allow scaling to large number of processes. Fusion hosts two quality parallel file systems, GPFS and PVFS. The wider HPC community often asks the Radix file system researchers to compare these file systems head-to-head. Fusion provides us the opportunity to do just that. For FY 2013 our file system research extends into data storage approaches suitable for exascale. Our research into object storage protocols, including high performance access and replication, fits well on Fusion, with the large amount of local storage on compute nodes. Our Exascale storage efforts will explore scalable algorithms for use in next generation HPC storage systems. The Fusion allocation will be used to evaluate programming models, fault detection algorithms, synchronization primitives, and fault tolerance strategies for a prototype object storage system. Our NoLoss project explores how to integrate in-system storage in the I/O software stack. Under the NoLoss project, an abstraction layer for in-system storage was developed. The SCR checkpointing library was modified to store checkpoints locally using this newly developed abstraction layer. The fusion allocation will be used to explore the performance characteristics of this new approach. Hosting two parallel file systems allows us to do more than just drag-race (so to speak). The two file systems have distinct characteristics, and act differently in the broader role of the I/O software stack. On Fusion, we can utilize a single software stack (application, parallel-netcdf, ROMIO), and compare the impact of a change in the underlying file system on overall application behavior. We know from other research, for example, that the two-phase optimization in ROMIO needs to adapt to the underlying file system, and that aligning variables to file system boundaries can yield performance improvements for parallel-netcdf. In FY 2013 we are continuing research into new high level I/O libraries. We anticipate Fusion and Surveyor will be our two main test platforms. Fusion compute nodes actually look fairly attractive to active storage research. These nodes have enough disk space to store non-trivial datasets, and powerful enough processors to carry out complex computations without much impact on storage performance. Radix projects also focus on deriving insight both about application behavior as well as scientific insight. 8 core nodes make efforts like in-situ visualization, where for example an application renders a frame of a movie, ever more feasible. Even without fancy graphics accelerators, the processing power on Fusion nodes makes such analysis and visualization feasible, especially if the visualization can take advantage of multiple cores and multiple nodes. Naturally, the radix visualization tools have demonstrated scalability to Fusion sizes and beyond on Blue Gene, so making full utilization of Fusion should not be a concern. As should be evident from the variety of planned experiments, the Radix group plans to make full use of Fusion. With Argonne employees and student collaborators, we expect around 20 members in FY2013. We plan on about 0% of our jobs being single core jobs. It's crazy that you even have to ask! While Fusion may contain a modest number of compute nodes, the machine still represents an attractive platform for testing: very fast, low latency interconnect; 8 cores per node; .25 TB of storage per compute node. The testing and experiments the Radix group can carry out on Fusion ensures the tools the group develops today will remain relevant even as supercomputers grow. Quite a few radix-developed projects are part of the Fusion software stack: CPU hours for this project yield improvements and benefits not just for the radix group, but for all users of Fusion and indeed users of high-end computational resources worldwide. Project URL: http://www.mcs.anl.gov/research/group_detail.php?id=1 Current FY Hours Used: undetermined amount New FY Requested allocation: 49500 Q1: 12375 Q2: 12375 Q3: 12375 Q4: 12375 Justification: Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov