[LCRC Accounts] Yearly Allocation Request for radix
Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Halim Amer Project Name: radix Division: MCS Project title: Scalable Parallel System Software Associated funding: DoE core, SciDAC, NSF, FASTOS Other Systems: breadboard (unlimited) ALCF Mira/Cetus/Tukey (16,000,000 cpu hours) Science: Our mission is to develop the technologies required to dramatically increase the productivity of scientists developing applications for parallel supercomputers. The focus is fourfold: integration of parallel programming tools, reuse of parallel program components, development of scientific computing toolkits and portable libraries, and exploration of requirements of future parallel computers. Project description: This allocation will be used for optimizations to various systems software components: (1) programming models and runtime systems, (2) data I/O and file systems, (3) fault tolerance and data compression, (4) operating systems, (5) data analysis and visualization, and (6) deep learning. Primary goal of the project is to perform research on improving the core systems software on high-end computing platforms. The radix allocation covers a large number of cross-cutting projects involving subsets of researchers in the group. A few core projects are presented here, but some of the techniques used are shared between them and are not easily separable (e.g., MPI improvements or resilience improvements): 1. Fault Tolerance and Data Compression - Local memory management * Advanced SDC Detection through Data Monitoring + Adapt the prediction techniques used in the SZ lossy compressor for SDC detection * Checkpoint size reduction to improve checkpoint/restart performance and energy efficiency + Adapt FTI for using lossy compression of checkpoints * Impact of lossy compression on Checkpoint/restart + Investigate fundamentally and experimentally the relation between discretization/truncation error with lossy compression error * End-to-end Detection of SDCs in workflows featuring a lossy compression stage + Compare combination of SDC detectors and lossy compressors with respect to the accuracy of SDC detection. - Distributed memory management * Integration of checkpointing in workflow environments + Adapt DECAF and/or SWIFT with FTI and perform performance evaluation * MPI Fault Tolerance + Improvements to MPICH to deal with faults * Resilient-MPI and FTI Integration + Experimentation and evaluation of FTI coupled with resilient-MPI - Improvements to lossy data compression quality assessment through Z-checker * Optimize offline quality assessment to better characterize the * various lossy compression techniques * Speed up Z-checker with parallel assessment algorithms * Augment Z-checker with online visualization of the in-situ lossy compression results. 2. Developpement and optimization of OpenMP and threading runtimes - Improvements to our BOLT OpenMP runtime and exploring new OpenMP capabilities - Improvements to the Argobots threading framework - Study of the above runtimes under fat Xeon cores as well as slow cores such as KNL 3. Lightweight, topology-aware, and thread-safe communication: - Improvements to MPI+threads * MPICH improvements to play well with threads, including improved locking and lock-free data structures. * Communication progress improvements in the presence of threads. - Lightweight communication * Low instruction-count and low-overhead irregular data movement. - Topology-aware communication * Optimizations to MPI virtual topology functionality 4. Data I/O, RMA and Active Messages: - Dynamic Locality-Aware Workflows * Dynamically distribute the tasks among nodes using Asynchronous Dynamic Load Balancer - Development and Evaluation of Composable Data Services * Exploration of lightweight, customized, domain-specific storage services to improve the performance of data-intensive HPC applications - High Performance Parallel I/O * The ROMIO MPI-IO implementation provides the foundation for many components of the HPC I/O stack. Continue tuning, debugging, and research efforts. 5. Data analysis and visualization: Testing and benchmarking of improvements to data movement infrastructure such as DIY and Decaf and to data model software such as MFA to prepare for next-generation machine architectures and to apply those improvements to applications across the lab that use those software platforms, such as cosmology, climate science, and materials science. The following areas will be tested on Blues and Bebop: - Relaxed BSP synchronization (e.g., for iterative algorithms, to overlap communication with computation and support irregular communication/computation patterns) * irregular communication patterns - Load-balancing algorithms (work stealing, adaptive decompositions) for unstructured time-varying mesh partitions (e.g., reactor design) * work stealing algorithms * adaptive (e.g. k-d tree) domain decompositions - Workflow management * User interfaces * Scalability, performance * Integration with distributed area workflows and system software * Application to HEP, climate, superconductivity, materials science - Data models * Multivariate functional approximation with adaptive knot insertion * Linear and nonlinear solutions * Domain segmentation 6. I/O enhancements for scalable deep learning - In FY17, we focused on optimizing the usage of mmap for deep learning. In FY18, we plan to move our focus on replacing mmap with direct I/O mechanisms. We will investigate mechanisms to use MPI I/O for direct I/O in an effort to improve performance through bulk I/O while maintaining correctness through fall-back mmap buffers. - Our work in FY17 focused on the Caffe framework. We plan to extend our work to newer deep learning frameworks such as Tensorflow and Caffe2. - We would like to investigate the impact of Intel Xeon Phi architecture on deep learning frameworks, and in particular how the improved computational resources would affect I/O performance. Industry partnership: Project URL: http://www.mcs.anl.gov/group/pmrs/ Current FY Hours Used: undetermined amount New FY Requested allocation: 5200000 Q1: 1300000 Q2: 1300000 Q3: 1300000 Q4: 1300000 Justification: 5.2M core hours requested, most of which would be using Bebop: We have six core projects. All projects have been demonstrated at scale in the previous years and form the core of the system software stack used in production on the Fusion and Blues machines. We expect to run around ~20 studies for each project on Bebop and on Blues, scaling from 1 node up to 512 nodes (average runtime of 1 hour) - Bebop (Broadwell): 6 (projects) x 10 (studies) x 1024 (1 to 512 nodes) x 36 (cores) x 1 (hour) = 2.2M - Bebop (KNL): 6 (projects) x 10 (studies) x 512 (1 to 256 nodes) x 64 (cores) x 1 (hour) = 2M - Blues: 6 (projects) x 20 (studies) x 512 (1 to 256 nodes) x 16 (cores) x 1 (hour) = 1M A small fraction of this would be used for debugging. Storage requirements: Persistent storage of the order of 100TB is required for our deep learning project. This project requires large datasets that take more than two weeks to generate. Recreating them every time on scratch space would be prohibitive; thus, the request for large storage space. Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov