Great! This time has been granted for radix (120K). -- John Roberts Argonne National Laboratory CELS Systems [email protected] On 2/10/17, 11:17 AM, "[email protected] on behalf of Bair, Raymond A." <[email protected] on behalf of [email protected]> wrote: Pavan says that 120K would be OK, so lets proceed with that. Ray ----------------------------------------- Ray Bair Argonne National Laboratory and the University of Chicago On 2/6/17, 7:50 AM, "[email protected] on behalf of [email protected]" <[email protected] on behalf of [email protected]> wrote: Hello, A change in allocation has been requested: Requester: balaji (Pavan Balaji) Project: radix Title: Scalable Parallel System Software Description: This allocation will be used for optimizations to various systems software components: (1) programming models and runtime systems, (2) data I/O and file systems, (3) fault tolerance, (4) operating systems, and (5) data analysis and visualization. Primary goal of the project is to perform research on improving the core systems software on high-end computing platforms. The radix allocation covers a large number of cross-cutting projects involving subsets of researchers in the group. A few core projects are presented here, but some of the techniques used are shared between them and are not easily separable (e.g., MPI improvements or resilience improvements): 1. Fault Tolerance Techniques: - Local memory management * Advanced SDC Detection through Data Monitoring + Adapt the prediction techniques used in the SZ lossy compressor for SDC detection * Checkpoint size reduction to improve checkpoint/restart performance and energy efficiency + Adapt FTI for using lossy compression of checkpoints * Impact of lossy compression on Checkpoint/restart + Investigate fundamentally and experimentally the relation between discretization/truncation error with lossy compression error * End-to-end Detection of SDCs in workflows featuring a lossy compression stage + Compare combination of SDC detectors and lossy compressors with respect to the accuracy of SDC detection. - Distributed memory management * Integration of checkpointing in workflow environments + Adapt DECAF and/or SWIFT with FTI and perform performance eveluation * MPI Fault Tolerance + Improvements to MPICH to deal with faults * Resilient-MPI and FTI Integration + Experimentation and evaluation of FTI coupled with resilient-MPI 2. Lightweight, topology-aware, and thread-safe communication: - Improvements to MPI+threads * MPICH improvements to play well with threads, including improved locking and lock-free data structures. * Communication progress improvements in the presence of threads. - Lightweight communication * Low instruction-count and low-overhead irregular data movement. - Topology-aware communication * Optimizations to MPI virtual topology functionality 3. Data I/O, RMA and Active Messages: - Dynamic Locality-Aware Workflows * Dynamically distribute the tasks among nodes using Asynchronous Dynamic Load Balancer - Development and Evaluation of Composable Data Services * Exploration of lightweight, customized, domain-specific storage services to improve the performance of data-intensive HPC applications - High Performance Parallel I/O * The ROMIO MPI-IO implementation provides the foundation for many components of the HPC I/O stack. Continue tuning, debugging, and research efforts. 4. Data analysis and visualization: Testing and benchmarking of improvements to data movement infrastructure such as DIY and Decaf software to prepare for next-generation machine architectures and to apply those improvements to applications across the lab that use those software platforms, such as cosmology, climate science, and materials science. The following areas will be tested on Fusion and Blues: - Relaxed BSP synchronization (e.g., for iterative algorithms, to overlap communication with computation and support irregular communication/computation patterns) * irregular communication patterns + short-circuit message completion - Load-balancing algorithms (work stealing, adaptive decompositions) for unstructured time-varying mesh partitions (e.g., reactor design) * work stealing algorithms + adaptive (e.g. k-d tree) domain decompositions - Integration with many-core thread models * Relaxed BSP communication with interchangeable thread models + Worker thread pool for executing block callback functions - Application domain decompositions * Add support for AMR (block-based and patch-based combustion and astrophysics codes) + Add support for directly using existing simulation decompositions (e.g., in climate codes) Current: undetermined amount Justification: 2M core hours requested: We have four core projects. All projects have been demonstrated at scale in the previous years and form the core of the system software stack used in production on the Fusion and Blues machines. We expect to run around ~40 studies for each project on Blues and on Fusion, scaling from 1 node to 256 nodes (average runtime of 1 hour) - Blues: 4 (projects) x 40 (studies) x 512 (1 to 256 nodes) x 16 (cores) x 1 (hour) = 1.3M - Fusion: 4 (projects) x 40 (studies) x 512 (1 to 256 nodes) x 8 (cores) x 1 (hour) = 655K A small fraction of this would be used for debugging. Requested: 200000 A specific reason has been given: We had several very productive students this year who made a lot of progress on their publications. We are still working on the papers, but seem to have run out of core hours. The additional core hours will help us complete the required work for our papers. This needs to be approved and the final allocation amount decided upon. Thank You, The LCRC Accounts System _______________________________________________ allocations-admins mailing list [email protected] https://lists.lcrc.anl.gov/mailman/listinfo/allocations-admins _______________________________________________ allocations-admins mailing list [email protected] https://lists.lcrc.anl.gov/mailman/listinfo/allocations-admins