Re: [allocations-admins] [LCRC Accounts] Project Allocation Request
Ray, Blues looks oversubscribed because projects with large allocations use up their allocations early in the quarter and get additional allocations to keep their research going. I think we should consider adding to allocations and expiring unused allocations on a monthly rather quarterly basis. Users won’t have to wait long (less than a month) for more time, if their allocations are used up early in the month. John J. Low Principal Computational Science Specialist Computing, Environment and Life Sciences Building 240, 2143 9700 South Cass Avenue Argonne National Laboratory Argonne, IL 60439. 630-252-0045 www.linkedin.com/pub/john-low/15/8b0/5aa/ -----Original Message----- From: allocations-admins <[email protected]> on behalf of Halim Amer <[email protected]> Reply-To: LCRC Allocations Admins <[email protected]> Date: Monday, July 10, 2017 at 12:03 PM To: "Raymond A. Bair" <[email protected]>, "Balaji, Pavan" <[email protected]> Cc: LCRC Allocations Admins <[email protected]> Subject: Re: [allocations-admins] [LCRC Accounts] Project Allocation Request Hi Ray, Thank you for considering our request. We think 150k would get us going the upcoming month, so let's do as you suggested. Best, Halim www.mcs.anl.gov/~aamer On 7/10/17 11:40 AM, Bair, Raymond A. wrote: > Dear Pavan and Halim, > > I propose that we add enough time to carry the radix project for the > next 30 days. By that time Bebop should be available, and we can > provide more. Would 150K core-hours suffice? > > Regards, > > Ray > > > ----------------------------------------- > Ray Bair > Argonne National Laboratory > > > On 7/10/17, 11:33 AM, "allocations-admins on behalf of > [email protected]" <[email protected] on > behalf of [email protected]> wrote: > > Hello, > > A change in allocation has been requested: > > Requester: aamer (Halim Amer) > Project: radix > Title: Scalable Parallel System Software > Description: This allocation will be used for optimizations to various > systems > software components: (1) programming models and runtime systems, (2) > data I/O and file systems, (3) fault tolerance, (4) operating systems, > and (5) data analysis and visualization. > > Primary goal of the project is to perform research on improving the > core systems software on high-end computing platforms. > > The radix allocation covers a large number of cross-cutting projects > involving subsets of researchers in the group. A few core projects > are presented here, but some of the techniques used are shared between > them and are not easily separable (e.g., MPI improvements or > resilience improvements): > > 1. Fault Tolerance Techniques: > > - Local memory management > > * Advanced SDC Detection through Data Monitoring > > + Adapt the prediction techniques used in the SZ lossy compressor > for SDC detection > > * Checkpoint size reduction to improve checkpoint/restart > performance and energy efficiency > > + Adapt FTI for using lossy compression of checkpoints > > * Impact of lossy compression on Checkpoint/restart > > + Investigate fundamentally and experimentally the relation > between discretization/truncation error with lossy compression > error > > * End-to-end Detection of SDCs in workflows featuring a lossy > compression stage > > + Compare combination of SDC detectors and lossy compressors with > respect to the accuracy of SDC detection. > > - Distributed memory management > > > * Integration of checkpointing in workflow environments > > + Adapt DECAF and/or SWIFT with FTI and perform performance > eveluation > > * MPI Fault Tolerance > > + Improvements to MPICH to deal with faults > > * Resilient-MPI and FTI Integration > > + Experimentation and evaluation of FTI coupled with > resilient-MPI > > 2. Lightweight, topology-aware, and thread-safe communication: > > - Improvements to MPI+threads > > * MPICH improvements to play well with threads, including improved > locking and lock-free data structures. > > * Communication progress improvements in the presence of threads. > > - Lightweight communication > > * Low instruction-count and low-overhead irregular data movement. > > - Topology-aware communication > > * Optimizations to MPI virtual topology functionality > > 3. Data I/O, RMA and Active Messages: > > - Dynamic Locality-Aware Workflows > > * Dynamically distribute the tasks among nodes using Asynchronous > Dynamic Load Balancer > > - Development and Evaluation of Composable Data Services > > * Exploration of lightweight, customized, domain-specific storage > services to improve the performance of data-intensive HPC > applications > > - High Performance Parallel I/O > > * The ROMIO MPI-IO implementation provides the foundation for many > components of the HPC I/O stack. Continue tuning, debugging, > and research efforts. > > 4. Data analysis and visualization: Testing and benchmarking of > improvements to data movement infrastructure such as DIY and Decaf > software to prepare for next-generation machine architectures and to > apply those improvements to applications across the lab that use those > software platforms, such as cosmology, climate science, and materials > science. The following areas will be tested on Fusion and Blues: > > - Relaxed BSP synchronization (e.g., for iterative algorithms, to > overlap communication with computation and support irregular > communication/computation patterns) > > * irregular communication patterns > > + short-circuit message completion > > - Load-balancing algorithms (work stealing, adaptive decompositions) > for unstructured time-varying mesh partitions (e.g., reactor > design) > > * work stealing algorithms > > + adaptive (e.g. k-d tree) domain decompositions > > - Integration with many-core thread models > > * Relaxed BSP communication with interchangeable thread models > > + Worker thread pool for executing block callback functions > > - Application domain decompositions > > * Add support for AMR (block-based and patch-based combustion and > astrophysics codes) > > + Add support for directly using existing simulation > decompositions (e.g., in climate codes) > Current: undetermined amount > Justification: 2M core hours requested: > > We have four core projects. All projects have been demonstrated at > scale in the previous years and form the core of the system software > stack used in production on the Fusion and Blues machines. > > We expect to run around ~40 studies for each project on Blues and on > Fusion, scaling from 1 node to 256 nodes (average runtime of 1 hour) > > - Blues: 4 (projects) x 40 (studies) x 512 (1 to 256 nodes) x > 16 (cores) x 1 (hour) = 1.3M > > - Fusion: 4 (projects) x 40 (studies) x 512 (1 to 256 nodes) x > 8 (cores) x 1 (hour) = 655K > > A small fraction of this would be used for debugging. > > Requested: 450000 > > A specific reason has been given: > Some of our projects consumed core hours at a higher rate than > anticipated. We also had a 300K core hours cut in the allocation we > anticipated. These factors contributed to this shortage we are > experiencing. Some of our important projects, such as MPI, critically > need this additional allocation. > > This needs to be approved and the final allocation amount decided upon. > > Thank You, > The LCRC Accounts System > _______________________________________________ > allocations-admins mailing list > [email protected] > https://lists.lcrc.anl.gov/mailman/listinfo/allocations-admins > > _______________________________________________ allocations-admins mailing list [email protected] https://lists.lcrc.anl.gov/mailman/listinfo/allocations-admins
participants (1)
-
Low, John J.