[LCRC Accounts] Project Request: hedm_midas
Hello, A new project on the LCRC cluster has been requested. Please forward the information on to the LCRC Allocation sub-committee. Applicant's name: Hemant Sharma Applicant's institution: ANL Applicant's division: XSD Project Name: hedm_midas Project title: Analysis of High Energy Diffraction Microscopy (HEDM) Data Associated funding: Other Systems: Local cluster at APS (Orthros) with 500 dedicated cores; NERSC (currently on descertionary time; ALCC proposal in preparation) Science: High Energy Diffraction Microscopy (HEDM) is used for non-destructive in-situ characterization of polycrystalline materials during thermo-mechanical treatments. HEDM has successfully been used to study a wide range of materials including: sand particles, nuclear materials, solder joints in electronics and turbine blades in jet engines. HEDM yields information about the position, crystallographic orientation and strain state of all crystallites (grains) in a surveyed region, which can be used to study material response down to the micrometer level. There are two modes of HEDM: near-field (NF) and far-field (FF). Crystallographic orientation within the material with a resolution of ~ 2 um can be obtained using NF-HEDM. FF-HEDM provides the strain state of grains by using a different acquisition mode and detector. The results from both NF-HEDM and FF-HEDM can be combined to get complete orientation and strain information for the material. The HEDM micro scope at 1-ID, Sector 1, APS, is specially tailored for HEDM experiments and currently generates up to 1TB of data per day. Project description: The MIDAS software package, developed at the APS, is highly parallelizable with low RAM/core requirement. The software is written in C99 and parallelization is achieved using SWIFT. The software is distributed using github and is available at https://github.com/marinerhemant/MIDAS FF-HEDM data reduction involves the following steps: the raw data stored at the local APS data server is first transferred to the compute machine file system. All the cores share a single copy of the data using shared memory in built in Linux. The raw data are then pre-processed to get information about the diffraction peaks from different grains using a peak fitting routine. Fitting of each peak can be run in parallel. This step takes between 1 min to 30 min using 500 cores, depending on the complexity of the data. The processed peaks are then fed into an indexing routine which assigns peaks to grains and optimizes the grain position, crystallographic orientation and strain. Analysis for each grain can be run in parallel and total execution time for ~2000 grains is ~2 min using 500 cores. RAM/core usage at any stage does not exceed ~200 MB. This full workflow has been tested on Bebop and we get 95% speedup up to 50 nodes. The final output from a starting raw dataset of ~15GB size is 30MB and is copied back to the beamline machine. Reduction of NF-HEDM data progresses as follows: once the raw data has been transferred to the compute machine, each diffraction image is processed in parallel to extract information regarding the detector pixels with positive intensity. This step reduces the dataset from ~15GB to ~200MB. A C-code then optimizes the crystallographic orientation at each position in a voxelized grid in the material. This step can be run in parallel for each voxel and it can be sped up further by using the FF-HEDM output as a starting estimate. Total execution time is ~3 min using 500 cores. By utilizing shared memory, RAM/core usage stays below ~200MB. The final output is 50MB and is copied back to the beamline machine. Initial runs on Bebop show ~90% speedup up to 50 nodes. A single dataset for a material state can be acquired in ~15 min (FF-HEDM) and ~40 min (NF-HEDM). The APS wishes to provide 1-ID users both a near-real-time analysis capability as well as a mechanism for users to reprocess their analyses after data collection, where parameters may be tailored. At present, beamline data are transferred to NERSC for automatic processing, where ~50 nodes on Edison are almost always with immediate access; this allows results to be obtained within a few minutes of data collection. We plan to use LCRC in collaboration with select users for off-line reprocessing of data, where computational throughput demands are not so stringent. We are also working with LCRC to implement on-demand scheduling and may utilize LCRC for near-real-time data analysis, should that capability be available in production. Industry partnership: Our users include GE Global Research, Air Force Research Laboratory and a number of universities and research institutes. Project URL: Requested allocation: 1000000 Q1: 250000 Q2: 250000 Q3: 250000 Q4: 250000 Justification: We have been running MIDAS on Orthros, Bebop, Blues, NERSC (Edison, Cori). MIDAS has been shown to efficiently utilize 1000s of cores on these machines. There are no active plans to port MIDAS to KNL processors. Storage requirements: A single experiment is ~2TB in size. A storage allocation of 20TB is requested so that several experiments may be processed at a single time. The requester has used undetermined amount hours of their initial startup project. In addition to approving an initial amount, please specify a Category and Subcategory for this project. For a list of the current selection of approved categories, please see: https://wiki.lcrc.anl.gov/wiki/Processes/Categories Once the Allocation committee has approved the project, please go to the Project Management page to create it: https://accounts.lcrc.anl.gov/projects.php Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov