[LCRC Accounts] Yearly Allocation Request for JDFT-Surfaces
Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Kendra Letchworth-Weaver Project Name: JDFT-Surfaces Division: NST Project title: Joint Density-Functional Theory Investigations of Realistic Surface Structure Associated funding: LDRD/Named Fellow Other Systems: Carbon at Center for Nanoscale Materials (375,000 hours) Science: The surface structure of a surface or nanoparticle in contact with liquid is often different from the structure in vacuum. This structure may be probed experimentally by techniques such as X-ray reflectivity or pair distribution functions, however the results often require interpretation and validation. Using Joint-Density-Functional Theory, which enables quantum-mechanical calculations of a solute system (such as an surface) in contact with a liquid environment, we will make predictions of surface structure in liquid which will provide unique insights into related experimental studies. By avoiding the statistical sampling over configurations of the liquid required by classical and ab initio MD, JDFT enables rapid screening of surface terminations in liquid environment for agreement with experimental data. In addition, JDFT naturally provides surface energies in solution and as a function of voltage, allowing us to synthesize experimental and theoretical informat ion to determine surface structure under realistic electrochemical and growth conditions (as in doi:10.1021/jacs.6b03338). Project description: We plan to consider several systems within this project: 1. FY2017 The surface structure of geological materials in water, such as alumina or hematite, for direct comparison to X-ray reflectivity data from the group of Paul Fenter in CSE. We will perform AIMD to sample the surface alone while treating the liquid through JDFT. For alumina and hematite each, we will consider roughly 5 terminations in AIMD, requiring 10 ps of statistics for each. Parallelized over 9 NVidia C2075 GPUs, 1 ps of statistics for these surfaces requires roughly 24 hours of walltime. We expect that parallelizing over 6 faster Tesla K40 GPUs will yield the same walltime (2 materials x 5 terminations x 10 ps x 24 hours x 6 GPUs x 32 core hours=460800 hours) FY2018 We encountered some challenges with the simulations we planned to do in FY2017. First, we determined that in order to provide sufficiently accurate X-ray Reflectivity signals, the system sizes required were not feasible to study on gpu. We also determined that additional full AIMD calculations are required to benchmark the accuracy of JDFT for the liquid at the interface. Such calculations are not feasible within JDFTx, however they are possible in highly parallelizable Qbox. We request 5 terminations*2 functionals* 10 ps *1000 cores*2 hours/ps= 200,000 core hours on KNL to run AIMD simulations of the Al2O3 surface in water, using both van der Waals (optB88) and PBE functionals 2. FY2017 The surface structure of SrTiO3 during growth of Ruddlesden Popper stacking faults, for comparison to resonant X-ray reflectivity data from the group of Dillon Fong in MSD. Though there is no liquid present, these measurements were taken at high temperatures so AIMD is also required to account for thermal motion of the surface. The number of calculations required is similar to the values for the geological materials above. (1 material x 5 terminations x 10 ps x 24 hours x 6 GPUs x 32 core hours= 230400 hours) FY2018 We have concluded the study of homoepitaxial growth of SrTiO3, which will be submitted to Advanced Materials this Fall. Now, we turn to the heteroepitaxial growth of LaTiO3 on a SrTiO3 substrate, in an attempt to understand the different timescales of layer rearrangement experimentally observed on LTO deposition vs STO deposition. Here, the smaller unit cells were sufficient to run on the GPU. (1 material x 5 terminations x 10 ps x 24 hours x 6 GPUs x 32 core hours= 230400 hours) 3. FY2017 How the shape of IrO2 nanoclusters for solar water-splitting applications changes in the presence of liquid (in collaboration with Theory and Modeling group at CNM). We will also consider how the binding energy of oxygen at different adsorption sites on the nanoparticle changes when liquid is included. The initial low-energy IrO2 nanoparticle shapes will be determined by a genetic algorithm search using force fields to represent the atomic interactions. We will then consider approximately 10 low energy nanoclusters in liquid using JDFT. Finally, we will study approximately 10 oxygen binding sites on the lowest-energy nanocluster in solution. These calculations are not currently MPI parallelizable (though they are OpenMP and GPU parallelized to run on a single node and MPI parallelization may be implemented before the end of the allocation). These geometry optimization calculations will each take 1 week of walltime to run in vacuum and another week in liquid. (2 solvation states x 20 calculations x 7 days x 24 hours x 32 core hours=215040 hours) FY2018 These calculations proved to be unreasonably slow to converge on the older hardware at LCRC (many degrees of freedom required >100 ionic steps to minimize and they required too much memory to fit on the GPU). For this reason, much of the time requested in FY2017 went unused These calculations should run 10x faster on new KNL nodes, making them now feasible. (2 solvation states x 20 calculations x 7 days x 24 hours x 32 core hours=215040 hours) FY2017 The total production run hours above add to ~900,000 hours. FY2018 The total production run hours above add to ~650,000 hours We estimate using roughly an additional 150,000 hours (~20% of the requested allocation) for debugging and convergence tests. We will use the electronic structure code JDFTx, an open-source plane-wave DFT code co-developed by the PI and written in C++. It is available for download at http://jdftx.org. We require the following packages to be installed: • cmake (>=2.8) • g++ (>=4.6) • libgsl0-dev • libopenmpi-dev (or equivalent for other MPI distributions) • libfftw3-dev • libatlas-base-dev (provides blas and cblas) • liblapack-dev (alternatively, math operations using Intel’s MKL are possible) See http://jdftx.org/Compiling.html) for details. We will also need an older version of Qbox which supports OPTB88 van der Waals functionals to be installed and compiled. I will be working in collaboration with Maria Chan as well as my group members Fatih Sen and Arun Mannodi on this project for 4 team members in total. Industry partnership: Project URL: http://jdftx.org Current FY Hours Used: undetermined amount New FY Requested allocation: 800000 Q1: 200000 Q2: 200000 Q3: 200000 Q4: 200000 Justification: JDFTx can be run in either a hybrid GPU-MPI mode for smaller memory jobs or an OpenMP-MPI mode for larger memory jobs, offering excellent performance and scaling. GPU parallelization offers a speedup of 3-5x over a single CPU. The JDFT liquid experiences near-linear scaling when MPI-parallelized over orientations (typically 24-144 orientations are required) The electronic DFT experiences near-linear scaling when MPI-parallelized over k-points. Coulomb truncation algorithms implemented in the code also reduce the size of the simulation grid for JDFTx wavefunctions. As an example, scaling data for a small Al2O3 surface supercell (34 atoms, 9 k points) on NERSC Edison is provided below in vacuum. For the vacuum calculations, we perform initialization and a single electronic wavefunction minimization. Each MPI process was run with 12 OpenMP threads (2 processes per 24-core node). MPI Processes time 1 449 s 2 249 s 3 165 s 5 114 s 9 71 s The large allocation is requested because GPU calculations are charged double compared to CPU calculations, yet we will be able to utilize MPI parallelization over GPU’s. Alternatively, if the GPU is not available, we could MPI parallelize over CPU’s, though we would require the additional factor of 3-5 in walltime. JDFTx does run in parallel on KNL nodes (verified on NERSC Cori). Further optimization and benchmarking of the JDFTx software to run in hybrid MPI/threaded mode for KNL nodes is in progress. The ability to run on the KNL nodes will increase the system sizes available for study with JDFTx, yet will continue to run at speeds comparable or better than the GPU. Storage requirements: 1 TB Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov