Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Donald Sinclair Project Name: Lattice-QCD Division: HEP Project title: Lattice simulations of QCD at finite baryon number density Associated funding: DOE Other Systems: NERSC: Cori (Cray XC), Edison (Cray XC30): Allocation 7,000,000 MPP-hours TACC: Stampede (cluster): Allocation 259,149 CPU-hours, Stampede 2 (KNL cluster): Allocation 15,299 node-hours. PSC: Bridges (cluster): Allocation 648,710 CPU-hours. SDSC: Comet (cluster): Allocation 100,000 CPU-hours. Science: Our goal is to study the physics of nuclear matter from first principles (QCD) both at zero and at finite temperature. We aim to study the phase diagram of QCD in the baryon-/quark-number density -- temperature plane or equivalently in the quark-number chemical potential (mu) -- temperature plane. QCD at finite mu and zero temperature describes hadronic and nuclear matter. This probes the physics of neutron stars and possibly heavy nuclei. From an HEP point of view, it helps us understand the dynamics of QCD. Our interest here is to observe the transition from hadronic matter to nuclear matter expected to occur for a mu value a little below one third of the nucleon mass. Later we will search for exotic states of nuclear matter such as colour-superconducting states. At finite temperature, we will study the phase transition from hadronic/nuclear matter to a quark-gluon plasma. Here we will be interested to determine the existence and position of the critical endpoint where this transition is expected to change from a crossover to a first-order phase transition. The finite temperature behaviour of hadronic/nuclear matter is relevant to the physics probed by relativistic heavy-ion colliders (RHIC, FAIR...). In addition nuclear matter at high temperatures was present in the early universe. QCD at finite mu has a complex fermion determinant. Hence standard methods for simulating Lattice QCD, which rely on importance sampling, fail. One method which can be applied to such problems is the Complex Langevin Equation. Earlier attempts at applying this method to QCD at finite mu had been frustrated by runaway behaviour which adaptive updating was unable to control. Recently there has been progress in preventing such behaviour by use of appropriate gauge fixing (gauge cooling). We are now simulating Lattice QCD at finite mu using complex-Langevin simulations with gauge cooling. Complex-Langevin methods are only guaranteed to converge to the correct limiting distribution when the domain over which the fields evolve is compact and the drift term (force term) derived from the action is holomorphic in the fields. For QCD, the zeros in the fermion determinant give rise to poles in the drift term so that it is meromorphic, not holomorphic in the fields. Hence such simulations should be considered explorative until we determine the domain of validity (if any) of this method. As we have discussed in our results section, our simulations to date have indicated that, for the lattice sizes and parameters we use, our results show qualitative agreement with expectations but fail in the details, so further investigations are needed. Project description: A field theory in the Euclidean time regime can be quantized in terms of a partition function which is the functional integral over the c-number fields of the exponential of minus the action. In the Langevin equation, the fields evolve in a fictitious time such that the derivative of each field with respect to this time is minus the derivative of the action with respect to that field -- the drift term -- plus a gaussian distributed random number appropriately normalized. It can be shown that in the large 'time' limit the fields are distributed with the desired (Boltzmann) weight exp(-S), where S is the action. The Langevin equation can be extended to the case where S is complex, by replacing the real fields by complex fields. Here it can only be demonstrated that the time averages of observables give the correct expectation values, when the region traversed by the fields is compact and the drift term is holomorphic in the fields. When this Complex Langevin Equation (CLE) is applied to Lattice QCD, the gauge fields must be extended from SU(3) to SL(3,C) and the action involves the gauge fields and their inverses, but not their hermitian conjugates. The fermions are integrated out producing a power of the determinant of the Dirac operator. We obtain an effective action for the gauge fields by using the identity, determinant(M)^q=exp{q*ln[determinant(M)]}=exp{q*Trace[ln(M)]}. The trace is replaced by a stochastic estimator. The derivative with respect to the fields introduces an inverse of the Dirac operator M, which is evaluated using the conjugate gradient method. Unfortunately, since the determinant has zeros, the inverse of M and thus the drift term have poles, so that the drift term is not holomorphic but meromorphic in the fields. It is for this reason that, although the domain over which the fields evolve appears compact, convergence of the observables to their correct values is not guaranteed, and must be checked. We integrate the CLE numerically, inverting the Dirac operator using the conjugate gradient method. The fermion fields are placed on the sites of a 4-dimensional hypercubic lattice and the gauge fields on the links, making the action gauge invariant. If we label the sites by the 4-vector (x,y,z,t), we parallelize using MPI giving each task 1 t coordinate and nz adjacent z coordinates. The x and y coordinates are internal to each task. Each task then performs the same number of operations, and the communications are all nearest neighbour and homogeneous. In particular M only connects nearest-neighbour sites and is sparse. M is discretized using the staggered fermion method. Gauge cooling is implemented to minimize the average distance of the gauge fields from the SU(3) manifold by minimizing the sum over links of Tr{UU^dagger+(UU^dagger)^[-1]-2}, where U are the lattice gauge fields. Because we use 64-bit floating point (our previous simulations used 32-bit floating point except places where precision is critical), performance is not as good as is in our previous work. In addition, since we can anticipate algorithmic changes, we have not optimized our conjugate gradient routine as aggressively as before. We plan to do this, as soon as we are satisfied that we have an optimum version. We only have absolute (Gflop) performances for 12^4 lattices measured on NERSC computers. The best such performance was 3.0 Gflops/core on Cori. We assume that the best performances on NERSC machines for 16^4 lattices would be similar. On 64 cores the performance of the code on a 16^4 lattice on the 32 core Haswell nodes was approximately 10% slower than on Cori. On the 16 core Blues nodes, the performance was approximately 25% slower than on Cori. We are running on 128 cores on both Blues and Bebop. Bebop. For 128 cores the Broadwell nodes of Bebop are about 30% faster than the Haswell nodes of Blues and 70% faster than the Sandy Bridge nodes. For scaling studies on Bebop, along with discussion of issues pertaining to scaling, see the Large Allocation Efficiency section. We have performed runs on a 12^4 lattice at beta=6/g^2=5.6 (g is the lattice QCD bare coupling) and quark mass m=0.025, over a wide range of mu values. The results of these runs produced results in qualitative agreement with what was expected. However, there were quantitative deviations from our expectations. We are now running at a weaker coupling (beta=5.7) and the same mass on a 16^4 lattice, in addition to continuing our 12^4 runs. While there is some improvement, the beta=5.7 results are disappointingly similar to the beta=5.6 results. These 16^4 runs are being performed on Blues, Bebop, Cori at NERSC and Bridges at PSC. The 12^4 runs are continuing on Comet at SDSC. We are also running simulations of the phase-quenched theory (where we use only the magnitude of the fermion determinant) for comparison. The 12^4 phase-quenched runs are being performed on Edison at NERSC and the 16^4 runs have been moved from Cori at NERSC to Bebop. We intend to complete the 12^4 and 16^4 runs on Bebop. This is important since the mu values we have yet to complete are in the large-mu domain, where we might expect the complex Langevin simulations to produce correct results. We also intend to complete our 16^4 phase-quenched runs on Bebop in FY2018. Our complex-Langevin simulations on Bebop currently take 48 hours on 4 Broadwell nodes while our phase-quenched runs take 24 hours on 8 Broadwell nodes. We are currently running 1 complex-Langevin and 2 phase-quenched runs on Bebop and 1 complex-Langevin run on Blues. When we lose access to Blues we will move those runs to Bebop. This means we will initially be using 24 Bebop nodes. This will be reduced when the phase-quenched runs finish at which point we plan to start 12^4 complex-Langevin runs each using a single (or at most 2) Bebop Broadwell nodes. Larger, 32^4 runs at beta=5.8,5.9 have started on Cori at NERSC and Stampede 2 at TACC. These projects are expected to finish in the middle of the year, to be replaced by finite temperature runs on 12^3*6 lattices, using 1 or 2 Broadwell nodes. We, however, will consider moving other projects to Bebop if this will lead to faster throughput. Included in these are projects where the codes have larger arithmetic intensity, which could make them suitable for running on the KNL nodes. When we are convinced that we have the code we wish to run in extensive production, we will apply optimizations that have proved effective with other codes, to the conjugate gradient routine where the jobs spend most of their time. Our project currently has 2 members only one of whom (me) will require access to Bebop. We estimate that the completion of the 16^4 complex-Langevin simulations will require 6200 wallclock-hours on 4 Broadwell nodes for completion. The 16^4 phase-quenched runs will require an estimated 3000 wallclock-hours on 8 Broadwell nodes to complete. The 12^4 complex-Langevin simulations will need a further 7500 wallclock-hours on a single Broadwell node. We are allotting 4000 hours (2 quarters) on 4 Broadwell nodes for the finite temperature (12^3*6) runs, for an estimated 2,500,000 core-hours for FY2018. We anticipate using 16 Broadwell nodes concurrently in the first quarter reducing this to 4 nodes by the end of FY2018. Industry partnership: Project URL: Current FY Hours Used: undetermined amount New FY Requested allocation: 2500000 Q1: 850000 Q2: 750000 Q3: 500000 Q4: 400000 Justification: The following are tables of performances of our complex-Langevin code on Blues. They each report the times taken for identical runs on different numbers of cores on a 16^4 lattice: This first table is of wallclock times for the 16-core nodes of Blues. 16 cores = 129m35.228s 32 cores = 52m37.305s 64 cores = 25m35.144s 128 cores = 17m55.801s 256 cores = 16m40.058s This second table is of the identical benchmark for the 32-core (Haswell) nodes of Blues. 32 cores haswell = 51m44.125s 64 cores haswell = 21m41.861s 128 cores haswell = 13m3.596s 256 cores haswell = 9m43.317s The following is a table of the identical benchmark for the Broadwell nodes of Bebop. 32 cores broadwell = 44m11.403s 64 cores broadwell = 19m 7.393s 128 cores broadwell = 9m46.931s 256 cores broadwell = 6m12.303s (Note here we used only 16 of the 18 cores per socket.) Finally we present benchmarks for the phase-quenched (RHMC) code on the Broadwell nodes of Bebop for a 16^4 lattice. 32 cores = 133m33.358s 64 cores = 67m28.117s 128 cores = 41m46.570s 256 cores = 29m17.960s We note that the complex-Langevin code shows good scaling properties, at least up to 128 cores. On 256 cores we are seeing the effect of a decrease in the ratio of number of floating point operations to the MPI message length and of having too little computation between MPI communication calls. Note that we are running with 1 MPI task/core. Unfortunately we have yet to find a good performance monitor to yield Gflop rates for these Intel-based machines. Storage requirements: 1 TB Thank You, The LCRC Accounts System