Hello,
A yearly allocation for the LCRC cluster has been requested with the
following updated information:
Submitter/PI: Donald Sinclair
Project Name: Lattice-QCD
Division: HEP
Project title: Lattice simulations of QCD at finite baryon number density
Associated funding: DOE
Other Systems: NERSC: Cori (Cray XC), Edison (Cray XC30): Allocation 7,000,000 MPP-hours
TACC: Stampede (cluster): Allocation 259,149 CPU-hours,
Stampede 2 (KNL cluster): Allocation 15,299 node-hours.
PSC: Bridges (cluster): Allocation 648,710 CPU-hours.
SDSC: Comet (cluster): Allocation 100,000 CPU-hours.
Science: Our goal is to study the physics of nuclear matter from first principles (QCD)
both at zero and at finite temperature. We aim to study the phase diagram of
QCD in the baryon-/quark-number density -- temperature plane or equivalently in
the quark-number chemical potential (mu) -- temperature plane. QCD at finite mu
and zero temperature describes hadronic and nuclear matter. This probes the
physics of neutron stars and possibly heavy nuclei. From an HEP point of view,
it helps us understand the dynamics of QCD. Our interest here is to observe
the transition from hadronic matter to nuclear matter expected to occur for a
mu value a little below one third of the nucleon mass. Later we will search
for exotic states of nuclear matter such as colour-superconducting states. At
finite temperature, we will study the phase transition from hadronic/nuclear
matter to a quark-gluon plasma. Here we will be interested to determine the
existence and position of the critical endpoint where this transition is
expected to change from a crossover to a first-order phase transition. The
finite temperature behaviour of hadronic/nuclear matter is relevant to the
physics probed by relativistic heavy-ion colliders (RHIC, FAIR...). In
addition nuclear matter at high temperatures was present in the early universe.
QCD at finite mu has a complex fermion determinant. Hence standard methods for
simulating Lattice QCD, which rely on importance sampling, fail. One method
which can be applied to such problems is the Complex Langevin Equation. Earlier
attempts at applying this method to QCD at finite mu had been frustrated by
runaway behaviour which adaptive updating was unable to control. Recently there
has been progress in preventing such behaviour by use of appropriate gauge
fixing (gauge cooling). We are now simulating Lattice QCD at finite mu using
complex-Langevin simulations with gauge cooling. Complex-Langevin methods are
only guaranteed to converge to the correct limiting distribution when the
domain over which the fields evolve is compact and the drift term (force term)
derived from the action is holomorphic in the fields. For QCD, the zeros in
the fermion determinant give rise to poles in the drift term so that it is
meromorphic, not holomorphic in the fields. Hence such simulations should be
considered explorative until we determine the domain of validity (if any) of
this method. As we have discussed in our results section, our simulations to
date have indicated that, for the lattice sizes and parameters we use, our
results show qualitative agreement with expectations but fail in the details,
so further investigations are needed.
Project description: A field theory in the Euclidean time regime can be quantized in terms of a
partition function which is the functional integral over the c-number fields of
the exponential of minus the action. In the Langevin equation, the fields evolve
in a fictitious time such that the derivative of each field with respect to
this time is minus the derivative of the action with respect to that field
-- the drift term -- plus a gaussian distributed random number appropriately
normalized. It can be shown that in the large 'time' limit the fields are
distributed with the desired (Boltzmann) weight exp(-S), where S is the action.
The Langevin equation can be extended to the case where S is complex, by
replacing the real fields by complex fields. Here it can only be demonstrated
that the time averages of observables give the correct expectation values,
when the region traversed by the fields is compact and the drift term is
holomorphic in the fields. When this Complex Langevin Equation (CLE) is
applied to Lattice QCD, the gauge fields must be extended from SU(3) to
SL(3,C) and the action involves the gauge fields and their inverses, but not
their hermitian conjugates.
The fermions are integrated out producing a power of the determinant of the
Dirac operator. We obtain an effective action for the gauge fields by using the
identity, determinant(M)^q=exp{q*ln[determinant(M)]}=exp{q*Trace[ln(M)]}. The
trace is replaced by a stochastic estimator. The derivative with respect to
the fields introduces an inverse of the Dirac operator M, which is evaluated
using the conjugate gradient method. Unfortunately, since the determinant has
zeros, the inverse of M and thus the drift term have poles, so that the drift
term is not holomorphic but meromorphic in the fields. It is for this reason
that, although the domain over which the fields evolve appears compact,
convergence of the observables to their correct values is not guaranteed, and
must be checked.
We integrate the CLE numerically, inverting the Dirac operator using the
conjugate gradient method. The fermion fields are placed on the sites of a
4-dimensional hypercubic lattice and the gauge fields on the links, making the
action gauge invariant. If we label the sites by the 4-vector (x,y,z,t), we
parallelize using MPI giving each task 1 t coordinate and nz adjacent z
coordinates. The x and y coordinates are internal to each task. Each task then
performs the same number of operations, and the communications are all nearest
neighbour and homogeneous. In particular M only connects nearest-neighbour sites
and is sparse. M is discretized using the staggered fermion method. Gauge
cooling is implemented to minimize the average distance of the gauge fields
from the SU(3) manifold by minimizing the sum over links of
Tr{UU^dagger+(UU^dagger)^[-1]-2}, where U are the lattice gauge fields.
Because we use 64-bit floating point (our previous simulations used 32-bit
floating point except places where precision is critical), performance is not as
good as is in our previous work. In addition, since we can anticipate
algorithmic changes, we have not optimized our conjugate gradient routine as
aggressively as before. We plan to do this, as soon as we are satisfied that we
have an optimum version. We only have absolute (Gflop) performances for 12^4
lattices measured on NERSC computers. The best such performance was
3.0 Gflops/core on Cori. We assume that the best performances on NERSC machines
for 16^4 lattices would be similar. On 64 cores the performance of the code on
a 16^4 lattice on the 32 core Haswell nodes was approximately 10% slower than
on Cori. On the 16 core Blues nodes, the performance was approximately 25%
slower than on Cori. We are running on 128 cores on both Blues and Bebop.
Bebop. For 128 cores the Broadwell nodes of Bebop are about 30% faster than
the Haswell nodes of Blues and 70% faster than the Sandy Bridge nodes. For
scaling studies on Bebop, along with discussion of issues pertaining to
scaling, see the Large Allocation Efficiency section.
We have performed runs on a 12^4 lattice at beta=6/g^2=5.6 (g is the lattice QCD
bare coupling) and quark mass m=0.025, over a wide range of mu values. The
results of these runs produced results in qualitative agreement with what was
expected. However, there were quantitative deviations from our expectations.
We are now running at a weaker coupling (beta=5.7) and the same mass on a
16^4 lattice, in addition to continuing our 12^4 runs. While there is some
improvement, the beta=5.7 results are disappointingly similar to the beta=5.6
results. These 16^4 runs are being performed on Blues, Bebop, Cori at NERSC
and Bridges at PSC. The 12^4 runs are continuing on Comet at SDSC. We are also
running simulations of the phase-quenched theory (where we use only the
magnitude of the fermion determinant) for comparison. The 12^4 phase-quenched
runs are being performed on Edison at NERSC and the 16^4 runs have been
moved from Cori at NERSC to Bebop.
We intend to complete the 12^4 and 16^4 runs on Bebop. This is important since
the mu values we have yet to complete are in the large-mu domain, where we
might expect the complex Langevin simulations to produce correct results. We
also intend to complete our 16^4 phase-quenched runs on Bebop in FY2018.
Our complex-Langevin simulations on Bebop currently take 48 hours on 4
Broadwell nodes while our phase-quenched runs take 24 hours on 8 Broadwell
nodes. We are currently running 1 complex-Langevin and 2 phase-quenched
runs on Bebop and 1 complex-Langevin run on Blues. When we lose access to
Blues we will move those runs to Bebop. This means we will initially be using
24 Bebop nodes. This will be reduced when the phase-quenched runs finish at
which point we plan to start 12^4 complex-Langevin runs each using a single
(or at most 2) Bebop Broadwell nodes. Larger, 32^4 runs at beta=5.8,5.9 have
started on Cori at NERSC and Stampede 2 at TACC.
These projects are expected to finish in the middle of the year, to be replaced
by finite temperature runs on 12^3*6 lattices, using 1 or 2 Broadwell nodes.
We, however, will consider moving other projects to Bebop if this will lead
to faster throughput. Included in these are projects where the codes have larger
arithmetic intensity, which could make them suitable for running on the KNL
nodes.
When we are convinced that we have the code we wish to run in extensive
production, we will apply optimizations that have proved effective with other
codes, to the conjugate gradient routine where the jobs spend most of their
time.
Our project currently has 2 members only one of whom (me) will require access
to Bebop.
We estimate that the completion of the 16^4 complex-Langevin simulations will
require 6200 wallclock-hours on 4 Broadwell nodes for completion. The 16^4
phase-quenched runs will require an estimated 3000 wallclock-hours on 8
Broadwell nodes to complete. The 12^4 complex-Langevin simulations will need
a further 7500 wallclock-hours on a single Broadwell node. We are allotting
4000 hours (2 quarters) on 4 Broadwell nodes for the finite temperature
(12^3*6) runs, for an estimated 2,500,000 core-hours for FY2018. We anticipate
using 16 Broadwell nodes concurrently in the first quarter reducing this to 4
nodes by the end of FY2018.
Industry partnership:
Project URL:
Current FY Hours Used: undetermined amount
New FY Requested allocation: 2500000
Q1: 850000
Q2: 750000
Q3: 500000
Q4: 400000
Justification: The following are tables of performances of our complex-Langevin code on Blues.
They each report the times taken for identical runs on different numbers of
cores on a 16^4 lattice:
This first table is of wallclock times for the 16-core nodes of Blues.
16 cores = 129m35.228s
32 cores = 52m37.305s
64 cores = 25m35.144s
128 cores = 17m55.801s
256 cores = 16m40.058s
This second table is of the identical benchmark for the 32-core (Haswell) nodes
of Blues.
32 cores haswell = 51m44.125s
64 cores haswell = 21m41.861s
128 cores haswell = 13m3.596s
256 cores haswell = 9m43.317s
The following is a table of the identical benchmark for the Broadwell nodes of
Bebop.
32 cores broadwell = 44m11.403s
64 cores broadwell = 19m 7.393s
128 cores broadwell = 9m46.931s
256 cores broadwell = 6m12.303s
(Note here we used only 16 of the 18 cores per socket.)
Finally we present benchmarks for the phase-quenched (RHMC) code on the
Broadwell nodes of Bebop for a 16^4 lattice.
32 cores = 133m33.358s
64 cores = 67m28.117s
128 cores = 41m46.570s
256 cores = 29m17.960s
We note that the complex-Langevin code shows good scaling properties, at least
up to 128 cores. On 256 cores we are seeing the effect of a decrease in the
ratio of number of floating point operations to the MPI message length and of
having too little computation between MPI communication calls. Note that we
are running with 1 MPI task/core.
Unfortunately we have yet to find a good performance monitor to yield Gflop
rates for these Intel-based machines.
Storage requirements: 1 TB
Thank You,
The LCRC Accounts System