[LCRC Accounts] Project Allocation Request
Hello, A change in allocation has been requested: Requester: sinclair (Donald Sinclair) Project: Lattice-QCD Title: Lattice simulations of QCD at finite baryon number density Description: A field theory in the Euclidean time regime can be quantized in terms of a partition function which is the functional integral over the c-number fields of the exponential of minus the action. In the Langevin equation, the fields evolve in a fictitious time such that the derivative of each field with respect to this time is minus the derivative of the action with respect to that field -- the drift term -- plus a gaussian distributed random number appropriately normalized. It can be shown that in the large 'time' limit the fields are distributed with weight exp(-S), where S is the action. The Langevin equation can be extended to the case where S is complex, by replacing the real fields by complex fields. Here it can only be demonstrated that the time averages of observables give the correct expectation values, when the region traversed by the fields is compact and S is holomorphic in the fields. When this Complex Langevin Equation (CLE) is applied to Lattice QCD, the gauge fields must be extended from SU(3) to SL(3,C) and the action involves the gauge fields and their inverses, but not their hermitian conjugates. The fermions are integrated out producing a power of the determinant of the Dirac operator. We obtain an effective action for the gauge fields by using the identity, determinant(M)^q=exp{q*ln[determinant(M)]} =exp{q*Trace[ln(M)]}. The trace is replaced by a stochastic estimator. The derivative with respect to the fields introduces an inverse of the Dirac operator M, which is evaluated using the conjugate gradient method. Unfortunately, since the determinant has zeros, the inverse of M and thus the drift term have poles, so that the drift term is not holomorphic but meromorphic in the fields. It is for this reason that, although the domain over which the fields evolve appears compact, convergence of the observables to their correct values is not guaranteed, and must be checked. We integrate the CLE numerically, inverting the Dirac operator using the conjugate gradient method. The fermion fields are placed on the sites of a 4-dimensional hypercubic lattice and the gauge fields on the links, making the action gauge invariant. If we label the sites by the 4-vector (x,y,z,t), we parallelize using MPI giving each task 1 t coordinate and nz adjacent z coordinates. The x and y coordinates are internal to each task. Each task then performs the same number of operations, and the communications are all nearest neighbour and homogeneous. In particular M only connects nearest-neighbour sites and is sparse. M is discretized using the staggered fermion method. Gauge cooling is implemented to minimize the sum over links of Tr{UU^dagger+(UU^dagger)^[-1]-2}, where U are the lattice gauge fields. Because we use 64-bit floating point (our previous simulations used 32-bit floating point except places where precision is critical), scaling is not as good as is in our previous work. In addition, since we can anticipate algorithmic changes, we have not optimized our conjugate gradient routine as aggressively as before. We plan to do this, as soon as we are satisfied that we have an optimum version. We only have absolute (Gflop) performances for 12^4 lattices measured on NERSC computers. The best such performance was 3.0 Gflops/core on Cori. We assume that the best performances on NERSC machines for 16^4 lattices would be similar. On 64 cores the performance of the code on a 16^4 lattice on the 32 core Haswell nodes of Blues was approximately 10% slower than on Cori. On the 16 core Blues nodes, the performance was approximately 25% slower than on Cori. For scaling studies on Blues, see the Large Allocation Efficiency section. We have performed runs on a 12^4 lattice at beta=6/g^2=5.6 (g is the lattice QCD bare coupling) and quark mass m=0.025, over a wide range of mu values. The results of these runs are in qualitative agreement with what is expected. However, there are quantitative deviations from our expectations. Running with the same beta and m values on a 16^4 lattice for a selection of mu values showed that these deviations are not a finite volume effect. On the 16^4 lattice we are able to simulate at weaker coupling (beta=5.7). At mu=0 we find far better agreement with known results than for beta=5.6. We are extending our beta=5.7 simulations to mu =/= 0 on Blues and Bridges at PSC to determine if there is similar improvement away from mu=0, and propose to continue these runs in FY2017. Note that we have less than 500,000 hours on Bridges until January, 2017. The current runs being performed on Blues are on 16^4 lattices, where we are using 64 cores (2 32-core Haswell nodes). Each job takes 50-60 wallclock-hours. We plan to run similar jobs in FY2017. We intend to run 2 such jobs concurrently, or 1 128 core job. When we are convinced that we have the code we wish to run in extensive production, we will apply optimizations that have proved effective with other codes, to the conjugate gradient routine where the jobs spend most of their time. Project website:www.hep.anl.gov/dks Current: undetermined amount Justification: We estimate that our code achieves approximately 3 Gflops per core on the Haswell nodes for the 64 cores we plan to use for production. 64 cores is the 'sweet spot' for 16^4 lattices where the per core performance is maximum. The following table shows wallclock times for identical benchmark runs on various numbers of cores on the 32-core Haswell nodes on Blues. Each run used all 32 cores on each node, and 1 task/core. 32 cores haswell = 51m44.125s 64 cores haswell = 21m41.861s 128 cores haswell = 13m3.596s 256 cores haswell = 9m43.317s The following table shows wallclock times for identical benchmark runs on various numbers of cores on the 16-core nodes on Blues. Each run used all 16 cores on each node, and 1 task/core. 16 cores = 129m35.228s 32 cores = 52m37.305s 64 cores = 25m35.144s 128 cores = 17m55.801s 256 cores = 16m40.058s As indicated above these numbers would be expected to improve when we apply further optimizations to the conjugate gradient code, which have proved effective on other, similar codes. Our request is based on performing 2 64-core or 1 128-core job continuously, assuming that Blues will be available for production 90% of the time. Requested: 100000 A specific reason has been given: Needed to complete present runs This needs to be approved and the final allocation amount decided upon. Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov