Since 80K is all that is needed to finish the work, then lets provide it. Ray ----------------------------------------- Ray Bair Computing, Environment, and Life Sciences Argonne National Laboratory and the University of Chicago TCS Building 240, Room 4122 9700 South Cass Avenue Argonne, IL 60439 email: rbair(at)anl.gov Phone: (630)252-5751 On 6/22/15, 11:17 AM, "[email protected]" <[email protected]> wrote:
Hello,
A change in allocation has been requested:
Requester: sinclair (Donald Sinclair) Project: Lattice-QCD Title: Lattice simulations of Conformal and Walking Technicolor. Description: We perform simulations to evaluate the functional integrals of QCD-like theories formulated on a discrete space-time lattice, to enable determination of the non-perturbative aspects these theories. These include the properties of these theories at non-zero temperature, including the scales of confinement and chiral symmetry breaking, and such zero temperature properties as spectra (including the Higgs mass), decay constants and the running of the gauge coupling constant. We are particularly interested in those theories where the coupling constant evolves very slowly, since these are candidate 'Walking Technicolor' theories. Related to these are theories with an infrared fixed point (conformal field theories). Our first goal is to differentiate between these two different types of behaviour for candidate theories.
We are performing simulations of QCD-like theories which are models for Walking or Conformal Technicolor. We have been studying theories which are essentially QCD but with colour-sextet rather than colour-triplet quarks. We hope to measure the running of the QCD coupling constant. For 2 or 3 flavours, 2-loop perturbation theory predicts an infrared fixed point.
For 2 flavours, it is possible that a chiral condensate forms before this fixed point is reached. If so, the fixed point is avoided, the theory is confining, and chiral symmetry breaks spontaneously. However, there is a region where the coupling constant evolves very slowly. These are the properties required for a walking technicolor theory. Simulations we have performed so far at finite temperature suggest that this theory might walk. So far the results are inconclusive.
For 3 flavours we know that the theory should be conformal. Simulations at N_t=6 (12^3 X 6 lattice) and N_t=8 (16^3 X 8 lattice), some of which used Fusion, did not yet show evidence of conformality. We are thus moving to N_t=12 (24^3 X 12 lattice). We started these simulations on Blues, and intend to continue them in FY2015 on Blues, while performing lower mass simulations on Edison at NERSC.
Our sextet quark codes are based on our earlier triplet quark codes and use the RHMC simulation method. The Rational Hybrid Monte Carlo (RHMC) is a stochastic molecular dynamics algorithm. The functional integral of QCD is written as a partition function of a classical field theory evolving in a fictitious time. The determinant of the Dirac operator raised to a fractional power is calculated by introducing bosonic fields (pseudofermions), and sandwiching this Dirac operator raised to minus said fractional power between them. This fractional power of the Dirac operator is approximated to machine accuracy by a rational approximation. After defining this this theory on a discrete space-time lattice, the inversions required by the partial-fraction expansion of the rational approximation are performed using Krylov space methods, in particular a multi-shift extension of the conjugate gradient algorithm. A global Metropolis Monte-Carlo accept/reject step applied at the end of each trajectory removes discretization errors introduced by the numerical integration of these stochastic equations of motion. We parallelize the code by assigning a fixed number of adjacent lattice sites to each MPI task. Network bandwidth ultimately limits how small a chunk of the lattice can be assigned to each task.
Our 24^3 X 12 runs will be performed on 288 cores (18 nodes) of Blues.
A short benchmark run of the 24^3 X 12 code yielded 667 Gflops or 2.3 Gflops/core. Since we have seen performances of > 3 Gflops/core on machines with similar processors, we suspect that at least 3 Gflops/core should be achievable on this machine. Since we have exhausted our allocation on Blues, we cannot perform more extensive tests. Scaling tests on Fusion and on Edison at NERSC for numbers of cores from 24 up to 288 cores indicate that the per core performance increases up to 144 or 288 cores, which we believe to be related to cache usage.
Our new project, simulating QCD at finite mu using Complex Langevin with Gauge Cooling, will require some code development. Our old codes do not as yet have gauge cooling and our parallel code uses SHMEM rather than MPI for communication. Since the real Langevin is the limit of our Hybrid Molecular Dynamics code (the forerunner of the RHMC method) in which there is a single update per trajectory, most of what we have discussed for the RHMC applies, and will not be repeated here. We plan to use Blues and Carver at NERSC for the early, small-lattice simulations, 8^3 X 4 on 32 cores and 8^4 on 64 cores. Because of the smaller memory requirements, allowing more efficient cache usage, we expect to get a minimum performance of 3-4 Gflops/core for this code.
Our requested allocation is based on running 1 288 core job half of the time at times when the NERSC machines are experiencing slowdowns for amall jobs.
Current: undetermined amount Justification: We have not yet performed a detailed scaling analysis on Blues, however, we have such analyses performed on Fusion as well as Edison at NERSC. For Fusion, for our QCD with sextet quarks code on a 24^3 X 12 lattice running on Fusion we observed the following performances: 24 cores = 45 Gflops = 1.9 Gflops/core 48 cores = 95 Gflops = 2.0 Gflops/core 72 cores = 148 Gflops = 2.1 Gflops/core 96 cores = 220 Gflops = 2.3 Gflops/core 144 cores = 369 Gflops = 2.6 Gflops/core 288 cores = 784 Gflops = 2.7 Gflops/core For the same code and lattice size running on Edison at NERSC we observed the following performances. 24 cores = 53 Gflops = 2.2 Gflops/core 48 cores = 115 Gflops = 2.4 Gflops/core 72 cores = 191 Gflops = 2.7 Gflops/core 96 cores = 245 Gflops = 2.6 Gflops/core 144 cores = 510 Gflops = 3.5 Gflops/core 288 cores = 959 Gflops = 3.3 Gflops/core We use a custom assignment of tasks to nodes in order to minimize communications. Earlier recoding reduced the number of global reductions in the routines, which use most of the CPU time, by a factor of 2. The node codes are transparently vectorizable, to allow use of SSE instructions.
Requested: 80000
A specific reason has been given: Needed to complete the runs at the current set of parameters. This set of runs is 90% completed
This needs to be approved and the final allocation amount decided upon.
Thank You, The LCRC Accounts System _______________________________________________ allocations-admins mailing list [email protected] https://lists.lcrc.anl.gov/mailman/listinfo/allocations-admins