Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Prasanna Balaprakash Project Name: perfopt Division: MCS Project title: Optimization Algorithms for Empirical Performance Associated funding: DOE ASCR Other Systems: Surveyor, ANL: 100000 core hours NERSC: 100000 core hours Science: Our scientific objectives are: i) To advance the state-of-the-art in empirical performance tuning and adaptation of scientific applications on extreme scale architectures and thereby significantly increasing scientific productivity ii) To improve our understanding of large-scale empirical performance tuning iii) To develop novel mathematical optimization algorithms that exploits the characteristics of the search problem in performance tuning Project description: The rapid rate of innovations in computing architectures has widened the gap between the theoretical peak and the achievable performance of scientific codes. Often, scientific application programmers address this issue by manually rewriting the code for the target machine, but this approach is neither scalable nor portable. Empirical performance tuning or automatic performance tuning (in short, autotuning) is a promising approach to address the limitations of manual tuning. This approach consists of identifying relevant code optimization techniques (such as loop unrolling, register tiling, and loop vectorization), assigning a range of parameter values using hardware expertise and application-specific knowledge, and then either enumerating or searching this parameter space to find the best-performing parameter configuration for the given machine. Using this approach, several researchers have achieved considerable success in tuning scientific kernels for b oth serial and multicore processors Our goal is to advance the state-of-the-art in empirical performance tuning and adaptation of scientific applications on extreme scale architectures and thereby significantly increasing scientific productivity. To achieve this goal, we will systematically evaluate the strengths and limitations of optimization algorithms in empirical performance tuning. This study will improve our understanding of large-scale empirical performance tuning and help us to identify open issues that need to be addressed which will require developing novel, flexible optimization approaches. A main aspect of this line of research will be the systematic integration and experimental analysis on multiple problem classes on different architectures. Currently, we are developing novel numerical optimization algorithms that exploits the characteristics of the search problem in performance tuning. We will investigate the effectiveness of the optimization algorithms and will conduct an experimental study of some popular optimization algorithms. Our hypothesis is that appropriately modified local search algorithms can find high-performing code variants in short computation times. This is based on the rationale that the exploration component of global search algorithms is less beneficial in empirical performance-tuning problems, where finding high-performing configurations in short computation time is more important than finding the optimal configuration independent of the computation time required. On Fusion, we conduct an experimental study of some global and local search algorithms on a number of problems from the previously developed SPAPT test suite. We analyze the impact of initial configuration from which a search algorithm starts, input size, and search time constraints on the effectiveness of global and local search algorithms. Another benefit Fusion brings to this research is the ability to test the portability of empirical performance tuning techniques on Intel-based clusters. Eventually, we hope to deploy our techniques for improving the performance of scientific applications running on Fusion. We estimate 100,000 core hours for our experimental study: evaluation of the strengths and limitations of empirical performance tuning = 15,000 core hours; experimental analysis on multiple problem classes on different architectures = 30,000 core hours; empirical analysis for new approaches = 30,000 core hours; tests on parallel performance tuning approaches = 25,000 core hours; We will strongly involve in research dissemination activities such as writing scientific articles and giving seminars at Argonne and in conferences. Concerning the percentage of serial jobs, in the the first and second quarter, our jobs will be 80% serial. For the third and fourth quarter, we will have 30%. Project URL: http://www.mcs.anl.gov/research/project_detail.php?id=90 Current FY Hours Used: undetermined amount New FY Requested allocation: 95000 Q1: 15000 Q2: 30000 Q3: 25000 Q4: 25000 Justification: Thank You, The LCRC Accounts System