[LCRC Accounts] Yearly Allocation Request from ScalaGAUSS
Hello, A yearly allocation for the LCRC cluster has been requested with the following updated information: Submitter/PI: Jie Chen Project Name: ScalaGAUSS Division: MCS Project title: Scalable Gaussian Process Data Analysis Associated funding: Scalable Gaussian Process Analysis of Spatio-Temporal Data (PI: Mihai Anitescu) Other Systems: MCS workstations Science: This project uses a maximum likelihood approach to perform Gaussian process analysis on both synthetic and real data. The observations considered for analysis are simulation and data from nuclear engineering and climate science. The resulting analysis is used to describe the spatial correlation of the sampled data, to predict and to understand uncertainty in very high fidelity phenomena. Project description: THIS PROJECT WAS STARTED CLOSE TO THE END OF FY2011 AND ONLY A SMALL NUMBER OF CORE-HOURS WERE ALLOCATED. WE ARE REQUESTING AN ADDITIONAL AMOUNT OF 500000 CORE-HOURS FOR FY2012. We model data as Gaussian process model using several covariance models, including a Matern covariance kernel. The goal of the computation is to estimate the parameters of the covariance kernel. We employ a maximum likelihood method, which is reformulated by using a sample average approximation approach. The novelty of the approach is that it allows for scalable analysis without compromise in the accuracy of the correlation. This reduces to solving linear systems that have the same size of the grid from which the data are sampled. The linear systems have multiple right-hand sides, and in the case of a regular grid, they are multilevel Toeplitz matrices. We employ a block CG solver with multilevel circulant preconditioners. Previous experiments indicate a linear scalability of computational costs for 2D grids of size up to $10^3times10^3$ on a single desktop machine in a Matlab environment. We have demonstrated that up to 1 million data points our approach scales virtually pe rfectly. We have developed parallel C++ codes based on the Trilinos project and will investigate the scalability of the algorithm for 4D data set with on the order of $10^9-10^{12}$ data points; beyond the $10^6$ data sets we can manipulate at the moment. For being able to compile our code, we need (in addition to standard compilers) Trilinos, and ideally also PETSC. The ScalaGauss code can be used in numerous other applications (chemistry, power grids, materials science) that require a Gaussian process statistical models. Project URL: Current FY Hours Used: undetermined amount New FY Requested allocation: 499999 Q1: 125000 Q2: 125000 Q3: 125000 Q4: 124999 Justification: Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov