[LCRC Accounts] Project Request: ScalaGAUSS
Hello, A new project on the LCRC cluster has been requested. Please forward the information on to the LCRC Allocation sub-committee. Applicant's name: Jie Chen Applicant's institution: ANL Applicant's division: MCS Project Name: ScalaGAUSS Project title: Scalable Gaussian Process Data Analysis Associated funding: Scalable Gaussian Process Analysis of Spatio-Temporal Data (PI: Mihai Anitescu) Other Systems: MCS workstations Science: This project uses a maximum likelihood approach to perform Gaussian process analysis on both synthetic and real data. The observations considered for analysis are simulation and data from nuclear engineering and climate science. The resulting analysis is used to describe the spatial correlation of the sampled data, to predict and to understand uncertainty in very high fidelity phenomena. Project description: We model data as Gaussian process model using several covariance models, including a Matern covariance kernel. The goal of the computation is to estimate the parameters of the covariance kernel. We employ a maximum likelihood method, which is reformulated by using a sample average approximation approach. The novelty of the approach is that it allows for scalable analysis without compromise in the accuracy of the correlation. This reduces to solving linear systems that have the same size of the grid from which the data are sampled. The linear systems have multiple right-hand sides, and in the case of a regular grid, they are multilevel Toeplitz matrices. We employ a block CG solver with multilevel circulant preconditioners. Previous experiments indicate a linear scalability of computational costs for 2D grids of size up to $10^3times10^3$ on a single desktop machine in a Matlab environment. We have demonstrated that up to 1 million data points our ap proach scales virtually perfectly. We have developed parallel C++ codes based on the Trilinos project and will investigate the scalability of the algorithm for 4D data set with on the order of $10^9-10^{12}$ data points; beyond the $10^6$ data sets we can manipulate at the moment. For being able to compile our code, we need (in addition to standard compilers) Trilinos, and ideally also PETSC. The ScalaGauss code can be used in numerous other applications (chemistry, power grids, materials science) that require a Gaussian process statistical models. Project URL: Requested allocation: 201000 Justification: The requester has used 0 hours of their initial startup project. In addition to approving an initial amount, please specify a Category and Subcategory for this project. For a list of the current selection of approved categories, please see: https://wiki.lcrc.anl.gov/wiki/Processes/Categories Once the Allocation committee has approved the project, please go to the Project Management page to create it: https://accounts.lcrc.anl.gov/projects.php Thank You, The LCRC Accounts System
participants (1)
-
accounts@lcrc.anl.gov