The default kokkos and cabana modules have been updated to point to those built with this SDK. The associated kokkos/Cabana source code versions are the most recent known to work with this SDK and SYCL execution space (backend).
This is the first time a cabana module has been built in /soft/restricted/CNDA on JLSE, note (there have been a few previous version on Sunspot). For historical consistency with Kokkos builds on JLSE, there are AOT and JIT builds for the {Serial,OpenMPOpenMPTarget} backend combination on JLSE, as well as AOT and JIT builds for the {Serial,OpenMP,SYCL} backend combination. On Sunspot, only the AOT build for the {Serial,OpenMP,SYCL} backend combination. Cabana is headers-only, so there’s no need for separate AOT and JIT builds. The only builds on both JLSE and Sunspot are those for the {Serial,OpenMP,SYCL} backend combination.
Best regards,
Tim
------- Timothy J. Williams, Ph.D. --- Argonne National Laboratory --- tjwilliams(a)anl.gov -------
From: Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss-bounces(a)lists.jlse.anl.gov> on behalf of Chan-nui, Christopher via Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss(a)lists.jlse.anl.gov>
Date: Friday, February 24, 2023 at 12:23 PM
To: Bertoni, Colleen <bertoni(a)anl.gov>, Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss(a)lists.jlse.anl.gov>
Subject: Re: [Ecp-aurora-sdk-discuss] Updating default engineering oneAPI module on JLSE/Sunspot
The oneapi/eng-compiler/2022.12.30.003 environment is now the default on the system.
--
Christopher Chan-Nui
christopher.chan-nui(a)intel.com
From: Bertoni, Colleen <bertoni(a)anl.gov>
Date: Wednesday, February 22, 2023 at 12:15 PM
To: Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss(a)lists.jlse.anl.gov>, Chan-nui, Christopher <christopher.chan-nui(a)intel.com>
Subject: Re: Updating default engineering oneAPI module on JLSE/Sunspot
Thanks a lot!
>From the ANL bug reproducer side, there were 32 bug fixes and 9 regressions on PVC (sunspot). The details are below. Please take a look if you submitted one of these bugs!
We've also been tracking performance for the applications and benchmarks where the reporter added an FOM. For these, there was 1 performance regression and 12 performance improvements of > 20%. See below for details.
Thanks,
Colleen
# Functionality:
## Fixes:
source/reproducers/dpcpp/allocate_vector:allocate_vector.b._cpu [OTFIP-608]
source/reproducers/dpcpp/ext_oneapi_barrier_issue:ext_oneapi_barrier_issue_barrier [CMPLRLLVM-37962,CMPLRLLVM-42541] P0
source/reproducers/dpcpp/ext_oneapi_barrier_issue:ext_oneapi_barrier_issue_barrier_no_parameters [CMPLRLLVM-37962,CMPLRLLVM-42541] P0
source/reproducers/dpcpp/ext_oneapi_barrier_issue:ext_oneapi_barrier_issue_barrier_no_parameters_workaround [CMPLRLLVM-37962,CMPLRLLVM-42541] P0
source/reproducers/dpcpp/marray_math [CMPLRLLVM-42382]
source/reproducers/dpcpp/reduce_over_group [CMPLRLLVM-41971,CMPLRLLVM-42157] P0
source/reproducers/dpcpp/reduction_syclrange_item [CMPLRLLVM-26213]
source/reproducers/dpcpp/test_xgc_get_acoef [CMPLRLLVM-40905]
source/reproducers/dpcpp/work_item_larger_than_uint:work_item_larger_than_uint_fit-in-int [CMPLRLLVM-38306,XDEPS-4586] P0
source/reproducers/hybrid/tile_as_device:tile_as_device_sycl [CMPLRLLVM-29858]
source/reproducers/ifx/CMPLRLLVM-41916
source/reproducers/igc/gamess_iot_fail
source/reproducers/l0/cache:cache_cache_ath_run
source/reproducers/mkl/fft1d_r2c_many_batches_fom [MKLD-13331] P0
source/reproducers/mkl/openmp_offload_dgemm_fortran_link [MKLD-10248,MKLD-13893] P0
source/reproducers/mpi/implicit_scaling:implicit_scaling_implicit_vs_explicit [CMPLRLLVM-41894]
source/reproducers/openmp/continuation_line:continuation_line_original [CMPLRLLVM-41203]
source/reproducers/openmp/continuation_line:continuation_line_simpler [CMPLRLLVM-41203]
source/reproducers/openmp/declare_varient_match [CMPLRLLVM-22836]
source/reproducers/openmp/first_private_this_wrong [CMPLRLLVM-42444] P0
source/reproducers/openmp/fortran_nested_function_calls [CMPLRLLVM-27380,CMPLRLLVM-39731]
source/reproducers/openmp/fortran_offload_print:fortran_offload_print_ze_binary_ath_run [CMPLRLLVM-24956,XDEPS-4365]
source/reproducers/openmp/function_ptr:function_ptr_ath_run [CMPLRLLVM-29485,XDEPS-2230]
source/reproducers/openmp/ifx_ice_thornado
source/reproducers/openmp/ldbl_true_min_issue [CMPLRLLVM-39265]
source/reproducers/openmp/omp_offload_isnan:omp_offload_isnan_ath_compile [CMPLRLLVM-25014,CMPLRLLVM-25019]
source/reproducers/openmp/one_reduction:one_reduction_host_thread [CMPLRLLVM-22031,CMPLRLLVM-42440,XDEPS-3875] P0
source/reproducers/openmp/printf_enum:printf_enum_hang [CMPLRLLVM-37260,CMPLRLLVM-37262] P0
source/reproducers/openmp/printf_enum:printf_enum_wrong_printf [CMPLRLLVM-37260,CMPLRLLVM-37262] P0
source/reproducers/openmp/slow_array_reduction [CMPLRLLVM-43486]
source/reproducers/openmp/thermo:thermo_ath_run [CMPLRLIBS-34271]
source/reproducers/tools/advisor_gflop:advisor_gflop.w_an_OpenMP_app
## Regressions:
source/reproducers/dpcpp/hypre_onedpl
source/reproducers/dpcpp_ct/dpct_constant_iterator
source/reproducers/hybrid/pci_bw_ze_bidirectional:pci_bw_ze_bidirectional_naive_ath_run
source/reproducers/openmp/gamess_mini_jit [XDEPS-2759]
source/reproducers/openmp/hang_glinetable_plus_large_register_file:hang_glinetable_plus_large_register_file_ath_run [CMPLRLLVM-36221]
source/reproducers/openmp/hanging_gamess:hanging_gamess_AOT_version [CMPLRLLVM-39824,GSD-2102,XDEPS-5609]
source/reproducers/openmp/hanging_gamess:hanging_gamess_ath_run [CMPLRLLVM-39824,GSD-2102,XDEPS-5609]
source/reproducers/openmp/long_double [CMPLRLLVM-10445]
source/reproducers/openmp/split-parallel-for-reduction:split-parallel-for-reduction_gemv-case-4-1 [CMPLRLLVM-10985]
# Performance:
## Improvement > 20%
source/applications/gtensor-bench:gtensor-bench_ath_run
- getrs.inverted_avg.float.210x210x1x512.32_32 0.0006409352 seconds | improved from 0.0012077637 (46.9%), lower is better
- getrs.inverted_avg.double.210x210x1x512.32_32 0.0008873758 seconds | improved from 0.0015019037 (40.9%), lower is better
source/applications/heFFTe:heFFTe_ath_run
- heFFTe-onemkl-float-64x64x64-GFlopsPerSec 174.12 GFlopsPerSec | improved from 66.959 (160.0%), higher is better
- heFFTe-onemkl-float-128x128x128-GFlopsPerSec 1142.84 GFlopsPerSec | improved from 468.78 (143.8%), higher is better
- heFFTe-onemkl-float-256x256x256-GFlopsPerSec 2608.64 GFlopsPerSec | improved from 469.02 (456.2%), higher is better
- heFFTe-onemkl-float-512x512x512-GFlopsPerSec 2339.37 GFlopsPerSec | improved from 612.039 (282.2%), higher is better
- heFFTe-onemkl-double-64x64x64-GFlopsPerSec 159.722 GFlopsPerSec | improved from 54.2324 (194.5%), higher is better
- heFFTe-onemkl-double-128x128x128-GFlopsPerSec 796.237 GFlopsPerSec | improved from 325.507 (144.6%), higher is better
- heFFTe-onemkl-double-256x256x256-GFlopsPerSec 1367.36 GFlopsPerSec | improved from 309.783 (341.4%), higher is better
- heFFTe-onemkl-double-512x512x512-GFlopsPerSec 1206.95 GFlopsPerSec | improved from 349.629 (245.2%), higher is better
source/applications/madgraph4gpu-SYCL:madgraph4gpu-SYCL_ee-mumu_ath_run
- MEs_per_sec 24172900.841834 sec^-1 | improved from 12222448.507893 (97.8%), higher is better
- time_in_SigmaKin 0.006778 sec | improved from 0.013405 (49.4%), lower is better
## Regressed > 20%
source/applications/gamess-rimp2-v2:gamess-rimp2-v2_fom2
InverseTimeforW60 0.0241074 1/s | regressed from 0.0380706 (36.7%), higher is better
________________________________
From: Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss-bounces(a)lists.jlse.anl.gov> on behalf of Chan-nui, Christopher via Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss(a)lists.jlse.anl.gov>
Sent: Wednesday, February 22, 2023 10:58 AM
To: Ecp-aurora-sdk-discuss <ecp-aurora-sdk-discuss(a)lists.jlse.anl.gov>
Subject: [Ecp-aurora-sdk-discuss] Updating default engineering oneAPI module on JLSE/Sunspot
This will update the engineering compiler, MPICH and GPU driver UMD version.
The new module file is currently available as:
module add oneapi/eng-compiler/2022.12.30.003
This will be made the default oneapi module on Friday, at noon Central.
--
Christopher Chan-Nui
christopher.chan-nui(a)intel.com