Hi,
Stating the obvious, the NRE model DL frameworks workloads take a long time to initialize and run when the SW stack components, including the Aurora NRE model frameworks conda environment are located on the shared Lustre FS /soft mount. Until a permanent solution for this issue is deployed on Aurora, the SW stack components have been archived in a 'local-frameworks' tarball located on the gecko FS.
This tarball can be untarred to compute node's /tmp directory. A script to setup the user's environment is then run. This will minimize the time to access the SW components when framework workloads are run.
Please see this wiki page for more information on how to use the tarball:
https://apps.cels.anl.gov/confluence/pages/viewpage.action?spaceKey=inteldg…
Please let me know if you have any questions or run into any issues.
Thanks,
Rebecca Ramer
rebecca.ramer(a)intel.com