14 Dec
2019
14 Dec
'19
6:51 p.m.
MPICH configure just asked me to choose ch4:ofi vs ch4:ucx vs ch3. Does anyone have an informed opinion on which one is the faster for shared-memory execution? My specific use case is NWChem on very large multi-socket Xeon nodes, where passive target RMA is the primary communication method, hence my primary concern is asynchronous progress and lack of serialization in RMA. In the past, I have observed significant performance issues due to serialization of RMA accumulate operations acting on non-overlapping memory regions. Jeff -- Jeff Hammond [email protected] http://jeffhammond.github.io/
2412
Age (days ago)
2412
Last active (days ago)
0 comments
1 participants
participants (1)
-
Jeff Hammond