On Fri, Apr 24, 2009 at 07:38:39AM -0500, Jeff Hammond wrote:
You're comparing ("scaled better than") apples ("an implementation on top of 2-sided MPI-1 send/recv") and oranges ("real one-sided hardware"). Please post the whitepaper (I cannot find it online) and clarify what you mean here.
QLogic seems to have hidden it here: http://www.qlogic.com/SiteCollectionDocuments/Education_and_Resource/whitepa... I used the Siosi6 benchmark, since I needed something that had public benchmark numbers available. Our implementation used 1/4 of the cores as "memory servers" which held the global memory and served as the counterparty to one-sided operations. Despite wasting 25% of the cores, it scaled great. Now this result won't apply to all one-sided programs; GA's model and usage is somewhat different from most other one-sided models. -- greg