Ticket 429 (https://svn.mpi-forum.org/trac/mpi-forum-web/ticket/429) is also connected with whether remote loads/stores are to be considered as local loads/stores or RMA operations. In the discussion on the mailing list about that ticket, some said they should not be considered as local loads/stores. That discussion is now linked from the ticket. Rajeev On Jul 2, 2014, at 10:31 AM, Rolf Rabenseifner <[email protected]> wrote:
In general I am opposed to decreasing the performance of applications by mandating strong synchronization semantics.
I will have to study this particular instance carefully. I agree.
If the user of a shared memory window only want to issue a memory fence, he/she can use the non-collective MPI_Win_sync.
I would expect that MPI_Win_fence and post-start-complete-wait should be used on a shared memory window only if the user wants process-to-process synchronization.
In all other cases, he/she would use lock/unlock or pure win_sync. The proposal does not touch lock/unlock and win_sync.
A possible alternative would define remote load/store as RMA and would terribly touch all synchronization routines. We all would be against this, although this was the first idea in 2012 with probably the historical remark "This is what I was referring to. I'm in favor of this proposal."
Rolf
----- Original Message -----
From: "William Gropp" <[email protected]> To: "Rolf Rabenseifner" <[email protected]> Cc: "Bill Gropp" <[email protected]>, "Rajeev Thakur" <[email protected]>, "Hubert Ritzdorf" <[email protected]>, "Pavan Balaji" <[email protected]>, "Torsten Hoefler" <[email protected]>, "MPI WG Remote Memory Access working group" <[email protected]> Sent: Wednesday, July 2, 2014 4:46:40 PM Subject: Re: Ticket 435 and Re: MPI_Win_allocate_shared and synchronization functions
In general I am opposed to decreasing the performance of applications by mandating strong synchronization semantics.
I will have to study this particular instance carefully.
Bill
William Gropp Director, Parallel Computing Institute Thomas M. Siebel Chair in Computer Science
University of Illinois Urbana-Champaign
On Jul 2, 2014, at 9:38 AM, Rolf Rabenseifner wrote:
Bill, Rajeev, and all other RMA WG members,
Hubert and Torsten discussed already in 2012 the meaning of the MPI one-sided synchronization routines for MPI-3.0 shared memory.
This question is still unresolved in the MPI-3.0 + errata.
Does the term "RMA Operation" include "a remote load/store from an origin process to the window Memory on a target"?
Or not?
The ticket #435 expects "not".
In this case, MPI_Win_fence and post-start-complete-wait cannot be used for synchronizing the sending process of data with the receiving process of data that use only local and remote load/stores on shared memory windows.
Ticket 435 extends the meaning of MPI_Win_fence and post-start-complete-wait that they provide sender-Receiver synchronization between processes that use local and remote load/stores on shared memory windows.
I hope that all RMA working group members, agree that - currently the behavior of these sync-routines for shared memory remote load/stores is undefined due to the undefined definition of "RMA Operation" - and that we need an errata that resolves this problem.
What is your opinion about the solution provided in https://svn.mpi-forum.org/trac/mpi-forum-web/ticket/435 ?
Best regards Rolf
PS: Ticket 435 is the result of a discussion of Pavan, Hubert and I at ISC2014.
----- Original Message -----
From: Hubert Ritzdorf
Sent: Tuesday, September 11, 2012 7:26 PM
Subject: MPI_Win_allocate_shared and synchronization functions
Hi,
it's quite unclear what Page 410, Lines 17-19
A consistent view can be created in the uni?fied
memory model (see Section 11.4) by utilizing the window
synchronization functions (see
Section 11.5)
really means. Section 11.5 doesn't mention any (load/store) access to
shared memory.
Thus, must
(*) RMA communication calls and RMA operations
be interpreted as RMA communication calls (MPI_GET, MPI_PUT,
...) and
ANY load/store access to shared
window
(*) put call as put call and any store to shared memory
(*) get call as get call and any load from shared memory
(*) accumulate call as accumulate call and any load or store access
to shared window ?
Example: Assertion MPI_MODE_NOPRECEDE
Does
the fence does not complete any sequence of locally issued RMA calls
mean for windows created by MPI_Win_Allocate_shared ()
the fence does not complete any sequence of locally issued RMA calls
or
any load/store access to the window memory ?
It's not clear to me. I will be probably not clear for the standard
MPI user.
RMA operations are defined only MPI functions for window objects
(as far as I can see).
But possibly I'm totally wrong and the synchronization functions
synchronize
only the RMA communication calls (MPI_GET, MPI_PUT, ...).
Hubert
-----------------------------------------------------------------------------
Wednesday, September 12, 2012 11:37 AM
Hubert,
This is what I was referring to. I'm in favor of this proposal.
Torsten
-- Dr. Rolf Rabenseifner . . . . . . . . . .. email [email protected] High Performance Computing Center (HLRS) . phone ++49(0)711/685-65530 University of Stuttgart . . . . . . . . .. fax ++49(0)711 / 685-65832 Head of Dpmt Parallel Computing . . . www.hlrs.de/people/rabenseifner Nobelstr. 19, D-70550 Stuttgart, Germany . . . . (Office: Room 1.307)
-- Dr. Rolf Rabenseifner . . . . . . . . . .. email [email protected] High Performance Computing Center (HLRS) . phone ++49(0)711/685-65530 University of Stuttgart . . . . . . . . .. fax ++49(0)711 / 685-65832 Head of Dpmt Parallel Computing . . . www.hlrs.de/people/rabenseifner Nobelstr. 19, D-70550 Stuttgart, Germany . . . . (Office: Room 1.307) _______________________________________________ mpiwg-rma mailing list [email protected] http://lists.mpi-forum.org/mailman/listinfo.cgi/mpiwg-rma