Begin forwarded message:

From: "Chinthamani, Sundaram" <sundaram.chinthamani@intel.com>
Subject: RE: Stream
Date: April 27, 2016 at 10:58:11 AM CDT
To: "Sodani, Avinash" <avinash.sodani@intel.com>, "Kumaran, Kalyan" <kumaran@alcf.anl.gov>
Cc: "Allison, CM" <cm.allison@intel.com>, "Kirkendall, Keith G" <keith.g.kirkendall@intel.com>, "Fromkin, Russ" <russ.fromkin@intel.com>, "Chinthamani, Sundaram" <sundaram.chinthamani@intel.com>

Thanks.   Then, All2All  (420 GB/s) and Quad mode (440 GB/s) are delivering performance as expected.   At this point, we are checking on whether we can re-produce this issue internally?  When I get time, I will also log-in into your system and see what is going on?
 
-Thanks,
-Sundaram
 
From: Sodani, Avinash 
Sent: Wednesday, April 27, 2016 8:32 AM
To: Kumaran, Kalyan <kumaran@alcf.anl.gov>; Chinthamani, Sundaram <sundaram.chinthamani@intel.com>
Cc: Allison, CM <cm.allison@intel.com>; Kirkendall, Keith G <keith.g.kirkendall@intel.com>; Fromkin, Russ <russ.fromkin@intel.com>; Sodani, Avinash <avinash.sodani@intel.com>
Subject: RE: Stream
 
Thanks Kumar. That at least confirms that quadrant mode bandwidth is as expected.
 
-Avinash
 
From: Kumaran, Kalyan [mailto:kumaran@alcf.anl.gov] 
Sent: Wednesday, April 27, 2016 7:24 AM
To: Sodani, Avinash; Chinthamani, Sundaram
Cc: Allison, CM; Kirkendall, Keith G; Fromkin, Russ
Subject: Stream
 
 

 

Begin forwarded message:
 
From: Scott Parker <sparker@anl.gov>
Subject: Re: [JLSE-Intel-NDA] Intel Meeting Notes
Date: April 26, 2016 at 5:50:32 PM CDT
 

Re-ran stream in Quadrant mode and I was able to get 435 GB/s with 1 thread per core, and a very slightly better 440 GB/s with 2 threads per core.
In SNC-4 mode I saw 311 GB/s with 1 thread and 442 GB/s with 2 threads. All2All hit 425 GB/s with 1 thread per core.

A peak of around 440 GB/s is basically what we should be expecting to see for a Bin 3 part, so the MCDRAM seems to be capable of delivering
the expected bandwidth. However, SNC-4 is an outlier in requiring 128 threads, and gives unexpected low (73 GB/s) bandwidth from one of the
clusters (vs 112 GB/s from each of the other three).

-Scott