Are you running in flat mode? Can you choose a size that fits in IPM in this mode? You are seeing only DDR bandwidth and not IPM.
This is the notes for KNH from SOW:
KNH node performance was projected for the STREAM benchmark for three cases:
A)���� IPM is configured in flat mode with all accesses targeting IPM memory addresses. In this mode, STREAM performance is within 90% of the peak IPM performance.
B)���� IPM is configured in cache mode. In this mode, non-temporal stores are used to minimize �Read-For-Ownership� cache coherency overhead. The use of non-temporal stores, however, requires an additional read for each store. This overhead causes a drop in BW from 603 GB/s to 404 GB/s (33%) for STREAM copy and 603 GB/s to 452 GB/s (25%) for STREAM Triad. Note that the additional read is required only for non-temporal stores. It is not needed for write-backs from the processor caches.
C)���� IPM is configured in flat mode with all accesses targeting NVM memory addresses: The controller reads 256B from the NVM and places it in the write data buffer. The subsequent writes to the same 256B does not take NVM BW. Eventually, the 256B data is written back to NVM, which consumes its BW. This results in 25% BW loss for STREAM Triad and 33% BW loss for STREAM copy.
On Mar 30, 2016, at 2:51 PM, Vitali A. Morozov <morozov@anl.gov> wrote:
<stream-knlb0-1core.png><stream-knlb0-8tiles.png><stream-knlb0-32tiles.png>_______________________________________________A quick update on STREAM triad on KNL B0 after BIOS upgrade.
The first figure is for 1 core. About 35 GB/s from L2 cache, drops down to 11.5 GB/s up to 64 MB, but then increases to a little over 13 GB/s. All exactly as expected.
The second figure is for 8 tiles (a quadrant). Near linear increase in bandwidth up to 284 GB/s when data fit into all L2 caches - 8 MB. This is near 80% of peak L2 bandwidth, which is expected at 357.6 GB/s then drops to about 78-80 GB/s, a little less then one forth of expected peak of 120 GB/s (67% of peak). Two points were obtained for the footprint outside MCDRAM size - both show 79 GB/s.
The last figure is for entire node. The accumulative L2 cache bandwidth is a little over 1 TB/s (1.4 TB/s is a peak) with 32 MB capacity. After this footprint, I observe the drop down of the bandwidth to 80 GB/s. This is strange as MCDRAM bandwidth is rated 480 GB/s.
Issues:
1. Full node MCDRAM bandwidth does not reach expected 420-430 GB/s bandwidth.
2. 75 GB memory footprint case failed all the the runs.
Vitali
On 03/25/2016 03:29 PM, Vitali A. Morozov wrote:
Hello,
here are the figures that I have obtained for STREAM triad. The peak numbers are:
L2 peak bandwidth 44.7 GB/s, based on the size of the request table, L2 hit latency, and core capabilities to issue memory requests clocked at 1.3 GHz.
MCDRAM peak bandwidth 480 GB/s scaled down to 1 port of L2 tile => 15 GB/s. Could be slightly less if the port was provisioned to 36 tiles (13.3 GB/s in this case).
DDR4 peak bandwidth 112 GB/s is based on 6 channel DDR4-2400. Being scaled down to 2134 GT/s, it will be 100 GB/s.
The first graph provides STREAM triad bandwidth for 1 core and 1 tile (2 cores, 1 thread is running on each core).
I do not look at L2 cache bandwidth, as 44.7 GB/s number is very hard to reach and there might be many factors that can reduce effective bandwidth (moreover, I was not carefully coding stream to get the full L2 bandwidth anyway). The measured 32 GB/s from a core and 37 GB/s for 2 cores on a tile are within 80% of expected tolerance. Running more then 2 threads on a tile severely degrades performance (20 GB/s at best), which is expected as the L2 request table gets quickly full.
When data fits into MCDRAM, I have measured 10-12 GB/s for a core and 10-13 GB/s for a tile with strange repeatable artifacts at around 64 MB and 640 MB. The tile performance is more uniform, but starts quickly degrade at around 900 MB, which is strange as we should have plenty of room in MCDRAM. Also, I am measuring triad, so the cache must be warmed out by the time I started it as triad is one of the last benchmarks in STREAM suite. In fact I do not expect any performance drop between about 8 MB up to 192 GB - the memory bandwidth should be capable to provide single port MCDRAM-available bandwidth at ease.
The second graph provides STREAM triad numbers for 8 tiles. Again L2 cache bandwidth is a kind of OK except that it quickly gets down not reaching the expected max of 8 MB combined L2 cache. The MCDRAM bandwidth is within the expected range (with a very strange dip in between 12 -� 40 MB. Then, the bandwidth decreases with a very unusual change of curvature at around 2 GB. Even if MCDRAM does not work as cache, I would expect to see around 90 GB/s - I am measuring 4 GB/s at 38 GB.
Sounds like Ben have installed CR instead of DDR :)
The second graph numbers can slightly be improved with careful mapping the processes to the cores, but not very much.
Vitali
_______________________________________________ Intel-nda mailing list Intel-nda@lists.jlse.anl.gov https://lists.jlse.anl.gov/mailman/listinfo/intel-nda
Intel-nda mailing list
Intel-nda@lists.jlse.anl.gov
https://lists.jlse.anl.gov/mailman/listinfo/intel-nda