On 3/30/16 2:51 PM, Vitali A. Morozov
wrote:
A quick update on STREAM triad on KNL
B0 after BIOS upgrade.
The first figure is for 1 core. About 35 GB/s from L2 cache,
drops down to 11.5 GB/s up to 64 MB, but then increases to a
little over 13 GB/s. All exactly as expected.
The second figure is for 8 tiles (a quadrant). Near linear
increase in bandwidth up to 284 GB/s when data fit into all L2
caches - 8 MB. This is near 80% of peak L2 bandwidth, which is
expected at 357.6 GB/s then drops to about 78-80 GB/s, a little
less then one forth of expected peak of 120 GB/s (67% of peak).
Two points were obtained for the footprint outside MCDRAM size -
both show 79 GB/s.
The last figure is for entire node. The accumulative L2 cache
bandwidth is a little over 1 TB/s (1.4 TB/s is a peak) with 32
MB capacity. After this footprint, I observe the drop down of
the bandwidth to 80 GB/s. This is strange as MCDRAM bandwidth is
rated 480 GB/s.
Issues:
1. Full node MCDRAM bandwidth does not reach expected 420-430
GB/s bandwidth.
2. 75 GB memory footprint case failed all the the runs.
Vitali
On 03/25/2016 03:29 PM, Vitali A. Morozov wrote:
Hello,
here are the figures that I have obtained for STREAM triad. The
peak numbers are:
L2 peak bandwidth 44.7 GB/s, based on the size of the request
table, L2 hit latency, and core capabilities to issue memory
requests clocked at 1.3 GHz.
MCDRAM peak bandwidth 480 GB/s scaled down to 1 port of L2 tile
=> 15 GB/s. Could be slightly less if the port was
provisioned to 36 tiles (13.3 GB/s in this case).
DDR4 peak bandwidth 112 GB/s is based on 6 channel DDR4-2400.
Being scaled down to 2134 GT/s, it will be 100 GB/s.
The first graph provides STREAM triad bandwidth for 1 core and 1
tile (2 cores, 1 thread is running on each core).
I do not look at L2 cache bandwidth, as 44.7 GB/s number is very
hard to reach and there might be many factors that can reduce
effective bandwidth (moreover, I was not carefully coding stream
to get the full L2 bandwidth anyway). The measured 32 GB/s from
a core and 37 GB/s for 2 cores on a tile are within 80% of
expected tolerance. Running more then 2 threads on a tile
severely degrades performance (20 GB/s at best), which is
expected as the L2 request table gets quickly full.
When data fits into MCDRAM, I have measured 10-12 GB/s for a
core and 10-13 GB/s for a tile with strange repeatable artifacts
at around 64 MB and 640 MB. The tile performance is more
uniform, but starts quickly degrade at around 900 MB, which is
strange as we should have plenty of room in MCDRAM. Also, I am
measuring triad, so the cache must be warmed out by the time I
started it as triad is one of the last benchmarks in STREAM
suite. In fact I do not expect any performance drop between
about 8 MB up to 192 GB - the memory bandwidth should be capable
to provide single port MCDRAM-available bandwidth at ease.
The second graph provides STREAM triad numbers for 8 tiles.
Again L2 cache bandwidth is a kind of OK except that it quickly
gets down not reaching the expected max of 8 MB combined L2
cache. The MCDRAM bandwidth is within the expected range (with a
very strange dip in between 12 -� 40 MB. Then, the bandwidth
decreases with a very unusual change of curvature at around 2
GB. Even if MCDRAM does not work as cache, I would expect to see
around 90 GB/s - I am measuring 4 GB/s at 38 GB.
Sounds like Ben have installed CR instead of DDR :)
The second graph numbers can slightly be improved with careful
mapping the processes to the cores, but not very much.
Vitali
_______________________________________________
Intel-nda mailing list
Intel-nda@lists.jlse.anl.gov
https://lists.jlse.anl.gov/mailman/listinfo/intel-nda
_______________________________________________
Intel-nda mailing list
Intel-nda@lists.jlse.anl.gov
https://lists.jlse.anl.gov/mailman/listinfo/intel-nda