Re: [JLSE-Admins] Storage-aware Large-scale Data Transfers Testbed
I think access to/from NERSC DTNs (dtn{01-04}.nersc.gov) and Mira DTNs should be fine for now. Eunsung: If you want access to/from other nodes, please let Ben know. Raj On Jul 14, 2014, at 11:53 AM, Allen, Benjamin S. <[email protected]> wrote:
Raj, et al.:
What firewall rules need to be put in place for external access to this testbed (both inbound and outbound)?
Do you need just the GridFTP set of ports? Do you have specific external DTNs you’re testing with, so we don’t have to open this testbed’s GridFTP services up to the entire public internet?
Thanks,
Ben
On Jul 8, 2014, at 10:59 AM, Jung, Eun Sung <[email protected]> wrote:
Ok, thank you.
-Eunsung
On 7/8/14, 10:54 AM, "Allen, Benjamin S." <[email protected]> wrote:
Hi Eun Sung,
* Questions regarding the testbed for baseline testing. - Is there any additional node which can play a role of DTN except two 8 file-server nodes?
In the proposed testbed no. Just to be clear I’m proposing 16 file servers, split into two groups. The 16 file servers are backed by two DDN S2A9900 racks of storage. The groups of 8 are to be defined by the file servers connected to the same backing DDN S2A9900
The idea for the baseline testing would be to use one or more of the 8 file servers from one group as clients across the network to the other group of 8 file servers. Thus simulating DTNs accessing a shared filesystem over a network.
- If so, all nodes including fs nodes and a DTN node share the same network?
All 16 file servers will be on the same 10Gb Myrinet network. Each file server will also have a 1GbE public interface, and a 1GbE internal management interface. The same configuration as the Petrel testbed.
The public interfaces will be heavily firewall’ed by default. We’ll explicitly allow inbound and outbound traffic as requested. Ideally we can get the list of firewall rules from Raj, Mike and you in by tomorrow. If you only need standard GridFTP ports opened, please let us know that.
* Questions regarding the testbed for storage-aware large data transfers - Do you mean a single file server by a single file system?
I mean a single shared filesystem (GPFS) across the 16 file servers. For testing using GPFS’ data locality hints.
- Once a file system created, can we adjust the number of file servers later? For example, up to 32 servers, 2 servers per fs node.
You will have 16 physical file servers. You’ll be able to reconfigure them how you see fit within reason. This includes destroying and recreating the GPFS if required.
By default all 51 LUNs (8+2p RAID6 LUNs) per DDN S2A9900 rack are presented to each of the 8 attached file servers via directly connected DDR Infiniband. We can restrict or group which LUNs are presented to which of the 8 file servers to force further data locality if you need to simulate a less ideal storage configuration.
Please let me know if you have further questions,
Ben
On 7/7/14, 11:44 PM, "Raj Kettimuthu" <[email protected]> wrote:
Thanks Ben. This sounds good to me.
Eunsung, Do you have any concerns/questions?
Raj
Sent from my iPhone
On Jul 7, 2014, at 4:05 PM, "Allen, Benjamin S." <[email protected]> wrote:
Hi Raj,
To summarize the requirements for everyone CC¹ed your two objectives are:
1. Baseline testing a traditional DTN which accesses the shared filesystem over a network fabric, compared to a DTN directly on a storage controller. The JLSE file servers provide reasonably facsimiles for embedded storage controllers as they have a point-to-point DDR IB connection to the D2A9900 controllers. They¹re also lower powered servers so should simulate the lower resources available on an embedded storage controller virtual machine.
2. Storage-aware Large Data Transfers: Using data locality information from GPFS to direct GridFTP to the DTN which has local block access to the data. Thus allowing GridFTP to be more efficient in its use of the storage network (Myrinet in this case). A good analogy is Hadoop¹s HDFS, but for data transfer instead of compute.
Please correct or expand the above as needed.
Due to the Petrel testbed being setup as a single file system, and being a pseudo production testbed for Globus Share work (i.e. we can't just destroy the existing globus-fs0 and rebuild it), I think the best course of action to support the above work is to setup two 8 file-server filesystems on a new Myrinet fabric. We¹ll initially create a file system on one (or optionally both) sets of 8 file servers. This will let you use opposing file servers for the baseline testing across a network fabric. This testbed would be similar in specs (same hardware, same networking) as the Petrel testbed.
Once that testing is complete, we will recreate the GPFS as a single file system so you can experiment with data locality. Without further configuration this will give you two silos of data locality (two S2A9900 racks). If you want to further worsen the test cases, we can limit which LUNs are presented to which file server, lowering the number of file-servers with block access to a given set of data.
It looks like we could have this testbed put together by early next week, as we have a near ready set of 16 file servers available. I¹m also suggesting we use this new testbed as a more general resource for those that need GPFS but can withstand the file system being destroyed. I believe this fits Bill¹s project needs, and he has similar goals in testing data locality with GPFS.
All, please let me know if you have any comments or concerns.
Thanks,
Ben
participants (1)
-
Raj Kettimuthu