Hello,
A change in allocation has been requested:
Requester: robl (Robert Latham)
Project: radix
Title: Scalable Parallel System Software
Description: This is a Computer Science request for time on Fusion for development
and evaluation of a family of tools and libraries with a direct benefit
to all users of high-performance computing resources. The Radix group
in Argonnne's Mathematics and Computer Science division conducts
research and development on facilities to better enable efficient
Parallel Programming, I/O, and Scientific Understanding.
The Fusion system represents an attractive milestone for our software
development: projects start out on our breadboard cluster and can
further explore scalability in the presence of a high speed network on
Fusion before making the jump to BlueGene-scale parallelism.
Radix experiments fall into two broad categories. Often, we are
concerned with latency. The infiniband network on Fusion provides an
excellent low-latency environment for message passing and storage
management research while also providing ample bandwidth for I/O and
message bandwidth testing.
Fusion hosts two quality parallel file systems, GPFS and PVFS. The
wider HPC community often asks the Radix file system researchers to
compare these file systems head-to-head. Fusion provides us the
opportunity to do just that.
Hosting two parallel file systems allows us to do more than just
drag-race (so to speak). The two file systems have distinct
characteristics, and act differently in the broader role of the I/O
software stack. On Fusion, we can utilize a single software stack
(application, parallel-netcdf, ROMIO), and compare the impact of a
change in the underlying file system on overall application behavior.
We know from other research, for example, that the two-phase
optimization in ROMIO needs to adapt to the underlying file system, and
that aligning variables to file system boundaries can yield performance
improvements for parallel-netcdf.
Fusion compute nodes actually look fairly attractive to active storage
research. These nodes have enough disk space to store non-trivial
datasets, and powerful enough processors to carry out complex
computations without much impact on storage performance.
Radix projects also focus on deriving insight both about application
behavior as well as scientific insight. 8 core nodes make efforts like
in-situ visualization, where for example an application renders a frame
of a movie, ever more feasible. Even without fancy graphics
accelerators, the processing power on Fusion nodes makes such analysis
and visualization feasible, especially if the visualization can take
advantage of multiple cores and multiple nodes. Naturally, the radix
visualization tools have demonstrated scalability to Fusion sizes and
beyond on BlueGene, so making full utilization of Fusion should not be a
concern.
As should be evident from the variety of planned experiments, the Radix
group plans to make full use of the new features Fusion brings to LCRC
and the flexibility of a Linux-based cluster.
With Argonne employees and student collaborators, we expect around 20
members in FY2010.
While Fusion may contain a modest number of compute nodes, the machine
still represents an attractive platform for testing: very fast, low
latency interconnect; 8 cores per node; .25 TB of storage per compute
node. The testing and experiments the Radix group can carry out on
Fusion ensures the tools the group develops today will remain relevant
even as supercomputers grow. Quite a few radix-developed projects are
part of the Fusion software stack: CPU hours for this project yield
improvements and benefits not just for the radix group, but for all
users of Fusion and indeed users of high-end computational resources
worldwide.
Current: undetermined amount
Justification: Our large allocation is primarily driven by a set of HFS (Hadoop File System) experiments. These experiments use the HFS interface to communicate to a set of PVFS servers. We have consumed a lot of hours getting the scaling right, but we now know that 12 hours on 201 nodes would give us the required 7.5 tb of scratch space, one name node and 200 HFS/MapReduce nodes.
We have tried to run this experiment on Magellan and have found it to be too unstable.
Requested: 4000000
A specific reason has been given:
This particular allocation request extension is for a long HFS (Hadoop File System) experiment. I have provided node counts and hours above in the "large allocation justification".
This needs to be approved and the final allocation amount decided upon.
Thank You,
The LCRC Accounts System