Rob,
Here is the pointer to the dissertation about a production grade memory database using ULFM as a communication substrate that can deal with different types of faults. The approach taken here is to add a layer of resilient collective to supplement the rest of the API, and to facilitate the integration of this code in the rest of the software infrastructure. The result section at the end shown very good results, basically 1/2 the time of the restart approach (even for small problems with 1GB / process).