error message while I tried to do a parallel job on a cluster
Hi all, I have an error message while I tried to do a parallel job on a cluster. The error is listed here: [mpiexec@M001-********] HYDU_sock_read (./utils/sock/sock.c:222): read errno (Input/output error) [mpiexec@M001-********] control_cb (./pm/pmiserv/pmiserv_cb.c:249): assert (!closed) failed [mpiexec@M001-********] HYDT_dmxu_poll_wait_for_event (./tools/demux/demux_poll.c:77): callback returned error status [mpiexec@M001-********] HYD_pmci_wait_for_completion (./pm/pmiserv/pmiserv_pmci.c:206): error waiting for event [mpiexec@M001-********] main (./ui/mpich/mpiexec.c:404): process manager error waiting for completion I'm using mpich2-1.3.2 to do the job.It seems the error is pointing to hydra? And I'm not sure about if hydra is installed in this cluster.So Maybe I need to install a same version hydra with mpich2 first? Also I'm confused how hydra is linked to mpich2-1.3.2 while installing them? Is there anyone who can give me some hint? Thanks a lot! Best, Chun
participants (1)
-
Chun Meng