Re: [JLSE-Intel-NDA] knl03 in trouble?
Looks like knl03 failed (likely kernel panic) sometime between its last job around 4:00 UTC and your jobs. Your first job should have trigger the node to be marked down however. We don't (today) have a way to reschedule a job when the node fails prologue, so in this case the first unlucky job will be sacrificed. Working on why the node wasn't marked down when the prologue failed. Ben
On Aug 16, 2016, at 10:26 AM, Tom LeCompte <[email protected]> wrote:
Hi everyone-
My jobs on knl03 immediately exit. The same jobs run OK on 0-2 and 6-8. Is this a known problem?
Cobalt doesn't say much:
Tue Aug 16 15:20:40 2016 +0000 (UTC) submitted with cwd set to: /home/lecompte/g4/Geant4HepExpMTBenchmark/run jobid 520 submitted from terminal /dev/pts/5 Tue Aug 16 15:23:55 2016 +0000 (UTC) Info: resources released; starting job epilogue scripts
Cheers,
Tom
_______________________________________________ Intel-nda mailing list [email protected] https://lists.jlse.anl.gov/mailman/listinfo/intel-nda
participants (1)
-
Allen, Benjamin S.