2017-11-18: Cluster login unavailable *RESOLVED*
We are having some infrastructure outages and it is not currently possible to log in to the clusters.
We're looking into the problem and will update this news item when more information is available.
We are having some infrastructure outages and it is not currently possible to log in to the clusters.
We're looking into the problem and will update this news item when more information is available.
We are experiencing some problems with our support address <support@hpc2n.umu.se> at the moment. This seems to be SNIC wide and they are working on it. This news item will be updated when more information is available.
Update 2017-11-15 14:20
Problem should be solved now and it should be possible to send in support problems to our support address again.
Today, 2017-10-18 13:40, we switched over the software stack on Kebnekaise to one that has been rebuilt from scratch.
All user level codes and most libraries and helper programs have been rebuilt from scratch.
We have tried to make sure nothing used by our users is missing, but should any job fail due to missing libraries or similar, please make sure to notify support@hpc2n.umu.se immediately and we will fix the problem.
On Sep 4th 08:00 CEST we will have a maintenance window to change some parameters for the lustre file system.
To be able to do this we need to empty the clusters from running jobs and reboot all nodes including the login nodes.
As we get closer to that point in time, jobs will not be allowed to start if their requested runtime is too long to fit before 2017-09-04 08:00.
In other words, submitting jobs with shorter runtimes will be a good idea.
Both Abisko and Kebnekaise have been experiencing connectivity issues with the pfs filesystem. Most jobs ought to be unaffected.
We are working to solve the problem as quickly as possible, and this news will be updated with information.
Update, 2017-08-02: problem should be resolved now.
Saturday morning (just after 10 minutes past 0700), there was a small power outage that made some compute nodes in the Abisko and Kebnekaise cluster power cycle.
Jobs running on the nodes that got rebooted did of course fail. No other issues has been detected.
Node reboots to activate security updates seems to have triggered strange disruptive behavior in the Lustre service affecting all compute resources.
We are currently diagnosing the issue in order to come up with a solution.
*UPDATE 2017-06-29 22:45*
Lustre now looks stable again. Queues have been re-enabled.
The login nodes will be rebooted at 12:00 today 2017-06-21, and will be unavailable during that timeframe.
On Wednesday, June 28 08:00 - 17:00, we will perform yet another maintenance on our Lustre filesystem.
This time two of the internal UPSes need to be replaced.
To be on the safe side we will take the Lustre file system offline during this process.
This means that no jobs will be allowed to run after 08:00.
*Update 2017-06-28 15:00*
The maintenance is finished and the queues are running again.
The PFS file system is having some server problems and is currently not accessible.
Due to this, it is not possible to log in to the login nodes.
The batch queues are suspended as we work on this.
We will update this news entry as we make progress and/or resolve the problem.
2017-06-13 17:45
The failing hardware have been reported to our hardware vendor.
2017-06-14 12:24
We are working with the vendor to resolve the problem.