Troubleshooting Common Job and Pod Failures
This section provides a comprehensive guide to identifying and resolving common issues that can lead to job and pod failures on our HPC cluster. By following these troubleshooting steps, you can effectively diagnose and rectify problems, ensuring the smooth execution of your computational tasks.
Even if your problem isn’t listed here, make sure to check your job’s logs, run runai logs <your_job>, or run runai logs -h for more help on watching logs.
Pod Failed Due to Out-of-Memory (OOM)
Problem Description
One common cause of pod failures is insufficient memory allocation.
When a pod’s memory usage exceeds its limit, Kubernetes may terminate the pod to prevent system instability. This can be particularly problematic for workspaces, which often require more resources than batch workloads.
How to Identify the Problem
While the Run:AI web interface and standard runai logs command may not always directly indicate an OOM condition, you can verify this by running the following command:
runai-bgu logs <workload_name> --loglevel debug
Look for the following message in the logs:
DEBU[0000] pod print status: OOMKilled
This message confirms that the pod was terminated due to excessive memory usage.
Solution
To address OOM issues, you can increase the memory limit allocated to your workload.
This can be done by adjusting the --memory parameter when sending workloads using the CLI or by adjusting the memory limit from the web interface when submitting the workload.
A general rule of thumb is that workspaces usually require more resources than the non-interactive workloads. For example, Pycharm requires 2GiB of free RAM and Visual Studio Code requires 1GiB of free RAM and at least 2 cores only for running the IDE, while a training will only use your resources to run your code. Keep in mind that these requirements are not set in stone and may be flexible to some point.