IKARUS HPC: User Documentation
Your complete guide to the IKARUS High-Performance Computing cluster: the IKARUS portal, interactive applications, job submission, and hardware specifications.
What is IKARUS?
IKARUS is KISR's High-Performance Computing (HPC) cluster, giving researchers and engineers shared access to large-scale computational resources. It is accessible through two methods:
- IKARUS Portal: browser-based access at khpc.kisr.edu.kw/ood/. No software installation required. Recommended for most users. This documentation is linked from the portal's Help menu.
- SSH (terminal): traditional command-line access via
hpc.kisr.edu.kwfor advanced or automated workflows.
Through the portal you can browse and transfer files, submit and monitor batch jobs, open a browser terminal, and launch fully-featured applications like JupyterLab, RStudio, and MATLAB all without leaving your browser.
Quick Access
Accessing the IKARUS Portal
The IKARUS portal gives you full access to the cluster from any web browser. No SSH client or special software required.
System Requirements
| Requirement | Details |
|---|---|
| Browser | Chrome 90+, Firefox 88+, Edge 90+, or Safari 14+. Chrome or Edge recommended. |
| JavaScript | Must be enabled (on by default in all browsers). |
| Pop-ups | Allow pop-ups from khpc.kisr.edu.kw, interactive app sessions open in new tabs. |
| Cookies | Must be enabled for login session to be maintained. |
Logging In
- Open your browser and go tohttps://khpc.kisr.edu.kw/ood/
- The IKARUS login page loads. Enter your SSH username and password provided by IKARUS administration.
- If it is your first time, you are asked to add Two-Factor Authentication to an authenticator app of your choice. Otherwise, enter the current code from your authenticator app.
- Click "Log In".
- On success you are redirected to the IKARUS Portal dashboard.
Dashboard Tour
The navigation bar is present on every page of the portal. It contains six menus:
| Menu | What it does |
|---|---|
| Files | Drop-down to open the file browser. Home Directory takes you to /NFS/scratch/homes/<username>. |
| Jobs | Two sub-items: Active Jobs (live queue view) and Job Composer (script editor and submission). |
| Clusters | Contains IKARUS Shell Access: a browser-based terminal directly on the login node. |
| Interactive Apps | Lists all available applications: Desktop, JupyterLab, RStudio, MATLAB...etc. |
| Tools | Additional research tools integrated into the portal, such as MLFlow for experiment tracking. |
| Help / Logout | Right side of the bar. Access documentation links and the logout option. |
The main body of the dashboard shows the My Interactive Sessions panel: cards for all your active or recently completed interactive app sessions. From here you connect to running sessions, check their status, and delete completed ones.
A Delete All Sessions button clears all completed/failed session records in one click, keeping the dashboard tidy.
khpc.kisr.edu.kw/ood/ for quick access. Your session stays active while the browser is open; after a period of inactivity you will be prompted to log in again.File Browser
Manage files on cluster storage directly in your browser. No SFTP client needed for everyday transfers.
/NFS/scratch/homes/<username>. This is the default location when you open the file browser. Store your scripts, data, and results here.Navigating Directories
- In the top bar click Files → Home Directory. The file browser opens to your home directory.
- The path bar at the top shows your current location. Click any segment to jump there, or type a path and press Enter.
- Click a folder name in the listing to enter it. Use the ← back button or path bar to navigate up.
Show Hidden Files
Toggle Show Dotfiles in the toolbar to reveal files beginning with . (such as .bashrc and application config files).
Filter / Search
Use the Filter box in the toolbar to narrow the file listing by name within the current directory.
Uploading Files
Method 1: Drag and Drop
- Navigate to the destination directory in the file browser.
- Drag one or more files from your local file manager and drop them into the portal file browser window.
- A progress indicator appears. Wait until each file shows 100% before navigating away.
Method 2: Upload Button
- Navigate to the destination directory.
- Click "Upload" in the toolbar.
- Select files in the file picker that appears. Uploads begin immediately.
sftp) connected to hpc.kisr.edu.kw for large data transfers instead.Downloading Files
- Navigate to the directory containing the files you want.
- Check the checkbox to the left of each file or folder you want to download.
- Click "Download" in the action toolbar that appears at the top. Files are saved to your browser's downloads folder.
.zip archive first. For large directories, compress with tar in the shell terminal first, then download the single archive.Create, Rename, Copy, Delete
New Directory
- Navigate to where you want the new folder.
- Click "New Dir" in the toolbar.
- Type the name and press Enter.
New File
- Navigate to the target directory.
- Click "New File" in the toolbar.
- Type the filename (including extension, e.g.
job.sh) and press Enter.
Actions on Existing Files
Check the checkbox next to any file or directory to reveal the action toolbar:
| Action | What it does |
|---|---|
| Edit | Open the file in the built-in web editor (text files only) |
| Rename/Move | Enter a new name or path to rename or move the item |
| Copy/Move | Duplicate or relocate selected files to a destination you specify |
| Delete | Permanently remove selected files/directories. Cannot be undone. |
| Download | Download selected items to your computer |
File Editor
The built-in web editor lets you edit job scripts, configuration files, and any text file directly in the browser.
- Check the checkbox next to the file, then click "Edit" or click the filename directly for text files.
- The editor opens in a new browser tab with syntax highlighting (detected from the file extension).
- Make your changes, then press "Ctrl+S" (or click Save) to save back to the cluster.
- Close the tab when done.
Active Jobs
A live, real-time view of all jobs in the SLURM queue across the entire cluster.
Click Jobs → Active Jobs in the navigation bar. Click Refresh to update, or enable Auto-refresh to poll automatically.
Reading the Queue
| Column | Description |
|---|---|
| Job ID | Unique SLURM identifier. Used with scancel, sacct, and sstat. |
| Job Name | Name set by --job-name in the script. |
| User | Owning user. You can see all users but can only cancel your own jobs. |
| Partition | Queue the job was submitted to (Res, def1, Dev). |
| State | RUNNING active · PENDING waiting · COMPLETED done · FAILED error · COMPLETING cleaning up |
| Time | Elapsed run time for running jobs; wait time for pending jobs. |
| Nodes | Number of nodes allocated. |
| Node List (Reason) | Allocated node names, or why the job is pending (e.g. Resources, Priority). |
Filtering
Use the Filter field to search by name, user, or ID. Toggle Your Jobs to show only your own activity.
Cancelling Jobs
- Find your job. Confirm the "User" column matches your username.
- Click the red trash icon to the right of the row.
- Confirm in the dialog that appears.
- The job moves to CANCELLED state and is removed from the queue.
Job Composer
A graphical SLURM script editor and submission tool. Create, edit, save, and submit job scripts without using the terminal.
Access via Jobs → Job Composer. The page shows saved jobs on the left and a detail/editor panel on the right.
Writing a Job Script
- Click + New Job then choose "From Default Template".
- A new job entry appears. Click it to open the detail panel.
- Click "Open Editor" in the detail panel. The script opens in a new tab.
- Write your
#SBATCHdirectives and commands. Save with "Ctrl+S". - Close the editor tab and return to the Job Composer.
ondemand/data/sys/myjobs/ and can also be edited directly from the file browser.#!/bin/bash
######## SLURM Options ########
#SBATCH --job-name=my_analysis
#SBATCH --partition=Res
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=32G
#SBATCH --time=02:00:00
#SBATCH --output=%j_output.log
#SBATCH --error=%j_error.log
######## End of Options ########
module load python/3.10
cd /NFS/scratch/homes/${USER}/my_project
python3 analysis.py --input data.csv --output results/Submitting a Job
- Select the job from the list.
- Optionally change the "Submit Directory", click the folder icon to browse.
- Click the green "Submit" button.
- On success a banner shows the assigned "Job ID". Monitor in Jobs → Active Jobs.
Using Templates
Save as Template
- Select a working job, then click "Copy to Templates" (star icon).
- Give it a descriptive name (e.g. Python 8-core 32G Res) and save.
Create from Template
- Click + New Job → From Template.
- Select your template. A copy is created and you can edit specific parameters before submitting.
Browser Terminal
A full Linux terminal inside your browser, connected directly to the IKARUS login node. No SSH client needed.
- Click Clusters → IKARUS Shell Access in the navigation bar.
- A new tab opens with a full terminal emulator. You are automatically logged in as your IKARUS user on
clavis. - Prompt confirms connection:
[hpcdemo@clavis2 ~]$ - Navigate the filesystem and manage files
- Load environment modules (
module avail,module load) - Submit and monitor SLURM jobs (
sbatch,squeue,scancel) - Start an interactive compute session with
srun - Compile code and run quick tests
- Transfer files with
scporrsync
squeue -me # your jobs only
sbatch my_job.sh # submit a batch job
srun --partition=Res --cpus-per-task=4 \
--mem=8G --time=01:00:00 bash -l # interactive compute session
module avail # list available modules
du -sh /NFS/scratch/homes/${USER} # check disk usagesbatch or srun.Interactive Applications
Thirteen applications running on real compute nodes, accessible in your browser with SLURM-allocated resources.
Launching an App
- Click Interactive Apps in the header and select the application you want.
- A launch form appears. Fill in the resource fields.
- Click "Launch".
- You are redirected to the My Interactive Sessions dashboard. A session card appears.
- Wait for the card to show "Running" (blue header).
- Click "Connect to [App Name]" to open the application in a new tab.
khpc.kisr.edu.kw in your browser settings.Resource Request Form
| Field | Description | Guidance |
|---|---|---|
| Partition | SLURM queue. Options shown depend on your group membership. | Use Res for standard work. Shorter time requests start sooner. |
| Number of Hours | Max session duration. Session is terminated when this expires. | Request only what you need, launch a new session if more time is needed. |
| Number of CPU Cores | Cores dedicated to the session. | 4–8 for most interactive work. Increase for parallel code. |
| Memory (GB) | RAM allocated to the session. | Match to your largest expected dataset. See individual app pages. |
| App-specific options | Some apps have extra fields (e.g. conda environment, version). | See the individual app page. |
Session States
Waiting for SLURM to allocate a node, or the app is initialising.
Action: Wait typically 30 sec to a few minutes.
Resources allocated, app is ready.
Action: Click the Connect button to open the application.
Session ended: time limit reached, you closed it, or an error occurred.
Action: Delete the card or review logs.
Managing Sessions
Reconnecting to a Running Session
If you close the app tab, the session continues running. Return to the IKARUS Portal dashboard, find your Running session card, and click "Connect" again.
Ending a Session
- Quit the application normally (e.g. File → Quit in MATLAB, close the Jupyter tab).
- Return to the IKARUS Portal dashboard. The session card should show "Completed".
- Click the "Delete" (trash) button on the card.
Delete All Sessions Button
Removes all completed and failed session cards at once. Running sessions must be deleted individually after ending the application.
Multiple Sessions
You can run multiple interactive sessions simultaneously (e.g. JupyterLab and a Desktop session at the same time), subject to your resource quota.
Linux Desktop (VNC)
A full graphical XFCE desktop running on a compute node, streamed to your browser via VNC. Ideal for GUI-based workflows and running graphical applications interactively.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | Choose based on your group access |
| Hours | 1–4 h | Extend if you need longer work sessions |
| CPU Cores | 4–8 | Increase for apps you run inside the desktop |
| Memory | 8–32 GB | Increase for memory-intensive graphical apps |
- Click Connect to Desktop on the session card. The desktop opens via noVNC in a new tab.
- An XFCE desktop is shown. Right-click the desktop for a context menu, or use the panel at the top.
- Open a Terminal Emulator (Applications → Terminal Emulator, or right-click → Open Terminal). Load modules and launch any application from here.
- Launch GUI apps with an ampersand to keep the terminal free:
matlab &,paraview &
noVNC Controls
| Control | Action |
|---|---|
| Ctrl+Alt+Shift | Open the noVNC side panel (clipboard, settings, fullscreen toggle) |
| Clipboard (noVNC panel) | Paste text from your local machine into the remote desktop |
| Fullscreen (noVNC panel) | Enter/exit fullscreen mode |
| Scroll wheel | Scroll within the VNC window |
JupyterLab
Interactive notebook-based computing with Python, R, and other kernels. Run code, visualise results, and document your analysis in the browser.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | |
| Hours | 2–8 h | Jupyter sessions can run long; plan ahead |
| CPU Cores | 4–16 | Increase for parallel Python or heavy data processing |
| Memory | 16–64 GB | Match to your largest expected dataset or array |
- After connecting, JupyterLab opens. The left panel shows a file browser rooted at your home directory. The right panel shows the Launcher.
- Click a kernel tile in the Launcher (e.g. Python 3) to open a new notebook, or use File → Open from Path to open an existing
.ipynbfile. - Write code in cells and press Shift+Enter to run. Output appears directly below.
- Use the file browser (left panel) to navigate directories and open additional files.
Available Kernels
IKARUS provides six shared pre-configured kernels accessible to all users from /NFS/shared/libraries/jupyter/kernels/. These appear in the JupyterLab Launcher automatically; no installation or configuration required.
| Kernel | Language | Notes |
|---|---|---|
| Python 3 | Python 3 (Anaconda base) | NumPy, SciPy, pandas, Matplotlib, scikit-learn and the full Anaconda scientific stack pre-loaded |
| R | R 4.5.3 (IRkernel) | Full R environment; install additional CRAN packages with install.packages() in a notebook cell |
| Bash | Bash shell | Run shell commands and SLURM submissions directly in notebook cells; useful for pipeline documentation |
| PyTorch (CPU) | Python 3 + PyTorch | torch, torchvision, torchaudio (CPU-only build). Use for model development and inference without GPU allocation |
| TensorFlow (CPU) | Python 3 + TensorFlow | tensorflow-cpu (CPU-only build). Suitable for training small models and running inference on the compute nodes |
| pyPDAF 1.0.4 | Python 3 + pyPDAF + mpi4py | Python interface to the PDAF Parallel Data Assimilation Framework. Supports all PDAF filters via MPI (Intel MPI 2021.12). Import submodules directly: from pyPDAF import PDAF, PDAF3, PDAFomi |
Installing Packages
# In a notebook cell: installs to your home directory
!pip install --user package_name
# For R:
install.packages("ggplot2")!pip install --user may not persist across different kernel environments. For reproducible work, request a dedicated Conda kernel from the administrators.RStudio
Full RStudio Server IDE for R statistical computing, running on a compute node and accessed in your browser.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | |
| Hours | 2–8 h | |
| CPU Cores | 4–8 | Increase for parallel R with parallel or future packages |
| Memory | 16–64 GB | R loads data entirely into RAM; match to dataset size |
- After connecting, RStudio opens with four panels: Source (top-left), Console (bottom-left), Environment/History (top-right), Files/Plots/Help (bottom-right).
- Use the Console for interactive R commands, or open/create
.Rscripts in the Source panel. - Set your working directory:
setwd("/NFS/scratch/homes/username/my_project"), or use File → New Project. - Browse cluster files in the Files tab (bottom-right).
Installing R Packages
install.packages("ggplot2")
install.packages(c("dplyr", "tidyr", "data.table"))Packages install to your personal R library in your home directory and persist across sessions.
.Rproj files): each project sets the working directory automatically and maintains a separate history and environment.MATLAB
The full MATLAB desktop running on a compute node and displayed via VNC. Suitable for numerical computation, signal processing, image analysis, and simulations.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | |
| Hours | 2–8 h | |
| CPU Cores | 8–16 | Increase for Parallel Computing Toolbox usage |
| Memory | 32–128 GB | Allocate enough for your largest arrays and datasets |
- After connecting, the MATLAB desktop opens in VNC: Command Window, Workspace, and file browser are visible.
- Use the Command Window for interactive commands, or open
.mfiles in the MATLAB Editor. - Set your working directory:
cd('/NFS/scratch/homes/username/my_project') - Run scripts with the green Run button or by typing the script name.
Parallel Computing
% Create a local parallel pool using the cores you requested
parpool('local', 8);
% Use parfor for parallelised loops
parfor i = 1:100
result(i) = my_function(i);
endParaView
Open-source scientific visualisation for large simulation and experimental datasets. Runs with software rendering on the compute node, displayed via VNC.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | |
| Hours | 1–4 h | |
| CPU Cores | 8–16 | More cores enable parallel rendering and processing |
| Memory | 32–128 GB | Large meshes require significant RAM; size to your dataset |
- After connecting, ParaView opens in VNC with its Pipeline Browser, Properties panel, and 3D render view.
- Open data: File → Open. Supported formats include VTK, EnSight, OpenFOAM, netCDF, HDF5, Exodus, CGNS, and more.
- Click Apply in the Properties panel to load and render the data.
- Use the Pipeline Browser to add filters (Clip, Threshold, Contour, Warp, Streamlines…).
- Navigate the 3D view: left-drag to rotate, middle-drag to pan, scroll to zoom.
QGIS
Professional GIS application for viewing, editing, and analysing geospatial data. Runs in a VNC session on the compute node.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | |
| Hours | 2–6 h | |
| CPU Cores | 4–8 | QGIS processing algorithms can use multiple cores |
| Memory | 16–64 GB | Increase for large rasters or complex vector datasets |
- After connecting, QGIS opens in VNC with its map canvas, Layers panel (left), and toolbars.
- Add data via Layer → Add Layer or drag files from the QGIS Browser panel. Supported: GeoTIFF, Shapefile, GeoPackage, WMS/WFS/WCS, PostGIS, and more.
- Use the Processing Toolbox (Processing → Toolbox) for hundreds of spatial analysis algorithms.
- Store your GIS data in
/NFS/scratch/homes/username/for best performance.
gdaladdo adds pyramid overviews that dramatically speed up rendering at different zoom levels.VS Code (Code Server)
Visual Studio Code in your browser, connected directly to the cluster filesystem. Full IDE: IntelliSense, debugging, integrated terminal, Git, and extensions, all running on the compute node.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res or Dev | Dev is ideal for development/testing workflows |
| Hours | 4–8 h | |
| CPU Cores | 4–8 | Increase if running code in the integrated terminal |
| Memory | 8–32 GB | Match to the data your code will process |
- After connecting, VS Code opens in a new browser tab, identical to the desktop application.
- Click File → Open Folder and navigate to your project at
/NFS/scratch/homes/username/my_project. - Edit files from the Explorer panel (left sidebar) with full syntax highlighting and IntelliSense.
- Open the Integrated Terminal (Ctrl+` or View → Terminal). This is a full shell on the compute node: load modules, run code, and submit jobs without leaving VS Code.
Recommended Extensions
Press Ctrl+Shift+X to browse and install. Extensions store in your home directory and persist across sessions:
- Python: IntelliSense, linting, debugging
- Jupyter: Run
.ipynbnotebooks inside VS Code - R: R language support and LSP integration
- GitLens: Enhanced Git history, blame, and diff views
- Remote-SSH: Connect to the cluster from your local VS Code outside the portal
module load, run scripts, and sbatch additional jobs all from within VS Code. It is the most complete all-in-one development environment on IKARUS.SIESTA
SIESTA is a density functional theory (DFT) code for electronic structure calculations and ab initio molecular dynamics. The app runs your SIESTA calculation as an MPI job and then automatically opens XCrySDen to visualise the result, all inside the same VNC session.
| Field | Recommended | Notes |
|---|---|---|
| Input FDF file | Full NFS path | Your main SIESTA input file. Must already exist on the cluster before launching, see below. |
| Pseudopotential / working directory | Full NFS path | A directory containing your pseudopotential files (*.psf or *.psml) and any other auxiliary input files |
| Nodes | 1 | 1–10 supported |
| MPI tasks per node | 4–20 | SIESTA's solver requires the orbital count to be at least the MPI task count; scale down for small test systems (see tip below) |
| Memory | 8–32 GB | 1–350 GB supported |
| Partition | Res | |
| Hours | 1–4 h | 1–72 h supported |
SIESTA does not accept file uploads through the launch form directly. Instead, you stage your files on the cluster first, then point the form at their paths:
- Upload your
.fdfinput file and all pseudopotential files to your home directory using the file browser. - In the launch form, enter the full path to your
.fdffile (e.g./NFS/scratch/homes/username/my_calc/input.fdf). - Enter the full path to the directory containing your pseudopotential files. Everything in that directory is copied into the job's working directory automatically, so pseudopotentials and any other auxiliary files can live together.
%block ChemicalSpeciesLabel. There's no need to name anything explicitly in the form beyond pointing to the directory.- After launching, the session runs SIESTA as a blocking MPI job for the duration of your calculation.
- On completion, your output files (
.out,.XV, and others) are written to a job-specific working directory under your home directory. - Retrieve your results via the file browser, or open a terminal in the VNC session to inspect them directly.
Quickplot
Quickplot (Deltares) is a visualisation and post-processing tool for Delft3D model output, supporting both structured and unstructured grid results. It opens as a standalone graphical application in a VNC session.
| Field | Recommended | Notes |
|---|---|---|
| Partition | Res | |
| Hours | 1–4 h | |
| Memory | 8–16 GB | Increase for very large output files |
- Stage your Delft3D output files on the cluster in advance, using the file browser.
- After connecting, Quickplot opens directly in the VNC session.
- Use Quickplot's own file open dialog to browse to and load your staged output file.
- Use Quickplot's plotting tools to inspect structured or unstructured grid results.
Delft3DFM
Delft3DFM is an ocean and coastal simulation pipeline. It runs your simulation across multiple compute nodes in the background while giving you a live desktop session to monitor progress, and launches Quickplot automatically to visualise results once the run finishes.
| Field | Recommended | Notes |
|---|---|---|
| Model folder | Full NFS path | Must contain dimr_config.xml and your dflowfm/<name>.mdu file. See requirements below. |
| MPI type | Depends on your build | MPICH or Intel MPI |
| Container version | Latest available | |
| Nodes / tasks per node | Depends on model size | Multi-node simulations are supported |
| Memory | 16–64 GB | Match to your model's grid size |
| Partition | Res | |
| Hours | Depends on simulation length |
| File / Folder | Required | Notes |
|---|---|---|
dimr_config.xml | Yes | Must exist at the model root |
dflowfm/<name>.mdu | Yes | Filename is read automatically from dimr_config.xml |
dflowfm/*.pli, *.nc, *.ext, *.bc, *.xyn | Varies | Referenced internally by your MDU file as needed |
dflowfm/output/ | Recommended | Create this directory before running, or the simulation will error |
- Prepare your model folder on the cluster with all required files, using the file browser.
- Fill in the launch form with your model folder path and desired resources, then launch.
- A desktop session opens with a live terminal showing simulation progress.
- On successful completion, Quickplot launches automatically for immediate visualisation.
- If the simulation fails, an error terminal opens showing the end of the error log so you can diagnose the issue; the session stays open for inspection.
FVCOM
FVCOM (Finite Volume Community Ocean Model) is a coastal ocean circulation model. IKARUS provides several physics variants of FVCOM as separate executables, so you can select a build with the physics your simulation needs without having to compile it yourself.
| Variant | Use Case |
|---|---|
| Base / General | 3D baroclinic simulations with water quality, data assimilation, heat flux, and dye tracers, all toggled at runtime |
| Wave-Current Interaction | Coupled wave-current and vortex force simulations. Requires external SWAN wave model output as forcing. |
| Spherical | Large-domain or regional runs using geographic coordinates |
| Sediment Transport (ORIG) | Suspended sediment transport, FVCOM's original scheme |
| Sediment Transport (CSTMS) | Suspended sediment transport, community CSTMS scheme |
| 2D Barotropic | Depth-averaged runs only, no vertical stratification |
| Field | Recommended | Notes |
|---|---|---|
| Physics variant | Base / General for most cases | See table above |
| Run directory | Full NFS path | Must contain your casename.nml file |
| Casename | Matches your .nml file | |
| Physics toggles | As needed | Water quality, dye release, data assimilation, and heat flux toggles, only shown for the Base variant |
| Post-processing visualiser | Quickplot or ParaView | Quickplot for Deltares-format unstructured results; ParaView for general 3D visualisation |
| Nodes | 1–56 | |
| MPI tasks per node | Up to 40 | Total tasks must equal Nprocs in your casename.nml |
| Partition | Res |
- Upload your FVCOM input files (grid, forcing data, and
casename.nml) to your home directory using the file browser. - Set
Nprocsin yourcasename.nmlto match Nodes × MPI tasks per node exactly. - Select the physics variant matching your simulation's requirements.
- If using the Base variant, set any physics toggles you need. The app patches your
casename.nmlautomatically before launch. - Choose a post-processing visualiser and optionally specify the output file to open automatically.
- Launch. A live terminal opens showing simulation progress.
- On success, your chosen visualiser opens automatically with the results. On failure, an error terminal opens showing the end of the error log.
casename_dep.dat) must already exist in your run directory before submitting, and its partition count must match your Nodes × tasks-per-node setting.WRF-Chem
WRF-Chem is an atmospheric chemistry and weather simulation pipeline. Unlike the other interactive apps, WRF-Chem runs as a background batch job with no live desktop; you submit it and monitor progress through Active Jobs and the file browser, and it runs the full preprocessing and simulation pipeline automatically.
| Field | Recommended | Notes |
|---|---|---|
| Run directory | Full NFS path | Must contain a pre-configured namelist.input and namelist.wps |
| Meteorological data path | Full NFS path | Directory of GRIB files from your chosen global model |
| Global model | GFS, IFS, or AIFS | Determines boundary condition interval and available start hours |
| Start date | yyyy/mm/dd | |
| Start hour | Model-dependent | GFS/AIFS: 00, 06, 12, 18 UTC. IFS: 00 or 12 UTC only. |
| Nodes | 2 or more | |
| MPI tasks per node | 40 | |
| Memory | 64 GB | Up to 350 GB |
| Hours | 24 h | Up to 120 h |
| Partition | Res |
- Run
geogrid.exeonce for your domain beforehand; the requiredgeo_em.d01.nc(and additional domains if nested) must already be present in your run directory. namelist.inputmust be fully configured for your domain (grid spacing, dimensions, physics, and chemistry options). The app only sets the start date/time and interval automatically.namelist.wpsmust be configured for your domain and projection. The app sets the start date, interval, and met data source automatically.- GRIB input files from your chosen global model must already be staged in your met data directory.
- Complete the prerequisites above, staging your run directory and meteorological data via the file browser.
- Fill in the launch form and submit.
- The pipeline runs automatically: ungrib, then metgrid, then real.exe, then wrf.exe, in sequence.
- Monitor progress via Active Jobs. There is no live desktop or terminal for this app.
- When the job completes, retrieve your output files via the file browser.
OSRA
OSRA (Oil Spill Risk Analysis) runs an ensemble of particle-tracking oil spill simulations across the Arabian Gulf, producing a probabilistic hazard atlas showing where spilled oil is likely to beach under different seasonal conditions. On completion, QGIS opens automatically with the hazard atlas loaded.
| Field | Recommended | Notes |
|---|---|---|
| Run type | Test, for exploration | Test: a single simulation task, runs in around 10–30 minutes. Full: submits a full 200-task ensemble. |
| Source location | Any of 10 Gulf locations | Test runs only |
| Season | Spring, summer, autumn, or winter | Test runs only |
| Year | 2018–2022 | Test runs only |
| Number of particles | 500 for a quick test; 10,000 for standard runs | |
| Partition | Res | |
| Hours | 4 h | |
| Memory | 32 GB |
Test Run
Choose a single source location, season, and year to run one simulation task interactively. A live terminal in the VNC session shows progress; on completion, trajectory plots open automatically.
Full Run
Submits a complete 200-task ensemble covering all source locations, seasons, and years. A monitoring terminal in the VNC session tracks task progress. Once the full ensemble and post-processing complete, QGIS opens automatically with the resulting hazard atlas loaded.
Tools
Additional research tools integrated directly into the portal, separate from the Interactive Apps. Accessed from the Tools menu in the top navigation bar.
- Log in to the portal at khpc.kisr.edu.kw/ood/
- Click Tools in the top navigation bar.
- Select the tool you want from the drop-down.
Available Tools
| Tool | Purpose |
|---|---|
| MLFlow | Experiment tracking, model registry, and artifact storage for machine learning and computational workflows |
MLFlow
MLFlow is an open-source platform for managing the full machine learning and computational experiment lifecycle: parameter and metric logging, model versioning, artifact storage, and reproducibility across runs.
| Capability | Description |
|---|---|
| Experiment tracking | Log parameters, metrics, and tags from any Python, R, or shell script. Compare runs across Slurm jobs without manual spreadsheets. |
| Artifact store | Save model files, plots, checkpoints, and datasets as versioned artifacts, stored on shared cluster storage and accessible from any node. |
| Model registry | Version, stage, and move models through a lifecycle (Staging → Production), giving you a central catalogue with lineage tracking. |
| Web UI | Browse experiments, compare runs side-by-side, and visualise metrics, accessible in your browser via the portal. |
| REST API & SDK | Full REST and Python SDK for programmatic access. Slurm jobs log directly to the tracking server with no manual steps. |
| Language support | Python (mlflow package), R (mlflow R package), and REST for any other language. |
Accessing MLFlow
| Interface | How to reach it |
|---|---|
| MLFlow UI | Click Tools → MLFlow in the portal, or go directly to khpc.kisr.edu.kw/mlflow/ |
| REST API | https://khpc.kisr.edu.kw/mlflow/api/2.0/, for use with the Python SDK or your own scripts from outside a Slurm job |
Logging Experiments from Slurm Jobs
Use the same public tracking URI everywhere, whether inside a Slurm job script or from your own machine:
https://khpc.kisr.edu.kw/mlflow/import mlflow
# Set tracking server (or use env var MLFLOW_TRACKING_URI)
mlflow.set_tracking_uri("https://khpc.kisr.edu.kw/mlflow/")
# Create or reuse an experiment
mlflow.set_experiment("water_quality_model_v2")
with mlflow.start_run():
# Log parameters
mlflow.log_param("learning_rate", 0.001)
mlflow.log_param("epochs", 100)
mlflow.log_param("batch_size", 32)
# --- your training code here ---
# Log metrics per epoch
for epoch in range(100):
loss = train_one_epoch(...)
mlflow.log_metric("loss", loss, step=epoch)
# Log artifacts (model file, plots)
mlflow.log_artifact("model.pkl")
mlflow.log_artifact("training_curve.png")
# Log the model itself with schema
mlflow.sklearn.log_model(model, "random_forest_model")library(mlflow)
mlflow_set_tracking_uri("https://khpc.kisr.edu.kw/mlflow/")
mlflow_set_experiment("air_quality_forecast")
with(mlflow_start_run(), {
mlflow_log_param("ntree", 500)
mlflow_log_param("mtry", 3)
# --- your model training ---
mlflow_log_metric("rmse", rmse_value)
mlflow_log_artifact("forecast_plot.pdf")
})For any language or tool that respects environment variables, set MLFLOW_TRACKING_URI in your Slurm job script. MLFlow's autologging can automatically capture parameters and metrics for supported frameworks (scikit-learn, TensorFlow, PyTorch, XGBoost, LightGBM, Keras):
export MLFLOW_TRACKING_URI="https://khpc.kisr.edu.kw/mlflow/"
export MLFLOW_EXPERIMENT_NAME="batch_simulation_run"
python3 -c "import mlflow; mlflow.autolog()"A pre-configured MLFlow Experiment template is available in the Job Composer. It sets MLFLOW_TRACKING_URI and a starting experiment name automatically. Click + New Job → From Template and select it to get started without writing the boilerplate yourself, then edit the script path and parameters for your own run.
Naming Convention
MLFlow experiment names are global on the shared tracking server. Two researchers using the same experiment name will have their runs merged into the same experiment. Adopt a naming convention to avoid collisions:
| Convention | Example | Notes |
|---|---|---|
<department>/<project>/<descriptive_name> | env/airquality/lstm_v3 | Hierarchical, maps to directory-like paths |
<username>_<project>_<date> | abduljalil_wq_20260501 | Simple, prevents collisions between users |
<PI_name>/<project> | hussain_lab/groundwater_model | Groups all experiments under a PI or group |
Viewing Results
- Click Tools → MLFlow in the portal, or open khpc.kisr.edu.kw/mlflow/ directly.
- Select an experiment from the left sidebar to see all its runs.
- Click a run to see its logged parameters, metrics, and artifacts.
- Use the comparison view to select multiple runs and compare metrics side-by-side.
- Artifacts are served through MLFlow directly. Model files and plots can be downloaded from the browser.
Cluster Hardware
Specifications for the IKARUS HPC compute nodes, storage systems, network, and partition structure.
| Cluster Name | IKARUS |
| Operator | Kuwait Institute for Scientific Research (KISR) |
| Job Scheduler | SLURM Workload Manager |
| Web Portal | IKARUS Portal: khpc.kisr.edu.kw/ood/ |
| SSH Access | hpc.kisr.edu.kw (port 22) → login nodes clavis1 / clavis2 |
| Operating System | Rocky Linux 8 (RHEL-compatible) |
| Total Compute Nodes | 57 (meteor1–meteor56 + fortis) |
| Partitions | Res · def1 · Dev |
| Primary Shared Filesystem | NFS: user homes at /NFS/scratch/homes/<username> |
| Software Modules | /NFS/shared/modules/, loaded via the module command |
Compute Nodes
| Quantity | 56 nodes |
| Node names | meteor1 through meteor56 |
| Processor | 2× Intel Xeon Gold 6248 @ 2.50 GHz |
| Cores per node | 40 cores (2 sockets × 20 cores) |
| RAM per node | 350 GB |
| Partitions | Res def1 |
| Purpose | Primary general-purpose compute nodes for all standard batch and interactive workloads |
| Quantity | 1 node |
| Node name | fortis |
| Processor | 4× Intel Xeon Gold 6252 @ 2.10 GHz |
| Cores | 96 cores (4 sockets × 24 cores) |
| RAM | 1.5 TB |
| Partitions | Dev |
| Purpose | Special-purpose node; contact administrators for use-case guidance |
sinfo -N -o "%-20N %-10c %-12m %-10T %-10P" in the shell terminal.Storage Systems
| Mount Point | Purpose | Access | Notes |
|---|---|---|---|
/NFS/scratch/homes/<username> | User home directory (Scratch) | Personal to each user | 460 TB total scratch storage. Default working location, shared across all nodes via NFS. Store scripts, data, and job output here. |
/NFS/shared/ | Shared software and libraries | Read-only for regular users | 230 TB shared storage. Contains modules, Conda environments, shared libraries, and Jupyter kernels. |
/NFS/shared/projects/ | Project storage | Personal to each project's members | Exclusive for permenant project data. Each project is partitioned from each other. Requested from IKARUS administration. |
Checking Your Disk Usage
du -sh /NFS/scratch/homes/${USER} # total home directory usage
du -sh /NFS/scratch/homes/${USER}/*/ # breakdown by subdirectory
find /NFS/scratch/homes/${USER} -type f \
-printf '%s %p\n' | sort -rn | head -20 # find largest filesNetwork
| Inter-node fabric | InfiniBand / 1GbE |
| Internet Routing |
Partitions & Access Levels
Which partitions you can access depends on your account. The Interactive Apps launch forms automatically show only the partitions available to you.
| Partition | Nodes | Cores | RAM | Max Walltime | Access |
|---|---|---|---|---|---|
| Res | 6 meteor nodes | 240 | 2.1 TB | 36 hours | All users |
| def1 | 50 meteor nodes | 2,000 | 17.5 TB | 4 hours | project, developer |
| Dev | fortis (1 node) | 192 | 1.5 TB | 72 hours | developer |
| Partition | Description & Objective |
|---|---|
| Res | Dedicated to experimental, early-stage, or small-scale scientific projects and proof-of-concept simulations. Uses less than 10% of total cluster resources. Aimed at providing a flexible, accessible platform for application development and testing before scaling to production, without impacting high-priority workloads. |
| def1 | Dedicated to KISR projects (for the duration of the project) and collaborators at other research institutions (limited lifespan). Uses approximately 85% of total cluster resources. Aimed at supporting integrated applications, ongoing KISR projects, and related workloads. |
| Dev | Dedicated to experimental or small-scale scientific developments and resource assessment. Uses less than 5% of total cluster resources. Aimed at providing a long-runtime platform for application development and testing before scaling to production, without impacting high-priority workloads. |
System Status
A live snapshot of overall cluster availability and load, visible directly on your portal dashboard.
The IKARUS Cluster Status panel is shown in two places, both displaying the same live information:
- Automatically on the portal dashboard, directly below the Recently Used Apps tiles, every time you log in.
- By clicking Clusters → System Status in the top navigation bar from anywhere in the portal.
What It Shows
The panel gives a quick, at-a-glance read of how busy the cluster currently is:
| Field | Description |
|---|---|
| Nodes Available | Total number of compute nodes in the cluster (57), with a bar showing the percentage currently allocated to jobs |
| Processors Available | Total CPU cores across the cluster (2,432), with a bar showing the percentage currently in use |
| GPUs Available | Always 0. IKARUS compute nodes do not have GPUs, so this bar has nothing to divide by and will show NaN% in use. This is expected and not an error. |
| Jobs Running | Number of jobs currently executing on compute nodes across the whole cluster |
| Jobs Queued | Number of jobs submitted and waiting for resources to become available |
SSH Access
Traditional terminal access, for advanced users and automated workflows.
Host: hpc.kisr.edu.kw · Port: 22 · Username: your KISR username
Windows (MobaXterm)
- Download MobaXterm (Portable edition) from mobaxterm.mobatek.net.
- Unzip and launch
MobaXterm_Personal_*.exe. - Click Session → SSH. Set Remote host:
hpc.kisr.edu.kw, Username: your KISR username, Port: 22. - Click OK and enter your password when prompted (input is hidden).
- A successful login shows a prompt like:
[hpcdemo@clavis1 ~]$ Connecting on Mac / Linux
- Open Terminal (Mac: Applications → Utilities → Terminal).
- Run the SSH command with X11 forwarding enabled:
$ ssh -Y hpcdemo@hpc.kisr.edu.kw- On first connection, type
yesto accept the RSA fingerprint. - Enter your password when prompted (not displayed on screen).
- On success you will see the shell prompt:
[hpcdemo@clavis1 ~]$
-Y flag enables X11 forwarding, allowing GUI applications launched in the terminal to render on your local screen.Linux Commands
Essential shell commands for working on IKARUS. The default shell is BASH (Bourne Again Shell). Linux paths and filenames are case-sensitive.
Navigation
pwd # print current directory path
cd /NFS/scratch/homes/hpcdemo/jobs # navigate to a directory
cd ~ # return to home directory
cd .. # go up one level (parent dir)
cd - # return to previous directory
ls # list directory contents
ls -l # long listing (permissions, size, date)
ls -a # include hidden files (dot files)
ls -la # combine long + hidden
ls *.py # list all Python files (wildcard)File Operations
mkdir my_dir # create a new directory
mkdir -p project/data/raw # create nested directories at once
touch file.txt # create an empty file
cp file.txt backup.txt # copy to a new filename
cp file.txt /path/to/dest/ # copy to a different directory
cp *.csv /data/ # copy all CSV files
cp -r project/ project_backup/ # copy directory recursively
mv file.txt newname.txt # rename a file
mv file.txt /path/to/dest/ # move a file
mv old_dir/ new_dir/ # rename a directory
rm file.txt # delete a file (PERMANENT)
rm -rf directory/ # delete directory + contents (PERMANENT)rm is immediate and cannot be undone. Always double-check before deleting.Viewing File Contents
cat file.txt # print entire file to screen
cat file.txt >> dest.txt # append file to another
head file.txt # show first 10 lines
head -n 30 file.txt # show first 30 lines
tail file.txt # show last 10 lines
tail -n 30 file.txt # show last 30 lines
tail -f output.log # follow file live (Ctrl+C to stop)
tail -F app.log # follow + handle log rotation
less file.txt # paginated viewer (q to quit, / to search)
more file.txt # simpler pager (Space=next page, q=quit)
more -5 file.log # show 5 lines at a time
more +10 file.log # start from line 10tail -f output.log is invaluable for watching a running job's output file in real time. Press Ctrl+C to stop following.Searching
grep "error" output.log # find lines containing "error"
grep -i "warning" output.log # case-insensitive search
grep -r "pattern" ./src/ # search recursively in directory
grep -v "debug" output.log # exclude lines containing "debug"
grep -n "error" output.log # show line numbers with matches
find . -name "results" # find a file/dir named "results"
find . -name "*.py" # find all Python files
find . -type f -name "*.log" # find only files (not dirs) matching *.log
wc -l file.txt # count lines
wc -w file.txt # count words
wc -c file.txt # count bytesOther Useful Commands
| Command | Description | Example |
|---|---|---|
date | Display current date, time, and timezone | date |
cal | Show calendar for the current month | cal |
time | Measure how long a command takes | time python3 script.py |
sleep | Pause for N seconds (useful in shell scripts) | sleep 5 |
sort | Sort lines of a text file | sort results.txt |
file | Determine file type from contents (not extension) | file data.bin |
tac | Print file contents in reverse line order | tac output.log |
echo | Print text to screen or redirect to file | echo "done" >> status.txt |
which | Show full path of an executable | which python3 |
env | Show all environment variables | env | grep PATH |
history | Show recent command history | history | tail -20 |
# Redirect and pipe examples
sort file.txt > sorted.txt # sort and save to new file
sort file.txt >> sorted.txt # sort and append to existing file
cat file.txt | grep "error" | wc -l # count error lines using pipes
ls -la | less # browse long directory listingsFile Permissions & Compression
Understanding and changing Linux file permissions, and working with compressed archives.
Permission Notation
Every file and directory has a 10-character permission string shown by ls -l. For example: drwxr-xr-x
| Position | Characters | Meaning |
|---|---|---|
| 1st | d or - | d = directory · - = regular file |
| 2nd–4th | rwx | Owner (user) permissions |
| 5th–7th | r-x | Group permissions |
| 8th–10th | r-x | Others (everyone else) permissions |
Each permission letter:
- r (Read): view or copy the file, or list directory contents
- w (Write): modify the file, or create/delete files in a directory
- x (Execute): run the file as a program, or enter a directory
- –: permission not granted
chmod: Changing Permissions
Format: chmod [who][action][permission] file
| Who | Action | Permission |
|---|---|---|
u: owner (user) | + add | r read |
g: group | - remove | w write |
o: others | = set exactly | x execute |
a: all (u+g+o) |
chmod u+x script.sh # make script executable by owner
chmod a+x script.sh # executable by everyone
chmod u+rwx,g+rx file.sh # owner: full access; group: read+execute
chmod o-x * # remove execute from others on all files
chmod g-w sensitive.txt # remove group write permission
chmod 755 script.sh # octal: rwxr-xr-x
chmod 644 data.csv # octal: rw-r--r--ls -l to inspect permissions and ownership, and chmod to correct them.Compression: tar, gzip
gzip / gunzip: Single File Compression
gzip filename.c # compress → filename.c.gz
gunzip filename.c.gz # decompresstar: Archive and Compress Directories
| Flag | Meaning |
|---|---|
-c | Create a new archive |
-x | Extract files from archive |
-t | List contents without extracting |
-z | Use gzip compression (.tar.gz) |
-v | Verbose: show files being processed |
-f | Specify archive filename (always last flag) |
tar -czvf archive.tar.gz my_project/ # create compressed archive of directory
tar -xzvf archive.tar.gz # extract archive here
tar -xzvf archive.tar.gz -C /target/ # extract to a specific directory
tar -tzvf archive.tar.gz # list contents without extracting
gzip -l archive.tar.gz # show compression ratio and sizestar -czvf mydata.tar.gz data/ in the shell terminal first. Then download the single archive file; much faster and more reliable than downloading many individual files.Environment Modules
IKARUS uses a module system so multiple versions of the same software can coexist without conflicts. Loading a module configures your environment for that software.
All modules are stored at /NFS/shared/modules/ and are accessible on every node.
Module Commands
module avail # list ALL available modules
module avail python # search for modules matching "python"
module avail jasper # find all versions of jasper
module show jasper/2.0.14 # see what the module sets (PATH, libs…)
module add jasper # load default version (same as module load)
module load jasper/2.0.14 # load a specific version
module list # show currently loaded modules
module rm jasper # unload a specific module
module rm jasper/2.0.14 # unload a specific version
module purge # unload ALL loaded modulesmodule add and module load are identical, both load a module. Without specifying a version, the highest alphanumeric version (or the tagged default, shown with D in module avail) is loaded.Using Modules in SLURM Scripts
Modules loaded in your interactive terminal session are not inherited by batch jobs. Always include module load commands inside your submission script:
#!/bin/bash
#SBATCH --job-name=python_job
#SBATCH --partition=Res
#SBATCH --cpus-per-task=8
#SBATCH --mem=32G
#SBATCH --time=01:00:00
# Load modules here (inside the script)
module purge # start from a clean state
module load python/3.10
module load openmpi/4.1.5
# Your code
python3 my_script.pySLURM Commands
SLURM is the workload manager on IKARUS. These commands are used from the login node (clavis) to submit, monitor, and manage jobs.
Submitting Jobs
sbatch: Submit a Batch Script
Submits a script to run asynchronously on compute nodes. Returns immediately with the assigned Job ID.
[hpcdemo@clavis2 ~]$ sbatch my_job.sh
Submitted batch job 568srun: Interactive Session on a Compute Node
Requests resources and gives you a shell on a compute node. Blocks until resources are available.
[hpcdemo@clavis2 ~]$ srun --partition=Res --nodes=1 \
--cpus-per-task=4 --mem=8G --time=01:00:00 bash -lMonitoring Jobs
squeue: View the Queue
squeue # all jobs across all users
squeue -me # your jobs only
squeue -u hpcdemo # jobs for a specific user
squeue -p Res # jobs in the Res partition
squeue --start # estimated start times for pending jobs
squeue -o "%.10i %.9P %.20j %.8u %.8T %.10M %.6D %R" # custom output formatsqueue State Codes
| Code | State | Meaning |
|---|---|---|
| R | RUNNING | Job is actively executing on compute nodes |
| PD | PENDING | Waiting for resources or priority; reason shown in last column |
| CG | COMPLETING | Job finishing, cleaning up processes |
| CD | COMPLETED | Finished successfully (exit code 0) |
| F | FAILED | Exited with non-zero status |
| CA | CANCELLED | Cancelled by user or administrator |
| TO | TIMEOUT | Exceeded the requested wall-clock time |
| OOM | OUT_OF_MEMORY | Job exceeded the requested memory allocation |
sacct: Job History and Accounting
Query completed and historical jobs from the SLURM accounting database.
sacct --user=$USER --starttime=today # your jobs from today
sacct --user=$USER --starttime=2025-08-01 # since a specific date
sacct --jobs=568 # details for job 568
sacct --jobs=568 --format=JobID,State,ExitCode,Elapsed,MaxRSS # custom fieldsscancel: Cancel a Job
scancel 568 # cancel job 568 (must be your job)
scancel -u $USER # cancel ALL your jobs
scancel -u $USER -t PENDING # cancel only your pending jobssinfo: Cluster State
sinfo # partition and node summary
sinfo -N -o "%-20N %-10c %-12m %-10P" # per-node: name, CPUs, RAM(MB), partition
sinfo -p Res # show only the Res partitionsstat: Live Resource Usage of a Running Job
sstat --jobs=568 # all resource stats for job 568
sstat --jobs=568 --format=JobID,AveCPU,AveRSS,MaxRSSMinimal Submission Script
#!/bin/bash
######## SLURM Directives ########
#SBATCH --job-name=my_job
#SBATCH --partition=Res
#SBATCH --nodes=1
#SBATCH --ntasks-per-node=1
#SBATCH --cpus-per-task=8
#SBATCH --mem=32G
#SBATCH --time=02:00:00
#SBATCH --output=%j.out # %j is replaced by Job ID
#SBATCH --error=%j.err
######## End of Directives ########
module purge
module load python/3.10
cd /NFS/scratch/homes/${USER}/my_project
python3 my_script.py#SBATCH Directive Reference
| Directive | Example Value | Description |
|---|---|---|
--job-name | my_sim | Name shown in squeue output |
--partition | Res | Target queue: Res, def1, or Dev |
--nodes | 1 | Minimum number of nodes to allocate |
--ntasks | 16 | Total number of MPI tasks |
--ntasks-per-node | 8 | MPI tasks per allocated node |
--cpus-per-task | 4 | CPU threads per MPI task (default: 1) |
--mem | 64G | RAM per node, units: K, M, G, T |
--mem-per-cpu | 4G | RAM per CPU core (alternative to --mem) |
--time | 12:00:00 | Max wall-clock time (D-HH:MM:SS or HH:MM:SS) |
--array | 1-50 | Submit as job array with indices 1 through 50 |
--output | %j.out | Standard output file (%j = Job ID, %a = array task ID) |
--error | %j.err | Standard error file |
--mail-type | END,FAIL | Send email on END, FAIL, BEGIN, or ALL |
--mail-user | user@kisr.edu.kw | Email address for notifications |
--test-only | (no value) | Validate script and show estimated start time without submitting |
--dependency | afterok:567 | Hold this job until job 567 completes successfully |
Job Array Example
Job arrays submit the same script N times, each with a unique SLURM_ARRAY_TASK_ID. Useful for running the same analysis on many input files.
#!/bin/bash
#SBATCH --job-name=array_example
#SBATCH --partition=Res
#SBATCH --array=1-16 # creates 16 jobs, IDs 1..16
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=4G
#SBATCH --time=00:30:00
#SBATCH --output=logs/job_%a.out # %a = array task ID
module load python/3.10
# Use SLURM_ARRAY_TASK_ID to select input
INPUT_FILE="data/input_${SLURM_ARRAY_TASK_ID}.csv"
OUTPUT_FILE="results/output_${SLURM_ARRAY_TASK_ID}.csv"
echo "Processing task ${SLURM_ARRAY_TASK_ID}: ${INPUT_FILE}"
python3 process.py --input ${INPUT_FILE} --output ${OUTPUT_FILE}MPI Parallel Job Example
#!/bin/bash
#SBATCH --job-name=mpi_job
#SBATCH --partition=Res
#SBATCH --nodes=2
#SBATCH --ntasks=40 # 20 tasks per node × 2 nodes
#SBATCH --cpus-per-task=1
#SBATCH --mem-per-cpu=4G
#SBATCH --time=00:30:00
module purge
module load openmpi/4.1.5
module load python/3.10
cd ${SLURM_SUBMIT_DIR}
# Run 5 parallel MPI jobs in background, each with 4 tasks
mpirun -n 4 python3 pi-digits.py 10 &
mpirun -n 4 python3 pi-digits.py 15 &
mpirun -n 4 python3 pi-digits.py 20 &
mpirun -n 4 python3 pi-digits.py 25 &
mpirun -n 4 python3 pi-digits.py 1000 &
# Wait for ALL background tasks to finish
wait
# Then run the aggregation step
mpirun -n 20 python3 sum-digit.py
exit 0wait command is critical in the MPI example. Without it, the aggregation step starts before all parallel tasks have completed.