Storage and quotas

Shared Storage

Currently users have 3 main storage areas share across every node. Each node has access to to this storage at all times and data is shared. Be careful with parallel jobs trying to write to the same filename!

  • /nfs/home/USERNAME - This is your Home Directory, each user has a 50 GB quota limit. The data is replicated off site and backed up regularly by Digital Solutions. Home directory tips

  • /nfs/scratch/USERNAME - This is your scratch space, each user has a 5 TB quota limit. This data is not backed up! Scratch directory tips

  • /nfs/scratch/noquota-volatile - This is a shared, quota-free scratch area with approximately 80 TB of storage, intended for large temporary datasets and workflows that exceed standard user quotas. There is no quota enforcement here. Data stored here is not backed up! and should be considered temporary. Files in this filesystem may be deleted periodically without notice, so important data should not be stored here.

Note: Home directory quotas cannot be increased, however if you need more space in your scratch folder let us know.

To view your current quota and usage use the vuw-quota command, for example:

<username@raapoi-login:~$ vuw-quota 

User Quotas

                       Storage  Usage (GB)  Quota (GB)     % Used 
            /nfs/home/<username>      18.32       50.00     36.63%

         /nfs/scratch/<username>       0.00     5000.00      0.00%

Per Node Storage

Each compute node has local storage you can use at /tmp.
This storage is not shared so a program running on amd01n02 will not be able to see data stored on node amd01n04's /tmp storage. Additionally, you can only access /tmp on any given node via a job running on that node.

On the quicktest, parallel, and GPU nodes the /tmp storage is very fast nvme storage with 1.7TB total space. On bigmem nodes this storage is 1.5TB. The longrun partition has two nodes, bigtmp01 and bigtmp02, each with 25TB of local /tmp storage.

Temp Disk Tips

If you use the /tmp storage it is your responsibility to copy data to the /tmp and clean it up when your job is done.

For more info see Temp Disk Tips.


Storage Performance

graph TD A(Home and Research Storage) --> B B[Scratch] --> D D[local /tmp on compute nodes]
Figure 1: Storage speed hierarchy. The slowest storage is your user home directory as well as any mounted research storage. The trade off for this is that this data is replicated off site as well as backed up by Digital Solutions. The fastest is the local /tmp space on the nodes - it is usually deleted shortly after you logout and only visible to the node it's on, but it is extremely fast with excellent IO performance.

Storage tips

Home Directory Tips

Home directories have a small quota and are on fairly slow storage. The data here is backed up. It is replicated off site live as well as periodically backed up to off site tape. In theory data here is fairly safe, even in a catastophic event it should be recoverable eventually.

If you accidentally delete something here it can be recovered with a service desk request.

While this storage is not performant, is is quite safe and is a good place for you scripts and code to live. Your data sets can also be on your home if they fit and the performance doesn't cause you any problems.

For bigger or faster storage see the Scratch page.


Scratch Tips

The scratch storage is provided by a dedicated ZFS storage system and is available to all Raapoi users at /nfs/scratch/<username>.

Scratch storage is intended for actively used research data and temporary working files. It is not backed up, so important data should not be stored exclusively on scratch storage. While the storage system includes RAIDZ2 disk redundancy and redundant power supplies, it remains a single storage system and cannot protect against all hardware failures.

Each user is allocated a default quota of 5TB of scratch space. You can check your quota and current usage by running vuw-quota. If your research requires additional storage, please contact the support team.

Scratch storage is a shared resource. Although users are allocated generous quotas, the total storage available is shared across all users. Researchers are encouraged to regularly remove data that is no longer required and archive completed work elsewhere. When storage usage becomes high, the Research Computing team may contact users with large allocations and ask them to review or clean up their data.

To check how much space is free on the scratch storage for all users, on Rāpoi:

df -h | grep scratch  #df -h is disk free with human units, | pipes the output to grep, which shows lines which contain the word scratch

This storage is not backed up at all. It is on a raid array so if a hard drive fails your data is safe. However in the event of a more dramatic hardware failure, earthquakes or fire - your data is gone forever. If you accidentally delete something, it's gone forever. If an Admin misconfigures something, your data is gone (we try not to do this!).

It is your responsiblilty to backup your data here - a good place is to use Digital Solutions Research Storage (see Connecting to SoLAR).

Scratch is also not a place for your old data to live forever, please clean up datasets you're no longer using!


Temp Disk Tips

This storage is very fast on the AMD nodes and GPU nodes. It is your job to move data to the tmp space and clean it up when done.

There is very little management of this space and currently it is not visible to slurm for fair use scheduling - in other words someone else might have used up most of the temp space on the node! This is generally not the case though.

A rough example of how you could use this in an sbatch script

#!/bin/bash
#
#SBATCH --job-name=bash_test
#
#SBATCH --partition=quicktest
#
#SBATCH --cpus-per-task=2 #Note: you are always allocated an even number of cpus
#SBATCH --mem=1G
#SBATCH --time=10:00

# Do the needed module loading for your use case
module load etc

#make a temporary directory with your usename so you don't tread on others
mkdir /tmp/<username>

#Copy dataset from scratch to the local tmp on the node (could also use rsync)
cp -r /nfs/scratch/<user>/dataset /tmp/<user>/dataset

Process data against /tmp/<user>/dataset
Lets say the data is output to /tmp/<user>/dataoutput/

# Copy data from output to your scratch - I suggest not overwriting your original dataset!
cp -r /tmp/<user>/dataoutput/* /nfs/scratch/<user>/dataset/dataoutput/

# Delete the data you copy to and created on tmp
rm -rf /tmp/<user>  #DANGER!!