ZFS
ZFS is a file system and logical volume manager.
Unless otherwise noted, this article will contain instructions which will apply to Linux, FreeBSD, and Solaris.
Installation
ZFS is included with Solaris and FreeBSD. Linux support is offered by various open source projects including the ZFS on Linux project.
Linux
Download the source or prebuilt packages from ZFS on Linux's website http://zfsonlinux.org/
Packages from ZFS on Linux will make use of dkms and can be rebuilt when a new kernel is used. (eg: dkms build -m zfs -v 0.6.5 -k 4.2.5-300.fc23.x86_64)
Introduction
ZFS stores data inside datasets which are contained inside storage pools (zpool) which are created on top of block devices (vdev). When creating zpools, the vdevs used can be configured similar to different RAID levels for different redundancies. There are also many other options for ZFS which will be covered later.
Creating vdevs - The Block Storage
To prevent write-amplification, when using hard drives with Advanced Format, partitions used for the vdevs should be aligned properly (either to the 512b or 4k sectors). For Advanced Format drives, partitions should be aligned to match the 4K boundary (8 times 512b) and typically 2048 sectors (equivalent to 1 MB of disk space) is used.
Furthermore, a safety margin created near the end of the disks in case replacement disks are slightly smaller than the current disks used. The examples used below will use a 200 MB safety margin. In sectors, this is 200 * 1024*1024 bytes / 512 bytes/sector = 200 * 2048 sectors = 409600 sectors.
To determine the total number of sectors your disks:
- On Linux, use
fdisk -l - On FreeBSD,
diskinfo -v
Create the data partitions using gpart for each of your data disks. In the example below, I will be creating a ZFS pool across 3 disks on a FreeBSD machine using ada1, ada2, ada3. Linux machines will do something similar, but the disk device names will be named differently (eg: /dev/sda1, /dev/sdb1, etc.).
# gpart create -s GPT /dev/ada1
# gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk01 /dev/ada1
# gpart create -s GPT /dev/ada2
# gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk02 /dev/ada2
# gpart create -s GPT /dev/ada3
# gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk03 /dev/ada3
Note: The partitions created above starts 2048 sectors from the beginning and ends 200 MB from the end. The total size of the data partition is 1953113520 sectors. This value may be different for you.
Creating ZFS Pool
With the data partitions created, you should be able to create your new ZFS volume. The most basic way to create a new zpool is to use the zpool create command.
In the example on this page, we will use a zpool named 'storage' across 3 disks using raidz1 (1 redundant disk).
# zpool create -f storage raidz1 gpt/disk01 gpt/disk02 gpt/disk03
Parameters specific to the zpool can also be passed using the -o option. For instance, if using Advanced Format disks with 4096 sector sizes, you might want to set ashift=12 so ZFS does not write blocks straddling sectors like so:
# zpool create -o ashift=12 -f storage raidz1 gpt/disk01 gpt/disk02 gpt/disk03
Once created, you can see the status of the pool using zfs status.
# zpool status
pool: pool-name
state: ONLINE
scan: none requested
config:
NAME STATE READ WRITE CKSUM
storage ONLINE 0 0 0
raidz1-0 ONLINE 0 0 0
gpt/disk01 ONLINE 0 0 0
gpt/disk02 ONLINE 0 0 0
gpt/disk03 ONLINE 0 0 0
errors: No known data errors
Zpool Parameters
Use the zdb command to look up a zpool's parameters. For instance, to see the ashift value of the newly created pool above:
# zdb | grep ashift
ashift: 12
Creating ZFS Filesystems
With a zpool in place, we can start creating ZFS filesystems, also called datasets, and are analogous to volumes in LVM.
The most basic usage is:
# zfs create storage
Parameters for this specific data set can be specified using the -o option. We'll talk more about options later.
# zfs create -o compress=lz4 storage/logs
Common ZFS Options
Here are some options that are pretty useful:compress=lz4 - For newer ZFS pools (versions >= 5000), lz4 compression should always be enabled because files that are not compressible will not be compressed which makes it highly efficient. Most CPUs should be able to keep up with IO.
copies=N - This tells ZFS to maintain multiple copies of the file for redundancy. In case of bad sectors, ZFS will have the ability to recover the bad data.
To list the data sets:
# zfs list
NAME USED AVAIL REFER MOUNTPOINT
storage 153K 1.11T 153K /storage
storage/logs 153K 1.11T 153K /storage/logs
Get/Set Volume Parameters
To get a ZFS parameter, use the zfs get command. Eg:
# zfs get storage compressratio
NAME PROPERTY VALUE SOURCE
storage compressratio 1.01x -
To set a ZFS parameter, use the zfs set command. Setting a parameter such as enabling/disabling compression will not alter the existing data. If you wish to compress a volume, you will need to move data out and then back in for it to take affect.
ZFS Snapshot
A ZFS snapshot is a read-only copy of the file system from a previous state. Because ZFS makes use of Copy-On-Write, snapshots take no additional space.
Listing Snapshots
It can be listed by running zfs list -t snapshot.
# zfs list -t snapshot
NAME USED AVAIL REFER MOUNTPOINT
storage@20120820 31.4G - 3.21T -
storage@20120924 134G - 4.15T -
storage@20121028 36.2G - 4.26T -
storage@20121201 33.2M - 4.55T -
The USED column shows the amount of space used by the snapshot. This amount will go up as files from the snapshot are deleted since the space freed cannot be reclaimed until the snapshot is deleted.
The REFER column shows the actual size of the pool at the snapshot's timepoint.
Creating Snapshots
To create a new snapshot, run zfs snapshot storage@snapshot-name. The snapshot name can be anything, but a datestamp is typically what I use.
No Additional Space is Used
ZFS snapshots make use of Copy-On-Write (COW) and will not use any additional space. However, deleting files that are part of an existing snapshot will not reclaim space. Instead, the storage capacity will appear to go down since the amount of available space for new data remains the same. Eg: A volume at 1TB of 2TB capacity has a 0.5TB file deleted. Because the 0.5TB file is still stored in a snapshot, the total capacity is now 2TB - 0.5TB and the reported disk usage will now be 0.5TB of 1.5TB.
If I were to run zfs snapshot storage@snapshot-name now, listing the snapshots will yield:
# zfs list -t snapshot
NAME USED AVAIL REFER MOUNTPOINT
storage@20120820 31.4G - 3.21T -
storage@20120924 134G - 4.15T -
storage@20121028 36.2G - 4.26T -
storage@20121201 33.2M - 4.55T -
storage@today 0 - 4.58T -
Accessing Snapshot Contents
Snapshot contents can be accessed through a special .zfs/snapshot/ directory. Each snapshot will contain a read-only copy of the data that existed when the snapshot was taken.
# ls /storage/.zfs/snapshot/
20120820/ 20120924/ 20121028/ 20121201/ today/
To roll back to a specific snapshot, run zfs rollback storage@yesterday. This will restore your volume to the snapshot state.
# zfs rollback storage@yesterday
To delete a specific snapshot, run zfs destroy storage@today. Note that this will not work if other volumes depend on it. eg: If you cloned it as another volume.
# zfs destroy storage@today
Data Integrity
One of the strengths of ZFS its resiliency thanks to its transactional file system. The only way data stored on a ZFS volume to be in an inconsistent state is through hardware failure or some sort of fault with the ZFS implementation. Similar to a fsck on ext file systems, a ZFS scrub provides a way to perform filesystem checking. To initiate a scrub, run:
# zpool scrub storage
Once the scrub process is underway, you can view its status by running:
# zpool status storage
pool: storage
state: ONLINE
scan: scrub in progress since Mon Dec 3 23:54:53 2012
18.3G scanned out of 6.05T at 211M/s, 8h20m to go
0 repaired, 0.30% done
config:
NAME STATE READ WRITE CKSUM
storage ONLINE 0 0 0
raidz1-0 ONLINE 0 0 0
gpt/disk00 ONLINE 0 0 0
gpt/disk01 ONLINE 0 0 0
gpt/disk02 ONLINE 0 0 0
gpt/disk03 ONLINE 0 0 0
gpt/disk04 ONLINE 0 0 0
errors: No known data errors
To stop a scrub process, run:
# zpool scrub -s storage
Drive Failure
On a drive failure, you will see something similar to:
# zpool status
pool: storage
state: DEGRADED
status: One or more devices could not be opened. Sufficient replicas exist for
the pool to continue functioning in a degraded state.
action: Attach the missing device and online it using 'zpool online'.
see: http://www.sun.com/msg/ZFS-8000-2Q
scan: scrub canceled on Tue Jun 4 21:45:26 2013
config:
NAME STATE READ WRITE CKSUM
storage DEGRADED 0 0 0
raidz1-0 DEGRADED 0 0 0
da1p1 ONLINE 0 0 0
da2p1 ONLINE 0 0 0
da3p1 ONLINE 0 0 0
6594991764398825070 UNAVAIL 39 148 0 was /dev/da4p1
da5p1 ONLINE 0 0 0
errors: No known data errors
In the above example, the disk /dev/da4p1 was disconnected after it started having read/write errors (39 read / 148 write errors). The solution to the above is to replace the dead drive (obviously) and resilver the data.
Once the new disk is added to the system, reinitialize the drive as you did with the other drives on the system. In my case, I reinitialized the disks using the geometry settings described above.
# gpart create -s GPT da0
da4 created
# gpart add -b 2048 -s 3906617520 -t freebsd-zfs -l disk03 da4
da4p1 added
To replace the offline disk, use zpool replace. Since the disk I replaced has the same device name (da4), I only need to tell ZFS to replace da4p1 with the same device. However, if your new drive has a new name (eg: da6), you will need to run zpool replace storage /dev/da4p1 /dev/da6p1 (or something similar).
[root@bsd /dev]# zpool replace storage /dev/da4p1
[root@bsd /dev]# zpool status
pool: storage
state: DEGRADED
status: One or more devices is currently being resilvered. The pool will
continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
scan: resilver in progress since Tue Jun 4 22:41:10 2013
40.5M scanned out of 8.12T at 5.78M/s, 409h1m to go
7.75M resilvered, 0.00% done
config:
NAME STATE READ WRITE CKSUM
storage DEGRADED 0 0 0
raidz1-0 DEGRADED 0 0 0
da1p1 ONLINE 0 0 0
da2p1 ONLINE 0 0 0
da3p1 ONLINE 0 0 0
replacing-3 UNAVAIL 0 0 0
6594991764398825070 UNAVAIL 0 0 0 was /dev/da4p1/old
da4p1 ONLINE 0 0 0 (resilvering)
da5p1 ONLINE 0 0 0
Tuning and Monitoring
TODO: ARC size information.
TODO: ashift information
the zdb utility.
To share a ZFS pool via NFS, ensure that you have the following in your /etc/rc.conf file.
mountd_enable="YES"
rpcbind_enable="YES"
nfs_server_enable="YES"
mountd_flags="-r -p 735"
zfs_enable="YES"
mountd is required for exports to be loaded from /etc/exports. You must also either reload or restart mountd everytime you make a change to the exports file in order to have it reread.
Set the sharenfs property using the zfs utility. To share it with anyone, set it to 'on', or to some other value. Eg:
# zfs sharenfs=off storage
# zfs sharenfs="-network 172.17.12.0/24" storage/linbuild
zfs get sharenfs
NAME PROPERTY VALUE SOURCE
data sharenfs off local
data/linbuild sharenfs -network 172.17.12.0/24 local
data/linbuild/centos sharenfs -network 172.17.12.0/24 inherited from data/linbuild
data/linbuild/fedora sharenfs -network 172.17.12.0/24 inherited from data/linbuild
data/linbuild/scientific sharenfs -network 172.17.12.0/24 inherited from data/linbuild
Add the paths you wish to export to /etc/exports and reload mountd.
Questionable stuff here:
By setting the sharenfs property, your system will automatically create an export for the zfs pool using mountd. By default, without specifying any options, the NFS share will only be accessible locally.
# showmount -e
Exports list on localhost:
/data Everyone
The above will create a share /storage that is available to all. If you want to restrict the share to a specific network, you can do something like:
# zfs sharenfs="-network 10.1.1.0/24" storage
# showmount -e
Exports list on localhost:
/storage 10.1.1.0
By the way, the exports are stored in /etc/zfs/exports and not in the usual /etc/exports. The ZFS and mountd service must be started for it to work. Therefore, you'll also need to append to /etc/rc.conf the following line:
mountd_enable="YES"
Other Notes
For your ZFS pool to be mounted on startup, you will need the zfs service enabled by having the following line in /etc/rc.conf:
zfs_enable="YES"
Transferring ZFS Data Sets
You can transfer a ZFS snapshot using the send and receive commands.
Using SSH:
# zfs send zones/UUID@snapshot
Using Netcat:
## On the source machine
# zfs send data/linbuild@B