ZFS: Difference between revisions

From Leo's Notes
This page was last edited on 29 May 2017, at 22:29.
Line 580: Line 580:
17:41:59  151    43    28    4    9    38  39    43  28  3.9G  3.9G
17:41:59  151    43    28    4    9    38  39    43  28  3.9G  3.9G
}}
}}
If you are running out of ARC, you might get arc_prune taking up all the CPU. See:
modprobe zfs zfs_arc_meta_strategy=0





Revision as of 22:29, 29 May 2017

ZFS is a file system and logical volume manager.

Unless otherwise noted, this article will contain instructions which will apply to Linux, FreeBSD, and Solaris.

Installation

ZFS is included with Solaris and FreeBSD. Linux support is offered by various open source projects including the ZFS on Linux project.

Linux

Download the source or prebuilt packages from ZFS on Linux's website http://zfsonlinux.org/

Packages from ZFS on Linux will make use of dkms and can be rebuilt when a new kernel is used. (eg: dkms build -m zfs -v 0.6.5 -k 4.2.5-300.fc23.x86_64)

Introduction

ZFS stores data inside datasets which are contained inside storage pools (zpool) which are created on top of block devices (vdev). When creating zpools, the vdevs used can be configured similar to different RAID levels for different redundancies. There are also many other options for ZFS which will be covered later.

Creating vdevs - The Block Storage

To prevent write-amplification, partitions created on disks with Advanced Format (ie. most modern disks) used for the vdevs should be aligned to 4 kilobyte boundaries, equivalent to 2048 sectors for disks with 512 byte sectors.

Optionally, partitions should be created slightly smaller than the total capacity of the drive in case replacement disks have a different size. The examples used below will be 200 MB smaller than the full size. In sectors, this is 200 * 1024*1024 bytes / 512 bytes/sector = 200 * 2048 sectors = 409600 sectors.

To determine the total number of sectors your disks:

  • On Linux, use fdisk -l
  • On FreeBSD, diskinfo -v

An example of a 3TB disk with 5860533168 sectors.

# fdisk -l /dev/sdf
Disk /dev/sdf: 2.7 TiB, 3000592982016 bytes, 5860533168 sectors
Units: sectors of 1 * 512 = 512 bytes
Sector size (logical/physical): 512 bytes / 4096 bytes
I/O size (minimum/optimal): 4096 bytes / 4096 bytes
Disklabel type: gpt
Disk identifier: 357F62D3-D1B9-11E3-BE34-00E04C801A50

Create the data partitions using gpart or fdisk for each of your data disks. Older versions of fdisk do not support drives larger than 2TB.

In the example below, I will be creating a ZFS pool across 3 disks on a FreeBSD machine using ada1, ada2, ada3. Linux machines will do something similar, but the disk device names will be named differently (eg: /dev/sda1, /dev/sdb1, etc.).

# gpart create -s GPT /dev/ada1
# gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk01 /dev/ada1
# gpart create -s GPT /dev/ada2
# gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk02 /dev/ada2
# gpart create -s GPT /dev/ada3
# gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk03 /dev/ada3

Note: The partitions created above starts at sector 2048 and ends 200 MB from the end.

Creating ZFS Pool

With the data partitions created, you should be able to create your new ZFS volume. The most basic way to create a new zpool is to use the zpool create command.

In the example on this page, we will use a zpool named 'storage' across 3 disks using raidz1 (1 redundant disk).

# zpool create -f storage raidz1 gpt/disk01 gpt/disk02 gpt/disk03

Parameters specific to the zpool can also be passed using the -o option. For instance, if using Advanced Format disks with 4096 sector sizes, you might want to set ashift=12 so ZFS does not write blocks straddling sectors like so:

# zpool create -o ashift=12 -f storage raidz1 gpt/disk01 gpt/disk02 gpt/disk03

Once created, you can see the status of the pool using zfs status.

# zpool status
  pool: pool-name
  state: ONLINE
  scan: none requested
 config:
 
         NAME            STATE     READ WRITE CKSUM
         storage         ONLINE       0     0     0
           raidz1-0      ONLINE       0     0     0
             gpt/disk01  ONLINE       0     0     0
             gpt/disk02  ONLINE       0     0     0
             gpt/disk03  ONLINE       0     0     0
 
 errors: No known data errors

Zpool Parameters

Use the zdb command to look up a zpool's parameters. For instance, to see the ashift value of the newly created pool above:

# zdb | grep ashift
            ashift: 12

Creating ZFS Filesystems

With a zpool in place, we can start creating ZFS filesystems, also called datasets, and are analogous to volumes in LVM.

The most basic usage is:

# zfs create storage

Parameters for this specific data set can be specified using the -o option. We'll talk more about options later.

# zfs create -o compress=lz4 storage/logs
Common ZFS Options
Here are some options that are pretty useful:

compress=lz4 - For newer ZFS pools (versions >= 5000), lz4 compression should always be enabled because files that are not compressible will not be compressed which makes it highly efficient. Most CPUs should be able to keep up with IO.

copies=N - This tells ZFS to maintain multiple copies of the file for redundancy. In case of bad sectors, ZFS will have the ability to recover the bad data.


To list the data sets:

# zfs list
NAME                         USED  AVAIL  REFER  MOUNTPOINT
storage                      153K  1.11T   153K  /storage
storage/logs                 153K  1.11T   153K  /storage/logs

Get/Set Volume Parameters

To get a ZFS parameter, use the zfs get command. Eg:

# zfs get storage compressratio
NAME                        PROPERTY       VALUE  SOURCE
storage                     compressratio  1.01x  -

To set a ZFS parameter, use the zfs set command. Setting a parameter such as enabling/disabling compression will not alter the existing data. If you wish to compress a volume, you will need to move data out and then back in for it to take affect.

ZFS Snapshot

A ZFS snapshot is a read-only copy of the file system from a previous state. Because ZFS makes use of Copy-On-Write, snapshots take no additional space.

Listing Snapshots

It can be listed by running zfs list -t snapshot.

# zfs list -t snapshot
NAME               USED  AVAIL  REFER  MOUNTPOINT
storage@20120820  31.4G      -  3.21T  -
storage@20120924   134G      -  4.15T  -
storage@20121028  36.2G      -  4.26T  -
storage@20121201  33.2M      -  4.55T  -

The USED column shows the amount of space used by the snapshot. This amount will go up as files from the snapshot are deleted since the space freed cannot be reclaimed until the snapshot is deleted.

The REFER column shows the actual size of the pool at the snapshot's timepoint.

To quickly get snapshots ordered by when they were made, use the -r option. Passing in the fields you need (name) will speed this operation up.

# zfs list -t snapshot -o name -s name -r data/home
NAME
data/home@zbk-daily-20170502-003001
data/home@zbk-daily-20170503-003001
data/home@zbk-daily-20170504-003001
data/home@zbk-daily-20170505-003001

## Get the most recent snapshot name
# zfs list -t snapshot -o name -s name -r data/home | tail -n 1 | awk -F@ '{print $2}'
zbk-daily-20170505-003001

Creating Snapshots

To create a new snapshot, run zfs snapshot storage@snapshot-name. The snapshot name can be anything, but a datestamp is typically what I use.

No Additional Space is Used
ZFS snapshots make use of Copy-On-Write (COW) and will not use any additional space. However, deleting files that are part of an existing snapshot will not reclaim space. Instead, the storage capacity will appear to go down since the amount of available space for new data remains the same. Eg: A volume at 1TB of 2TB capacity has a 0.5TB file deleted. Because the 0.5TB file is still stored in a snapshot, the total capacity is now 2TB - 0.5TB and the reported disk usage will now be 0.5TB of 1.5TB.


If I were to run zfs snapshot storage@snapshot-name now, listing the snapshots will yield:

# zfs list -t snapshot
NAME               USED  AVAIL  REFER  MOUNTPOINT
storage@20120820  31.4G      -  3.21T  -
storage@20120924   134G      -  4.15T  -
storage@20121028  36.2G      -  4.26T  -
storage@20121201  33.2M      -  4.55T  -
storage@today         0      -  4.58T  -

Accessing Snapshot Contents

Snapshot contents can be accessed through a special .zfs/snapshot/ directory. Each snapshot will contain a read-only copy of the data that existed when the snapshot was taken.

# ls /storage/.zfs/snapshot/
20120820/ 20120924/ 20121028/ 20121201/ today/

To roll back to a specific snapshot, run zfs rollback storage@yesterday. This will restore your volume to the snapshot state.

# zfs rollback storage@yesterday

To delete a specific snapshot, run zfs destroy storage@today. Note that this will not work if other volumes depend on it. eg: If you cloned it as another volume.

# zfs destroy storage@today

Data Integrity

One of the strengths of ZFS its resiliency thanks to its transactional file system. The only way data stored on a ZFS volume to be in an inconsistent state is through hardware failure or some sort of fault with the ZFS implementation. Similar to a fsck on ext file systems, a ZFS scrub provides a way to perform filesystem checking. To initiate a scrub, run:

# zpool scrub storage

Once the scrub process is underway, you can view its status by running:

# zpool status storage
   pool: storage
  state: ONLINE
  scan: scrub in progress since Mon Dec  3 23:54:53 2012
     18.3G scanned out of 6.05T at 211M/s, 8h20m to go
     0 repaired, 0.30% done
 config:
 
         NAME            STATE     READ WRITE CKSUM
         storage         ONLINE       0     0     0
           raidz1-0      ONLINE       0     0     0
             gpt/disk00  ONLINE       0     0     0
             gpt/disk01  ONLINE       0     0     0
             gpt/disk02  ONLINE       0     0     0
             gpt/disk03  ONLINE       0     0     0
             gpt/disk04  ONLINE       0     0     0
 
 errors: No known data errors

To stop a scrub process, run:

# zpool scrub -s storage


Administration

Listing Zpools and Filesystems

List pools using zpool list:

# zpool list
NAME      SIZE  ALLOC   FREE  EXPANDSZ   FRAG    CAP  DEDUP  HEALTH  ALTROOT
data     8.12T  5.42T  2.71T         -    50%    66%  1.00x  ONLINE  -
storage  9.06T  5.93T  3.13T         -    13%    65%  1.00x  ONLINE  -

The FRAG value is the average fragmentation of available space.

To list ZFS filesystems, use zfs list

# zfs list
NAME                         USED  AVAIL  REFER  MOUNTPOINT
data                        3.61T  1.63T   139K  /data
data/backups                 944G  1.63T   901G  /data/backups
data/crypt                   125G  1.63T   125G  /data/crypt
data/dashcam                84.5G  1.63T  84.5G  /data/dashcam
data/home                    169G  1.63T   153G  /data/home
data/images                  999G  1.63T   999G  /data/images
data/public                 1.11T  1.63T  1.03T  /data/public
data/scratch                 167G  1.63T   158G  /data/scratch
data/vm                     71.2G  1.63T  57.8G  /data/vm
storage                     4.74T  2.27T  26.2G  /storage

To list snapshots, pass -t snapshot

# zfs list -t snapshot
NAME                                                   USED  AVAIL  REFER  MOUNTPOINT
data/backups@2015-Dec-15                               658M      -   588G  -
data/backups@2017-Jan-01                                  0      -   927G  -
data/backups@2017-Jan-08                                  0      -   927G  -
data/backups@2017-Jan-15                              74.6K      -   927G  -
data/backups@2017-Jan-22                               416K      -   927G  -
data/backups@2017-Jan-29                               767K      -   927G  -
data/backups@2017-Feb-05                                  0      -   927G  -
data/backups@2017-Feb-12                                  0      -   927G  -
data/backups@2017-Feb-19                               202K      -   927G  -


To get properties of your ZFS filesystems, use the zfs get PROPERTYNAME command:

# zfs get compressratio
NAME  PROPERTY       VALUE  SOURCE
data  compressratio  1.02x  -


Transferring ZFS Data Sets

You can transfer a ZFS snapshot using the send and receive commands.

Using SSH:

# zfs send  zones/UUID@snapshot

Using Netcat:

## On the source machine
# zfs send data/linbuild@B


NFS Share

FreeBSD

To share a ZFS pool via NFS on a FreeBSD system, ensure that you have the following in your /etc/rc.conf file.

mountd_enable="YES"
rpcbind_enable="YES"
nfs_server_enable="YES"
mountd_flags="-r -p 735"
zfs_enable="YES"

mountd is required for exports to be loaded from /etc/exports. You must also either reload or restart mountd everytime you make a change to the exports file in order to have it reread.

Set the sharenfs property using the zfs utility. To share it with anyone, set it to 'on', or to some other value. Eg:

# zfs sharenfs=off storage
# zfs sharenfs="-network 172.17.12.0/24" storage/linbuild
zfs get sharenfs
NAME                      PROPERTY  VALUE                    SOURCE
data                      sharenfs  off                      local
data/linbuild             sharenfs  -network 172.17.12.0/24  local
data/linbuild/centos      sharenfs  -network 172.17.12.0/24  inherited from data/linbuild
data/linbuild/fedora      sharenfs  -network 172.17.12.0/24  inherited from data/linbuild
data/linbuild/scientific  sharenfs  -network 172.17.12.0/24  inherited from data/linbuild

Add the paths you wish to export to /etc/exports and reload mountd.

Questionable stuff here:

By setting the sharenfs property, your system will automatically create an export for the zfs pool using mountd. By default, without specifying any options, the NFS share will only be accessible locally.

# showmount -e
Exports list on localhost:
/data                           Everyone


The above will create a share /storage that is available to all. If you want to restrict the share to a specific network, you can do something like:

# zfs sharenfs="-network 10.1.1.0/24" storage
# showmount -e
Exports list on localhost:
/storage                           10.1.1.0

By the way, the exports are stored in /etc/zfs/exports and not in the usual /etc/exports. The ZFS and mountd service must be started for it to work. Therefore, you'll also need to append to /etc/rc.conf the following line:

mountd_enable="YES"

On Linux

To share ZFS pools on Linux, do the same stuff as you would on a FreeBSD system (see above). However, you may need to update the NFS export file manually by running the systemd zfs-share script in /usr/lib/systemd/system.


Enabling on Startup

FreeBSD

For your ZFS pool to be mounted on startup, you will need the zfs service enabled by having the following line in /etc/rc.conf:

zfs_enable="YES"

Linux

Loading Kernel Modules on Boot

After installing ZFS on Linux, you may need to create a file in either /etc/modules-load.d or /etc/modprobe.d so that the ZFS module gets loaded.

To only load the module, create a file at /etc/modules-load.d/zfs.conf containing:

zfs

To load the module with options, create a file at /etc/modprobe.d/zfs.conf containing (for example):

options zfs zfs_arc_max=4294967296
Enabling Systemd Services

On Systemd based systems, you will need to enable the following services in order to have the zpool loaded and mounted on start up. If you have sharing enabled, you will want to also enable the zfs-share service too.

  • zfs-import-cache
  • zfs-mount
  • zfs-share

You may want to look at what these service files do by looking at them at /usr/lib/systemd/system/.

If you don't enable these services, your zpools will be loaded but not mounted.


Handling Drive Failure

When a drive fails, you will see something similar to:

# zpool status
  pool: data
 state: DEGRADED
status: One or more devices could not be used because the label is missing or
        invalid.  Sufficient replicas exist for the pool to continue
        functioning in a degraded state.
action: Replace the device using 'zpool replace'.
   see: http://zfsonlinux.org/msg/ZFS-8000-4J
  scan: scrub repaired 0 in 2h31m with 0 errors on Fri Jun  3 19:20:15 2016
config:

        NAME        STATE     READ WRITE CKSUM
        data        DEGRADED     0     0     0
          raidz2-0  DEGRADED     0     0     0
            sdb     ONLINE       0     0     0
            sdc     ONLINE       0     0     0
            sdd     ONLINE       0     0     0
            sde     ONLINE       0     0     0
            sdf     ONLINE       0     0     0
            sdg     UNAVAIL      3   222     0  corrupted data

errors: No known data errors

The device sdg failed and was removed from the pool.

Dell Server Info
Since this happened on a server, I removed the failed disk from the machine and replaced it with another one.

Because this server uses a Dell RAID controller, and these Dell RAID controllers can't do a pass-through, each disk that is part of the ZFS array is its own Raid0 vdisk. When inserting a new drive, the vdisk information needs to be re-created via OpenManage before the disk comes online.

Follow the steps on the Dell OpenManage after inserting the replacement drive if this applies to you.

Once the vdisk is recreated, it should show up as the old device name again.


Once the replacement disk is installed on the system, reinitialize the drive with the GPT label.

On a FreeBSD system: I reinitialized the disks using the geometry settings described above.

# gpart create -s GPT da0
da4 created
# gpart add -b 2048 -s 3906617520 -t freebsd-zfs -l disk03 da4
da4p1 added

On a ZFS on Linux system:

# parted /dev/sdg
GNU Parted 2.1
Using /dev/sdg
Welcome to GNU Parted! Type 'help' to view a list of commands.
(parted) mklabel GPT
Warning: The existing disk label on /dev/sdg will be destroyed and all data on this disk will be lost. Do
you want to continue?
Yes/No? y
(parted) quit

To replace the offline disk, use zpool replace pool old-device new-device.

# zpool replace data sdg /dev/sdg
# zpool status
  pool: data
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
        continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
  scan: resilver in progress since Mon Apr 10 16:36:35 2017
    35.5M scanned out of 3.03T at 3.55M/s, 249h19m to go
    5.45M resilvered, 0.00% done
config:

        NAME             STATE     READ WRITE CKSUM
        data             DEGRADED     0     0     0
          raidz2-0       DEGRADED     0     0     0
            sdb          ONLINE       0     0     0
            sdc          ONLINE       0     0     0
            sdd          ONLINE       0     0     0
            sde          ONLINE       0     0     0
            sdf          ONLINE       0     0     0
            replacing-5  UNAVAIL      0     0     0
              old        UNAVAIL      3   222     0  corrupted data
              sdg        ONLINE       0     0     0  (resilvering)

errors: No known data errors

Tuning and Monitoring

ZFS ARC

The ZFS Adaptive Replacement Cache (ARC) is an in-memory cached managed by ZFS to help improve speed duplicate reads.

The ZFS ARC size from the ZFS on Linux implementation defaults to half the host's available memory and may decrease when available memory gets too low. Depending on the system setup, you may want to change the maximum amount of memory allocated to ARC.

To change the maximum ARC size, edit the ZFS zfs_arc_max kernel module parameter:

# cat /etc/modprobe.d/zfs.conf
options zfs zfs_arc_max=4294967296

You may check the current ARC usage by checking cat /proc/spl/kstat/zfs/arcstats or use the arcstat.py script that's part of the zfs package.

# cat /proc/spl/kstat/zfs/arcstats
p                               4    1851836123
c                               4    4105979840
c_min                           4    33554432
c_max                           4    4294967296
size                            4    4105591928
hdr_size                        4    55529696
data_size                       4    3027917312
metadata_size                   4    747323904
other_size                      4    274821016
anon_size                       4    4979712
anon_evictable_data             4    0
anon_evictable_metadata         4    0
mru_size                        4    848103424
mru_evictable_data              4    605918208
mru_evictable_metadata          4    97727488
mru_ghost_size                  4    3224668160
mru_ghost_evictable_data        4    2267575296
mru_ghost_evictable_metadata    4    957092864
mfu_size                        4    2922158080
mfu_evictable_data              4    2421089280
mfu_evictable_metadata          4    495401984
mfu_ghost_size                  4    835275264
mfu_ghost_evictable_data        4    787349504
mfu_ghost_evictable_metadata    4    47925760
l2_hits                         4    0
l2_misses                       4    0
l2_feeds                        4    0
l2_rw_clash                     4    0
l2_read_bytes                   4    0
l2_write_bytes                  4    0
l2_writes_sent                  4    0
l2_writes_done                  4    0
l2_writes_error                 4    0
l2_writes_lock_retry            4    0
l2_evict_lock_retry             4    0
l2_evict_reading                4    0
l2_evict_l1cached               4    0
l2_free_on_write                4    0
l2_cdata_free_on_write          4    0
l2_abort_lowmem                 4    0
l2_cksum_bad                    4    0
l2_io_error                     4    0
l2_size                         4    0
l2_asize                        4    0
l2_hdr_size                     4    0
l2_compress_successes           4    0
l2_compress_zeros               4    0
l2_compress_failures            4    0
memory_throttle_count           4    0
duplicate_buffers               4    0
duplicate_buffers_size          4    0
duplicate_reads                 4    0
memory_direct_count             4    46505
memory_indirect_count           4    30062
arc_no_grow                     4    0
arc_tempreserve                 4    0
arc_loaned_bytes                4    0
arc_prune                       4    0
arc_meta_used                   4    1077674616
arc_meta_limit                  4    3113852928
arc_meta_max                    4    1077674616
arc_meta_min                    4    16777216
arc_need_free                   4    0
arc_sys_free                    4    129740800

# arcstat.py 10 10
    time  read  miss  miss%  dmis  dm%  pmis  pm%  mmis  mm%  arcsz     c
17:41:09     0     0      0     0    0     0    0     0    0   3.9G  3.9G
17:41:19   211    68     32     6    7    62   50    68   32   3.9G  3.9G
17:41:29   146    32     21     6   10    25   30    32   21   3.9G  3.9G
17:41:39   170    50     29     6    9    44   42    50   31   3.9G  3.9G
17:41:49   150    33     22     6    9    27   31    33   22   3.9G  3.9G
17:41:59   151    43     28     4    9    38   39    43   28   3.9G  3.9G


If you are running out of ARC, you might get arc_prune taking up all the CPU. See:

modprobe zfs zfs_arc_meta_strategy=0


See Also:

ZFS L2ARC

The ZFS L2ARC provides caching by extending the main memory cache with faster-than-storage-pool disks such as SSDs. Cached data that normally wouldn't fit in the ARC would be cached in the L2ARC.

ZFS ZIL

The ZFS Intent Log (ZIL) is how ZFS keeps track of synchronous write operations so that they can be completed or rolled back after a crash or failure. The ZIL is not used or asynchronous writes - those are still done through system caches.

Because ZIL is stored in the data pool, writing to a ZFS pool involves duplicate writes to the pool: Once to the ZIL and again for the actual data. This is detrimental to performance since one write operation now requires two or more writes. To improve performance, the ZIL could be moved to a separate device called a Separate Intent Log (SLOG), typically a SSD. With a SLOG, a write operation will write once to the SLOG and once for the actual data.

See Also:

Alignment Shift

Make sure that your vdevs have the proper alignment shift.

Device Alignment
Hard Drives with 512B sectors 9
Flash Media / HDD with 4K sectors 12
Flash Media / HDD with 8K sectors 13
Amazon EC2 12

See Also:

See Also