ZFS: Difference between revisions

From Leo's Notes
This page was last edited on 4 November 2015, at 22:46.
snapshot cleanup
Line 199: Line 199:
== NFS Share ==
== NFS Share ==


On FreeBSD, to share a ZFS pool over NFS, run the following:
On FreeBSD, to enable NFS exports, run the following:
# showmount -e
{{highlight|lang=bash|code=
Exports list on localhost:
# showmount -e
# service mountd start
Exports list on localhost:
# zfs sharenfs=on storage
# touch /etc/exports
  # showmount -e
# service mountd start
Exports list on localhost:
# zfs sharenfs=on storage
/storage                           Everyone
}}
 
{{code|mountd}} requires {{code|/etc/exports}} to exist. By setting the {{code|sharenfs}} property, your system will automatically create an export for the zfs pool using {{code|mountd}}. By default, without specifying any options, the NFS share will only be accessible locally.
 
You can look at the value of any zfs property using {{code|zfs get property}}. Eg:
{{highlight|lang=bash|code=
# zfs get sharenfs
NAME  PROPERTY  VALUE    SOURCE
data  sharenfs on        local
}}
 
{{highlight|lang=bash|code=
# showmount -e
Exports list on localhost:
/data                           Everyone
}}
 


The above will create a share <code>/storage</code> that is available to all. If you want to restrict the share to a specific network, you can do something like:
The above will create a share <code>/storage</code> that is available to all. If you want to restrict the share to a specific network, you can do something like:

Revision as of 22:46, 4 November 2015

This will be a quick overview on how to manage ZFS snapshots and volumes on a FreeBSD system, but it should apply to Solaris as well.

Creating a ZFS Volume

Creating Data Partitions

For alignment and safety space, when creating a new volume, we want to have it start after 1 MB from the start and end 200 MB from the end. Therefore, the data partition should be 201 MB less than the total size of the disk.

To calculate 201 MB in number of sectors, taking the following into account:

  • 512 bytes/sector
  • 1 MB = 1024 * 1024 bytes

1 MB in sectors is calculated by: 1024*1024 bytes / 512 bytes/sector = 2048 sectors. Therefore, 201 MB in sectors will be 201 * 2048, or 411648.

Use diskinfo -v to determine the total number of sectors your disks contains and subtract 411648.

Create the data partitions using gpart for each of your data disks. In the example below, I will be creating a ZFS pool across 3 disks (at ada1, ada2, ada3).

gpart create -s GPT /dev/ada1
gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk01 /dev/ada1
gpart create -s GPT /dev/ada2
gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk02 /dev/ada2
gpart create -s GPT /dev/ada3
gpart add -b 2048 -s 1953113520 -t freebsd-zfs -l disk03 /dev/ada3

The partitions created above starts 2048 sectors from the beginning and ends 200 MB from the end.

Creating ZFS Pool

With the data partitions created, you should be able to create your new ZFS volume. The example below will create a pool named pool-name using raidz1 (1 redundant disk).

zpool create -f pool-name raidz1 gpt/disk01 gpt/disk02 gpt/disk03

Once created, you can poll for the status using zfs status.

zpool status
  pool: pool-name
  state: ONLINE
  scan: none requested
 config:
 
         NAME            STATE     READ WRITE CKSUM
         storage         ONLINE       0     0     0
           raidz1-0      ONLINE       0     0     0
             gpt/disk01  ONLINE       0     0     0
             gpt/disk02  ONLINE       0     0     0
             gpt/disk03  ONLINE       0     0     0
 
 errors: No known data errors


ZFS Snapshot

Listing Snapshots

A ZFS snapshot is a read-only copy of the file system from a past state. It can be listed by running zfs list -t snapshot.

# zfs list -t snapshot
NAME               USED  AVAIL  REFER  MOUNTPOINT
storage@20120820  31.4G      -  3.21T  -
storage@20120924   134G      -  4.15T  -
storage@20121028  36.2G      -  4.26T  -
storage@20121201  33.2M      -  4.55T  -

The USED column shows the amount of space used by the snapshot. This amount will go up as files from the snapshot are deleted since the space freed cannot be reclaimed until the snapshot is deleted.

The REFER column shows the actual size of the pool at the snapshot's timepoint.

Creating Snapshots

To create a new snapshot, run zfs snapshot storage@snapshot-name. The snapshot name can be anything, but a datestamp is typically what I use.

No Additional Space is Used
ZFS snapshots make use of Copy-On-Write (COW) and will not use any additional space. However, deleting files that are part of an existing snapshot will not reclaim space. Instead, the storage capacity will appear to go down.


If I were to run zfs snapshot storage@snapshot-name now, listing the snapshots will yield:

# zfs list -t snapshot
NAME               USED  AVAIL  REFER  MOUNTPOINT
storage@20120820  31.4G      -  3.21T  -
storage@20120924   134G      -  4.15T  -
storage@20121028  36.2G      -  4.26T  -
storage@20121201  33.2M      -  4.55T  -
storage@today         0      -  4.58T  -

Accessing Snapshot Contents

Snapshot contents can be accessed through a special .zfs/snapshot/ directory. Each snapshot will contain a read-only copy of the data that existed when the snapshot was taken.

# ls /storage/.zfs/snapshot/
20120820/ 20120924/ 20121028/ 20121201/ today/

To roll back to a specific snapshot, run zfs rollback storage@yesterday. This will restore your volume to the snapshot state.

# zfs rollback storage@yesterday

To delete a specific snapshot, run zfs destroy storage@today. Note that this will not work if other volumes depend on it. eg: If you cloned it as another volume.

# zfs destroy storage@today

Data Integrity

One of the strengths of ZFS its resiliency thanks to its transactional file system. The only way data stored on a ZFS volume to be in an inconsistent state is through hardware failure or some sort of fault with the ZFS implementation. Similar to a fsck on ext file systems, a ZFS scrub provides a way to perform filesystem checking. To initiate a scrub, run:

# zpool scrub storage

Once the scrub process is underway, you can view its status by running:

# zpool status storage
  pool: storage
 state: ONLINE
 scan: scrub in progress since Mon Dec  3 23:54:53 2012
    18.3G scanned out of 6.05T at 211M/s, 8h20m to go
    0 repaired, 0.30% done
config:

        NAME            STATE     READ WRITE CKSUM
        storage         ONLINE       0     0     0
          raidz1-0      ONLINE       0     0     0
            gpt/disk00  ONLINE       0     0     0
            gpt/disk01  ONLINE       0     0     0
            gpt/disk02  ONLINE       0     0     0
            gpt/disk03  ONLINE       0     0     0
            gpt/disk04  ONLINE       0     0     0

errors: No known data errors

To stop a scrub process, run:

# zpool scrub -s storage


Drive Failure

On a drive failure, you will see something similar to:

zpool status
  pool: storage
 state: DEGRADED
status: One or more devices could not be opened.  Sufficient replicas exist for
        the pool to continue functioning in a degraded state.
action: Attach the missing device and online it using 'zpool online'.
   see: http://www.sun.com/msg/ZFS-8000-2Q
 scan: scrub canceled on Tue Jun  4 21:45:26 2013
config:

        NAME                     STATE     READ WRITE CKSUM
        storage                  DEGRADED     0     0     0
          raidz1-0               DEGRADED     0     0     0
            da1p1                ONLINE       0     0     0
            da2p1                ONLINE       0     0     0
            da3p1                ONLINE       0     0     0
            6594991764398825070  UNAVAIL     39   148     0  was /dev/da4p1
            da5p1                ONLINE       0     0     0

errors: No known data errors

In the above example, the disk /dev/da4p1 was disconnected after it started having read/write errors (39 read / 148 write errors). The solution to the above is to replace the dead drive (obviously) and readd and resliver the data.

Once the new disk is added to the system, reinitialize the drive as you did with the other drives on the system. In my case, I reinitialized the disks using the geometry settings described above.

# gpart create -s GPT da0
da4 created
# gpart add -b 2048 -s 3906617520 -t freebsd-zfs -l disk03 da4
da4p1 added

To replace the offline disk, use zpool replace. Since the disk I replaced has the same device name (da4), I only need to tell ZFS to replace da4p1 with the same device. However, if your new drive has a new name (eg: da6), you will need to run zpool replace storage /dev/da4p1 /dev/da6p1 (or something similar).

[root@bsd /dev]# zpool replace storage /dev/da4p1
[root@bsd /dev]# zpool status
  pool: storage
 state: DEGRADED
status: One or more devices is currently being resilvered.  The pool will
        continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
 scan: resilver in progress since Tue Jun  4 22:41:10 2013
    40.5M scanned out of 8.12T at 5.78M/s, 409h1m to go
    7.75M resilvered, 0.00% done
config:

        NAME                       STATE     READ WRITE CKSUM
        storage                    DEGRADED     0     0     0
          raidz1-0                 DEGRADED     0     0     0
            da1p1                  ONLINE       0     0     0
            da2p1                  ONLINE       0     0     0
            da3p1                  ONLINE       0     0     0
            replacing-3            UNAVAIL      0     0     0
              6594991764398825070  UNAVAIL      0     0     0  was /dev/da4p1/old
              da4p1                ONLINE       0     0     0  (resilvering)
            da5p1                  ONLINE       0     0     0

NFS Share

On FreeBSD, to enable NFS exports, run the following:

# showmount -e
Exports list on localhost:
# touch /etc/exports
# service mountd start
# zfs sharenfs=on storage

mountd requires /etc/exports to exist. By setting the sharenfs property, your system will automatically create an export for the zfs pool using mountd. By default, without specifying any options, the NFS share will only be accessible locally.

You can look at the value of any zfs property using zfs get property. Eg:

# zfs get sharenfs
NAME  PROPERTY  VALUE     SOURCE
data  sharenfs  on        local
# showmount -e
Exports list on localhost:
/data                           Everyone


The above will create a share /storage that is available to all. If you want to restrict the share to a specific network, you can do something like:

# zfs sharenfs="-network 10.1.1.0/24" storage
# showmount -e
Exports list on localhost:
/storage                           10.1.1.0

By the way, the exports are stored in /etc/zfs/exports and not in the usual /etc/exports. The ZFS and mountd service must be started for it to work. Therefore, you'll also need to append to /etc/rc.conf the following line:

mountd_enable="YES"

Other Notes

For your ZFS pool to be mounted on startup, you will need the zfs service enabled by having the following line in /etc/rc.conf:

 zfs_enable="YES"