Awk

From Leo's Notes
Revision as of 23:33, 23 October 2020 by Leo (talk | contribs)
This page was last edited on 23 October 2020, at 23:33.

Random Notes

Insert a single quote within single quotes

Awk script is almost always passed to awk using single quotes since it allows us to use $1 ... $NF variables without needing to escape the $. However, this makes inserting a single quote tricky. To add a single quote, you will need to use '"'"' which inserts a single ' by switching from single to double quotes. This is useful if you need to insert a single quote in order to trigger a subshell call, for instance.

Variables

Awk has built-in variables:

  • NR - current line number
  • NF - number of fields in current line
  • OFS - output field separator
  • FS - field separator, specified with -F
  • RS - record separator

You can also pass your own with the -v var=value parameter. Eg:

$ echo | awk \
   -v hostname=`hostname` \
   -v timestamp=`date +%s` '
{
   printf("Hostname is %s at time %s", hostname, timestamp")
}
'

Tasks

Line Matching

With Awk, it's easy to do something for each matching line using the regex matching operator.

Print every line before a line match

Simple awk code:

# cat list.txt | awk '/PATTERN/ { exit } { print $0 }'

Basically, the code does: match against the given pattern. If it matches, awk exits. Otherwise, print the line and continue.

Print the line number on matching lines

To print the number of lines up to a matching line, we do something similar to the previous example but now we keep an accumulator (n):

awk 'BEGIN { n = 0 } /PATTERN/ { print n; exit } { n++ }'

Eg: Suppose I have a file with the contents:

leo
spoon
cake
fork

To find the number of lines up until 'cake', do:

# cat list.txt | awk 'BEGIN { n = 0 } /cake/ { print n; exit } { n++ }'

Retrieve a section of text with line matching

To grab a specific section from a .spec file (where sections begin with %):

if [ $# -ne 2 ] ; then
	echo "Usage: $0 file.spec section"
	exit
fi

Section="$2"

cat $1 | awk -v Section="$Section" '
BEGIN {
	InSection=false
} 
/^%.*/ { 
	if ($0 ~ Section) {
		InSection=1
	} else {
		InSection=0
	}

	next
}
{
	if (InSection) {
		print $0
	}
}

String Matching

Use the match(string, regex, output_array) command to parse out specific values with regex.

# Given a string Job <87010>, User <asdf>, Project <default>
# we can parse out the Job ID with:

match($0, /^Job <([^>]+)>.*/, arr)
print "Job ID: " arr[1]

Converting Values

Convert Number as Bytes to Human Readable Value

I wanted to sort size in reverse order, but in order to do that properly, the value from du needs to be in kilobytes. I also didn't want to run this through du again just to get the human readable value. So, I did this:

$ du -s ./*/ ./*/*/ ./*/*/*/ \ 
   | sort -rn \ 
   | awk 'BEGIN { \ 
      split("KB MB GB TB PB", type) \ 
   } \ 
   { \ 
      y = 0; \                                                                                                  
      x = $1; \ 
      for (i = 4; y &lt; 1 ; i--) \ 
         y = x / (2 ** (10 * i)); \ 
      print y type[i+2]" "$2 \ 
   }'

If you want to count using bytes instead of kilobytes, just add K to the split function and replace i=4 with i=5.

Convert ps Elapsed Time to Seconds

To convert the elapsed time (in POSIX locale formatted as [[dd-]hh:]mm:ss) to seconds using awk:

Elapsed=`ps -p $Pid -o etime=`
Elapsed=`echo $Elapsed | tr - : | tr : ' '` 

Seconds=`echo $Elapsed | awk ' 
	NF == 2 { print ($1 * 60) + $2 } 
	NF == 3 { print ((($1 * 60) + $2) * 60) + $3 } 
	NF == 4 { print ((((($1 * 24) + $2) * 60) + $3) * 60) + $4 } 
	{}'` 
	
echo $Seconds

For example, a process running for 66-00:12:58 has been running for 5703178 seconds. The Awk command will match how many columns were found and then do the proper calculation.

String Manipulation

String Upper/Lower Casing

There is a tolower() function that can lowercase an entire string.

You could use this to mass-rename a bunch of files to lower case for instance.

## Lower case every file in the current directory
$ for i in `ls` ; do mv $i `echo $i | awk '{print tolower($1)}'`; done

## Alternatively, use `tr 'A-Z' 'a-z'` to do the lowercasing.
$ for i in `ls` ; do mv $i `echo $i | tr 'A-Z 'a-z'`; done

Retrieving all columns after column N

You can retrieve a specific column using $N, where N is the column number. If you wish to get all values after a specific column, you can use this function which concatenates all strings after the specified column together and returns it:

function after(x) {
        out=""
        for (i=x; i<=NF; i++) out=out" "$i
        return out
}

printf("After column 11: %s\n", after(11))

Executing Commands

If your awk script needs to call an external process, pass the command to getline followed by the variable name used to store stdout output.

Eg. To get the path of a executable from a given PID, use:

"readlink /proc/" $2 "/exe" | getline proc
printf("%s\n", proc)

If you intend to run many commands, you should close the pipe or else you will get a fatal: cannot open pipe 'xyz' (Too many open files) error. Do so by using the close function:

cmd="date -d\""$1" "$2"\" \"+%s\""; cmd | getline timestamp; close(cmd)