Logstash: Difference between revisions

From Leo's Notes
This page was last edited on 1 August 2014, at 17:14.
No edit summary
Line 27: Line 27:
For example, my current configuration is:
For example, my current configuration is:
<syntaxhighlight lang="text" line start="1" enclose="div">
<syntaxhighlight lang="text" line start="1" enclose="div">
input {
input {
# Import syslog messages
# Import syslog messages
tcp {
tcp {
Line 40: Line 42:
}
}
}
}


filter {
filter {
if [type] == "syslog" {
# grok {
# match => { "message" => "%{SYSLOGTIMESTAMP:syslog_timestamp} %{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }
# }
# Does the syslog parsing.
syslog_pri { }
if !("_grokparsefailure" in [tags]) {
mutate {
replace => [ "@source", "%{logsource}" ]
replace => [ "@message", "%{message}" ]
replace => [ "@program", "%{program}" ]
replace => [ "@type", "syslog" ]
}
}
# Date is parsed and placed into @timstamp.
date {
match => [ "syslog_timestamp", "MMM  d HH:mm:ss", "MMM dd HH:mm:ss", "ISO8601" ]
}
# Clean up the extra syslog_ fields generated above from grok.
mutate {
remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type", "message", "logsource", "program"]
}
}
# For imported syslog messages...
# For imported syslog messages...
if [type] == "syslog_import" {
if [type] == "syslog_import" {
if [message] =~ /last message repeated.*/ {
drop {
}
}
if [message] == "" {
drop {
}
}
# Parse with grok
# Parse with grok
grok {
grok {
Line 64: Line 109:
if !("_grokparsefailure" in [tags]) {
if !("_grokparsefailure" in [tags]) {
mutate {
mutate {
replace => [ "@source_host", "%{syslog_hostname}" ]
replace => [ "@source", "%{syslog_hostname}" ]
replace => [ "@message", "%{syslog_message}" ]
replace => [ "@message", "%{syslog_message}" ]
replace => [ "@program", "%{syslog_program}" ]
replace => [ "@program", "%{syslog_program}" ]
replace => [ "@type", "syslog imported" ]
}
}
}
}
Line 78: Line 124:
# Clean up the extra syslog_ fields generated above from grok.
# Clean up the extra syslog_ fields generated above from grok.
mutate {
mutate {
remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type" ]
remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type", "message", "host" ]
}
}
}
}
}
}




output {
output {
# Print each event to stdout.
# stdout {
#  stdout {
# codec => json
# Enabling 'debug' on the stdout output will make logstash pretty-print the
# }
# entire event as something similar to a JSON representation.
 
#    debug => true
 
# }
# You can have multiple outputs. All events generally to all outputs.
# Output events to elasticsearch
elasticsearch {
elasticsearch {
# Setting 'embedded' will run  a real elasticsearch server inside logstash.
# Define our own... hosted on my computer
# This option below saves you from having to run a separate process just
# bind_host => "leo-linux"
# for ElasticSearch, so you can get started quicker!
# bind_port => 9200
embedded => true
host => "127.0.0.1"
port => 9300
 
cluster => "logstash_es"
node_name => "logstash_0"
 
# Index defaults to 'logstash-%{+YYYY.MM.dd}'
# The templates being used can be defined using:
template => "/etc/logstash/template/logstash.json"
}
}
}
}


</syntaxhighlight>


I do some processing on the messages coming in from port 4401, tagged as <code>syslog_import</code>. In the filter step, any messages tagged as <code>syslog_import</code> will be processed by grok, the parser. Let's take a closer look at the grok configuration:
I do some processing on the messages coming in from port 4401, tagged as <code>syslog_import</code>. In the filter step, any messages tagged as <code>syslog_import</code> will be processed by grok, the parser. Let's take a closer look at the grok configuration:

Revision as of 17:14, 1 August 2014

Logstash is the open source version of splunk, using ElasticSearch as its search engine.

Installation

For detailed information, consult logstash's tutorial (http://logstash.net/docs/). Prior to logstash 1.4.0, the logstash package comes as a monolithic .jar file. To get started, install java and run the jar file.

Running Logstash

I made a wrapper script called run.sh which launches the jar file with my configuration.

#!/bin/bash
java -jar logstash-1.3.3-flatjar.jar agent -f configuration.conf -- web

The core of logstash is the agent. The web site (formerly known as a separate package called kibana) is built in and can be started by appending -- web to the command line.

Configuration

The configuration you provide logstash defines how logstash deals with incoming messages. There are three main parts to the configuration:

  1. Input
  2. Filter
  3. Output

Input defines what ports logstash listens on and how to tag the incoming messages. Filter defines what logstash needs to do on the incoming messages, based on the tags defined from the input step. Output defines how these messages are stored.

For example, my current configuration is:

input {

	# Import syslog messages
	tcp {
		type => syslog_import
		port => 4401
	}

	# Accept syslog messages from hosts
	syslog {
		type => syslog
		port => 5544
	}
}



filter {

	if [type] == "syslog" {
		# grok {
		# 	match => { "message" => "%{SYSLOGTIMESTAMP:syslog_timestamp} %{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }
		# }

		# Does the syslog parsing.
		syslog_pri { }

		if !("_grokparsefailure" in [tags]) {
			mutate {
				replace => [ "@source", "%{logsource}" ]
				replace => [ "@message", "%{message}" ]
				replace => [ "@program", "%{program}" ]
				replace => [ "@type", "syslog" ]
			}
		}

		# Date is parsed and placed into @timstamp.
		date {
			match => [ "syslog_timestamp", "MMM  d HH:mm:ss", "MMM dd HH:mm:ss", "ISO8601" ]
		}

		# Clean up the extra syslog_ fields generated above from grok.
		mutate {
			remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type", "message", "logsource", "program"]
		}
	}

	# For imported syslog messages...
	if [type] == "syslog_import" {
		
		if [message] =~ /last message repeated.*/ {
			drop {
			}
		}
		
		if [message] == "" {
			drop {
			}
		}


		# Parse with grok
		grok {
			# Use the custom SYSLOGYEARTIMESTAMP pattern from the patterns
			# directory. We need this to define year.
			patterns_dir => "./patterns"

			# The pattern to match.
			# This is the standard syslog pattern.
			match => { "message" => "%{SYSLOGYEARTIMESTAMP:syslog_timestamp} (%{USER:syslog_user}\@)?%{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }

			# Add a few intermediate fields
			add_field => [ "received_at", "%{@timestamp}" ]
			add_field => [ "received_from", "%{host}" ]
		}
		
		# When the above grok parsing fails, a '_grokparsefailure' tag gets
		# added to the message. In that case, we attempt to update some fields.
		# Why? Beats me.
		if !("_grokparsefailure" in [tags]) {
			mutate {
				replace => [ "@source", "%{syslog_hostname}" ]
				replace => [ "@message", "%{syslog_message}" ]
				replace => [ "@program", "%{syslog_program}" ]
				replace => [ "@type", "syslog imported" ]
			}
		}

		# Parse the date. This puts it into the @timestamp field on a successful
		# parse.
		date {
			match => [ "syslog_timestamp", "MMM  d HH:mm:ss", "MMM dd HH:mm:ss", "YYYY MMM  d HH:mm:ss", "YYYY MMM dd HH:mm:ss" ]
		}

		# Clean up the extra syslog_ fields generated above from grok.
		mutate {
			remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type", "message", "host" ]
		}

	}

}


output {
	# stdout {
	# 	codec => json
	# }


	elasticsearch {
		# Define our own... hosted on my computer
		# bind_host => "leo-linux"
		# bind_port => 9200
		host => "127.0.0.1"
		port => 9300

		cluster => "logstash_es"
		node_name => "logstash_0"

		# Index defaults to 'logstash-%{+YYYY.MM.dd}'
		# The templates being used can be defined using:
		template => "/etc/logstash/template/logstash.json"
		
	}
}

I do some processing on the messages coming in from port 4401, tagged as syslog_import. In the filter step, any messages tagged as syslog_import will be processed by grok, the parser. Let's take a closer look at the grok configuration:

grok {
	# Use the custom SYSLOGYEARTIMESTAMP pattern from the patterns
	# directory. We need this to define year.
	patterns_dir => "./patterns"

	# The pattern to match.
	# This is the standard syslog pattern.
	match => { "message" => "%{SYSLOGYEARTIMESTAMP:syslog_timestamp} (%{USER:syslog_user}\@)?%{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }

	# Add a few intermediate fields
	add_field => [ "received_at", "%{@timestamp}" ]
	add_field => [ "received_from", "%{host}" ]
}

This grok instance attempts to match the incoming message to the defined pattern. The syntax defining matched strings is %{PATTERN_NAME:variable_name} where PATTERN_NAME is a grok-pattern defined in /patterns/* in the .jar file and also in the directory defined in the patterns_dir directory, and variable_name is the name that can be used to referenced the matched value later on in the grok instance.

The SYSLOGYEARTIMESTAMP pattern is a custom pattern defined in my ./patterns directory.

cat patterns/extra 
SYSLOGYEARTIMESTAMP %{YEAR} %{MONTH} +%{MONTHDAY} %{TIME}

In the case above, syslog messages being imported whose date field matches the format given in SYSLOGYEARTIMESTAMP will be placed in the variable syslog_timestamp.



</syntaxhighlight >

Integration with Clients

https://groups.google.com/forum/#!msg/logstash-users/X6kNHU0alBg/j95HZkTLo-EJ