Logstash: Difference between revisions

From Leo's Notes
This page was last edited on 3 March 2014, at 23:21.
No edit summary
No edit summary
Line 3: Line 3:
== Installation ==
== Installation ==


For detailed information, consult logstash's tutorial at http://logstash.net/docs/1.1.10/tutorials/getting-started-centralized
For detailed information, consult logstash's tutorial (http://logstash.net/docs/). Prior to logstash 1.4.0, the logstash package comes as a monolithic .jar file. To get started, install java and run the jar file.


=== Elastic Search ===
== Running Logstash ==
Download and extract the archive. Install java:
 
wget jre-7u21-linux-x64.rpm
I made a wrapper script called run.sh which launches the jar file with my configuration.  
rpm -ivh jre-7u21-linux-x64.rpm
Then run elastic search. You probably want to specify a configuration file.  


<syntaxhighlight lang="bash" line start="1" enclose="div">
<syntaxhighlight lang="bash" line start="1" enclose="div">
#!/bin/bash
java -jar logstash-1.3.3-flatjar.jar agent -f configuration.conf -- web
</syntaxhighlight>
The core of logstash is the agent. The web site (formerly known as a separate package called kibana) is built in and can be started by appending <code> -- web</code> to the command line.
=== Configuration ===
The configuration you provide logstash defines how logstash deals with incoming messages. There are three main parts to the configuration:
# Input
# Filter
# Output
Input defines what ports logstash listens on and how to tag the incoming messages. Filter defines what logstash needs to do on the incoming messages, based on the tags defined from the input step. Output defines how these messages are stored.
For example, my current configuration is:
<syntaxhighlight lang="text" line start="1" enclose="div">
input {
input {
  stdin {
# Import syslog messages
    # A type is a label applied to an event. It is used later with filters
tcp {
    # to restrict what filters are run against each event.
type => syslog_import
    type => "human"
port => 4401
  }
}


  syslog {
# Accept syslog messages from hosts
        type => syslog
syslog {
        port => 5544
type => syslog
  }
port => 5544
}
}
}
filter {
# For imported syslog messages...
if [type] == "syslog_import" {
# Parse with grok
grok {
# Use the custom SYSLOGYEARTIMESTAMP pattern from the patterns
# directory. We need this to define year.
patterns_dir => "./patterns"
# The pattern to match.
# This is the standard syslog pattern.
match => { "message" => "%{SYSLOGYEARTIMESTAMP:syslog_timestamp} (%{USER:syslog_user}\@)?%{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }
# Add a few intermediate fields
add_field => [ "received_at", "%{@timestamp}" ]
add_field => [ "received_from", "%{host}" ]
}
# When the above grok parsing fails, a '_grokparsefailure' tag gets
# added to the message. In that case, we attempt to update some fields.
# Why? Beats me.
if !("_grokparsefailure" in [tags]) {
mutate {
replace => [ "@source_host", "%{syslog_hostname}" ]
replace => [ "@message", "%{syslog_message}" ]
replace => [ "@program", "%{syslog_program}" ]
}
}
# Parse the date. This puts it into the @timestamp field on a successful
# parse.
date {
match => [ "syslog_timestamp", "MMM  d HH:mm:ss", "MMM dd HH:mm:ss", "YYYY MMM  d HH:mm:ss", "YYYY MMM dd HH:mm:ss" ]
}
# Clean up the extra syslog_ fields generated above from grok.
mutate {
remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type" ]
}
}
}


output {
output {
  # Print each event to stdout.
# Print each event to stdout.
  stdout {
#  stdout {
    # Enabling 'debug' on the stdout output will make logstash pretty-print the
# Enabling 'debug' on the stdout output will make logstash pretty-print the
    # entire event as something similar to a JSON representation.
# entire event as something similar to a JSON representation.
#    debug => true
#    debug => true
  }
#  }
  # You can have multiple outputs. All events generally to all outputs.
  # Output events to elasticsearch
# You can have multiple outputs. All events generally to all outputs.
  elasticsearch {
# Output events to elasticsearch
    # Setting 'embedded' will run  a real elasticsearch server inside logstash.
elasticsearch {
    # This option below saves you from having to run a separate process just
# Setting 'embedded' will run  a real elasticsearch server inside logstash.
    # for ElasticSearch, so you can get started quicker!
# This option below saves you from having to run a separate process just
    embedded => true
# for ElasticSearch, so you can get started quicker!
  }
embedded => true
}
}
}
</syntaxhighlight>


Then, just run the monolithic jar with the agent and web enabled like so:
</syntaxhighlight >
 


java -jar logstash-1.2.1-flatjar.jar agent -f hello.conf -- web





Revision as of 23:21, 3 March 2014

Logstash is the open source version of splunk, using ElasticSearch as its search engine.

Installation

For detailed information, consult logstash's tutorial (http://logstash.net/docs/). Prior to logstash 1.4.0, the logstash package comes as a monolithic .jar file. To get started, install java and run the jar file.

Running Logstash

I made a wrapper script called run.sh which launches the jar file with my configuration.

#!/bin/bash
java -jar logstash-1.3.3-flatjar.jar agent -f configuration.conf -- web

The core of logstash is the agent. The web site (formerly known as a separate package called kibana) is built in and can be started by appending -- web to the command line.

Configuration

The configuration you provide logstash defines how logstash deals with incoming messages. There are three main parts to the configuration:

  1. Input
  2. Filter
  3. Output

Input defines what ports logstash listens on and how to tag the incoming messages. Filter defines what logstash needs to do on the incoming messages, based on the tags defined from the input step. Output defines how these messages are stored.

For example, my current configuration is:


input {
	# Import syslog messages
	tcp {
		type => syslog_import
		port => 4401
	}

	# Accept syslog messages from hosts
	syslog {
		type => syslog
		port => 5544
	}
}

filter {

	# For imported syslog messages...
	if [type] == "syslog_import" {
		# Parse with grok
		grok {
			# Use the custom SYSLOGYEARTIMESTAMP pattern from the patterns
			# directory. We need this to define year.
			patterns_dir => "./patterns"

			# The pattern to match.
			# This is the standard syslog pattern.
			match => { "message" => "%{SYSLOGYEARTIMESTAMP:syslog_timestamp} (%{USER:syslog_user}\@)?%{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }

			# Add a few intermediate fields
			add_field => [ "received_at", "%{@timestamp}" ]
			add_field => [ "received_from", "%{host}" ]
		}
		
		# When the above grok parsing fails, a '_grokparsefailure' tag gets
		# added to the message. In that case, we attempt to update some fields.
		# Why? Beats me.
		if !("_grokparsefailure" in [tags]) {
			mutate {
				replace => [ "@source_host", "%{syslog_hostname}" ]
				replace => [ "@message", "%{syslog_message}" ]
				replace => [ "@program", "%{syslog_program}" ]
			}
		}

		# Parse the date. This puts it into the @timestamp field on a successful
		# parse.
		date {
			match => [ "syslog_timestamp", "MMM  d HH:mm:ss", "MMM dd HH:mm:ss", "YYYY MMM  d HH:mm:ss", "YYYY MMM dd HH:mm:ss" ]
		}

		# Clean up the extra syslog_ fields generated above from grok.
		mutate {
			remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type" ]
		}
	}
}


output {
	# Print each event to stdout.
	#  stdout {
	# Enabling 'debug' on the stdout output will make logstash pretty-print the
	# entire event as something similar to a JSON representation.
	#    debug => true
	#  }
	
	# You can have multiple outputs. All events generally to all outputs.
	# Output events to elasticsearch
	elasticsearch {
		# Setting 'embedded' will run  a real elasticsearch server inside logstash.
		# This option below saves you from having to run a separate process just
		# for ElasticSearch, so you can get started quicker!
		embedded => true
	}
}



Integration with Clients

https://groups.google.com/forum/#!msg/logstash-users/X6kNHU0alBg/j95HZkTLo-EJ