Logstash
Logstash is the open source copy of splunk - a log capturing, indexing, and searching service - using ElasticSearch as its search engine.
Installation
For detailed information, consult logstash's tutorial (http://logstash.net/docs/). Prior to logstash 1.4.0, the logstash package comes as a monolithic .jar file. To get started, install java and run the jar file.
If you are using logstash >= 1.4.0, it's probably easier to just install logstash and ElasticSearch from their repository:
[logstash-1.4]
name=logstash repository for 1.4.x packages
baseurl=http://packages.elasticsearch.org/logstash/1.4/centos
gpgcheck=1
gpgkey=http://packages.elasticsearch.org/GPG-KEY-elasticsearch
enabled=1
[elasticsearch-1.3]
name=Elasticsearch repository for 1.3.x packages
baseurl=http://packages.elasticsearch.org/elasticsearch/1.3/centos
gpgcheck=1
gpgkey=http://packages.elasticsearch.org/GPG-KEY-elasticsearch
enabled=1
Then run yum install logstash elasticsearch
Configuration
The configuration files for logstash are located at:
/etc/sysconfig/logstash/etc/logstash/conf.d/
The files in the conf.d directory will be treated as a single configuration file in sequence by their filenames by logstash.
You may want to change the DATA_DIR path in the logstash configuration. The actual input/parse/output configurations will be placed in the conf.d directory. More on this below.
ElasticSearch's files are at:
/etc/elasticsearch/elasticsearch.yml
You will need to configure ElasticSearch based on how you want to set up your search. For replication/sharding, you should ideally have more than one server. If you do have more than one server, make sure you have the node names set and have the autodiscoverer configured.
Configuration
The configuration you provide logstash defines how logstash deals with incoming messages. There are three main parts to the configuration:
- Input
- Filter
- Output
Input defines what ports logstash listens on and how to tag the incoming messages. Filter defines what logstash needs to do on the incoming messages, based on the tags defined from the input step. Output defines how these messages are stored.
For example, my current configuration is:
input {
# Import syslog messages
tcp {
type => syslog_import
port => 4401
}
# Accept syslog messages from hosts
syslog {
type => syslog
port => 5544
}
}
filter {
if [type] == "syslog" {
# Does the syslog parsing.
syslog_pri { }
mutate {
replace => [ "@source", "%{logsource}" ]
replace => [ "@message", "%{message}" ]
replace => [ "@program", "%{program}" ]
replace => [ "@type", "syslog" ]
}
# Date is parsed and placed into @timstamp.
date {
match => [ "syslog_timestamp", "MMM d HH:mm:ss", "MMM dd HH:mm:ss", "ISO8601" ]
}
# Clean up the extra syslog_ fields generated above from grok.
mutate {
remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type", "message", "logsource", "program"]
}
}
# For imported syslog messages...
if [type] == "syslog_import" {
if [message] =~ /last message repeated.*/ {
drop {
}
}
if [message] == "" {
drop {
}
}
# Parse with grok
grok {
# Use the custom SYSLOGYEARTIMESTAMP pattern from the patterns
# directory. We need this to define year.
patterns_dir => "./patterns"
# The pattern to match.
# This is the standard syslog pattern.
match => { "message" => "%{SYSLOGYEARTIMESTAMP:syslog_timestamp} (%{USER:syslog_user}\@)?%{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }
# Add a few intermediate fields
add_field => [ "received_at", "%{@timestamp}" ]
add_field => [ "received_from", "%{host}" ]
}
# When the above grok parsing fails, a '_grokparsefailure' tag gets
# added to the message. In that case, we attempt to update some fields.
# Why? Beats me.
if !("_grokparsefailure" in [tags]) {
mutate {
replace => [ "@source", "%{syslog_hostname}" ]
replace => [ "@message", "%{syslog_message}" ]
replace => [ "@program", "%{syslog_program}" ]
replace => [ "@type", "syslog imported" ]
}
}
# Parse the date. This puts it into the @timestamp field on a successful
# parse.
date {
match => [ "syslog_timestamp", "MMM d HH:mm:ss", "MMM dd HH:mm:ss", "YYYY MMM d HH:mm:ss", "YYYY MMM dd HH:mm:ss" ]
}
# Clean up the extra syslog_ fields generated above from grok.
mutate {
remove_field => [ "syslog_hostname", "syslog_message", "syslog_program", "syslog_timestamp", "type", "message", "host" ]
}
}
}
output {
# Debugging
# stdout {
# codec => json
# }
elasticsearch {
# Define our own... hosted on my computer
# bind_host => "leo-linux"
# bind_port => 9200
host => "127.0.0.1"
port => 9300
cluster => "logstash_es"
node_name => "logstash_0"
# Index defaults to 'logstash-%{+YYYY.MM.dd}'
# The templates being used can be defined using:
template => "/etc/logstash/template/logstash.json"
}
}
The major sections of the configuration file are:
- Inputs
- Parsing/Filtering
- Outputs
You can see the entire list of available plugins for each of these sections at: http://logstash.net/docs/1.4.2/
Inputs
The inputs define what logstash will listen to for information. The configuration above listens on a tcp port (for backlog imports) for raw text and another port for syslog.
Parsing / Filtering
For each of the inputs, certain things are done in order to parse the input data into variables which are then passed to the output section.
Operations to the data are done through more sets of plugins (think: functions). For example, text parsing is done through grok. Parameters to these 'functions' are passed as parameters in the grok block. When grok fails at parsing a certain string, it will add an additional tag (ie: variable) called _grokparsefailure which can be used later on in the parse section.
Variables starting with a '@' are used by ElasticSearch to denote mandatory fields (... I think?) which are defined in the ElasticSearch template file.
Example
grok {
# Use the custom SYSLOGYEARTIMESTAMP pattern from the patterns
# directory. We need this to define year.
patterns_dir => "./patterns"
# The pattern to match.
# This is the standard syslog pattern.
match => { "message" => "%{SYSLOGYEARTIMESTAMP:syslog_timestamp} (%{USER:syslog_user}\@)?%{SYSLOGHOST:syslog_hostname} %{DATA:syslog_program}(?:\[%{POSINT:syslog_pid}\])?: %{GREEDYDATA:syslog_message}" }
# Add a few intermediate fields
add_field => [ "received_at", "%{@timestamp}" ]
add_field => [ "received_from", "%{host}" ]
}
This grok instance attempts to match the incoming message to the defined pattern. The syntax defining matched strings is %{PATTERN_NAME:variable_name} where PATTERN_NAME is a grok-pattern defined in /patterns/* in the .jar file and also in the directory defined in the patterns_dir directory, and variable_name is the name that can be used to referenced the matched value later on in the grok instance.
The SYSLOGYEARTIMESTAMP pattern is a custom pattern defined in my ./patterns directory.
cat patterns/extra
SYSLOGYEARTIMESTAMP %{YEAR} %{MONTH} +%{MONTHDAY} %{TIME}
In the case above, syslog messages being imported whose date field matches the format given in SYSLOGYEARTIMESTAMP will be placed in the variable syslog_timestamp.
Outputs
The outputs section defines what LogStash will do with the variables generated from the parsing/filtering section.
To debug the inputs/filtering section, you can do:
stdout {
codec => json
}
Variables / tags generated can be seen as part of a json object.
ElasticSearch takes in a template which defines the schema of indexes generated by logstash. This template is optional, since logstash will use a default template by default.
In the configuration example above, a template was defined for the ElasticSearch output.
{
"template": "logstash-*",
"settings" : {
"index.query.default_field" : "@message"
},
"mappings": {
"_default_": {
"_all": { "enabled": false },
"_source": { "compress": false },
"dynamic_templates": [
{
"fields_template" : {
"mapping": { "type": "string", "index": "not_analyzed" },
"path_match": "@fields.*"
}
},
{
"tags_template" : {
"mapping": { "type": "string", "index": "not_analyzed" },
"path_match": "@tags.*"
}
}
],
"properties" : {
"@fields": { "type": "object", "dynamic": true, "path": "full" },
"@timestamp" : { "type" : "date", "index" : "not_analyzed" },
"@program" : { "type" : "string", "index" : "not_analyzed" },
"@source" : { "type" : "string", "index" : "not_analyzed" },
"@message" : { "type" : "string", "analyzer" : "whitespace" },
"@type" : { "type" : "string", "index" : "not_analyzed" }
}
}
}
}
The variables/tags that were generated from the parse field should match the property names defined in the template file. Depending on what you want out of ElasticSearch, you may or may not want to have every field analyzed.
Be careful with templates though. If the properties defined in the template file are not provided by the filtering/parsing section, the log entry will not be added to ElasticSearch.
See Also
- https://groups.google.com/forum/#!msg/logstash-users/X6kNHU0alBg/j95HZkTLo-EJ
- http://www.chriscowley.me.uk/blog/2014/03/21/logstash-on-centos-6/
- https://www.digitalocean.com/community/tutorials/how-to-use-logstash-and-kibana-to-centralize-logs-on-centos-6
Various Linux Notes - /etc/fstab
- Access.conf
- ACL
- Apache Proxy to Internal Server
- APM X-C1 (Mustang)
- ARP
- Authselect
- Bash Scripting
- Blockparser
- Booting Linux without a Graphics Card
- Building Container Images
- Burning CD/DVD in Linux
- Clear RAID Signatures on Linux
- Cobbler
- Colorized Terminal Outputs
- Compiling MIPS
- Configure Sendmail
- CPanel
- CPanel Fork Bomb Protection
- CPU Frequency Scaling
- Create a Linux User with an Empty Password
- Cron and PAM Issues
- Dell OpenManage
- Diff Two Command Outputs
- DirectAdmin
- Disable Filesystem Check on Startup
- DNS Ad Blocker
- Driver Disk
- Drop caches
- End / Home keys don't work in Terminal
- Entropy in the Linux Kernel
- Entropy Source using RTL-SDR
- Exit Codes
- Extract .exe Resources with dd
- File Attributes
- Fixing ixgbe unsupported SFP+ module type was detected
- Get Active Linux Virtual Console
- Getting Hardware UUID
- Hosts.deny
- How to change Linux desktop user directories
- How to hot-swap SATA disks on Linux
- HP Smart Storage Administrator
- Hyper-threading
- IBM Spectrum Archive
- IBM Spectrum Protect
- IBM Tape
- IBM Tape Diagnostic Tool
- InterWorx
- Kerberize NFS
- Kerberize SSH
- Linux Clustering
- Linux Fonts
- Linux Namespaces
- Linux Network Interface Naming
- Linux Nvidia Driver
- Linux Process Accounting
- Linux Uptime in Seconds
- Linux UTF-8 Font
- Mainline Kernel on CentOS 7
- Missing Fonts
- Mod fastcgi Install on Apache 2 / cPanel
- Mod fcgid
- Monitoring network traffic in Linux
- Mounting / Unmounting KVM Image
- Mounting Samba (CIFS) shares
- Multiple Networks on Linux
- MySQL Database with Hash Sign
- No Console Output
- Number of Files Opened
- Open OnDemand
- Packing and unpacking initrd
- PAM Issues
- Partition Alignment
- Patching a binary file with dd
- Perl Module Location
- Raspberry Pi
- Red Hat kickstart
- Red Hat to Debian
- Reverse SSH Tunnel
- Ruby on Rails under cPanel
- Rutorrent + rtorrent Installation Guide on CentOS 6.4
- Self Signed SSL Certificates
- Service Management
- Sick Beard
- StartSSL Free Certificate
- Symlink
- Taking a Screenshot in X11
- Timezone
- Tor
- TOR Transparent Proxy
- Traefik
- Troubleshooting a Slow Linux System
- Turning on swap with a page file
- Udev Rules
- Verify SSL Certificate matches Private Key
- VMware Workstation
- Webcam
- X Display Manipulation
- X Forwarding
Linux Tools and Utilites - Anaconda
- Ansible
- Aria2
- Autofs
- Awk
- Badblocks
- Bash Shell
- Binwalk
- Bosh
- Ceph
- Chntpw
- Chrony
- Clonezilla
- Cloud-init
- CloudStack
- Column
- Cron
- Curl
- Cvs2git
- Date
- Dbus
- Dd
- Dm-crypt
- Dovecot
- DRBD
- ElasticSearch
- Enroot
- Environment Module
- Envsubst
- Fail2ban
- FFmpeg
- Find
- Firecracker
- Flashrom
- Foreman
- FortiClient
- Fswebcam
- Galaxy
- Git
- Gnome
- Gobetween
- GPFS
- Grafana
- Grub
- Hdparm
- Home Assistant
- How to disable SELinux
- Htaccess
- Infiniband
- InfluxDB
- InfluxDB 1.x
- Insert a kickstart file into a iso image
- Inspircd
- IOzone
- Iperf3
- Ipmitool
- Irqbalance
- John The Ripper
- Lightdm
- Lm sensors
- Logrotate
- LSF
- LVM
- Lynx
- Mailx
- Md5sum
- Mdadm
- Midnight Commander
- Motion
- Mount
- Mutt
- Nomad
- OpenLDAP
- Openocd
- OpenSSL
- OpenVPN
- Packer
- PHP
- Pi-hole
- Postgres
- PowerBroker Identity Service
- Proxmox
- Pueue
- PulseAudio
- Puppet
- Quota
- Red Hat Satellite
- Restic
- Rsync
- Rtorrent
- Ruby
- Sabnzbd
- Sage
- Samba
- Screen
- Sed
- SELinux
- Sendmail
- Shell Configs
- Singularity
- Sleep
- Slurm
- SMART
- Sonarr
- Sqlite
- SquashFS
- Squid
- SSH
- Steam
- Stoken
- Strace
- Sudo
- Sync
- Sysctl
- Syslog
- Sysrq
- System Security Services Daemon (SSSD)
- Systemd
- Tar
- Tcsh
- Telegraf
- Terraform
- Thttpd
- Tmux
- Tomcat
- Top
- Umask
- Unix2dos
- Vim
- Virsh
- Virt-customize
- VirtualBox
- VirtualGL
- Visidata
- Vnstat
- Weechat
- Wget
- XFS
- Youtube-dl
- ZFS
- Zram
Package Management Linux Distributions Networking - Blazemeter
- CSF/LFD
- Exim
- Firewall
- FreeIPA
- Get DHCP Network Settings
- IP Aliasing
- Ipset
- IPTables
- IPv6
- IPXE
- Link Aggregation
- Linux Network Namespaces
- MTU
- Net-tools to iproute2
- Netcat
- Open vSwitch
- OpenWRT
- Postfix
- Raspberry Pi Torified Wifi
- Socat
- Static Routes
- StrongSwan
- Tcpdump
- Traffic Forwarder using IPTables
- Wi-Fi
- WireGuard
Containers Linux