Monitoring#

You can debug system issues in many different ways.

Each Squirro service maintains its own log file, which helps with finding the source of an error. For the list of services, see the Services page.

Additionally, systemd automatically restarts services that fail completely. The following sections also explain a few basic operating system commands, which help you find out whether the base system is running smoothly.

Log Files#

Most log files are located under the /var/log/ directory. The following table shows the relevant log files and their usage.

Log File

Service

OS

Additional Notes

/var/log/squirro/SERVICE_NAME/SERVICE_NAME.log

Squirro services

All

Detailed log file about each service.

/var/log/squirro/SERVICE_NAME/stdout.log

Squirro services

All

Messages sent to the standard output stream of the service and not logged in the main service log file.

/var/log/squirro/SERVICE_NAME/stderr.log

Squirro services

All

Messages sent to the standard error stream of the service (typically error messages) and not logged in the main service log file. Might contain useful information when a service is unable to boot up.

/var/log/squirro/SERVICE_NAME/nginx-access.log

Nginx (Squirro services)

All

Every request to the web services is recorded in this log file in a line-by-line format.

/var/log/squirro/SERVICE_NAME/nginx-error.log

Nginx (Squirro services)

All

Records errors on the HTTP level. When a service is stopped, errors may show up here indicating that the service is not reachable.

/var/log/squirro/update-cluster-node.log

/var/log/squirro/update-storage-node.log

All

Update log for Squirro Cluster/Storage node.

/var/log/messages

All

General system log. Serious system failures are recorded here.

/var/log/elasticsearch/ES_CLUSTER_NAME*

Elasticsearch

All

ES_CLUSTER_NAME.log: Records cluster information and major failures.

ES_CLUSTER_NAME_index_indexing_slowlog.json: Records indexing operations that exceed the configured slow log thresholds.

ES_CLUSTER_NAME_index_search_slowlog.json: Records search queries that exceed the configured slow log thresholds.

/var/log/redis/*.log

Redis

All

/var/log/mysqld.log

MySQL

RHEL

/var/log/mariadb/mariadb.log

MariaDB

RHEL

/var/log/cron

Cron

/var/log/secure

/var/log/audit/audit.log

System

RHEL

Used to debug connection issues

Additionally, the /var/lib/squirro/ directory contains the following log files:

Log File

Service

OS

Additional Notes

/var/lib/squirro/datasource/job_logs/*.log

datasource (sqdatasourced)

All

Contains rotated log files for the created data sources. Any logs written while the source is created and data is loaded into the system appear here. Those are the dataloader logs, written before the pipeline transforms the data. For the pipeline, see the ingester logs.

/var/lib/squirro/machinelearning/job_logs/*.log

machinelearning (sqmachinelearningd)

All

Contains log files for the machinelearning jobs that run on the server. Each machinelearning job uses its own log file and any output during its execution is logged there (for example, output during the training of a model).

You can change the log level for each service, in the ini file corresponding to that service under /etc/squirro/.

For any of the services, add the following to the ini file to adjust the log level:

[logger_root]
level = INFO

Restart the affected service afterwards for the new log level to take effect. For instructions, see the Services page.

Studio Plugin Logs#

Studio plugins, also referred to as Studio components, run inside the frontend service and do not have log files of their own. Everything a Studio plugin logs is written to /var/log/squirro/frontend/frontend.log, alongside the rest of the frontend service output.

The main file of a plugin logs under a name of the form studio_plugin_PLUGIN_NAME, where PLUGIN_NAME is the name the plugin was uploaded under. That name is not always identical to the title shown in the user interface. For example, the main file of the plugin titled Log Files logs under studio_plugin_log_files.

To follow a single plugin, filter the frontend log on that name:

tail -f /var/log/squirro/frontend/frontend.log | grep studio_plugin_monitoring

To follow every Studio plugin at once, filter on the shared prefix:

tail -f /var/log/squirro/frontend/frontend.log | grep studio_plugin_

Three kinds of output do not appear under that prefix, so a filter on the prefix alone does not capture everything a plugin writes:

  • Output from another file of the same plugin

    A plugin can be split across several files, and not all of them log under the studio_plugin_ prefix. A file that does not is named after itself, for example helpers for a module named helpers.py. A plugin can also name its logger explicitly, in which case that name appears instead. Search for those names as well when following a plugin that is split across several files.

  • A plugin that fails to load

    The failure, including the traceback, is recorded in the same file, but against the frontend service rather than the plugin, so the plugin name shows up only as part of the reported file path. Search the log for the plugin name instead of the studio_plugin_ prefix.

  • Output that bypasses the logging system

    Messages written directly to the standard output or standard error streams, such as the output of a print() call in plugin code, land in /var/log/squirro/frontend/stdout.log and /var/log/squirro/frontend/stderr.log.

To change how much detail a plugin records, set level in the [logger_root] section of /etc/squirro/frontend.ini, using one of the log levels listed on the common.ini page. Restart the sqfrontendd service afterwards for the change to take effect. For instructions, see the Services page.

In a deployment with more than one frontend node, a plugin logs on whichever node handled the request, so check the frontend log on each node.

You can also read the frontend log from the Squirro user interface, without server access. In the Server space, select Log Files, then choose frontend.log. That view is restricted to cluster administrators and lists the service log files only, so reading stdout.log or stderr.log still requires server access. For more information, see the Cluster Status page.

RHEL Service Monitoring with systemctl#

With Red Hat Enterprise Linux (RHEL), Squirro relies on systemd to control and manage the Squirro services.

To check all the services, use the following command:

systemctl list-units --type service --all

To inspect a single service, use:

systemctl status SERVICE_NAME

That command also returns fundamental information about the service, such as its current status and process ID. Called with root permissions, it also returns the last lines of the logs.

To restart a particular service, run the following command:

systemctl restart SERVICE_NAME

When Squirro services go down, the systemd daemon automatically attempts to restart the service. If the service remains inactive, inspect the logs belonging to that service. These log files consist of:

  • /var/log/squirro/SERVICE_NAME/SERVICE_NAME.log

  • /var/log/squirro/SERVICE_NAME/stderr.log

Monitoring Services from the Web Interface#

As a server administrator, you can also inspect the status of the services from the web interface.

That feature is available as a plugin in the Server space, as shown below. For more information, see the Monitoring Plugin section of the Cluster Status page.

image1

System Commands#

The Squirro services are standard Unix daemons, so you can use standard Linux utilities to debug any issues that arise.

Processor Usage#

Two standard commands report the current processor usage: uptime and top.

uptime#

Next to some uptime information, the uptime command outputs the load average for the past 1, 5, and 15 minutes. The load average is a simple metric showing how many processes had to wait for processing. It is usually close to or below 1.0. A value above 5.0 means the load is quite high, and higher values are unusual.

When the load average is high, top usually shows the processes that are generating the load. When the CPU usage shown by top is low despite a high load average, that may indicate issues with I/O, such as disk performance.

top#

The top command shows a list of all processes on the system, sorted by current CPU usage. Press M on the keyboard (upper case, so use Shift+m) to sort the list by memory usage.

Memory Usage#

To debug the memory usage of individual processes, use the top command. To see the memory usage of the system as a whole, use free.

free#

The free command outputs statistics on how much RAM the system uses. The most useful value to consider is the available column, which estimates how much memory can still be given to new processes without swapping.

By default, free outputs all values in kibibytes. Called with the -m parameter (free -m), it outputs all values in mebibytes instead.

When free memory is very low, the system may run into issues with memory usage. In some cases, the kernel needs to end processes to make space. Those instances appear in the standard system log /var/log/messages, in lines such as Out of memory: kill process 23123.

Disk Usage#

A full disk prevents the system from working. The df command helps with finding those issues.

df#

Use the df command to see a list of all partitions and their disk usage. The Use% column shows the usage as a percentage. Anything above 95% counts as full and usually hinders the system from working well.

When you are experiencing full disks, consider enlarging the corresponding disk, or visit the Squirro Support website and submit a technical support request for ways to remove extra data.

Following Log Files#

tail#

Log files capture a lot of information. Follow these files with the tail command, using its -f parameter to follow all updates on a file.

For example:

tail -f /var/log/squirro/topic/topic.log

That command shows a real-time view of what is written into the topic service log file.

tail also accepts multiple file names and wildcards. Monitor all Squirro service log files as follows:

tail -f /var/log/squirro/*/*.log

grep#

The grep command searches files for occurrences of a specific text. For example, if Squirro reports errors but you are unsure where they come from, the following command helps you pin down the responsible service:

grep ERROR /var/log/squirro/*/*.log

That command outputs a list of all Squirro log files that contain the text ERROR, together with the lines that contain this text.

Squirro Logs#

All the Squirro logs are stored in /var/log/squirro. Each service has its own directory, which you query as follows:

tail -f /var/log/squirro/SERVICE_NAME/*.log

For example, to inspect the topic service:

tail -f /var/log/squirro/topic/*.log

To check the logs of all services:

tail -f /var/log/squirro/*/*.log

Squirro Log Utilities#

For convenience, Squirro provides shell functions that combine common log monitoring operations.

squirro_tail_errors#

The squirro_tail_errors function monitors all Squirro service logs and filters for warnings and errors only. That is particularly useful for troubleshooting issues across all services.

squirro_tail_errors

This command is equivalent to:

tail -f /var/log/squirro/*/*.log /var/log/squirro/ingester/processor_*/*.log | grep -E 'WARNING|ERROR'

The function monitors log files from:

  • All Squirro service directories under /var/log/squirro/*/.

  • All ingester processor directories under /var/log/squirro/ingester/processor_*/.

Only log entries containing WARNING or ERROR appear, which makes it easier to spot issues without being overwhelmed by informational messages.

squirro_tail_logs#

For complete log monitoring without filtering, use:

squirro_tail_logs

That function shows all log entries from all Squirro services in real time.

Note

These functions come from the Squirro shell aliases, which Squirro servers load automatically. If the functions are not available, load them manually by running:

source /etc/profile.d/squirro-aliases.sh

To make them available automatically in future shell sessions, add this line to your shell profile (for example, ~/.bashrc or ~/.profile).

The Ingester Service#

Due to its complexity, the ingester service has a different log structure. The service manages a set of processes, named processor_1 through processor_N. Each process maintains its own log in its own directory. The easiest way to debug is to merge their content with the following command:

tail -f /var/log/squirro/ingester/processor_*/*.log