Cluster Expansion#
Note for administrators
End of support for expanding a Squirro cluster on Linux manually. This page remains available for reference only and is no longer maintained. Use Ansible for all installations and upgrades. See the Install and Manage Squirro with Ansible page.
Note
Squirro removed the Cluster Service (sqclusterd) in version 3.16.2, together with its cluster.ini configuration file and the endpoint_discovery and db_endpoint_discovery settings. The information on this page applies only to Squirro versions earlier than 3.16.2. For high-availability and multi-node deployments on Squirro 3.16.2 and later, visit the Squirro Support website and submit a technical support request.
This page details how to add cluster nodes to a Squirro installation. For base Linux installation, see the Installing Squirro on Linux page.
For background on Squirro cluster setups, see the How Squirro Scales page. That page provides an overview of Squirro components and their scaling considerations.
Prerequisites#
To add cluster nodes, first ensure that:
The Linux machines are up and running, see the System Requirements page for the list of supported operating systems.
The network is set up and all the machines that are to become part of the Squirro cluster can talk to each other.
The firewalls are open between the machines with the documented ports accessible.
The Squirro YUM repository is configured and accessible. In an enterprise environment, where this poses a problem, offline installation is available. In this case, visit the Squirro Support website and submit a technical support request.
Overview#
Any expansion of the cluster requires some work on the old and new nodes. This is outlined in the processes below by splitting the work up into sections, based on where the work is to be executed.
The process as described here involves some cluster downtime. It is possible to expand a Squirro cluster without any downtime involved - but that process requires a bit more planning and orchestration. If you need a downtime-free expansion, visit the Squirro Support website and submit a technical support request.
Scaling Considerations#
Cluster nodes and storage nodes can be scaled independently. You can add storage nodes without modifying the cluster node configuration, and vice versa.
Adding storage nodes
Add storage nodes one at a time. Adding multiple nodes simultaneously increases the risk of discovery and replication issues during cluster formation.
Master-eligible node count
To avoid split-brain situations, keep an odd number of master-eligible storage nodes (1, 3, 5, and so on). With an even number of nodes, the cluster may fail to elect a leader if the nodes become equally divided across a network partition. Squirro recommends following the Elasticsearch guidance on limiting the number of master-eligible nodes to a small, odd-numbered set. For more information, see the Elasticsearch high-availability guide on the Elasticsearch website.
Storage Node Installation#
The Squirro storage nodes are based on Elasticsearch. As such, some of the configurations needed for adding a storage node are found in the Elasticsearch configuration files.
Process for New Storage Nodes
Follow the steps below to add a new storage node.
Install/update the storage node package, as described in Storage Node Installation.
Apply the configuration from the previous storage nodes to the new one. Copy over the following setting from
/etc/elasticsearch/elasticsearch.yml:
cluster.name
Make sure that this setting value is copied from the previous storage nodes to the new one - and not the other way around.
Process for All Storage Nodes (Existing as Well as New)
Allow the hosts to discover each other. Again in
/etc/elasticsearch/elasticsearch.ymlchange the following settings:Allow Elasticsearch to bind to a network interface by using the following config:
/etc/elasticsearch/elasticsearch.ymlnetwork.host: <server ip>,127.0.0.1
This is a list of the server’s own IP addresses (as can be retrieved with
ip addrfor example), for example:10.1.4.5.Set
discovery.seed_hostsandcluster.initial_master_nodesto a list of all the storage nodes that have been set up. For example:/etc/elasticsearch/elasticsearch.ymldiscovery.seed_hosts: ["<storagenode1 ip>", "<storagenode2 ip>", "<storagenode3 ip>"] cluster.initial_master_nodes: ["<storagenode1 ip>", "<storagenode2 ip>", "<storagenode3 ip>"]
This is the easiest way to set up discovery and ensure all the Elasticsearch nodes can see each other. However, there are also other ways of configuring the discovery of the Elasticsearch nodes.
This is documented by Elasticsearch in the Discovery section of the Elasticsearch Manual.
Also, optionally, you can set the node name to a friendlier name:
/etc/elasticsearch/elasticsearch.yml# User friendly node name node.name: test-node1
For new nodes remove the current Elasticsearch state:
mv /var/lib/elasticsearch/nodes /tmp/
Note
You can also remove this folder instead of moving it to
/tmp. Moving it allows you to recover if you ran this on the wrong node.Restart the service for the settings to take effect.
systemctl restart elasticsearch
To verify the nodes discovered each other and formed a cluster, you can debug with this command:
es_curl https://localhost:9200/_nodes?pretty | less
If successful, you should see the correct number of nodes in the output:
{ "_nodes" : { "total" : 3, "successful" : 3, "failed" : 0 } }
Set up the number of shards and number of replicas.
Modify
number_of_shardsandnumber_of_replicasin the templates. For more information, see the Configuring Elasticsearch Templates page. For multi-node storage setups, setnumber_of_replicasto1andnumber_of_shardsto the number of Elasticsearch nodes.Now update the shards and replicas settings of indices that were already present on the cluster before updating the templates using the curl command below. You can also selectively update the shards and replica settings of a particular index (instead of all indices) by replacing the
*with the name of the index. This is a cluster wide setting and only needs to be done on one of the nodes of ES cluster.es_curl -XPUT https://127.0.0.1:9200/*/_settings -H "Content-Type: application/json" -d '{"index": {"number_of_replicas": 1}}'
Cluster Node Installation#
Process for Each New Squirro Cluster Node Server#
Install the cluster node package, as described in Cluster Node Installation.
Ensure that each of the cluster node can talk to the Elasticsearch cluster (Squirro Storage Node). Change the config at
/etc/squirro/common.inito[index] es_index_servers = <storagenode1 ip>:9200,<storagenode2 ip>:9200,<storagenode3 ip>:9200
Allowlist all the Squirro Cluster nodes in the following nginx ipfilter files:
/etc/nginx/conf.d/ipfilter-cluster-nodes.inc/etc/nginx/conf.d/ipfilter-api-clients.inc/etc/nginx/conf.d/ipfilter-monitoring.inc.In each of these files include each IP address as follows:
allow <clusternode1 ip>; allow <clusternode2 ip>; allow <clusternode2 ip>;
Alternatively, the
allowdirective also accepts network addresses, for example10.1.4.0/24to allowlist an entire network.
Reload nginx at each of the cluster nodes.
$ systemctl reload nginx
Set up the shared storage that the cluster nodes rely on, as described in the Setting up Cluster Node Storage section below.
Configure the relational database, as described in the Relational Database Configuration section below.
Configure Redis, as described in the Redis Configuration section below.
Start all Squirro services on the new cluster node.
$ squirro_start
For high-availability deployments where the relational database (MariaDB or PostgreSQL) or Redis is replicated across nodes, visit the Squirro Support website and submit a technical support request.
Relational Database Configuration#
The settings below apply to MariaDB, the default relational database. When using PostgreSQL instead, see the Configure PostgreSQL as the Database Backend page.
Give the node a unique identity within the cluster. The
/etc/mysql/conf.d/replication.cnffile is installed with the following two settings commented out, so set both of them on every node:server_idAn integer that must be unique across the whole cluster. For example, use
10for the first server in the cluster,11for the second, and so on.report_hostThe name of the server as it is reported to the other hosts, for example
node01.
Raise the limits on open files and maximum connections. Create the
/etc/mysql/conf.d/maxconnections.cnffile with the following content:[mysqld] open_files_limit = 8192 max_connections = 500
Set
max_connectionshigher depending on the number of cluster nodes. Squirro recommends at least 150 connections for each cluster node.Restart MariaDB to apply both changes.
$ systemctl restart mariadb
Redis Configuration#
All cluster nodes must reach the same Redis instances, one for storage and one for the cache. The /etc/redis/redis.conf and /etc/redis/cache.conf files are installed to listen on the local interface only, so extend the listening addresses on the server that runs Redis.
Edit both
/etc/redis/redis.confand/etc/redis/cache.conf, listing every cluster node:bind 127.0.0.1 <clusternode1 ip> <clusternode2 ip> <clusternode3 ip>
Where the cluster node addresses are not fixed, use
bind 0.0.0.0instead, which configures Redis to listen on all network interfaces and IP addresses on the server.Warning
Redis must never be reachable from client or public networks. Restrict ports 6379 and 6380 to the cluster nodes at the network or security-group level, as described on the System Requirements page. Squirro configures password authentication for both Redis instances, but network isolation is the primary defense layer.
Restart both Redis services.
$ systemctl restart redis-server $ systemctl restart redis-server-cache
Point every cluster node at that server by setting the
hostoption of each[redis*]section in the/etc/squirro/common.inifile.
Setting up Cluster Node Storage#
In a multi-node cluster deployment, several directories must reside on a shared filesystem accessible from all cluster nodes, so the platform can read and write shared data regardless of which node handles a request.
For what is stored and how storage requirements vary by workload, see the System Requirements page. This section describes the directories to share and how to mount them.
SELinux Configuration for NFS#
On systems with SELinux in enforcing mode, for example RHEL or CentOS, nginx requires an additional SELinux policy module to access files served from an NFS mount.
Run the following script on each cluster node to create and install the policy module:
mkdir -p /etc/squirro/selinux
cat > /etc/squirro/selinux/squirro-nginx-nfs.te << 'EOF'
module squirro-nginx-nfs 1.0;
require {
type httpd_t;
type nfs_t;
class file { getattr open read };
}
allow httpd_t nfs_t:file { getattr open read };
EOF
checkmodule -M -m -o /etc/squirro/selinux/squirro-nginx-nfs.mod \
/etc/squirro/selinux/squirro-nginx-nfs.te
semodule_package -o /etc/squirro/selinux/squirro-nginx-nfs.pp \
-m /etc/squirro/selinux/squirro-nginx-nfs.mod
semodule -i /etc/squirro/selinux/squirro-nginx-nfs.pp
Troubleshooting#
Network Drop Between Servers#
This could be caused by a network monitoring tool closing all idle connections at periodic interval. In this cases, try lowering the TCP keep-alive used by the system and services:
Example, setting the value to 600s:
Change
/proc/sys/net/ipv4/tcp_keepalive_timevalue to 600:
# echo 600 > /proc/sys/net/ipv4/tcp_keepalive_time
Change “tcp-keepalive” value to 600 in
/etc/redis/redis.conf- Add a new file
/etc/sysctl.d/98-elasticsearch.confwith the following content:# lower keepalive settings to avoid elasticsearch cluster disconnects net.ipv4.tcp_keepalive_time = 600 net.ipv4.tcp_keepalive_intvl = 60 net.ipv4.tcp_keepalive_probes = 20