- Overview
- Requirements
- Deployment templates
- Manual: Preparing the installation
- Manual: Preparing the installation
- Step 2: Configuring the OCI-compliant registry for offline installations
- Step 3: Configuring the external objectstore
- Step 4: Configuring High Availability Add-on
- Step 5: Configuring SQL databases
- Step 7: Configuring the DNS
- Step 8: Configuring the disks
- Step 9: Configuring kernel and OS level settings
- Step 10: Configuring the node ports
- Step 11: Applying miscellaneous settings
- Step 12: Validating and installing the required RPM packages
- Step 13: Generating cluster_config.json
- Cluster_config.json Sample
- General configuration
- Profile configuration
- Certificate configuration
- Database configuration
- External Objectstore configuration
- Pre-signed URL configuration
- ArgoCD configuration
- Kerberos authentication configuration
- External OCI-compliant registry configuration
- Disaster recovery: Active/Passive and Active/Active configurations
- High Availability Add-on configuration
- Orchestrator-specific configuration
- Insights-specific configuration
- Process Mining-specific configuration
- Document Understanding-specific configuration
- Automation Suite Robots-specific configuration
- Monitoring configuration
- Optional: Configuring the proxy server
- Optional: Enabling resilience to zonal failures in a multi-node HA-ready production cluster
- Optional: Passing custom resolv.conf
- Optional: Increasing fault tolerance
- Adding a dedicated agent node with GPU support
- Adding a Dedicated Agent Node for Automation Suite Robots
- Step 15: Configuring the temporary Docker registry for offline installations
- Step 16: Validating the prerequisites for the installation
- Running uipathctl
- Manual: Performing the installation
- Post-installation
- Cluster administration
- Managing products
- Getting Started with the Cluster Administration portal
- Adding a new node to the cluster
- Removing a node from the cluster
- Repaving a cluster node
- Starting and shutting down a node
- Renaming a node
- Migrating Redis from in-cluster to external High Availability Add-on
- Migrating data between objectstores
- Migrating in-cluster objectstore to external objectstore
- Migrating from in-cluster registry to an external OCI-compliant registry
- Switching to the secondary cluster manually in an Active/Passive setup
- Disaster Recovery: Performing post-installation operations
- Converting an existing installation to multi-site setup
- Guidelines on upgrading an Active/Passive or Active/Active deployment
- Guidelines on backing up and restoring an Active/Passive or Active/Active deployment
- Scaling a single-node (evaluation) deployment to a multi-node (HA) deployment
- Monitoring and alerting
- Migration and upgrade
- Migrating between Automation Suite clusters
- Upgrading Automation Suite
- Downloading the installation packages and getting all the files on the first server node
- Retrieving the latest applied configuration from the cluster
- Updating the cluster configuration
- Configuring the OCI-compliant registry for offline installations
- Executing the upgrade
- Performing post-upgrade operations
- Product-specific configuration
- Orchestrator advanced configuration
- Configuring Orchestrator parameters
- Configuring appSettings
- Configuring the maximum request size
- Overriding cluster-level storage configuration
- Configuring NLog
- Saving robot logs to Elasticsearch
- Configuring credential stores
- Configuring encryption key per tenant
- Cleaning up the Orchestrator database
- Skipping host library installation
- Best practices and maintenance
- Troubleshooting
- How to troubleshoot services during installation
- How to reduce permissions for an NFS backup directory
- How to uninstall the cluster
- How to clean up offline artifacts to improve disk space
- How to clear Redis data
- How to enable Istio logging
- How to manually clean up logs
- Moving Ceph out of read-only mode
- How to clean up old logs stored in the sf-logs bucket
- How to disable streaming logs for AI Center
- How to debug failed Automation Suite installations
- How to delete images from the old installer after upgrade
- How to disable TX checksum offloading
- How to manually set the ArgoCD log level to Info
- How to expand AI Center storage
- How to generate the encoded pull_secret_value for external registries
- How to address weak ciphers in TLS 1.2
- How to check the TLS version
- How to work with certificates
- How to schedule Ceph backup and restore data
- How to collect DU usage data with in-cluster objectstore (Ceph)
- How to install RKE2 SELinux on air-gapped environments
- How to clean up old differential backups on an NFS server
- How to deploy Insights in a FIPS-enabled cluster
- How to migrate to cgroup v2
- How to recover Kerberos authentication after a VM restart
- How to push a local Docker image to the in-cluster registry
- How to exclude buckets from backup
- Error in downloading the bundle
- Offline installation fails because of missing binary
- Azure disk not marked as SSD
- Failure after certificate update
- TLS certificate validation errors
- Antivirus causes installation issues
- Automation Suite not working after OS upgrade
- Automation Suite requires backlog_wait_time to be set to 0
- Temporary registry installation fails on RHEL 8.9
- Frequent restart issue in uipath namespace deployments during offline installations
- DNS settings not honored by CoreDNS
- In-cluster registry seeding fails due to insufficient memory
- Prerequisite checks fail when Document Understanding modern projects is enabled and AI Center is disabled
- Upgrade fails due to unhealthy Ceph
- Upgrade fails due to classic objects in the Orchestrator database
- Ceph cluster found in a degraded state after side-by-side upgrade
- Service upgrade fails for Apps
- In-place upgrade timeouts
- Upgrade fails in offline environments
- snapshot-controller-crds pod in CrashLoopBackOff state after upgrade
- Upgrade fails due to overridden Insights PVC sizes
- Upgrade failure due to uppercase hostname
- Setting a timeout interval for the management portals
- Authentication not working after migration
- Kinit: Cannot find KDC for realm <AD Domain> while getting initial credentials
- Kinit: Keytab contains no suitable keys for *** while getting initial credentials
- GSSAPI operation failed due to invalid status code
- Alarm received for failed Kerberos-tgt-update job
- SSPI provider: Server not found in Kerberos database
- Login failed for AD user due to disabled account
- ArgoCD login failed
- Failure to get the sandbox image
- Pods not showing in ArgoCD UI
- Redis probe failure
- RKE2 server fails to start
- ArgoCD goes into progressing state after first installation
- ArgoCD repo-server pod in CrashLoopBackOff
- Manual ArgoCD NetworkPolicy mitigation (GHSA-47m3-95c7-g2g8)
- Missing Ceph-rook metrics from monitoring dashboards
- Mismatch in reported errors during diagnostic health checks
- Configuring resource requests and limits for uipathctl-created workloads
- No healthy upstream issue
- Redis startup blocked by antivirus
- AI Center and Document Understanding pods fail to start with TLS certificate verification enabled
- Fluentd does not export logs in IPv6 environments
- Studio Desktop cannot load Integration Service connectors and activities
- Running High Availability with Process Mining
- Process Mining ingestion failed when logged in using Kerberos
- Unable to connect to AutomationSuite_ProcessMining_Warehouse database using a pyodbc format connection string
- Airflow installation fails with sqlalchemy.exc.ArgumentError: Could not parse rfc1738 URL from string ''
- How to add an IP table rule to use SQL Server port 1433
- Automation Suite certificate is not trusted from the server where CData Sync is running
- Process Mining fails to load after disabling and re-enabling it
- Running the diagnostics tool
- Using the Automation Suite support bundle
- Exploring Logs
Start and shut down nodes safely in Automation Suite, covering manual and automatic startup and shutdown behavior.
This page explains the manual and automatic startup and shutdown behavior of Automation Suite.
You must always proceed by shutting down one node, performing the required operation, waiting until the node is healthy, and then taking down the other node to perform the same operation.
The following table describes different scenarios you may experience when shutting down cluster services or nodes. The table provides detailed actions you must take for each situation, alongside guidance on understanding the expected behavior in response to these actions.
| Scenario | Action | Expected behavior |
|---|---|---|
| Shutting down cluster services on one node without turning off the node, for maintenance or any other reason. |
| In an HA scenario, most services will remain up. The node should start up without any issue and any down services should restart. |
| Shutting down all cluster services without turning off nodes, for maintenance or any other reason. |
| Services will become unavailable. Nodes should startup without issue. |
| Shutting down all nodes. | If your hypervisor management portal (such as VMware, AWS) allows for services to graceful shutdown without force terminating the machine, carry out a normal shutdown. By default, systemd subsystem allows a grace period for services to shutdown before they are forcefully terminated. However, if your system overwrites configured shutdown times, it may interfere with a graceful shutdown. For example, on AWS, the platform can force terminate a VM after two minutes. As such, the services must be shut down manually as a node drain can take up to 5 minutes (this is a requirement of a graceful shutdown). | If the shutdown is graceful, the nodes should start up without issue. If the cluster was shut down for more than 6 hours and Kerberos authentication is configured, you may need to renew Kerberos tickets after restart. For details, refer to Recovering Kerberos Authentication After VM Restart. |
| Shutting down an individual node. | If your hypervisor management portal (such as VMware, AWS) allows for services to graceful shutdown without force terminating the machine, carry out a normal shutdown. By default, systemd subsystem allows a grace period for services to shutdown before they are forcefully terminated. However, if your system overwrites configured shutdown times, it may interfere with a graceful shutdown.For example, on AWS, the platform can force terminate a VM after two minutes. As such, the services must be shut down manually as a node drain can take up to 5 minutes (this is a requirement of a graceful shutdown). | If the shutdown process is not forceful, the node should reboot without any issues. |
| Forcefully terminating a server node. | Not applicable. | In most cases the node will start up, but there may be problems with some services that use persistent data. Although these issues are typically recoverable, setting up backups is strongly recommended. The insights pod will not restart until the original node is back online, in order to prevent potential data loss. If the node is not recoverable, contact the support team. |
Shutdown behavior
During shutdown, systemd stops the services in the order they were started. Since the node-drain service has the After=rke2-server.service or After=rke2-agent.service directive, it executes its shutdown sequence before the rke2-service shutdown. This means that, in a properly configured system, simply gracefully shutting down the node is a safe operation.
Manual restart
If you plan to stop the rke2 service and reboot the machine, take the following steps:
-
To ensure that the cluster is healthy while performing node maintenance activity, you must drain the workloads running on that node to other nodes. To drain the node, run the following command:
systemctl stop node-drain.servicesystemctl stop node-drain.service -
Stop the Kubernetes process on the node, depending on the node type:
- On a server node:
systemctl stop rke2-serversystemctl stop rke2-server - On an agent node:
systemctl stop rke2-agentsystemctl stop rke2-agent
- On a server node:
-
Terminate the rke2 services and containerd and all child processes:
rke2-killall.shrke2-killall.sh
To download rke2-killall.sh script, refer to Installation packages download links.
Startup behavior
The rke2-service starts and is followed by node-drainer and node-uncordon. node-drainer does not do any action at startup, just returns confirmation that the service is up.
The node-uncordon only runs once and starts /opt/node-drain.sh nodestart, which uncordons the node. As part of the drain procedure that occurs at shutdown, this cordons the node, making it unschedulable. This state persists when the rke2 service starts. As such, the node must be uncordoned after rke2-service restarts.
Manual startup
The service starts automatically with Automation Suite. However, if rke2-service was manually stopped, you must start the service again by running the following commands:
-
Start the Kubernetes process on the node, depending on the node type:
- On a server node:
systemctl start rke2-serversystemctl start rke2-server - On an agent node:
systemctl start rke2-agentsystemctl start rke2-agent
- On a server node:
-
Once the
rke2service is started, uncordon the node to ensure Kubernetes can now schedule workloads on this node:systemctl restart node-uncordonsystemctl restart node-uncordon -
Once the node is started, you must drain the node:
systemctl start node-drain.servicesystemctl start node-drain.serviceImportant:Skipping this step could cause the Kubelet service to shut down in an unhealthy way if the system is restarted.
Patching cluster nodes
When patching or restarting server nodes, the order in which you apply changes directly affects cluster stability.
Pre-patch checks
Before touching any node, confirm the cluster is healthy:
-
Confirm all three etcd members are healthy:
ETCD_CONTAINER="$(/var/lib/rancher/rke2/bin/crictl ps --name etcd --state Running -q | head -n1)" /var/lib/rancher/rke2/bin/crictl exec "$ETCD_CONTAINER" etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/rke2/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/rke2/server/tls/etcd/server-client.key \ endpoint health --clusterETCD_CONTAINER="$(/var/lib/rancher/rke2/bin/crictl ps --name etcd --state Running -q | head -n1)" /var/lib/rancher/rke2/bin/crictl exec "$ETCD_CONTAINER" etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/rke2/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/rke2/server/tls/etcd/server-client.key \ endpoint health --cluster -
Check for crash-loops on any node:
journalctl -u rke2-server --since "2 hours ago" | grep -i "FAILURE\|left-over\|unclean"journalctl -u rke2-server --since "2 hours ago" | grep -i "FAILURE\|left-over\|unclean" -
Verify a recent etcd snapshot exists:
ls -lth /var/lib/rancher/rke2/server/db/snapshots/ | head -5ls -lth /var/lib/rancher/rke2/server/db/snapshots/ | head -5
Do not start patching if any node is already crash-looping or if etcd health checks fail. Resolve the instability before proceeding.
Identify the bootstrap node
The bootstrap node is typically the current etcd leader. To identify it, run the following command and look for IS LEADER: true:
ETCD_CONTAINER="$(/var/lib/rancher/rke2/bin/crictl ps --name etcd --state Running -q | head -n1)"
/var/lib/rancher/rke2/bin/crictl exec "$ETCD_CONTAINER" etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \
--cert=/var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
--key=/var/lib/rancher/rke2/server/tls/etcd/server-client.key \
endpoint status --cluster
ETCD_CONTAINER="$(/var/lib/rancher/rke2/bin/crictl ps --name etcd --state Running -q | head -n1)"
/var/lib/rancher/rke2/bin/crictl exec "$ETCD_CONTAINER" etcdctl \
--endpoints=https://127.0.0.1:2379 \
--cacert=/var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \
--cert=/var/lib/rancher/rke2/server/tls/etcd/server-client.crt \
--key=/var/lib/rancher/rke2/server/tls/etcd/server-client.key \
endpoint status --cluster
Patch order
Always patch the bootstrap node first, then secondary nodes one at a time.
Patching the bootstrap node first ensures that when it goes down, the two secondary nodes maintain quorum and elect a new leader. The bootstrap node then rejoins as a follower in a clean state. Patching secondary nodes first and leaving the bootstrap node for last can cause wider instability if the bootstrap node fails to rejoin cleanly.
After each node patches and restarts, confirm it is fully back before proceeding to the next node:
-
Verify the node status is
Ready:kubectl get nodeskubectl get nodes -
Confirm all three etcd members are healthy:
ETCD_CONTAINER="$(/var/lib/rancher/rke2/bin/crictl ps --name etcd --state Running -q | head -n1)" /var/lib/rancher/rke2/bin/crictl exec "$ETCD_CONTAINER" etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/rke2/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/rke2/server/tls/etcd/server-client.key \ endpoint health --clusterETCD_CONTAINER="$(/var/lib/rancher/rke2/bin/crictl ps --name etcd --state Running -q | head -n1)" /var/lib/rancher/rke2/bin/crictl exec "$ETCD_CONTAINER" etcdctl \ --endpoints=https://127.0.0.1:2379 \ --cacert=/var/lib/rancher/rke2/server/tls/etcd/server-ca.crt \ --cert=/var/lib/rancher/rke2/server/tls/etcd/server-client.crt \ --key=/var/lib/rancher/rke2/server/tls/etcd/server-client.key \ endpoint health --cluster -
Check for errors in the RKE2 server logs:
journalctl -u rke2-server -n 30journalctl -u rke2-server -n 30
Kernel upgrades
When performing a kernel upgrade, wipe the containerd snapshot store after running rke2-killall.sh and before applying the kernel upgrade. This prevents containerd state corruption on any node, regardless of patch order.
Files created during installation
The following unit files are created during installation:
rke2-server.service(server only) - Starts therke2-server, which starts the server node.rke2-agent.service(agent only) - Starts therke2-agent, which starts the agent node.node-drain.service- Used at shutdown time. Executed before shutting downrke2-agentorrke2-serverand performs a drain. Has a timeout of 300 seconds.node-uncordon.service- Used at startup to uncordon a node.var-lib-kubelet.mount- Autogenerated by fstab generator.var-lib-rancher-rke2-server-db.mount- Autogenerated by fstab generator.var-lib-rancher.mount- Autogenerated by fstab generator.
There are no strong dependencies between the unit files. However, node-drain and node-uncordon have the After=rke2-server.service or After=rke2-agent.service directive. This means that those services will start after the rke2-server.service.