- Overview
- Requirements
- Deployment templates
- Manual: Preparing the installation
- Manual: Preparing the installation
- Step 2: Configuring the OCI-compliant registry for offline installations
- Step 3: Configuring the external objectstore
- Step 4: Configuring High Availability Add-on
- Step 5: Configuring SQL databases
- Step 6: Configuring the load balancer
- Step 7: Configuring the DNS
- Step 8: Configuring the disks
- Step 9: Configuring kernel and OS level settings
- Step 10: Configuring the node ports
- Step 11: Applying miscellaneous settings
- Step 12: Validating and installing the required RPM packages
- Step 13: Generating cluster_config.json
- Cluster_config.json Sample
- General configuration
- Profile configuration
- Certificate configuration
- Database configuration
- External Objectstore configuration
- Pre-signed URL configuration
- ArgoCD configuration
- Kerberos authentication configuration
- External OCI-compliant registry configuration
- Disaster recovery: Active/Passive and Active/Active configurations
- High Availability Add-on configuration
- Orchestrator-specific configuration
- Insights-specific configuration
- Process Mining-specific configuration
- Document Understanding-specific configuration
- Automation Suite Robots-specific configuration
- Monitoring configuration
- Optional: Configuring the proxy server
- Optional: Enabling resilience to zonal failures in a multi-node HA-ready production cluster
- Optional: Passing custom resolv.conf
- Optional: Increasing fault tolerance
- Adding a dedicated agent node with GPU support
- Adding a dedicated agent Node for Task Mining
- Connecting Task Mining application
- Adding a Dedicated Agent Node for Automation Suite Robots
- Step 15: Configuring the temporary Docker registry for offline installations
- Step 16: Validating the prerequisites for the installation
- Running uipathctl
- Manual: Performing the installation
- Post-installation
- Cluster administration
- Managing products
- Getting Started with the Cluster Administration portal
- Migrating Redis from in-cluster to external High Availability Add-on
- Migrating data between objectstores
- Migrating in-cluster objectstore to external objectstore
- Migrating from in-cluster registry to an external OCI-compliant registry
- Switching to the secondary cluster manually in an Active/Passive setup
- Disaster Recovery: Performing post-installation operations
- Converting an existing installation to multi-site setup
- Guidelines on upgrading an Active/Passive or Active/Active deployment
- Guidelines on backing up and restoring an Active/Passive or Active/Active deployment
- Scaling a single-node (evaluation) deployment to a multi-node (HA) deployment
- Monitoring and alerting
- Migration and upgrade
- Migrating standalone products to Automation Suite
- Step 1: Restoring the standalone product database
- Step 2: Updating the schema of the restored product database
- Step 3: Moving the Identity organization data from standalone to Automation Suite
- Step 4: Backing up the platform database in Automation Suite
- Step 5: Merging organizations in Automation Suite
- Step 6: Updating the migrated product connection strings
- Step 7: Migrating standalone Orchestrator
- Step 8: Migrating standalone Insights
- Step 9: Migrating standalone Test Manager
- Step 10: Deleting the default tenant
- Performing a single tenant migration
- Migrating between Automation Suite clusters
- Upgrading Automation Suite
- Downloading the installation packages and getting all the files on the first server node
- Retrieving the latest applied configuration from the cluster
- Updating the cluster configuration
- Configuring the OCI-compliant registry for offline installations
- Executing the upgrade
- Performing post-upgrade operations
- Product-specific configuration
- Orchestrator advanced configuration
- Configuring Orchestrator parameters
- Configuring appSettings
- Configuring the maximum request size
- Overriding cluster-level storage configuration
- Configuring NLog
- Saving robot logs to Elasticsearch
- Configuring credential stores
- Configuring encryption key per tenant
- Cleaning up the Orchestrator database
- Skipping host library installation
- Best practices and maintenance
- Troubleshooting
- How to troubleshoot services during installation
- How to uninstall the cluster
- How to clean up offline artifacts to improve disk space
- How to clear Redis data
- How to enable Istio logging
- How to manually clean up logs
- How to clean up old logs stored in the sf-logs bucket
- How to disable streaming logs for AI Center
- How to debug failed Automation Suite installations
- How to delete images from the old installer after upgrade
- How to disable TX checksum offloading
- How to manually set the ArgoCD log level to Info
- How to expand AI Center storage
- How to generate the encoded pull_secret_value for external registries
- How to address weak ciphers in TLS 1.2
- How to check the TLS version
- How to reduce permissions for an NFS backup directory
- How to work with certificates
- How to schedule Ceph backup and restore data
- How to clean up unused Docker images from registry pods
- How to collect DU usage data with in-cluster objectstore (Ceph)
- How to install RKE2 SELinux on air-gapped environments
- How to clean up old differential backups on an NFS server
- How to deploy Insights in a FIPS-enabled cluster
- How to migrate to cgroup v2
- How to push a local Docker image to the in-cluster registry
- How to exclude buckets from backup
- Unable to run an offline installation on RHEL 8.4 OS
- Error in downloading the bundle
- Offline installation fails because of missing binary
- Certificate issue in offline installation
- SQL connection string validation error
- Azure disk not marked as SSD
- Failure after certificate update
- TLS certificate validation errors
- Antivirus causes installation issues
- Automation Suite not working after OS upgrade
- Automation Suite requires backlog_wait_time to be set to 0
- Volume unable to mount due to not being ready for workloads
- Support bundle log collection failure
- Temporary registry installation fails on RHEL 8.9
- Frequent restart issue in uipath namespace deployments during offline installations
- DNS settings not honored by CoreDNS
- Unable to install temporary registry
- In-cluster registry seeding fails due to insufficient memory
- Prerequisite checks fail when Document Understanding modern projects is enabled and AI Center is disabled
- Data loss when reinstalling or upgrading Insights following Automation Suite upgrade
- Unable to access Automation Hub following upgrade to Automation Suite 2024.10.0
- Upgrade failure during posthook import
- Single-node upgrade fails at the fabric stage
- Upgrade fails due to unhealthy Ceph
- RKE2 not getting started due to space issue
- Volume unable to mount and remains in attach/detach loop state
- Upgrade fails due to classic objects in the Orchestrator database
- Ceph cluster found in a degraded state after side-by-side upgrade
- Unhealthy Insights component causes the migration to fail
- Service upgrade fails for Apps
- In-place upgrade timeouts
- Docker registry migration stuck in PVC deletion stage
- AI Center provisioning failure after upgrading to 2023.10 or later
- Upgrade fails in offline environments
- SQL validation fails during upgrade
- snapshot-controller-crds pod in CrashLoopBackOff state after upgrade
- Upgrade fails due to overridden Insights PVC sizes
- Failure to upgrade to Automation Suite 2024.10.1
- Upgrade fails due to Velero migration issue
- Upgrade stuck on rook-ceph application deletion
- Setting a timeout interval for the management portals
- Authentication not working after migration
- Kinit: Cannot find KDC for realm <AD Domain> while getting initial credentials
- Kinit: Keytab contains no suitable keys for *** while getting initial credentials
- GSSAPI operation failed due to invalid status code
- Alarm received for failed Kerberos-tgt-update job
- SSPI provider: Server not found in Kerberos database
- Login failed for AD user due to disabled account
- ArgoCD login failed
- Update the underlying directory connections
- Robot cannot connect to an Automation Suite Orchestrator instance
- Partial failure to restore backup in Automation Suite 2024.10.0
- Failure to get the sandbox image
- Pods not showing in ArgoCD UI
- Accessing FQDN returns RBAC access denied error
- Redis probe failure
- RKE2 server fails to start
- Secret not found in UiPath namespace
- ArgoCD goes into progressing state after first installation
- ArgoCD repo-server pod in CrashLoopBackOff
- Manual ArgoCD NetworkPolicy mitigation (GHSA-47m3-95c7-g2g8)
- Pods stuck in Init:0/X
- Missing Ceph-rook metrics from monitoring dashboards
- Mismatch in reported errors during diagnostic health checks
- No healthy upstream issue
- Log streaming does not work in proxy setups
- Failure to add agent nodes in offline environments
- Node becomes unresponsive (OOM) during large Document Understanding bundle upload
- Backup operations fail with PartiallyFailed status
- Running High Availability with Process Mining
- Process Mining ingestion failed when logged in using Kerberos
- After Disaster Recovery Dapr is not working properly for Process Mining
- Unable to connect to AutomationSuite_ProcessMining_Warehouse database using a pyodbc format connection string
- Airflow installation fails with sqlalchemy.exc.ArgumentError: Could not parse rfc1738 URL from string ''
- How to add an IP table rule to use SQL Server port 1433
- Automation Suite certificate is not trusted from the server where CData Sync is running
- Running the diagnostics tool
- Using the Automation Suite support bundle
- Exploring Logs
Reference for required node port and firewall configuration for online and offline Automation Suite installations.
Changes to IP tables are not recommended or supported.
Make sure to enable the following ports on your firewall for each source.
The following ports require inbound or outbound network access between cluster nodes, external clients, or the load balancer:
| Port | Protocol | Source | Purpose | Requirements |
|---|---|---|---|---|
22 | TCP | Jump Server / client machine | For SSH (installation, cluster management debugging) | Do not open this port to the internet. Allow access to the client machine or jump server. |
443 | TCP | All nodes in a cluster and the load balancer | For HTTPS (accessing Automation Suite) | This port must have inbound and outbound connectivity from all the nodes in the cluster and the load balancer. |
2334 | TCP | All nodes in a cluster | Transient HTTP file server (hostNetwork: true, 0.0.0.0) used by uipathctl support-bundle to collect logs and metrics. Open only during support-bundle, historical-logs, or service-metrics collection. | Do not open this port to the internet. Allow node-to-node connectivity inside the cluster CIDR. |
2379 | TCP | All nodes in a cluster | etcd client port | Do not open this port to the internet. Access between nodes must be ensured over a private IP address. |
2380 | TCP | All nodes in a cluster | etcd peer port | Do not open this port to the internet. Access between nodes must be ensured over a private IP address. |
4240 | TCP | Peer Cilium agent pods on other cluster nodes | Cilium cluster-wide health probe for verifying L3 reachability between nodes; binds to the node IP, not 0.0.0.0. Blocking inter-node connectivity causes false-positive node-down alerts. | Do not open this port to the internet. Allow node-to-node connectivity inside the cluster CIDR. |
5001 | TCP | RKE2 internal (kubelet/containerd image pull) | RKE2 embedded internal image registry (separate from offline-registry port 30070), hosting pre-loaded images pulled during cluster startup. Present only on the bootstrap server node. | Binds to 0.0.0.0:5001 on the bootstrap server; allow intra-cluster access from peer nodes. Do not open this port to the internet. |
6443 | TCP | All nodes in a cluster | For accessing Kube API using HTTPS, and required for node joining. On worker nodes, rke2 agent also binds 127.0.0.1:6443 as a local load-balancer for apiserver access. | This port requires inbound and outbound connectivity from all nodes. Do not block 6443 on the loopback interface on worker nodes; kubelet and kube-proxy will lose apiserver connectivity. |
8472 | UDP | All nodes in a cluster | Required for Cilium. | Do not open this port to the internet. Access between nodes must be ensured over a private IP address. |
9090 | TCP | All nodes in the cluster | Used by Cilium for monitoring and handling pod crashes | This port must have inbound and outbound connectivity from all the nodes in the cluster. |
9100 | TCP | Cluster Prometheus pods | Prometheus node_exporter (monitoring-prometheus-node-exporter DaemonSet): host-level metrics (CPU, memory, filesystem, network, process counts). Binds to 0.0.0.0:9100. | Do not open this port to the internet; exposes sensitive host telemetry. Allow intra-cluster scrape from Prometheus pods. |
9345 | TCP | All nodes in a cluster and the load balancer | For accessing Kube API using HTTPS, required for node joining | This port must have inbound and outbound connectivity from all nodes in the cluster and the load balancer. |
9963 | TCP | Cluster Prometheus pods | cilium-operator Prometheus metrics endpoint: operator-level metrics (leader-election, identity GC, CNP processing) | Binds to 0.0.0.0:9963; do not expose to the internet. Allow intra-cluster scrape from Prometheus pods. |
10250 | TCP | All nodes in a cluster | kubelet / metrics server | Do not open this port to the internet. Access between nodes must be ensured over a private IP address. |
30000–32767 | TCP | All nodes in a cluster | NodePort range for Istio Ingress Gateway; required for routing external traffic into the cluster via Kubernetes NodePort services | Allow inbound connectivity from the load balancer to all cluster nodes. Do not expose this range to the internet; route traffic only through the load balancer. |
30071 | TCP | All nodes in a cluster | NodePort port for internal communication between nodes in a cluster | Do not open this port to the internet. Access between nodes must be ensured over a private IP address. |
The following ports bind to 127.0.0.1 (loopback) only and are not reachable across the network. Do not block them on the loopback interface:
| Port | Protocol | Source | Purpose | Requirements |
|---|---|---|---|---|
2381 | TCP | Local node only (loopback) | etcd metrics endpoint used for etcd readiness probes | Binds to 127.0.0.1 only; no cross-cluster scrape is configured by default. Do not block on the loopback interface. |
2382 | TCP | Local node only (loopback) | etcd loopback peer endpoint used by RKE2's internal etcd controller for member-list and cluster-management operations | Binds to 127.0.0.1 only. Do not block on the loopback interface. |
6444 | TCP | Local node only (loopback) | On worker nodes: RKE2 agent supervisor load-balancer; routes local agent components (kubelet, kube-proxy, helm-controller) to healthy supervisor servers on port 9345 | Binds to 127.0.0.1 only. Do not block on the loopback interface; kubelet and kube-proxy will lose control plane connectivity and the node will go NotReady. |
9234 | TCP | Local node only (loopback) | Cilium operator REST API (/healthz); target of the cilium-operator pod liveness probe | Binds to 127.0.0.1 only. Do not block on the loopback interface; the liveness probe will fail, stopping CRD reconciliation, identity GC, and BGP routing. |
9879 | TCP | Local node only (loopback) | cilium-agent /healthz endpoint; target of the cilium-agent liveness, readiness, and startup probes | Binds to 127.0.0.1 only. Do not block on the loopback interface; kubelet probes will fail and cilium-agent will be restarted, causing CNI disruption cluster-wide. |
10256 | TCP | Local node only (loopback) | kube-proxy /livez healthz endpoint; target of the kube-proxy static-pod liveness probe | Binds to 127.0.0.1 only. Do not block on the loopback interface; the liveness probe will fail and kubelet will restart kube-proxy. |
10257 | TCP | Local node only (loopback) | kube-controller-manager secure-port endpoint (/healthz, /metrics, leader-election). Control-plane nodes only. | Binds to 127.0.0.1 only. |
10259 | TCP | Local node only (loopback) | kube-scheduler secure-port endpoint (/healthz, /metrics, leader-election). Control-plane nodes only. | Binds to 127.0.0.1 only. |
The following additional ports are required in offline installations:
| Port | Protocol | Source | Purpose | Requirements |
|---|---|---|---|---|
80 | TCP | All nodes in the cluster | Required for sending system email notifications | Do not open this port to the internet. Access between nodes and the SMTP server must be ensured over a private IP address. |
587 | TCP | All nodes in the cluster | Required for sending system email notifications | Do not open this port to the internet. Access between nodes and the SMTP server must be ensured over a private IP address. |
300701 | TCP | The machine on which you plan to trigger the installation or upgrade. | For accessing the temporary registry during installation and upgrade using HTTP. | Traffic on this port must be forwarded to the Temporary Registry Pool. |
1 If an external registry is not available in the offline installation, open port 30070 on the machine on which you plan to trigger the installation or upgrade.
Exposing port 6443 outside the cluster boundary is mandatory if there is a direct connection to the Kerberos API.
Port 9345 is used by nodes to discover existing nodes and join the cluster in the multi-node deployment. To keep the high availability discovery mechanisms running, we recommend exposing it via the load balancer with health check.
Additionally, make sure you have connectivity from all nodes to the SQL server. Do not expose the SQL server on one of the Istio reserved ports, as it may lead to connection failures.
If you have a firewall set up in the network, make sure that it has these ports open and allows traffic according to the aforementioned requirements.