- 概述
- 要求
- 部署模板
- 手动:准备安装
- 手动:准备安装
- 步骤 2:为离线安装配置符合 OCI 的注册表
- 步骤 3:配置外部对象存储
- 步骤 4:配置 High Availability Add-on
- 步骤 5:配置 SQL 数据库
- 步骤 7:配置 DNS
- 步骤 8:配置磁盘
- 步骤 9:配置内核和操作系统级别设置
- 步骤 10:配置节点端口
- 步骤 11:应用其他设置
- 步骤 12:验证并安装所需的 RPM 包
- 步骤 13:生成 cluster_config.json
- Cluster_config.json 示例
- 常规配置
- 配置文件配置
- 证书配置
- 数据库配置
- 外部对象存储配置
- 预签名 URL 配置
- ArgoCD 配置
- Kerberos 身份验证配置
- 符合 OCI 的外部注册表配置
- Disaster Recovery:主动/被动和主动/主动配置
- High Availability Add-on 配置
- 特定于 Orchestrator 的配置
- Insights 特定配置
- Process Mining 特定配置
- Document Understanding 特定配置
- Automation Suite Robot 特定配置
- 监控配置
- 可选:配置代理服务器
- 可选:在多节点 HA 就绪生产集群中启用区域故障恢复
- 可选:传递自定义 resolv.conf
- 可选:提高容错能力
- 添加具有 GPU 支持的专用代理节点
- 为 Automation Suite Robot 添加专用代理节点
- 步骤 15:为离线安装配置临时 Docker 注册表
- 步骤 16:验证安装的先决条件
- 正在运行 uipathctl
- 手动:执行安装
- 安装后
- 集群管理
- 监控和警示
- 迁移和升级
- 特定于产品的配置
- 最佳实践和维护
- 故障排除
- 如何在安装过程中对服务进行故障排除
- 如何减少 NFS 备份目录的权限
- 如何卸载集群
- 如何清理离线工件以改善磁盘空间
- 如何清除 Redis 数据
- 如何启用 Istio 日志记录
- 如何手动清理日志
- 将 Ceph 退出只读模式
- 如何清理存储在 sf-logs 存储桶中的旧日志
- 如何禁用 AI Center 的流日志
- 如何对失败的 Automation Suite 安装进行调试
- 如何在升级后从旧安装程序中删除映像
- 如何禁用 TX 校验和卸载
- 如何手动将 ArgoCD 日志级别设置为 Info
- 如何扩展 AI Center 存储
- 如何为外部注册表生成已编码的 pull_secret_value
- 如何解决 TLS 1.2 中的弱密码问题
- 如何查看 TLS 版本
- 如何使用证书
- 如何计划 Ceph 备份和还原数据
- 如何使用集群内对象存储 (Ceph) 收集 DU 使用情况数据
- 如何在离线环境中安装 RKE2 SELinux
- 如何清理 NFS 服务器上的旧差异备份
- 如何在已启用 FIPS 的集群中部署 Insights
- 如何迁移到 cgroup v2
- 如何在虚拟机重新启动后恢复 Kerberos 身份验证
- 如何将本地 Docker 映像推送到集群内注册表
- 如何从备份中排除存储桶
- 无法获取沙盒映像
- Pod 未显示在 ArgoCD 用户界面中
- Redis 探测器失败
- RKE2 服务器无法启动
- ArgoCD 在首次安装后进入“进行中”状态
- 处于 CrashLoopBackOff 状态的 ArgoCD 存储库服务器 Pod
- 手动 ArgoCD 网络策略缓解措施 (MHSA-47m3-95c7-g2g8)
- 监控仪表板中缺少 Ceph-rook 指标
- 诊断性运行状况检查期间报告的错误不匹配
- 为 uipathctl 创建的工作负载配置资源请求和限制
- 无正常的上游问题
- 杀毒软件阻止了 Redis 启动
- 无法在启用 TLS 证书验证的情况下启动 AI Center 和 Document Understanding Pod
- Fluentd 不会在 IPv6 环境中导出日志
- Studio 桌面版无法加载 Integration Service 连接器和活动
- 使用 Process Mining 运行高可用性
- 使用 Kerberos 登录时 Process Mining 挖掘失败
- 无法使用 pyodbc 格式连接字符串连接到 AutomationSuite_ProcessMining_Warehouse 数据库
- Airflow 安装失败,并显示 sqlalchemy.exc.ArgumentError:无法从字符串“”中解析 rfc1738 URL
- 如何添加 IP 表格规则以使用 SQL Server 端口 1433
- 运行 CData Sync 的服务器不信任 Automation Suite 证书
- Process Mining fails to load after disabling and re-enabling it
- 运行诊断工具
- 使用 Automation Suite 支持捆绑包
- 探索日志
解决 Automation Suite 中与 Alertmanager 配置不一致相关的管理警示。
alertmanager.rules
AlertmanagerConfigInconsistent
当同一集群中的 Alertmanager 实例具有不同配置时,将触发此警示。 这可能表明配置存在问题,该问题在 Alertmanager 的所有实例中不一致。
要解决此问题,请执行以下步骤:
- 在已部署的所有
alertmanager.yml之间运行diff工具以识别问题。 - 删除不正确的密码并部署正确的密码。
如果问题仍然存在,请联系 UiPath™ 支持团队。
AlertmanagerFailedReload
警示管理器加载或重新加载配置失败。请检查任何自定义警示管理器配置中是否存在输入错误,否则请联系 UiPath™ 支持团队并提供支持包。有关详细信息,请参阅使用 Automation Suite 支持捆绑包。
AlertmanagerMembersInconsistent
这些是具有多个警示管理器副本的 HA 集群的内部警示管理器错误。警示可能会间歇性地出现和消失。暂时缩小规模,然后扩大警示管理器副本可能会解决此问题。
要解决此问题,请执行以下步骤:
-
缩放至零。请注意,Pod 需要一段时间才能关闭:
statefulset.apps/alertmanager-monitoring-kube-prometheus-alertmanager scaledstatefulset.apps/alertmanager-monitoring-kube-prometheus-alertmanager scaled -
缩小到 2:
kubectl scale statefulset -n monitoring alertmanager-monitoring-kube-prometheus-alertmanager --replicas=2kubectl scale statefulset -n monitoring alertmanager-monitoring-kube-prometheus-alertmanager --replicas=2 -
检查 Alertmanager Pod 是否已启动以及是否处于正在运行状态:
kubectl get po -n monitoringkubectl get po -n monitoring
如果问题仍然存在,请联系 UiPath™ 支持团队。
常规.规则
TargetDown
Prometheus 无法从警示中的目标收集指标,这意味着 Grafana 仪表板和基于该目标的指标的进一步警示不可用。检查与该目标相关的其他警示。
Watchdog
这是一个警示,用于确保整个警示管道正常运行。此警示始终处于触发状态。因此,它应始终在“警示管理器”中针对接收器触发。有各种通知机制的集成,可在此警示未触发时通知您。例如,PagerDuty 中的 DeadMansSnitch 集成。
prometheus-operator
PrometheusOperatorListErrors, PrometheusOperatorWatchErrors, PrometheusOperatorSyncFailed, PrometheusOperatorReconcileErrors, PrometheusOperatorNodeLookupErrors, PrometheusOperatorNotReady, PrometheusOperatorRejectedResources
控制 Prometheus 资源的 Prometheus 运算符的内部错误。存在这些错误时,Prometheus 本身可能仍然运行良好;但是,此错误表示监控可配置性下降。请联系 UiPath™ 支持团队。
Prometheus
PrometheusBadConfig
Prometheus 加载或重新加载配置失败。请检查任何自定义 Prometheus 配置是否存在输入错误。否则,请联系 UiPath™ 支持团队。
PrometheusErrorSendingAlertsToSomeAlertmanagers, PrometheusErrorSendingAlertsToAnyAlertmanager, PrometheusNotConnectedToAlertmanagers
从 Prometheus 到警示管理器的连接不正常。指标仍可查询,并且 Grafana 仪表板可能仍会显示指标,但不会触发警示。检查警示管理器的任何自定义配置是否存在输入错误,否则请联系 UiPath™ 支持团队。
PrometheusNotificationQueueRunningFull, PrometheusTSDBReloadsFailing, PrometheusTSDBCompactionsFailing, PrometheusNotIngestingSamples, PrometheusDuplicateTimestamps, PrometheusOutOfOrderTimestamps, PrometheusRemoteStorageFailures, PrometheusRemoteWriteBehind, PrometheusRemoteWriteDesiredShards
表示可能无法按预期收集指标的内部 Prometheus 错误。请联系 UiPath™ 支持团队。
PrometheusRuleFailures
如果存在基于不存在的指标或不正确的 PromQL 语法的格式错误警示,则可能会发生这种情况。如果未添加自定义警示,请联系 UiPath™ 支持团队。
PrometheusMissingRuleEvaluations
Prometheus 无法评估是否应触发警示。如果警示太多,可能会发生这种情况。请删除昂贵的自定义警示评估和/或查看有关增加 Prometheus CPU 限制的文档。如果未添加自定义警示,请联系 UiPath™ 支持团队。
PrometheusTargetLimitHit
Prometheus 要收集的目标过多。如果添加了额外的 ServiceMonitor(请参阅监控控制台),您可以将其删除。
uipath.prometheus.resource.provisioning.alerts
PrometheusMemoryUsage, PrometheusStorageUsage
当集群接近配置的内存和存储限制时,这些警示会发出警告。 这可能发生在最近使用量大幅增加的集群上(通常来自机器人而不是用户),或者在未调整 Prometheus 资源的情况下将节点添加到集群中时。 这是因为要收集的指标数量增加。
这也可能是由于触发了大量警示所致,因此请务必检查触发大量警示的原因。
如果此问题仍然存在,请联系 UiPath™ 支持团队并提供生成的支持捆绑包。
uipath.availability.alerts
UiPathAvailabilityHighTrafficUserFacing
来自 UiPath™ 服务的 HTTP 500 响应数量超过给定阈值。下表描述了用于评估错误率的流量级别阈值。
| 流量级别 | 20 分钟内的请求数 | 错误阈值(适用于 HTTP 500) |
|---|---|---|
| 高 | >100,000 | 0.1% |
| 中 | 10,000 到 100,000 之间 | 1% |
| 低 | < 10,000 | 5% |
面向用户的服务中的错误可能会导致可在 Automation Suite 用户界面中直接观察到的功能降级,而后端服务中的错误则不会产生明显的后果。
警示会指明哪个服务的错误率较高。要了解报告服务所依赖的其他服务可能存在哪些级联问题,您可以使用 Istio 工作负载仪表板,该仪表板会显示服务之间的错误。
请仔细检查所有最近重新配置的 Automation Suite 产品。还可以使用 kubectl logs 命令获取详细日志。如果错误仍然存在,请联系 UiPath™ 支持团队。
备份
NFSServerDisconnected
此警示表示 NFS 服务器连接已丢失。
您需要检查 NFS 服务器连接和装载路径。
VolumeBackupFailed
此警示表示 PVC 的备份失败。
要解决此问题,请执行以下步骤:
-
检查 PVC 的状态,以确保对于持久卷 (PV),它是
Bound。kubectl get pvc --namespace <namespace>kubectl get pvc --namespace <namespace>该命令会列出所有 PVC 及其当前状态。 PVC 应具有
Bound状态,以指示其已成功声明 PV。如果状态为
Pending,则表示 PVC 仍在等待合适的 PV,需要进一步调查。 -
如果 PVC 不处于
Bound状态或者您需要更详细的信息,请使用describe命令:kubectl describe pvc <pvc-name> --namespace <namespace>kubectl describe pvc <pvc-name> --namespace <namespace>查找有关状态、事件和任何错误消息的信息。 例如,某个问题可能与存储类配置错误或配额限制有关。
-
检查绑定到 PVC 的持久卷 (PV) 的运行状况:
kubectl get pv <pv-name>kubectl get pv <pv-name>状态应为
Bound。 如果 PV 处于Released或Failed状态,则可能表明基础存储存在问题。 -
如果 Pod 使用 PVC,请检查 Pod 是否已成功装载卷:
kubectl get pod <pod-name> --namespace <namespace>kubectl get pod <pod-name> --namespace <namespace>如果 Pod 处于
Running状态,则表示 PVC 已成功装载。 如果 Pod 处于错误状态(例如InitBackOff),则可能表明卷装载存在问题。 -
如果装载 PVC 存在问题,请描述 Pod 以检查是否存在任何安装错误:
kubectl describe pod <pod-name> --namespace <namespace>kubectl describe pod <pod-name> --namespace <namespace>
已禁用备份
此警示表示备份已禁用。
您需要启用备份功能。
备份部分失败
此警示表示 Velero 备份失败。
您需要联系 UiPath™ 支持团队。
cronjob-alerts
CronJobSuspended
uipath-infra/istio-configure-script-cronjob cronjob 处于挂起状态。
要解决此问题,请执行以下步骤来启用 cronjob:
export KUBECONFIG="/etc/rancher/rke2/rke2.yaml" && export PATH="$PATH:/usr/local/bin:/var/lib/rancher/rke2/bin"
kubectl -n uipath-infra patch cronjob istio-configure-script-cronjob -p '{"spec":{"suspend":false}}'
epoch=$(date +"%s")
kubectl -n uipath-infra create job istio-configure-script-cronjob-manual-$epoch --from=cronjob/istio-configure-script-cronjob
kubectl -n uipath-infra wait --for=condition=complete --timeout=300s job/istio-configure-script-cronjob-manual-$epoch
kubectl get node -o wide
#Verif if all the IP's listed by the previous command are part of output of the following command
kubectl -n istio-system get svc istio-ingressgateway -o json | jq '.spec.externalIPs'
export KUBECONFIG="/etc/rancher/rke2/rke2.yaml" && export PATH="$PATH:/usr/local/bin:/var/lib/rancher/rke2/bin"
kubectl -n uipath-infra patch cronjob istio-configure-script-cronjob -p '{"spec":{"suspend":false}}'
epoch=$(date +"%s")
kubectl -n uipath-infra create job istio-configure-script-cronjob-manual-$epoch --from=cronjob/istio-configure-script-cronjob
kubectl -n uipath-infra wait --for=condition=complete --timeout=300s job/istio-configure-script-cronjob-manual-$epoch
kubectl get node -o wide
#Verif if all the IP's listed by the previous command are part of output of the following command
kubectl -n istio-system get svc istio-ingressgateway -o json | jq '.spec.externalIPs'
IdentityKerberosTgtUpdateFailed
此作业将最新的 Kerberos 票证更新为所有 UiPath™ 服务。此作业失败将导致 SQL Server 身份验证失败。请联系 UiPath™ 支持团队。
- alertmanager.rules
- AlertmanagerConfigInconsistent
- AlertmanagerFailedReload
- AlertmanagerMembersInconsistent
- 常规.规则
- TargetDown
- Watchdog
- prometheus-operator
- PrometheusOperatorListErrors, PrometheusOperatorWatchErrors, PrometheusOperatorSyncFailed, PrometheusOperatorReconcileErrors, PrometheusOperatorNodeLookupErrors, PrometheusOperatorNotReady, PrometheusOperatorRejectedResources
- Prometheus
- PrometheusBadConfig
- PrometheusErrorSendingAlertsToSomeAlertmanagers, PrometheusErrorSendingAlertsToAnyAlertmanager, PrometheusNotConnectedToAlertmanagers
- PrometheusNotificationQueueRunningFull, PrometheusTSDBReloadsFailing, PrometheusTSDBCompactionsFailing, PrometheusNotIngestingSamples, PrometheusDuplicateTimestamps, PrometheusOutOfOrderTimestamps, PrometheusRemoteStorageFailures, PrometheusRemoteWriteBehind, PrometheusRemoteWriteDesiredShards
- PrometheusRuleFailures
- PrometheusMissingRuleEvaluations
- PrometheusTargetLimitHit
- uipath.prometheus.resource.provisioning.alerts
- PrometheusMemoryUsage, PrometheusStorageUsage
- uipath.availability.alerts
- UiPathAvailabilityHighTrafficUserFacing
- 备份
- NFSServerDisconnected
- VolumeBackupFailed
- BackupDisabled
- 备份部分失败
- cronjob-alerts
- CronJobSuspended
- IdentityKerberosTgtUpdateFailed