UiPath Documentation
automation-suite
2.2510
true
Linux 版 Automation Suite 安装指南
重要 :
请注意,此内容已使用机器翻译进行了部分本地化。 新发布内容的本地化可能需要 1-2 周的时间才能完成。

管理警示

解决 Automation Suite 中与 Alertmanager 配置不一致相关的管理警示。

alertmanager.rules

AlertmanagerConfigInconsistent

当同一集群中的 Alertmanager 实例具有不同配置时,将触发此警示。 这可能表明配置存在问题,该问题在 Alertmanager 的所有实例中不一致。

要解决此问题,请执行以下步骤:

  1. 在已部署的所有 alertmanager.yml 之间运行 diff 工具以识别问题。
  2. 删除不正确的密码并部署正确的密码。

如果问题仍然存在,请联系 UiPath™ 支持团队。

AlertmanagerFailedReload

警示管理器加载或重新加载配置失败。请检查任何自定义警示管理器配置中是否存在输入错误,否则请联系 UiPath™ 支持团队并提供支持包。有关详细信息,请参阅使用 Automation Suite 支持捆绑包

AlertmanagerMembersInconsistent

这些是具有多个警示管理器副本的 HA 集群的内部警示管理器错误。警示可能会间歇性地出现和消失。暂时缩小规模,然后扩大警示管理器副本可能会解决此问题。

要解决此问题,请执行以下步骤:

  1. 缩放至零。请注意,Pod 需要一段时间才能关闭:

    statefulset.apps/alertmanager-monitoring-kube-prometheus-alertmanager scaled
    statefulset.apps/alertmanager-monitoring-kube-prometheus-alertmanager scaled
    
  2. 缩小到 2:

    kubectl scale statefulset -n monitoring alertmanager-monitoring-kube-prometheus-alertmanager --replicas=2
    kubectl scale statefulset -n monitoring alertmanager-monitoring-kube-prometheus-alertmanager --replicas=2
    
  3. 检查 Alertmanager Pod 是否已启动以及是否处于正在运行状态:

    kubectl get po -n monitoring
    kubectl get po -n monitoring
    

如果问题仍然存在,请联系 UiPath™ 支持团队。

常规.规则

TargetDown

Prometheus 无法从警示中的目标收集指标,这意味着 Grafana 仪表板和基于该目标的指标的进一步警示不可用。检查与该目标相关的其他警示。

Watchdog

这是一个警示,用于确保整个警示管道正常运行。此警示始终处于触发状态。因此,它应始终在“警示管理器”中针对接收器触发。有各种通知机制的集成,可在此警示未触发时通知您。例如,PagerDuty 中的 DeadMansSnitch 集成。

prometheus-operator

PrometheusOperatorListErrors, PrometheusOperatorWatchErrors, PrometheusOperatorSyncFailed, PrometheusOperatorReconcileErrors, PrometheusOperatorNodeLookupErrors, PrometheusOperatorNotReady, PrometheusOperatorRejectedResources

控制 Prometheus 资源的 Prometheus 运算符的内部错误。存在这些错误时,Prometheus 本身可能仍然运行良好;但是,此错误表示监控可配置性下降。请联系 UiPath™ 支持团队。

Prometheus

PrometheusBadConfig

Prometheus 加载或重新加载配置失败。请检查任何自定义 Prometheus 配置是否存在输入错误。否则,请联系 UiPath™ 支持团队。

PrometheusErrorSendingAlertsToSomeAlertmanagers, PrometheusErrorSendingAlertsToAnyAlertmanager, PrometheusNotConnectedToAlertmanagers

从 Prometheus 到警示管理器的连接不正常。指标仍可查询,并且 Grafana 仪表板可能仍会显示指标,但不会触发警示。检查警示管理器的任何自定义配置是否存在输入错误,否则请联系 UiPath™ 支持团队。

PrometheusNotificationQueueRunningFull, PrometheusTSDBReloadsFailing, PrometheusTSDBCompactionsFailing, PrometheusNotIngestingSamples, PrometheusDuplicateTimestamps, PrometheusOutOfOrderTimestamps, PrometheusRemoteStorageFailures, PrometheusRemoteWriteBehind, PrometheusRemoteWriteDesiredShards

表示可能无法按预期收集指标的内部 Prometheus 错误。请联系 UiPath™ 支持团队。

PrometheusRuleFailures

如果存在基于不存在的指标或不正确的 PromQL 语法的格式错误警示,则可能会发生这种情况。如果未添加自定义警示,请联系 UiPath™ 支持团队。

PrometheusMissingRuleEvaluations

Prometheus 无法评估是否应触发警示。如果警示太多,可能会发生这种情况。请删除昂贵的自定义警示评估和/或查看有关增加 Prometheus CPU 限制的文档。如果未添加自定义警示,请联系 UiPath™ 支持团队。

PrometheusTargetLimitHit

Prometheus 要收集的目标过多。如果添加了额外的 ServiceMonitor(请参阅监控控制台),您可以将其删除。

uipath.prometheus.resource.provisioning.alerts

PrometheusMemoryUsage, PrometheusStorageUsage

当集群接近配置的内存和存储限制时,这些警示会发出警告。 这可能发生在最近使用量大幅增加的集群上(通常来自机器人而不是用户),或者在未调整 Prometheus 资源的情况下将节点添加到集群中时。 这是因为要收集的指标数量增加。

这也可能是由于触发了大量警示所致,因此请务必检查触发大量警示的原因。

如果此问题仍然存在,请联系 UiPath™ 支持团队并提供生成的支持捆绑包

uipath.availability.alerts

UiPathAvailabilityHighTrafficUserFacing

来自 UiPath™ 服务的 HTTP 500 响应数量超过给定阈值。下表描述了用于评估错误率的流量级别阈值。

流量级别20 分钟内的请求数错误阈值(适用于 HTTP 500)
>100,0000.1%
10,000 到 100,000 之间1%
< 10,0005%

面向用户的服务中的错误可能会导致可在 Automation Suite 用户界面中直接观察到的功能降级,而后端服务中的错误则不会产生明显的后果。

警示会指明哪个服务的错误率较高。要了解报告服务所依赖的其他服务可能存在哪些级联问题,您可以使用 Istio 工作负载仪表板,该仪表板会显示服务之间的错误。

请仔细检查所有最近重新配置的 Automation Suite 产品。还可以使用 kubectl logs 命令获取详细日志。如果错误仍然存在,请联系 UiPath™ 支持团队。

备份

NFSServerDisconnected

此警示表示 NFS 服务器连接已丢失。

您需要检查 NFS 服务器连接和装载路径。

VolumeBackupFailed

此警示表示 PVC 的备份失败。

要解决此问题,请执行以下步骤:

  1. 检查 PVC 的状态,以确保对于持久卷 (PV),它是 Bound

    kubectl get pvc --namespace <namespace>
    kubectl get pvc --namespace <namespace>
    

    该命令会列出所有 PVC 及其当前状态。 PVC 应具有 Bound 状态,以指示其已成功声明 PV。

    如果状态为 Pending,则表示 PVC 仍在等待合适的 PV,需要进一步调查。

  2. 如果 PVC 不处于 Bound 状态或者您需要更详细的信息,请使用 describe 命令:

    kubectl describe pvc <pvc-name> --namespace <namespace>
    kubectl describe pvc <pvc-name> --namespace <namespace>
    

    查找有关状态、事件和任何错误消息的信息。 例如,某个问题可能与存储类配置错误或配额限制有关。

  3. 检查绑定到 PVC 的持久卷 (PV) 的运行状况:

    kubectl get pv <pv-name>
    kubectl get pv <pv-name>
    

    状态应为 Bound。 如果 PV 处于 ReleasedFailed 状态,则可能表明基础存储存在问题。

  4. 如果 Pod 使用 PVC,请检查 Pod 是否已成功装载卷:

    kubectl get pod <pod-name> --namespace <namespace>
    kubectl get pod <pod-name> --namespace <namespace>
    

    如果 Pod 处于 Running 状态,则表示 PVC 已成功装载。 如果 Pod 处于错误状态(例如 InitBackOff),则可能表明卷装载存在问题。

  5. 如果装载 PVC 存在问题,请描述 Pod 以检查是否存在任何安装错误:

    kubectl describe pod <pod-name> --namespace <namespace>
    kubectl describe pod <pod-name> --namespace <namespace>
    

已禁用备份

此警示表示备份已禁用。

您需要启用备份功能。

备份部分失败

此警示表示 Velero 备份失败。

您需要联系 UiPath™ 支持团队。

cronjob-alerts

CronJobSuspended

uipath-infra/istio-configure-script-cronjob cronjob 处于挂起状态。

要解决此问题,请执行以下步骤来启用 cronjob:

export KUBECONFIG="/etc/rancher/rke2/rke2.yaml" && export PATH="$PATH:/usr/local/bin:/var/lib/rancher/rke2/bin"
kubectl -n uipath-infra patch cronjob istio-configure-script-cronjob -p '{"spec":{"suspend":false}}'
epoch=$(date +"%s")
kubectl -n uipath-infra create job istio-configure-script-cronjob-manual-$epoch --from=cronjob/istio-configure-script-cronjob
kubectl -n uipath-infra wait --for=condition=complete --timeout=300s job/istio-configure-script-cronjob-manual-$epoch
kubectl get node -o wide
#Verif if all the IP's listed by the previous command are part of output of the following command
kubectl -n istio-system get svc istio-ingressgateway -o json | jq '.spec.externalIPs'
export KUBECONFIG="/etc/rancher/rke2/rke2.yaml" && export PATH="$PATH:/usr/local/bin:/var/lib/rancher/rke2/bin"
kubectl -n uipath-infra patch cronjob istio-configure-script-cronjob -p '{"spec":{"suspend":false}}'
epoch=$(date +"%s")
kubectl -n uipath-infra create job istio-configure-script-cronjob-manual-$epoch --from=cronjob/istio-configure-script-cronjob
kubectl -n uipath-infra wait --for=condition=complete --timeout=300s job/istio-configure-script-cronjob-manual-$epoch
kubectl get node -o wide
#Verif if all the IP's listed by the previous command are part of output of the following command
kubectl -n istio-system get svc istio-ingressgateway -o json | jq '.spec.externalIPs'

IdentityKerberosTgtUpdateFailed

此作业将最新的 Kerberos 票证更新为所有 UiPath™ 服务。此作业失败将导致 SQL Server 身份验证失败。请联系 UiPath™ 支持团队。

此页面有帮助吗?

连接

需要帮助? 支持

想要了解详细内容? UiPath Academy

有问题? UiPath 论坛

保持更新