- Introduction
- Commencer avec les agents UiPath
- Premiers pas avec les agents UiPath utilisant LangGraph
- Créer un agent low-code dans Studio Web
- Ajouter des outils à votre agent UiPath
- Getting Started with UiPath Maestro Flow
Exécutez les traçages d'évaluation localement et examinez les résultats entrant dans Studio Web.
Step 8 - Create evaluation tests
Evaluations test how well your agent performs across a range of inputs, including whether it calls your new tool at the right moments. The uipath-agents skill includes the complete evaluation framework reference: evaluator types, eval set schema, directory structure conventions, and best practices like using gpt-4.1 (not mini) for LLM judge evaluators. Your coding agent uses this to produce correct evaluator configs and test sets from a short prompt.
Demandez à votre agent de codage:
Create an evaluation set for the intake classifier agent with 5 test cases:
1. A clearly trivial request (e.g., deliver a letter) - no creature named, get_challenge_rating should not be called
2. A standard request (e.g., escort a caravan) - no creature named, get_challenge_rating should not be called
3. A heroic request naming a goblin (e.g., clear a goblin stronghold) - get_challenge_rating should be called exactly once, querying for a goblin, and no other creature
4. A legendary request naming a dragon (e.g., slay a dragon) - get_challenge_rating should be called exactly once, querying for a dragon, and no other creature
5. An edge case that's ambiguous on difficulty but also names no specific creature - get_challenge_rating should not be called; this tests that the agent doesn't over-call the tool just because a case is hard to classify
Use both a semantic similarity evaluator (to check the output) and a trajectory evaluator (to check whether get_challenge_rating was called, and with what search term, matching the expectations above).
Include evaluator config files in evaluations/evaluators/ and the eval set, named smoke-test.json, in evaluations/eval-sets/. Use gpt-4.1-2025-04-14 as the model in the evaluator configs. Each evaluator config must include a populated defaultEvaluationCriteria - use {"expectedOutput": {}} for the semantic evaluator and {"expectedAgentBehavior": ""} for the trajectory evaluator. Empty {} fails schema validation.
Create an evaluation set for the intake classifier agent with 5 test cases:
1. A clearly trivial request (e.g., deliver a letter) - no creature named, get_challenge_rating should not be called
2. A standard request (e.g., escort a caravan) - no creature named, get_challenge_rating should not be called
3. A heroic request naming a goblin (e.g., clear a goblin stronghold) - get_challenge_rating should be called exactly once, querying for a goblin, and no other creature
4. A legendary request naming a dragon (e.g., slay a dragon) - get_challenge_rating should be called exactly once, querying for a dragon, and no other creature
5. An edge case that's ambiguous on difficulty but also names no specific creature - get_challenge_rating should not be called; this tests that the agent doesn't over-call the tool just because a case is hard to classify
Use both a semantic similarity evaluator (to check the output) and a trajectory evaluator (to check whether get_challenge_rating was called, and with what search term, matching the expectations above).
Include evaluator config files in evaluations/evaluators/ and the eval set, named smoke-test.json, in evaluations/eval-sets/. Use gpt-4.1-2025-04-14 as the model in the evaluator configs. Each evaluator config must include a populated defaultEvaluationCriteria - use {"expectedOutput": {}} for the semantic evaluator and {"expectedAgentBehavior": ""} for the trajectory evaluator. Empty {} fails schema validation.
Step 9 - Run evaluations
Exécutez l'ensemble d'évaluation localement:
uip codedagent eval agent evaluations/eval-sets/smoke-test.json --workers 3 --output-file eval-results.json
uip codedagent eval agent evaluations/eval-sets/smoke-test.json --workers 3 --output-file eval-results.json
Le cadre d'évaluation exécute chaque cas de test via votre agent et donne un score aux résultats.
| Score | Ce qu'il mesure |
|---|---|
| Similitude sémantique | La proximité de la sortie de l’agent avec la sortie attendue |
| Trajectoire de l'agent | Whether the agent called get_challenge_rating when (and only when) it should have |
Trajectory now means something here. With the tool in place, expect trajectory scores close to 1.0 across all five cases: no tool call on the trivial, standard, and ambiguous cases, and exactly one correctly-targeted tool call on the goblin and dragon cases. A low score tells you the agent called the tool when it should not have, skipped a call it should have made, or looked up the wrong creature, not just whether the final tier happens to be right.
For semantic similarity, scores above 0.8 are generally solid; expect the same for trajectory now that it is tracking something specific. Review eval-results.json to see how your agent performed.
Après vous être connecté à Studio Web lors de l'étape suivante, l'exécution de uip codedagent eval run depuis la CLI charge automatiquement les résultats dans Studio Web; elles apparaissent dans l'onglet Ensembles d'évaluation sous Exécutions.
Le bouton Exécuter les évaluations de Studio Web n’est pas identique. Ce bouton déclenche une exécution de robot cloud nécessitant une prise en charge du runtime Python — une configuration plus impliquée en dehors du champ d'application de ce laboratoire. Utilisez uip codedagent eval run à partir de la CLI à la place; les résultats apparaissent dans Studio Web dans l'une ou l'autre des cas.
Les résultats de l'évaluation locale ayant été confirmés, vous êtes prêt à connecter le projet à Studio Web dans la section suivante.