- Einleitung
- Erste Schritte mit UiPath Agents
- Erste Schritte mit UiPath Agents mit LangGraph
- Erstellen eines Low-Code-Agents in Studio Web
- Hinzufügen von Tools zu Ihrem UiPath Agent
- Getting Started with UiPath Maestro Flow
Führen Sie Evaluierungsablaufverfolgungen lokal aus und überprüfen Sie die Ergebnisse, die in Studio Web einfließen.
Step 8 - Create evaluation tests
Evaluations test how well your agent performs across a range of inputs, including whether it calls your new tool at the right moments. The uipath-agents skill includes the complete evaluation framework reference: evaluator types, eval set schema, directory structure conventions, and best practices like using gpt-4.1 (not mini) for LLM judge evaluators. Your coding agent uses this to produce correct evaluator configs and test sets from a short prompt.
Fragen Sie Ihren Programmier-Agent:
Create an evaluation set for the intake classifier agent with 5 test cases:
1. A clearly trivial request (e.g., deliver a letter) - no creature named, get_challenge_rating should not be called
2. A standard request (e.g., escort a caravan) - no creature named, get_challenge_rating should not be called
3. A heroic request naming a goblin (e.g., clear a goblin stronghold) - get_challenge_rating should be called exactly once, querying for a goblin, and no other creature
4. A legendary request naming a dragon (e.g., slay a dragon) - get_challenge_rating should be called exactly once, querying for a dragon, and no other creature
5. An edge case that's ambiguous on difficulty but also names no specific creature - get_challenge_rating should not be called; this tests that the agent doesn't over-call the tool just because a case is hard to classify
Use both a semantic similarity evaluator (to check the output) and a trajectory evaluator (to check whether get_challenge_rating was called, and with what search term, matching the expectations above).
Include evaluator config files in evaluations/evaluators/ and the eval set, named smoke-test.json, in evaluations/eval-sets/. Use gpt-4.1-2025-04-14 as the model in the evaluator configs. Each evaluator config must include a populated defaultEvaluationCriteria - use {"expectedOutput": {}} for the semantic evaluator and {"expectedAgentBehavior": ""} for the trajectory evaluator. Empty {} fails schema validation.
Create an evaluation set for the intake classifier agent with 5 test cases:
1. A clearly trivial request (e.g., deliver a letter) - no creature named, get_challenge_rating should not be called
2. A standard request (e.g., escort a caravan) - no creature named, get_challenge_rating should not be called
3. A heroic request naming a goblin (e.g., clear a goblin stronghold) - get_challenge_rating should be called exactly once, querying for a goblin, and no other creature
4. A legendary request naming a dragon (e.g., slay a dragon) - get_challenge_rating should be called exactly once, querying for a dragon, and no other creature
5. An edge case that's ambiguous on difficulty but also names no specific creature - get_challenge_rating should not be called; this tests that the agent doesn't over-call the tool just because a case is hard to classify
Use both a semantic similarity evaluator (to check the output) and a trajectory evaluator (to check whether get_challenge_rating was called, and with what search term, matching the expectations above).
Include evaluator config files in evaluations/evaluators/ and the eval set, named smoke-test.json, in evaluations/eval-sets/. Use gpt-4.1-2025-04-14 as the model in the evaluator configs. Each evaluator config must include a populated defaultEvaluationCriteria - use {"expectedOutput": {}} for the semantic evaluator and {"expectedAgentBehavior": ""} for the trajectory evaluator. Empty {} fails schema validation.
Step 9 - Run evaluations
Führen Sie den Auswertungssatz lokal aus:
uip codedagent eval agent evaluations/eval-sets/smoke-test.json --workers 3 --output-file eval-results.json
uip codedagent eval agent evaluations/eval-sets/smoke-test.json --workers 3 --output-file eval-results.json
Das Auswertungsframework führt jeden Testfall über Ihren Agent aus und bewertet die Ergebnisse.
| Bewertung | Was gemessen wird |
|---|---|
| Semantische Ähnlichkeit | Wie genau die Ausgabe des Agents der erwarteten Ausgabe entspricht |
| Agent-Verlauf | Whether the agent called get_challenge_rating when (and only when) it should have |
Trajectory now means something here. With the tool in place, expect trajectory scores close to 1.0 across all five cases: no tool call on the trivial, standard, and ambiguous cases, and exactly one correctly-targeted tool call on the goblin and dragon cases. A low score tells you the agent called the tool when it should not have, skipped a call it should have made, or looked up the wrong creature, not just whether the final tier happens to be right.
For semantic similarity, scores above 0.8 are generally solid; expect the same for trajectory now that it is tracking something specific. Review eval-results.json to see how your agent performed.
Nachdem Sie im nächsten Schritt eine Verbindung mit Studio Web hergestellt haben, werden die Ergebnisse durch uip codedagent eval run über die CLI automatisch in Studio Web hochgeladen; Sie werden auf der Registerkarte Evaluierungssätze unter Ausführungen angezeigt.
Die Schaltfläche „Evals ausführen“ in Studio Web ist nicht dasselbe. Diese Schaltfläche löst eine Cloud Robot-Ausführung aus, die Python-Laufzeitunterstützung erfordert – ein aufwendigeres Setup außerhalb des Rahmens dieses Projekts. Verwenden Sie stattdessen uip codedagent eval run aus der CLI; Die Ergebnisse werden in beiden Versionen in Studio Web angezeigt.
Wenn die lokalen Auswertungsergebnisse bestätigt wurden, können Sie das Projekt im nächsten Abschnitt mit Studio Web verbinden.