- はじめに
- UiPath Agents の利用を開始する
- LangGraph を使用した UiPath Agents の利用を開始する
- Studio Web でローコード エージェントを構築する
- UiPath エージェントにツールを追加する
- Getting Started with UiPath Maestro Flow
評価のトレースをローカルで実行し、Studio Web に流入する結果を確認する
Step 8 - Create evaluation tests
Evaluations test how well your agent performs across a range of inputs, including whether it calls your new tool at the right moments. The uipath-agents skill includes the complete evaluation framework reference: evaluator types, eval set schema, directory structure conventions, and best practices like using gpt-4.1 (not mini) for LLM judge evaluators. Your coding agent uses this to produce correct evaluator configs and test sets from a short prompt.
コーディング エージェントに依頼してください。
Create an evaluation set for the intake classifier agent with 5 test cases:
1. A clearly trivial request (e.g., deliver a letter) - no creature named, get_challenge_rating should not be called
2. A standard request (e.g., escort a caravan) - no creature named, get_challenge_rating should not be called
3. A heroic request naming a goblin (e.g., clear a goblin stronghold) - get_challenge_rating should be called exactly once, querying for a goblin, and no other creature
4. A legendary request naming a dragon (e.g., slay a dragon) - get_challenge_rating should be called exactly once, querying for a dragon, and no other creature
5. An edge case that's ambiguous on difficulty but also names no specific creature - get_challenge_rating should not be called; this tests that the agent doesn't over-call the tool just because a case is hard to classify
Use both a semantic similarity evaluator (to check the output) and a trajectory evaluator (to check whether get_challenge_rating was called, and with what search term, matching the expectations above).
Include evaluator config files in evaluations/evaluators/ and the eval set, named smoke-test.json, in evaluations/eval-sets/. Use gpt-4.1-2025-04-14 as the model in the evaluator configs. Each evaluator config must include a populated defaultEvaluationCriteria - use {"expectedOutput": {}} for the semantic evaluator and {"expectedAgentBehavior": ""} for the trajectory evaluator. Empty {} fails schema validation.
Create an evaluation set for the intake classifier agent with 5 test cases:
1. A clearly trivial request (e.g., deliver a letter) - no creature named, get_challenge_rating should not be called
2. A standard request (e.g., escort a caravan) - no creature named, get_challenge_rating should not be called
3. A heroic request naming a goblin (e.g., clear a goblin stronghold) - get_challenge_rating should be called exactly once, querying for a goblin, and no other creature
4. A legendary request naming a dragon (e.g., slay a dragon) - get_challenge_rating should be called exactly once, querying for a dragon, and no other creature
5. An edge case that's ambiguous on difficulty but also names no specific creature - get_challenge_rating should not be called; this tests that the agent doesn't over-call the tool just because a case is hard to classify
Use both a semantic similarity evaluator (to check the output) and a trajectory evaluator (to check whether get_challenge_rating was called, and with what search term, matching the expectations above).
Include evaluator config files in evaluations/evaluators/ and the eval set, named smoke-test.json, in evaluations/eval-sets/. Use gpt-4.1-2025-04-14 as the model in the evaluator configs. Each evaluator config must include a populated defaultEvaluationCriteria - use {"expectedOutput": {}} for the semantic evaluator and {"expectedAgentBehavior": ""} for the trajectory evaluator. Empty {} fails schema validation.
Step 9 - Run evaluations
評価セットをローカルで実行します。
uip codedagent eval agent evaluations/eval-sets/smoke-test.json --workers 3 --output-file eval-results.json
uip codedagent eval agent evaluations/eval-sets/smoke-test.json --workers 3 --output-file eval-results.json
評価フレームワークによって、各テスト ケースがエージェントを介して実行され、結果がスコアリングされます。
| スコア | 評価項目 |
|---|---|
| Semantic similarity (意味的類似性) | エージェントの出力が期待される出力にどの程度一致しているか |
| エージェントの軌跡 | Whether the agent called get_challenge_rating when (and only when) it should have |
Trajectory now means something here. With the tool in place, expect trajectory scores close to 1.0 across all five cases: no tool call on the trivial, standard, and ambiguous cases, and exactly one correctly-targeted tool call on the goblin and dragon cases. A low score tells you the agent called the tool when it should not have, skipped a call it should have made, or looked up the wrong creature, not just whether the final tier happens to be right.
For semantic similarity, scores above 0.8 are generally solid; expect the same for trajectory now that it is tracking something specific. Review eval-results.json to see how your agent performed.
次の手順で Studio Web に接続した後、CLI から uip codedagent eval run を実行すると、結果が Studio Web に自動的にアップロードされます。評価セットは [ 評価セット ] タブの [実行] に表示されます。
Studio Web の [評価を実行] ボタンは同じものではありません。このボタンをクリックすると、Python ランタイムのサポートが必要な Cloud ロボットの実行がトリガーされます。このラボの範囲外にある、より複雑な設定です。代わりに CLI からの uip codedagent eval run を使用してください。どちらの方法でも、結果は Studio Web に表示されます。
ローカルでの評価結果を確認したら、次のセクションに進むでプロジェクトを Studio Web に接続する準備が整いました。