Create and run a Batch Test
Build a reusable Batch Test dataset manually, from CSV, or from the catalog, then run it against an Agent.
Use Batch Test to evaluate several repeatable scenarios together instead of relying only on an informal Playground conversation. A dataset becomes a stable regression checklist for important customer questions.
An Agent combines instructions, customer-facing interface settings, Knowledge, catalog data, and optional tools. A reliable setup is tested with realistic customer questions before deployment. Tests should cover expected answers, missing information, product scenarios, escalation, and any optional interaction that the team has enabled.
Before you begin
Access: Open Agent setup and testing with permission to run tests for the selected Agent.
- Choose the Agent under test.
- Define the customer journeys the dataset should cover.
- Prepare questions manually, in a compatible CSV, or from catalog products.
Work in the smallest owning area described below and keep the current customer-facing state available while you prepare the change. Before clicking any final action, confirm the active company, Agent, store, language, and market shown in Humind. A missing control can indicate read-only access or a capability that is not configured for this company. In that case, record the intended task and ask an administrator to review the exact permission or dependency. Do not bypass the boundary by sharing an account, copying data into another area, or promising a capability that the workspace does not expose.
Step-by-step workflow
Define the test objective
Create a focused dataset for one release, risk, or journey. Mix expected success cases with missing-information and boundary cases so a high score cannot hide unsafe behavior.
- Give the dataset a specific, reusable name.
- Write the acceptance expectation beside each planned question.
Add test cases
Enter questions manually, import a CSV with up to 50 cases, or generate product-oriented cases from the catalog. Review imported text before running it.
- Remove duplicates and customer personal data.
- Keep each case understandable without hidden context.
Select the Agent and run
Confirm the Agent and dataset, then start the run. Let the run reach its final state before interpreting partial results.
- Do not edit the target configuration during the run.
- Record the run time and configuration being evaluated.
Preserve the dataset for regression
After the run, keep useful cases and refine ambiguous ones. Reuse the same dataset after a Knowledge, guidance, catalog, or tool change.
- Add newly discovered failure cases.
- Avoid rewriting old cases simply to improve the score.
Important limits and operating notes
- CSV import is limited to 50 test cases per dataset.
- Generated product cases depend on the available catalog.
- A batch result does not replace live channel validation.
- Do not include real customer secrets or personal information in test prompts.
Verify the result
- The dataset contains the intended number of distinct cases.
- The run reaches a completed or clearly reported final state.
- Each response can be opened and reviewed.
- The dataset can be selected again for a later regression run.
Keep a short record of what you tested, which customer scenario you used, and what changed. This makes later troubleshooting more precise and helps another teammate reproduce the result without relying on memory.
Troubleshooting
CSV cases do not import as expected
Check the file structure, encoding, row count, and required question content. Start with a small clean file, verify it imports, then add the remaining cases.
The run remains queued or fails
Keep the dataset unchanged, note the reported status, and retry only after confirming the Agent and service are available. Repeated failure should be reported with the run context.