Review and rate Batch Test results

Interpret Batch Test responses, apply Good, Acceptable, or Poor ratings, add notes, and export evidence.

Turn a completed Batch Test into an actionable review. Ratings summarize quality, while notes preserve why a response passed or failed and what should change.

An Agent combines instructions, customer-facing interface settings, Knowledge, catalog data, and optional tools. A reliable setup is tested with realistic customer questions before deployment. Tests should cover expected answers, missing information, product scenarios, escalation, and any optional interaction that the team has enabled.

Before you begin

Access: Open a completed Batch Test run with permission to review Agent tests.

  • Use a completed run with stable results.
  • Keep the expected outcome for each case available.
  • Agree on how the team distinguishes Good, Acceptable, and Poor.

Work in the smallest owning area described below and keep the current customer-facing state available while you prepare the change. Before clicking any final action, confirm the active company, Agent, store, language, and market shown in Humind. A missing control can indicate read-only access or a capability that is not configured for this company. In that case, record the intended task and ask an administrator to review the exact permission or dependency. Do not bypass the boundary by sharing an account, copying data into another area, or promising a capability that the workspace does not expose.

Step-by-step workflow

  1. Review the response and evidence

    Read the complete response, not only its opening sentence. Check whether it answers the question, respects boundaries, and relies on appropriate Knowledge or product information.

    • Compare the answer with the expected outcome.
    • Open relevant source material when the result is surprising.
  2. Apply a consistent rating

    Use Good for a ready response, Acceptable for a useful response with a non-blocking issue, and Poor for an incorrect, unsupported, unsafe, or materially incomplete response.

    • Rate the customer impact, not writing style alone.
    • Use the same standard across similar cases.
  3. Add a diagnostic note

    Record the specific reason and likely owning layer, such as Knowledge, catalog, guidance, tool configuration, or unsupported scope. A note should make the next action obvious.

    • Quote only the minimum relevant phrase.
    • Name the correction owner or follow-up test.
  4. Summarize and export

    Group poor and acceptable cases by root cause, then export the report when it needs to be shared outside the review screen. Retest after focused corrections.

    • Prioritize repeated customer-impacting failures.
    • Keep the original run as before-change evidence.

Important limits and operating notes

  • A high aggregate rating can hide one severe failure.
  • Ratings reflect the agreed reviewer standard and need calibration.
  • Changing a source after the run does not change the recorded response.
  • An exported report is evidence, not a live configuration.

Verify the result

  • Every critical case has a rating and diagnostic note.
  • Poor cases are grouped by an owning layer.
  • The export contains the reviewed run rather than a different dataset.
  • A follow-up run is planned for corrected blocking issues.

Keep a short record of what you tested, which customer scenario you used, and what changed. This makes later troubleshooting more precise and helps another teammate reproduce the result without relying on memory.

Troubleshooting

Reviewers disagree on a rating

Return to the expected customer outcome and classify the impact. If the expectation itself is unclear, fix the test definition before using its rating in a release decision.

The answer looks plausible but has no support

Rate the evidence problem explicitly and investigate the source layer. Do not accept a confident answer merely because its wording is polished.

Related guides

Was this article helpful?