Understand question quality, and check the match rate of AutoQA evaluations against benchmark scores to help you decide if the scorecard is of acceptable standards.
This article supports Super User and Supervisor roles only.
You can view AutoQA’s performance against results from benchmark scores to gauge if the question is up to par. The quality score is only computed if benchmark scores are available.
If scores were added in the Manual QA form that was uploaded while adding questions, the quality page will have full details. Otherwise, you will see this message and you can begin adding the benchmark scores by clicking ‘Start Manual QA’.
Listen to the audio samples and input the answers for the quality scores to be calculated and displayed.
Now, let’s go through the quality page, one component at a time.
‘Overall AutoQA quality‘
- How AutoQA matched up to Manual QA evaluations across all the samples is computed.
- If you had previously tested the scorecard, the difference between the current test and the most recent will be shown.
- You can download the entire page as a reference or to be shared easily with collaborators.
‘Test attempt’
If this is not the first round of testing, you will be able to view quality scores from previous attempts.
‘Average scores‘
The average score is derived from the results of all the audio samples tested, by both AutoQA and Manual QA. Average scores according to sections are included — it helps you decide if a section requires revision or if AutoQA is too far off from Manual QA.
‘Question quality‘
These indicate the number of times samples tested returned a TP, TN, FP, or FN classification when the results of AutoQA are compared to the benchmark scores. The tabulated accuracy rate across the samples will determine the question’s quality.
Compare AutoQA and Manual QA
When you switch to the ‘Comparison’ tab, you can view the AutoQA classification (TP/TN/FP/FN) for each question, on each audio sample. Access the evaluation preview by:
- clicking on any audio to preview its evaluation in full.
- selecting any of the classifications to be directed to the specific question.
‘Refine question’
This floating action button allows you to tweak any questions that do not meet the expected quality. You will be redirected to the adding question phase.
‘Evaluation preview’
This is a snapshot of the most likely result when the evaluation is applied to conversations using AutoQA.
Audio selection
Use the panel on the left to pick the audio sample you’d like to assess.
‘AutoQA Grading’
This shows the overall grade of the conversation after being evaluated by AutoQA. The grades will appear as Pass, Fail, or Critical Failure, with the score breakdown according to the sections.
‘AutoQA quality’
This percentage denotes how well AutoQA matched against Manual QA evaluation.
‘Manual QA evaluation’
Should you disagree with AutoQA’s evaluation, you can change the result manually by selecting from the drop-down button. It revises the match rate so that you’ll know whether questions need to be tweaked if the rate dips.
Remember to ‘Save’ the changes otherwise they will not be applied.
You’re almost ready to apply the scorecard for evaluation. Take a look at the indicative metrics to decide if you’d want to proceed to apply for evaluation or refine questions.
- ‘Save question’ allows you to ensure questions with at least 98% accuracy remain in the system to be reused and shared with other builders. Switch off for questions that you don’t want to keep.
- ‘Apply for evaluation’ when you’re ready.
You will have the last opportunity to revise the scorecard but if you’re confident, go ahead and send the scorecard to be queued for AutoQA evaluation to start.
Great work, you’ll be directed back to the ‘Scorecards’ listing page and when the evaluations are complete, the status will read ‘Ready’. Do catch up on the first two phases of adding questions and testing questions (if you haven’t done so already).
However, should there be issues with the questions, the apply for evaluation button is disabled and you’ll be prompted to review the error. Possible reasons include:
- Empty or invalid input in the uploaded Manual QA form containing pre-written questions and/or benchmark scores.
- Sections were reordered and conditions of question logic are no longer valid.