Follow these guidelines when building scorecards to better optimize your questions for effective, accurate, and efficient AI scoring.
To build an effective scorecard, you first need to construct effective prompts – in this case, these refer to the questions, or the beginning of building your scorecard, which the AI will use to evaluate agent performance.
In order for the AI to properly evaluate your scorecard questions, your prompts will have to be written in a way that is straightforward, clear, and easy for the AI to understand. This is paramount in the AI scoring process and is used to improve the quality, relevance, and accuracy of your responses.
This is especially true when it comes to conversational analysis since each exchange might contain certain nuanced, linguistic, or tonal differences depending on both the caller and agent. However, there are a few best practices to be followed to leverage your AI for effective scoring.
Rule #1: Be specific
To get the most accurate and relevant answers, your question should be as specific as possible to minimize uncertainty and allow the AI to properly identify, understand, and process the prompt.
To do this, your questions should be to the point, and have a clear directive in order to achieve the right answer.
⭐️ Be sure to avoid leading questions, as this might incorrectly influence the AI’s outcome.
✅ Correct: “Did the agent greet the customer at the start of the call?”
❌ Incorrect: “Was a “hello” said by the agent to the customer when the call started?”
Rule #2: Optimize your prompt length
The length of your questions is another significant part of prompting.
Here, it’s important to strike a good balance, which means your question prompt shouldn’t be too short, which could result in an inaccurate answer due to ambiguity or vagueness.
However, it shouldn’t be too long either. Although we encourage you to be specific, this doesn’t mean you should go above and beyond when doing so, as this could easily overwhelm or confuse the AI and produce inaccurate information as a result. Break your questions up into smaller, more manageable prompts for your AI.
✅ Correct:
“Did the agent ask for the customer’s name during the call?”
“Did the agent mention the customer’s name at least 3 times during the conversation?”
❌ Incorrect:
“Did the agent mention the customer by name at least three times during the call, spanning from either the beginning of the call, the middle of the call, or right at the end of the call, to acknowledge the customer individually?”
Rule #3: Provide context and examples
Context is vital when it comes to prompting, and can often be considered one of the key elements of a prompt to demonstrate clarity and relevance.
To build an efficient scorecard, the easiest way would be to segregate the call into distinct parts: you’d have the greeting at the beginning of the call, the customer issue or problem, the agent’s solution delivery, as well as the overall professionalism of the agent.
Following this sequence, your prompt should also correspond to these parts when it comes to contextually prompting your AI scorecard. Would this be part of the greeting, the probing for the issue, the solution delivery, or the agent’s mannerisms?
Examples are a helpful tool in contextualizing your prompt further: for instance, examples of an agent greeting could be “Good morning”, or “Hello”.
✅ Correct: “Did the agent introduce himself? For example, “I am…”, “My name is…”?”
❌ Incorrect: “Did the agent say what their name is?”
Rule #4: Focus on the “Do’s” rather than the “Don'ts”
Focusing on the desired outcome or answer is a lot more effective than detailing what you’re not looking for when it comes to prompting. This helps reduce confusion when it comes to evaluating the answers and guides the AI in a more constructive manner.
This rule is especially pertinent to PeopleQX since we’re trying to get to the bottom of a particular set of actions, rather than trying to contradict them. This helps both you and the AI get a straightforward result instead of having to spend more time thinking about the phrasing of each response.
✅ Correct: “Did the agent provide a reason for the call transition?”
❌ Incorrect: “Was the reason for the transition not given?”
Rule #5: Iterate and evaluate
The advancement in AI has been remarkable to witness, but it’s necessary to keep in mind that it’s a still-emerging sector that requires further refinement and experimentation.
As such, it’s always important to test your prompts and refine or modify them iteratively by starting with a basic prompt rather than hitting the ground running with a complex question. Always remember to follow the rule of thumb of PeopleQX to test each scorecard with at least three audio files to ensure you obtain the most accurate results. Do keep in mind that you should have a benchmark score, or a fundamental ground truth, to test your scorecards against, and the audio files should be of good quality with relevant questions to fully optimize this testing process.
However, if you find that the scorecard might not work in your favor, there’s no need to worry – PeopleQX also has a manual QA that overrides the AI scoring in cases where users need to navigate through conversational nuance, tone, or slang better.
Curious about how PeopleQX scorecards work? You can find out all about them here.