Get a better understanding of what goes into building a scorecard and writing a prompt, as well as how the AI model works to optimize your scorecard accuracy.
With PeopleQX’s main feature being quality assurance (QA) scorecards, the end goal is to help you optimize the evaluation experience and, ultimately, foster better communication between your brand and your customers.
Therefore, rather than focusing on prompts that generate content or gather information from external sources, AI scoring is used in PeopleQX’s scorecards, created through adding a series of relevant questions divided into different sections. This is known as the AutoQA process and aims to evaluate calls undertaken by agents in order to highlight areas of success or improvement.
Prompt engineering is the first step in building these scorecards, helping the AI to process, understand, and evaluate based on your brand’s customer-facing requirements. This is what ultimately will accelerate and streamline the entire QA process.
To help understand the components of how these prompts work, let’s take a look at the key elements of our prompts. Here’s an example of a prompt and response on PeopleQX:
There are three key elements from the above:
1. Input data (Question)
This is the common term for the data you’re inputting into the system for the AI to process – in this case, it is the prompt.
Prompts on PeopleQX will always be in the form of a close-ended question, as only a “yes” or “no” answer will be shown.
2. Prompt context (Section name)
Context refers to external information or additional context that can steer the model to better responses. While this isn’t a tangible metric, it is extremely crucial in prompting your AI, and is one of the best practices when it comes to prompt engineering.
In the example above, the context is the introduction of the agent, and there are specifications of the context through giving examples to better help the AI understand what you’re trying to ask it to evaluate.
3. Output data (Answer)
This is the result of your prompt.
Our answers will be either a clear-cut and concise “yes” or “no”.
Finally, in order to optimize your AI’s accuracy before launching your scorecards, it’s essential to test your prompts to verify their strength. It is recommended to try uploading a minimum of three sample files to test your scorecard during this process. This should then be tested against a benchmark score, based on a ground truth, done through manual evaluation on the Manual QA tab – you can learn how to do this here.
This will also enhance reliability and maintain a continuous feedback loop for both agents and AI scoring, improving its overall quality through iterative testing and improvement.