Testing
We built measurable framework to compare how well each voice assistant could understand, respond to, and follow up on a wide range of real-world questions.
We asked 50 questions to each tool: 15 commerce, 20 informational, and 15 reasoning prompts.
We graded the voice assistants on their understanding of the question, need for clarifications, helpfulness, clarity, and conciseness of the response, could it contextually handle a follow-up question, and the speed of response.
Since all voice assistants were fully capable of understanding each question without the need for clarifications, we used the following scoring system:
- 2 points: Was the answer helpful, clear, and concise?
- 2 points: Was the voice assistant able to handle a follow-up question?
- 1 point: Were the responses fast? (i.e., responded within 5 seconds)
If the assistant crushed it, it was awarded 5 points. If it completely failed, 0 points. The range of scores ended up falling between 1-5 points.



