We've developed and open sourced an extensive test suite for HA voice assistants that will run any LLM or conversation agent through a wide battery of tests. You may view and download the Github repo and run tests yourself at:
View on GitHubLeaderboard
Below shows the current leaderboard of conversation agents and their accuracy mapping messages to HA intents. This only currently measures the "family_home_medium_two_storey" test suite, as the others need to be manually vetted for hallucinations, improper config / sentences, etc.
| Model Name | # Tests | Success | Failed | Duration | % Success |
|---|---|---|---|---|---|
| Sophia NLU | 3715 | 3655 | 60 | 10.1s | 98.4% |
| Gemma4 e4B | 3715 | 2399 | 1316 | 1188m 35s | 64.6% |
| qwen3.5-9b-q4km | 3715 | 2096 | 1619 | 40m 59s | 56.4% |
| Llama 3.1 8B | 3715 | 1493 | 2222 | 437m 14s | 40.2% |
| Qwen3 4B Instruct | 3715 | 1338 | 2377 | 99m 13s | 36.0% |
| HA Assist | 3715 | 540 | 3175 | 1m 6s | 14.5% |
NOTE: Both Sophia NLU and built-in HA Assist are deterministic, meaning runnning the tests against them will always result in the exact same results shown above. LLMs are probabilistic, so the results above will differ slightly but will generally the be the same.
Run the Tests Yourself
You may run the tests yourself against any LLM or conversation agent by grabbing a copy of the test suite repository via Github at:
View on GitHub