Sophia Test Suite

We've developed and open sourced an extensive test suite for HA voice assistants that will run any LLM or conversation agent through a wide battery of tests. You may view and download the Github repo and run tests yourself at:

View on GitHub

Leaderboard

Below shows the current leaderboard of conversation agents and their accuracy mapping messages to HA intents. This only currently measures the "family_home_medium_two_storey" test suite, as the others need to be manually vetted for hallucinations, improper config / sentences, etc.

Model Name # Tests Success Failed Duration % Success
Sophia NLU 3715 3655 60 10.1s 98.4%
Gemma4 e4B 3715 2399 1316 1188m 35s 64.6%
qwen3.5-9b-q4km 3715 2096 1619 40m 59s 56.4%
Llama 3.1 8B 3715 1493 2222 437m 14s 40.2%
Qwen3 4B Instruct 3715 1338 2377 99m 13s 36.0%
HA Assist 3715 540 3175 1m 6s 14.5%

NOTE: Both Sophia NLU and built-in HA Assist are deterministic, meaning runnning the tests against them will always result in the exact same results shown above. LLMs are probabilistic, so the results above will differ slightly but will generally the be the same.

Run the Tests Yourself

You may run the tests yourself against any LLM or conversation agent by grabbing a copy of the test suite repository via Github at:

View on GitHub