Need Advice - How to Move Sophia Forward?

Matt
Apr 2025
43 posts
0 likes
Member
Need Advice - How to Move Sophia Forward?
Posted Wed Jul 08th 17:55• Last edited Sat Jul 11th 17:00

If you wish to respond but don't currently have a nlu.to account, please just grab a free trial as that's same as registering.

I'm genuinely looking for advice because I'm at a loss atm, so bare with me.

I'm developer of Sophia: https://nlu.to/ha/

On device, low compute NLU engine that acts as a conversation agent. Requires only 160MB RAM, no GPU, millisecond latency, handles unlimited devices, multiple intents per-message, state persistance, requests clarification when needed, doesn't connect to the internet and never calls home.

I'm debating, do I open this up for free, or what do I do here?

For a quick backstory, years ago all within 16 months I went suddenly and totally blind, my primary business colleague of 9 years was murdered via professional hit, and I was forced to relocate back to Canada resulting in the loss of my fiance and dogs. All three of those things individually are life altering, but combined in such a short time span understandably set you back to even less than square one, because now you're stuck rebuilding yourself from scratch without vision.

Developed out Apex (https://apexpl.io/) with hopes of modernizing the Wordpress eco-system, but that was a no go. Then I turned my sights to on device, low compute, deterministic NLU mainly as a push back against big tech. The HA edition is nothing more than a stepping stone for me -- a vital stepping stone, but a stepping stone nonetheless.

I honestly thought this was going to be fairly straight forward, and I thought it solved a real pain point. Built-in Assist is very rigid and not nice, local LLMs require dedicated GPU and come with latency and context window issues, and nobody wants to send their home device config over the wire to big tech. Sophia resolves this pain point.

However, getting traction is seemingly impossible. I have thankfully gotten it to the point where I can quite confidently say it works across all device configs without issue, but even that was way more painful than necessary. Just giving away licenses and getting feedback was in and of itself a full time job.

The only two things I care about in this life right now are giving my partner a hug who has been very patiently waiting years for me on the other side of the world, and getting necessary resources to complete Sophia v2 development which is the version I really care about.

That's when it goes from a narrow domain specific NLU engine into a general purpose one. It will handle multi-turn dialog across any domain, knowledge compression into both relational and time series searchable data stores, surface buried links within large data volumes, and more.

You know how current AI can't even reliably take a Taco Bell order? This will, all on device, low compute, and deterministic so zero hallucinations or probabistic mismatches. I've been at this 2.5 years full time, am 100% confident in the b2 architecture, and this HA edition should be more than enough of a proof of concept.

If desired, quick 1 page more corporate brief at: https://nlu.to/aquila_brief.pdf

If you have any wisdom, perspective or advice to share on how to move forward with this, I would greatly appreciate it. Right now, I'm contemplating just opening the HA edition up free of charge, then using those install numbers as validation to source commercial clients. For example, approach https://orcam.com/ saying xxx number of smart homes are voice controlled with my NLU tech, want a test pilot, kind of thing?

Is there even any appetite for that, or not really? Do I just delete everything, give up, become blind and homeless and call it a life, or? Any advice you have would be greatly appreciated, and if desired you may reach me personally at [email protected]. I can have this thing opened to the world within a few hours if desired.

deastlake
Apr 2026
0 posts
0 likes
Member
Posted Thu Jul 09th 18:22

My 2 cents:

1 - To get more traction in the HA space, there should be VERY frequent posts on the HA subreddit, the HA forum, some of the FB groups, etc. Also, see if you can get some of the more well known HA YouTubers to do a video about Sophia, I think you'll see way more engagement from the community.

2 - Sophia still does some things that are not expected. I really like what you've added where there's a way to report back issues, but I almost feel like every interaction should be logged, with the option to give feedback.

3 - Maybe this is too much, but at my house we have a standard set of voice commands that we've used for A LONG time with the existing voice assistant (Siri). I don't really mind changing what we say to the voice box to turn lights on and off, but I don't even know what vocabulary to use with Sophia. It's all guess work. The internal HA intents are documented, and Claude can understand what's going on with them. If there was documentation on how to interact with Sophia, we'd use it much more at my house.

homeiswherethesmartis
Jun 2026
0 posts
0 likes
Member
Posted Fri Jul 10th 12:13

My initial feeling at this stage of your development is that you need data (as the other person has commented).

Opening up the HA edition would help you get a lot of data, as long as you make it SUPER easy to report feedback, and possibly even give people an opt-in to report errors automatically, as "the ONLY way that it ever phones home" per se.

I'm seeing other projects do this pretty well, like WLED recently added a pop-up when you first flash a device, asking if you mind sharing your config once, never or always. And I'm happy to report that back to them because it's nothing I mind them having.

If you take a similar approach and make it very clear that you would never receive any actual recordings (because you are downstream of that part) you only receive the request text, and only when an action could NOT be confidently completed, then you should quickly highlight the shortfalls and build a better working product.

As an example, I asked Sophia to add something to my shopping list recently and it could not handle that, but the default HA intent engine could. And even slightly differently worded commands to control lights gave differing levels of success.

Keep your head up though, you're clearly very talented, and I look forward to seeing what you achieve with this :)

Matt
Apr 2025
43 posts
0 likes
Member
Posted Sun Jul 12th 02:21

Thanks to both of you for responding, much appreciated. I think it just boils down to the fact I missed the mark when I thought this was a pain point within the HA eco-system.

As for the accuracy issues, I', genuinely unsure what to do there. I can't fix problems I can't see and I'm unable to get feedback.

There were definitely some initial hiccups, mainly because I was unaware of what real world device configs looked like. Confident I've been through enough configs now though that the base software is solid, although some edge cases remain unresolved.

Testing pipeline is about as robust and ambiguous as I can make it. Have multitude of configs, some spanning hundreds of devices and others with thousands. Run a large loop which grabs 2+ entities at random along with random intents, and sends that to a LLM asking for a natural and contextually ambiguous sentence that encompasses them all. I then run those thousands of sentences through Sophia for verification.

So I'm confident the base software is now solid, but also confident there's still a bunch of edge cases out there. Right now, I have that one shopping cart bug on my pending list and thanks to homeiswherethesmartis for sending it in.

That's the only pending issue I have, yet no misinterpretations have been uploaded via the semi-automated feedback pipeline. There's no way I can guess or emulate the remaining edge issues, and need that feedback coming in to understand the issues. When I e-mail folks, 95% of the time it's ignored.

Understandably, I'm not going to bother spending my weeks acting like an aggressive panhandler with hopes of getting the feedback I need. This brings me back to my initial point, I quite clearly mistook what I thought was a solid solution to a real pain point.

I'll leave the site up as a proof of concept while I pursue other venues and verticals in regards to NLU tech. Will probably release an upgrade with built-in STT / TTS models as I'll need to do that anyway, and occasionally release upgrades if feedback comes in, but drop it as a priority. Will probably also add another option to obtan a free license of upload 15 misinterpretations plus your device config, get a free license.

Thanks again for the responses, appreciate it.

Best regards, Matt

Matt
Apr 2025
43 posts
0 likes
Member
Posted Mon Jul 13th 16:50

See, this Reddit thread titled LLM as HA assistant - dumber than sand is what I find frustrating. There's clearly a pain point here, but everyone seems absolutely adamant on using LLMs.

Whatever, changed things so 5 uploaded misinterpretations gets you a free license.

bearded
Jun 2026
0 posts
0 likes
Member
Posted Mon Jul 13th 18:20

I've continued testing Sophia in Flexible mode, and overall my experience has been positive.

When using relatively simple vocabulary and one command per prompt, recognition is now extremely reliable — close to 100% in the scenarios I tested.

For example: "Set all lights in master bedroom to 95%" works correctly.

However, "Set all master bedroom lights to 35%" fails with: "Uh oh, unable to find appropriate response, please let my developers know so they can fix me."

The bigger issue is that when this happens, I haven't been able to successfully report it. Saying something like: "You got the last message wrong." doesn't seem to trigger the expected feedback or logging flow.

My impression so far is that Sophia performs extremely well when the sentence structure closely follows the way Home Assistant intents are typically phrased. Once more natural spoken variations are introduced, it can occasionally fail on wording that I think most users would reasonably expect to work. Another thing I've noticed is that when an entity or area is misinterpreted—for example interpreting "living" instead of "living room"—Sophia may still respond as though the action was completed, even though nothing actually happened. Personally, I think returning an explicit failure would be much better than giving false confirmation.

Overall, when the request isn't an edge case, Sophia is genuinely impressive. It's fast, local, and feels very responsive. But it still occasionally struggles with surprisingly ordinary language variations. From my perspective, the biggest areas to improve are: - making it clearer which sentence structures currently work best, - improving support for natural language variations, - ensuring the feedback mechanism always works when interpretation fails, - and verifying successful execution before confirming that an action was actually performed.

One additional observation that may or may not be useful. I don't think you're only competing on technical capability. Right now, the AI ecosystem seems to be split into roughly three groups: -people who already use cloud AI services such as ChatGPT, Claude or Gemini every day, - people who enjoy building everything locally with Llama, Qwen and other self-hosted models, - and a much larger group that is simply waiting for the ecosystem to mature before committing to anything.

Across all three groups, community size and visibility matter far more than many developers expect. People don't necessarily avoid smaller projects because they think the technology is worse. They hesitate because they don't know whether there will be enough documentation, enough community knowledge, and enough long-term support when they inevitably run into problems.

Then there's the financial side. Many Home Assistant users already pay for ChatGPT, Claude or Gemini because they use those services outside Home Assistant as well. Once somebody already has one of those subscriptions, asking them to invest into another AI ecosystem—even one with clear technical advantages—is a much bigger ask than it first appears. Because of that, I don't actually think Sophia is solving the wrong problem.

I think the challenge is convincing people to adopt another AI ecosystem before they've experienced why it's worth doing. One suggestion I'd seriously consider is engaging more directly with some of the key people in the Home Assistant ecosystem. Conversations with Paulus, Frenck, or maintainers of major community projects such as Music Assistant—or even contributors like the original ESPHome creator—could provide valuable feedback and perhaps create opportunities for wider visibility if they see value in what you're building.

Finally—and this is probably the hardest part—I think it's important to be patient. I completely understand that Sophia is your project and that you've invested an enormous amount of time into it. From your perspective it's easy to see how capable it already is. But the community needs time to reach that conclusion on its own. Home Assistant users are generally very conservative about adopting new core technologies. They rarely jump on version 1.0 of anything. They wait until a project has proven itself, built a community, and demonstrated long-term stability. That may sound frustrating, but I don't think it's a reflection of Sophia's quality.

If the product continues improving, if the community grows, and if people keep seeing successful real-world use cases, the users will come. It just takes longer than any of us would like. I hope that's useful, and I'll keep testing as I continue using it in my own Home Assistant setup.

Matt
Apr 2025
43 posts
0 likes
Member
Posted Wed Jul 15th 15:33

Hey bearded,

Thanks so much for the post, and it's much appreciated. I'm actually going to come out with a much better reply in a day or two, but first, thanks very much for uploading misinterpretations. Congrats, you popped the feedback pipeline cherry!

Quick glance though and I already know, the reason you're having issues is because you're still on flexible mode. Switch over to context only mode, and it will immediately start working much better. Sorry about that, but I did mention in each of the previous updates everyone should switch to context only mode.

Barring a major security reason, I will never release an upgrade that modifies user settings. It's simply bad practice, erodes user's trust, and causes more problems than it's worth so I just don't do it.

However, I will completely remove a setting when it's no longer needed, which I just did with the new v1.3.4 upgrade released. It removes the mode setting altogether and puts everyone on context only. Been meaning to do this, but was just waiting for enough feedback to ensure it's the right decision. It's now done.

Again, there were definitely some initial hiccups as I was unaware of how device configs in the wild looked. Once I managed to get my hands on some larger configs of influencers, I seen what I was working with, and my answer to that was this context only mode in v1.3.

Thanks again for your response, I'll reply in a day or two with a new and large test suite.

Quick Reply

You must be logged in to reply to this thread.

Login Register