LLM fallback intermittently includes “invalid” in spoken responses

chkulakowski
Sep 2026
2 posts
0 likes
Member
LLM fallback intermittently includes “invalid” in spoken responses
Posted 18 hours 44 mins ago• Last edited 18 hours 44 mins ago

LLM fallback intermittently includes “invalid” in spoken responses

Environment

  • Sophia NLU: v2.2.0
  • Home Assistant integration: conversation.sophia_nlu
  • LLM provider: Ollama
  • Ollama version: 0.33.2
  • Model configured in Sophia: sophia-fast
  • Base model: qwen2.5:3b
  • Hardware: Intel Core i3-1215U, CPU-only inference
  • Home Assistant Core version: not recorded
  • Sophia HACS integration version: not recorded separately from engine version

Custom Ollama model definition:

FROM qwen2.5:3b
PARAMETER num_thread 4
PARAMETER num_ctx 4096

Sophia connects directly to the Ollama /api/chat endpoint.

Summary

When Sophia forwards a general-knowledge question to Ollama, the response sometimes contains the standalone word invalid, either before or after an otherwise normal answer.

Home Assistant receives this word as part of the conversation response and passes it to text-to-speech, so it is spoken aloud.

The issue occurs intermittently and has been observed in a new conversation as well as repeated requests in an existing conversation.

Steps to reproduce

  1. Configure Sophia’s LLM fallback to use Ollama with the model above.
  2. Submit general questions through Sophia, using the “Ask AI” trigger.
  3. Repeat several times, including in a new Assist conversation.

Example questions:

  • “Ask AI how many feet are in 1 mile.”
  • “Ask AI what color is the sky?”
  • “Ask AI whats the distance in miles between Austin Texas and Dallas Texas?”

Expected behavior

General questions receive a conversational answer without an internal classification marker.

If invalid is used to signal a home-control request, it should not appear alongside an answer or be spoken to the user.

Actual behavior

Examples copied from Home Assistant pipeline traces:

Example 1: Prefix

invalid

Okay, let me help with that! There are 5280 feet in 1 mile.

Sophia processing time: approximately 3.96 seconds.

Example 2: Suffix

The sky is typically blue, but it can also appear other colors
like pink during sunrise/sunset or yellow during a sunset.
What time of day are you asking about?

Invalid

Sophia processing time: approximately 4.27 seconds.

Example 3: New conversation

Conversation ID: 01M3TKK4VA6FEED5PRAFRYE5E9

Input:

ask ai whats the distance in miles between austin texas and dallas texas?

Response:

invalid

Austin and Dallas are both in Texas, but they're not too close
to be considered just as "not far away" rather than a specific
distance in miles. The actual driving distance is about 110
miles. Would you like to know more about something else?

Sophia processing time: approximately 6.73 seconds.

The factual quality of this answer is a separate concern; this report focuses on the leaked invalid marker.

Relevant Home Assistant trace

For the sky question, both the conversation response and TTS input contained the unwanted word:

intent_output:
  response:
    speech:
      plain:
        speech: >-
          The sky is typically blue, but it can also appear other colors
          like pink during sunrise/sunset or yellow during a sunset.
          What time of day are you asking about?

          Invalid
    response_type: action_done
    data:
      success: []
      failed: []

tts_input: >-
  The sky is typically blue, but it can also appear other colors
  like pink during sunrise/sunset or yellow during a sunset.
  What time of day are you asking about?

  Invalid

This establishes that the word is present before TTS, rather than being introduced by speech synthesis.

Captured Sophia → Ollama request

A packet capture confirmed that Sophia includes the instruction to emit invalid in the same user message as the question.

Captured request body, pretty-printed:

{
  "model": "sophia-fast",
  "messages": [
    {
      "role": "user",
      "content": "\nReply with \"invalid\" if the message  is asking you to perform house operations such as turning lights on or off.\n\nIf not, answer the question in a helpful, concise and conversational way.\n\n-----\nask ai how many inches are in 1 mile?\n\n    "
    }
  ],
  "stream": false,
  "temperature": 0.5,
  "max_tokens": null
}

Additional observations:

  • No separate system message was included in this captured request.
  • The “ask ai” trigger text remained in the forwarded question.
  • temperature and max_tokens were sent at the top level. Please verify their handling for Ollama’s native API, which uses options.temperature and options.num_predict.

Captured Ollama response

The captured exchange happened to return a clean answer:

{
  "model": "sophia-fast",
  "created_at": "2026-10-01T02:22:02.243881102Z",
  "message": {
    "role": "assistant",
    "content": "To answer your question: There are 63,360 inches in 1 mile. \n\nWould you like to know something else about miles or inches, or is this all set?"
  },
  "done": true,
  "done_reason": "stop",
  "total_duration": 4231825201,
  "load_duration": 683380,
  "prompt_eval_count": 82,
  "prompt_eval_duration": 119801000,
  "eval_count": 40,
  "eval_duration": 4109232000
}

This confirms the prompt construction, but it does not conclusively establish where the word was introduced during the failing exchanges. A request/response capture of an affected exchange would help confirm that.

Direct Ollama comparison

A direct request without Sophia’s classification instruction returned:

There are 5280 feet in one mile.

Timing:

  • Total: approximately 1.52 seconds
  • Generation: approximately 9.8 tokens/second
  • No invalid marker

This demonstrates that the model can answer cleanly. It does not exclude prompt-dependent model behavior.

Suspected cause

The fallback prompt asks the same small language model to both:

  1. Classify home-control requests using the literal word invalid.
  2. Generate conversational answers for other requests.

The model may occasionally blend these instructions and include the classification marker alongside an answer.

This is a hypothesis supported by the captured prompt, not a confirmed root cause.

Separate timeout observation

Earlier requests sometimes returned “Sorry, I didn’t understand that” after approximately 10 seconds. However, the marker-containing examples above completed in approximately 4–7 seconds.

The invalid issue therefore also occurs on completed requests below that timeout and should be investigated separately.

Matt
Apr 2025
36 posts
0 likes
Member
Posted 18 hours 14 mins ago

Thanks for letting me know, and I've added it to my todo list. There will be a small patch upgrade coming out this weekend probably and I'll include it in that.

It's happening because the LLM messed up and isn't providing the proper response that Sophia asked for and expects. I'll modify it so it retries 3 times and if th eLLM still doesn't produce a properly formatted response it will output a better error than just "invalid".

Thanks again for letting me know.

Quick Reply

You must be logged in to reply to this thread.

Login Register