Skip to content
    technical EN

    My AI assistant reported success. Half of it was invented.

    An audit of my own AI stack showed that roughly half of the success messages could not be traced to a real action. Three of the four causes were my own build mistakes. This is what I rebuilt, and the rule that remained.

    My AI assistant reported success. Half of it was invented.

    There is a moment when you stop believing your own dashboard. For me it was the moment I wanted to show off a "successful" deploy and the URL didn't exist. Not broken, not slow. Non-existent. So I started digging through my own AI stack. What I found was worse than a bug: roughly half of my AI agent's success messages could not be traced back to a real action.

    One thing up front. This did not turn into a story about a lying language model. Three of the four causes turned out to be my own build mistakes. That doesn't make it better. It does make it more useful.

    A dashboard without bad news

    On paper my setup looked mature. My own AI agent on a VPS, n8n for the workflows, Supabase as the database, a dashboard with status lights. Every run neatly wrote a line: completed. Errors that month: zero.

    That should have been the red flag. Systems that never bring bad news don't bring good news. They bring no news.

    So I started translating every success message back into something checkable. A URL that exists. A row in the database. An HTTP status. That audit of my own AI stack, spring 2026, showed that about 50% of my AI agent's success messages could not be traced to a verifiable action: no URL, no database row, no status. The dashboard showed zero errors for the same period. I never tallied the exact numbers; the ratio was painful enough.

    ~50%
    Share of success messages that could not be traced to a real action
    300 s
    The timeout that was counted as success
    1
    Exit path per workflow, so failure had nowhere to go
    FIGUUR 01 What was left of the success messages once I started demanding evidence.

    How a timeout became a success

    The biggest culprit was painfully banal, and entirely mine. The frontend of my assistant had a 300-second timeout. If no answer came within that time, the operation was closed. Logical so far. Except: it closed with status "success".

    Five minutes of silence were booked as a result. And because nobody double-checks a successful operation, that stayed invisible for months. Every morning I read invented history and called it reporting.

    Four patterns, zero error messages

    The audit produced four variants of the same problem. What they shared: none of them ever raised an error.

    You already know the timeout. The second was worse. Every n8n workflow ended in the same block that wrote "completed", even if something had gone wrong halfway. Failure literally had nowhere to go. Number three was subtler: AI nodes that finished in under 100 milliseconds, with empty output and a green status. A language model doing real work is never done that fast. And the fourth, the only one the model itself was to blame for: textual claims like "deploy succeeded" without any artefact underneath.

    There is a tidy explanation for that fourth one. Language models are trained to guess rather than admit uncertainty: guessing gets rewarded. This is the operational side of AI hallucinations. An AI agent that doesn't know whether something worked says "done" rather than "no idea". Not malice, the default setting. The other three patterns I had simply built myself.

    The fix: a duty of proof

    The rebuild was less work than the discovery. Three boring interventions:

    1. A timeout is an error. Always, without exception. Silence is not evidence.

    2. Every workflow got a second exit path. Failure now has its own route to the database, with its own status. n8n plainly describes this pattern in its error handling documentation; I had just never treated it as mandatory.

    3. A duty of proof. No layer passes a status upward without an artefact from the layer below. Verification as an architecture rule, not a good intention.

    The rule that remains is short: a claim without a URL, status or database row does not exist.

    2
    Exit paths per workflow since the rebuild
    0
    Statuses that can still turn green without an artefact
    <100 ms
    Runtime that now triggers a warning instead of a tick
    FIGUUR 02 Small interventions, structurally different behaviour.

    What this means for your webshop automation

    This looks like a story for people who build their own AI stack. But every e-commerce manager has been running on the same promise for years. Your repricer reports adjusted prices and your feed tool reports updated listings. How many of those messages have you ever checked on the channel itself?

    My rule of thumb since the audit: automation may report what it did, but you only believe what you see at the endpoint. Spot-check on the marketplace itself, not in the tool. Writing about designing with LLMs I arrived at the same place from the UX side: you build trust with verifiable output, not with green ticks.

    The nuance: half of it was right

    Did that make my AI agent worthless? No. The other half of the messages was simply correct, and the system did do useful work. Fabrication is not a unique flaw of my build either. It is the predictable behaviour of any system in which success is the default status. Anyone who has never audited their own stack should not assume it is different there.

    And I deliberately don't make the big claim. Whether nothing has failed unnoticed since, I cannot know; that is exactly what unnoticed means. What I do know: every status now has an artefact underneath it, and what I miss, I no longer miss silently. How I check that this checking itself keeps running became a story of its own: who watches the watchman.

    Build systems that have to prove, not systems that are allowed to claim.

    Sources and further viewing

    Let's talk

    Want to spar about your marketplace strategy?

    No hype. A sober look at where your growth is and where margin leaks away.

    Get in touch