BernyFlow

Multi-tenant WhatsApp-first CRM SaaS · my own product · 742 commits over 10 months

Context

Businesses connect their WhatsApp number and an LLM agent with retrieval over their own documents answers customers, hands the conversation to a human when it should, and runs follow-up flows. The same tenant also gets scheduling, billing and a consignment module, shaped for a few service verticals.

I built and run all of it alone — 69 data models, 228 endpoints, 13 releases. The CRUD is not the interesting part. The interesting part is the message pipeline, where every hard problem is about time: a gateway that retries if you answer slowly, a person who sends three fragments instead of one sentence, and a model whose context you have to assemble correctly on every turn.

Engineering decisions

  1. 01

    Acknowledge the WhatsApp webhook immediately and process it off a queue

    WhyAn LLM call takes seconds. The gateway treats a slow response as a failure and retries it, so answering inline would have duplicated messages under exactly the load I wanted to handle.

    What it costThe queue became a hard dependency: if it is unreachable the endpoint returns 503 and logs the payload as lost. That is an accepted loss window, not a durable outbox — the honest name for a trade-off I have not paid down yet.

  2. 02

    Remove the similarity cutoff from retrieval instead of tuning it

    WhyA weak chunk in the prompt gets judged by the model, which reads it. A cosine-distance threshold decides before reading, and it was throwing away context that turned out to matter.

    What it costNothing short-circuits on a weak match, so more tokens go out on every call — and I have no automated evaluation proving the model-judges approach beats the threshold. I believe it does; I have not measured it.

  3. 03

    Coalesce message bursts with an eight-second window per conversation

    WhyPeople type the way they speak: three fragments in a row. Answering each one separately reads as a broken bot, and costs three LLM calls for one question.

    What it costThe window lives in one process's memory. A worker restart mid-window drops the pending reply, and running a second worker would break the coalescing outright — this is the piece that blocks horizontal scaling.

  4. 04

    Normalise money to integer cents before applying a percentage

    WhyIn floating point, seventy per cent of eighty-five reais evaluates to 59.49999999999999. One cent, every time, compounding across a consignment ledger.

    What it costThe store's share is derived as the total minus the commission and never rounded on its own, so the two always sum exactly — one more invariant to keep. And the older sales modules still do float arithmetic on money, so the discipline is not uniform across the codebase yet.

  5. 05

    Enforce tenant isolation with a filter on every query rather than database row-level security

    WhyIt was simpler with the ORM's query API and kept the schema portable across environments.

    What it costIt is discipline, not a structural guarantee — every new query is a chance to forget. It failed once in production, which is the third incident below.

What broke, and what I learned

The agent stopped seeing new messages after the fifteenth of each conversation

What happenedThe history query ordered ascending and took fifteen rows, so it returned the fifteen *oldest* messages. Past that point the model's view of the conversation was frozen at the beginning. Worse, the customer's current message was never added to the prompt at all — it was only used to extract intent for retrieval, so the model was answering the past while reading nothing of the present.

The fixCorrect the window and inject the current turn explicitly. I found it live from three symptoms that only make sense together: replies repeated word for word, replies that were generically empty, and replies that answered a question from several turns earlier.

A deploy that reported success and did nothing

What happenedThe deploy script was piped to the server over SSH. A health check in the middle of it used `docker compose exec -T`, which also reads standard input — it swallowed the rest of the script, the shell reached end of input and exited zero. The pipeline went green, the release tag was never written, image pruning never ran, and the automatic rollback had no previous version to return to.

The fixCopy the script to the server and run it as a file instead of piping it, and redirect the health check's stdin from /dev/null as a second barrier. The lesson stuck harder than the fix: a deploy that cannot fail loudly will fail quietly, and a green pipeline is a claim, not evidence.

A handoff that could write into another tenant's data

What happenedThe routing code updated a record by id with no tenant filter — the gap that per-query discipline always eventually leaves. Separately, two simultaneous handoffs collided on a composite unique key and the losing write disappeared into a caught error that logged nothing useful.

The fixScope the update by tenant, and replace the read-then-write with an atomic upsert on the composite key. The swallowed error was the more dangerous half: the isolation bug had a fix, but the silence meant nobody would have known to look.