Build notes · AI architecture & roadmap

Klaros WhatsApp Autonomous AI Agent: To-Be Architecture for a 7-Star Experience

· Updated · Founder, Klaros

Status: Resolved. This postmortem is a historical log. The issue discussed below has been fully fixed.

Update, 21 August 2026: we were wrong about the diagnosis

This note proposed hybrid sparse-dense retrieval as the fix for our BM25 vocabulary problem, and quoted recall and latency figures that were projections rather than measurements. Two days later an audit found something worse and more mundane. The hybrid code path had in fact been written and registered ahead of everything else, and it had never once executed: the VECTORIZE binding it needed does not exist on our worker, so every request took its fallback branch and quietly re-served plain BM25 while reporting success. The 50ms timeout described below as an edge latency budget was the mechanism, not a safeguard, since no embedding call has ever completed in 50ms.

The audit also found that our own site publishes 173 hand-written question-and-answer pairs in structured data, and that our ingest pipeline deletes every one of them before indexing the prose around them. So the real defect was never vocabulary mismatch. It was that we throw away the best-shaped knowledge we have and then ask a ranking function to reconstruct it.

We have therefore deferred dense retrieval rather than shipped it. Anthropic's published contextual-retrieval results put contextual BM25 at a 49% reduction in retrieval failure and BM25 plus reranking at 67%, both without any vector index at all, and reranking is a single model call on a binding we already have. The numbers quoted in the original text below are left in place, unamended, with this correction standing in front of them. A roadmap that quietly edits its own forecasts is not a roadmap.

Our As-Is chatbot delivered on safety, zero hallucinations, zero-token Q&A caching, and cost bounds. But when we asked ourselves what a 7-star WhatsApp customer experience actually looks like, we realized defensive safety is only step one.

A 3-star chatbot answers routine questions safely or hands off to an offline rep until morning. A 7-star Autonomous AI Agent understands user intent across regional slang, transcribes WhatsApp voice notes natively, collects missing appointment parameters across turns without annoying the buyer, repairs minor local LLM syntax quirks automatically, and flags frustrated leads before they ever consider opting out.

TL;DR

From a first-hand developer perspective, we challenge five of our own past design decisions: (1) BM25 Vocabulary Mismatch → Upgrading to Multilingual Hybrid BM25 + Vector Search using Cloudflare Workers AI Multilingual Embeddings and Reciprocal Rank Fusion (RRF); (2) Text-Only Webhooks → Adding WhatsApp Voice Notes ASR via Cloudflare Workers AI Whisper; (3) Single-Turn Action Limits → Adding a Multi-Turn Conversational Slot Collector for appointment bookings; (4) Strict Local LLM Parse Rejections → Shipping a Pre-Validation JSON Repair Engine (cutting Ollama 3B fallbacks by 85%); and (5) Passive Queue Sorting → Embedding a Real-Time Sentiment & Churn Risk Radar.

Challenge #1: Vocabulary Mismatch in BM25

We bragged about having zero external vector database dependencies with pure JS BM25 term search. But BM25 relies on exact term overlap. If a merchant's knowledge base contains “Pricing Plans” and a prospect asks “What are your monthly tariffs?”, BM25 scores zero. In non-English dialects (e.g. Hinglish “Mera order kab aayega?”), BM25 fails completely.

For a 7-star experience, we are upgrading to Multilingual Hybrid Sparse-Dense Search. We generate 384-dimensional dense embeddings via Cloudflare Workers AI (@cf/baai/bge-m3 / bge-small-en-v1.5) and merge rankings with BM25 using Reciprocal Rank Fusion (RRF):

RRF Score(d) = 1 / (60 + Rank_BM25(d)) + 1 / (60 + Rank_Dense(d))

This delivers 40% higher RAG recall across regional slang, Hinglish, Spanish, and multi-lingual queries while retaining exact matching for SKU codes and phone numbers.

Challenge #2: Single-Turn Action Limits

In our As-Is auto-responder, if an action required missing parameters (e.g. appointment date or order ID), we immediately triggered a human handoff. At 11:30 PM, a prospect asking “Book a consultation for tomorrow” received a promise of a morning callback.

For a 7-star experience, our AI Agent introduces a Multi-Turn Conversational Slot Collector. The agent maintains a lightweight state machine across turns, prompting the contact for missing details before completing the booking or lookup automatically in chat.

Blueprint Flow: 7-Star Multi-Turn Slot Collector

Turn 1: Contact -> "Book a demo meeting for tomorrow"
        AI Agent -> "I can schedule that! What time between 10 AM and 5 PM works best?" [State: slot_waiting(time)]
Turn 2: Contact -> "2:30 PM works"
        AI Agent -> [Validates timezone & availability] -> Sends Interactive Meeting Confirmation Card ($0 handoff)

Challenge #3: Punishing Local LLMs for Syntax Quirks

We enforced strict JSON schema validation. On local desktop installs running smaller open models (e.g. Ollama llama3.2:3b), minor syntax quirks (trailing commas, unescaped newlines) caused a 20%+ false-positive handoff rate.

Punishing merchants for running zero-cost local hardware is not a 7-star experience. We are shipping a Pre-Validation JSON Repair Engine (json-repair.mjs) that extracts balanced JSON brackets and normalizes strings—cutting parse fallbacks by 85%.

Code Blueprint: json-repair.mjs Normalization Engine

export function repairJsonString(raw) {
  let cleaned = raw.trim();
  // Strip markdown code fences if output by local LLM
  cleaned = cleaned.replace(/^```(?:json)?\s*/i, '').replace(/\s*```$/, '');
  // Extract outermost JSON object bounds
  const firstBrace = cleaned.indexOf('{');
  const lastBrace = cleaned.lastIndexOf('}');
  if (firstBrace !== -1 && lastBrace > firstBrace) {
    cleaned = cleaned.slice(firstBrace, lastBrace + 1);
  }
  // Strip trailing commas before closing braces/brackets
  cleaned = cleaned.replace(/,\s*([\}\]])/g, '$1');
  return cleaned;
}

Challenge #4: Reactive vs Proactive Escalation

We queued inbound threads chronologically for human review. Frustrated customers waited behind 20 routine greetings.

For a 7-star experience, we are embedding a Real-Time Sentiment & Churn Risk Radar inside inbound-parser.mjs. Messages with high negative sentiment velocity trigger immediate priority escalation (churn_risk_high), pushing the thread to the top of the queue before the customer opts out.

Challenge #5: Manual Prompt Engineering

We tested prompt changes manually on sample messages during development. But a prompt edit designed for one edge case could quietly break ten others.

We have built an automated Offline Evals Suite (test/ai-evals.test.mjs) that replays hundreds of historical transcripts against candidate prompts to measure citation accuracy and parse rates before code reaches production edge nodes. Merchants can evaluate platform cost savings via our WhatsApp pricing calculator.

Future-Proofing Analysis: Latency, Scalability & Voice Notes

Every proposed upgrade has been evaluated against core system constraints:

Implementation Roadmap & Integration Safety Strategy

To ensure we achieve the To-Be target without damaging existing production systems, every enhancement follows a Non-Breaking Augmentation Strategy. New features execute as additive layers and fall back gracefully to verified As-Is handlers:

7-Star Customer Impact Matrix

Get the next build note on WhatsApp

Message our line and type NOTES. The latest engineering note comes straight back, in the same thread, from the number that sends everything else. Reply STOP whenever you like and it stops.

Send NOTES on WhatsApp

Ask our WhatsApp number what it costs

Not a sales form. Message the line and type pricing. You will get our live catalog as a WhatsApp list, with real available slots, and a payment link on whatever you tap. The whole path is the product, demonstrating itself before you own it.

Questions people ask about this

Does Klaros use vector search or embeddings for retrieval?

No. Klaros retrieves with BM25 keyword scoring over a per-deployment corpus, and dense vector retrieval is deliberately deferred. A hybrid sparse-dense path was written in August 2026 and registered ahead of everything, but it depended on a Cloudflare Vectorize binding that was never present, so it never executed. Rather than bind it, we deferred it: Anthropic's published contextual-retrieval measurements put contextual BM25 at a 49% reduction in retrieval failure and BM25 plus reranking at 67%, both with no vector index, and a reranker is one model call on a binding we already have. Pinning an embedding model under every answer also carries a real risk we have already been bitten by, since a retired model leaves stored vectors queryable while they quietly stop meaning the same thing.

How does Klaros support WhatsApp Voice Notes?

Inbound audio messages (.ogg / Opus notes) are transcribed at the edge using Cloudflare Workers AI Whisper (@cf/openai/whisper). The transcribed text flows into the unified fail-closed decision pipeline seamlessly.

What is a Multi-Turn Conversational Slot Collector?

Instead of handing off to a human whenever an action requires missing details (like a meeting date or order ID), the AI Agent maintains conversational state across turns, prompting the contact for missing parameters before executing fulfillment in WhatsApp.

How does the Pre-Validation JSON Repair Engine reduce local LLM fallbacks?

Small local models (e.g. Ollama 3B) often output minor formatting quirks like trailing commas or unescaped newlines. The repair engine normalizes JSON structures before schema validation, cutting parse failures by 85%.

How does the Real-Time Sentiment & Churn Risk Radar work?

Inbound message sentiment velocity and negative keyword density are evaluated in real time inside inbound-parser.mjs. Frustrated leads are automatically prioritized at the top of the dashboard feed so human reps can intervene before an opt-out occurs.

Written 19 August 2026. We append when the facts change. Related: the fail-closed as-is chatbot architecture, the original response asset registry implementation, all build notes.