Skip to content
chatAgent
Esc
↑↓navigate↵open⌘Jpreview
On this page

Turn Router Middleware

Turn Router Middleware

What it is

Every inbound message is judged once by Jev (TypeSafe System One) before the shopping agent runs. The same request answers three things: which handler should take the turn, whether answering needs store data, and which keywords the catalog search should run — and the handler is now one of three:

  • direct — a conversational turn; Gemini answers it in one short call with no tools and no thinking.
  • browse — the customer wants to look through the catalog; code answers it from Jev’s keywords and the store’s own names and prices. No model writes anything.
  • agent — everything else: Gemini runs the full shopping agent with tools.

Code, not the model, owns the policy. Jev answers; the thresholds below decide what may be done with the answers.

flowchart TD
    M["inbound message"] --> G{"code gate: text, no audio, no reply-to product"}
    G -- "no" --> A["shopping agent"]
    G -- "yes" --> J["Jev: handler? store data? catalog keywords?"]
    J -- "direct_reply, confident, store data unlikely" --> F["fast lane: same model, no tools, no thinking"]
    J -- "browse_catalog + scored keywords" --> B["code: search the keywords, write the list"]
    J -- "anything else" --> A
    F -- "empty, failed, or claims a price" --> A
    B -- "nothing in stock, store down, no keywords" --> A
    F -- "usable text" --> R["reply to customer"]
    B -- "products found" --> R
    A --> R

Why

The full agent turn sends a ~3.7 KB system prompt (persona + security rules + CTA skill) plus seven tool schemas before the model says a word. The fast lane sends ~1.4 KB and no tools. For a message that needs no store data that is the whole cost of the turn, so greetings and acknowledgements get materially faster and cheaper without changing what customers see.

Browsing goes further: the model is not called at all, so a “show me hoodies” turn costs one Jev judgment and one Shopify search instead of a full agent turn with tool round trips.

Files

File Role
runtime/src/agent-worker/turn-router.ts Config, candidate gate, Jev call, thresholds, log line
runtime/src/agent-worker/search-query.ts Catalog keyword candidates, the Jev question, the query ladder
runtime/src/agent-worker/direct-reply.ts The fast lane: one pi agent run, no tools, capped
runtime/src/agent-worker/model-runtime.ts Provider/model/call wiring shared by both routes
runtime/src/app/prompts.ts DIRECT_REPLY_PROMPT (fast lane) and SHOPPING_AGENT_PROMPT
runtime/src/benchmark/router-probe.ts Run real messages through the judgment and print the routes

Browsing is code’s job

“Show me shirts”, “what hoodies do you have”, “something for the beach” are lookups, not conversations: Jev already picked the query, the catalog returns the products, and the reply is a list. So code writes it (browse-reply.ts):

  • No model generation at all. A browse turn is one Jev judgment (shared with the routing decision, so no extra request) plus one Shopify search — no Gemini call, no tool round trips.
  • Nothing can be invented. Names, prices and stock come verbatim from the store; a price range renders as a range; out-of-stock products are not offered, because a browse that lists a sold-out item sends the customer into a dead end.
  • The turn falls back to the agent whenever code cannot answer it well: no keywords, no matches, nothing in stock, a store outage, a long message, or Jev unsure whether the customer means a category or one specific item.
  • The list carries the store’s photos for the products it names, and the add-to-cart CTA.

Query terms match their singular too (hoodies → hoodie), because a customer typing the plural was silently getting nothing from a store whose titles are singular — Shopify’s own search stems that far, the title filter here did not.

The judgment

One Jev request per candidate message carries the assistant’s role and capabilities, the last 8 conversation messages, and the new message. Three questions are asked together (they run in parallel server-side):

  1. reply_mode (Choice) — direct_reply, browse_catalog, or agent_turn. agent_turn includes “you are not sure”; browse_catalog is only for a kind of product, never for one particular item.
  2. needs_store_data (Noul) — probability that answering needs the catalog, prices, stock, variants, policies, shipping, returns or the cart, or an action on the store.
  3. search_query (Choice) — which candidate phrase should search the catalog (see below). Omitted when code found nothing to choose between, or when every candidate is a request word.

The fast lane requires all of:

  • reply_mode == direct_reply
  • confidence >= LLM_ROUTER_MIN_CONFIDENCE (default 0.6)
  • needs_store_data <= LLM_ROUTER_MAX_STORE_PROB (default 0.2)
  • the code gate passed (text, no audio, no reply-to product) and the message is at most LLM_ROUTER_MAX_CHARS long

Missing or malformed answers fail safe: unknown mode → agent, missing confidence → 0, missing store-data probability → 1.

Catalog search keywords

Shopify search only works well with short product-name phrases; a customer sentence is a bad query. Code builds the candidate phrases, Jev picks one, and the catalog tool searches it:

  • Candidates come from the customer’s message (runs of content words, then longest-first n-grams: carnage retro short, carnage retro, retro short, …) plus the product names in the assistant’s previous reply, so a follow-up like “is it available in black?” still resolves to the product under discussion. Request words, sizes, prices and function words are stripped; the list is capped at 10 phrases.
  • The question offers those phrases plus an explicit (no catalog search) option, so a social message is never forced into a product query.
  • The answer’s probabilities rank every candidate: the pick is the primary query, the runners-up become fallbacks. Below LLM_ROUTER_MIN_QUERY_PROB (default 0.25) no phrase is trusted and the router produces no keywords at all.

Where the keywords are used:

Consumer Behaviour
Speculative prefetch The catalog lookup started in parallel with the model’s first call now runs Jev’s primary phrase (previously the raw customer text, and skipped entirely for messages over 80 characters). Skipped for cart/checkout/policy/small-talk turns.
search_shop_catalog Searches Jev’s phrase when the model’s own query adds no term of its own (it is restating the customer); when the model adds a term — a broader category, an alternative after a miss — its query stays primary. Either way the other rung and Jev’s runners-up are merged into the same lookup as extra attempts, so one tool call still answers both.
Result filtering Unchanged: results must match the effective query’s significant terms, so a fallback can add recall but never noise.

With the router off (or no keywords), the tool behaves exactly as before: the model’s query is the query.

Configuration

Env Default Meaning
LLM_ROUTER_ENABLED 1 Master switch. 0 sends every turn to the agent.
TYPESAFE_API_KEY — Jev credential. Without it the judge is skipped (route not-configured) and every turn runs the agent.
TYPESAFE_BASE_URL — Optional TypeSafe API root (proxy/self-hosted).
LLM_ROUTER_TIMEOUT_MS 1500 Per-attempt deadline. One attempt only — a slow judge must not stack retries onto a waiting customer.
LLM_ROUTER_MAX_CHARS 160 Longest message that may take the fast lane.
LLM_ROUTER_MIN_CONFIDENCE 0.6 Minimum Jev confidence in direct_reply.
LLM_ROUTER_MAX_STORE_PROB 0.2 Highest store-data probability that may still skip the agent.
LLM_ROUTER_MIN_QUERY_PROB 0.25 Minimum probability behind Jev’s catalog keyword phrase before it is used as a query.

All of them are in the control-plane allowlist (runtime/src/app/control.ts), so they can be changed at runtime without a redeploy — use /control/env.

Failure modes

The fast lane is an optimisation; it never owns correctness. Every path below ends in a complete answer from the shopping agent:

Situation Route
No TYPESAFE_API_KEY, or LLM_ROUTER_ENABLED=0 agent (not-configured / router-off)
Audio message, reply-to-product context, or empty message agent (not-a-candidate), no Jev call
Jev unreachable, timed out, or unparsable agent (judge-error)
Jev unsure, or store data likely needed agent (low-confidence / needs-store-data / agent-chosen)
Message too long for a one-shot reply agent (long-message)
Browsing with no usable keywords agent (browse-without-keywords)
Browsing, but the search returns nothing in stock agent (browse falls through)
Jev’s keyword answer is flat, missing, or (no catalog search) no keywords: the tool searches the model’s query, as before
Fast lane returned nothing or a provider error agent
Fast lane text claims a price/currency amount agent — the fast lane has no store data, so such a claim is a hallucination

Observability

One line per turn, plus route fields on the reply timings:

[agent-worker] turn route=browse reason=browse jev=browse_catalog conf=0.99 storeProb=0.98 judgeMs=476 query="shirts" p=1.00 user=9477…
[agent-worker] browse reply for 9477… from Jev keywords "shirts": 2 product(s)
[agent-worker] turn route=agent reason=agent-chosen jev=agent_turn conf=1.00 storeProb=0.98 judgeMs=299 jevTokens=1017/143 query="carnage retro short" p=0.98 user=9477…
[agent-worker] catalog search query="carnage retro short" source=jev modelQuery="do you have the carnage retro short in black size M please?"
[agent-worker] direct reply unusable for 9477… (empty or unsafe text); falling back to the agent turn

timings.route (direct | browse | agent) and timings.routerMs are returned to the actor, which logs them, so latency and cost can be split per route.

Verifying and tuning

If turns you expected on the fast lane run the agent, or a fast-lane reply answers a product question, the thresholds need attention:

# Real messages through the real judgment (needs TYPESAFE_API_KEY).
TYPESAFE_API_KEY=ts_… bun run router:probe "thanks!" "ok" "what shirts do you have?"

# Existing thresholds are read from the same env vars as production.
LLM_ROUTER_MIN_CONFIDENCE=0.8 bun run router:probe "hey"

The probe prints the route code would take plus Jev’s probabilities, which is the evidence to move LLM_ROUTER_MIN_CONFIDENCE / LLM_ROUTER_MAX_STORE_PROB. The unit tests (runtime/test/turn-router.test.ts, runtime/test/direct-reply.test.ts) cover the policy and the fast lane’s prompt/cap/thinking behaviour; they do not judge Jev’s quality — that is what the probe and the route logs are for.

A threshold can also be moved on the running service, without a redeploy — the router keys are in the control-plane allowlist:

# Apply an override (payload is {"values": {...}}), then read it back
curl -X POST https://$DOMAIN/control/env \
  -H "authorization: Bearer $CONTROL_TOKEN" -H "content-type: application/json" \
  --data '{"values":{"LLM_ROUTER_MIN_CONFIDENCE":"0.8"}}'

# Back to the .env defaults
curl -X POST https://$DOMAIN/control/env/reset -H "authorization: Bearer $CONTROL_TOKEN"

On the deployment host the same thing is one command:

docker logs -f chatagent-runtime | grep -E "turn route=|catalog search"

Back to Agent Architecture · Agent Lifecycle

Last updated on October 2, 2026