---
title: Turn Router Middleware
---
# Turn Router Middleware

## What it is

Every inbound message is judged once by **Jev** (TypeSafe System One) before the
shopping agent runs. The same request answers three things: which handler should
take the turn, whether answering needs store data, and **which keywords the
catalog search should run** — and the handler is now one of three:

- **direct** — a conversational turn; Gemini answers it in one short call with no
  tools and no thinking.
- **browse** — the customer wants to look through the catalog; **code answers it**
  from Jev's keywords and the store's own names and prices. No model writes
  anything.
- **agent** — everything else: Gemini runs the full shopping agent with tools.

Code, not the model, owns the policy. Jev answers; the thresholds below decide
what may be done with the answers.

```mermaid
flowchart TD
    M["inbound message"] --> G{"code gate: text, no audio, no reply-to product"}
    G -- "no" --> A["shopping agent"]
    G -- "yes" --> J["Jev: handler? store data? catalog keywords?"]
    J -- "direct_reply, confident, store data unlikely" --> F["fast lane: same model, no tools, no thinking"]
    J -- "browse_catalog + scored keywords" --> B["code: search the keywords, write the list"]
    J -- "anything else" --> A
    F -- "empty, failed, or claims a price" --> A
    B -- "nothing in stock, store down, no keywords" --> A
    F -- "usable text" --> R["reply to customer"]
    B -- "products found" --> R
    A --> R
```

## Why

The full agent turn sends a ~3.7 KB system prompt (persona + security rules +
CTA skill) plus seven tool schemas before the model says a word. The fast lane
sends ~1.4 KB and no tools. For a message that needs no store data that is the
whole cost of the turn, so greetings and acknowledgements get materially faster
and cheaper without changing what customers see.

Browsing goes further: the model is not called at all, so a "show me hoodies"
turn costs one Jev judgment and one Shopify search instead of a full agent turn
with tool round trips.

## Files

| File | Role |
|------|------|
| `runtime/src/agent-worker/turn-router.ts` | Config, candidate gate, Jev call, thresholds, log line |
| `runtime/src/agent-worker/search-query.ts` | Catalog keyword candidates, the Jev question, the query ladder |
| `runtime/src/agent-worker/direct-reply.ts` | The fast lane: one pi agent run, no tools, capped |
| `runtime/src/agent-worker/model-runtime.ts` | Provider/model/call wiring shared by both routes |
| `runtime/src/app/prompts.ts` | `DIRECT_REPLY_PROMPT` (fast lane) and `SHOPPING_AGENT_PROMPT` |
| `runtime/src/benchmark/router-probe.ts` | Run real messages through the judgment and print the routes |

## Browsing is code's job

"Show me shirts", "what hoodies do you have", "something for the beach" are
lookups, not conversations: Jev already picked the query, the catalog returns the
products, and the reply is a list. So code writes it
([browse-reply.ts](https://github.com/MalinrRuwan/chatAgent/blob/main/runtime/src/agent-worker/browse-reply.ts)):

- **No model generation at all.** A browse turn is one Jev judgment (shared with
  the routing decision, so no extra request) plus one Shopify search — no Gemini
  call, no tool round trips.
- **Nothing can be invented.** Names, prices and stock come verbatim from the
  store; a price range renders as a range; out-of-stock products are not offered,
  because a browse that lists a sold-out item sends the customer into a dead end.
- **The turn falls back to the agent** whenever code cannot answer it well: no
  keywords, no matches, nothing in stock, a store outage, a long message, or Jev
  unsure whether the customer means a category or one specific item.
- The list carries the store's photos for the products it names, and the
  add-to-cart CTA.

Query terms match their singular too (`hoodies` → `hoodie`), because a customer
typing the plural was silently getting nothing from a store whose titles are
singular — Shopify's own search stems that far, the title filter here did not.

## The judgment

One Jev request per candidate message carries the assistant's role and
capabilities, the last 8 conversation messages, and the new message. Three
questions are asked together (they run in parallel server-side):

1. `reply_mode` (**Choice**) — `direct_reply`, `browse_catalog`, or `agent_turn`.
   `agent_turn` includes "you are not sure"; `browse_catalog` is only for a kind
   of product, never for one particular item.
2. `needs_store_data` (**Noul**) — probability that answering needs the
   catalog, prices, stock, variants, policies, shipping, returns or the cart,
   or an action on the store.
3. `search_query` (**Choice**) — which candidate phrase should search the
   catalog (see below). Omitted when code found nothing to choose between, or
   when every candidate is a request word.

The fast lane requires **all** of:

- `reply_mode == direct_reply`
- `confidence >= LLM_ROUTER_MIN_CONFIDENCE` (default `0.6`)
- `needs_store_data <= LLM_ROUTER_MAX_STORE_PROB` (default `0.2`)
- the code gate passed (text, no audio, no reply-to product) **and** the message
  is at most `LLM_ROUTER_MAX_CHARS` long

Missing or malformed answers fail safe: unknown mode → agent, missing
confidence → 0, missing store-data probability → 1.

## Catalog search keywords

Shopify search only works well with short product-name phrases; a customer
sentence is a bad query. Code builds the candidate phrases, Jev picks one, and
the catalog tool searches it:

- **Candidates** come from the customer's message (runs of content words, then
  longest-first n-grams: `carnage retro short`, `carnage retro`, `retro short`,
  …) plus the product names in the assistant's previous reply, so a follow-up
  like "is it available in black?" still resolves to the product under
  discussion. Request words, sizes, prices and function words are stripped; the
  list is capped at 10 phrases.
- **The question** offers those phrases plus an explicit `(no catalog search)`
  option, so a social message is never forced into a product query.
- **The answer's probabilities rank every candidate**: the pick is the primary
  query, the runners-up become fallbacks. Below
  `LLM_ROUTER_MIN_QUERY_PROB` (default `0.25`) no phrase is trusted and the
  router produces no keywords at all.

Where the keywords are used:

| Consumer | Behaviour |
|----------|-----------|
| Speculative prefetch | The catalog lookup started in parallel with the model's first call now runs Jev's primary phrase (previously the raw customer text, and skipped entirely for messages over 80 characters). Skipped for cart/checkout/policy/small-talk turns. |
| `search_shop_catalog` | Searches Jev's phrase when the model's own query adds no term of its own (it is restating the customer); when the model adds a term — a broader category, an alternative after a miss — its query stays primary. Either way the other rung and Jev's runners-up are merged into the same lookup as extra attempts, so one tool call still answers both. |
| Result filtering | Unchanged: results must match the effective query's significant terms, so a fallback can add recall but never noise. |

With the router off (or no keywords), the tool behaves exactly as before: the
model's query is the query.

## Configuration

| Env | Default | Meaning |
|-----|---------|---------|
| `LLM_ROUTER_ENABLED` | `1` | Master switch. `0` sends every turn to the agent. |
| `TYPESAFE_API_KEY` | — | Jev credential. **Without it the judge is skipped** (route `not-configured`) and every turn runs the agent. |
| `TYPESAFE_BASE_URL` | — | Optional TypeSafe API root (proxy/self-hosted). |
| `LLM_ROUTER_TIMEOUT_MS` | `1500` | Per-attempt deadline. One attempt only — a slow judge must not stack retries onto a waiting customer. |
| `LLM_ROUTER_MAX_CHARS` | `160` | Longest message that may take the fast lane. |
| `LLM_ROUTER_MIN_CONFIDENCE` | `0.6` | Minimum Jev confidence in `direct_reply`. |
| `LLM_ROUTER_MAX_STORE_PROB` | `0.2` | Highest store-data probability that may still skip the agent. |
| `LLM_ROUTER_MIN_QUERY_PROB` | `0.25` | Minimum probability behind Jev's catalog keyword phrase before it is used as a query. |

All of them are in the control-plane allowlist (`runtime/src/app/control.ts`),
so they can be changed at runtime without a redeploy — use `/control/env`.

## Failure modes

The fast lane is an optimisation; it never owns correctness. Every path below
ends in a complete answer from the shopping agent:

| Situation | Route |
|-----------|-------|
| No `TYPESAFE_API_KEY`, or `LLM_ROUTER_ENABLED=0` | agent (`not-configured` / `router-off`) |
| Audio message, reply-to-product context, or empty message | agent (`not-a-candidate`), no Jev call |
| Jev unreachable, timed out, or unparsable | agent (`judge-error`) |
| Jev unsure, or store data likely needed | agent (`low-confidence` / `needs-store-data` / `agent-chosen`) |
| Message too long for a one-shot reply | agent (`long-message`) |
| Browsing with no usable keywords | agent (`browse-without-keywords`) |
| Browsing, but the search returns nothing in stock | agent (browse falls through) |
| Jev's keyword answer is flat, missing, or `(no catalog search)` | no keywords: the tool searches the model's query, as before |
| Fast lane returned nothing or a provider error | agent |
| Fast lane text claims a price/currency amount | agent — the fast lane has no store data, so such a claim is a hallucination |

## Observability

One line per turn, plus route fields on the reply timings:

```
[agent-worker] turn route=browse reason=browse jev=browse_catalog conf=0.99 storeProb=0.98 judgeMs=476 query="shirts" p=1.00 user=9477…
[agent-worker] browse reply for 9477… from Jev keywords "shirts": 2 product(s)
[agent-worker] turn route=agent reason=agent-chosen jev=agent_turn conf=1.00 storeProb=0.98 judgeMs=299 jevTokens=1017/143 query="carnage retro short" p=0.98 user=9477…
[agent-worker] catalog search query="carnage retro short" source=jev modelQuery="do you have the carnage retro short in black size M please?"
[agent-worker] direct reply unusable for 9477… (empty or unsafe text); falling back to the agent turn
```

`timings.route` (`direct` | `browse` | `agent`) and `timings.routerMs` are
returned to the actor, which logs them, so latency and cost can be split per
route.

## Verifying and tuning

If turns you expected on the fast lane run the agent, or a fast-lane reply
answers a product question, the thresholds need attention:

```bash
# Real messages through the real judgment (needs TYPESAFE_API_KEY).
TYPESAFE_API_KEY=ts_… bun run router:probe "thanks!" "ok" "what shirts do you have?"

# Existing thresholds are read from the same env vars as production.
LLM_ROUTER_MIN_CONFIDENCE=0.8 bun run router:probe "hey"
```

The probe prints the route code would take plus Jev's probabilities, which is
the evidence to move `LLM_ROUTER_MIN_CONFIDENCE` / `LLM_ROUTER_MAX_STORE_PROB`.
The unit tests (`runtime/test/turn-router.test.ts`, `runtime/test/direct-reply.test.ts`)
cover the policy and the fast lane's prompt/cap/thinking behaviour; they do not
judge Jev's quality — that is what the probe and the route logs are for.

A threshold can also be moved on the running service, without a redeploy — the
router keys are in the control-plane allowlist:

```bash
# Apply an override (payload is {"values": {...}}), then read it back
curl -X POST https://$DOMAIN/control/env \
  -H "authorization: Bearer $CONTROL_TOKEN" -H "content-type: application/json" \
  --data '{"values":{"LLM_ROUTER_MIN_CONFIDENCE":"0.8"}}'

# Back to the .env defaults
curl -X POST https://$DOMAIN/control/env/reset -H "authorization: Bearer $CONTROL_TOKEN"
```

On the deployment host the same thing is one command:

```bash
docker logs -f chatagent-runtime | grep -E "turn route=|catalog search"
```

---

Back to [Agent Architecture](/agent-architecture) · [Agent Lifecycle](/agent-lifecycle)
