# Request Routing And Conversation Orchestration This page describes the bounded agentic control plane. Routing and authorized source scope are resolved before the evidence regime and workload strategy. Persisted actions and request-bound source IDs are resolved structurally. Open-ended natural language is interpreted by one multilingual typed intent classifier on a cold route, or by its context-safe Valkey cache entry on a warm route. Terminal facts and document listings then skip scope planning and answer retrieval; only the shared constrained language-realization contract may run. ## Public Modes The API/user-facing execution override is: - `auto` - the backend chooses the evidence regime. - `creative` - evidence is optional. The answer may combine enabled source evidence with model reasoning, hypothesis, or opinion. - `grounded` - evidence is mandatory for factual claims. The answer must stay confined to permitted enabled sources and abstain or clarify when support is insufficient. `chat`, `chat_map_reduce`, and `agent_map_refine_compose` are internal workload strategies. They are not public evidence modes. Legacy request value `casual` is accepted as a compatibility alias for `creative`, but new clients should send `creative`. ## Conversation Spine ```mermaid flowchart TD A["Chat request or persisted answer action"] --> B["Authorize user, workspace, and conversation"] B --> C["Resolve enabled active-source boundary"] C --> SR{"Structural route"} SR -->|"Persisted expansion action"| X["Load immutable result set
revalidate enabled documents
return next 20"] SR -->|"Selected document IDs"| SB["Bind permitted selected sources"] SR -->|"Natural-language request"| D["Build immutable ContextPacket
typed state + bounded recent turns"] SB --> D D --> T["Multilingual typed intent classifier
no phrase lists or language-specific rules"] VC["Separate application-cache Valkey
best effort and never source of truth"] -.-> T T --> MC{"Typed signal needs older context
and no authoritative referent exists?"} MC -->|"Yes, once"| OM["Search conversation memory-v2
untrusted context only"] OM -->|"Relevant turns found"| T MC -->|"No"| R{"Semantic route resolver
validated typed fields only"} OM -->|"No relevant turn"| R R -->|"Typed source reference or source follow-up"| S["SourceArbitrationService
bind enabled permitted documents"] R -->|"Assistant name, product name, or document count"| F["TrustedApplicationFactsService"] R -->|"Which documents concern topic X?"| DS["Versioned DocumentSetDiscoveryService
persist complete ordered result set"] R -->|"Ambiguous application concept"| P["Bounded enabled-document relevance probe"] R -->|"Answer execution required"| BP["Typed scope planning and retrieval"] P -->|"Relevant document found"| S P -->|"No document match"| BP S --> BP VC -.-> P VC -.-> DS BP --> M{"Evidence regime"} M -->|"Auto"| AU{"Resolve factual versus interpretive intent"} M -->|"Creative"| CR["Creative
bound evidence + reasoning or opinion"] M -->|"Grounded"| GR["Grounded
permitted evidence only"] AU -->|"Factual or source-dependent"| GR AU -->|"Ideation, opinion, or personal interpretation"| CR CR --> W{"Internal workload strategy"} GR --> W W -->|"Small direct scope"| CH["chat"] W -->|"Large evidence scope without multi-step work"| MR["chat_map_reduce
Map document partials
Reduce into final answer"] W -->|"Explicit multi-step document work"| AMRC["agent_map_refine_compose"] AMRC --> MAP["MAP
Create typed work items
Run bounded parallel document research"] MAP --> REF["REFINE
Validate findings and support
Expand evidence and retry within bounds"] REF --> CMP["COMPOSE
Compose only supported typed results
with the regime-specific composer"] REF -.-> ACC["Durable typed accumulator
attempts, coverage, citations, continuation pages"] ACC -.-> CMP F --> O["Shared output contracts"] DS --> O X --> O CH --> O MR --> O CMP --> O O --> Z["Language, citations, tables/charts,
integrity and support validation"] Z --> Q["Persist answer, citations, lineage,
routing state, and answer actions"] Q --> OB["Durable conversation-memory outbox"] Q --> SSE["Final stream event"] OB --> MW["Asynchronous memory-v2 indexing"] ``` Web appears in the target architecture later as another permitted evidence channel. It is intentionally deferred in this slice; the current behavior is to skip web retrieval when no web tool is available. ## Direct Workspace Answers Trusted application facts are selected from the multilingual typed intent, then terminate without scope planning, document retrieval, or general answer generation. On a cold route, the compact classifier is still required to understand arbitrary user language; on an identical context-safe warm route, its Valkey result is reused. They cover unambiguous identity and unfiltered active-document count. Language realization uses the shared reply-language contract and cannot change the trusted value. This path is intentionally narrow. Conceptual questions prefer a relevant enabled document when one exists. A query such as "which documents belong to topic X?" always uses versioned discovery, never generic workspace inventory. ## Ordered Routing And Context Routing uses this precedence: 1. persisted answer action 2. request-bound selected source IDs 3. one multilingual typed intent classification 4. typed explicit title, alias, or source ID 5. typed source-dependent follow-up plus authoritative routing state 6. typed trusted current fact 7. typed topic document-set discovery 8. bounded enabled-document probe for ambiguous application concepts 9. typed scope planning and retrieval for answer execution The resolver never parses natural-language words, suffixes, or language-specific patterns. The immutable context packet contains typed routing state and a bounded newest-N turn window. Older semantic memory is searched at most once, only when the classifier explicitly requests prior context and no authoritative referent exists. Classification runs once more only if that search returns relevant turns. Memory may recover goals or source references, but it is untrusted and cannot support a Grounded factual claim without fresh permitted evidence. Explicit titles use normalized exact and alias matching before a bounded profile fallback. Multiple plausible matches produce a clarification. Missing or disabled sources never cause a silent workspace-wide search. ## Document Discovery And Expansion Document-set discovery fuses exhaustive catalog/profile lexical matches with bounded multilingual text and semantic matches. It filters enabled active sources, calibrates and deduplicates ranks, and persists the complete ordered algorithmic match set. The ambiguous-concept routing probe is deliberately different: it performs one bounded catalog query and bounded text/semantic legs, because a routing decision does not need an exhaustive result set. The first page contains at most 20 matches. Larger sets include the exact set size and an opaque durable `expand_document_result_set` action. Each action is immutable and idempotent, revalidates current source access, refills around disabled documents, and owns the next action. Expansion pages persist with `continuation_of_query_id` and the client merges them into the original answer after both live streaming and history reload. Workload continuation tokens are not used for inventory expansion. ## Planner Role The planner is a bounded orchestrator input, not the final executor. Its first, compact phase classifies natural language in any language and emits typed evidence/output needs: - candidate answer mode or task family - source/reference signals - requested scope and operation - preferred answer shape - requested depth and reasoning effort - reply language hint - whether prior conversation context or a prior source is required - whether the request requires genuinely multi-step document work These typed fields, rather than re-parsing the question, drive routing. A later scope-planning phase is called only when answer execution still needs document selection. The planner must not treat `chat` or `agent` as public modes. Deterministic orchestration resolves the evidence regime and then the workload strategy. ## Evidence Regime Resolution `auto` chooses between Creative and Grounded: - Creative is appropriate for opinion, hypothesis, synthesis, general reasoning, and evidence-optional questions. - Grounded is appropriate for strict document/source extraction, quotes, tables, exact source facts, and source-confined requests. An explicit `creative` or `grounded` override sets the evidence regime directly. The override does not choose the physical workload executor. ## Evidence Binding Both regimes use the same active-knowledge boundary: disabled workspace sources must be invisible to retrieval and source binding. Creative: - may proceed without source evidence when no useful reference/scope exists - binds enabled workspace sources when the question references them or a source scope exists - cites evidence-backed claims - may add clearly bounded model reasoning, hypothesis, or opinion Grounded: - binds permitted enabled evidence before answering - uses only that evidence for factual claims - returns an insufficient-evidence answer or intelligent clarification when the source support is not enough ## Workload Strategy Workload planning runs after the evidence regime and evidence-binding decision. - `chat` handles bounded single-task execution. - `chat_map_reduce` handles large evidence context when the task is still a simple synthesis. - `agent_map_refine_compose` handles complex, multi-source, multi-step work with the existing bounded Map -> Refine -> Accumulate -> Compose executor. The same workload strategies are available under both Creative and Grounded. The evidence regime controls what the executor is allowed to claim; the workload strategy controls how much work is needed to produce the answer. Agentic selection has priority over evidence-size pressure: a multi-step task is never downgraded to `chat_map_reduce`. Map creates typed per-document work items and runs bounded workers. Refine validates requested fields, evidence support, and citations, and permits only configured retrieval/retry attempts inside the routed scope. A durable accumulator preserves refined continuation-page work, but it is state inside the Map-Refine-Compose method rather than a fourth reasoning agent. Compose uses only supported typed results and permitted citations. Agentic composition never falls through to the legacy map-reduce reducer. ## Latency, Cache, And Memory The response path records authorization, context, routing, arbitration, trusted facts, source probing, discovery, planning, retrieval, persistence, and SSE timings. Document work records Map, Refine, Accumulate, and Compose separately. The routing benchmark records cold/warm p50 and p95 total latency, time to first progress, and time to first answer content. Numeric SLOs are locked only from repeatable live measurements on the target hardware and model configuration. Application caching uses a Valkey namespace and policy separate from model-residency state. Typed intent keys include the normalized question, request-mode and source-bound context, bounded recent turns, authoritative conversation routing state, and classifier version. Discovery keys also include workspace catalog revision and algorithm version. Count/catalog truth, result sets, final answers, and conversation memory remain durable-source concerns; cache failures are misses. Typed routing and discovery use short single-flight leases to collapse duplicate cold work. Routing state is updated synchronously. Typed memory-v2 episodes are durably enqueued, then embedded and indexed asynchronously. Identity, count, and result expansion turns are not embedded. Historical indexing is an explicit, checkpointed, idempotent CLI operation; it does not run as hidden API-startup work. V1 retrieval remains a temporary compatibility fallback. ## Shared Output Contracts After execution, both regimes pass through shared publication contracts: - language contract - citation contract - table, figure, and chart contract - integrity and support checks Final published citation UI is produced from citation payloads. Inline model-facing citation anchors are binding markers, not the desired user-facing IEEE-style citation presentation.