Most RAG prototypes start the same way: a script that reads a file, chunks it, embeds the chunks, and writes vectors into an index. It works great in a demo. It falls over in production, usually within the first few weeks of real traffic — not because the embedding model is wrong, but because the architecture never separated two workloads that behave nothing alike: ingesting content and answering queries against it.
This is the architecture I've converged on after rebuilding a RAG platform from a single-pipeline design into something that survives multi-tenant production load. I'm calling it RAG architecture 2026 here, since the point isn't a specific release — it's a set of structural decisions that any team building a serious retrieval platform in this generation of tooling will probably end up making.
The failure mode that forces the redesign
The original design had a simple shape: one message in a queue equals one document equals one job. A worker pulls a file from storage, parses it, chunks it, embeds it, and upserts it into the vector store — all inside a single function call, on a single queue, with no boundary between stages. It works fine for a 3-page memo. It breaks in three specific ways once real documents show up:
- A 500-page scanned PDF turns into one 40-minute job. There is no timeout you can set that's both safe for a 3-page file and tolerant of a 500-page one. Whatever you pick, something times out.
- Partial failure has no vocabulary. If page 240 of 500 fails to parse, the only two states available are "the whole document succeeded" or "the whole document failed." Neither is true.
- Retrying means retrying everything. A transient failure on chunk 480 out of 500 forces you to redo pages 1 through 479 as well, because the unit of retry is the whole document, not the unit that actually failed.
The fix isn't a bigger worker or a longer timeout. It's shrinking the unit of work until retry granularity matches failure granularity, and separating the two workloads — ingestion and querying — that were never structurally alike in the first place.
Two planes with opposite performance profiles
Ingestion is:
- Asynchronous — nobody is blocking on it in real time.
- Throughput-bound — the goal is to process the most documents per hour at the lowest cost.
- Bursty and fan-out heavy — one upload can turn into hundreds of downstream work units (pages, segments, chunks, embeddings).
- Tolerant of partial failure — if 2 out of 30 segments in a document fail, that's a data-quality issue, not an outage.
Querying is:
- Synchronous — a human or an agent is waiting on the response.
- Latency-bound — p99 matters more than total throughput.
- Fan-in, not fan-out — one query reads from a huge, mostly-static index.
- Binary in the moment — a query either returns a usable, grounded answer within budget or it doesn't.
Forcing both workloads through the same metrics, the same queue, and the same worker pool hides exactly where the system is breaking, because the two have categorically different bottlenecks: worker saturation on one side, p99 latency on the other. When a single pipeline tries to serve both, ingestion jobs get starved by latency-sensitive query traffic sharing the same connection pool, or query paths get stuck behind a backlog of embedding jobs competing for the same accelerator queue. The fix is splitting the system into two planes that scale, fail, and get operated independently.
An eight-layer decomposition
flowchart TB
subgraph INGEST["INGESTION PLANE — async, throughput-bound"]
L1[L1 Admission & Intake]
L2[L2 Triage & Routing]
L3[L3 Extraction Swarm]
L4[L4 Semantic Construction]
L1 --> L2 --> L3 --> L4
end
subgraph SHARED["SHARED SEAM"]
IDX[(Vector store + metadata store)]
end
subgraph QUERY["QUERY PLANE — sync, latency-bound"]
L5[L5 Query Understanding & Access Resolution]
L6[L6 Retrieval & Evidence Assembly]
L7[L7 Generation & Egress Guardrails]
L5 --> L6 --> L7
end
L0[L0 Control Plane & Exception Bus<br/>wraps every layer]
L4 --> IDX --> L6
L0 -.-> L1
L0 -.-> L2
L0 -.-> L3
L0 -.-> L4
- L0 — Control Plane & Exception Bus. A sidecar concern that wraps every other layer: a job state machine in Postgres, a reconciler that sweeps stalled jobs, admission quotas with ETA responses, a centralized exception classifier, and document lifecycle management. No individual layer decides a document's fate — it throws a classified error and L0 decides what happens next.
- L1 — Admission & Intake. Validate the upload, content-hash dedup against what's already indexed, persist the job row to Postgres, and only after that commit publish to the queue. Publishing before the commit is a classic bug: a worker can pick up the message before the row exists to update.
- L2 — Triage & Routing. Cheap, no-model signal extraction — text coverage, image density, page count, table complexity, language — picks a lane (plain document, structured spreadsheet, visual/scanned) and a parser engine, then splits the document into segments.
- L3 — Extraction Swarm. Fans out N segments in parallel. Each segment's output is scored for quality and escalated to a stronger engine if the score is poor — not just on exception, on measured bad output. A dedicated sub-lane handles spreadsheets through a fingerprint → discovery → sandbox → validation → columnar-artifact pipeline (its own full escalation ladder, covered in a separate post).
- L4 — Semantic Construction. Assembles segments back into a single document — this is the only place absolute page offsets are known, after concatenation — then does hierarchical parent/child chunking, contextual enrichment (prepending file/folder/section-hierarchy headers so a chunk is understandable out of context), batch embedding, batch indexing, and a verify step that counts actual chunks landed against the expected count before flipping status to "ready." Expensive per-chunk summarization, when needed, runs as a separate sub-stage specifically so an LLM call never sits on this stage's critical path.
- L5 — Query Understanding & Access Resolution. Builds the access-control filter expression from the caller's identity before any retrieval happens, routes to the right retrieval lane, and optionally expands the query.
- L6 — Retrieval & Evidence Assembly. Hybrid search (dense + sparse) fused to a shortlist, reranked down to a handful, expanded back out to full parent chunks, then an access-control post-check that re-verifies the final candidates against the database — not just the pre-filter, because a pre-filter can go stale between index time and query time. A confidence gate does exactly one rewrite-and-retry before returning "insufficient grounding" instead of letting the model hallucinate past a weak retrieval.
- L7 — Generation & Egress Guardrails. An access-scoped semantic cache (unknown-provenance answers get no semantic cache — exact-match only, never fuzzy), generation with cost tracking and PII redaction, a grounding check that traces every factual claim back to a retrieved chunk with exactly one regeneration attempt on failure, and a mandatory citation schema.
The design rule underneath all eight layers: a layer boundary is a retry boundary, and a layer boundary is a measurement boundary. Since retry granularity equals work-unit size, the single biggest lever is making L3 operate on segments (roughly 10–25 pages) rather than whole documents. That's what turns a 500-page PDF into 25 parallel five-minute jobs instead of one 40-minute job that can't fit any sane timeout and can't be partially retried.
The state machine, not a boolean
Document status moves strictly forward through queued → admitted → triaging → parsing → assembling → chunking → embedding → indexing → verifying → ready, with branch states that exist specifically to stop the system from lying about partial outcomes:
| State | Meaning |
|---|---|
partial | Some segments failed, the rest indexed. A flag, not a terminal state — a document can be ready and partial at once. |
needs_review | Escalation exhausted, routed to a human queue. Does not serve queries yet. |
unsupported | Zero chunks after every available engine. Terminal, with a user-facing explanation instead of a silent disappearance. |
inconsistent | The verify step's count mismatched expectations; the reconciler picks it up for inspection. |
superseded | An older version after a re-ingest of the same source. |
archived | Soft-deleted; chunks are retained for restore. |
A document that finishes with 28 of 30 segments embedded is partial, never silently rounded to complete or failed. That distinction matters to three different consumers: retrieval should still search the 28 that succeeded instead of hiding the whole document; an operator dashboard needs partial to show up as "retry 2 segments," distinct from genuine failures that need a human; and a usage or billing system that counts "documents processed" has to decide explicitly whether partial counts, instead of a rounding error nobody can audit.
Fan-in without a watchdog
Tracking "did all N segments of this document finish" without polling is one atomic statement:
UPDATE ingestion_job
SET segments_done = segments_done + 1
WHERE id = $1
RETURNING segments_done, segments_total;
Only the caller that observes segments_done == segments_total in the row it just updated is allowed to trigger the next stage. That single invariant — exactly one caller ever sees "I am the last one," no matter how many workers race or how many times a message gets redelivered — is what the entire fan-in correctness rests on. No cron job polling for "is this document done yet," no separate watchdog process; the database's own atomicity gives you the coordination for free.
Why there's no heartbeat
An earlier version of this system did have a heartbeat, and it caused real double-billing. A job could run long enough to outlive any plausible lease, so ownership had to be continuously re-asserted by a background heartbeat. The failure mode: a CPU-bound parsing step blocked the event loop long enough that the heartbeat stopped firing, a second worker concluded the job was abandoned, reclaimed it, and the same document got processed — and billed — twice.
The fix wasn't a smarter heartbeat. It was sizing the unit of work so a stage simply finishes before its lease expires, collapsing two independent clocks (a Postgres lease_until column and the queue's own visibility timeout) into one invariant: every lease duration is shorter than its stage's timeout, and every queue visibility timeout is longer than the lease. The IAM policy for queue consumers deliberately omits the permission to extend a message's visibility window — not an oversight, a guardrail against anyone quietly reintroducing heartbeat-style lease extension and recreating the exact bug this design removed.
Classifying exceptions centrally instead of per-stage
Every failure gets mapped to exactly one of six branches, decided by L0, never by the stage that threw it:
transient network blip, 429, 5xx → retry same stage, exponential backoff + jitter
escalate output quality below gate → retry with a DIFFERENT engine, max 2 tiers
degrade one segment failed, most OK → mark that unit failed, continue → doc ends up partial
human automation exhausted → needs_review queue, does not serve queries
poison no engine produced output → terminal `unsupported`, user-facing explanation
capacity over quota / queue too deep → explicitly NOT an error — stays `queued`, returns an ETA
Centralizing this is what lets you change retry policy in one place instead of hunting down dozens of scattered try/except blocks that each made a slightly different, undocumented decision about what "failed" means.
Six worker pools, not eight layers
The eight logical layers don't map one-to-one onto deployment. In practice the system runs as six worker pools, because the pool boundary has to equal the queue boundary — an autoscaler that scales a pool from one queue's depth can't scale correctly if two unrelated stages share that queue.
| Pool | Stages | Bounded by | Concurrency / pod | Visibility timeout |
|---|---|---|---|---|
| light | triage, assemble, chunk, verify | database + storage round-trips | 16 | 660s |
| parse | parse | memory (structure-preserving parsers peak in GB) | 2 | 960s |
| sheet | spreadsheet normalize + assemble | wall-clock of a model+sandbox call | 4 | 960s |
| embed | embed | embedding-provider rate limit | 8 | 360s |
| index | index | vector-store write capacity | 4 (deliberately less than embed) | 360s |
| summary | summarize, render | nothing — nobody blocks on it | 4 | 480s |
Two details worth calling out. First, index runs at lower concurrency than embed on purpose — deliberate backpressure, because the vector store's write capacity is the tightest resource in the whole pipeline, and letting index concurrency match or exceed embed concurrency just moves the bottleneck downstream without fixing it. Second, these pools are never merged even though that would simplify deployment: a 15-minute parse job would head-of-line-block a 200ms status write if they shared a queue, scaling to survive a parsing burst would multiply concurrent writes into the one component that can't absorb them, and a single out-of-memory event in one pool would otherwise take down unrelated in-flight work in another.
Batch sizing follows the same "respect the real constraint" logic: chunk batches for embedding settled at 96, not a rounder number like 128, because 96 is the embedding provider's actual per-call cap. At 128, every batch needed two API calls — one full, one a third full — on the one pool whose entire cost model is round-trips, and a retry after a partial failure risked double-sending the first call.
Three tables, and the queue is transport-only
All ingest state lives in Postgres, never in the queue message itself:
- jobs — one row per document version: routing decision, triage signals, fan-out counters (segments expected/done/failed, batches expected/done), ownership (
attempt,lease_until), and the document metadata — file name, folder path, access-control principals, visibility — that gets written onto every chunk at index time and can never be backfilled later without a full re-index. - segments — one row per page range (or per spreadsheet sheet, same table, reused). Tracks
engine_usedand a counter for engine escalations that is kept separate from the transient-retry counter — conflating the two would burn a document's entire retry budget before it even got a chance to try a different engine. - batches — one row per chunk batch, spanning both the embed and the index stage, so fan-in only has to track one counter instead of two that can drift apart.
The queue's only job is to move a reference to work from one pool to the next. Every fact about whether that work succeeded, how many times it's been tried, and what should happen next lives in a table that survives a redeploy, a dead consumer, or a message that got redelivered three times before anyone noticed.
What this buys in production
Once ingestion and query are separate planes with an explicit, ratio-aware handoff at the vector-store seam: embedding workers can scale up for a bulk backfill without touching the query service's autoscaling policy at all. A query-path incident — a bad reranker deploy, a slow model endpoint — can't starve ingestion, because the two no longer share a worker pool or a queue. An ingestion backlog during a large onboarding doesn't degrade search latency for existing tenants, because the query plane only ever reads from the shared index and never waits on the ingest plane directly. And partial-failure data becomes an operational signal instead of noise, because it's a first-class state instead of a rounding error forced into a boolean that was never going to hold.
None of the eight layers is novel in isolation — extraction, chunking, and embedding are textbook RAG stages. The part that's easy to skip, and expensive to retrofit later, is drawing a hard boundary between the plane that produces the index and the plane that reads it, sizing the unit of retry to match the unit of failure, and giving fan-out, partial failure, and exception handling their own explicit vocabulary instead of squeezing all three into a boolean that was never going to hold.

