"Self-hosted is cheaper" is true exactly up until the first outage that costs more engineering time than the managed-service bill you were avoiding would have. A production RAG system ran its vector database self-hosted, single-node, non-clustered, on Kubernetes — and in a three-month stretch it generated seven documented incidents, including one 43-hour search outage and one disaster-recovery restore that failed on the one occasion it was actually needed. The honest accounting here isn't "self-hosted vs. managed." It's "the configuration that was actually running vs. the configuration that was assumed to be running" — and those turned out to be two very different things, discoverable only by reading what the database was actually doing, not what the deployment manifest said it should be doing.
The outage that teaches the real lesson
The most severe incident is worth walking through in detail, because the root cause sits one layer below where almost anyone would think to look. The vector database's single-node deployment ran with an embedded metadata store rather than an external one — convenient, one fewer component to operate — but the configuration never overrode the metadata store's data directory away from its default. That default path was relative to the container's own working directory, which lands on the container's ephemeral writable layer, not on the mounted persistent volume.
The persistent volume was attached, had plenty of free space, and looked completely healthy on every dashboard that checked for it. It just wasn't holding the thing that actually mattered. It held mmap cache and scratch files; the metadata that maps every vector to its actual meaning lived somewhere that vanishes the instant the container is rescheduled onto a different node.
Assumed persistence boundary: container ──► persistent volume ──► survives reschedule
Actual persistence boundary: metadata ──► container's ephemeral layer ──► gone on reschedule
vectors ──► persistent volume ──► survives reschedule
The trigger was almost comically unrelated: a prior, seemingly unconnected cost-optimization pass had shrunk the node group's maximum size, so when this pod needed rescheduling, the cluster autoscaler had nowhere to place it. The pod sat Pending for over forty hours. When it finally got a node, the embedded metadata store started fresh — and the vector data sitting safely in object storage was now orphaned: intact, byte-for-byte correct, and completely unreachable, because the index that pointed to it was gone.
Recovery hit three compounding failures of its own, each one unrelated to the original bug: the pinned container image for the external metadata store the team stood up as a replacement had been removed from its registry during an unrelated vendor restructuring, forcing a same-day switch to a different base image; that replacement image had no shell, which broke an exec-based health probe that assumed one existed; and a Python dependency pinned transitively by the database client released a breaking version mid-recovery, breaking the client library at the worst possible moment. None of these three would have mattered on a normal day. All three mattered during an active outage, which is exactly when unrelated infrastructure fragility tends to surface — because it's the first time in months anyone has touched that code path under pressure.
The honest recovery cost wasn't just the 43 hours of downtime. It was roughly ten hours of ingestions between the last good backup and the eviction that were gone for good — recoverable only by re-ingesting from the original source documents, not from any backup.
"A backup you haven't restored from is a hypothesis"
A separate incident, on its own a minor footnote, turned out to be the highest-leverage lesson in the entire log: a disaster-recovery rehearsal — restoring a production backup into a scratch environment, specifically to verify the backup path still worked — simply failed. Not during a real outage. During a drill, which is the cheapest possible time to discover that your safety net has a hole in it.
That drill failure is what should have prevented the 43-hour outage from being as bad as it was, if it had been caught earlier. A backup job that exits with status zero tells you the job ran. It tells you nothing about whether the thing it produced can actually be restored.
# What existed: a backup job.
0 2 * * * /usr/local/bin/snapshot-vectordb.sh >> /var/log/snapshot.log
# What was missing: automated proof the backup is restorable.
0 3 * * 0 /usr/local/bin/restore-to-scratch-and-verify.sh || page_oncall
Adding that second line — restore to a scratch environment on a schedule, diff the result against expected row counts, page someone if it doesn't match — is the single change in this whole incident log that would have surfaced the problem weeks before an outage needed it, independent of every other fix made afterward.
A second, completely different failure class: write amplification nobody designed for
The third major incident has nothing to do with persistence at all, and it's a useful lesson in how an access-control decision made early can turn into an infrastructure incident much later, in a part of the system nobody was looking at. A bulk permission change — sharing a large folder containing thousands of documents — triggered the application layer to issue one HTTP call per document to update that document's access-control field on its vector rows. Two amplification effects compounded:
- Read amplification. Each update first queried the vector store for the row's every field, including the full embedding vector, just to read and then flip a single scalar access-control flag. A single query that only needed one small field was pulling back the entire high-dimensional vector along with it, every time.
- Write amplification. The vector store's upsert operation is a full delete-and-reinsert under the hood, not an in-place field update. Flipping one scalar on a row forced a complete rewrite of the full vector alongside it, for every single row touched.
Multiplied across however many worker processes, jobs per process, and batch size were running concurrently, this produced on the order of dozens to a hundred simultaneous full-row rewrites competing for the same memory-constrained write path — and the database fell over under memory pressure, for a bulk operation that was conceptually "update one small field on a few thousand rows."
The architectural lesson generalizes well past vector databases: access control should be resolved at query time through a lightweight lookup, never stored redundantly on every row that needs filtering. Keep access control as a side table or lookup keyed by document ID, joined or filtered at retrieval time, and a share operation becomes a write proportional to the number of documents — cheap, a metadata update. Store it denormalized on every vector row instead, and the exact same operation becomes a write proportional to the number of vectors times the size of everything stored alongside them — which is precisely the kind of amplification that looks fine in a demo with ten documents and falls over in production with ten thousand.
The failure mode that no configuration fix was going to solve
A later, separate incident looked at first like a repeat of the same memory problem, triggered this time by a routine cost-cutting pass that reduced the database's memory limit. It crashed immediately with an explicit out-of-memory error stating the corpus needed to load more memory than the configured limit allowed. The easy read is "the cost cut broke it." The accurate read, confirmed by checking the actual memory footprint against the original, pre-cut limit: the corpus needed more memory to load than even the original, more generous limit provided. The cut didn't cause the problem. It only exposed a pre-existing ceiling sooner than an unrelated event — a routine restart, a security patch, a host replacement — eventually would have anyway. This specific failure mode, memory exhaustion under concurrent write load on a single unclustered node, turned out to match a known, publicly documented limitation of that database's single-node deployment mode specifically — not a misconfiguration fixable from the operator's side, a structural property of running that mode at all.
Comparing the wrong two numbers
The original cost justification for self-hosting was straightforward: self-hosted compute cost versus a managed service's list price, self-hosted wins. That comparison was real, and also comparing the wrong two things. The actual comparison has to be between two self-hosted configurations, not between a self-hosted sticker price and a managed sticker price:
Config A — what was actually running:
single node, no replicas
backups: scheduled, never restore-tested
monitoring: "is the process up," nothing deeper
Config B — what reliability at this scale actually requires:
multi-node cluster with replication
backups: scheduled + automated restore-verification
monitoring: replication lag, disk headroom, backup verification status
Config A isn't "self-hosted." It's self-hosted without the operational investment that makes self-hosting a viable choice in the first place. Once the real engineering hours burned across seven incidents are priced in — one of them a 43-hour outage involving a failed disaster-recovery attempt at the worst possible moment — Config A was never actually the cheap option. It was deferred cost wearing the costume of avoided cost, and the bill arrives all at once, during the worst possible week, instead of showing up predictably on a monthly invoice.
There's a real, uncomfortable cost category underneath all of this that never appears on an infrastructure bill at all: every incident here produced its own postmortem, and the operational burden of keeping one stateful component alive self-hosted generated more internal documentation — scaling plans, operational runbooks, incident writeups — than the rest of the surrounding platform combined. That's engineering time with a real cost, it simply never shows up next to a dollar figure where anyone doing a cost comparison would think to look for it.
What actually has to be verified before moving, not assumed
The decision to move to a managed equivalent wasn't purely a cost-and-reliability call — it included an explicit capability-parity check before committing, because migrating onto a managed service that silently lacked a feature the system depended on would have forced a reversal, and discovering that after cutover is far more expensive than checking beforehand. The checklist covered hybrid sparse-plus-dense search support, native per-tenant partitioning so one tenant's data is physically isolated rather than merely filtered, specialized indexing for low-cardinality fields versus high-cardinality ones, and — flagged explicitly as the single highest-risk item, because failing it would have forced a reversal back to self-hosting — full parity for a specific language's text-analysis pipeline used in lexical search. All of it checked out before the migration proceeded, which is the point: verify the thing that would force you to reverse course before you've already committed, not after.
Moving to a managed, multi-tenant service also introduces a residual risk worth naming honestly rather than hand-waving: document content now lives inside a vendor's infrastructure rather than entirely within infrastructure the team directly controls. This is not actually a new category of exposure if embedding and reranking calls were already routing through third-party model providers — the same content was already leaving the perimeter for those calls — but it does mean that after the move, none of the stack is "fully at home" anymore, where some of it used to be. The practical mitigation was starting on a private-networking tier that keeps traffic off the public internet, while explicitly keeping a bring-your-own-cloud option on the table for whenever a customer's contractual requirements demand it, rather than discovering that requirement for the first time during a contract negotiation.
The migration was just normal ingestion, pointed somewhere else
One pragmatic detail worth keeping: the new schema needed fields — language, visibility, a tenant partition key — that couldn't be backfilled onto the existing collection no matter which vendor ended up hosting it. A full re-index was mandatory regardless of the self-hosted-versus-managed decision. That turned the migration into something simpler than a dedicated data-migration project: point the re-index that was already required directly at the new destination, run both systems live in parallel throughout, and treat the old deployment purely as a two-week rollback option rather than something that needed a careful, separate cutover plan.
The real lesson
Self-hosting a stateful system and paying for a managed version of the same system aren't "cheap" and "expensive." They're two different allocations of the same underlying cost — engineering time instead of a subscription line item — and the engineering-time allocation is easy to systematically underprice because it doesn't show up as a single number anywhere. A seven-incident stretch over three months isn't evidence that self-hosting is a bad idea in general. It's evidence that a single unreplicated node with an unverified backup path and no read/write separation is a bad idea regardless of who operates the software — and that the real price of that specific configuration shows up later, all at once, usually during the one week you can least afford it.

