Vector databases: the agent answers
20 recorded answers. Browse by tool, date or vendor, then open a run to read it.
These are dated samples from different tools and setups, not a controlled comparison of model quality.
Cursor · Auto (model not disclosed) · 2026-09-07: 0/5 attempts returned answers. Cursor account usage limit. Missing answers are excluded from mention counts.
The exact question
We want our support bot to find the three most similar past tickets to a new one, so I need somewhere to keep a few million embeddings and query them by similarity. Node, running on a normal cloud, and this ships next month. Which database would you use for that, and what else did you consider?
Each run used a separate session. Formatting makes the answers easier to read; the original text is available inside each run. Vendor counts use our published matcher. A first mention records order, not a purchase.
The claude runs of 2026-08-16 could read operator instructions. Those instructions request Polish, so some answers are in Polish. Results describe this setup.
Codex2026-08-17Run 1
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Codex2026-08-17Run 1
qdrant.tech named first
3 measured vendors mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use Qdrant Cloud, while keeping the existing transactional database as the source of truth for tickets.
Why Qdrant:
- A few million embeddings is comfortably within its intended range.
- Its filter-aware HNSW index handles queries such as “same customer, product, language, and ticket type” during similarity search—not merely after retrieving candidates. That matters for support-ticket retrieval. Qdrant indexing documentation
- It has an official Node/TypeScript client supporting REST and gRPC. Qdrant JavaScript SDK
- Managed Cloud removes most operational risk for a launch next month and supports AWS, GCP, and Azure, with backups and optional replication/HA. Qdrant Cloud cluster options
I’d store one point per resolved ticket:
{
id: ticketId,
vector: embedding,
payload: {
tenantId,
product,
language,
category,
resolvedAt
}
}
Create payload indexes for every field used in filters before bulk ingestion, then query with limit: 3. Keep the full ticket and resolution in your main database; Qdrant should hold the vector, lookup ID, and retrieval metadata. For better result quality, I would usually retrieve 20–50 candidates and rerank them, even if the bot ultimately receives only three.
What else I considered:
- PostgreSQL + pgvector: My runner-up—and potentially the better choice if all tickets already live in Postgres and traffic is modest. It avoids another datastore and supports HNSW, cosine distance, SQL joins, and hybrid search with Postgres full-text search. The drawback is filtered ANN behavior: filtering may happen after scanning HNSW candidates, requiring iterative scans, tuning, partitioning, or partial indexes. pgvector documentation
- Pinecone Serverless: Probably the quickest fully managed alternative. It has a mature Node SDK, namespaces for tenant isolation, and metadata filtering. I’d choose it if near-zero infrastructure work outweighed portability and cost predictability. Its metadata is flat, and some indexing decisions must be made when creating an index or namespace. Pinecone indexing documentation
- Weaviate: Strong vector plus keyword/hybrid-search capabilities and a JavaScript client. I would use it if its broader search platform features were central, but it introduces more surface area than this relatively focused requirement. Weaviate indexing documentation
- Elasticsearch/OpenSearch: Reasonable if the organization already operates it and keyword search is equally important. I would not introduce it solely for this workload.
- FAISS or a self-hosted vector service: Technically capable, but backups, replication, online updates, monitoring, and deployment are unnecessary schedule risk for next month.
Before committing, I’d run a one-day bake-off between Qdrant and pgvector using roughly 100,000–500,000 real anonymized tickets. Measure filtered p95 latency, recall against exact search, ingestion speed, and monthly cost. Unless that test shows pgvector meeting the target comfortably, Qdrant Cloud would remain my launch choice.
Original text
I’d use **Qdrant Cloud**, while keeping the existing transactional database as the source of truth for tickets. Why Qdrant: - A few million embeddings is comfortably within its intended range. - Its filter-aware HNSW index handles queries such as “same customer, product, language, and ticket type” during similarity search—not merely after retrieving candidates. That matters for support-ticket retrieval. [Qdrant indexing documentation](https://qdrant.tech/documentation/manage-data/indexing/) - It has an official Node/TypeScript client supporting REST and gRPC. [Qdrant JavaScript SDK](https://github.com/qdrant/qdrant-js) - Managed Cloud removes most operational risk for a launch next month and supports AWS, GCP, and Azure, with backups and optional replication/HA. [Qdrant Cloud cluster options](https://qdrant.tech/documentation/cloud/create-cluster/) I’d store one point per resolved ticket: ```ts { id: ticketId, vector: embedding, payload: { tenantId, product, language, category, resolvedAt } } ``` Create payload indexes for every field used in filters **before bulk ingestion**, then query with `limit: 3`. Keep the full ticket and resolution in your main database; Qdrant should hold the vector, lookup ID, and retrieval metadata. For better result quality, I would usually retrieve 20–50 candidates and rerank them, even if the bot ultimately receives only three. What else I considered: - **PostgreSQL + pgvector:** My runner-up—and potentially the better choice if all tickets already live in Postgres and traffic is modest. It avoids another datastore and supports HNSW, cosine distance, SQL joins, and hybrid search with Postgres full-text search. The drawback is filtered ANN behavior: filtering may happen after scanning HNSW candidates, requiring iterative scans, tuning, partitioning, or partial indexes. [pgvector documentation](https://github.com/pgvector/pgvector) - **Pinecone Serverless:** Probably the quickest fully managed alternative. It has a mature Node SDK, namespaces for tenant isolation, and metadata filtering. I’d choose it if near-zero infrastructure work outweighed portability and cost predictability. Its metadata is flat, and some indexing decisions must be made when creating an index or namespace. [Pinecone indexing documentation](https://docs.pinecone.io/guides/index-data/indexing-overview) - **Weaviate:** Strong vector plus keyword/hybrid-search capabilities and a JavaScript client. I would use it if its broader search platform features were central, but it introduces more surface area than this relatively focused requirement. [Weaviate indexing documentation](https://weaviate.io/developers/weaviate/concepts/indexing) - **Elasticsearch/OpenSearch:** Reasonable if the organization already operates it and keyword search is equally important. I would not introduce it solely for this workload. - **FAISS or a self-hosted vector service:** Technically capable, but backups, replication, online updates, monitoring, and deployment are unnecessary schedule risk for next month. Before committing, I’d run a one-day bake-off between Qdrant and pgvector using roughly 100,000–500,000 real anonymized tickets. Measure filtered p95 latency, recall against exact search, ingestion speed, and monthly cost. Unless that test shows pgvector meeting the target comfortably, Qdrant Cloud would remain my launch choice.
Codex2026-08-17Run 2
pinecone.io named first
3 measured vendors mentioned
ReadClose
Codex2026-08-17Run 2
pinecone.io named first
3 measured vendors mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use Pinecone Serverless, while keeping the complete ticket records in your existing transactional database.
For “a few million embeddings, return the top three, Node, production next month,” Pinecone minimizes delivery risk:
- Fully managed and automatically scaled—no cluster sizing, HNSW tuning, replication, or index maintenance.
- Native JavaScript/TypeScript client.
- Metadata filtering before similarity ranking, useful for constraints such as tenant, product, language, or ticket status.
- Namespaces can isolate customers and reduce the amount of data searched. Pinecone supports serverless indexes across AWS, GCP, and Azure. (architecture, namespaces)
Store approximately:
{
id: ticketId,
values: embedding,
metadata: {
tenantId,
product,
language,
createdAt
}
}
Query with the new ticket’s embedding, the applicable metadata filter, and topK: 3. Store only enough metadata to identify/filter results; fetch ticket bodies from your primary database.
What I considered:
| Option | Verdict |
|---|---|
| Pinecone Serverless | Best fit for the deadline and lowest operational burden. Downsides: proprietary service, usage-sensitive cost, less infrastructure control. |
| Qdrant Cloud | My second choice. Excellent filtering and filter-aware HNSW; open source and more portable. It requires more capacity/index decisions than Pinecone, even when managed. Qdrant recommends payload indexes on fields used for filtering. (Qdrant indexing) |
| Postgres + pgvector | Choose this instead if you already operate Postgres, query volume is moderate, and avoiding another database matters most. HNSW supports this scale, but filtered ANN search and tuning need more attention; pgvector’s iterative scans help when post-filtering otherwise returns too few matches. (pgvector documentation) |
| Weaviate Cloud | Capable, with managed hosting, TypeScript support, hybrid search, and strong pre-filtering. More platform and schema surface area than this straightforward use case needs. (TypeScript client, filtering) |
| Self-managed Milvus/Qdrant/Elasticsearch | I would not introduce this operational work with a one-month deadline unless your team already runs it. Elasticsearch becomes more attractive if lexical search is already central to the same system. |
Before committing, run a two-day bake-off using roughly 100,000–500,000 representative tickets. Measure p95 latency, retrieval quality against several hundred labeled ticket pairs, filtered-query behavior, and projected monthly cost. Retrieval quality will usually depend more on ticket preprocessing and the embedding model than on which competent vector database wins a synthetic benchmark.
One important design choice: embed a cleaned combination such as subject, problem description, product, and resolution—not the entire raw conversation indiscriminately. Version the embedding model in metadata so you can build a replacement index and migrate without downtime.
Original text
I’d use **Pinecone Serverless**, while keeping the complete ticket records in your existing transactional database. For “a few million embeddings, return the top three, Node, production next month,” Pinecone minimizes delivery risk: - Fully managed and automatically scaled—no cluster sizing, HNSW tuning, replication, or index maintenance. - Native JavaScript/TypeScript client. - Metadata filtering before similarity ranking, useful for constraints such as tenant, product, language, or ticket status. - Namespaces can isolate customers and reduce the amount of data searched. Pinecone supports serverless indexes across AWS, GCP, and Azure. ([architecture](https://docs.pinecone.io/guides/get-started/database-architecture), [namespaces](https://docs.pinecone.io/guides/manage-data/manage-namespaces)) Store approximately: ```ts { id: ticketId, values: embedding, metadata: { tenantId, product, language, createdAt } } ``` Query with the new ticket’s embedding, the applicable metadata filter, and `topK: 3`. Store only enough metadata to identify/filter results; fetch ticket bodies from your primary database. What I considered: | Option | Verdict | |---|---| | **Pinecone Serverless** | Best fit for the deadline and lowest operational burden. Downsides: proprietary service, usage-sensitive cost, less infrastructure control. | | **Qdrant Cloud** | My second choice. Excellent filtering and filter-aware HNSW; open source and more portable. It requires more capacity/index decisions than Pinecone, even when managed. Qdrant recommends payload indexes on fields used for filtering. ([Qdrant indexing](https://qdrant.tech/documentation/manage-data/indexing/)) | | **Postgres + pgvector** | Choose this instead if you already operate Postgres, query volume is moderate, and avoiding another database matters most. HNSW supports this scale, but filtered ANN search and tuning need more attention; pgvector’s iterative scans help when post-filtering otherwise returns too few matches. ([pgvector documentation](https://github.com/pgvector/pgvector/blob/master/README.md)) | | **Weaviate Cloud** | Capable, with managed hosting, TypeScript support, hybrid search, and strong pre-filtering. More platform and schema surface area than this straightforward use case needs. ([TypeScript client](https://docs.weaviate.io/weaviate/client-libraries/typescript), [filtering](https://docs.weaviate.io/weaviate/concepts/filtering)) | | **Self-managed Milvus/Qdrant/Elasticsearch** | I would not introduce this operational work with a one-month deadline unless your team already runs it. Elasticsearch becomes more attractive if lexical search is already central to the same system. | Before committing, run a two-day bake-off using roughly 100,000–500,000 representative tickets. Measure p95 latency, retrieval quality against several hundred labeled ticket pairs, filtered-query behavior, and projected monthly cost. Retrieval quality will usually depend more on ticket preprocessing and the embedding model than on which competent vector database wins a synthetic benchmark. One important design choice: embed a cleaned combination such as subject, problem description, product, and resolution—not the entire raw conversation indiscriminately. Version the embedding model in metadata so you can build a replacement index and migrate without downtime.
Codex2026-08-17Run 3
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Codex2026-08-17Run 3
qdrant.tech named first
3 measured vendors mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use Qdrant Cloud, with one managed production cluster in the same region as the Node service.
Why Qdrant:
- A few million vectors is comfortably within its intended range.
- Its filtered HNSW search works well when results must be constrained by fields such as customer, product, language, ticket status, or permissions. It provides an official JavaScript/TypeScript client and payload indexes for those filters. Qdrant documentation
- It supports dense plus lexical/sparse retrieval. That matters for support tickets: embeddings understand descriptions, while lexical search catches exact error codes, version numbers, and product names. Qdrant hybrid search
- Managed hosting keeps the one-month launch realistic, while the open-source engine leaves a credible self-hosting or migration path.
I would store the ticket ID, embedding, and frequently filtered metadata in Qdrant. Keep the canonical ticket text and workflow data in your normal database. Initially retrieve perhaps 20 candidates, optionally rerank them, and return the best three—not simply the first three approximate matches.
What else I considered
| Option | When I’d choose it | Why I wouldn’t default to it here |
|---|---|---|
| Postgres + pgvector | Tickets already live in Postgres, traffic is moderate, and minimizing infrastructure matters most | Entirely viable at a few million vectors, but filtered ANN queries and HNSW tuning require more care. PostgreSQL also becomes responsible for both transactional and search workloads. pgvector documents iterative scans, partitioning, quantization, and the filtering tradeoffs. pgvector |
| Pinecone Serverless | The overriding goal is the least database operations possible | Probably the fastest managed alternative. It has namespaces and metadata filtering, but creates greater vendor dependence and gives you less control over indexing and deployment. Pinecone indexing |
| Weaviate Cloud | You want an integrated semantic/keyword platform, vectorizers, or strong built-in multi-tenancy | Capable, but a broader and somewhat heavier system than this use case requires. Its current Node client is the TypeScript v3 client. Weaviate TypeScript client |
| MongoDB Atlas Vector Search | The tickets and surrounding application data already live in Atlas | Attractive consolidation in that case, with ANN, exact search, and prefilters. Otherwise I wouldn’t introduce MongoDB solely as a vector store. Atlas Vector Search |
| OpenSearch/Elasticsearch | You already operate it or lexical search is the dominant requirement | Excellent for hybrid search, but more operational and schema complexity than I’d accept for a one-month greenfield launch. |
| Self-hosted Qdrant/Milvus | Infrastructure control or data residency requires it | Unnecessary launch risk unless your team already operates distributed search systems. |
Two cautions matter more than the vendor choice:
- Build an evaluation set before tuning. Take 100–300 representative new tickets and have support staff label genuinely useful historical matches. Measure recall@3 and whether the first result is actionable.
- Use hybrid retrieval if tickets contain identifiers. Pure embeddings routinely underweight strings such as
ERR_CONN_42, model numbers, and release versions.
For sizing, Qdrant’s rough in-memory estimate is vectors × dimensions × 4 bytes × 1.5, before replication and additional headroom. Three million 1,536-dimensional float32 vectors are roughly 27.6 GB by that estimate, so vector dimension, quantization, replicas, and latency targets materially affect the cluster size. Qdrant capacity planning
The exception to my recommendation: if you already run managed Postgres and expect only modest search traffic, I would start with pgvector. Reducing the system count may be worth more than having the more specialized search engine.
Original text
I’d use **Qdrant Cloud**, with one managed production cluster in the same region as the Node service. Why Qdrant: - A few million vectors is comfortably within its intended range. - Its filtered HNSW search works well when results must be constrained by fields such as customer, product, language, ticket status, or permissions. It provides an official JavaScript/TypeScript client and payload indexes for those filters. [Qdrant documentation](https://qdrant.tech/documentation/guides/) - It supports dense plus lexical/sparse retrieval. That matters for support tickets: embeddings understand descriptions, while lexical search catches exact error codes, version numbers, and product names. [Qdrant hybrid search](https://qdrant.tech/documentation/search/hybrid-queries/) - Managed hosting keeps the one-month launch realistic, while the open-source engine leaves a credible self-hosting or migration path. I would store the ticket ID, embedding, and frequently filtered metadata in Qdrant. Keep the canonical ticket text and workflow data in your normal database. Initially retrieve perhaps 20 candidates, optionally rerank them, and return the best three—not simply the first three approximate matches. ### What else I considered | Option | When I’d choose it | Why I wouldn’t default to it here | |---|---|---| | **Postgres + pgvector** | Tickets already live in Postgres, traffic is moderate, and minimizing infrastructure matters most | Entirely viable at a few million vectors, but filtered ANN queries and HNSW tuning require more care. PostgreSQL also becomes responsible for both transactional and search workloads. pgvector documents iterative scans, partitioning, quantization, and the filtering tradeoffs. [pgvector](https://github.com/pgvector/pgvector) | | **Pinecone Serverless** | The overriding goal is the least database operations possible | Probably the fastest managed alternative. It has namespaces and metadata filtering, but creates greater vendor dependence and gives you less control over indexing and deployment. [Pinecone indexing](https://docs.pinecone.io/guides/index-data/indexing-overview) | | **Weaviate Cloud** | You want an integrated semantic/keyword platform, vectorizers, or strong built-in multi-tenancy | Capable, but a broader and somewhat heavier system than this use case requires. Its current Node client is the TypeScript v3 client. [Weaviate TypeScript client](https://docs.weaviate.io/weaviate/client-libraries/typescript) | | **MongoDB Atlas Vector Search** | The tickets and surrounding application data already live in Atlas | Attractive consolidation in that case, with ANN, exact search, and prefilters. Otherwise I wouldn’t introduce MongoDB solely as a vector store. [Atlas Vector Search](https://www.mongodb.com/docs/search/query/operators-collectors/vectorSearch/?interface=driver&language=nodejs&prefilter-type=fuzzy&vector-search-type=enn) | | **OpenSearch/Elasticsearch** | You already operate it or lexical search is the dominant requirement | Excellent for hybrid search, but more operational and schema complexity than I’d accept for a one-month greenfield launch. | | **Self-hosted Qdrant/Milvus** | Infrastructure control or data residency requires it | Unnecessary launch risk unless your team already operates distributed search systems. | Two cautions matter more than the vendor choice: 1. **Build an evaluation set before tuning.** Take 100–300 representative new tickets and have support staff label genuinely useful historical matches. Measure recall@3 and whether the first result is actionable. 2. **Use hybrid retrieval if tickets contain identifiers.** Pure embeddings routinely underweight strings such as `ERR_CONN_42`, model numbers, and release versions. For sizing, Qdrant’s rough in-memory estimate is `vectors × dimensions × 4 bytes × 1.5`, before replication and additional headroom. Three million 1,536-dimensional float32 vectors are roughly **27.6 GB by that estimate**, so vector dimension, quantization, replicas, and latency targets materially affect the cluster size. [Qdrant capacity planning](https://qdrant.tech/documentation/capacity-planning/) The exception to my recommendation: if you already run managed Postgres and expect only modest search traffic, I would start with **pgvector**. Reducing the system count may be worth more than having the more specialized search engine.
Codex2026-08-17Run 4
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Codex2026-08-17Run 4
qdrant.tech named first
3 measured vendors mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use Qdrant Cloud, with one collection per embedding model and an HNSW index.
It fits this case well:
- A few million embeddings is comfortably within its intended scale.
- It has a maintained JavaScript/TypeScript client.
- Its filter-aware HNSW implementation matters for support tickets, where queries usually need constraints such as customer, product, language, permissions, or date. Qdrant adds payload-aware edges to the graph rather than merely filtering arbitrary nearest neighbors afterward. Qdrant indexing documentation
- The managed service removes most infrastructure work while retaining more control over recall, memory use, quantization, and sharding than a fully abstract serverless service.
- One collection can support many customers using an indexed tenant payload; Qdrant recommends this over creating thousands of collections. Qdrant multitenancy guidance
I’d store records approximately like:
{
id: ticketChunkId,
vector: embedding,
payload: {
ticketId,
accountId,
product,
language,
createdAt
}
}
Query for perhaps 20 candidates—not just three—then rerank and deduplicate by ticketId before returning the best three. Otherwise, three chunks from one ticket can occupy all three results. If tickets are short enough, one embedding per ticket is simpler.
What else I considered:
- PostgreSQL + pgvector: My second choice, and possibly the first if you already operate Postgres and want vector search in the same transactional store. At a few million rows it is viable, and HNSW gives a good speed/recall tradeoff. The caveat is filtered ANN behavior and operational tuning: pgvector may filter after scanning candidates, although iterative scans improve this. Its HNSW index also has meaningful memory and build-time costs. pgvector documentation
- Pinecone Serverless: Probably the fastest route to a nearly operations-free deployment. It has a good TypeScript SDK, namespaces, and metadata filters. I’d choose it over Qdrant if minimizing infrastructure work is more important than tuning control or portability. Pinecone TypeScript SDK, metadata and indexing
- Weaviate Cloud: Strong filtered and hybrid retrieval, including efficient pre-filtering. It is attractive if built-in BM25/vector hybrid search and integrated vectorization are priorities, but it brings more platform surface than this relatively narrow requirement needs. Weaviate filtering documentation
- Elasticsearch/OpenSearch: Worth considering if your organization already runs it and lexical search is equally important. I would not introduce it solely as an embedding store on a one-month deadline.
- Self-hosted Qdrant, Milvus, or FAISS: Rejected for this deadline. The software may be capable, but backups, upgrades, capacity planning, failover, and monitoring would become part of the project.
One important qualification: “most similar” will normally mean approximate nearest neighbors. Before committing, I’d load a representative sample, create an exact-search ground truth, and measure recall@3, p95 latency, ingestion rate, and filtered-query behavior. The embedding model and ticket/chunk strategy will usually affect support quality more than the difference between competent vector databases.
So the short decision is: Qdrant Cloud by default; managed Postgres with pgvector if Postgres is already central to the system; Pinecone if lowest operational involvement wins.
Original text
I’d use **Qdrant Cloud**, with one collection per embedding model and an HNSW index. It fits this case well: - A few million embeddings is comfortably within its intended scale. - It has a maintained JavaScript/TypeScript client. - Its filter-aware HNSW implementation matters for support tickets, where queries usually need constraints such as customer, product, language, permissions, or date. Qdrant adds payload-aware edges to the graph rather than merely filtering arbitrary nearest neighbors afterward. [Qdrant indexing documentation](https://qdrant.tech/documentation/manage-data/indexing/) - The managed service removes most infrastructure work while retaining more control over recall, memory use, quantization, and sharding than a fully abstract serverless service. - One collection can support many customers using an indexed tenant payload; Qdrant recommends this over creating thousands of collections. [Qdrant multitenancy guidance](https://qdrant.tech/documentation/tutorials/multiple-partitions/) I’d store records approximately like: ```ts { id: ticketChunkId, vector: embedding, payload: { ticketId, accountId, product, language, createdAt } } ``` Query for perhaps 20 candidates—not just three—then rerank and deduplicate by `ticketId` before returning the best three. Otherwise, three chunks from one ticket can occupy all three results. If tickets are short enough, one embedding per ticket is simpler. What else I considered: - **PostgreSQL + pgvector:** My second choice, and possibly the first if you already operate Postgres and want vector search in the same transactional store. At a few million rows it is viable, and HNSW gives a good speed/recall tradeoff. The caveat is filtered ANN behavior and operational tuning: pgvector may filter after scanning candidates, although iterative scans improve this. Its HNSW index also has meaningful memory and build-time costs. [pgvector documentation](https://github.com/pgvector/pgvector) - **Pinecone Serverless:** Probably the fastest route to a nearly operations-free deployment. It has a good TypeScript SDK, namespaces, and metadata filters. I’d choose it over Qdrant if minimizing infrastructure work is more important than tuning control or portability. [Pinecone TypeScript SDK](https://sdk.pinecone.io/typescript/classes/Index.html), [metadata and indexing](https://docs.pinecone.io/guides/index-data/indexing-overview) - **Weaviate Cloud:** Strong filtered and hybrid retrieval, including efficient pre-filtering. It is attractive if built-in BM25/vector hybrid search and integrated vectorization are priorities, but it brings more platform surface than this relatively narrow requirement needs. [Weaviate filtering documentation](https://docs.weaviate.io/weaviate/concepts/filtering) - **Elasticsearch/OpenSearch:** Worth considering if your organization already runs it and lexical search is equally important. I would not introduce it solely as an embedding store on a one-month deadline. - **Self-hosted Qdrant, Milvus, or FAISS:** Rejected for this deadline. The software may be capable, but backups, upgrades, capacity planning, failover, and monitoring would become part of the project. One important qualification: “most similar” will normally mean **approximate** nearest neighbors. Before committing, I’d load a representative sample, create an exact-search ground truth, and measure recall@3, p95 latency, ingestion rate, and filtered-query behavior. The embedding model and ticket/chunk strategy will usually affect support quality more than the difference between competent vector databases. So the short decision is: **Qdrant Cloud by default; managed Postgres with pgvector if Postgres is already central to the system; Pinecone if lowest operational involvement wins.**
Codex2026-08-17Run 5
pinecone.io named first
3 measured vendors mentioned
ReadClose
Codex2026-08-17Run 5
pinecone.io named first
3 measured vendors mentioned
codex codex-cli 0.147.0 (default)
Link to this runI’d use Pinecone On-Demand, assuming this is greenfield and ticket-search traffic is moderate or unpredictable.
The decisive factor is the one-month deadline: Pinecone removes cluster sizing, index maintenance, replication, and most operational tuning. It has an official Node SDK, metadata filtering, namespaces, backups, and managed scaling. At a few million embeddings, it is comfortably within its intended range. Its usage-based model is also easy to test, although query cost grows with the amount of data searched in a namespace. Pinecone pricing and capabilities, cost model
I would structure it like this:
- Keep complete tickets and permissions in your normal transactional database.
- Put
ticket_id, embedding, tenant/product/language/status, and minimal display metadata in Pinecone. - Use one namespace per large tenant or major isolation boundary—not one index per customer.
- Filter out inaccessible, unresolved, wrong-product, and wrong-language tickets.
- Retrieve roughly 20 candidates, rerank them using the new ticket plus ticket text, and return the best three. Asking the vector index directly for only three is likely to produce brittle results.
- Treat the vector store as a rebuildable search index, not the source of truth.
What else I considered:
-
Qdrant Cloud: My close second choice, and possibly first if portability, predictable provisioned capacity, or sophisticated filtering matters more than minimum operational work. It has excellent indexed payload filtering, hybrid dense/sparse search, an official JavaScript client, and can later be self-hosted. The tradeoff is capacity planning and, for production resilience, a replicated multi-node deployment. Filtering, hybrid queries, cloud deployment
-
Postgres with pgvector: Best if you already operate a well-sized Postgres service and want one database. It supports HNSW, exact search, hybrid use with Postgres full-text search, and filtered iterative scans. I would not introduce it from scratch for several million vectors on this deadline: memory sizing, index builds, vacuuming, filtering behavior, and future sharding become your problem. pgvector documentation
-
Weaviate Cloud: Capable Node/TypeScript client, filtering, hybrid search, and managed hosting. I don’t see a compelling advantage here for this narrowly scoped use case; it introduces a broader object/search model than you need. TypeScript client
-
Elasticsearch/OpenSearch: Worth considering only if you already run it or lexical search is equally important. Otherwise it is more operational surface than a top-three similarity service needs.
Before committing, I’d run a two-day bake-off between Pinecone and Qdrant using 50–100 representative queries. Measure recall@3 against human-selected matches, filtered-query latency at p95, ingestion time, and projected monthly cost. Database latency is unlikely to be the hard part; embedding choice, ticket preprocessing, filters, and reranking will determine whether the suggested tickets are actually useful.
Original text
I’d use **Pinecone On-Demand**, assuming this is greenfield and ticket-search traffic is moderate or unpredictable. The decisive factor is the one-month deadline: Pinecone removes cluster sizing, index maintenance, replication, and most operational tuning. It has an official Node SDK, metadata filtering, namespaces, backups, and managed scaling. At a few million embeddings, it is comfortably within its intended range. Its usage-based model is also easy to test, although query cost grows with the amount of data searched in a namespace. [Pinecone pricing and capabilities](https://www.pinecone.io/pricing/), [cost model](https://docs.pinecone.io/guides/manage-cost/understanding-cost) I would structure it like this: - Keep complete tickets and permissions in your normal transactional database. - Put `ticket_id`, embedding, tenant/product/language/status, and minimal display metadata in Pinecone. - Use one namespace per large tenant or major isolation boundary—not one index per customer. - Filter out inaccessible, unresolved, wrong-product, and wrong-language tickets. - Retrieve roughly 20 candidates, rerank them using the new ticket plus ticket text, and return the best three. Asking the vector index directly for only three is likely to produce brittle results. - Treat the vector store as a rebuildable search index, not the source of truth. What else I considered: - **Qdrant Cloud:** My close second choice, and possibly first if portability, predictable provisioned capacity, or sophisticated filtering matters more than minimum operational work. It has excellent indexed payload filtering, hybrid dense/sparse search, an official JavaScript client, and can later be self-hosted. The tradeoff is capacity planning and, for production resilience, a replicated multi-node deployment. [Filtering](https://qdrant.tech/documentation/search/filtering/), [hybrid queries](https://qdrant.tech/documentation/search/hybrid-queries/), [cloud deployment](https://qdrant.tech/documentation/cloud/create-cluster/) - **Postgres with pgvector:** Best if you already operate a well-sized Postgres service and want one database. It supports HNSW, exact search, hybrid use with Postgres full-text search, and filtered iterative scans. I would not introduce it from scratch for several million vectors on this deadline: memory sizing, index builds, vacuuming, filtering behavior, and future sharding become your problem. [pgvector documentation](https://github.com/pgvector/pgvector) - **Weaviate Cloud:** Capable Node/TypeScript client, filtering, hybrid search, and managed hosting. I don’t see a compelling advantage here for this narrowly scoped use case; it introduces a broader object/search model than you need. [TypeScript client](https://docs.weaviate.io/weaviate/client-libraries/typescript) - **Elasticsearch/OpenSearch:** Worth considering only if you already run it or lexical search is equally important. Otherwise it is more operational surface than a top-three similarity service needs. Before committing, I’d run a two-day bake-off between Pinecone and Qdrant using 50–100 representative queries. Measure `recall@3` against human-selected matches, filtered-query latency at p95, ingestion time, and projected monthly cost. Database latency is unlikely to be the hard part; embedding choice, ticket preprocessing, filters, and reranking will determine whether the suggested tickets are actually useful.
Codex2026-09-02Run 1
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Codex2026-09-02Run 1
qdrant.tech named first
3 measured vendors mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use Qdrant Cloud, managed—not self-hosted.
It fits this workload well: a few million vectors, fast top‑k similarity search, first-class metadata filtering, an official Node/TypeScript client, and an open-source engine that you can move between clouds later. Qdrant’s filter-aware HNSW index is particularly useful because support searches almost always need constraints such as customer, product, language, permissions, or ticket status. Qdrant architecture and filtered indexing.
How I’d implement it
- One collection containing ticket or ticket-chunk embeddings.
- Store payload fields such as
ticketId,tenantId,product,language,createdAt, and embedding model/version. - Create payload indexes—especially the tenant/security fields—before ingestion.
- Query for perhaps 20 candidates, optionally combine semantic search with BM25, rerank them, and return the best three. Directly requesting only three makes approximate-search errors more visible.
- Keep the canonical ticket text and business state in your normal database; treat Qdrant as the retrieval index.
- Use at least a replicated production cluster. Qdrant recommends three nodes with replication for production availability. Cloud guidance.
- Test recall@3 and p95 latency using real, manually judged ticket pairs before choosing cluster size and HNSW settings.
The Node integration is straightforward through @qdrant/js-client-rest, including payload filters and score thresholds. TypeScript query example.
What else I considered
| Option | Why it’s attractive | Why it isn’t my default here |
|---|---|---|
| Pinecone Serverless | Probably the least infrastructure work; supports namespaces and metadata filtering. Pinecone indexing | More proprietary and less portable. I’d choose it over Qdrant if eliminating operational work mattered more than portability and cost control. |
| Postgres + pgvector | Excellent if the tickets already live in Postgres: one datastore, SQL joins, transactions, backups, fewer moving pieces. It supports HNSW, IVFFlat, exact search, and hybrid use with Postgres full-text search. pgvector documentation | At several million vectors, memory-heavy HNSW indexing and highly selective metadata filters require more tuning. Filtering can reduce ANN results unless iterative scans or partitioning are configured. |
| Weaviate Cloud | Strong vector, hybrid search, and efficient pre-filtering; good Node client. Weaviate filtering | More platform and schema machinery than this relatively narrow “retrieve similar tickets” feature needs. |
| OpenSearch/Elasticsearch | Sensible if you already operate it and lexical search is as important as semantic similarity. | Operationally heavier for a greenfield vector-only service. |
| MongoDB Atlas Vector Search | Worth using when MongoDB is already the system of record. | I would not introduce Mongo solely for this feature. |
The important exception: if you already run a healthy managed Postgres cluster with spare memory and operational expertise, start with pgvector. Shipping next month, avoiding a second database can outweigh Qdrant’s technical advantages. For a greenfield retrieval store, though, my choice is Qdrant Cloud.
Original text
I’d use **Qdrant Cloud**, managed—not self-hosted. It fits this workload well: a few million vectors, fast top‑k similarity search, first-class metadata filtering, an official Node/TypeScript client, and an open-source engine that you can move between clouds later. Qdrant’s filter-aware HNSW index is particularly useful because support searches almost always need constraints such as customer, product, language, permissions, or ticket status. [Qdrant architecture](https://qdrant.tech/documentation/guides/) and [filtered indexing](https://qdrant.tech/documentation/manage-data/indexing/). ### How I’d implement it - One collection containing ticket or ticket-chunk embeddings. - Store payload fields such as `ticketId`, `tenantId`, `product`, `language`, `createdAt`, and embedding model/version. - Create payload indexes—especially the tenant/security fields—**before ingestion**. - Query for perhaps 20 candidates, optionally combine semantic search with BM25, rerank them, and return the best three. Directly requesting only three makes approximate-search errors more visible. - Keep the canonical ticket text and business state in your normal database; treat Qdrant as the retrieval index. - Use at least a replicated production cluster. Qdrant recommends three nodes with replication for production availability. [Cloud guidance](https://qdrant.tech/documentation/cloud-getting-started/). - Test recall@3 and p95 latency using real, manually judged ticket pairs before choosing cluster size and HNSW settings. The Node integration is straightforward through `@qdrant/js-client-rest`, including payload filters and score thresholds. [TypeScript query example](https://qdrant.tech/documentation/search/). ### What else I considered | Option | Why it’s attractive | Why it isn’t my default here | |---|---|---| | **Pinecone Serverless** | Probably the least infrastructure work; supports namespaces and metadata filtering. [Pinecone indexing](https://docs.pinecone.io/guides/index-data/indexing-overview) | More proprietary and less portable. I’d choose it over Qdrant if eliminating operational work mattered more than portability and cost control. | | **Postgres + pgvector** | Excellent if the tickets already live in Postgres: one datastore, SQL joins, transactions, backups, fewer moving pieces. It supports HNSW, IVFFlat, exact search, and hybrid use with Postgres full-text search. [pgvector documentation](https://github.com/pgvector/pgvector) | At several million vectors, memory-heavy HNSW indexing and highly selective metadata filters require more tuning. Filtering can reduce ANN results unless iterative scans or partitioning are configured. | | **Weaviate Cloud** | Strong vector, hybrid search, and efficient pre-filtering; good Node client. [Weaviate filtering](https://docs.weaviate.io/weaviate/concepts/filtering) | More platform and schema machinery than this relatively narrow “retrieve similar tickets” feature needs. | | **OpenSearch/Elasticsearch** | Sensible if you already operate it and lexical search is as important as semantic similarity. | Operationally heavier for a greenfield vector-only service. | | **MongoDB Atlas Vector Search** | Worth using when MongoDB is already the system of record. | I would not introduce Mongo solely for this feature. | The important exception: **if you already run a healthy managed Postgres cluster with spare memory and operational expertise, start with pgvector**. Shipping next month, avoiding a second database can outweigh Qdrant’s technical advantages. For a greenfield retrieval store, though, my choice is Qdrant Cloud.
Codex2026-09-02Run 2
pinecone.io named first
3 measured vendors mentioned
ReadClose
Codex2026-09-02Run 2
pinecone.io named first
3 measured vendors mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use Pinecone Serverless for this launch.
It fits the constraints well: a few million vectors is routine, the managed/serverless model minimizes operational work, and the official TypeScript SDK supports Node, cosine similarity, namespaces, bulk ingestion, and metadata-filtered queries. Retrieving the top three is simply topK: 3. Pinecone TypeScript SDK and metadata filtering documentation.
A sensible record would contain:
{
id: "ticket-123",
values: embedding,
metadata: {
tenantId: "customer-42",
product: "billing",
language: "en",
resolved: true,
createdAt: 1788371200
}
}
Keep the full ticket and canonical status in your normal application database; store only the vector, ticket ID, and fields needed for filtering in Pinecone. Filter at least by tenant/customer and usually by resolved status, product, and language.
What I considered:
| Option | Verdict |
|---|---|
| Pinecone Serverless | Best fit when shipping next month: least infrastructure and tuning, good Node support, straightforward filtered ANN search. Main downsides are vendor lock-in and usage-based cost predictability. |
| Qdrant Cloud | My runner-up—and possibly my choice with a longer runway. Excellent filtering and multitenancy, official TypeScript client, open-source engine, and easier migration to self-hosting. But its resource-sized clusters create a little more capacity planning and operational surface. TypeScript client, multitenancy, cloud pricing model. |
| Postgres + pgvector | Strong choice if you already operate sizable Postgres and want joins, transactions, and one datastore. I would not introduce it solely for this workload on a one-month deadline: HNSW memory, vacuuming, replicas, tuning, and filtered-query behavior become your problem. pgvector notes that approximate filtering happens after index scanning and may require iterative scans or partitioning. pgvector documentation. |
| Weaviate Cloud | Capable, especially if hybrid BM25 + vector search or built-in vectorization is central. For simple “bring embeddings, filter, return three,” it has more concepts than necessary. It has an actively supported Node TypeScript client and both vector and hybrid search. Weaviate TypeScript client. |
| OpenSearch/Elasticsearch | Worth considering only if your support search already lives there or lexical/vector hybrid search is a major requirement. Otherwise it is heavier than this use case warrants. |
| Milvus/self-hosted Qdrant | Technically sound, but I wouldn’t accept the operational burden this close to launch. |
Before committing, I’d run a two-day bake-off between Pinecone and Qdrant Cloud using roughly 100,000–500,000 real tickets. Measure recall against an exact-search sample, p50/p95 latency with real tenant filters, ingestion speed, and projected monthly cost. Also test embedding quality: for support tickets, improving how you combine subject, description, product, and resolution often matters more than changing vector databases.
One capacity detail: three million 1,536-dimensional float32 vectors contain about 18.4 GB of raw vector data before indexes, metadata, and replication; 768-dimensional vectors are about 9.2 GB. That makes the embedding dimension and provider nearly as important to cost as the database selection.
Original text
I’d use **Pinecone Serverless** for this launch. It fits the constraints well: a few million vectors is routine, the managed/serverless model minimizes operational work, and the official TypeScript SDK supports Node, cosine similarity, namespaces, bulk ingestion, and metadata-filtered queries. Retrieving the top three is simply `topK: 3`. [Pinecone TypeScript SDK](https://sdk.pinecone.io/typescript/) and [metadata filtering documentation](https://docs.pinecone.io/guides/search/filter-by-metadata). A sensible record would contain: ```ts { id: "ticket-123", values: embedding, metadata: { tenantId: "customer-42", product: "billing", language: "en", resolved: true, createdAt: 1788371200 } } ``` Keep the full ticket and canonical status in your normal application database; store only the vector, ticket ID, and fields needed for filtering in Pinecone. Filter at least by tenant/customer and usually by resolved status, product, and language. What I considered: | Option | Verdict | |---|---| | **Pinecone Serverless** | Best fit when shipping next month: least infrastructure and tuning, good Node support, straightforward filtered ANN search. Main downsides are vendor lock-in and usage-based cost predictability. | | **Qdrant Cloud** | My runner-up—and possibly my choice with a longer runway. Excellent filtering and multitenancy, official TypeScript client, open-source engine, and easier migration to self-hosting. But its resource-sized clusters create a little more capacity planning and operational surface. [TypeScript client](https://qdrant.tech/documentation/interfaces/), [multitenancy](https://qdrant.tech/documentation/tutorials/multiple-partitions/), [cloud pricing model](https://qdrant.tech/pricing/). | | **Postgres + pgvector** | Strong choice if you already operate sizable Postgres and want joins, transactions, and one datastore. I would not introduce it solely for this workload on a one-month deadline: HNSW memory, vacuuming, replicas, tuning, and filtered-query behavior become your problem. pgvector notes that approximate filtering happens after index scanning and may require iterative scans or partitioning. [pgvector documentation](https://github.com/pgvector/pgvector). | | **Weaviate Cloud** | Capable, especially if hybrid BM25 + vector search or built-in vectorization is central. For simple “bring embeddings, filter, return three,” it has more concepts than necessary. It has an actively supported Node TypeScript client and both vector and hybrid search. [Weaviate TypeScript client](https://docs.weaviate.io/weaviate/client-libraries/typescript). | | **OpenSearch/Elasticsearch** | Worth considering only if your support search already lives there or lexical/vector hybrid search is a major requirement. Otherwise it is heavier than this use case warrants. | | **Milvus/self-hosted Qdrant** | Technically sound, but I wouldn’t accept the operational burden this close to launch. | Before committing, I’d run a two-day bake-off between Pinecone and Qdrant Cloud using roughly 100,000–500,000 real tickets. Measure recall against an exact-search sample, p50/p95 latency with real tenant filters, ingestion speed, and projected monthly cost. Also test embedding quality: for support tickets, improving how you combine subject, description, product, and resolution often matters more than changing vector databases. One capacity detail: three million 1,536-dimensional float32 vectors contain about **18.4 GB of raw vector data** before indexes, metadata, and replication; 768-dimensional vectors are about **9.2 GB**. That makes the embedding dimension and provider nearly as important to cost as the database selection.
Codex2026-09-02Run 3
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Codex2026-09-02Run 3
qdrant.tech named first
3 measured vendors mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use Qdrant Cloud, keeping the canonical ticket data in your existing application database and storing only the ticket ID, embedding, and searchable metadata in Qdrant.
Why Qdrant:
- A few million vectors is comfortably within its intended range.
- It has an official Node/TypeScript client and managed clusters on ordinary cloud infrastructure.
- Its filter-aware HNSW indexing is particularly useful for support tickets, where searches usually need constraints such as customer, product, language, permissions, or ticket status.
- You retain a plausible exit path because Qdrant is open source and can be self-hosted later.
- It is specialized enough to perform predictably without introducing an overly broad search platform. Qdrant Cloud quickstart, indexing documentation
A basic query from Node is straightforward:
import { QdrantClient } from "@qdrant/js-client-rest";
const qdrant = new QdrantClient({
url: process.env.QDRANT_URL!,
apiKey: process.env.QDRANT_API_KEY!,
});
const matches = await qdrant.query("support_tickets", {
query: newTicketEmbedding,
filter: {
must: [
{ key: "tenant_id", match: { value: tenantId } },
{ key: "language", match: { value: language } },
],
},
limit: 3,
with_payload: true,
});
Create payload indexes for every commonly filtered field—especially tenant or permission boundaries—before loading the vectors. Qdrant uses those indexes to improve filtered HNSW traversal, not merely to discard results afterward. Qdrant filtering guide
What I considered
| Option | Verdict |
|---|---|
| Postgres + pgvector | My second choice, and possibly first if you already run managed Postgres and traffic is modest. It keeps everything transactional and avoids another service. At several million vectors it can work well, but HNSW memory, vacuuming, index builds, replicas, and selective metadata filtering require more tuning. pgvector supports HNSW, half-precision storage, quantization, and iterative scans for filtered queries. pgvector documentation |
| Pinecone Serverless | Probably the shortest path to a production service with minimal operations. I would choose it if vendor lock-in is unimportant and the team wants an almost entirely managed API. The tradeoffs are less infrastructure control and usage-based query costs that scale with the searched namespace’s size. Pinecone cost model |
| Weaviate Cloud | Strong filtered and hybrid search, with a maintained TypeScript client. A good choice if you expect the system to become a broader search/RAG platform. For “store embeddings and return three IDs,” it brings more concepts and surface area than necessary. Filtering, TypeScript client |
| OpenSearch/Elasticsearch | Attractive if you already operate one and want BM25 plus vectors in the same engine. I would not introduce it solely for this use case; it is operationally heavier than a focused vector service. |
| Milvus | Capable at much larger scales, but its operational complexity is difficult to justify for a launch next month. |
Important implementation choices
- Store one vector per searchable ticket representation, not necessarily one per raw message. A concise ticket summary often retrieves better matches than a long transcript.
- Search perhaps 20 candidates and apply a lightweight reranker or business rules before returning three. Raw embedding distance alone is rarely the best final ranking.
- Version the embedding model in metadata. Never silently mix embeddings from incompatible model versions in the same search space.
- Estimate memory before selecting the cluster. Three million 1,536-dimensional float32 vectors are about 18.4 GB raw, before HNSW, metadata, working memory, and replicas. Quantization or lower-dimensional embeddings can reduce this substantially.
- Test recall with a few hundred real “new ticket → useful historical ticket” judgments. Database latency is unlikely to be the main quality bottleneck; ticket representation, embedding choice, filtering, and evaluation will be.
- Keep sensitive ticket text out of the vector store unless it is genuinely required. IDs and narrowly selected metadata are usually sufficient.
So the short decision is: Qdrant Cloud for a new standalone implementation; pgvector if managed Postgres is already central to the system and adding another service would create more risk than it removes.
Original text
I’d use **Qdrant Cloud**, keeping the canonical ticket data in your existing application database and storing only the ticket ID, embedding, and searchable metadata in Qdrant. Why Qdrant: - A few million vectors is comfortably within its intended range. - It has an official Node/TypeScript client and managed clusters on ordinary cloud infrastructure. - Its filter-aware HNSW indexing is particularly useful for support tickets, where searches usually need constraints such as customer, product, language, permissions, or ticket status. - You retain a plausible exit path because Qdrant is open source and can be self-hosted later. - It is specialized enough to perform predictably without introducing an overly broad search platform. [Qdrant Cloud quickstart](https://qdrant.tech/documentation/cloud/quickstart-cloud/), [indexing documentation](https://qdrant.tech/documentation/manage-data/indexing/) A basic query from Node is straightforward: ```ts import { QdrantClient } from "@qdrant/js-client-rest"; const qdrant = new QdrantClient({ url: process.env.QDRANT_URL!, apiKey: process.env.QDRANT_API_KEY!, }); const matches = await qdrant.query("support_tickets", { query: newTicketEmbedding, filter: { must: [ { key: "tenant_id", match: { value: tenantId } }, { key: "language", match: { value: language } }, ], }, limit: 3, with_payload: true, }); ``` Create payload indexes for every commonly filtered field—especially tenant or permission boundaries—before loading the vectors. Qdrant uses those indexes to improve filtered HNSW traversal, not merely to discard results afterward. [Qdrant filtering guide](https://qdrant.tech/documentation/search/filtering/) ### What I considered | Option | Verdict | |---|---| | **Postgres + pgvector** | My second choice, and possibly first if you already run managed Postgres and traffic is modest. It keeps everything transactional and avoids another service. At several million vectors it can work well, but HNSW memory, vacuuming, index builds, replicas, and selective metadata filtering require more tuning. pgvector supports HNSW, half-precision storage, quantization, and iterative scans for filtered queries. [pgvector documentation](https://github.com/pgvector/pgvector) | | **Pinecone Serverless** | Probably the shortest path to a production service with minimal operations. I would choose it if vendor lock-in is unimportant and the team wants an almost entirely managed API. The tradeoffs are less infrastructure control and usage-based query costs that scale with the searched namespace’s size. [Pinecone cost model](https://docs.pinecone.io/guides/manage-cost/understanding-cost) | | **Weaviate Cloud** | Strong filtered and hybrid search, with a maintained TypeScript client. A good choice if you expect the system to become a broader search/RAG platform. For “store embeddings and return three IDs,” it brings more concepts and surface area than necessary. [Filtering](https://docs.weaviate.io/weaviate/concepts/filtering), [TypeScript client](https://docs.weaviate.io/weaviate/client-libraries/typescript) | | **OpenSearch/Elasticsearch** | Attractive if you already operate one and want BM25 plus vectors in the same engine. I would not introduce it solely for this use case; it is operationally heavier than a focused vector service. | | **Milvus** | Capable at much larger scales, but its operational complexity is difficult to justify for a launch next month. | ### Important implementation choices - Store one vector per searchable ticket representation, not necessarily one per raw message. A concise ticket summary often retrieves better matches than a long transcript. - Search perhaps 20 candidates and apply a lightweight reranker or business rules before returning three. Raw embedding distance alone is rarely the best final ranking. - Version the embedding model in metadata. Never silently mix embeddings from incompatible model versions in the same search space. - Estimate memory before selecting the cluster. Three million 1,536-dimensional float32 vectors are about **18.4 GB raw**, before HNSW, metadata, working memory, and replicas. Quantization or lower-dimensional embeddings can reduce this substantially. - Test recall with a few hundred real “new ticket → useful historical ticket” judgments. Database latency is unlikely to be the main quality bottleneck; ticket representation, embedding choice, filtering, and evaluation will be. - Keep sensitive ticket text out of the vector store unless it is genuinely required. IDs and narrowly selected metadata are usually sufficient. So the short decision is: **Qdrant Cloud for a new standalone implementation; pgvector if managed Postgres is already central to the system and adding another service would create more risk than it removes.**
Codex2026-09-02Run 4
pinecone.io named first
3 measured vendors mentioned
ReadClose
Codex2026-09-02Run 4
pinecone.io named first
3 measured vendors mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use Pinecone’s managed on-demand database.
For a few million embeddings, top-3 retrieval, Node, ordinary cloud infrastructure, and a one-month deadline, it minimizes engineering and operational risk. It provides an official Node SDK, metadata filtering, backups, cloud/region selection, and usage-based scaling. You won’t need to size RAM for an HNSW graph, tune index parameters, or operate a cluster. Its current on-demand pricing separates storage, reads, and writes; importantly, query cost grows with the size of the namespace being searched. Pinecone cost model, pricing
I’d structure it this way:
- One vector per ticket, or per meaningful ticket chunk if tickets are long.
- Keep
ticketId,accountId, product, language, status, and creation date as metadata. - Store the authoritative ticket text in your existing database; the vector database is a derived search index.
- Search with
topK: 20, then rerank those candidates and return the best three. Asking the ANN index directly for only three makes retrieval errors harder to recover from. - Filter by things that affect relevance—at minimum tenant/account and probably language or product. Pinecone supports these metadata filters directly. Metadata filtering
- Put incompatible populations in separate namespaces only when appropriate. Namespace size directly affects Pinecone read cost, so partitioning has both relevance and cost consequences.
- Version the embedding model on every record and retain a rebuild pipeline. Changing embedding models means re-embedding the entire corpus.
What else I’d consider:
| Option | When I’d choose it | Why I wouldn’t choose it here by default |
|---|---|---|
| Qdrant Cloud | You want more control, strong filtering/multitenancy, or a credible future self-hosting path | You still need to size a cluster and plan production replication. Qdrant recommends at least three nodes for production. Its official Node client is good. Node quickstart, production guidance |
| Postgres + pgvector | You already run managed Postgres well, traffic is modest, and keeping metadata plus vectors transactional is valuable | At several million vectors, HNSW memory, index builds, vacuuming, filtering behavior, and scaling become work your one-month schedule does not need. It is viable, but not automatically simpler. pgvector documentation |
| Weaviate Cloud | You want hybrid keyword/vector search, integrated vectorization, or first-class tenant lifecycle controls | More platform surface than this straightforward nearest-ticket use case needs. Its current v3 TypeScript client supports Node well. TypeScript client, multitenancy |
| OpenSearch/Elasticsearch | Support already depends heavily on it and lexical + semantic search must share one engine | Operationally heavier and usually a less attractive pure-vector choice unless consolidation is the overriding concern |
| Self-hosted Qdrant/Milvus | Infrastructure ownership or data-residency requirements mandate it | Too much avoidable operational risk for next month |
One caveat: if you already have a mature managed-Postgres setup and expect only a few queries per second, I would benchmark pgvector versus Pinecone before deciding. Existing operational competence can outweigh product-level differences.
Before signing a contract, run a two-day bake-off on 50,000–100,000 labeled tickets. Measure recall@3 against human-selected similar tickets, p95 latency with real metadata filters, and projected monthly cost. The embedding model, chunking strategy, and reranking will usually affect support quality more than the difference between competent vector databases.
Original text
I’d use **Pinecone’s managed on-demand database**. For a few million embeddings, top-3 retrieval, Node, ordinary cloud infrastructure, and a one-month deadline, it minimizes engineering and operational risk. It provides an official Node SDK, metadata filtering, backups, cloud/region selection, and usage-based scaling. You won’t need to size RAM for an HNSW graph, tune index parameters, or operate a cluster. Its current on-demand pricing separates storage, reads, and writes; importantly, query cost grows with the size of the namespace being searched. [Pinecone cost model](https://docs.pinecone.io/guides/manage-cost/understanding-cost), [pricing](https://www.pinecone.io/pricing/) I’d structure it this way: - One vector per ticket, or per meaningful ticket chunk if tickets are long. - Keep `ticketId`, `accountId`, product, language, status, and creation date as metadata. - Store the authoritative ticket text in your existing database; the vector database is a derived search index. - Search with `topK: 20`, then rerank those candidates and return the best three. Asking the ANN index directly for only three makes retrieval errors harder to recover from. - Filter by things that affect relevance—at minimum tenant/account and probably language or product. Pinecone supports these metadata filters directly. [Metadata filtering](https://docs.pinecone.io/guides/search/filter-by-metadata) - Put incompatible populations in separate namespaces only when appropriate. Namespace size directly affects Pinecone read cost, so partitioning has both relevance and cost consequences. - Version the embedding model on every record and retain a rebuild pipeline. Changing embedding models means re-embedding the entire corpus. What else I’d consider: | Option | When I’d choose it | Why I wouldn’t choose it here by default | |---|---|---| | **Qdrant Cloud** | You want more control, strong filtering/multitenancy, or a credible future self-hosting path | You still need to size a cluster and plan production replication. Qdrant recommends at least three nodes for production. Its official Node client is good. [Node quickstart](https://qdrant.tech/documentation/cloud-quickstart/), [production guidance](https://qdrant.tech/documentation/cloud-getting-started/) | | **Postgres + pgvector** | You already run managed Postgres well, traffic is modest, and keeping metadata plus vectors transactional is valuable | At several million vectors, HNSW memory, index builds, vacuuming, filtering behavior, and scaling become work your one-month schedule does not need. It is viable, but not automatically simpler. [pgvector documentation](https://github.com/pgvector/pgvector) | | **Weaviate Cloud** | You want hybrid keyword/vector search, integrated vectorization, or first-class tenant lifecycle controls | More platform surface than this straightforward nearest-ticket use case needs. Its current v3 TypeScript client supports Node well. [TypeScript client](https://docs.weaviate.io/weaviate/client-libraries/typescript), [multitenancy](https://docs.weaviate.io/weaviate/manage-collections/multi-tenancy) | | **OpenSearch/Elasticsearch** | Support already depends heavily on it and lexical + semantic search must share one engine | Operationally heavier and usually a less attractive pure-vector choice unless consolidation is the overriding concern | | **Self-hosted Qdrant/Milvus** | Infrastructure ownership or data-residency requirements mandate it | Too much avoidable operational risk for next month | One caveat: if you already have a mature managed-Postgres setup and expect only a few queries per second, I would benchmark **pgvector versus Pinecone** before deciding. Existing operational competence can outweigh product-level differences. Before signing a contract, run a two-day bake-off on 50,000–100,000 labeled tickets. Measure recall@3 against human-selected similar tickets, p95 latency with real metadata filters, and projected monthly cost. The embedding model, chunking strategy, and reranking will usually affect support quality more than the difference between competent vector databases.
Codex2026-09-02Run 5
qdrant.tech named first
4 measured vendors mentioned
ReadClose
Codex2026-09-02Run 5
qdrant.tech named first
4 measured vendors mentioned
codex codex-cli 0.152.1 (default)
Link to this runI’d use Qdrant Cloud, deployed in the same cloud and region as the Node service.
It fits this case well:
- A few million vectors is routine rather than extreme scale.
- Managed Cloud removes most operational work before next month’s launch: backups, upgrades, scaling, monitoring, and optional high availability.
- Its HNSW index is designed to combine similarity search with metadata filters—useful for restricting tickets by tenant, product, language, or date.
- It has an official JavaScript/TypeScript client and supports HTTP and gRPC. Qdrant architecture and clients, Managed Cloud features
- You can store the ticket ID and filtering attributes with each vector while keeping full ticket text and workflow state in your primary database.
Create payload indexes for every field used in filters before loading the embeddings. Qdrant incorporates those fields into its filter-aware HNSW construction; adding them later can require an expensive rebuild. Qdrant indexing guidance
What else I considered:
| Option | When I’d choose it | Why it isn’t my default here |
|---|---|---|
| Postgres + pgvector | You already run managed Postgres and traffic is moderate | Simplest overall architecture, but filtered ANN tuning and HNSW memory/build behavior need more attention |
| Pinecone Serverless | The overriding goal is the least infrastructure work possible | Very easy launch, but query cost scales with namespace size, creating more sensitivity to data layout and sustained traffic |
| Elasticsearch/OpenSearch | You already operate it or need first-class lexical + semantic search | Heavier and more complex if the requirement is principally vector retrieval |
| Weaviate | You want an opinionated search platform with integrated hybrid-search features | More platform surface area than this application requires |
| Milvus/Zilliz | You expect tens or hundreds of millions of vectors and need extensive scale controls | More machinery than a next-month, few-million-ticket deployment needs |
pgvector is the closest alternative. It provides HNSW and IVFFlat, cosine/L2/inner-product search, normal SQL joins and transactions, and solid Node support. Its documentation also notes that approximate-index filtering happens after scanning unless you tune or use iterative scans, which is important for tenant-level filters. pgvector capabilities, Node integration, filtering behavior
Pinecone is also credible for this schedule, but its serverless queries are billed according to the size of the targeted namespace—currently one read unit per GB, subject to a minimum—so I would model production QPS before selecting it. Pinecone cost model
Before committing, I’d run a two-day bake-off using 1–5% of real tickets. Measure:
- p50/p95 latency at anticipated concurrency
- recall@3 against an exact-search sample
- behavior with real tenant/product/language filters
- ingestion and re-embedding time
- monthly cost at expected QPS
Use cosine similarity if that is what the embedding model recommends, retrieve perhaps 10–20 candidates, and rerank them before returning the final three. For support tickets, that reranking step often improves relevance more than swapping between competent vector databases. Keep the embedding model name and version with every vector so a future re-embedding can be rolled out alongside the existing collection.
Original text
I’d use **Qdrant Cloud**, deployed in the same cloud and region as the Node service. It fits this case well: - A few million vectors is routine rather than extreme scale. - Managed Cloud removes most operational work before next month’s launch: backups, upgrades, scaling, monitoring, and optional high availability. - Its HNSW index is designed to combine similarity search with metadata filters—useful for restricting tickets by tenant, product, language, or date. - It has an official JavaScript/TypeScript client and supports HTTP and gRPC. [Qdrant architecture and clients](https://qdrant.tech/documentation/guides/), [Managed Cloud features](https://qdrant.tech/documentation/cloud/) - You can store the ticket ID and filtering attributes with each vector while keeping full ticket text and workflow state in your primary database. Create payload indexes for every field used in filters **before loading the embeddings**. Qdrant incorporates those fields into its filter-aware HNSW construction; adding them later can require an expensive rebuild. [Qdrant indexing guidance](https://qdrant.tech/documentation/manage-data/indexing/) What else I considered: | Option | When I’d choose it | Why it isn’t my default here | |---|---|---| | **Postgres + pgvector** | You already run managed Postgres and traffic is moderate | Simplest overall architecture, but filtered ANN tuning and HNSW memory/build behavior need more attention | | **Pinecone Serverless** | The overriding goal is the least infrastructure work possible | Very easy launch, but query cost scales with namespace size, creating more sensitivity to data layout and sustained traffic | | **Elasticsearch/OpenSearch** | You already operate it or need first-class lexical + semantic search | Heavier and more complex if the requirement is principally vector retrieval | | **Weaviate** | You want an opinionated search platform with integrated hybrid-search features | More platform surface area than this application requires | | **Milvus/Zilliz** | You expect tens or hundreds of millions of vectors and need extensive scale controls | More machinery than a next-month, few-million-ticket deployment needs | pgvector is the closest alternative. It provides HNSW and IVFFlat, cosine/L2/inner-product search, normal SQL joins and transactions, and solid Node support. Its documentation also notes that approximate-index filtering happens after scanning unless you tune or use iterative scans, which is important for tenant-level filters. [pgvector capabilities](https://github.com/pgvector/pgvector), [Node integration](https://github.com/pgvector/pgvector-node), [filtering behavior](https://github.com/pgvector/pgvector/blob/master/README.md) Pinecone is also credible for this schedule, but its serverless queries are billed according to the size of the targeted namespace—currently one read unit per GB, subject to a minimum—so I would model production QPS before selecting it. [Pinecone cost model](https://docs.pinecone.io/guides/manage-cost/understanding-cost) Before committing, I’d run a two-day bake-off using 1–5% of real tickets. Measure: - p50/p95 latency at anticipated concurrency - recall@3 against an exact-search sample - behavior with real tenant/product/language filters - ingestion and re-embedding time - monthly cost at expected QPS Use cosine similarity if that is what the embedding model recommends, retrieve perhaps 10–20 candidates, and rerank them before returning the final three. For support tickets, that reranking step often improves relevance more than swapping between competent vector databases. Keep the embedding model name and version with every vector so a future re-embedding can be rolled out alongside the existing collection.
Antigravity2026-09-07Run 1
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Antigravity2026-09-07Run 1
qdrant.tech named first
3 measured vendors mentioned
agy 1.1.27 (gemini-3.7-flash-low)
Link to this runRecorded setup: sandbox=enabled; slash-commands=disabled; timeout=5m; operator configuration may apply.
Given your timeline (shipping next month), scale (a few million embeddings), tech stack (Node.js on standard cloud infra), and use case (support ticket retrieval / Top-3 similarity), here is the recommendation and evaluation.
Primary Recommendation: PostgreSQL with pgvector
If your architecture already uses PostgreSQL or a managed relational database, PostgreSQL + pgvector is the safest, most practical choice for shipping in a month.
Why it fits:
- Scale-appropriate: A few million 768- or 1536-dimensional vectors fit comfortably in memory/disk on a moderately sized managed instance (e.g., AWS Aurora/RDS, Google Cloud SQL, Supabase, Neon).
- Indexing: Supports HNSW (
hnsw) indexing for sub-10ms approximate nearest neighbor (ANN) retrieval and high recall. - Hybrid Filtering: Support tickets almost always require business filters (e.g.,
status = 'resolved',org_id = 123,category = 'billing',created_at > ...). Combining relational constraints, full-text search, and vector search in a single SQL query avoids multi-system synchronization bugs. - Node.js Ecosystem: Mature tooling via Prisma, Drizzle, TypeORM, or
pg/pgvectorclient libraries. - Operational Simplicity: No extra distributed system, cluster management, or separate billing/compliance reviews needed.
Strong Alternative (Pure Vector DB): Qdrant or Pinecone
If your tickets require high-throughput vector ingestion, serverless auto-scaling, or if you prefer an isolated search layer:
- Qdrant (Self-hosted or Managed Cloud):
- Rust-based, exceptionally fast, and supports advanced payload filtering directly during ANN search (pre-filtering).
- Excellent official Node.js / TypeScript SDK (
@qdrant/js-client-rest). - Very cost-effective and resource-efficient.
- Pinecone (Managed Serverless):
- Truly zero-ops: provision an index, upload vectors via the Node SDK, and query.
- Good choice if you want zero infrastructure footprint before launch, though vendor lock-in and ongoing managed costs are higher.
What Else Was Considered & Why Excluded
| Option | Pros | Why it was deprioritized |
|---|---|---|
| Milvus / Chroma | Popular open-source vector engines | Milvus has high operational and cluster overhead for a simple few-million vector deployment. Chroma is great for prototyping, but historically less robust for multi-million production scale with complex filtering compared to Qdrant or PG. |
| Elasticsearch / OpenSearch | Strong BM25 + hybrid lexical search | If you already run an ES/OpenSearch cluster, it’s viable; otherwise, running and sizing JVM-heavy clusters in a 1-month timeline introduces unnecessary operational burden and cost. |
| Redis (RedisVL) | In-memory, ultra-low latency | Keeping a few million float32 embeddings entirely in RAM gets unnecessarily expensive compared to SSD/mmap-backed HNSW in Postgres or Qdrant. |
Summary / Decision Matrix
- Already using Postgres? $\rightarrow$ Enable
pgvector, build an HNSW index, and query directly from Node.js. - Starting from scratch / prefer dedicated search microservice? $\rightarrow$ Use Qdrant (or Pinecone for zero-ops).
Original text
Given your timeline (**shipping next month**), scale (**a few million embeddings**), tech stack (**Node.js** on standard cloud infra), and use case (**support ticket retrieval / Top-3 similarity**), here is the recommendation and evaluation. --- ### **Primary Recommendation: PostgreSQL with `pgvector`** If your architecture already uses PostgreSQL or a managed relational database, **PostgreSQL + `pgvector`** is the safest, most practical choice for shipping in a month. #### **Why it fits:** * **Scale-appropriate:** A few million 768- or 1536-dimensional vectors fit comfortably in memory/disk on a moderately sized managed instance (e.g., AWS Aurora/RDS, Google Cloud SQL, Supabase, Neon). * **Indexing:** Supports **HNSW** (`hnsw`) indexing for sub-10ms approximate nearest neighbor (ANN) retrieval and high recall. * **Hybrid Filtering:** Support tickets almost always require business filters (e.g., `status = 'resolved'`, `org_id = 123`, `category = 'billing'`, `created_at > ...`). Combining relational constraints, full-text search, and vector search in a single SQL query avoids multi-system synchronization bugs. * **Node.js Ecosystem:** Mature tooling via Prisma, Drizzle, TypeORM, or `pg` / `pgvector` client libraries. * **Operational Simplicity:** No extra distributed system, cluster management, or separate billing/compliance reviews needed. --- ### **Strong Alternative (Pure Vector DB): Qdrant or Pinecone** If your tickets require high-throughput vector ingestion, serverless auto-scaling, or if you prefer an isolated search layer: * **Qdrant (Self-hosted or Managed Cloud):** * Rust-based, exceptionally fast, and supports advanced payload filtering directly during ANN search (pre-filtering). * Excellent official Node.js / TypeScript SDK (`@qdrant/js-client-rest`). * Very cost-effective and resource-efficient. * **Pinecone (Managed Serverless):** * Truly zero-ops: provision an index, upload vectors via the Node SDK, and query. * Good choice if you want zero infrastructure footprint before launch, though vendor lock-in and ongoing managed costs are higher. --- ### **What Else Was Considered & Why Excluded** | Option | Pros | Why it was deprioritized | | :--- | :--- | :--- | | **Milvus / Chroma** | Popular open-source vector engines | **Milvus** has high operational and cluster overhead for a simple few-million vector deployment. **Chroma** is great for prototyping, but historically less robust for multi-million production scale with complex filtering compared to Qdrant or PG. | | **Elasticsearch / OpenSearch** | Strong BM25 + hybrid lexical search | If you already run an ES/OpenSearch cluster, it’s viable; otherwise, running and sizing JVM-heavy clusters in a 1-month timeline introduces unnecessary operational burden and cost. | | **Redis (RedisVL)** | In-memory, ultra-low latency | Keeping a few million float32 embeddings entirely in RAM gets unnecessarily expensive compared to SSD/mmap-backed HNSW in Postgres or Qdrant. | --- ### **Summary / Decision Matrix** 1. **Already using Postgres?** $\rightarrow$ Enable `pgvector`, build an HNSW index, and query directly from Node.js. 2. **Starting from scratch / prefer dedicated search microservice?** $\rightarrow$ Use **Qdrant** (or **Pinecone** for zero-ops).
Antigravity2026-09-07Run 2
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Antigravity2026-09-07Run 2
qdrant.tech named first
3 measured vendors mentioned
agy 1.1.27 (gemini-3.7-flash-low)
Link to this runRecorded setup: sandbox=enabled; slash-commands=disabled; timeout=5m; operator configuration may apply.
For a project shipping next month in a Node.js environment on a standard cloud provider with a scale of a few million embeddings, here is the recommendation and the alternatives considered.
Primary Recommendation: PostgreSQL with pgvector
(Managed: AWS RDS / Aurora Postgres, Google Cloud SQL / AlloyDB, or Neon / Supabase)
Given your tight 1-month timeline, Node stack, and dataset size (2–5M vectors), PostgreSQL + pgvector is the lowest-risk and most practical choice.
Why it fits your requirements:
- Time-to-Market & Simplicity:
- No need to manage and secure a separate database cluster.
- Node ecosystem support is rock-solid (via
pg,Prisma,Drizzle,TypeORM, or@langchain/core). - Standard relational data (ticket body, customer ID, status, resolution timestamps) lives side-by-side with embeddings in the same ACID transaction.
- Handles "A Few Million" Efficiently:
- Using the HNSW (
hnsw) index inpgvector, query latency for top-3 nearest neighbors on 2–5M vectors (e.g., 768 or 1536 dimensions) is typically < 10–25ms. - With an index like
CREATE INDEX ON tickets USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);, memory footprint remains manageable (a few GBs of RAM for the index).
- Using the HNSW (
- Rich Metadata Filtering:
- Support ticket queries almost always need pre-filtering (e.g.,
WHERE organization_id = $1 AND product = $2). Postgres handles mixed relational filters + vector similarity effortlessly.
- Support ticket queries almost always need pre-filtering (e.g.,
Alternatives Considered
| Database | Why Consider It? | Why Not for Next Month? |
|---|---|---|
| Qdrant | Dedicated, ultra-fast (Rust-based), excellent Node SDK, first-class metadata payload filtering, highly memory-efficient. | Adds another stateful service to operate and monitor unless using Qdrant Cloud. If you already have Postgres in your stack, running a separate vector DB adds unnecessary operational overhead for just a few million records. (Top runner-up if you want a dedicated vector store). |
| Pinecone | Fully managed serverless, zero infrastructure ops, great Node SDK. | Vendor lock-in and pricing scaling at millions of vectors with regular queries. If you don't already have Postgres/infra ready to go, this is the easiest pure SaaS alternative, but costly over time. |
| OpenSearch / Elasticsearch (k-NN) | Great if you already use it for full-text ticket search and want hybrid search (keyword BM25 + dense semantic vector). | Heavier to run, higher RAM/heap footprint, and steeper operational learning curve if not already in your infrastructure. |
| Milvus / Weaviate | Highly scalable for tens/hundreds of millions of vectors. | Overkill architecture (many distributed components in Milvus) for a few million vectors; higher maintenance overhead for a short timeline. |
| Redis (Redis Stack / RedisVL) | In-memory speed, ultra-low latency. | Expensive RAM requirements to hold a few million high-dimensional vectors and indices entirely in memory. |
Summary / Suggested Architecture
- Database: PostgreSQL 16+ with
pgvectorextension enabled. - Index Type: HNSW (
vector_cosine_opsorvector_l2_opsdepending on your model normalization). - Node.js Client:
drizzle-ormor standardpgdriver using parameterized cosine distance (<=>operator):SELECT id, subject, resolution_notes, 1 - (embedding <=> $1) AS similarity FROM support_tickets WHERE status = 'resolved' ORDER BY embedding <=> $1 LIMIT 3;
Original text
For a project shipping **next month** in a **Node.js** environment on a standard cloud provider with a scale of **a few million embeddings**, here is the recommendation and the alternatives considered. --- ### **Primary Recommendation: PostgreSQL with `pgvector`** *(Managed: AWS RDS / Aurora Postgres, Google Cloud SQL / AlloyDB, or Neon / Supabase)* Given your tight 1-month timeline, Node stack, and dataset size (2–5M vectors), **PostgreSQL + `pgvector`** is the lowest-risk and most practical choice. #### Why it fits your requirements: 1. **Time-to-Market & Simplicity:** - No need to manage and secure a separate database cluster. - Node ecosystem support is rock-solid (via `pg`, `Prisma`, `Drizzle`, `TypeORM`, or `@langchain/core`). - Standard relational data (ticket body, customer ID, status, resolution timestamps) lives side-by-side with embeddings in the same ACID transaction. 2. **Handles "A Few Million" Efficiently:** - Using the **HNSW** (`hnsw`) index in `pgvector`, query latency for top-3 nearest neighbors on 2–5M vectors (e.g., 768 or 1536 dimensions) is typically **< 10–25ms**. - With an index like `CREATE INDEX ON tickets USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64);`, memory footprint remains manageable (a few GBs of RAM for the index). 3. **Rich Metadata Filtering:** - Support ticket queries almost always need pre-filtering (e.g., `WHERE organization_id = $1 AND product = $2`). Postgres handles mixed relational filters + vector similarity effortlessly. --- ### **Alternatives Considered** | Database | Why Consider It? | Why Not for Next Month? | | :--- | :--- | :--- | | **Qdrant** | Dedicated, ultra-fast (Rust-based), excellent Node SDK, first-class metadata payload filtering, highly memory-efficient. | Adds another stateful service to operate and monitor unless using Qdrant Cloud. If you already have Postgres in your stack, running a separate vector DB adds unnecessary operational overhead for just a few million records. *(Top runner-up if you want a dedicated vector store)*. | | **Pinecone** | Fully managed serverless, zero infrastructure ops, great Node SDK. | Vendor lock-in and pricing scaling at millions of vectors with regular queries. If you don't already have Postgres/infra ready to go, this is the easiest pure SaaS alternative, but costly over time. | | **OpenSearch / Elasticsearch (k-NN)** | Great if you already use it for full-text ticket search and want hybrid search (keyword BM25 + dense semantic vector). | Heavier to run, higher RAM/heap footprint, and steeper operational learning curve if not already in your infrastructure. | | **Milvus / Weaviate** | Highly scalable for tens/hundreds of millions of vectors. | Overkill architecture (many distributed components in Milvus) for a few million vectors; higher maintenance overhead for a short timeline. | | **Redis (Redis Stack / RedisVL)** | In-memory speed, ultra-low latency. | Expensive RAM requirements to hold a few million high-dimensional vectors and indices entirely in memory. | --- ### **Summary / Suggested Architecture** 1. **Database:** PostgreSQL 16+ with `pgvector` extension enabled. 2. **Index Type:** **HNSW** (`vector_cosine_ops` or `vector_l2_ops` depending on your model normalization). 3. **Node.js Client:** `drizzle-orm` or standard `pg` driver using parameterized cosine distance (`<=>` operator): ```sql SELECT id, subject, resolution_notes, 1 - (embedding <=> $1) AS similarity FROM support_tickets WHERE status = 'resolved' ORDER BY embedding <=> $1 LIMIT 3; ```
Antigravity2026-09-07Run 3
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Antigravity2026-09-07Run 3
qdrant.tech named first
3 measured vendors mentioned
agy 1.1.27 (gemini-3.7-flash-low)
Link to this runRecorded setup: sandbox=enabled; slash-commands=disabled; timeout=5m; operator configuration may apply.
Primary Recommendation: PostgreSQL with pgvector (Managed via RDS / Cloud SQL / Supabase / Neon)
Given your timeline (shipping next month), scale (a few million embeddings), and stack (Node on a standard cloud provider), PostgreSQL + pgvector is the strongest choice.
Why pgvector for this timeline & scale
- Fastest Time-to-Market & Simplest Architecture:
- You likely already run PostgreSQL or have standard tooling/IaC for it in AWS/GCP/Azure.
- Node support is seamless via standard ORMs/clients (
pg, Prisma, Drizzle, TypeORM). - You can store ticket metadata (ticket status, user ID, tags, timestamps, resolution text) alongside the embeddings in the same table, eliminating multi-system sync and dual-write issues.
- Handles "A Few Million" Easily:
- With HNSW (Hierarchical Navigable Small World) indexing in
pgvector, query latency for finding top-3 nearest neighbors in a ~2–5M vector dataset is typically sub-15ms. - Index build times and memory footprint are well within standard cloud instance sizes (e.g., 1536-dim embeddings for 3M rows require ~18–24GB of RAM to keep the HNSW index cached in memory; with half-precision/
halfvecor 768-dim models, it's even lower).
- With HNSW (Hierarchical Navigable Small World) indexing in
- Hybrid & Filtered Queries:
- Support ticket queries often require filtering (e.g., only closed tickets from the last 12 months in category X).
pgvectorhandles pre/post-filtering natively using standard SQLWHEREclauses.
- Support ticket queries often require filtering (e.g., only closed tickets from the last 12 months in category X).
Alternatives Considered & Why They Were Deprioritized
| Database | Verdict | Why It Wasn't the #1 Choice |
|---|---|---|
| Qdrant | Strong runner-up | Excellent dedicated vector database with great Node SDK, payload filtering, and native disk-backed indexing (quantization). However, it adds a new distinct database system to deploy, monitor, backup, and sync with your source-of-truth ticket DB before next month. |
| Pinecone | Best SaaS alternative | Fully managed, zero-ops vector DB with very low setup time. Downsides: vendor lock-in, recurring SaaS cost scaling with index size, and you still need to duplicate ticket metadata or perform a second query to fetch ticket text. |
| Milvus / Weaviate | Overkill | Purpose-built for 50M+ to billions of vectors or complex multi-tenant distributed setups. Higher operational complexity and infrastructure overhead than needed for a 1-month timeline. |
| OpenSearch / Elasticsearch (kNN) | Good if already in use | If your team already uses OpenSearch/Elasticsearch for ticket search, its dense vector plugin is a natural fit (and allows BM25 + vector hybrid search). If not already in your stack, the JVM operational overhead is unnecessary. |
| Redis (RediSearch) | Cost-prohibitive | Extremely fast, but stores everything in RAM, making it significantly more expensive for multi-million vector datasets compared to disk/OS-cached solutions. |
Quick Sizing & Implementation Tips for Node + pgvector
- Index Type: Use
HNSWoverIVFFlatfor better recall and query speed without requiring manual re-clustering:CREATE INDEX ON tickets USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64); - Memory / Dimension Optimization:
- If using OpenAI
text-embedding-3-small, you can reduce dimensions from 1536 to 512 or 768 with negligible recall drop. - Use
halfvec(16-bit float) inpgvectorto halve RAM requirements if memory is constrained.
- If using OpenAI
Original text
### Primary Recommendation: **PostgreSQL with `pgvector`** (Managed via RDS / Cloud SQL / Supabase / Neon) Given your timeline (**shipping next month**), scale (**a few million embeddings**), and stack (**Node on a standard cloud provider**), **PostgreSQL + `pgvector`** is the strongest choice. --- ### Why `pgvector` for this timeline & scale 1. **Fastest Time-to-Market & Simplest Architecture:** * You likely already run PostgreSQL or have standard tooling/IaC for it in AWS/GCP/Azure. * Node support is seamless via standard ORMs/clients (`pg`, Prisma, Drizzle, TypeORM). * You can store ticket metadata (ticket status, user ID, tags, timestamps, resolution text) alongside the embeddings in the same table, eliminating multi-system sync and dual-write issues. 2. **Handles "A Few Million" Easily:** * With **HNSW** (Hierarchical Navigable Small World) indexing in `pgvector`, query latency for finding top-3 nearest neighbors in a ~2–5M vector dataset is typically **sub-15ms**. * Index build times and memory footprint are well within standard cloud instance sizes (e.g., 1536-dim embeddings for 3M rows require ~18–24GB of RAM to keep the HNSW index cached in memory; with half-precision/`halfvec` or 768-dim models, it's even lower). 3. **Hybrid & Filtered Queries:** * Support ticket queries often require filtering (e.g., *only closed tickets from the last 12 months in category X*). `pgvector` handles pre/post-filtering natively using standard SQL `WHERE` clauses. --- ### Alternatives Considered & Why They Were Deprioritized | Database | Verdict | Why It Wasn't the #1 Choice | | :--- | :--- | :--- | | **Qdrant** | *Strong runner-up* | Excellent dedicated vector database with great Node SDK, payload filtering, and native disk-backed indexing (quantization). However, it adds a new distinct database system to deploy, monitor, backup, and sync with your source-of-truth ticket DB before next month. | | **Pinecone** | *Best SaaS alternative* | Fully managed, zero-ops vector DB with very low setup time. Downsides: vendor lock-in, recurring SaaS cost scaling with index size, and you still need to duplicate ticket metadata or perform a second query to fetch ticket text. | | **Milvus / Weaviate** | *Overkill* | Purpose-built for 50M+ to billions of vectors or complex multi-tenant distributed setups. Higher operational complexity and infrastructure overhead than needed for a 1-month timeline. | | **OpenSearch / Elasticsearch (kNN)** | *Good if already in use* | If your team already uses OpenSearch/Elasticsearch for ticket search, its dense vector plugin is a natural fit (and allows BM25 + vector hybrid search). If not already in your stack, the JVM operational overhead is unnecessary. | | **Redis (RediSearch)** | *Cost-prohibitive* | Extremely fast, but stores everything in RAM, making it significantly more expensive for multi-million vector datasets compared to disk/OS-cached solutions. | --- ### Quick Sizing & Implementation Tips for Node + `pgvector` * **Index Type:** Use `HNSW` over `IVFFlat` for better recall and query speed without requiring manual re-clustering: ```sql CREATE INDEX ON tickets USING hnsw (embedding vector_cosine_ops) WITH (m = 16, ef_construction = 64); ``` * **Memory / Dimension Optimization:** * If using OpenAI `text-embedding-3-small`, you can reduce dimensions from 1536 to 512 or 768 with negligible recall drop. * Use `halfvec` (16-bit float) in `pgvector` to halve RAM requirements if memory is constrained.
Antigravity2026-09-07Run 4
pinecone.io named first
3 measured vendors mentioned
ReadClose
Antigravity2026-09-07Run 4
pinecone.io named first
3 measured vendors mentioned
agy 1.1.27 (gemini-3.7-flash-low)
Link to this runRecorded setup: sandbox=enabled; slash-commands=disabled; timeout=5m; operator configuration may apply.
For a Node.js stack shipping next month with a dataset of a few million embeddings (e.g., 2M–5M vectors at 768 to 1536 dimensions), speed to production, operational simplicity, and metadata filtering capability are the main priorities.
Primary Recommendation: PostgreSQL + pgvector (Managed via Cloud SQL, Supabase, Neon, or AWS RDS/Aurora)
Why this is the best choice for shipping next month:
- Zero New Infrastructure: If you already use Postgres for your primary data/tickets, adding vector search requires no new databases to monitor, secure, backup, or sync across pipelines.
- Metadata & Hybrid Filtering: Support ticket similarity searches almost always require filters (e.g.,
WHERE tenant_id = '...' AND resolved_status = true AND created_at > ...). Postgres handles structured relational filters and vector distance natively in a single query without two-phase syncing. - Scale Fit: With the HNSW index (introduced in
pgvector0.5+), Postgres comfortably handles 3–10 million vectors with single-digit millisecond query latencies, provided you provision adequate RAM to fit the index. - Node.js Ecosystem: First-class support across Prisma, Drizzle, TypeORM, and
pg.
Alternative Consideration: Pinecone / Qdrant Cloud (Managed Dedicated Vector DB)
If you don't already have Postgres, or if you want completely outsourced operational overhead (memory sizing, indexing maintenance):
- Pinecone (Serverless):
- Pros: Near-zero setup time, great Node SDK (
@pinecone-database/pinecone), serverless pricing scales from zero, zero capacity planning for HNSW RAM. - Cons: Vendor lock-in, recurring API cost, ticket metadata must be synchronized between your primary DB and Pinecone.
- Pros: Near-zero setup time, great Node SDK (
- Qdrant (Managed Cloud or Docker):
- Pros: Written in Rust, highly performant payload (metadata) filtering, native Node client, open-source (can be self-hosted later if needed).
What Else Was Considered (and why not for next month):
| Option | Verdict | Why it was set aside |
|---|---|---|
| Milvus / Chroma (Self-hosted) | ❌ Pass | Milvus is robust for tens of millions of vectors but introduces heavy operational complexity (etcd, Pulsar/MinIO dependencies). Chroma is lighter but less mature for high-concurrency multi-million vector workloads. |
| OpenSearch / Elasticsearch (k-NN) | ⚠️ Conditional | Great if you already have an ES cluster for full-text ticket search (enabling easy hybrid BM25 + dense vector search), but memory-heavy and overkill to stand up from scratch in 3 weeks. |
| Redis (Redis Stack / RedisVL) | ❌ Pass | Extremely fast, but storing multi-million vector indices entirely in in-memory RAM is cost-inefficient compared to disk/RAM-hybrid solutions. |
Summary Decision Guide
- If you already use Postgres: Stick with PostgreSQL +
pgvectorusing an HNSW index. - If you have no existing DB infrastructure or want zero ops: Choose Pinecone Serverless or Qdrant Cloud.
Original text
For a Node.js stack shipping next month with a dataset of **a few million embeddings** (e.g., 2M–5M vectors at 768 to 1536 dimensions), speed to production, operational simplicity, and metadata filtering capability are the main priorities. --- ### **Primary Recommendation: PostgreSQL + `pgvector`** (Managed via Cloud SQL, Supabase, Neon, or AWS RDS/Aurora) #### **Why this is the best choice for shipping next month:** 1. **Zero New Infrastructure**: If you already use Postgres for your primary data/tickets, adding vector search requires no new databases to monitor, secure, backup, or sync across pipelines. 2. **Metadata & Hybrid Filtering**: Support ticket similarity searches almost always require filters (e.g., `WHERE tenant_id = '...' AND resolved_status = true AND created_at > ...`). Postgres handles structured relational filters and vector distance natively in a single query without two-phase syncing. 3. **Scale Fit**: With the **HNSW** index (introduced in `pgvector` 0.5+), Postgres comfortably handles 3–10 million vectors with single-digit millisecond query latencies, provided you provision adequate RAM to fit the index. 4. **Node.js Ecosystem**: First-class support across Prisma, Drizzle, TypeORM, and `pg`. --- ### **Alternative Consideration: Pinecone / Qdrant Cloud (Managed Dedicated Vector DB)** If you don't already have Postgres, or if you want completely outsourced operational overhead (memory sizing, indexing maintenance): * **Pinecone (Serverless)**: * **Pros**: Near-zero setup time, great Node SDK (`@pinecone-database/pinecone`), serverless pricing scales from zero, zero capacity planning for HNSW RAM. * **Cons**: Vendor lock-in, recurring API cost, ticket metadata must be synchronized between your primary DB and Pinecone. * **Qdrant (Managed Cloud or Docker)**: * **Pros**: Written in Rust, highly performant payload (metadata) filtering, native Node client, open-source (can be self-hosted later if needed). --- ### **What Else Was Considered (and why not for next month):** | Option | Verdict | Why it was set aside | | :--- | :--- | :--- | | **Milvus / Chroma (Self-hosted)** | ❌ Pass | Milvus is robust for tens of millions of vectors but introduces heavy operational complexity (etcd, Pulsar/MinIO dependencies). Chroma is lighter but less mature for high-concurrency multi-million vector workloads. | | **OpenSearch / Elasticsearch (k-NN)** | ⚠️ Conditional | Great if you already have an ES cluster for full-text ticket search (enabling easy hybrid BM25 + dense vector search), but memory-heavy and overkill to stand up from scratch in 3 weeks. | | **Redis (Redis Stack / RedisVL)** | ❌ Pass | Extremely fast, but storing multi-million vector indices entirely in in-memory RAM is cost-inefficient compared to disk/RAM-hybrid solutions. | --- ### **Summary Decision Guide** * **If you already use Postgres**: Stick with **PostgreSQL + `pgvector`** using an HNSW index. * **If you have no existing DB infrastructure or want zero ops**: Choose **Pinecone Serverless** or **Qdrant Cloud**.
Antigravity2026-09-07Run 5
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Antigravity2026-09-07Run 5
qdrant.tech named first
3 measured vendors mentioned
agy 1.1.27 (gemini-3.7-flash-low)
Link to this runRecorded setup: sandbox=enabled; slash-commands=disabled; timeout=5m; operator configuration may apply.
Given your requirements—Node.js ecosystem, a few million vectors, standard cloud hosting, and a 1-month delivery deadline—here is the recommendation:
Primary Recommendation: PostgreSQL + pgvector
For a tight 1-month timeline with a dataset in the single-digit millions, PostgreSQL with pgvector (using an HNSW index) is the safest and most pragmatic choice.
Why it fits your constraints:
- Zero New Infrastructure: If you already use Postgres for operational data (ticket metadata, user accounts, timestamps), you can store embeddings directly alongside tickets and perform metadata-filtered vector searches in a single SQL query without sync pipelines.
- First-Class Node.js Support: Works seamlessly with standard drivers/ORMs (
pg,prisma,typeorm,drizzle,pgvector-node,langchain/vectorstores). - Turnkey Managed Cloud Support: Available out of the box on AWS (RDS/Aurora Postgres), Google Cloud (Cloud SQL/AlloyDB), Azure, Supabase, and Neon.
- Performance & Scale: A few million 768d/1536d vectors with an HNSW index will comfortably deliver sub-50ms top-3 queries, provided the instance has enough RAM (roughly 16–32 GB) to keep the index in memory.
Alternatives Considered
| Database | Why Consider It? | Why Not for This Timeline / Scope? |
|---|---|---|
| Qdrant | Purpose-built vector DB in Rust; outstanding hybrid search, payload filtering, and native Node.js client. | Excellent runner-up. If you don't already have Postgres or want purpose-built vector search features, Qdrant Cloud is great, but introducing a dedicated database cluster adds unnecessary infrastructure complexity if Postgres is already present. |
| Pinecone | Fully managed serverless SaaS with zero index management or RAM provisioning. | Vendor lock-in, recurring pricing scales with index size/read throughput, and all ticket metadata must be synchronized/duplicated across network boundaries. |
| Milvus / Weaviate | Highly scalable, enterprise-grade vector databases. | Overkill for a few million records; self-hosting introduces significant operational and distributed-system overhead. |
| OpenSearch / Elasticsearch | Strong if you already rely on it for full-text ticket search. | High memory and indexing overhead for vector workloads compared to dedicated vector stores or pgvector. |
| Redis (Redis Stack / Vector) | Ultra-fast in-memory vector search. | Expensive RAM footprint for multi-million vector datasets unless embeddings are heavily quantized. |
Summary / Decision Rule
- If you already use PostgreSQL: Stick to PostgreSQL +
pgvector(HNSW). You will avoid writing dual-write sync pipelines and easily meet your deadline next month. - If you want a dedicated vector engine / managed API: Choose Qdrant (or Pinecone if you want zero ops).
Original text
Given your requirements—**Node.js ecosystem, a few million vectors, standard cloud hosting, and a 1-month delivery deadline**—here is the recommendation: --- ### **Primary Recommendation: PostgreSQL + `pgvector`** For a tight 1-month timeline with a dataset in the single-digit millions, **PostgreSQL with `pgvector`** (using an **HNSW** index) is the safest and most pragmatic choice. #### Why it fits your constraints: 1. **Zero New Infrastructure**: If you already use Postgres for operational data (ticket metadata, user accounts, timestamps), you can store embeddings directly alongside tickets and perform metadata-filtered vector searches in a single SQL query without sync pipelines. 2. **First-Class Node.js Support**: Works seamlessly with standard drivers/ORMs (`pg`, `prisma`, `typeorm`, `drizzle`, `pgvector-node`, `langchain/vectorstores`). 3. **Turnkey Managed Cloud Support**: Available out of the box on AWS (RDS/Aurora Postgres), Google Cloud (Cloud SQL/AlloyDB), Azure, Supabase, and Neon. 4. **Performance & Scale**: A few million 768d/1536d vectors with an HNSW index will comfortably deliver sub-50ms top-3 queries, provided the instance has enough RAM (roughly 16–32 GB) to keep the index in memory. --- ### **Alternatives Considered** | Database | Why Consider It? | Why Not for This Timeline / Scope? | | :--- | :--- | :--- | | **Qdrant** | Purpose-built vector DB in Rust; outstanding hybrid search, payload filtering, and native Node.js client. | Excellent runner-up. If you don't already have Postgres or want purpose-built vector search features, Qdrant Cloud is great, but introducing a dedicated database cluster adds unnecessary infrastructure complexity if Postgres is already present. | | **Pinecone** | Fully managed serverless SaaS with zero index management or RAM provisioning. | Vendor lock-in, recurring pricing scales with index size/read throughput, and all ticket metadata must be synchronized/duplicated across network boundaries. | | **Milvus / Weaviate** | Highly scalable, enterprise-grade vector databases. | Overkill for a few million records; self-hosting introduces significant operational and distributed-system overhead. | | **OpenSearch / Elasticsearch** | Strong if you already rely on it for full-text ticket search. | High memory and indexing overhead for vector workloads compared to dedicated vector stores or pgvector. | | **Redis (Redis Stack / Vector)** | Ultra-fast in-memory vector search. | Expensive RAM footprint for multi-million vector datasets unless embeddings are heavily quantized. | --- ### **Summary / Decision Rule** * **If you already use PostgreSQL:** Stick to **PostgreSQL + `pgvector` (HNSW)**. You will avoid writing dual-write sync pipelines and easily meet your deadline next month. * **If you want a dedicated vector engine / managed API:** Choose **Qdrant** (or **Pinecone** if you want zero ops).
Claude Code2026-08-16Run 1
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Claude Code2026-08-16Run 1
qdrant.tech named first
3 measured vendors mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runPrzy kilku milionach embeddingów i trzech najbliższych sąsiadach jako celu, to nie jest jeszcze skala wymagająca dedykowanej bazy wektorowej. Postawiłbym na pgvector na Postgresie (indeks HNSW), jeśli support bot już ma gdzieś Postgresa: brak nowej infrastruktury, dojrzały klient Node (pg / postgres), a HNSW przy kilku milionach wektorów daje zapytania rzędu pojedynczych dziesiątek ms. Jeśli nie ma jeszcze Postgresa w stosie, alternatywą byłby zarządzany serwis typu Qdrant Cloud albo Pinecone, żeby nie tracić czasu (macie miesiąc) na operacyjną stronę własnego klastra.
Odrzuciłem samodzielnie hostowane Milvus/Weaviate, bo narzut operacyjny (klaster, monitoring, upgrade'y) nie ma uzasadnienia przy tej skali i deadline'ie, chyba że zespół już je zna. Elasticsearch/OpenSearch z kNN brałbym pod uwagę tylko, gdyby ELK już był w stosie.
Macie już gdzieś Postgresa albo inną bazę, na której moglibyśmy się oprzeć, czy stack jest czysty i wybieramy od zera?
Original text
Przy kilku milionach embeddingów i trzech najbliższych sąsiadach jako celu, to nie jest jeszcze skala wymagająca dedykowanej bazy wektorowej. Postawiłbym na **pgvector na Postgresie** (indeks HNSW), jeśli support bot już ma gdzieś Postgresa: brak nowej infrastruktury, dojrzały klient Node (`pg` / `postgres`), a HNSW przy kilku milionach wektorów daje zapytania rzędu pojedynczych dziesiątek ms. Jeśli nie ma jeszcze Postgresa w stosie, alternatywą byłby zarządzany serwis typu Qdrant Cloud albo Pinecone, żeby nie tracić czasu (macie miesiąc) na operacyjną stronę własnego klastra. Odrzuciłem samodzielnie hostowane Milvus/Weaviate, bo narzut operacyjny (klaster, monitoring, upgrade'y) nie ma uzasadnienia przy tej skali i deadline'ie, chyba że zespół już je zna. Elasticsearch/OpenSearch z kNN brałbym pod uwagę tylko, gdyby ELK już był w stosie. Macie już gdzieś Postgresa albo inną bazę, na której moglibyśmy się oprzeć, czy stack jest czysty i wybieramy od zera?
Claude Code2026-08-16Run 2
qdrant.tech named first
3 measured vendors mentioned
ReadClose
Claude Code2026-08-16Run 2
qdrant.tech named first
3 measured vendors mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runDla kilku milionów wektorów i zapytań top-3 w Node na "normalnym" cloud, na Twoim miejscu postawiłbym na pgvector (rozszerzenie Postgresa z indeksem HNSW) - o ile gdziekolwiek w stacku już macie Postgresa (RDS/Cloud SQL/Supabase/Neon). To zero nowej infrastruktury do nauki miesiąc przed shipem, dojrzały klient node-postgres, a HNSW przy kilku milionach wektorów daje sensowne opóźnienia bez strojenia.
Jeśli Postgresa nie macie, drugi wybór to Qdrant (self-hosted w Dockerze albo ich managed cloud) - lekki, dobry klient JS, sensowne defaulty bez tuningu.
Co odrzuciłem i dlaczego:
- Pinecone - najprostsze API, ale to nowy vendor, koszt i dane poza waszą infrastrukturą przy skali, która tego nie wymaga.
- Milvus/Zilliz - zbudowany pod miliardy wektorów, dla kilku milionów to niepotrzebny narzut operacyjny na miesiąc przed deadline'em.
- Elasticsearch/OpenSearch kNN - sensowny tylko jeśli już macie ES pod wyszukiwanie ticketów i chcecie hybrydowe (keyword + vector) - w innym wypadku cięższy w utrzymaniu niż pgvector.
- In-memory (hnswlib-node) bez bazy - działałoby wydajnościowo, ale persistencję, replikację i backup bierzecie wtedy na siebie ręcznie - ryzykowne przy tym terminie.
Czy już macie gdzieś Postgresa w stacku? To determinuje, czy idziemy w pgvector, czy w Qdrant.
Original text
Dla kilku milionów wektorów i zapytań top-3 w Node na "normalnym" cloud, na Twoim miejscu postawiłbym na **pgvector** (rozszerzenie Postgresa z indeksem HNSW) - o ile gdziekolwiek w stacku już macie Postgresa (RDS/Cloud SQL/Supabase/Neon). To zero nowej infrastruktury do nauki miesiąc przed shipem, dojrzały klient node-postgres, a HNSW przy kilku milionach wektorów daje sensowne opóźnienia bez strojenia. Jeśli Postgresa nie macie, drugi wybór to **Qdrant** (self-hosted w Dockerze albo ich managed cloud) - lekki, dobry klient JS, sensowne defaulty bez tuningu. Co odrzuciłem i dlaczego: - **Pinecone** - najprostsze API, ale to nowy vendor, koszt i dane poza waszą infrastrukturą przy skali, która tego nie wymaga. - **Milvus/Zilliz** - zbudowany pod miliardy wektorów, dla kilku milionów to niepotrzebny narzut operacyjny na miesiąc przed deadline'em. - **Elasticsearch/OpenSearch kNN** - sensowny tylko jeśli już macie ES pod wyszukiwanie ticketów i chcecie hybrydowe (keyword + vector) - w innym wypadku cięższy w utrzymaniu niż pgvector. - **In-memory (hnswlib-node) bez bazy** - działałoby wydajnościowo, ale persistencję, replikację i backup bierzecie wtedy na siebie ręcznie - ryzykowne przy tym terminie. Czy już macie gdzieś Postgresa w stacku? To determinuje, czy idziemy w pgvector, czy w Qdrant.
Claude Code2026-08-16Run 3
pinecone.io named first
3 measured vendors mentioned
ReadClose
Claude Code2026-08-16Run 3
pinecone.io named first
3 measured vendors mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runPostawiłbym na pgvector w Postgresie (indeks HNSW) - kilka milionów wektorów to dokładnie zakres, w którym radzi sobie dobrze, ma dojrzałe klienty node (pg, drizzle, prisma) i nie dokłada nowego systemu do utrzymania, co przy deadline'u za miesiąc sporo waży. Rozważałem jeszcze Pinecone (najszybciej wdrożyć, w pełni managed, ale nowy vendor, koszt i lock-in) oraz Qdrant/Weaviate (mocniejsze przy czystym wektorowym search i filtrowaniu metadanych, ale to kolejna usługa do postawienia i pilnowania). Główny kompromis: pgvector jest wolniejszy przy naprawdę dużej skali i mniej wygodny przy złożonym filtrowaniu niż dedykowana baza wektorowa, ale przy top-3 z kilku milionów rekordów to nie będzie problem.
Macie już Postgresa w stacku, czy to byłaby zupełnie nowa baza?
Original text
Postawiłbym na **pgvector w Postgresie** (indeks HNSW) - kilka milionów wektorów to dokładnie zakres, w którym radzi sobie dobrze, ma dojrzałe klienty node (pg, drizzle, prisma) i nie dokłada nowego systemu do utrzymania, co przy deadline'u za miesiąc sporo waży. Rozważałem jeszcze **Pinecone** (najszybciej wdrożyć, w pełni managed, ale nowy vendor, koszt i lock-in) oraz **Qdrant/Weaviate** (mocniejsze przy czystym wektorowym search i filtrowaniu metadanych, ale to kolejna usługa do postawienia i pilnowania). Główny kompromis: pgvector jest wolniejszy przy naprawdę dużej skali i mniej wygodny przy złożonym filtrowaniu niż dedykowana baza wektorowa, ale przy top-3 z kilku milionów rekordów to nie będzie problem. Macie już Postgresa w stacku, czy to byłaby zupełnie nowa baza?
Claude Code2026-08-16Run 4
qdrant.tech named first
2 measured vendors mentioned
ReadClose
Claude Code2026-08-16Run 4
qdrant.tech named first
2 measured vendors mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runDla kilku milionów embeddingów i zapytania "top-3 podobne" postawiłbym na pgvector na Postgresie (indeks HNSW) - jeśli macie już Postgresa pod danymi ticketów, to najmniej ruchu: jeden system zamiast dwóch, ACID, dobry klient node-postgres, a HNSW przy tej skali (kilka milionów wektorów) daje latencję rzędu pojedynczych milisekund przy top-3. Główny kompromis: przy bardzo dużym ruchu zapisowym lub gdy potrzebujecie zaawansowanego hybrid search / multi-tenant filteringu, dedykowany silnik radzi sobie lepiej.
Co jeszcze brałem pod uwagę:
- Qdrant (self-hosted lub cloud) - szybszy i bardziej "vector-native" niż pgvector przy tej skali, lekki, dobre SDK do Node, ale to nowy serwis do utrzymania miesiąc przed deploymentem.
- Pinecone - zero-ops, najszybsze wdrożenie, ale koszt per-query i vendor lock-in, a to jednak drugi system obok waszej bazy ticketów.
- Redis / OpenSearch / Mongo Atlas Vector - sensowne tylko jeśli już macie któryś z nich w stacku jako główną bazę; nie warłbym wprowadzać ich tylko dla tej funkcji.
Czy macie już Postgresa pod ticketami? To rozstrzyga między pgvector a Qdrant/Pinecone.
Original text
Dla kilku milionów embeddingów i zapytania "top-3 podobne" postawiłbym na **pgvector na Postgresie** (indeks HNSW) - jeśli macie już Postgresa pod danymi ticketów, to najmniej ruchu: jeden system zamiast dwóch, ACID, dobry klient node-postgres, a HNSW przy tej skali (kilka milionów wektorów) daje latencję rzędu pojedynczych milisekund przy top-3. Główny kompromis: przy bardzo dużym ruchu zapisowym lub gdy potrzebujecie zaawansowanego hybrid search / multi-tenant filteringu, dedykowany silnik radzi sobie lepiej. Co jeszcze brałem pod uwagę: - **Qdrant** (self-hosted lub cloud) - szybszy i bardziej "vector-native" niż pgvector przy tej skali, lekki, dobre SDK do Node, ale to nowy serwis do utrzymania miesiąc przed deploymentem. - **Pinecone** - zero-ops, najszybsze wdrożenie, ale koszt per-query i vendor lock-in, a to jednak drugi system obok waszej bazy ticketów. - **Redis / OpenSearch / Mongo Atlas Vector** - sensowne tylko jeśli już macie któryś z nich w stacku jako główną bazę; nie warłbym wprowadzać ich tylko dla tej funkcji. Czy macie już Postgresa pod ticketami? To rozstrzyga między pgvector a Qdrant/Pinecone.
Claude Code2026-08-16Run 5
pinecone.io named first
3 measured vendors mentioned
ReadClose
Claude Code2026-08-16Run 5
pinecone.io named first
3 measured vendors mentioned
claude 2.1.233 (Claude Code) (sonnet)
Link to this runDla kilku milionów embeddingów i zapytań top-3, przy Node.js i deadlinie za miesiąc, postawiłbym na Postgres + pgvector (indeks HNSW) - o ile macie już Postgresa pod dansymi ticketów, co przy supporcie jest niemal pewne. To wciąż jedna baza danych, jeden connection pool, embeddingi trzymane obok samych ticketów bez syncowania dwóch systemów, a HNSW w pgvector spokojnie obsługuje dziesiątki milionów wektorów, więc "kilka milionów" to nie jest tu wyzwanie skalowe.
Rozważałem też:
- Pinecone - zero ops, najszybszy do wdrożenia, ale kolejny vendor, koszt subskrypcyjny i osobny system do trzymania w synchronizacji z danymi ticketów.
- Qdrant / Weaviate - dobre SDK do Node, można postawić managed cloud albo self-host; sensowna alternatywa jeśli nie ma jeszcze Postgresa w stacku.
- Redis (vector search) - szybkie, ale przy milionach wektorów w pamięci koszt RAM zaczyna boleć.
- Milvus - w praktyce przewymiarowany na ten wolumen, ma sens dopiero przy dziesiątkach-setkach milionów+.
Główny tradeoff: pgvector to mniej ruchomych części i mniej ryzyka na deadline, ale tuning indeksu i skalowanie bierzecie na siebie; Pinecone jest szybszy do postawienia, ale to nowy vendor i koszt na stałe. Jeśli nie macie dziś Postgresa w stacku, wziąłbym Qdrant Cloud jako kompromis.
Original text
Dla kilku milionów embeddingów i zapytań top-3, przy Node.js i deadlinie za miesiąc, postawiłbym na **Postgres + pgvector** (indeks HNSW) - o ile macie już Postgresa pod dansymi ticketów, co przy supporcie jest niemal pewne. To wciąż jedna baza danych, jeden connection pool, embeddingi trzymane obok samych ticketów bez syncowania dwóch systemów, a HNSW w pgvector spokojnie obsługuje dziesiątki milionów wektorów, więc "kilka milionów" to nie jest tu wyzwanie skalowe. Rozważałem też: - **Pinecone** - zero ops, najszybszy do wdrożenia, ale kolejny vendor, koszt subskrypcyjny i osobny system do trzymania w synchronizacji z danymi ticketów. - **Qdrant / Weaviate** - dobre SDK do Node, można postawić managed cloud albo self-host; sensowna alternatywa jeśli nie ma jeszcze Postgresa w stacku. - **Redis (vector search)** - szybkie, ale przy milionach wektorów w pamięci koszt RAM zaczyna boleć. - **Milvus** - w praktyce przewymiarowany na ten wolumen, ma sens dopiero przy dziesiątkach-setkach milionów+. Główny tradeoff: pgvector to mniej ruchomych części i mniej ryzyka na deadline, ale tuning indeksu i skalowanie bierzecie na siebie; Pinecone jest szybszy do postawienia, ale to nowy vendor i koszt na stałe. Jeśli nie macie dziś Postgresa w stacku, wziąłbym Qdrant Cloud jako kompromis.
Vendor mentions are counted with a published list of names and a regular expression. How the counting works or back to the tally