Vector database: A knowledge compendium on AI assistant memory
In short
This compendium explains what a vector database is and how it serves as the memory of AI assistants. It covers embeddings, indexing methods (HNSW, IVF+PQ), chunking, metadata filters, and update and versioning strategies that determine the quality of system responses.
What is a vector database and how does it build the memory of AI assistants?
Imagination suggests that the memory of a digital assistant must look like an enormous filing cabinet, where every document waits for its label. In reality, a vector database (vector database) operates on an entirely different principle - closer to a library organized not by title or date, but by the meaning of the stored material. Two documents with similar content sit next to each other regardless of their physical form, filename, or creation date.
This distinction has deeply practical consequences for AI system architecture. A classical relational database (relational database, or RDBMS) stores structured records and retrieves them through exact matching of column values. A query "find documents about credit risk" in a relational environment requires those exact words to be present in a specific column. A vector database solves this problem differently: instead of labels, it stores a mathematical representation of a document's intent - a multidimensional numerical vector (a sequence of numbers describing the meaning of a word or sentence) that encodes semantics (meaning), not syntax (how a word is written).
This infrastructure has become the foundation of the long-term AI memory of modern assistants - not because it is modern and fashionable, but because it solves a concrete engineering problem: how to find, in a fraction of a second, among millions of stored fragments, those that are semantically closest to a given question.
Brief definition and key role in RAG architecture
A vector database is a data management system designed to store, index, and search multidimensional mathematical objects called embedding vectors (embedding vectors), enabling similarity search (similarity search) based on semantics rather than exact matching.
In the RAG (Retrieval-Augmented Generation) pipeline, the vector database serves as the external memory of a language model (Large Language Model, LLM). The RAG architecture resolves a fundamental limitation of LLMs: a model has a fixed, frozen knowledge horizon and cannot reach beyond its training data. RAG breaks through this limitation by building a three-stage pipeline: a user query goes to the vector database, which returns semantically related knowledge fragments, and those fragments then feed the LLM's context before a response is generated.
The vector database is therefore an intermediary between the query and the model - it is responsible for the quality of the information the model will ever see. If the retrieval mechanism delivers irrelevant or outdated fragments, the model generates a partial or incorrect response, regardless of its own quality. This is precisely why the technical choices surrounding the vector database have a direct bearing on the visible competence of the assistant.

Vectors and similarity search
An embedding is a numerical representation of text - a sequence of hundreds or thousands of floating-point values that an embedding model (separate from the LLM) assigns to a given fragment. Semantically similar phrases - "operational risk management" and "risk control in a company" - receive vectors whose values resemble one another. Geometrically, these points in multidimensional space lie close together.
Similarity search (similarity search) involves finding the points in vector space that lie closest to the vector representing the query. The most commonly used measure is cosine similarity (cosine similarity), which measures the angle between vectors regardless of their length - which means differences in document length do not distort the assessment of topical proximity.
The contrast with classical full-text search (keyword search, e.g. BM25) is fundamental. Keyword search operates on lexical matching: the phrase "return on equity" will not return a document containing only the expression "ROE" or "return on assets", even though all three concepts semantically overlap. Vector search identifies that proximity automatically - provided the embedding model was trained on domain-relevant data.

Physical storage model: Anatomy of a vector record and the role of metadata
Every unit of knowledge in a vector database is stored as a vector record (vector record) - a structure composed of three independent but related components: a record identifier, an embedding vector, and a set of metadata.
The identifier acts as a primary key - a unique label that allows the management system to reference a specific object without scanning the vector index. The embedding vector is the heart of the record, containing the semantic digital fingerprint of the knowledge fragment. Metadata consists of descriptive attributes - creation date, document source, content type, language, version - that enable filtering without engaging the semantic engine.
The design of this record is not a trivial decision. An impoverished metadata set will later force a full scan of the entire vector index in search of relevant fragments, directly slowing system responses. A well-considered metadata structure makes it possible to narrow the search space before the costly similarity computations are even triggered.
Numerical representation and serialization of information
A standard embedding vector generated by popular embedding models contains between 384 and 3,072 floating-point values (floating-point numbers) in 32-bit format (float32), though for performance reasons some systems move to 16-bit (float16) or even 8-bit representations with lossy quantization (quantization - a compression technique that approximates vector values using fewer bits at the cost of a slight loss of precision).
Each vector therefore occupies approximately 1.5 KB to 12 KB of raw memory, which across collections of tens of millions of records quickly determines the total RAM budget of the server. Writing this data to disk is most commonly done in binary columnar formats (compact binary files optimized for fast sequential reads), which perform more efficiently than the text-based JSON format - as a result, the index loads faster after each service restart.
It is worth distinguishing the raw vector from its indexed representation. A raw vector is a static record - it does not search for anything on its own. The index is a separate data structure built from many vectors, which establishes shortcuts (pre-computed paths) needed for fast searching. It is the index - not the raw vectors - that determines how quickly the system responds to queries.
Metadata as traditional filters in a semantic world
Storing vectors without accompanying metadata creates a system that is semantically powerful but operationally blind. An AI assistant handling a company's legal documentation needs to know not only that a given fragment concerns rental agreements, but also that it comes from the revision in force since the current quarter - not from a version three years old.
The pre-filtering strategy (pre-filtering) means narrowing the candidate set using metadata conditions before running vector similarity computations. The analogy: before reaching for the semantic magnifying glass, all documents predating 2022, in languages other than the target language, and outside the finance department, are filtered out of the catalog first. The similarity engine then operates only on the already-selected set, drastically reducing its load.
Post-filtering (post-filtering) reverses the order: the system first executes the vector query and returns the N nearest neighbors, then discards those that do not meet the metadata conditions. This method is straightforward to implement, but carries a risk: if the filter rejects too many results from the returned set, the assistant receives too few useful fragments - even though they exist in the database.
The choice between pre- and post-filtering is one of the first design trade-offs that arise when building a RAG system. Pre-filtering is more efficient when metadata selectivity is high; post-filtering is easier to implement but susceptible to the precision-recall trade-off at the retrieval stage.

The algorithmic core: Mechanics of data indexing
Vector indexes are data structures built in advance that enable fast similarity search without needing to compare the query against every record individually.
Without an index, searching a vector database would proceed by brute-force comparison of the query against every single record - a method known as brute-force or exact nearest neighbor search (ENN). With a collection of one million records and vectors of 1,536 dimensions, every individual query would require a billion multiplication operations. Indexes solve this problem by constructing data structures in advance that drastically restrict the search space at the cost of an acceptable, controlled loss of precision - the algorithm becomes an Approximate Nearest Neighbor (ANN) search.
No single ANN algorithm is optimal for all cases. Designing indexes is a matter of consciously managing a trade-off triangle: query speed (latency, throughput), memory consumption, and accuracy (recall) - and these three dimensions cannot all be maximized simultaneously.

HNSW: Navigating multi-layer highway networks
HNSW (Hierarchical Navigable Small World) is currently the most widely used indexing method in production vector systems, combining high accuracy with low query latency.
The intuition behind the algorithm: imagine a road network where motorways connect distant nodes at a coarse grain, regional roads connect neighboring towns, and local streets connect individual addresses. An HNSW search starts at the highest layer of the graph - the motorway equivalent - identifying the approximate neighborhood of the target. It then descends layer by layer, narrowing the search to an increasingly dense grid of local connections, until it reaches the nearest neighbors at the level of individual records.
The layer hierarchy allows the algorithm to skip comparisons with millions of distant records - the motorway layer identifies the right region before the algorithm descends to local connections. This means search time grows very slowly as the database expands - the algorithm is therefore many times faster than scanning all records sequentially. The cost is memory: the HNSW graph must keep all connections between nodes in RAM, which at hundreds of millions of records represents a serious hardware requirement.
IVF and PQ: Clustering objects and compressing memory
IVF (Inverted File Index) approaches the scaling problem through an analogy to a filing system. Before any actual search, the entire vector collection is grouped into clusters (typically using k-means), and each cluster receives its centroid. When a query arrives, the system first identifies a few nearest centroids, then conducts detailed search exclusively within those selected clusters.
IVF alone does not solve the working memory problem - each vector still needs to be stored in full. This is where PQ (Product Quantization) comes in, a lossy compression technique (one that reduces data size at the cost of a minor loss of precision). PQ splits a vector into smaller sub-vectors, and each sub-vector is replaced by a simplified code from a pre-prepared codebook. As a result, a 1,536-dimension vector originally weighing 6 KB can be stored in just a handful of bytes.
The IVF+PQ combination is particularly valuable when RAM availability is limited: a collection that would not fit in working memory in full-dimensional format becomes indexable. The price is a certain loss of accuracy measured by recall - vectors are no longer compared exactly, but based on their compressed approximations.
The trade-off triangle: Speed, memory consumption, and accuracy
Every decision about the choice of indexing algorithm is, in essence, a business decision. The three axes of this triangle - query speed (latency, throughput), memory consumption, and accuracy (recall) - exist in irresolvable tension. A brief comparison of the main approaches:
-
HNSW - low latency, high recall, high RAM consumption; the right choice for systems requiring responses under 10 ms with collection sizes up to a few hundred million records.
-
IVF+PQ - moderate latency, memory compression, some recall degradation; justified for billion-scale collections or heavily constrained hardware budgets.
-
Tree-based algorithms (e.g. Annoy, Approximate Nearest Neighbors Oh Yeah) - deterministic build time and straightforward implementation, but weaker scaling on large vector dimensions.
In other words: if a system serves a brand knowledge base where an incorrect response is unacceptable (e.g. legal support, regulatory documentation), result accuracy must be the priority - in which case HNSW with an appropriately tuned parameter controlling how broadly the algorithm explores the graph is the right choice. If infrastructure cost matters more and accuracy at the 90-95% level is sufficient, IVF+PQ with a well-chosen nprobe parameter (which determines how many clusters the algorithm checks during a search) will prove the more economical solution.
Processing information: The lifecycle of a vector from ingestion to deletion
Before any knowledge fragment reaches the vector database as an indexed record, it passes through a series of deterministic transformations that together determine how relevant later search results will be. This ingest pipeline is - contrary to appearances - equally important as the choice of indexing algorithm. A perfect index fed with poorly prepared fragments will deliver inconsistent or incomplete information to the AI assistant.
The main stages of the database ingestion process include: text segmentation (chunking), embedding generation, metadata assignment, validation and writing to the database, and then - as knowledge evolves - updating and controlled deletion of records.

Text segmentation (chunking) and transformation
A raw document - a financial report, a product specification, a collection of legal articles - cannot be processed as a single object. Long documents generate long vectors that average out the semantics of the whole, losing precision on detailed queries. Chunking (chunking, segmentation) involves dividing a document into smaller, semantically coherent fragments before they are processed by the embedding model.
An example: a detailed financial report of 80 pages may contain sections on capital structure, currency risk, segment analysis, and management forecasts. Chunking with a 400-token window and 50-token overlap (overlap) will ensure that sentences at split boundaries appear in both the preceding and following fragment - which prevents context loss at critical junctions. A chunk that is too short loses context; one that is too long dilutes semantics.
After segmentation, each chunk passes through the embedding model, which returns an embedding vector. Before writing to the database, basic validation is worthwhile: detecting empty fragments, duplicate chunks, text extraction errors, and metadata inconsistencies. Poor validation at this stage directly reduces recall (the proportion of relevant results the system actually retrieves) and introduces noise into retrieval - even with a well-configured index. The vector is then paired with the fragment's metadata (source document identifier, page number, date, document type) and written to the database.
Updates, versioning, and deletion
A vector database becomes operationally complex the moment the knowledge it stores begins to evolve. Modifying a vector record is a non-trivial operation: simply overwriting floating-point values is not sufficient - it requires rebuilding the relevant portion of the HNSW graph or re-indexing the IVF clusters that referenced that record.
A vector update in practice is implemented as a delete-then-insert operation: the old record is marked as deleted (tombstone), and the new one - with the updated vector and metadata - is inserted as a separate entry. Many vector systems perform periodic compaction operations (compaction), which physically remove orphaned records and rebuild the index, restoring the optimal graph structure.
Hard deletion (hard delete) is particularly important from a compliance perspective: if a source document has legally expired or contained erroneous information, its vector representations must be removed from the database with a guarantee of irreversibility - not merely flagged for skipping.
Memory versioning of the assistant is a strategy for preventing unintentional degradation of responses during knowledge base updates. Maintaining parallel collections - a production version and a newly prepared one - allows for atomic switching, with the ability to immediately roll back to the previous state if the new version proves less accurate in recall testing.
How vector architecture determines brand perception in the AI era
All the layers described - from the vector serialization format, through the HNSW graph architecture, the precision of chunking, to the configuration of metadata filters - combine into one coherent mechanism: the operational memory of the AI assistant. And that memory is not a transparent background technology. Its quality directly determines what the assistant knows, what it does not know, and with what degree of confidence it can provide an answer.
An assistant using a vector database poorly prepared in terms of chunking will return fragments cut off mid-argument - and the LLM will fill in the missing context with its own generalization, often incorrect. An assistant using an index that too frequently misses the right results will regularly lose important knowledge fragments, making its responses appear confident while actually being incomplete. An assistant with outdated records - without a versioning mechanism and controlled deletion - will give answers based on stale data, undermining user trust in the entire brand.
The practical consequence is straightforward: if public mentions of a brand are concentrated around a specific product category, the assistant will locate it there precisely, regardless of how broad the actual offering is. If the dominant narratives are negative or ambivalent, the brand's embedding will find itself in the vicinity of vectors associated with those connotations - and those are the ones that will be retrieved as context during user queries. Control over what appears in public sources, and how, therefore becomes not merely a matter of reputation in the traditional sense, but literally an entry point into the semantic memory of AI systems.