Back to blog
·7 min read·BitAtlas Team

Encrypted Vector Embeddings: Semantic Search Without Exposing Your Data

Learn how to store and query AI vector embeddings without exposing raw data, using encryption techniques that preserve search semantics while keeping sensitive content private.

vector embeddingsencrypted searchsemantic searchprivacy-preserving AIhomomorphic encryptionclient-side encryptionAI infrastructure

Vector databases have become the backbone of modern AI applications: semantic search, RAG pipelines, recommendation engines, and agent memory all rely on storing and querying high-dimensional embeddings. But most deployments send raw embeddings to a hosted vector store — and those embeddings are not as opaque as they appear.

Research has repeatedly shown that embeddings can be inverted: given enough samples, an attacker can reconstruct the original text with surprising fidelity. For applications handling personal health records, legal documents, or proprietary code, that is an unacceptable privacy risk. This post covers practical strategies for encrypting vector embeddings while preserving enough structure for useful semantic search.

Why Embeddings Are Sensitive

An embedding is a lossy representation, but loss is not the same as privacy. Two documents about the same topic will have similar embeddings — that geometric relationship is precisely what makes similarity search work. An adversary with access to your embedding store can:

  • Cluster embeddings to infer topic distributions across your corpus
  • Use membership inference to determine whether a specific document was indexed
  • Reconstruct text using embedding inversion models trained on your encoder family
  • Link records across datasets by comparing embeddings from different stores

Standard at-rest encryption (AES-256 on disk) protects against a compromised hard drive but does nothing while the database processes queries. The threat model for embedding privacy requires protecting data during similarity computation, not just at rest.

Strategy 1: Client-Side Embedding + Trusted Execution

The simplest defence is ensuring the vector database never sees the embedding in cleartext. Generate embeddings in a trusted environment — your server, not a SaaS API — and encrypt each vector before storing it.

from cryptography.fernet import Fernet
import numpy as np
import json

key = Fernet.generate_key()
cipher = Fernet(key)

def encrypt_embedding(vector: np.ndarray) -> bytes:
    raw = json.dumps(vector.tolist()).encode()
    return cipher.encrypt(raw)

def decrypt_embedding(ciphertext: bytes) -> np.ndarray:
    raw = cipher.decrypt(ciphertext)
    return np.array(json.loads(raw))

This approach is simple and fast. The tradeoff: you cannot run similarity queries on the encrypted store. Every search requires fetching all encrypted vectors, decrypting them locally, and computing cosine similarity yourself. For corpora under ~100k documents that fits in memory and completes in seconds; beyond that, you need a smarter approach.

Strategy 2: Approximate Locality-Sensitive Hashing

Locality-sensitive hashing (LSH) maps similar vectors to the same hash buckets with high probability. You can apply LSH before encryption to create an index structure that preserves approximate neighbourhood relationships without revealing exact embeddings.

import hashlib
import numpy as np

def lsh_bucket(vector: np.ndarray, planes: np.ndarray) -> str:
    """Project vector onto random hyperplanes, return binary signature."""
    projections = np.dot(planes, vector)
    bits = (projections > 0).astype(int)
    return ''.join(map(str, bits))

# Generate random projection planes once, store them securely
rng = np.random.default_rng(seed=42)
planes = rng.standard_normal((64, embedding_dim))  # 64-bit signature

bucket = lsh_bucket(my_embedding, planes)
# Store: { bucket_id: bucket, encrypted_vector: encrypt_embedding(my_embedding) }

At query time, compute the LSH bucket of the query vector, retrieve only the vectors in matching buckets, decrypt those, and compute exact similarity within that reduced set. You leak only bucket membership — a coarse topological fact — rather than exact positions. The random projection planes are your secret; keep them in a key management system, not the database.

Strategy 3: Structured Encryption for Vector Similarity

Structured encryption schemes provide formal privacy guarantees while supporting specific query patterns. For inner-product similarity (the basis of cosine similarity after L2 normalization), recent work on inner-product functional encryption (IPFE) allows a server to compute ⟨query, stored_vector⟩ without learning either vector.

A practical construction uses the MCFE (Multi-Client Functional Encryption) scheme:

  1. Setup: Generate a master secret key msk and public parameters pp.
  2. Encryption: For each embedding x, compute ct = Enc(msk, x).
  3. Key derivation: For a query vector y, derive a functional key sk_y = KeyGen(msk, y).
  4. Evaluation: The server computes Dec(sk_y, ct) = ⟨x, y⟩ without learning x or y.

Libraries like OpenFHE implement variants of this for production use. The overhead is roughly 10-100x compared to plaintext inner products — acceptable for high-value search over modest corpora, but challenging at billion-vector scale.

Strategy 4: Differential Privacy Noise Injection

If exact reconstruction is the threat rather than aggregate topology, you can add calibrated Laplace or Gaussian noise to embeddings before storage. This is the differential privacy (DP) approach applied to embedding spaces.

def dp_embed(vector: np.ndarray, epsilon: float, sensitivity: float = 1.0) -> np.ndarray:
    """Add Laplace noise calibrated to (epsilon, delta=0) DP."""
    scale = sensitivity / epsilon
    noise = np.random.laplace(0, scale, vector.shape)
    noisy = vector + noise
    # Re-normalize to unit sphere (cosine similarity is scale-invariant)
    return noisy / np.linalg.norm(noisy)

The privacy-utility tradeoff is direct: a smaller epsilon provides stronger privacy guarantees but degrades recall. Empirically, epsilon values between 1.0 and 3.0 preserve enough similarity structure for useful top-k retrieval while providing meaningful inversion resistance against black-box adversaries.

The key advantage over encryption: DP-noised embeddings can be stored and queried in any standard vector database (pgvector, Pinecone, Qdrant) with zero infrastructure changes. The key limitation: it is a probabilistic guarantee, not a cryptographic one. Sophisticated inversion models trained on your embedding family can partially overcome DP noise with enough queries.

Operational Patterns

Regardless of the strategy you choose, several operational practices apply:

Separate your encoder from your store. Never send raw text to a hosted embedding API and then store the result in a hosted vector database you do not control. Generate embeddings in your own infrastructure. If you use a hosted LLM for embedding generation, ensure the API provider's data retention policy aligns with your compliance requirements.

Rotate projection planes and functional keys regularly. LSH planes and IPFE master keys should be rotated on the same schedule as your other long-lived cryptographic material. Build your storage schema to support versioned keys from the start — retrofitting key rotation to an existing 50M-vector store is painful.

Audit similarity queries. A stream of queries that systematically probes your vector space is an inversion attack in progress. Log query vectors (or their hashes), set rate limits, and alert on abnormal query distributions. This is analogous to rate-limiting authentication endpoints.

Use embeddings only for their intended purpose. An embedding generated for semantic search should not be reused as a general-purpose document fingerprint. Purpose limitation reduces the information leakage surface: an attacker who steals your search index should not automatically gain a linkage oracle across your other systems.

Choosing the Right Approach

ApproachPrivacy strengthSearch performanceInfrastructure complexity
Client-side decryptionHighLow (full scan)Low
LSH + encryptionMediumGoodLow-Medium
Functional encryptionHighMediumHigh
Differential privacy noiseMedium (probabilistic)GoodLow

For most teams starting out, LSH + client-side encryption hits the best balance: it is straightforward to implement, works with any vector database, and provides meaningful protection against the most likely threats (a compromised database, a vendor breach, or an over-curious SaaS provider). Graduate to functional encryption if you require formal cryptographic guarantees and can absorb the performance overhead.

The goal is not to make semantic search impossible — it is to ensure that a breach of your vector store yields nothing actionable. Embeddings are not opaque by default; making them opaque requires deliberate design. The patterns above give you the tools to do that without abandoning the AI capabilities your application depends on.

Encrypt your agent's data today

BitAtlas gives your AI agents AES-256-GCM encrypted storage with zero-knowledge guarantees. Free tier, no credit card required.