Skip to main content

Milvus

Milvus

Milvus is an open-source vector database built for AI applications, enabling efficient similarity search and analytics on embedding vectors.

Overview​

  • Versions: 2.6.3, 2.5.19, 2.4.11 (default: 2.6.3)
  • Default Port: 19530 (gRPC)
  • Cluster Support: Yes, version 2.6 or later
  • Use Cases: Vector databases, AI embeddings, semantic search, RAG systems
  • Features: Similarity search, multiple index types, hybrid search with scalar filtering

Key Features​

  • Billion-Scale Vector Search: Handle massive vector datasets efficiently
  • Multiple Index Types: IVF, HNSW, FLAT, ANNOY, and more for different use cases
  • Hybrid Search: Combine vector similarity with attribute filtering
  • Dynamic Schema: Flexible schema with collections and partitions
  • Multiple Distance Metrics: L2, IP (Inner Product), Cosine similarity
  • Data Consistency: Tunable consistency levels
  • Scalar Filtering: Filter vectors by metadata attributes

Deployment Modes​

Single Node​

  • Milvus standalone with its metadata store and object storage alongside, for development and moderate data sizes

Cluster (High Availability)​

  • Distributed Milvus: search runs on query nodes (one per Data Node, 3-10) behind a proxy, with redundant coordination and storage
  • Replication Factor: how many in-memory copies of each loaded collection are kept across query nodes (1-10)
  • Requires version 2.6 or later
  • Connect exactly as to a single node: the same host and port

Resources​

Choose the add-on's resources on the create form:

SettingOptionsDefault
CPU (vCPU)Any number of cores, e.g. 0.5, 1, 20.5
MemoryAny amount in GB, at least the type's minimum1 GB
DiskAny amount in GB10 GB
GPU Count0-8 (0 for CPU-only)0
Sizing

Milvus keeps vector indexes in memory, so give it more memory than a traditional database of the same data size.

Creating a Milvus Add-on​

  1. Navigate to Add-ons and click Create Add-on
  2. On the Create New Add-on page, select Milvus as the type
  3. Choose a version (2.6.3, 2.5.19, or 2.4.11)
  4. Select deployment mode:
    • Single Node
    • Cluster (High Availability): 3-10 data nodes and a replication factor (version 2.6 or later)
  5. Configure:
    • Add-on Label (required): descriptive name (e.g., "vector-store")
    • Description (optional): purpose and notes
    • Resources: CPU, memory, and disk for your workload
  6. Optionally enable automatic backups:
    • Schedule: Hourly, Daily, Weekly, or Monthly
    • Retention: number of backups to keep (1-30, default 7)
  7. Click Create Add-on

Connection Information​

Once the add-on is running, the Connection tab of the add-on details page shows the internal host (for apps), port, username, and password. The same details are exposed to your apps via STRONGLY_SERVICES.

Connection Format​

Milvus uses gRPC for client connections. The connection_string for Milvus is a plain host:port pair (no URI scheme):

<internal-host>:19530

A username and password are generated and shown with the add-on. Milvus add-ons require sign-in: every connection must pass this username and password, as the examples below do. A connection without them, or with a wrong password, is refused. The add-on is reachable only from your organization's apps, workspaces and workflows, never from outside the platform.

The Milvus source, Milvus destination and Semantic Memory workflow nodes sign in with the add-on's credentials automatically. An add-on created before sign-in was required starts requiring it the next time it is started from a new deployment (for example after Recover); until then it still accepts connections without credentials, and connections that pass them keep working either way.

These nodes are safe to run side by side: each call opens its own connection, so a store or search step can run inside a Loop or Map on several parallel workers and across pods. They use PyMilvus 2.6, the client release for the Milvus 2.4, 2.5 and 2.6 versions the add-on offers.

Accessing Connection Details​

In STRONGLY_SERVICES, add-ons are grouped by type under services.addons, and each entry is one provisioned instance:

{
"id": "addon-abc123defg",
"name": "vector-store",
"type": "milvus",
"category": "add-on",
"status": "running",
"version": "2.6.3",
"connection": {
"connection_string": "<internal-host>:19530",
"uri": "<internal-host>:19530",
"host": "<internal-host>",
"port": 19530
},
"auth": {
"method": "username_password",
"credentials": { "username": "user_a1b2c3d4", "password": "<password>" }
},
"limits": { "max_connections": 100, "storage_gb": 10 },
"metadata": { "cpu": "0.5", "memory": "1GB", "disk": "10GB", "backup_enabled": false }
}
import os
import json
from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType

# Parse STRONGLY_SERVICES
services = json.loads(os.environ['STRONGLY_SERVICES'])

# Pick your Milvus add-on by name (the label you gave it)
milvus_addon = next(
a for a in services['services']['addons']['milvus']
if a['name'] == 'vector-store'
)

# Connect using host and port
connections.connect(
alias='default',
host=milvus_addon['connection']['host'],
port=milvus_addon['connection']['port'],
user=milvus_addon['auth']['credentials']['username'],
password=milvus_addon['auth']['credentials']['password']
)

# Verify connection
print("Connected to Milvus")

# Disconnect when done
connections.disconnect('default')

Core Concepts​

Collections​

Collections are similar to tables in relational databases, storing vectors and metadata.

from pymilvus import Collection, FieldSchema, CollectionSchema, DataType

# Define schema
fields = [
FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True),
FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=128),
FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=1000),
FieldSchema(name="category", dtype=DataType.VARCHAR, max_length=100)
]

schema = CollectionSchema(fields, description="Document embeddings")

# Create collection
collection = Collection(name="documents", schema=schema)

print(f"Collection created: {collection.name}")

Vectors and Fields​

Vectors represent your embeddings, and fields store metadata.

# Insert data
data = [
# embeddings (128-dimensional vectors)
[[0.1] * 128, [0.2] * 128, [0.3] * 128],
# text metadata
["Document 1", "Document 2", "Document 3"],
# category metadata
["tech", "science", "tech"]
]

collection.insert(data)
collection.flush()

print(f"Inserted {collection.num_entities} entities")

Indexes​

Indexes accelerate vector similarity search.

# Create IVF_FLAT index
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "L2",
"params": {"nlist": 128}
}

collection.create_index(
field_name="embedding",
index_params=index_params
)

print("Index created")

Common Operations​

Creating a Collection​

from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType

connections.connect(host='host', port=19530, user='username', password='password')

# Define schema with multiple field types
fields = [
FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=False),
FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=768),
FieldSchema(name="title", dtype=DataType.VARCHAR, max_length=500),
FieldSchema(name="content", dtype=DataType.VARCHAR, max_length=5000),
FieldSchema(name="timestamp", dtype=DataType.INT64),
FieldSchema(name="score", dtype=DataType.FLOAT)
]

schema = CollectionSchema(
fields,
description="Article embeddings",
enable_dynamic_field=True # Allow dynamic fields
)

collection = Collection(name="articles", schema=schema)

Inserting Vectors​

# Prepare data
import numpy as np

num_entities = 1000
embeddings = np.random.random((num_entities, 768)).tolist()
ids = list(range(num_entities))
titles = [f"Article {i}" for i in range(num_entities)]
contents = [f"Content of article {i}" for i in range(num_entities)]
timestamps = [1234567890 + i for i in range(num_entities)]
scores = np.random.random(num_entities).tolist()

# Insert
data = [ids, embeddings, titles, contents, timestamps, scores]
collection.insert(data)
collection.flush()

print(f"Total entities: {collection.num_entities}")

Creating Indexes​

Different index types for different use cases:

# IVF_FLAT: Good balance of speed and accuracy
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "L2",
"params": {"nlist": 1024}
}

# HNSW: High accuracy, memory intensive
index_params = {
"index_type": "HNSW",
"metric_type": "L2",
"params": {
"M": 16, # Max connections per layer
"efConstruction": 200 # Search quality during build
}
}

# IVF_PQ: Memory efficient, lower accuracy
index_params = {
"index_type": "IVF_PQ",
"metric_type": "L2",
"params": {
"nlist": 1024,
"m": 8, # PQ compression factor
"nbits": 8
}
}

collection.create_index(
field_name="embedding",
index_params=index_params
)

Searching Vectors​

# Load collection to memory
collection.load()

# Prepare search vectors
search_vectors = [[0.1] * 768]

# Basic search
search_params = {
"metric_type": "L2",
"params": {"nprobe": 10} # Number of clusters to search
}

results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10, # Top 10 results
output_fields=["title", "content", "score"]
)

for hits in results:
for hit in hits:
print(f"ID: {hit.id}, Distance: {hit.distance}, Title: {hit.entity.get('title')}")

Hybrid Search (Vector + Scalar Filtering)​

# Search with metadata filtering
expr = "score > 0.5 and timestamp > 1234567890"

results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10,
expr=expr, # Filter expression
output_fields=["title", "score", "timestamp"]
)

Querying by ID​

# Query specific entities by ID
ids = [1, 5, 10, 15]
results = collection.query(
expr=f"id in {ids}",
output_fields=["id", "title", "content"]
)

for result in results:
print(result)

Deleting Vectors​

# Delete by expression
expr = "id in [1, 2, 3]"
collection.delete(expr)

# Delete by range
expr = "timestamp < 1234567890"
collection.delete(expr)

Distance Metrics​

Milvus supports multiple distance metrics:

  • L2 (Euclidean): sqrt(sum((x-y)^2)) - Smaller is more similar
  • IP (Inner Product): sum(x*y) - Larger is more similar
  • COSINE: 1 - (xy)/(||x||||y||) - Smaller is more similar
# L2 distance (Euclidean)
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "L2",
"params": {"nlist": 128}
}

# Inner Product
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "IP",
"params": {"nlist": 128}
}

# Cosine similarity
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "COSINE",
"params": {"nlist": 128}
}

Use Cases​

from sentence_transformers import SentenceTransformer

# Initialize embedding model
model = SentenceTransformer('all-MiniLM-L6-v2')

# Encode documents
documents = [
"Machine learning is a subset of artificial intelligence",
"Deep learning uses neural networks with multiple layers",
"Python is a popular programming language"
]

embeddings = model.encode(documents)

# Insert into Milvus
data = [
list(range(len(documents))), # IDs
embeddings.tolist(), # Vectors
documents # Text
]
collection.insert(data)
collection.flush()

# Search with query
query = "What is AI and machine learning?"
query_vector = model.encode([query])

collection.load()
results = collection.search(
data=query_vector.tolist(),
anns_field="embedding",
param={"metric_type": "COSINE", "params": {"nprobe": 10}},
limit=3,
output_fields=["text"]
)

for hits in results:
for hit in hits:
print(f"Distance: {hit.distance:.4f}, Text: {hit.entity.get('text')}")
from PIL import Image
import torch
from torchvision import transforms, models

# Load pre-trained model
model = models.resnet50(pretrained=True)
model.eval()

# Remove classification layer to get embeddings
model = torch.nn.Sequential(*(list(model.children())[:-1]))

# Image preprocessing
transform = transforms.Compose([
transforms.Resize(256),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
])

def get_image_embedding(image_path):
image = Image.open(image_path).convert('RGB')
image_tensor = transform(image).unsqueeze(0)

with torch.no_grad():
embedding = model(image_tensor)

return embedding.squeeze().numpy()

# Index images
image_paths = ["img1.jpg", "img2.jpg", "img3.jpg"]
embeddings = [get_image_embedding(path) for path in image_paths]

data = [
list(range(len(image_paths))),
embeddings,
image_paths
]
collection.insert(data)

# Search for similar images
query_embedding = get_image_embedding("query.jpg")
results = collection.search(
data=[query_embedding.tolist()],
anns_field="embedding",
param={"metric_type": "L2", "params": {"nprobe": 10}},
limit=5,
output_fields=["image_path"]
)

RAG (Retrieval-Augmented Generation)​

from openai import OpenAI
from sentence_transformers import SentenceTransformer

# Initialize models
embedder = SentenceTransformer('all-MiniLM-L6-v2')
llm = OpenAI()

# Index knowledge base
knowledge_base = [
"The Eiffel Tower is located in Paris, France.",
"It was built in 1889 for the World's Fair.",
"The tower is 330 meters tall.",
]

embeddings = embedder.encode(knowledge_base)
data = [list(range(len(knowledge_base))), embeddings.tolist(), knowledge_base]
collection.insert(data)
collection.flush()
collection.load()

# RAG query
def rag_query(question):
# Retrieve relevant context
query_vector = embedder.encode([question])
results = collection.search(
data=query_vector.tolist(),
anns_field="embedding",
param={"metric_type": "COSINE", "params": {"nprobe": 10}},
limit=3,
output_fields=["text"]
)

# Extract context
context = "\n".join([hit.entity.get('text') for hit in results[0]])

# Generate answer with LLM
response = llm.chat.completions.create(
model="gpt-6-luna",
messages=[
{"role": "system", "content": "Answer based on the following context:\n" + context},
{"role": "user", "content": question}
]
)

return response.choices[0].message.content

# Usage
answer = rag_query("How tall is the Eiffel Tower?")
print(answer)

Partitions​

Partitions divide a collection for better organization and query performance.

# Create partitions
collection.create_partition("2024")
collection.create_partition("2023")

# Insert into specific partition
collection.insert(data, partition_name="2024")

# Search in specific partition
results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10,
partition_names=["2024"]
)

# List partitions
partitions = collection.partitions
for partition in partitions:
print(f"Partition: {partition.name}, Entities: {partition.num_entities}")

Backups​

A Milvus backup holds every collection with its data and index settings, taken with the official Milvus backup tool while Milvus keeps serving, and a record of which collections were loaded. Databases with no collections are not included.

  • Back up now: click Backup Now on the status card, or Back Up Now on the Backup tab, while the add-on is running.

  • Automatic: on the Backup tab turn on Enable Automatic Backups, choose a Backup Schedule (Hourly, Daily, Weekly or Monthly) and a Retention (3, 7, 14 or 30 backups), and click Save Configuration. Older backups beyond the retention count are deleted automatically.

  • History: the Backup tab lists every backup with its status, size and any error.

  • Restore: click Restore next to a succeeded backup in Backup History and confirm. The backup is loaded back into this add-on while it keeps running: every collection is dropped, the collections in the backup are restored with their indexes, and the ones that were loaded are loaded again. Searches fail until then. Data written after the backup is lost. See Restoring a backup.

Performance Optimization​

Index Selection​

Choose index based on your use case:

Index TypeSpeedMemoryAccuracyUse Case
FLATSlowHigh100%Small datasets, highest accuracy needed
IVF_FLATFastMedium~99%General purpose
IVF_PQFastestLow~95%Large datasets, memory constrained
HNSWFastHigh~99%High QPS, low latency

Search Parameters​

Tune search parameters for performance:

# Lower nprobe = faster but less accurate
search_params = {"metric_type": "L2", "params": {"nprobe": 10}}

# Higher nprobe = slower but more accurate
search_params = {"metric_type": "L2", "params": {"nprobe": 64}}

# HNSW search parameters
search_params = {"metric_type": "L2", "params": {"ef": 64}} # Higher ef = better accuracy

Batch Operations​

Batch inserts and searches for better throughput:

# Batch insert
batch_size = 1000
for i in range(0, len(all_data), batch_size):
batch = all_data[i:i + batch_size]
collection.insert(batch)

# Batch search
search_vectors = [[0.1] * 768 for _ in range(100)]
results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10
)

Monitoring​

The Metrics tab on the add-on details page measures the running add-on live: CPU, memory and disk use against its size, network traffic, open and new connections, response time, and instance health and uptime. See Metrics. The Logs tab shows its recent log output.

Collection Statistics​

# Collection info
print(f"Entities: {collection.num_entities}")

# Index info
indexes = collection.indexes
for index in indexes:
print(f"Index: {index.field_name}, Type: {index.params}")

# Partition info
for partition in collection.partitions:
print(f"Partition: {partition.name}, Entities: {partition.num_entities}")

Best Practices​

  1. Choose Right Index: Select index type based on dataset size and accuracy needs
  2. Batch Operations: Insert and search in batches for better performance
  3. Use Partitions: Organize data by time or category for faster queries
  4. Normalize Vectors: Normalize embeddings when using IP or COSINE metrics
  5. Monitor Memory: Large indexes require significant memory
  6. Tune Search Params: Balance nprobe/ef for speed vs accuracy
  7. Regular Backups: Enable daily backups for production
  8. Load Collections: Load collections to memory before searching
  9. Release Collections: Release unused collections to free memory
  10. Use Hybrid Search: Combine vector search with scalar filtering for better results

Troubleshooting​

Connection Issues​

from pymilvus import connections

try:
connections.connect(host='host', port=19530, user='username', password='password')
print("Connected successfully")
except Exception as e:
print(f"Connection failed: {e}")

Collection Not Found​

from pymilvus import utility

# List all collections
collections = utility.list_collections()
print(f"Available collections: {collections}")

# Check if collection exists
exists = utility.has_collection("collection_name")
print(f"Collection exists: {exists}")

Out of Memory​

# Release collection from memory
collection.release()

# Drop unused index
collection.drop_index()

# Use more memory-efficient index (IVF_PQ)
# Reduce nlist parameter
# Use partitions to limit search scope

Support​

For issues or questions:

  • Check add-on logs in the Logs tab of the add-on details page
  • Review Milvus official documentation
  • Contact Strongly support through the platform