Milvus
Milvus is an open-source vector database built for AI applications, enabling efficient similarity search and analytics on embedding vectors.
Overview
- Versions: 2.6.3, 2.5.19, 2.4.11 (default: 2.6.3)
- Default Port: 19530 (gRPC)
- Cluster Support: Yes, version 2.6 or later
- Use Cases: Vector databases, AI embeddings, semantic search, RAG systems
- Features: Similarity search, multiple index types, hybrid search with scalar filtering
Key Features
- Billion-Scale Vector Search: Handle massive vector datasets efficiently
- Multiple Index Types: IVF, HNSW, FLAT, ANNOY, and more for different use cases
- Hybrid Search: Combine vector similarity with attribute filtering
- Dynamic Schema: Flexible schema with collections and partitions
- Multiple Distance Metrics: L2, IP (Inner Product), Cosine similarity
- Data Consistency: Tunable consistency levels
- Scalar Filtering: Filter vectors by metadata attributes
Deployment Modes
Single Node
- Milvus standalone with its metadata store and object storage alongside, for development and moderate data sizes
Cluster (High Availability)
- Distributed Milvus: search runs on query nodes (one per Data Node, 3-10) behind a proxy, with redundant coordination and storage
- Replication Factor: how many in-memory copies of each loaded collection are kept across query nodes (1-10)
- Requires version 2.6 or later
- Connect exactly as to a single node: the same host and port
Resources
Choose the add-on's resources on the create form:
| Setting | Options | Default |
|---|---|---|
| CPU (vCPU) | Any number of cores, e.g. 0.5, 1, 2 | 0.5 |
| Memory | Any amount in GB, at least the type's minimum | 1 GB |
| Disk | Any amount in GB | 10 GB |
| GPU Count | 0-8 (0 for CPU-only) | 0 |
Milvus keeps vector indexes in memory, so give it more memory than a traditional database of the same data size.
Creating a Milvus Add-on
- Navigate to Add-ons and click Create Add-on
- On the Create New Add-on page, select Milvus as the type
- Choose a version (2.6.3, 2.5.19, or 2.4.11)
- Select deployment mode:
- Single Node
- Cluster (High Availability): 3-10 data nodes and a replication factor (version 2.6 or later)
- Configure:
- Add-on Label (required): descriptive name (e.g., "vector-store")
- Description (optional): purpose and notes
- Resources: CPU, memory, and disk for your workload
- Optionally enable automatic backups:
- Schedule: Hourly, Daily, Weekly, or Monthly
- Retention: number of backups to keep (1-30, default 7)
- Click Create Add-on
Connection Information
Once the add-on is running, the Connection tab of the add-on details page shows the internal host (for apps), port, username, and password. The same details are exposed to your apps via STRONGLY_SERVICES.
Connection Format
Milvus uses gRPC for client connections. The connection_string for Milvus is a plain host:port pair (no URI scheme):
<internal-host>:19530
A username and password are generated and shown with the add-on. Milvus add-ons require sign-in: every connection must pass this username and password, as the examples below do. A connection without them, or with a wrong password, is refused. The add-on is reachable only from your organization's apps, workspaces and workflows, never from outside the platform.
The Milvus source, Milvus destination and Semantic Memory workflow nodes sign in with the add-on's credentials automatically. An add-on created before sign-in was required starts requiring it the next time it is started from a new deployment (for example after Recover); until then it still accepts connections without credentials, and connections that pass them keep working either way.
These nodes are safe to run side by side: each call opens its own connection, so a store or search step can run inside a Loop or Map on several parallel workers and across pods. They use PyMilvus 2.6, the client release for the Milvus 2.4, 2.5 and 2.6 versions the add-on offers.
Accessing Connection Details
In STRONGLY_SERVICES, add-ons are grouped by type under services.addons, and each entry is one provisioned instance:
{
"id": "addon-abc123defg",
"name": "vector-store",
"type": "milvus",
"category": "add-on",
"status": "running",
"version": "2.6.3",
"connection": {
"connection_string": "<internal-host>:19530",
"uri": "<internal-host>:19530",
"host": "<internal-host>",
"port": 19530
},
"auth": {
"method": "username_password",
"credentials": { "username": "user_a1b2c3d4", "password": "<password>" }
},
"limits": { "max_connections": 100, "storage_gb": 10 },
"metadata": { "cpu": "0.5", "memory": "1GB", "disk": "10GB", "backup_enabled": false }
}
- Python
- Node.js
- Go
import os
import json
from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType
# Parse STRONGLY_SERVICES
services = json.loads(os.environ['STRONGLY_SERVICES'])
# Pick your Milvus add-on by name (the label you gave it)
milvus_addon = next(
a for a in services['services']['addons']['milvus']
if a['name'] == 'vector-store'
)
# Connect using host and port
connections.connect(
alias='default',
host=milvus_addon['connection']['host'],
port=milvus_addon['connection']['port'],
user=milvus_addon['auth']['credentials']['username'],
password=milvus_addon['auth']['credentials']['password']
)
# Verify connection
print("Connected to Milvus")
# Disconnect when done
connections.disconnect('default')
const { MilvusClient } = require('@zilliz/milvus2-sdk-node');
// Parse STRONGLY_SERVICES
const services = JSON.parse(process.env.STRONGLY_SERVICES);
const milvusAddon = services.services.addons.milvus
.find(a => a.name === 'vector-store');
// Connect using host and port
const client = new MilvusClient({
address: `${milvusAddon.connection.host}:${milvusAddon.connection.port}`,
username: milvusAddon.auth.credentials.username,
password: milvusAddon.auth.credentials.password
});
// Check connection
const health = await client.checkHealth();
console.log('Milvus health:', health);
package main
import (
"context"
"encoding/json"
"fmt"
"os"
"github.com/milvus-io/milvus-sdk-go/v2/client"
)
type Connection struct {
Host string `json:"host"`
Port int `json:"port"`
}
type Credentials struct {
Username string `json:"username"`
Password string `json:"password"`
}
type Addon struct {
Name string `json:"name"`
Connection Connection `json:"connection"`
Auth struct {
Credentials Credentials `json:"credentials"`
} `json:"auth"`
}
type Services struct {
Services struct {
Addons map[string][]Addon `json:"addons"`
} `json:"services"`
}
func main() {
var services Services
json.Unmarshal([]byte(os.Getenv("STRONGLY_SERVICES")), &services)
milvusAddon := services.Services.Addons["milvus"][0]
ctx := context.Background()
// Connect using host, port and the add-on's credentials
c, err := client.NewDefaultGrpcClientWithAuth(
ctx,
fmt.Sprintf("%s:%d", milvusAddon.Connection.Host, milvusAddon.Connection.Port),
milvusAddon.Auth.Credentials.Username,
milvusAddon.Auth.Credentials.Password,
)
if err != nil {
panic(err)
}
defer c.Close()
fmt.Println("Connected to Milvus")
}
Core Concepts
Collections
Collections are similar to tables in relational databases, storing vectors and metadata.
from pymilvus import Collection, FieldSchema, CollectionSchema, DataType
# Define schema
fields = [
FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=True),
FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=128),
FieldSchema(name="text", dtype=DataType.VARCHAR, max_length=1000),
FieldSchema(name="category", dtype=DataType.VARCHAR, max_length=100)
]
schema = CollectionSchema(fields, description="Document embeddings")
# Create collection
collection = Collection(name="documents", schema=schema)
print(f"Collection created: {collection.name}")
Vectors and Fields
Vectors represent your embeddings, and fields store metadata.
# Insert data
data = [
# embeddings (128-dimensional vectors)
[[0.1] * 128, [0.2] * 128, [0.3] * 128],
# text metadata
["Document 1", "Document 2", "Document 3"],
# category metadata
["tech", "science", "tech"]
]
collection.insert(data)
collection.flush()
print(f"Inserted {collection.num_entities} entities")
Indexes
Indexes accelerate vector similarity search.
# Create IVF_FLAT index
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "L2",
"params": {"nlist": 128}
}
collection.create_index(
field_name="embedding",
index_params=index_params
)
print("Index created")
Common Operations
Creating a Collection
from pymilvus import connections, Collection, FieldSchema, CollectionSchema, DataType
connections.connect(host='host', port=19530, user='username', password='password')
# Define schema with multiple field types
fields = [
FieldSchema(name="id", dtype=DataType.INT64, is_primary=True, auto_id=False),
FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=768),
FieldSchema(name="title", dtype=DataType.VARCHAR, max_length=500),
FieldSchema(name="content", dtype=DataType.VARCHAR, max_length=5000),
FieldSchema(name="timestamp", dtype=DataType.INT64),
FieldSchema(name="score", dtype=DataType.FLOAT)
]
schema = CollectionSchema(
fields,
description="Article embeddings",
enable_dynamic_field=True # Allow dynamic fields
)
collection = Collection(name="articles", schema=schema)
Inserting Vectors
# Prepare data
import numpy as np
num_entities = 1000
embeddings = np.random.random((num_entities, 768)).tolist()
ids = list(range(num_entities))
titles = [f"Article {i}" for i in range(num_entities)]
contents = [f"Content of article {i}" for i in range(num_entities)]
timestamps = [1234567890 + i for i in range(num_entities)]
scores = np.random.random(num_entities).tolist()
# Insert
data = [ids, embeddings, titles, contents, timestamps, scores]
collection.insert(data)
collection.flush()
print(f"Total entities: {collection.num_entities}")
Creating Indexes
Different index types for different use cases:
# IVF_FLAT: Good balance of speed and accuracy
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "L2",
"params": {"nlist": 1024}
}
# HNSW: High accuracy, memory intensive
index_params = {
"index_type": "HNSW",
"metric_type": "L2",
"params": {
"M": 16, # Max connections per layer
"efConstruction": 200 # Search quality during build
}
}
# IVF_PQ: Memory efficient, lower accuracy
index_params = {
"index_type": "IVF_PQ",
"metric_type": "L2",
"params": {
"nlist": 1024,
"m": 8, # PQ compression factor
"nbits": 8
}
}
collection.create_index(
field_name="embedding",
index_params=index_params
)
Searching Vectors
# Load collection to memory
collection.load()
# Prepare search vectors
search_vectors = [[0.1] * 768]
# Basic search
search_params = {
"metric_type": "L2",
"params": {"nprobe": 10} # Number of clusters to search
}
results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10, # Top 10 results
output_fields=["title", "content", "score"]
)
for hits in results:
for hit in hits:
print(f"ID: {hit.id}, Distance: {hit.distance}, Title: {hit.entity.get('title')}")
Hybrid Search (Vector + Scalar Filtering)
# Search with metadata filtering
expr = "score > 0.5 and timestamp > 1234567890"
results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10,
expr=expr, # Filter expression
output_fields=["title", "score", "timestamp"]
)
Querying by ID
# Query specific entities by ID
ids = [1, 5, 10, 15]
results = collection.query(
expr=f"id in {ids}",
output_fields=["id", "title", "content"]
)
for result in results:
print(result)
Deleting Vectors
# Delete by expression
expr = "id in [1, 2, 3]"
collection.delete(expr)
# Delete by range
expr = "timestamp < 1234567890"
collection.delete(expr)
Distance Metrics
Milvus supports multiple distance metrics:
- L2 (Euclidean): sqrt(sum((x-y)^2)) - Smaller is more similar
- IP (Inner Product): sum(x*y) - Larger is more similar
- COSINE: 1 - (xy)/(||x||||y||) - Smaller is more similar
# L2 distance (Euclidean)
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "L2",
"params": {"nlist": 128}
}
# Inner Product
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "IP",
"params": {"nlist": 128}
}
# Cosine similarity
index_params = {
"index_type": "IVF_FLAT",
"metric_type": "COSINE",
"params": {"nlist": 128}
}
Use Cases
Semantic Search
from sentence_transformers import SentenceTransformer
# Initialize embedding model
model = SentenceTransformer('all-MiniLM-L6-v2')
# Encode documents
documents = [
"Machine learning is a subset of artificial intelligence",
"Deep learning uses neural networks with multiple layers",
"Python is a popular programming language"
]
embeddings = model.encode(documents)
# Insert into Milvus
data = [
list(range(len(documents))), # IDs
embeddings.tolist(), # Vectors
documents # Text
]
collection.insert(data)
collection.flush()
# Search with query
query = "What is AI and machine learning?"
query_vector = model.encode([query])
collection.load()
results = collection.search(
data=query_vector.tolist(),
anns_field="embedding",
param={"metric_type": "COSINE", "params": {"nprobe": 10}},
limit=3,
output_fields=["text"]
)
for hits in results:
for hit in hits:
print(f"Distance: {hit.distance:.4f}, Text: {hit.entity.get('text')}")
Image Similarity Search
from PIL import Image
import torch
from torchvision import transforms, models
# Load pre-trained model
model = models.resnet50(pretrained=True)
model.eval()
# Remove classification layer to get embeddings
model = torch.nn.Sequential(*(list(model.children())[:-1]))
# Image preprocessing
transform = transforms.Compose([
transforms.Resize(256),
transforms.CenterCrop(224),
transforms.ToTensor(),
transforms.Normalize(mean=[0.485, 0.456, 0.406], std=[0.229, 0.224, 0.225])
])
def get_image_embedding(image_path):
image = Image.open(image_path).convert('RGB')
image_tensor = transform(image).unsqueeze(0)
with torch.no_grad():
embedding = model(image_tensor)
return embedding.squeeze().numpy()
# Index images
image_paths = ["img1.jpg", "img2.jpg", "img3.jpg"]
embeddings = [get_image_embedding(path) for path in image_paths]
data = [
list(range(len(image_paths))),
embeddings,
image_paths
]
collection.insert(data)
# Search for similar images
query_embedding = get_image_embedding("query.jpg")
results = collection.search(
data=[query_embedding.tolist()],
anns_field="embedding",
param={"metric_type": "L2", "params": {"nprobe": 10}},
limit=5,
output_fields=["image_path"]
)
RAG (Retrieval-Augmented Generation)
from openai import OpenAI
from sentence_transformers import SentenceTransformer
# Initialize models
embedder = SentenceTransformer('all-MiniLM-L6-v2')
llm = OpenAI()
# Index knowledge base
knowledge_base = [
"The Eiffel Tower is located in Paris, France.",
"It was built in 1889 for the World's Fair.",
"The tower is 330 meters tall.",
]
embeddings = embedder.encode(knowledge_base)
data = [list(range(len(knowledge_base))), embeddings.tolist(), knowledge_base]
collection.insert(data)
collection.flush()
collection.load()
# RAG query
def rag_query(question):
# Retrieve relevant context
query_vector = embedder.encode([question])
results = collection.search(
data=query_vector.tolist(),
anns_field="embedding",
param={"metric_type": "COSINE", "params": {"nprobe": 10}},
limit=3,
output_fields=["text"]
)
# Extract context
context = "\n".join([hit.entity.get('text') for hit in results[0]])
# Generate answer with LLM
response = llm.chat.completions.create(
model="gpt-6-luna",
messages=[
{"role": "system", "content": "Answer based on the following context:\n" + context},
{"role": "user", "content": question}
]
)
return response.choices[0].message.content
# Usage
answer = rag_query("How tall is the Eiffel Tower?")
print(answer)
Partitions
Partitions divide a collection for better organization and query performance.
# Create partitions
collection.create_partition("2024")
collection.create_partition("2023")
# Insert into specific partition
collection.insert(data, partition_name="2024")
# Search in specific partition
results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10,
partition_names=["2024"]
)
# List partitions
partitions = collection.partitions
for partition in partitions:
print(f"Partition: {partition.name}, Entities: {partition.num_entities}")
Backups
A Milvus backup holds every collection with its data and index settings, taken with the official Milvus backup tool while Milvus keeps serving, and a record of which collections were loaded. Databases with no collections are not included.
-
Back up now: click Backup Now on the status card, or Back Up Now on the Backup tab, while the add-on is running.
-
Automatic: on the Backup tab turn on Enable Automatic Backups, choose a Backup Schedule (Hourly, Daily, Weekly or Monthly) and a Retention (3, 7, 14 or 30 backups), and click Save Configuration. Older backups beyond the retention count are deleted automatically.
-
History: the Backup tab lists every backup with its status, size and any error.
-
Restore: click Restore next to a succeeded backup in Backup History and confirm. The backup is loaded back into this add-on while it keeps running: every collection is dropped, the collections in the backup are restored with their indexes, and the ones that were loaded are loaded again. Searches fail until then. Data written after the backup is lost. See Restoring a backup.
Performance Optimization
Index Selection
Choose index based on your use case:
| Index Type | Speed | Memory | Accuracy | Use Case |
|---|---|---|---|---|
| FLAT | Slow | High | 100% | Small datasets, highest accuracy needed |
| IVF_FLAT | Fast | Medium | ~99% | General purpose |
| IVF_PQ | Fastest | Low | ~95% | Large datasets, memory constrained |
| HNSW | Fast | High | ~99% | High QPS, low latency |
Search Parameters
Tune search parameters for performance:
# Lower nprobe = faster but less accurate
search_params = {"metric_type": "L2", "params": {"nprobe": 10}}
# Higher nprobe = slower but more accurate
search_params = {"metric_type": "L2", "params": {"nprobe": 64}}
# HNSW search parameters
search_params = {"metric_type": "L2", "params": {"ef": 64}} # Higher ef = better accuracy
Batch Operations
Batch inserts and searches for better throughput:
# Batch insert
batch_size = 1000
for i in range(0, len(all_data), batch_size):
batch = all_data[i:i + batch_size]
collection.insert(batch)
# Batch search
search_vectors = [[0.1] * 768 for _ in range(100)]
results = collection.search(
data=search_vectors,
anns_field="embedding",
param=search_params,
limit=10
)
Monitoring
The Metrics tab on the add-on details page measures the running add-on live: CPU, memory and disk use against its size, network traffic, open and new connections, response time, and instance health and uptime. See Metrics. The Logs tab shows its recent log output.
Collection Statistics
# Collection info
print(f"Entities: {collection.num_entities}")
# Index info
indexes = collection.indexes
for index in indexes:
print(f"Index: {index.field_name}, Type: {index.params}")
# Partition info
for partition in collection.partitions:
print(f"Partition: {partition.name}, Entities: {partition.num_entities}")
Best Practices
- Choose Right Index: Select index type based on dataset size and accuracy needs
- Batch Operations: Insert and search in batches for better performance
- Use Partitions: Organize data by time or category for faster queries
- Normalize Vectors: Normalize embeddings when using IP or COSINE metrics
- Monitor Memory: Large indexes require significant memory
- Tune Search Params: Balance nprobe/ef for speed vs accuracy
- Regular Backups: Enable daily backups for production
- Load Collections: Load collections to memory before searching
- Release Collections: Release unused collections to free memory
- Use Hybrid Search: Combine vector search with scalar filtering for better results
Troubleshooting
Connection Issues
from pymilvus import connections
try:
connections.connect(host='host', port=19530, user='username', password='password')
print("Connected successfully")
except Exception as e:
print(f"Connection failed: {e}")
Collection Not Found
from pymilvus import utility
# List all collections
collections = utility.list_collections()
print(f"Available collections: {collections}")
# Check if collection exists
exists = utility.has_collection("collection_name")
print(f"Collection exists: {exists}")
Out of Memory
# Release collection from memory
collection.release()
# Drop unused index
collection.drop_index()
# Use more memory-efficient index (IVF_PQ)
# Reduce nlist parameter
# Use partitions to limit search scope
Support
For issues or questions:
- Check add-on logs in the Logs tab of the add-on details page
- Review Milvus official documentation
- Contact Strongly support through the platform