Apache Kafka
Apache Kafka is an event streaming platform: producers append records to topics, and any number of consumers read them at their own pace, as often as they like, for as long as the topic keeps them. On Strongly it is the natural feed for streaming workflows.
Overview
- Versions: 4.3.1, 4.2.2, 4.1.2 (default: 4.3.1)
- Default Port: 9092
- Cluster Support: No (one broker)
- Use Cases: Event streams between services, feeding streaming workflows, change and activity logs, fan-out to many consumers
- Sign-in: SASL/PLAIN with the add-on's username and password; a client without them is refused
Key Features
- Durable, replayable topics: records stay for the topic's retention (7 days unless you set it), and a consumer can start from the beginning or from the latest record
- Consumer groups: the consumers of a group share a topic's partitions and resume where the group left off
- Ordering per partition: records with the same key land in the same partition, in order
- Topics on first use: a topic is created the first time a record is written to it (one partition); create it yourself first for more partitions or a set retention
Resources
Choose the add-on's resources on the create form:
| Setting | Options | Default |
|---|---|---|
| CPU (vCPU) | Any number of cores, e.g. 0.5, 1, 2 | 0.5 |
| Memory | Any amount in GB, at least 1 GB | 2 GB |
| Disk | Any amount in GB | 10 GB |
Half of the add-on's memory is Kafka's heap; the rest is the cache that serves recent records. Disk holds every topic's records for as long as their retention keeps them.
Creating a Kafka Add-on
- Navigate to Add-ons and click Create Add-on
- Select Apache Kafka as the type
- Choose a version (4.3.1, 4.2.2 or 4.1.2)
- Configure:
- Add-on Label (required): descriptive name (e.g., "order-events")
- Description (optional)
- Resources: CPU, memory and disk
- Optionally enable automatic backups (schedule and retention)
- Click Create Add-on
The add-on is RUNNING once its broker accepts the add-on's credentials and refuses a client without them.
Connection Information
The Connection tab shows the bootstrap address, port, username and password. Apps, workspaces and workflows get the same in STRONGLY_SERVICES, under services.addons.kafka:
{
"id": "addon-abc123defg",
"name": "order-events",
"type": "kafka",
"category": "add-on",
"status": "running",
"version": "4.3.1",
"connection": {
"connection_string": "<internal-host>:9092",
"host": "<internal-host>",
"port": 9092,
"brokers": "<internal-host>:9092",
"saslMechanism": "plain",
"ssl": false
},
"auth": {
"method": "username_password",
"credentials": { "username": "user_a1b2c3d4", "password": "<password>" }
}
}
brokers, saslMechanism and ssl have the same meaning as on a Kafka data source, so code (and workflow nodes) that read one read the other.
- Python
- Node.js
import json, os
from kafka import KafkaProducer, KafkaConsumer
services = json.loads(os.environ['STRONGLY_SERVICES'])
kafka = next(a for a in services['services']['addons']['kafka'] if a['name'] == 'order-events')
auth = dict(
bootstrap_servers=kafka['connection']['brokers'],
security_protocol='SASL_PLAINTEXT',
sasl_mechanism='PLAIN',
sasl_plain_username=kafka['auth']['credentials']['username'],
sasl_plain_password=kafka['auth']['credentials']['password'],
)
producer = KafkaProducer(**auth)
producer.send('orders', key=b'order-42', value=b'{"total": 12.5}')
producer.flush()
consumer = KafkaConsumer('orders', group_id='billing', auto_offset_reset='earliest', **auth)
for record in consumer:
print(record.key, record.value)
const { Kafka } = require('kafkajs');
const services = JSON.parse(process.env.STRONGLY_SERVICES);
const addon = services.services.addons.kafka.find(a => a.name === 'order-events');
const kafka = new Kafka({
clientId: 'my-app',
brokers: addon.connection.brokers.split(','),
sasl: {
mechanism: 'plain',
username: addon.auth.credentials.username,
password: addon.auth.credentials.password,
},
});
const producer = kafka.producer();
await producer.connect();
await producer.send({ topic: 'orders', messages: [{ key: 'order-42', value: '{"total": 12.5}' }] });
In Streaming Workflows
The Streaming Kafka Source subscribes to topics and turns each record into a frame; the Streaming Kafka Producer writes frames to a topic. Each node's Connection Type picks a Kafka add-on or a Kafka data source (a cluster you run elsewhere); the node reads either the same way.
Marketplace templates that need Kafka (for example Live Transcription Analytics and Redis Stream to LLM Enrichment) offer three choices on the install form: create a new Kafka add-on, use one you have, or use a Kafka data source.
Backups
A Kafka backup holds every topic (its partitions and the configs set on it), every record exactly as written (keys, values, headers and timestamps), and each consumer group's position. It reads each partition up to where it ended when the backup started; records written while it runs are not included.
-
Back up now: click Back Up Now on the Backup tab while the add-on is running.
-
Automatic: on the Backup tab turn on Enable Automatic Backups, choose a Backup Schedule and a Retention, and click Save Configuration.
-
History: the Backup tab lists every backup with its status, size and any error.
-
Restore: stop the consumers first; a restore refuses while a consumer group has members, before changing anything. Then click Restore next to a succeeded backup and confirm. Every topic and consumer group is deleted, the backup's topics are recreated with their records, and each consumer group resumes from the same position as when the backup was taken. Producers and consumers fail until the restore finishes. See Restoring a backup.
Monitoring
The Metrics tab measures the running add-on live: CPU, memory and disk use against its size, network traffic, open and new connections, response time (a metadata request) and health. The Logs tab shows the broker's log.
Best Practices
- Pick a key for records that must stay in order (an order id, a device id): records with the same key go to the same partition
- Create busy topics yourself with the partitions your consumers need; partitions are what a consumer group spreads over
- Set retention per topic (
retention.ms,retention.bytes) so the disk holds what you need and no more - Use one consumer group per purpose: each group reads the whole topic independently
- Watch disk on the Metrics tab and resize before it fills