Data Forge
Generate synthetic fine-tuning datasets from your documents. Data Forge lets you upload source documents, parse and chunk them, generate Q&A training pairs using a teacher LLM, review and curate the output, and export it for fine-tuning.
All endpoints require authentication via X-API-Key header.
DataForgeProject Object
{
"projectId": "proj_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"name": "Customer Support FAQ",
"description": "Generate training data from support documentation",
"status": "created",
"config": {
"chunkStrategy": "semantic",
"embeddingModelId": "165ab9b829cb6d776b350b32",
"chunkSize": 1024,
"chunkOverlap": 128,
"generationModel": null,
"generationPromptTemplate": null,
"pairsPerChunk": 3,
"outputFormat": "chatml",
"systemPrompt": "",
"temperature": 0.7,
"maxTokens": 2048
},
"stats": {
"totalDocuments": 12,
"totalChunks": 340,
"totalPairs": 1020,
"acceptedPairs": 850,
"rejectedPairs": 45,
"pendingPairs": 125,
"totalGenerations": 3,
"avgQualityScore": 0.87
},
"createdAt": "2025-06-15T10:00:00Z",
"updatedAt": "2025-06-15T14:30:00Z"
}
config.chunk_strategy is how documents are split into chunks: semantic (a chunk ends where the meaning shifts, judged by the embeddings of embeddingModelId), heading (one chunk per section), paragraph (whole paragraphs, the default) or sliding_window (fixed-size windows overlapping by chunkOverlap). Every strategy keeps a chunk within chunkSize characters. embeddingModelId is required with semantic and null otherwise; it must be an embedding model you may use, from GET /api/v1/ai-gateway/data-forge/embedding-models.
status is created for a new project and moves to parsing, parsed and generating as you work on it. stats.avg_quality_score is null while no pair has been scored. A project that has been exported also carries exports, one record per export.
Projects
POST /api/v1/ai-gateway/data-forge/projects
Create a new Data Forge project.
Request Body
{
"name": "Customer Support FAQ",
"description": "Generate training data from support documentation",
"chunkStrategy": "semantic",
"embeddingModelId": "165ab9b829cb6d776b350b32"
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Project name (1-200 characters) |
description | string | No | Project description (max 2000 characters) |
chunkStrategy | string | No | semantic, heading, paragraph (default) or sliding_window |
embeddingModelId | string | With semantic | The embedding model semantic chunking splits by |
Returns 400 validation-error if name is missing or not a string, the strategy is not one of these, semantic has no embeddingModelId, or the model is not an embedding model you may use. The new project starts with the default config shown in the DataForgeProject object; change it with PUT.
Response 201 Created
Returns the full DataForgeProject object.
GET /api/v1/ai-gateway/data-forge/projects
List all Data Forge projects for the authenticated user, scoped to the organization.
Response 200 OK
Returns a list of DataForgeProject objects.
GET /api/v1/ai-gateway/data-forge/projects/:id
Get details of a specific Data Forge project.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Response 200 OK
Returns the full DataForgeProject object.
PATCH /api/v1/ai-gateway/data-forge/projects/:id
Update a Data Forge project.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Request Body
{
"name": "Updated Project Name",
"description": "Updated description",
"config": {}
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | No | Updated project name (1-200 characters) |
description | string | No | Updated description (max 2000 characters) |
config | object | No | The config fields to change; the others keep their values. Chunking changes apply to the documents parsed afterwards. |
Response 200 OK
Returns the updated DataForgeProject object.
DELETE /api/v1/ai-gateway/data-forge/projects/:id
Delete a Data Forge project and clean up associated S3 objects, documents, chunks, and pairs.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Response 200 OK
{
"data": { "success": true },
"meta": { "requestId": "req_abc123" }
}
Documents
POST /api/v1/ai-gateway/data-forge/projects/:id/uploads
Get a presigned PUT URL for uploading a document directly to S3.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Request Body
{
"fileName": "product-manual.pdf",
"mimeType": "application/pdf"
}
| Field | Type | Required | Description |
|---|---|---|---|
fileName | string | Yes | Name of the file to upload (1-500 characters) |
mimeType | string | Yes | MIME type of the file (e.g. application/pdf, text/plain) |
Returns 400 validation-error if fileName or mimeType is missing.
Response 200 OK
{
"data": {
"uploadUrl": "https://s3.amazonaws.com/bucket/data-forge/...",
"s3Key": "data-forge/user_456/proj_abc123/sources/product-manual.pdf",
"documentId": "doc_xyz789",
"filename": "product-manual.pdf",
"expiresIn": 3600
},
"meta": { "requestId": "req_abc123" }
}
PUT the file bytes to uploadUrl (valid for expiresIn seconds), then register the document with s3Key.
POST /api/v1/ai-gateway/data-forge/projects/:id/documents
Register a document that has been uploaded to S3. Call this after the browser finishes uploading to the presigned URL.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Request Body
{
"name": "product-manual.pdf",
"mimeType": "application/pdf",
"fileSize": 2048576,
"s3Key": "data-forge/user_456/proj_abc123/sources/product-manual.pdf"
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Document file name (1-500 characters) |
mimeType | string | Yes | MIME type of the uploaded file |
fileSize | number | Yes | File size in bytes (at least 1) |
s3Key | string | Yes | The s3Key returned by upload-url |
Returns 400 validation-error if any field is missing or fileSize is not a number.
Response 201 Created
{
"data": {
"documentId": "doc_xyz789",
"projectId": "proj_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"filename": "product-manual.pdf",
"name": "product-manual.pdf",
"s3Key": "data-forge/user_456/proj_abc123/sources/product-manual.pdf",
"fileSize": 2048576,
"contentType": "application/pdf",
"mimeType": "application/pdf",
"status": "uploaded",
"parseStatus": "pending",
"parsingStatus": "pending",
"chunkCount": 0,
"metadata": {},
"parsedMetadata": {},
"createdAt": "2025-06-15T10:05:00Z",
"updatedAt": "2025-06-15T10:05:00Z"
},
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/data-forge/projects/:id/documents
List all documents in a Data Forge project.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Response 200 OK
Returns a list of document objects.
DELETE /api/v1/ai-gateway/data-forge/projects/:projectId/documents/:id
Delete a document from a project and remove from S3.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
projectId | string | Yes | Project ID |
id | string | Yes | Document ID |
Response 200 OK
{
"data": { "success": true },
"meta": { "requestId": "req_abc123" }
}
Chunks
GET /api/v1/ai-gateway/data-forge/projects/:id/chunks
Get paginated chunks for a Data Forge project. Chunks are created when documents are parsed.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
page | integer | No | Page number, 1-based (default: 1) |
limit | integer | No | Items per page, 1-500 (default: 50) |
Response 200 OK
skip is the number of chunks before this page and limit the page size. contentType is the MIME type of the source document and position is the chunk's place within it.
{
"data": {
"chunks": [
{
"chunkId": "chunk_001",
"projectId": "proj_abc123",
"documentId": "doc_xyz789",
"userId": "user_456",
"organizationId": "org_xyz",
"content": "To reset your password, navigate to Settings > Security...",
"heading": "Password Reset",
"position": 12,
"contentType": "application/pdf",
"pairsGenerated": 3,
"status": "ready",
"createdAt": "2025-06-15T10:10:00Z"
}
],
"total": 340,
"skip": 0,
"limit": 50
},
"meta": { "requestId": "req_abc123" }
}
PATCH /api/v1/ai-gateway/data-forge/projects/:projectId/chunks/:id
Update a specific chunk's content or metadata.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
projectId | string | Yes | Project ID |
id | string | Yes | Chunk ID |
Request Body
{
"content": "Updated chunk text content",
"metadata": {},
"excluded": false
}
| Field | Type | Required | Description |
|---|---|---|---|
content | string | No | Updated chunk text content |
metadata | object | No | Updated chunk metadata |
excluded | boolean | No | Whether to exclude this chunk from generation |
Response 200 OK
Returns the updated chunk object.
Pipeline
POST /api/v1/ai-gateway/data-forge/projects/:id/parse
Start a document parsing job. Parses the uploaded documents not yet parsed (or whose parse failed) into text chunks using a K8s Job, with the project's chunking settings. Semantic chunking embeds each document's sentences with the project's embedding model, as you. Returns 400 if the project chunks semantically without an embedding model, or with one you may no longer use.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Response 200 OK
{
"data": {
"jobName": "df-parse-gen_pars",
"mode": "parse",
"projectId": "proj_abc123",
"documentCount": 12,
"status": "launched"
},
"meta": { "requestId": "req_abc123" }
}
POST /api/v1/ai-gateway/data-forge/projects/:id/generations
Start a data generation job. Uses an AI model to generate training pairs from parsed chunks.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Request Body
{
"teacherModelId": "gpt-4o",
"outputFormat": "qa",
"pairsPerChunk": 3,
"difficultyDistribution": {"easy": 0.3, "medium": 0.5, "hard": 0.2},
"systemPrompt": "Generate high-quality Q&A pairs from the provided context.",
"temperature": 0.7,
"styleTemplate": "mixed",
"thresholds": {"minQualityScore": 0.7, "minGroundingScore": 0.7}
}
| Field | Type | Required | Description |
|---|---|---|---|
teacherModelId | string | Yes | ID of the teacher model that generates the pairs (from available models) |
outputFormat | string | Yes | Generation type: qa, instruction, conversation, summary, or classification |
pairsPerChunk | integer | No | Pairs to generate per chunk (default: the project's setting, 3 unless changed) |
difficultyDistribution | object | No | Difficulty mix as fractions, e.g. {"easy": 0.3, "medium": 0.5, "hard": 0.2} |
systemPrompt | string | No | Custom system prompt for the teacher model (max 10,000 characters; default: the project's setting) |
temperature | float | No | Sampling temperature (0.0-2.0, default: 0.7) |
styleTemplate | string | No | Style template applied to generated pairs |
thresholds | object | No | Auto-filter thresholds, e.g. {"minQualityScore": 0.7, "minGroundingScore": 0.7} |
Returns 400 validation-error if teacherModelId or outputFormat is missing.
Response 201 Created
{
"data": {
"jobName": "df-generate-gen_abc1",
"mode": "generate",
"projectId": "proj_abc123",
"generationId": "gen_abc123",
"chunkCount": 340,
"status": "launched"
},
"meta": { "requestId": "req_abc123" }
}
Poll GET /ai-gateway/data-forge/projects/-/generations/:id with generationId for progress.
POST /api/v1/ai-gateway/data-forge/projects/-/generations/:id/cancel
Cancel a running generation job.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Generation job ID |
Response 200 OK: the generation run, as GET /ai-gateway/data-forge/projects/-/generations/:id shows it, status cancelled. Pairs it already generated are kept.
Generations
GET /api/v1/ai-gateway/data-forge/projects/:id/generations
List all generation jobs for a project.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
limit | integer | No | Number of results to return (default: 50, max: 200) |
cursor | string | No | meta.nextCursor of the previous page; omit for the first page |
Response 200 OK
{
"data": [
{
"generationId": "gen_abc123",
"projectId": "proj_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"status": "completed",
"config": {
"modelId": "gpt-4o",
"generationType": "qa",
"temperature": 0.7,
"pairsPerChunk": 3
},
"progress": 100,
"jobName": "df-generate-gen_abc1",
"stats": {
"totalChunksProcessed": 340,
"totalPairsGenerated": 1020,
"failedChunks": 0
},
"results": {
"chunksProcessed": 340,
"pairsGenerated": 1020,
"pairsValid": 980,
"avgQualityScore": 0.87,
"avgGroundingScore": 0.91,
"duplicateCount": 12,
"tokensUsed": 450000
},
"error": null,
"startedAt": "2025-06-15T11:00:30Z",
"completedAt": "2025-06-15T11:45:00Z",
"createdAt": "2025-06-15T11:00:00Z",
"updatedAt": "2025-06-15T11:45:00Z"
}
],
"meta": {
"total": 3,
"limit": 50,
"nextCursor": null,
"requestId": "req_abc123"
}
}
GET /api/v1/ai-gateway/data-forge/projects/-/generations/:id
Get details of a specific generation job.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Generation job ID |
Response 200 OK
data is the generation record, with the same fields as an entry in the generations list.
GET /api/v1/ai-gateway/data-forge/projects/-/generations/:id/logs
Get logs for a specific generation job. Logs are streamed in real-time during active jobs. Returns the most recent 1,000 entries, oldest first.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Generation job ID |
Response 200 OK
{
"data": [
{
"timestamp": "2025-06-15T11:00:30Z",
"level": "info",
"stage": "generation",
"message": "Starting generation for 340 chunks"
},
{
"timestamp": "2025-06-15T11:01:00Z",
"level": "info",
"stage": "generation",
"message": "Processing chunk 1/340: Password Reset"
},
{
"timestamp": "2025-06-15T11:01:05Z",
"level": "info",
"stage": "generation",
"message": "Generated 3 pairs for chunk 1 (avg quality: 0.92)"
}
],
"meta": { "requestId": "req_abc123" }
}
Pairs
GET /api/v1/ai-gateway/data-forge/projects/:id/pairs
Get paginated training pairs for a project with optional filtering.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
page | integer | No | Page number, 1-based (default: 1) |
limit | integer | No | Items per page, 1-500 (default: 50) |
status | string | No | Filter by review status: pending, accepted, rejected, edited |
documentId | string | No | Filter by generation ID (only pairs produced by that generation) |
q | string | No | Search within pair input/output text |
Response 200 OK
skip is the number of pairs before this page and limit the page size.
{
"data": {
"pairs": [
{
"pairId": "pair_001",
"projectId": "proj_abc123",
"chunkId": "chunk_001",
"documentId": "doc_xyz789",
"generationId": "gen_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"type": "single_turn",
"systemPrompt": "You are a helpful customer support assistant.",
"messages": [
{"role": "system", "content": "You are a helpful customer support assistant."},
{"role": "user", "content": "How do I reset my password?"},
{"role": "assistant", "content": "To reset your password, navigate to Settings > Security..."}
],
"question": "How do I reset my password?",
"answer": "To reset your password, navigate to Settings > Security...",
"qualityScore": 0.92,
"groundingScore": 0.95,
"complexityLevel": "easy",
"questionType": "factual",
"contentHash": null,
"isDuplicate": false,
"duplicateOf": null,
"status": "accepted",
"reviewerNotes": null,
"editHistory": [],
"createdAt": "2025-06-15T11:01:05Z",
"updatedAt": "2025-06-15T11:20:00Z"
}
],
"total": 1020,
"skip": 0,
"limit": 50
},
"meta": { "requestId": "req_abc123" }
}
PATCH /api/v1/ai-gateway/data-forge/projects/-/pairs/:id
Edit a training pair during review: rewrite its question, answer or system prompt, set its review status, or add reviewer notes. Only the fields given change.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Pair ID |
Request Body
{
"question": "Updated question text",
"answer": "Updated answer text",
"status": "accepted"
}
| Field | Type | Required | Description |
|---|---|---|---|
question | string | No | The pair's question (its input) |
answer | string | No | The pair's answer (its output) |
systemPrompt | string | No | The system prompt stored with the pair |
status | string | No | Review status: pending, accepted, rejected, edited |
reviewerNotes | string | No | Free-form reviewer notes |
An edited question, answer or system prompt marks the pair manually_edited and rewrites its chat-format messages to match (with a system message only where the pair had one). The edit is recorded as yours.
Response 200 OK
data is the pair record after the update, with the fields shown in the pairs list.
POST /api/v1/ai-gateway/data-forge/projects/:id/pairs/bulk
Perform a bulk action on the listed training pairs.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Request Body
{
"action": "accept",
"pairIds": ["pair_001", "pair_002", "pair_003"]
}
| Field | Type | Required | Description |
|---|---|---|---|
action | string | Yes | Action to perform: accept, reject, delete, reset |
pairIds | string[] | Yes | IDs of the pairs to act on |
Returns 400 validation-error if action is missing or pairIds is not an array.
Response 200 OK
{
"data": {
"action": "accept",
"affectedCount": 3
},
"meta": { "requestId": "req_abc123" }
}
Export
POST /api/v1/ai-gateway/data-forge/projects/:id/exports
Export the dataset as a JSONL file in the specified format, written into the data half of a shared volume. Only accepted pairs are included.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Request Body
{
"format": "chatml",
"includeSystemPrompt": true,
"minQualityScore": 0.8,
"volumeId": "vol_abc123"
}
| Field | Type | Required | Description |
|---|---|---|---|
format | string | Yes | Export format: chatml or alpaca |
includeSystemPrompt | boolean | No | Include system prompt in ChatML output (default: true) |
minQualityScore | float | No | Minimum quality score filter for pairs (0.0-1.0) |
volumeId | string | One of volumeId / newVolumeName | An existing shared volume you can edit; the export lands in its data half |
newVolumeName | string | One of volumeId / newVolumeName | Create a new shared volume with this name and land the export in it |
Returns 400 validation-error if format is missing or neither volumeId nor newVolumeName is given, and 404 not-found if the volume does not exist or you cannot edit it.
Export Formats
ChatML -- OpenAI-compatible chat format:
{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}
Alpaca -- Instruction-following format:
{"instruction": "...", "input": "", "output": "..."}
Response 201 Created
{
"data": {
"version": 3,
"volumeId": "vol_abc123",
"path": "dataforge-export-v3.jsonl",
"format": "chatml",
"pairCount": 850,
"fileSizeBytes": 4521890
},
"meta": { "requestId": "req_abc123" }
}
path is the file's path in the volume's data half.
GET /api/v1/ai-gateway/data-forge/projects/:projectId/exports/:version
Get a download URL for a specific export version.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
projectId | string | Yes | Project ID |
version | integer | Yes | Export version number |
Response 200 OK
Returns the export object with a presigned download URL.
Analytics
GET /api/v1/ai-gateway/data-forge/projects/:id/analytics
Get analytics and statistics for a Data Forge project.
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Project ID |
Response 200 OK
{
"data": {
"projectId": "proj_abc123",
"projectName": "Customer Support FAQ",
"projectStatus": "generating",
"documents": {
"total": 12,
"totalSizeBytes": 24576000,
"parsed": 12,
"pending": 0,
"processing": 0,
"failed": 0
},
"chunks": {
"total": 340,
"avgLength": 812,
"minLength": 120,
"maxLength": 1024
},
"pairs": {
"total": 1020,
"accepted": 850,
"rejected": 45,
"pending": 125,
"acceptanceRate": 83.3
},
"qualityDistribution": [
{ "range": "0.8-1.0", "count": 800, "avgScore": 0.91 },
{ "range": "0.6-0.8", "count": 150, "avgScore": 0.72 }
],
"generations": {
"total": 3,
"completed": 2,
"running": 1,
"pending": 0,
"failed": 0
},
"pairsOverTime": [
{ "date": "2025-06-15", "total": 1020, "accepted": 850, "rejected": 45 }
],
"exportsCount": 2,
"lastExport": { /* the most recent export record, or null */ },
"createdAt": "2025-06-15T10:00:00Z",
"updatedAt": "2025-06-15T14:30:00Z"
},
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/data-forge/models
The teacher models a generation may use: the chat and multimodal models you may use (your own, shared with you, or open to all users) that are serving.
Response 200 OK
{
"data": [
{
"modelId": "6a502004c1c6457701e39093",
"name": "GPT-4o",
"provider": "openai",
"vendorModelId": "gpt-4o",
"type": "third-party",
"modelType": "chat",
"status": "active"
}
],
"meta": { "requestId": "req_abc123" }
}
GET /api/v1/ai-gateway/data-forge/embedding-models
The embedding models a project's semantic chunking may use: the embedding models you may use that are serving, including the platform's embedding model, which every user may use.
Response 200 OK
{
"data": [
{
"modelId": "165ab9b829cb6d776b350b32",
"name": "Snowflake Arctic Embed S",
"provider": "self_hosted",
"vendorModelId": "Snowflake/snowflake-arctic-embed-s",
"type": "self-hosted",
"modelType": "embedding",
"status": "active"
}
],
"meta": { "requestId": "req_abc123" }
}
Generation Status Values
| Status | Description |
|---|---|
pending | Job created, waiting to start |
parsing | Parsing documents into chunks |
generating | Generating Q&A pairs from chunks |
validating | Running quality validation and deduplication |
completed | Job finished successfully |
failed | Error occurred (check logs) |
cancelled | Stopped by user |
Pair Review Status Values
| Status | Description |
|---|---|
pending | Not yet reviewed |
accepted | Approved for export |
rejected | Excluded from export |
edited | Modified by reviewer, approved for export |