Skip to main content

Data Forge

Generate synthetic fine-tuning datasets from your documents. Data Forge lets you upload source documents, parse and chunk them, generate Q&A training pairs using a teacher LLM, review and curate the output, and export it for fine-tuning.

All endpoints require authentication via X-API-Key header.


DataForgeProject Object​

{
"projectId": "proj_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"name": "Customer Support FAQ",
"description": "Generate training data from support documentation",
"status": "created",
"config": {
"chunkStrategy": "semantic",
"embeddingModelId": "165ab9b829cb6d776b350b32",
"chunkSize": 1024,
"chunkOverlap": 128,
"generationModel": null,
"generationPromptTemplate": null,
"pairsPerChunk": 3,
"outputFormat": "chatml",
"systemPrompt": "",
"temperature": 0.7,
"maxTokens": 2048
},
"stats": {
"totalDocuments": 12,
"totalChunks": 340,
"totalPairs": 1020,
"acceptedPairs": 850,
"rejectedPairs": 45,
"pendingPairs": 125,
"totalGenerations": 3,
"avgQualityScore": 0.87
},
"createdAt": "2025-06-15T10:00:00Z",
"updatedAt": "2025-06-15T14:30:00Z"
}

config.chunk_strategy is how documents are split into chunks: semantic (a chunk ends where the meaning shifts, judged by the embeddings of embeddingModelId), heading (one chunk per section), paragraph (whole paragraphs, the default) or sliding_window (fixed-size windows overlapping by chunkOverlap). Every strategy keeps a chunk within chunkSize characters. embeddingModelId is required with semantic and null otherwise; it must be an embedding model you may use, from GET /api/v1/ai-gateway/data-forge/embedding-models.

status is created for a new project and moves to parsing, parsed and generating as you work on it. stats.avg_quality_score is null while no pair has been scored. A project that has been exported also carries exports, one record per export.


Projects​

POST /api/v1/ai-gateway/data-forge/projects​

Create a new Data Forge project.

Request Body

{
"name": "Customer Support FAQ",
"description": "Generate training data from support documentation",
"chunkStrategy": "semantic",
"embeddingModelId": "165ab9b829cb6d776b350b32"
}
FieldTypeRequiredDescription
namestringYesProject name (1-200 characters)
descriptionstringNoProject description (max 2000 characters)
chunkStrategystringNosemantic, heading, paragraph (default) or sliding_window
embeddingModelIdstringWith semanticThe embedding model semantic chunking splits by

Returns 400 validation-error if name is missing or not a string, the strategy is not one of these, semantic has no embeddingModelId, or the model is not an embedding model you may use. The new project starts with the default config shown in the DataForgeProject object; change it with PUT.

Response 201 Created

Returns the full DataForgeProject object.


GET /api/v1/ai-gateway/data-forge/projects​

List all Data Forge projects for the authenticated user, scoped to the organization.

Response 200 OK

Returns a list of DataForgeProject objects.


GET /api/v1/ai-gateway/data-forge/projects/:id​

Get details of a specific Data Forge project.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Response 200 OK

Returns the full DataForgeProject object.


PATCH /api/v1/ai-gateway/data-forge/projects/:id​

Update a Data Forge project.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Request Body

{
"name": "Updated Project Name",
"description": "Updated description",
"config": {}
}
FieldTypeRequiredDescription
namestringNoUpdated project name (1-200 characters)
descriptionstringNoUpdated description (max 2000 characters)
configobjectNoThe config fields to change; the others keep their values. Chunking changes apply to the documents parsed afterwards.

Response 200 OK

Returns the updated DataForgeProject object.


DELETE /api/v1/ai-gateway/data-forge/projects/:id​

Delete a Data Forge project and clean up associated S3 objects, documents, chunks, and pairs.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Response 200 OK

{
"data": { "success": true },
"meta": { "requestId": "req_abc123" }
}

Documents​

POST /api/v1/ai-gateway/data-forge/projects/:id/uploads​

Get a presigned PUT URL for uploading a document directly to S3.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Request Body

{
"fileName": "product-manual.pdf",
"mimeType": "application/pdf"
}
FieldTypeRequiredDescription
fileNamestringYesName of the file to upload (1-500 characters)
mimeTypestringYesMIME type of the file (e.g. application/pdf, text/plain)

Returns 400 validation-error if fileName or mimeType is missing.

Response 200 OK

{
"data": {
"uploadUrl": "https://s3.amazonaws.com/bucket/data-forge/...",
"s3Key": "data-forge/user_456/proj_abc123/sources/product-manual.pdf",
"documentId": "doc_xyz789",
"filename": "product-manual.pdf",
"expiresIn": 3600
},
"meta": { "requestId": "req_abc123" }
}

PUT the file bytes to uploadUrl (valid for expiresIn seconds), then register the document with s3Key.


POST /api/v1/ai-gateway/data-forge/projects/:id/documents​

Register a document that has been uploaded to S3. Call this after the browser finishes uploading to the presigned URL.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Request Body

{
"name": "product-manual.pdf",
"mimeType": "application/pdf",
"fileSize": 2048576,
"s3Key": "data-forge/user_456/proj_abc123/sources/product-manual.pdf"
}
FieldTypeRequiredDescription
namestringYesDocument file name (1-500 characters)
mimeTypestringYesMIME type of the uploaded file
fileSizenumberYesFile size in bytes (at least 1)
s3KeystringYesThe s3Key returned by upload-url

Returns 400 validation-error if any field is missing or fileSize is not a number.

Response 201 Created

{
"data": {
"documentId": "doc_xyz789",
"projectId": "proj_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"filename": "product-manual.pdf",
"name": "product-manual.pdf",
"s3Key": "data-forge/user_456/proj_abc123/sources/product-manual.pdf",
"fileSize": 2048576,
"contentType": "application/pdf",
"mimeType": "application/pdf",
"status": "uploaded",
"parseStatus": "pending",
"parsingStatus": "pending",
"chunkCount": 0,
"metadata": {},
"parsedMetadata": {},
"createdAt": "2025-06-15T10:05:00Z",
"updatedAt": "2025-06-15T10:05:00Z"
},
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/data-forge/projects/:id/documents​

List all documents in a Data Forge project.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Response 200 OK

Returns a list of document objects.


DELETE /api/v1/ai-gateway/data-forge/projects/:projectId/documents/:id​

Delete a document from a project and remove from S3.

Path Parameters

ParameterTypeRequiredDescription
projectIdstringYesProject ID
idstringYesDocument ID

Response 200 OK

{
"data": { "success": true },
"meta": { "requestId": "req_abc123" }
}

Chunks​

GET /api/v1/ai-gateway/data-forge/projects/:id/chunks​

Get paginated chunks for a Data Forge project. Chunks are created when documents are parsed.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Query Parameters

ParameterTypeRequiredDescription
pageintegerNoPage number, 1-based (default: 1)
limitintegerNoItems per page, 1-500 (default: 50)

Response 200 OK

skip is the number of chunks before this page and limit the page size. contentType is the MIME type of the source document and position is the chunk's place within it.

{
"data": {
"chunks": [
{
"chunkId": "chunk_001",
"projectId": "proj_abc123",
"documentId": "doc_xyz789",
"userId": "user_456",
"organizationId": "org_xyz",
"content": "To reset your password, navigate to Settings > Security...",
"heading": "Password Reset",
"position": 12,
"contentType": "application/pdf",
"pairsGenerated": 3,
"status": "ready",
"createdAt": "2025-06-15T10:10:00Z"
}
],
"total": 340,
"skip": 0,
"limit": 50
},
"meta": { "requestId": "req_abc123" }
}

PATCH /api/v1/ai-gateway/data-forge/projects/:projectId/chunks/:id​

Update a specific chunk's content or metadata.

Path Parameters

ParameterTypeRequiredDescription
projectIdstringYesProject ID
idstringYesChunk ID

Request Body

{
"content": "Updated chunk text content",
"metadata": {},
"excluded": false
}
FieldTypeRequiredDescription
contentstringNoUpdated chunk text content
metadataobjectNoUpdated chunk metadata
excludedbooleanNoWhether to exclude this chunk from generation

Response 200 OK

Returns the updated chunk object.


Pipeline​

POST /api/v1/ai-gateway/data-forge/projects/:id/parse​

Start a document parsing job. Parses the uploaded documents not yet parsed (or whose parse failed) into text chunks using a K8s Job, with the project's chunking settings. Semantic chunking embeds each document's sentences with the project's embedding model, as you. Returns 400 if the project chunks semantically without an embedding model, or with one you may no longer use.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Response 200 OK

{
"data": {
"jobName": "df-parse-gen_pars",
"mode": "parse",
"projectId": "proj_abc123",
"documentCount": 12,
"status": "launched"
},
"meta": { "requestId": "req_abc123" }
}

POST /api/v1/ai-gateway/data-forge/projects/:id/generations​

Start a data generation job. Uses an AI model to generate training pairs from parsed chunks.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Request Body

{
"teacherModelId": "gpt-4o",
"outputFormat": "qa",
"pairsPerChunk": 3,
"difficultyDistribution": {"easy": 0.3, "medium": 0.5, "hard": 0.2},
"systemPrompt": "Generate high-quality Q&A pairs from the provided context.",
"temperature": 0.7,
"styleTemplate": "mixed",
"thresholds": {"minQualityScore": 0.7, "minGroundingScore": 0.7}
}
FieldTypeRequiredDescription
teacherModelIdstringYesID of the teacher model that generates the pairs (from available models)
outputFormatstringYesGeneration type: qa, instruction, conversation, summary, or classification
pairsPerChunkintegerNoPairs to generate per chunk (default: the project's setting, 3 unless changed)
difficultyDistributionobjectNoDifficulty mix as fractions, e.g. {"easy": 0.3, "medium": 0.5, "hard": 0.2}
systemPromptstringNoCustom system prompt for the teacher model (max 10,000 characters; default: the project's setting)
temperaturefloatNoSampling temperature (0.0-2.0, default: 0.7)
styleTemplatestringNoStyle template applied to generated pairs
thresholdsobjectNoAuto-filter thresholds, e.g. {"minQualityScore": 0.7, "minGroundingScore": 0.7}

Returns 400 validation-error if teacherModelId or outputFormat is missing.

Response 201 Created

{
"data": {
"jobName": "df-generate-gen_abc1",
"mode": "generate",
"projectId": "proj_abc123",
"generationId": "gen_abc123",
"chunkCount": 340,
"status": "launched"
},
"meta": { "requestId": "req_abc123" }
}

Poll GET /ai-gateway/data-forge/projects/-/generations/:id with generationId for progress.


POST /api/v1/ai-gateway/data-forge/projects/-/generations/:id/cancel​

Cancel a running generation job.

Path Parameters

ParameterTypeRequiredDescription
idstringYesGeneration job ID

Response 200 OK: the generation run, as GET /ai-gateway/data-forge/projects/-/generations/:id shows it, status cancelled. Pairs it already generated are kept.


Generations​

GET /api/v1/ai-gateway/data-forge/projects/:id/generations​

List all generation jobs for a project.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Query Parameters

ParameterTypeRequiredDescription
limitintegerNoNumber of results to return (default: 50, max: 200)
cursorstringNometa.nextCursor of the previous page; omit for the first page

Response 200 OK

{
"data": [
{
"generationId": "gen_abc123",
"projectId": "proj_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"status": "completed",
"config": {
"modelId": "gpt-4o",
"generationType": "qa",
"temperature": 0.7,
"pairsPerChunk": 3
},
"progress": 100,
"jobName": "df-generate-gen_abc1",
"stats": {
"totalChunksProcessed": 340,
"totalPairsGenerated": 1020,
"failedChunks": 0
},
"results": {
"chunksProcessed": 340,
"pairsGenerated": 1020,
"pairsValid": 980,
"avgQualityScore": 0.87,
"avgGroundingScore": 0.91,
"duplicateCount": 12,
"tokensUsed": 450000
},
"error": null,
"startedAt": "2025-06-15T11:00:30Z",
"completedAt": "2025-06-15T11:45:00Z",
"createdAt": "2025-06-15T11:00:00Z",
"updatedAt": "2025-06-15T11:45:00Z"
}
],
"meta": {
"total": 3,
"limit": 50,
"nextCursor": null,
"requestId": "req_abc123"
}
}

GET /api/v1/ai-gateway/data-forge/projects/-/generations/:id​

Get details of a specific generation job.

Path Parameters

ParameterTypeRequiredDescription
idstringYesGeneration job ID

Response 200 OK

data is the generation record, with the same fields as an entry in the generations list.


GET /api/v1/ai-gateway/data-forge/projects/-/generations/:id/logs​

Get logs for a specific generation job. Logs are streamed in real-time during active jobs. Returns the most recent 1,000 entries, oldest first.

Path Parameters

ParameterTypeRequiredDescription
idstringYesGeneration job ID

Response 200 OK

{
"data": [
{
"timestamp": "2025-06-15T11:00:30Z",
"level": "info",
"stage": "generation",
"message": "Starting generation for 340 chunks"
},
{
"timestamp": "2025-06-15T11:01:00Z",
"level": "info",
"stage": "generation",
"message": "Processing chunk 1/340: Password Reset"
},
{
"timestamp": "2025-06-15T11:01:05Z",
"level": "info",
"stage": "generation",
"message": "Generated 3 pairs for chunk 1 (avg quality: 0.92)"
}
],
"meta": { "requestId": "req_abc123" }
}

Pairs​

GET /api/v1/ai-gateway/data-forge/projects/:id/pairs​

Get paginated training pairs for a project with optional filtering.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Query Parameters

ParameterTypeRequiredDescription
pageintegerNoPage number, 1-based (default: 1)
limitintegerNoItems per page, 1-500 (default: 50)
statusstringNoFilter by review status: pending, accepted, rejected, edited
documentIdstringNoFilter by generation ID (only pairs produced by that generation)
qstringNoSearch within pair input/output text

Response 200 OK

skip is the number of pairs before this page and limit the page size.

{
"data": {
"pairs": [
{
"pairId": "pair_001",
"projectId": "proj_abc123",
"chunkId": "chunk_001",
"documentId": "doc_xyz789",
"generationId": "gen_abc123",
"userId": "user_456",
"organizationId": "org_xyz",
"type": "single_turn",
"systemPrompt": "You are a helpful customer support assistant.",
"messages": [
{"role": "system", "content": "You are a helpful customer support assistant."},
{"role": "user", "content": "How do I reset my password?"},
{"role": "assistant", "content": "To reset your password, navigate to Settings > Security..."}
],
"question": "How do I reset my password?",
"answer": "To reset your password, navigate to Settings > Security...",
"qualityScore": 0.92,
"groundingScore": 0.95,
"complexityLevel": "easy",
"questionType": "factual",
"contentHash": null,
"isDuplicate": false,
"duplicateOf": null,
"status": "accepted",
"reviewerNotes": null,
"editHistory": [],
"createdAt": "2025-06-15T11:01:05Z",
"updatedAt": "2025-06-15T11:20:00Z"
}
],
"total": 1020,
"skip": 0,
"limit": 50
},
"meta": { "requestId": "req_abc123" }
}

PATCH /api/v1/ai-gateway/data-forge/projects/-/pairs/:id​

Edit a training pair during review: rewrite its question, answer or system prompt, set its review status, or add reviewer notes. Only the fields given change.

Path Parameters

ParameterTypeRequiredDescription
idstringYesPair ID

Request Body

{
"question": "Updated question text",
"answer": "Updated answer text",
"status": "accepted"
}
FieldTypeRequiredDescription
questionstringNoThe pair's question (its input)
answerstringNoThe pair's answer (its output)
systemPromptstringNoThe system prompt stored with the pair
statusstringNoReview status: pending, accepted, rejected, edited
reviewerNotesstringNoFree-form reviewer notes

An edited question, answer or system prompt marks the pair manually_edited and rewrites its chat-format messages to match (with a system message only where the pair had one). The edit is recorded as yours.

Response 200 OK

data is the pair record after the update, with the fields shown in the pairs list.


POST /api/v1/ai-gateway/data-forge/projects/:id/pairs/bulk​

Perform a bulk action on the listed training pairs.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Request Body

{
"action": "accept",
"pairIds": ["pair_001", "pair_002", "pair_003"]
}
FieldTypeRequiredDescription
actionstringYesAction to perform: accept, reject, delete, reset
pairIdsstring[]YesIDs of the pairs to act on

Returns 400 validation-error if action is missing or pairIds is not an array.

Response 200 OK

{
"data": {
"action": "accept",
"affectedCount": 3
},
"meta": { "requestId": "req_abc123" }
}

Export​

POST /api/v1/ai-gateway/data-forge/projects/:id/exports​

Export the dataset as a JSONL file in the specified format, written into the data half of a shared volume. Only accepted pairs are included.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Request Body

{
"format": "chatml",
"includeSystemPrompt": true,
"minQualityScore": 0.8,
"volumeId": "vol_abc123"
}
FieldTypeRequiredDescription
formatstringYesExport format: chatml or alpaca
includeSystemPromptbooleanNoInclude system prompt in ChatML output (default: true)
minQualityScorefloatNoMinimum quality score filter for pairs (0.0-1.0)
volumeIdstringOne of volumeId / newVolumeNameAn existing shared volume you can edit; the export lands in its data half
newVolumeNamestringOne of volumeId / newVolumeNameCreate a new shared volume with this name and land the export in it

Returns 400 validation-error if format is missing or neither volumeId nor newVolumeName is given, and 404 not-found if the volume does not exist or you cannot edit it.

Export Formats

ChatML -- OpenAI-compatible chat format:

{"messages": [{"role": "system", "content": "..."}, {"role": "user", "content": "..."}, {"role": "assistant", "content": "..."}]}

Alpaca -- Instruction-following format:

{"instruction": "...", "input": "", "output": "..."}

Response 201 Created

{
"data": {
"version": 3,
"volumeId": "vol_abc123",
"path": "dataforge-export-v3.jsonl",
"format": "chatml",
"pairCount": 850,
"fileSizeBytes": 4521890
},
"meta": { "requestId": "req_abc123" }
}

path is the file's path in the volume's data half.


GET /api/v1/ai-gateway/data-forge/projects/:projectId/exports/:version​

Get a download URL for a specific export version.

Path Parameters

ParameterTypeRequiredDescription
projectIdstringYesProject ID
versionintegerYesExport version number

Response 200 OK

Returns the export object with a presigned download URL.


Analytics​

GET /api/v1/ai-gateway/data-forge/projects/:id/analytics​

Get analytics and statistics for a Data Forge project.

Path Parameters

ParameterTypeRequiredDescription
idstringYesProject ID

Response 200 OK

{
"data": {
"projectId": "proj_abc123",
"projectName": "Customer Support FAQ",
"projectStatus": "generating",
"documents": {
"total": 12,
"totalSizeBytes": 24576000,
"parsed": 12,
"pending": 0,
"processing": 0,
"failed": 0
},
"chunks": {
"total": 340,
"avgLength": 812,
"minLength": 120,
"maxLength": 1024
},
"pairs": {
"total": 1020,
"accepted": 850,
"rejected": 45,
"pending": 125,
"acceptanceRate": 83.3
},
"qualityDistribution": [
{ "range": "0.8-1.0", "count": 800, "avgScore": 0.91 },
{ "range": "0.6-0.8", "count": 150, "avgScore": 0.72 }
],
"generations": {
"total": 3,
"completed": 2,
"running": 1,
"pending": 0,
"failed": 0
},
"pairsOverTime": [
{ "date": "2025-06-15", "total": 1020, "accepted": 850, "rejected": 45 }
],
"exportsCount": 2,
"lastExport": { /* the most recent export record, or null */ },
"createdAt": "2025-06-15T10:00:00Z",
"updatedAt": "2025-06-15T14:30:00Z"
},
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/data-forge/models​

The teacher models a generation may use: the chat and multimodal models you may use (your own, shared with you, or open to all users) that are serving.

Response 200 OK

{
"data": [
{
"modelId": "6a502004c1c6457701e39093",
"name": "GPT-4o",
"provider": "openai",
"vendorModelId": "gpt-4o",
"type": "third-party",
"modelType": "chat",
"status": "active"
}
],
"meta": { "requestId": "req_abc123" }
}

GET /api/v1/ai-gateway/data-forge/embedding-models​

The embedding models a project's semantic chunking may use: the embedding models you may use that are serving, including the platform's embedding model, which every user may use.

Response 200 OK

{
"data": [
{
"modelId": "165ab9b829cb6d776b350b32",
"name": "Snowflake Arctic Embed S",
"provider": "self_hosted",
"vendorModelId": "Snowflake/snowflake-arctic-embed-s",
"type": "self-hosted",
"modelType": "embedding",
"status": "active"
}
],
"meta": { "requestId": "req_abc123" }
}

Generation Status Values​

StatusDescription
pendingJob created, waiting to start
parsingParsing documents into chunks
generatingGenerating Q&A pairs from chunks
validatingRunning quality validation and deduplication
completedJob finished successfully
failedError occurred (check logs)
cancelledStopped by user

Pair Review Status Values​

StatusDescription
pendingNot yet reviewed
acceptedApproved for export
rejectedExcluded from export
editedModified by reviewer, approved for export