Fine-Tuning
Create, manage, and monitor fine-tuning jobs for large language models. Fine-tuning allows you to customize a base model on your own training data to improve performance for specific tasks.
All endpoints require authentication via X-API-Key header and the appropriate scope.
FineTuningJob Object
{
"_id": "67a1b9e0e4b0a1b2c3d4e5f0",
"jobId": "ftjob_abc123",
"name": "Customer Support Classifier",
"description": "Fine-tuned on 10k support tickets for intent classification",
"userId": "user_456",
"organizationId": "org_xyz",
"baseModel": "meta-llama/Llama-3.1-8B",
"method": "lora",
"status": "training",
"progress": 40,
"statusMessage": "Epoch 2/3",
"currentEpoch": 2,
"totalEpochs": 3,
"trainingConfig": {
"learningRate": 0.0002,
"batchSize": 4,
"epochs": 3,
"warmupSteps": 10,
"weightDecay": 0.01,
"gradientAccumulationSteps": 4,
"maxSequenceLength": 2048,
"loraConfig": { "r": 16, "alpha": 32, "dropout": 0.05, "targetModules": ["q_proj", "v_proj"] }
},
"datasetInfo": {
"name": "training_data.jsonl",
"size": 5242880,
"trainingSamples": 9000,
"validationSamples": 500
},
"hardware": { "instanceType": "g4dn.xlarge", "gpuCount": 1, "gpuType": "nvidia-t4", "storageGb": 50 },
"modelArtifacts": { "s3Bucket": null, "s3Key": null, "downloadUrl": null },
"createdAt": "2025-02-01T09:55:00Z",
"updatedAt": "2025-02-01T12:30:00Z",
"startedAt": "2025-02-01T10:00:00Z"
}
status is one of pending, preparing, pending_resources, building, deploying, training, evaluating, validating, validation_failed, completed, failed, cancelled, stopping, timeout. A failed job carries error (message, details, timestamp). A completed job has completed_at, the trained model's location in output_model, and its results in training_summary. Every job route's :id is the job's _id; jobId is the AI Gateway's own name for it.
GET /api/v1/ai-gateway/fine-tuning-jobs
List all fine-tuning jobs accessible to the authenticated user.
Scope: fine-tuning:read
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
status | string | No | Filter by status (see the FineTuningJob object for the values) |
baseModel | string | No | Filter by base model name |
q | string | No | Search by job name or description |
limit | integer | No | Number of results to return (default: 50, max: 200) |
cursor | string | No | meta.nextCursor of the previous page; omit for the first page |
sort | string | No | Sort field, prefix with - for descending (default: -createdAt) |
Response 200 OK
{
"data": [
{
"_id": "67a1b9e0e4b0a1b2c3d4e5f0",
"jobId": "ftjob_abc123",
"name": "Customer Support Classifier",
"description": "Fine-tuned on 10k support tickets for intent classification",
"userId": "user_456",
"organizationId": "org_xyz",
"baseModel": "meta-llama/Llama-3.1-8B",
"method": "lora",
"status": "training",
"progress": 40,
"trainingConfig": { /* see the FineTuningJob object */ },
"datasetInfo": { /* see the FineTuningJob object */ },
"createdAt": "2025-02-01T09:55:00Z",
"updatedAt": "2025-02-01T12:30:00Z"
}
],
"meta": {
"total": 15,
"limit": 50,
"nextCursor": null,
"requestId": "req_abc123"
}
}
POST /api/v1/ai-gateway/fine-tuning-jobs
Create a new fine-tuning job. Only self-hosted HuggingFace base models can be fine-tuned; pick one from base models.
The trained model is written to a shared volume: pass outputVolumeId (an
existing shared volume) or outputVolumeName (a new one the platform creates).
Without either the request is refused.
Scope: fine-tuning:write
Request Body
{
"name": "Customer Support Classifier",
"baseModel": "meta-llama/Llama-2-7b-hf",
"trainingDataset": "hf://acme/support-tickets",
"outputVolumeName": "support-classifier-models",
"description": "Fine-tuned on 10k support tickets for intent classification",
"trainingMethod": "lora",
"epochs": 3,
"batchSize": 8,
"datasetFormat": "jsonl"
}
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Human-readable job name |
baseModel | string | Yes | HuggingFace base model id (e.g. meta-llama/Llama-2-7b-hf). Vendor models such as gpt-4o cannot be fine-tuned |
trainingDataset | string | Yes | The training data: an hf:// dataset id, a URL, or the file named with volumeId / volumeFilePath. Also used as the dataset URL unless datasetUrl is given |
outputVolumeId | string | One of the two | The shared volume the trained model is written to |
outputVolumeName | string | One of the two | The name of a new shared volume to create for the trained model |
description | string | No | Job description (default: empty) |
trainingMethod | string | No | Training method (default: lora): lora, qlora, dora, rslora, adalora, pissa, longlora, full, freeze, oft, boft, ia3, prefix_tuning, prompt_tuning, dpo, kto, orpo, reward_modeling or ppo. The LoRA family takes its settings from loraConfig (and qloraConfig for qlora); the others from their own block, which starts from the Create Fine-Tuning Job page's defaults and takes your overrides: freezeConfig, oftConfig (also boft), ia3Config, longloraConfig, dpoConfig, ktoConfig, orpoConfig, ppoConfig (needs rewardModelPath) and rewardConfig |
epochs | number | No | Training epochs (default: 3) |
batchSize | number | No | Batch size (default: 4) |
datasetFormat | string | No | Dataset format (default: jsonl): chatml, jsonl, json, csv, parquet, arrow, alpaca, sharegpt, openai, llama_factory, huggingface, dpo, kto, reward, llava, qwen_vl, vlm_chat, image_caption, vqa or interleaved |
datasetUrl | string | No | Where the training data is read from, in place of trainingDataset |
volumeId | string | No | Shared volume whose data half holds the training file |
volumeFilePath | string | No | The training file's path in that volume's data half |
dataVersion | number | No | Version of that file (default: latest) |
training | object | No | Overrides for the training settings below |
environment | object | No | Overrides for the compute settings below, including useSpot (train on spot capacity; when it is reclaimed, training resumes from its last checkpoint) and spotFallback (default true: on-demand when no spot is available) |
dataset | object | No | Overrides for the dataset settings (for example urlVolumeId, split_ratio, cutoff_length) |
loraConfig | object | No | Overrides for the LoRA settings below |
qloraConfig | object | No | Overrides for the QLoRA settings below |
advanced | object | No | Overrides for advanced training settings (for example bf16, gradientCheckpointing, flashAttention) |
monitoring | object | No | Overrides for logging and checkpoint settings (for example loggingSteps, evalSteps, saveSteps) |
validation | object | No | Data validation before training: { enabled, sample_size, checks } |
Defaults the endpoint applies before your overrides:
{
"training": {
"optimizer": "adamw_torch", "learningRate": 0.0002, "weightDecay": 0.01,
"scheduler": "cosine", "warmupSteps": 10, "warmupRatio": 0.03,
"numEpochs": 3, "batchSize": 4, "gradientAccumulationSteps": 4,
"maxGradNorm": 1.0, "seed": 42
},
"environment": {
"environmentId": "custom", "cpu": "4", "memory": "16GB", "storageGb": 50,
"gpuCount": 0, "gpuType": "nvidia-t4", "useSpot": false, "spotFallback": true
},
"loraConfig": {
"rank": 16, "alpha": 32, "dropout": 0.05, "targetModules": ["q_proj", "v_proj"],
"useRslora": false, "useDora": false, "loraPlusLrRatio": 1.0
},
"qloraConfig": {
"bits": 4, "computeDtype": "float16", "quantizationType": "nf4", "doubleQuantization": true
}
}
A dataset URL that is not an hf:// id is downloaded into the shared volume given as dataset.urlVolumeId before training.
An hf://org/dataset id is read from the HuggingFace Hub, both when the dataset is validated before training and when it is trained on, with the job owner's HuggingFace token (set in Profile > Integrations > HuggingFace). A gated or private dataset needs that token, with access to the dataset granted on huggingface.co: without a token, validation fails saying to add one; with a token that has no access, it fails saying to request access or update the token.
Response 201 Created
When a job is created:
{
"data": {
"jobId": "ftjob_abc123",
"name": "Customer Support Classifier",
"status": "pending",
"modelType": "chat",
"createdAt": "2025-02-01T09:55:00",
"message": "Fine-tuning job created successfully. Container will be deployed shortly."
},
"meta": {
"requestId": "req_abc123"
}
}
Errors, in the order they are checked:
400 validation-error:name,baseModelortrainingDatasetis missing (listed indetails).422 invalid-state: your profile has no HuggingFace token (add one under Profile > Integrations).400 validation-error: the dataset URL is not anhf://id and nodataset.urlVolumeIdis given (Choose the shared volume the dataset downloaded from the URL is stored in).400 validation-error: the configuration is invalid; the message lists each problem (for example an unknowndatasetFormat).400 validation-error:An output shared volume is required for the trained model. The trained model is written to a shared volume, and this endpoint has no field to name one, so every request that passes the checks above currently ends with this error.
GET /api/v1/ai-gateway/fine-tuning-jobs/:id
Get a single fine-tuning job by ID.
Scope: fine-tuning:read
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Response 200 OK
Returns the full FineTuningJob object.
DELETE /api/v1/ai-gateway/fine-tuning-jobs/:id
Delete a fine-tuning job. Running jobs must be stopped before deletion.
Scope: fine-tuning:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Response 204 No Content
POST /api/v1/ai-gateway/fine-tuning-jobs/:id/stop
Stop a running fine-tuning job. The job will be marked as stopped and training will be terminated.
Scope: fine-tuning:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Response 200 OK
The job, as GET /ai-gateway/fine-tuning-jobs/:id shows it.
POST /api/v1/ai-gateway/fine-tuning-jobs/:id/restart
Restart a stopped or failed fine-tuning job. The job will be re-queued for execution.
Scope: fine-tuning:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Response 200 OK
The job, as GET /ai-gateway/fine-tuning-jobs/:id shows it.
POST /api/v1/ai-gateway/fine-tuning-jobs/:id/deploy
Deploy the resulting model from a completed fine-tuning job. Makes the fine-tuned model available for inference.
Scope: fine-tuning:write
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Body Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
deploymentName | string | Yes | Name of the self-hosted model |
useSpot | boolean | No | Run on spot capacity (cheaper, can be reclaimed: the model restarts on a new node). Default false |
spotFallback | boolean | No | With useSpot: fall back to on-demand when no spot capacity is available (default true); false waits for spot |
Response 202 Accepted
The self-hosted model the deploy made, as GET /ai-gateway/models/:id shows it. Its status reports the deploy: poll the model until it is running.
GET /api/v1/ai-gateway/fine-tuning-jobs/:id/logs
Retrieve a fine-tuning job's logs from one source, oldest first.
Scope: fine-tuning:read
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Query Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
logType | string | No | The logs' source: build, validation (the dataset checks before training), training (default), deploy, or all |
Response 200 OK
{
"data": [
{
"_id": "67a1c0f2e4b0a1b2c3d4e5f6",
"jobId": "ftjob_abc123",
"timestamp": "2025-02-01T10:00:05",
"level": "info",
"source": "training",
"message": "Training started with 3 epochs"
},
{
"_id": "67a1c23ae4b0a1b2c3d4e5f7",
"jobId": "ftjob_abc123",
"timestamp": "2025-02-01T10:05:30",
"level": "info",
"source": "training",
"message": "Epoch 1/3 completed - loss: 0.542"
}
],
"meta": {
"requestId": "req_abc123"
}
}
GET /api/v1/ai-gateway/fine-tuning-jobs/:id/metrics
Retrieve training metrics for a fine-tuning job, including loss curves and evaluation results.
Scope: fine-tuning:read
Path Parameters
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job ID |
Response 200 OK
One entry per recorded training step, ordered by step:
{
"data": [
{
"jobId": "ftjob_abc123",
"epoch": 1.0,
"step": 100,
"globalStep": 100,
"percentComplete": 33.3,
"batchSize": 4,
"loss": 0.85,
"learningRate": 0.0002,
"gradientNorm": 1.12,
"timestamp": "2025-02-01T10:20:00"
},
{
"jobId": "ftjob_abc123",
"epoch": 2.0,
"step": 200,
"globalStep": 200,
"percentComplete": 66.7,
"batchSize": 4,
"loss": 0.54,
"evalLoss": 0.62,
"learningRate": 0.0002,
"gradientNorm": 0.87,
"timestamp": "2025-02-01T10:40:00"
}
],
"meta": {
"requestId": "req_abc123"
}
}
GET /api/v1/ai-gateway/fine-tuning-jobs/stats
Get aggregate statistics for fine-tuning jobs in the organization.
Scope: fine-tuning:read
Response 200 OK
{
"data": {
"totalJobs": 15,
"activeJobs": 2,
"completedJobs": 8
},
"meta": {
"requestId": "req_abc123"
}
}
GET /api/v1/ai-gateway/fine-tuning/base-models
List available base models that can be used for fine-tuning.
Scope: fine-tuning:read
Response 200 OK
{
"data": [
{
"id": "mistralai/Mistral-7B-v0.1",
"name": "Mistral 7B",
"size": "7B",
"type": "causal_lm",
"description": "Mistral AI's 7B model with excellent performance",
"supportedMethods": ["lora", "qlora", "full"],
"recommendedHardware": ["g4dn.xlarge", "g5.xlarge"]
},
{
"id": "google/flan-t5-base",
"name": "FLAN-T5 Base",
"size": "248M",
"type": "seq2seq_lm",
"description": "Google's instruction-tuned T5 model",
"supportedMethods": ["lora", "full"],
"recommendedHardware": ["g4dn.large", "t3.xlarge"]
}
],
"meta": {
"requestId": "req_abc123"
}
}
GET /api/v1/ai-gateway/fine-tuning/hardware
List available hardware options for fine-tuning jobs.
Scope: fine-tuning:read
Response 200 OK
{
"data": [
{
"instanceType": "g4dn.xlarge",
"gpuType": "T4",
"gpuCount": 1,
"memoryGb": 16,
"vcpu": 4,
"recommendedFor": ["7B models", "LoRA", "QLoRA"]
},
{
"instanceType": "g5.2xlarge",
"gpuType": "A10G",
"gpuCount": 1,
"memoryGb": 32,
"vcpu": 8,
"recommendedFor": ["13B models", "Full fine-tuning"]
}
],
"meta": {
"requestId": "req_abc123"
}
}
Planning a job
Check a configuration before you submit it: a job that does not fit its GPU fails after it has started, at a cost.
GET /api/v1/ai-gateway/fine-tuning/model-requirements
The GPU memory and hardware a base model needs to be fine-tuned, for a method when you give one.
Scope: fine-tuning:read
| Parameter | Type | Required | Description |
|---|---|---|---|
baseModel | string | Yes | A base model id, from the base model list |
method | string | No | lora, qlora or full |
Response 200 OK
POST /api/v1/ai-gateway/fine-tuning/recommend
A recommended configuration (method, hyperparameters, hardware) for a base model, a use case and a dataset size.
Scope: fine-tuning:read
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
baseModel | string | Yes | A base model id |
useCase | string | No | chat (default), instruction or classification |
datasetSize | integer | No | Training examples in the dataset |
Response 200 OK
POST /api/v1/ai-gateway/fine-tuning-jobs/validate
Check a configuration against its base model and hardware: { valid, errors, warnings } (for
example, a GPU without enough memory for the model and method).
Scope: fine-tuning:read
Request Body
| Field | Type | Required | Description |
|---|---|---|---|
baseModel | string | Yes | A base model id |
method | string | Yes | lora, qlora or full |
hardware | object | No | The GPU, for example { "gpu_type": "A100", "gpu_count": 1 } |
methodConfig | object | No | The method's settings, for example { "bits": 4 } for qlora |
Response 200 OK
GET /api/v1/ai-gateway/fine-tuning-jobs/:id/status
A job's status, read from the gateway on each call: jobId, status (queued, running,
completed, failed), phase, progress, error, startedAt, completedAt and updatedAt.
Lighter than the whole job.
Scope: fine-tuning:read
| Parameter | Type | Required | Description |
|---|---|---|---|
id | string | Yes | Fine-tuning job id |
Response 200 OK