Skip to main content

Fine-Tuning

Create, manage, and monitor fine-tuning jobs for large language models. Fine-tuning allows you to customize a base model on your own training data to improve performance for specific tasks.

All endpoints require authentication via X-API-Key header and the appropriate scope.


FineTuningJob Object​

{
"_id": "67a1b9e0e4b0a1b2c3d4e5f0",
"jobId": "ftjob_abc123",
"name": "Customer Support Classifier",
"description": "Fine-tuned on 10k support tickets for intent classification",
"userId": "user_456",
"organizationId": "org_xyz",
"baseModel": "meta-llama/Llama-3.1-8B",
"method": "lora",
"status": "training",
"progress": 40,
"statusMessage": "Epoch 2/3",
"currentEpoch": 2,
"totalEpochs": 3,
"trainingConfig": {
"learningRate": 0.0002,
"batchSize": 4,
"epochs": 3,
"warmupSteps": 10,
"weightDecay": 0.01,
"gradientAccumulationSteps": 4,
"maxSequenceLength": 2048,
"loraConfig": { "r": 16, "alpha": 32, "dropout": 0.05, "targetModules": ["q_proj", "v_proj"] }
},
"datasetInfo": {
"name": "training_data.jsonl",
"size": 5242880,
"trainingSamples": 9000,
"validationSamples": 500
},
"hardware": { "instanceType": "g4dn.xlarge", "gpuCount": 1, "gpuType": "nvidia-t4", "storageGb": 50 },
"modelArtifacts": { "s3Bucket": null, "s3Key": null, "downloadUrl": null },
"createdAt": "2025-02-01T09:55:00Z",
"updatedAt": "2025-02-01T12:30:00Z",
"startedAt": "2025-02-01T10:00:00Z"
}

status is one of pending, preparing, pending_resources, building, deploying, training, evaluating, validating, validation_failed, completed, failed, cancelled, stopping, timeout. A failed job carries error (message, details, timestamp). A completed job has completed_at, the trained model's location in output_model, and its results in training_summary. Every job route's :id is the job's _id; jobId is the AI Gateway's own name for it.


GET /api/v1/ai-gateway/fine-tuning-jobs​

List all fine-tuning jobs accessible to the authenticated user.

Scope: fine-tuning:read

Query Parameters

ParameterTypeRequiredDescription
statusstringNoFilter by status (see the FineTuningJob object for the values)
baseModelstringNoFilter by base model name
qstringNoSearch by job name or description
limitintegerNoNumber of results to return (default: 50, max: 200)
cursorstringNometa.nextCursor of the previous page; omit for the first page
sortstringNoSort field, prefix with - for descending (default: -createdAt)

Response 200 OK

{
"data": [
{
"_id": "67a1b9e0e4b0a1b2c3d4e5f0",
"jobId": "ftjob_abc123",
"name": "Customer Support Classifier",
"description": "Fine-tuned on 10k support tickets for intent classification",
"userId": "user_456",
"organizationId": "org_xyz",
"baseModel": "meta-llama/Llama-3.1-8B",
"method": "lora",
"status": "training",
"progress": 40,
"trainingConfig": { /* see the FineTuningJob object */ },
"datasetInfo": { /* see the FineTuningJob object */ },
"createdAt": "2025-02-01T09:55:00Z",
"updatedAt": "2025-02-01T12:30:00Z"
}
],
"meta": {
"total": 15,
"limit": 50,
"nextCursor": null,
"requestId": "req_abc123"
}
}

POST /api/v1/ai-gateway/fine-tuning-jobs​

Create a new fine-tuning job. Only self-hosted HuggingFace base models can be fine-tuned; pick one from base models.

The trained model is written to a shared volume: pass outputVolumeId (an existing shared volume) or outputVolumeName (a new one the platform creates). Without either the request is refused.

Scope: fine-tuning:write

Request Body

{
"name": "Customer Support Classifier",
"baseModel": "meta-llama/Llama-2-7b-hf",
"trainingDataset": "hf://acme/support-tickets",
"outputVolumeName": "support-classifier-models",
"description": "Fine-tuned on 10k support tickets for intent classification",
"trainingMethod": "lora",
"epochs": 3,
"batchSize": 8,
"datasetFormat": "jsonl"
}
FieldTypeRequiredDescription
namestringYesHuman-readable job name
baseModelstringYesHuggingFace base model id (e.g. meta-llama/Llama-2-7b-hf). Vendor models such as gpt-4o cannot be fine-tuned
trainingDatasetstringYesThe training data: an hf:// dataset id, a URL, or the file named with volumeId / volumeFilePath. Also used as the dataset URL unless datasetUrl is given
outputVolumeIdstringOne of the twoThe shared volume the trained model is written to
outputVolumeNamestringOne of the twoThe name of a new shared volume to create for the trained model
descriptionstringNoJob description (default: empty)
trainingMethodstringNoTraining method (default: lora): lora, qlora, dora, rslora, adalora, pissa, longlora, full, freeze, oft, boft, ia3, prefix_tuning, prompt_tuning, dpo, kto, orpo, reward_modeling or ppo. The LoRA family takes its settings from loraConfig (and qloraConfig for qlora); the others from their own block, which starts from the Create Fine-Tuning Job page's defaults and takes your overrides: freezeConfig, oftConfig (also boft), ia3Config, longloraConfig, dpoConfig, ktoConfig, orpoConfig, ppoConfig (needs rewardModelPath) and rewardConfig
epochsnumberNoTraining epochs (default: 3)
batchSizenumberNoBatch size (default: 4)
datasetFormatstringNoDataset format (default: jsonl): chatml, jsonl, json, csv, parquet, arrow, alpaca, sharegpt, openai, llama_factory, huggingface, dpo, kto, reward, llava, qwen_vl, vlm_chat, image_caption, vqa or interleaved
datasetUrlstringNoWhere the training data is read from, in place of trainingDataset
volumeIdstringNoShared volume whose data half holds the training file
volumeFilePathstringNoThe training file's path in that volume's data half
dataVersionnumberNoVersion of that file (default: latest)
trainingobjectNoOverrides for the training settings below
environmentobjectNoOverrides for the compute settings below, including useSpot (train on spot capacity; when it is reclaimed, training resumes from its last checkpoint) and spotFallback (default true: on-demand when no spot is available)
datasetobjectNoOverrides for the dataset settings (for example urlVolumeId, split_ratio, cutoff_length)
loraConfigobjectNoOverrides for the LoRA settings below
qloraConfigobjectNoOverrides for the QLoRA settings below
advancedobjectNoOverrides for advanced training settings (for example bf16, gradientCheckpointing, flashAttention)
monitoringobjectNoOverrides for logging and checkpoint settings (for example loggingSteps, evalSteps, saveSteps)
validationobjectNoData validation before training: { enabled, sample_size, checks }

Defaults the endpoint applies before your overrides:

{
"training": {
"optimizer": "adamw_torch", "learningRate": 0.0002, "weightDecay": 0.01,
"scheduler": "cosine", "warmupSteps": 10, "warmupRatio": 0.03,
"numEpochs": 3, "batchSize": 4, "gradientAccumulationSteps": 4,
"maxGradNorm": 1.0, "seed": 42
},
"environment": {
"environmentId": "custom", "cpu": "4", "memory": "16GB", "storageGb": 50,
"gpuCount": 0, "gpuType": "nvidia-t4", "useSpot": false, "spotFallback": true
},
"loraConfig": {
"rank": 16, "alpha": 32, "dropout": 0.05, "targetModules": ["q_proj", "v_proj"],
"useRslora": false, "useDora": false, "loraPlusLrRatio": 1.0
},
"qloraConfig": {
"bits": 4, "computeDtype": "float16", "quantizationType": "nf4", "doubleQuantization": true
}
}

A dataset URL that is not an hf:// id is downloaded into the shared volume given as dataset.urlVolumeId before training.

An hf://org/dataset id is read from the HuggingFace Hub, both when the dataset is validated before training and when it is trained on, with the job owner's HuggingFace token (set in Profile > Integrations > HuggingFace). A gated or private dataset needs that token, with access to the dataset granted on huggingface.co: without a token, validation fails saying to add one; with a token that has no access, it fails saying to request access or update the token.

Response 201 Created

When a job is created:

{
"data": {
"jobId": "ftjob_abc123",
"name": "Customer Support Classifier",
"status": "pending",
"modelType": "chat",
"createdAt": "2025-02-01T09:55:00",
"message": "Fine-tuning job created successfully. Container will be deployed shortly."
},
"meta": {
"requestId": "req_abc123"
}
}

Errors, in the order they are checked:

  • 400 validation-error: name, baseModel or trainingDataset is missing (listed in details).
  • 422 invalid-state: your profile has no HuggingFace token (add one under Profile > Integrations).
  • 400 validation-error: the dataset URL is not an hf:// id and no dataset.urlVolumeId is given (Choose the shared volume the dataset downloaded from the URL is stored in).
  • 400 validation-error: the configuration is invalid; the message lists each problem (for example an unknown datasetFormat).
  • 400 validation-error: An output shared volume is required for the trained model. The trained model is written to a shared volume, and this endpoint has no field to name one, so every request that passes the checks above currently ends with this error.

GET /api/v1/ai-gateway/fine-tuning-jobs/:id​

Get a single fine-tuning job by ID.

Scope: fine-tuning:read

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Response 200 OK

Returns the full FineTuningJob object.


DELETE /api/v1/ai-gateway/fine-tuning-jobs/:id​

Delete a fine-tuning job. Running jobs must be stopped before deletion.

Scope: fine-tuning:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Response 204 No Content


POST /api/v1/ai-gateway/fine-tuning-jobs/:id/stop​

Stop a running fine-tuning job. The job will be marked as stopped and training will be terminated.

Scope: fine-tuning:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Response 200 OK​

The job, as GET /ai-gateway/fine-tuning-jobs/:id shows it.


POST /api/v1/ai-gateway/fine-tuning-jobs/:id/restart​

Restart a stopped or failed fine-tuning job. The job will be re-queued for execution.

Scope: fine-tuning:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Response 200 OK​

The job, as GET /ai-gateway/fine-tuning-jobs/:id shows it.


POST /api/v1/ai-gateway/fine-tuning-jobs/:id/deploy​

Deploy the resulting model from a completed fine-tuning job. Makes the fine-tuned model available for inference.

Scope: fine-tuning:write

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Body Parameters

ParameterTypeRequiredDescription
deploymentNamestringYesName of the self-hosted model
useSpotbooleanNoRun on spot capacity (cheaper, can be reclaimed: the model restarts on a new node). Default false
spotFallbackbooleanNoWith useSpot: fall back to on-demand when no spot capacity is available (default true); false waits for spot

Response 202 Accepted​

The self-hosted model the deploy made, as GET /ai-gateway/models/:id shows it. Its status reports the deploy: poll the model until it is running.


GET /api/v1/ai-gateway/fine-tuning-jobs/:id/logs​

Retrieve a fine-tuning job's logs from one source, oldest first.

Scope: fine-tuning:read

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Query Parameters

ParameterTypeRequiredDescription
logTypestringNoThe logs' source: build, validation (the dataset checks before training), training (default), deploy, or all

Response 200 OK

{
"data": [
{
"_id": "67a1c0f2e4b0a1b2c3d4e5f6",
"jobId": "ftjob_abc123",
"timestamp": "2025-02-01T10:00:05",
"level": "info",
"source": "training",
"message": "Training started with 3 epochs"
},
{
"_id": "67a1c23ae4b0a1b2c3d4e5f7",
"jobId": "ftjob_abc123",
"timestamp": "2025-02-01T10:05:30",
"level": "info",
"source": "training",
"message": "Epoch 1/3 completed - loss: 0.542"
}
],
"meta": {
"requestId": "req_abc123"
}
}

GET /api/v1/ai-gateway/fine-tuning-jobs/:id/metrics​

Retrieve training metrics for a fine-tuning job, including loss curves and evaluation results.

Scope: fine-tuning:read

Path Parameters

ParameterTypeRequiredDescription
idstringYesFine-tuning job ID

Response 200 OK

One entry per recorded training step, ordered by step:

{
"data": [
{
"jobId": "ftjob_abc123",
"epoch": 1.0,
"step": 100,
"globalStep": 100,
"percentComplete": 33.3,
"batchSize": 4,
"loss": 0.85,
"learningRate": 0.0002,
"gradientNorm": 1.12,
"timestamp": "2025-02-01T10:20:00"
},
{
"jobId": "ftjob_abc123",
"epoch": 2.0,
"step": 200,
"globalStep": 200,
"percentComplete": 66.7,
"batchSize": 4,
"loss": 0.54,
"evalLoss": 0.62,
"learningRate": 0.0002,
"gradientNorm": 0.87,
"timestamp": "2025-02-01T10:40:00"
}
],
"meta": {
"requestId": "req_abc123"
}
}

GET /api/v1/ai-gateway/fine-tuning-jobs/stats​

Get aggregate statistics for fine-tuning jobs in the organization.

Scope: fine-tuning:read

Response 200 OK

{
"data": {
"totalJobs": 15,
"activeJobs": 2,
"completedJobs": 8
},
"meta": {
"requestId": "req_abc123"
}
}

GET /api/v1/ai-gateway/fine-tuning/base-models​

List available base models that can be used for fine-tuning.

Scope: fine-tuning:read

Response 200 OK

{
"data": [
{
"id": "mistralai/Mistral-7B-v0.1",
"name": "Mistral 7B",
"size": "7B",
"type": "causal_lm",
"description": "Mistral AI's 7B model with excellent performance",
"supportedMethods": ["lora", "qlora", "full"],
"recommendedHardware": ["g4dn.xlarge", "g5.xlarge"]
},
{
"id": "google/flan-t5-base",
"name": "FLAN-T5 Base",
"size": "248M",
"type": "seq2seq_lm",
"description": "Google's instruction-tuned T5 model",
"supportedMethods": ["lora", "full"],
"recommendedHardware": ["g4dn.large", "t3.xlarge"]
}
],
"meta": {
"requestId": "req_abc123"
}
}

GET /api/v1/ai-gateway/fine-tuning/hardware​

List available hardware options for fine-tuning jobs.

Scope: fine-tuning:read

Response 200 OK

{
"data": [
{
"instanceType": "g4dn.xlarge",
"gpuType": "T4",
"gpuCount": 1,
"memoryGb": 16,
"vcpu": 4,
"recommendedFor": ["7B models", "LoRA", "QLoRA"]
},
{
"instanceType": "g5.2xlarge",
"gpuType": "A10G",
"gpuCount": 1,
"memoryGb": 32,
"vcpu": 8,
"recommendedFor": ["13B models", "Full fine-tuning"]
}
],
"meta": {
"requestId": "req_abc123"
}
}

Planning a job​

Check a configuration before you submit it: a job that does not fit its GPU fails after it has started, at a cost.

GET /api/v1/ai-gateway/fine-tuning/model-requirements​

The GPU memory and hardware a base model needs to be fine-tuned, for a method when you give one.

Scope: fine-tuning:read

ParameterTypeRequiredDescription
baseModelstringYesA base model id, from the base model list
methodstringNolora, qlora or full

Response 200 OK


POST /api/v1/ai-gateway/fine-tuning/recommend​

A recommended configuration (method, hyperparameters, hardware) for a base model, a use case and a dataset size.

Scope: fine-tuning:read

Request Body

FieldTypeRequiredDescription
baseModelstringYesA base model id
useCasestringNochat (default), instruction or classification
datasetSizeintegerNoTraining examples in the dataset

Response 200 OK


POST /api/v1/ai-gateway/fine-tuning-jobs/validate​

Check a configuration against its base model and hardware: { valid, errors, warnings } (for example, a GPU without enough memory for the model and method).

Scope: fine-tuning:read

Request Body

FieldTypeRequiredDescription
baseModelstringYesA base model id
methodstringYeslora, qlora or full
hardwareobjectNoThe GPU, for example { "gpu_type": "A100", "gpu_count": 1 }
methodConfigobjectNoThe method's settings, for example { "bits": 4 } for qlora

Response 200 OK


GET /api/v1/ai-gateway/fine-tuning-jobs/:id/status​

A job's status, read from the gateway on each call: jobId, status (queued, running, completed, failed), phase, progress, error, startedAt, completedAt and updatedAt. Lighter than the whole job.

Scope: fine-tuning:read

ParameterTypeRequiredDescription
idstringYesFine-tuning job id

Response 200 OK