Skip to main content

Model Comparison

Library's comparison feature lets you test the same prompt against multiple AI Gateway models simultaneously, helping you choose the best model for your use case.

Getting Started​

From a Prompt​

  1. Open any prompt's detail page
  2. Click the Compare button
  3. Click Add Model to add models from your AI Gateway: the prompt's content is filled in as each model's system prompt (a system prompt) or user prompt (a user prompt or template)
  4. Write the user prompt each model is sent, if the prompt filled the system prompt
  5. If the prompt has {{variables}}, enter their values under Variables (a blank one takes its default; one marked * has no default and needs a value). Every model's system and user prompt is sent with them filled in
  6. Click Run Comparison

To compare without a saved prompt, go to /library/prompts/compare and write the prompts yourself.

Only chat models you may use that are serving (or idle and woken by the request) are offered. With none, the model picker says so: add a model in the AI Gateway first.

Configuring Models​

You can compare up to 4 models simultaneously. For each model, configure:

ParameterRangeDefaultDescription
Temperature0.0 - 2.00.7Higher values = more creative, lower = more deterministic
Max Tokens1 - model limit1000Maximum response length
Top P0.0 - 1.01.0Nucleus sampling threshold

Each model can have different parameters, allowing you to test variations like:

  • Same prompt, different models
  • Same model, different temperatures
  • Different system prompts per model

Reading Results​

After a comparison runs, each model card shows:

Response​

The model's generated text output, displayed in a scrollable area.

Metrics​

MetricDescription
Response TimeTime in milliseconds from request to complete response
Input TokensNumber of tokens in the prompt (system + user)
Output TokensNumber of tokens in the response
Total TokensSum of input and output tokens

The best value for each metric is highlighted in green across all models, making it easy to identify the fastest or most efficient model.

Saving & Exporting​

Export as JSON​

Click Export JSON to download the complete comparison data including prompts, parameters, responses, and metrics. Useful for offline analysis or sharing with teammates.

Save to History​

Click Save Comparison to store the results, with the variable values they ran with and optional notes. Saved comparisons appear in the History tab, newest first, ten to a page: search them by date, model, notes or variable value. Opening one restores its models and variable values.

Comparison History​

The History tab shows every comparison saved from this page, newest first: a prompt's comparisons on its own Compare page, and those written without a saved prompt on /library/prompts/compare. Each shows:

  • Date and time
  • Models used (as badges)
  • Notes
  • Actions: View details or Delete

Tips​

  • Start with temperature 0.7 as a baseline, then test 0.3 (factual) and 1.0 (creative)
  • Compare cost vs quality: A smaller model at lower cost might be sufficient for simple tasks
  • Test edge cases: Try unusual inputs to see which model handles them best
  • Save comparisons before and after prompt changes to track improvements
  • Use the same user prompt across models for fair comparison