Model Comparison
Library's comparison feature lets you test the same prompt against multiple AI Gateway models simultaneously, helping you choose the best model for your use case.
Getting Started
From a Prompt
- Open any prompt's detail page
- Click the Compare button
- Click Add Model to add models from your AI Gateway: the prompt's content is filled in as each model's system prompt (a system prompt) or user prompt (a user prompt or template)
- Write the user prompt each model is sent, if the prompt filled the system prompt
- If the prompt has
{{variables}}, enter their values under Variables (a blank one takes its default; one marked * has no default and needs a value). Every model's system and user prompt is sent with them filled in - Click Run Comparison
To compare without a saved prompt, go to /library/prompts/compare and write the prompts yourself.
Only chat models you may use that are serving (or idle and woken by the request) are offered. With none, the model picker says so: add a model in the AI Gateway first.
Configuring Models
You can compare up to 4 models simultaneously. For each model, configure:
| Parameter | Range | Default | Description |
|---|---|---|---|
| Temperature | 0.0 - 2.0 | 0.7 | Higher values = more creative, lower = more deterministic |
| Max Tokens | 1 - model limit | 1000 | Maximum response length |
| Top P | 0.0 - 1.0 | 1.0 | Nucleus sampling threshold |
Each model can have different parameters, allowing you to test variations like:
- Same prompt, different models
- Same model, different temperatures
- Different system prompts per model
Reading Results
After a comparison runs, each model card shows:
Response
The model's generated text output, displayed in a scrollable area.
Metrics
| Metric | Description |
|---|---|
| Response Time | Time in milliseconds from request to complete response |
| Input Tokens | Number of tokens in the prompt (system + user) |
| Output Tokens | Number of tokens in the response |
| Total Tokens | Sum of input and output tokens |
The best value for each metric is highlighted in green across all models, making it easy to identify the fastest or most efficient model.
Saving & Exporting
Export as JSON
Click Export JSON to download the complete comparison data including prompts, parameters, responses, and metrics. Useful for offline analysis or sharing with teammates.
Save to History
Click Save Comparison to store the results, with the variable values they ran with and optional notes. Saved comparisons appear in the History tab, newest first, ten to a page: search them by date, model, notes or variable value. Opening one restores its models and variable values.
Comparison History
The History tab shows every comparison saved from this page, newest first: a prompt's comparisons on its own Compare page, and those written without a saved prompt on /library/prompts/compare. Each shows:
- Date and time
- Models used (as badges)
- Notes
- Actions: View details or Delete
Tips
- Start with temperature 0.7 as a baseline, then test 0.3 (factual) and 1.0 (creative)
- Compare cost vs quality: A smaller model at lower cost might be sufficient for simple tasks
- Test edge cases: Try unusual inputs to see which model handles them best
- Save comparisons before and after prompt changes to track improvements
- Use the same user prompt across models for fair comparison