Features
- Side-by-side comparison — run multiple LLMs against the same prompt simultaneously and view their responses in parallel
- Performance metrics — compare response time and token usage (prompt tokens, completion tokens) for each model
- Cost analysis — see the estimated cost of each request to help you balance budget and performance
- Response quality — read and compare the full text of each model’s response for the same input
- Experiment history — all experiments are saved automatically so you can revisit and share previous comparisons
Access OpenGround
In the OpenLIT UI, click OpenGround in the left navigation sidebar. The landing page shows all previously created experiments. Click any experiment to open its results, or click Create new to start a fresh comparison.Run an experiment
1
List existing experiments
When you open OpenGround, you see an overview of all previously created experiments. Each row shows the experiment name, the models compared, and when it was last run.Browse existing experiments to find prior comparisons before creating a new one — you may have already run the comparison you need.
2
Create a new experiment
Click Create new to open the experiment editor.Configure your first LLM:
- Choose a provider (e.g., OpenAI, Anthropic, Mistral)
- Select a model from that provider
- Set model parameters such as temperature and max tokens
- Enter the provider API key to authorize requests
3
Enter your prompt and compare
Once both models are configured, type your prompt into the prompt editor and click Compare Response.OpenGround sends the prompt to both models in parallel and displays the results side-by-side as they arrive.
4
Review the results
For each model, OpenGround shows:
Use these metrics together to evaluate the trade-off between response quality, latency, and cost for your specific use case.
Comparison metrics
OpenGround surfaces three categories of metrics to help you choose between models: Cost — estimated cost per request based on each provider’s per-token pricing. Use this to identify which model fits your budget, especially when running high-volume workloads. Duration — end-to-end response time from request to final token. Use this to evaluate latency-sensitive applications such as chatbots or real-time assistants. Response tokens — the number of output tokens generated. Longer responses consume more tokens and cost more. Compare response length alongside response quality to find models that are concise without sacrificing accuracy.Cost estimates in OpenGround are calculated from publicly listed pricing at the time the experiment runs. Actual costs may differ based on your provider agreement, caching, or pricing changes.
Experiment history
Every experiment you run is saved automatically. You can return to OpenGround at any time to review past results, share them with teammates, or re-run the same prompt against updated model versions.Evaluations
Score LLM outputs for hallucination, bias, and toxicity automatically
Manage LLM secrets
Centrally store LLM API keys for use across experiments without re-entering them
