Skip to main content
Automated AI evaluation to assess and monitor the quality, safety, and performance of your LLM outputs across development and production environments. Open Monitor → Evaluations for Analytics, Evaluators, and Configuration. For evaluation types and custom evaluators, see Evaluators.

Find the right feature

Setup

1

Pick a judge model

Go to Evaluations → Settings, choose a provider and model to act as the judge (OpenAI, Anthropic, Google, Mistral, and 7+ others), and add its API key from Vault.
2

Enable evaluators

Switch to the Evaluation Types tab and turn on the evaluators you want. Hallucination, Bias, and Toxicity are enabled by default; Relevance, Coherence, Safety, and 5 others are opt-in.
3

Turn on Auto Evaluation

Back in Settings, enable Auto Evaluation with a cron schedule so every new trace gets scored automatically - or skip this and click Run Evaluation from any trace’s Evaluation tab to score it on demand.
4

Review results

Open any trace’s Evaluation tab for its score, classification, and reasoning, or check the Evaluations dashboard for aggregate trends across models and time.

Evaluators

Built-in and custom evaluation types

Configuration

Auto Evaluation schedule and judge model

LLM-as-a-Judge

Use advanced LLMs to evaluate AI application quality with automated scoring

Programmatic evaluations

Quick start guide for implementing custom evaluations in your code