Skip to main content
Evaluations (/evaluations) is OpenLIT’s LLM evaluation surface under Monitor. Score production traces with an LLM-as-a-judge, enable built-in or custom evaluators, schedule Auto Evaluation, and review pass-rate analytics - or run the same criteria programmatically via the SDK for offline CI/CD gates. Open it from Monitor → Evaluations. The page has three tabs:

Analytics

Pass rates, executions, cost, and evaluator results over time

Evaluators

Enable built-in types or create custom evaluators

Configuration

Judge model, Vault API key, Auto Evaluation schedule and sampling
Default tab is Analytics. Use ?tab=evaluators or ?tab=configuration for the others.

Find the right feature

Methods

LLM-as-a-Judge

Automated scoring on live traces

Programmatic evals

SDK / CI offline evaluation

Manual Feedback

Good / Bad / Neutral human ratings