Skip to main content
OpenLIT uses OpenTelemetry to help you monitor NVIDIA and AMD GPUs for AI applications. Track GPU metrics like utilization, temperature, memory usage, and power consumption during AI training and inference workloads.

Choose your method

GPU monitoring can be implemented in two ways depending on your setup and requirements:

OpenLIT SDK

It is useful if you already have an AI application running on GPU that’s instrumented with OpenLIT.It extends your existing observability to include GPU metrics alongside your LLM traces.

OpenTelemetry GPU Collector

It is useful for remote GPUs with only LLM models hosted, containerized deployments.This approach allows you to get GPU metrics without modifying application code.

Supported Parameters

SDK Configuration Options

Environment Variables


Deploy OpenLIT

Deployment options for scalable LLM monitoring infrastructure

Integrations

60+ AI integrations with automatic instrumentation and performance tracking

Destinations

Send telemetry to Datadog, Grafana, New Relic, and other observability stacks