Directory · Observability
Weights & Biases
ML experiment tracking platform whose Weave product adds tracing and evaluations for LLM applications.
Best for
Teams already tracking ML experiments who need LLM evals and tracing
More observability tools
Arize AI
Observability · Arize AI · United States
AI observability and evaluation platform for production LLM and ML systems; maker of the open-source Phoenix library.
Best for
Enterprise-scale monitoring, tracing, and evals of deployed AI systems
Braintrust
Observability · Braintrust · United States
AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.
Best for
Eval-first teams, prompt A/B testing, quality loops tied to production traces
Pros
- Traces to evals in one product
- Strong experiment and scoring UX
- Good for PM + eng collaboration
Cons
- Not self-hostable (managed only)
- Pricier at scale than OSS Langfuse
- Less infra-flexible than OpenTelemetry
Alternatives
Giskard
Observability · Giskard · France
Open-source testing framework from a French startup that scans LLM apps and ML models for bias, errors, and vulnerabilities.
Best for
Automated evaluation and red-teaming of LLM applications before release