Skip to content

Directory · Observability

Weights & Biases

ML experiment tracking platform whose Weave product adds tracing and evaluations for LLM applications.

FreemiumWeights & Biases (CoreWeave) · 🇺🇸 United States

Best for

Teams already tracking ML experiments who need LLM evals and tracing

More observability tools

Arize AI

Observability · Arize AI · United States

Freemium

AI observability and evaluation platform for production LLM and ML systems; maker of the open-source Phoenix library.

Best for

Enterprise-scale monitoring, tracing, and evals of deployed AI systems

Braintrust

Observability · Braintrust · United States

Freemium

AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.

Best for

Eval-first teams, prompt A/B testing, quality loops tied to production traces

Pros

  • Traces to evals in one product
  • Strong experiment and scoring UX
  • Good for PM + eng collaboration

Cons

  • Not self-hostable (managed only)
  • Pricier at scale than OSS Langfuse
  • Less infra-flexible than OpenTelemetry

Alternatives

LangfuseLangSmithPhoenix

Giskard

Observability · Giskard · France

Open Source

Open-source testing framework from a French startup that scans LLM apps and ML models for bias, errors, and vulnerabilities.

Best for

Automated evaluation and red-teaming of LLM applications before release

Browse the full directory →