Skip to content

Directory · Observability

Langfuse

Open-source LLM observability — traces, evals, prompt versioning, and cost dashboards for production AI apps.

Open SourceLangfuse · 🇩🇪 Germany

Best for

Debugging agent runs, tracking prompt changes, eval datasets, team visibility

Pros and cons

Pros

  • Self-hostable with generous OSS tier
  • Traces tie prompts to outputs and cost
  • Eval and dataset workflows built in

Cons

  • Self-host needs Postgres + ClickHouse appetite
  • Competes with a crowded observability market
  • Instrumentation takes engineering time upfront

Alternatives

More observability tools

Arize AI

Observability · Arize AI · United States

Freemium

AI observability and evaluation platform for production LLM and ML systems; maker of the open-source Phoenix library.

Best for

Enterprise-scale monitoring, tracing, and evals of deployed AI systems

Braintrust

Observability · Braintrust · United States

Freemium

AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.

Best for

Eval-first teams, prompt A/B testing, quality loops tied to production traces

Pros

  • Traces to evals in one product
  • Strong experiment and scoring UX
  • Good for PM + eng collaboration

Cons

  • Not self-hostable (managed only)
  • Pricier at scale than OSS Langfuse
  • Less infra-flexible than OpenTelemetry

Alternatives

LangfuseLangSmithPhoenix

Giskard

Observability · Giskard · France

Open Source

Open-source testing framework from a French startup that scans LLM apps and ML models for bias, errors, and vulnerabilities.

Best for

Automated evaluation and red-teaming of LLM applications before release

Browse the full directory →