Skip to content

Directory · Observability

Phoenix (Arize)

Open-source LLM tracing and evaluation from Arize — inspect spans, embeddings, and retrieval quality locally.

Open SourceArize AI · 🇺🇸 United States

Best for

RAG debugging, embedding drift analysis, OSS tracing without vendor lock-in

Pros and cons

Pros

  • Free and open source
  • Strong RAG and retrieval visualization
  • Pairs with Arize cloud for production scale

Cons

  • Smaller community than Langfuse
  • Eval story less mature than Braintrust
  • Self-host setup takes effort

Alternatives

More observability tools

Arize AI

Observability · Arize AI · United States

Freemium

AI observability and evaluation platform for production LLM and ML systems; maker of the open-source Phoenix library.

Best for

Enterprise-scale monitoring, tracing, and evals of deployed AI systems

Braintrust

Observability · Braintrust · United States

Freemium

AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.

Best for

Eval-first teams, prompt A/B testing, quality loops tied to production traces

Pros

  • Traces to evals in one product
  • Strong experiment and scoring UX
  • Good for PM + eng collaboration

Cons

  • Not self-hostable (managed only)
  • Pricier at scale than OSS Langfuse
  • Less infra-flexible than OpenTelemetry

Alternatives

LangfuseLangSmithPhoenix

Giskard

Observability · Giskard · France

Open Source

Open-source testing framework from a French startup that scans LLM apps and ML models for bias, errors, and vulnerabilities.

Best for

Automated evaluation and red-teaming of LLM applications before release

Browse the full directory →