Directory · Observability
Phoenix (Arize)
Open-source LLM tracing and evaluation from Arize — inspect spans, embeddings, and retrieval quality locally.
Best for
RAG debugging, embedding drift analysis, OSS tracing without vendor lock-in
Pros and cons
Pros
- Free and open source
- Strong RAG and retrieval visualization
- Pairs with Arize cloud for production scale
Cons
- Smaller community than Langfuse
- Eval story less mature than Braintrust
- Self-host setup takes effort
Alternatives
More observability tools
Arize AI
Observability · Arize AI · United States
AI observability and evaluation platform for production LLM and ML systems; maker of the open-source Phoenix library.
Best for
Enterprise-scale monitoring, tracing, and evals of deployed AI systems
Braintrust
Observability · Braintrust · United States
AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.
Best for
Eval-first teams, prompt A/B testing, quality loops tied to production traces
Pros
- Traces to evals in one product
- Strong experiment and scoring UX
- Good for PM + eng collaboration
Cons
- Not self-hostable (managed only)
- Pricier at scale than OSS Langfuse
- Less infra-flexible than OpenTelemetry
Alternatives
Giskard
Observability · Giskard · France
Open-source testing framework from a French startup that scans LLM apps and ML models for bias, errors, and vulnerabilities.
Best for
Automated evaluation and red-teaming of LLM applications before release