Directory · Observability
Langfuse
Open-source LLM observability — traces, evals, prompt versioning, and cost dashboards for production AI apps.
Best for
Debugging agent runs, tracking prompt changes, eval datasets, team visibility
Pros and cons
Pros
- Self-hostable with generous OSS tier
- Traces tie prompts to outputs and cost
- Eval and dataset workflows built in
Cons
- Self-host needs Postgres + ClickHouse appetite
- Competes with a crowded observability market
- Instrumentation takes engineering time upfront
Alternatives
More observability tools
Arize AI
Observability · Arize AI · United States
AI observability and evaluation platform for production LLM and ML systems; maker of the open-source Phoenix library.
Best for
Enterprise-scale monitoring, tracing, and evals of deployed AI systems
Braintrust
Observability · Braintrust · United States
AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.
Best for
Eval-first teams, prompt A/B testing, quality loops tied to production traces
Pros
- Traces to evals in one product
- Strong experiment and scoring UX
- Good for PM + eng collaboration
Cons
- Not self-hostable (managed only)
- Pricier at scale than OSS Langfuse
- Less infra-flexible than OpenTelemetry
Alternatives
Giskard
Observability · Giskard · France
Open-source testing framework from a French startup that scans LLM apps and ML models for bias, errors, and vulnerabilities.
Best for
Automated evaluation and red-teaming of LLM applications before release