Skip to content

Directory · Observability

Braintrust

AI evaluation and observability platform — traces feed directly into eval datasets, regression tests, and prompt experiments.

Freemium

Best for

Eval-first teams, prompt A/B testing, quality loops tied to production traces

Pros and cons

Pros

  • Traces to evals in one product
  • Strong experiment and scoring UX
  • Good for PM + eng collaboration

Cons

  • Not self-hostable (managed only)
  • Pricier at scale than OSS Langfuse
  • Less infra-flexible than OpenTelemetry

Alternatives

LangfuseLangSmithPhoenix

More observability tools

Langfuse

Observability

Open Source

Open-source LLM observability — traces, evals, prompt versioning, and cost dashboards for production AI apps.

Best for

Debugging agent runs, tracking prompt changes, eval datasets, team visibility

Pros

  • Self-hostable with generous OSS tier
  • Traces tie prompts to outputs and cost
  • Eval and dataset workflows built in

Cons

  • Self-host needs Postgres + ClickHouse appetite
  • Competes with a crowded observability market
  • Instrumentation takes engineering time upfront

Alternatives

BraintrustPhoenix (Arize)LangSmith

Phoenix (Arize)

Observability

Open Source

Open-source LLM tracing and evaluation from Arize — inspect spans, embeddings, and retrieval quality locally.

Best for

RAG debugging, embedding drift analysis, OSS tracing without vendor lock-in

Pros

  • Free and open source
  • Strong RAG and retrieval visualization
  • Pairs with Arize cloud for production scale

Cons

  • Smaller community than Langfuse
  • Eval story less mature than Braintrust
  • Self-host setup takes effort

Alternatives

LangfuseBraintrustLangSmith

Browse the full directory →