
Phoenix
by Arize Phoenix · Open-source tracing, evaluation and experimentation for AI agents
77.3 — BenchRank score out of 100
Screenshots of Phoenix
Homepage Pricing page
Overview
Phoenix is an open-source platform for developing and evaluating AI agents. It traces each step an agent takes — prompts, retrievals, tool calls and outputs — then lets you annotate runs, build datasets from traces, run experiments and score results on cost, latency and performance. It uses OpenTelemetry and runs locally, in Docker, on Kubernetes via Helm, or as a hosted Cloud instance.
- Best for
- AI engineers who need to trace, evaluate and iterate on LLM agents on their own infrastructure
- Pricing
- No prices are published for Phoenix itself — it is ELv2-licensed and self-hostable with 2 free Phoenix Cloud instances; the captured pricing page prices the separate Arize AX product instead (Free, $50/month Pro, custom Enterprise).
- Runs on
- WebSelf-hostedCLI
Strengths and trade-offs
Strengths
- Self-host so traces stay on your own infrastructure
- Native OpenTelemetry; works with any model, framework or language
- Runs locally, via Docker, on Kubernetes with Helm, or Phoenix Cloud
- Tracing, evals, datasets, experiments and a prompt IDE in one tool
Trade-offs
- Free Phoenix Cloud is capped at 2 instances
- Self-hosting means running, upgrading and scaling it yourself
- Pages steer scaled use cases to the paid Arize AX product
- No published pricing or limits for Phoenix itself
How this score is made up
Each dimension is scored out of 100 and combined into the headline score using fixed weights.
- MCP support
- 60 out of 100
- API quality
- 70 out of 100
- Documentation
- 90 out of 100
- Agent friendliness
- 98 out of 100
- Pricing transparency
- 100 out of 100
- Changelog
- 100 out of 100
- Marketing site structure
- 60 out of 100
- Operational trust
- 25 out of 100
Measured, but not part of the score
Useful to know, but not a mark for or against the product — so these do not affect the ranking.
- Openness
- 55 out of 100
- Maintenance
- 100 out of 100
This doesn’t look right — report a problem with Phoenix’s score
Alternatives in Model Hosting & Inference
Ranked 2
72.5 — BenchRank score out of 100Helicone
Helicone · AI gateway and LLM observability for routing, debugging and analysing apps
Best for: AI engineering teams routing, debugging and monitoring LLM calls across many providers
Ranked 3
72.2 — BenchRank score out of 100Replicate
Replicate · Run, fine-tune and deploy AI models through a cloud API
Best for: Developers who want to run, fine-tune or deploy AI models via an API without managing GPUs
Ranked 4
70.2 — BenchRank score out of 100Langfuse
Langfuse · Open-source tracing, evaluation and prompt management for LLM apps
Best for: Engineering teams tracing, evaluating and improving LLM apps who want open source and self-hosting.

