Product overview
What is Arize AI?
Arize AI is an engineering platform for observing, evaluating, and improving AI agents and applications. It helps development teams inspect how an AI system behaves in production, evaluate outputs against defined criteria, and use those findings during further development and experimentation. The platform is organized around a continual improvement loop described as trace, evaluate, and learn.
Arize offers Arize AX as a managed AI engineering platform and Phoenix as an open-source option for observability and evaluation. Its capabilities cover tracing agent execution, analyzing production behavior, running evaluations, comparing experiments, and using production signals to guide changes to prompts, models, or application logic.
Core Features
- Agent tracing: Captures execution details so teams can inspect multi-step agent behavior and identify where a result went wrong.
- Observability: Monitors AI application behavior and production signals over time.
- Evaluations: Measures outputs with configured evaluators to identify quality, safety, or performance issues.
- Experimentation: Supports comparing development changes before they are promoted to production.
- Production learning loop: Connects observed outcomes back to the process of improving agents.
- Managed and open-source products: Provides the managed Arize AX platform and the open-source Phoenix project.
Use Cases
- Agent debugging: Engineers can follow traces to understand tool calls, intermediate steps, and unexpected outputs.
- AI quality evaluation: Teams can run repeatable checks across prompts, models, and application versions.
- Production monitoring: Operators can detect changes or failures in deployed AI behavior.
- Development experiments: AI teams can compare proposed changes using datasets and evaluation results.
Frequently Asked Questions
What is Phoenix?
Phoenix is Arize's open-source product for AI observability and evaluation.
Is Arize AI only for production monitoring?
No. The official site also describes evaluation, development, and experimentation capabilities used before and after deployment.


