Product overview
What is LangWatch?
LangWatch is an open-source LLMOps platform for testing, evaluating, observing, and governing AI agents. It helps teams turn production failures into repeatable tests, inspect complete execution traces, compare prompts and models, and monitor quality, cost, latency, and policy risks.
Its Scenario framework simulates multi-turn text or voice users from written specifications. A judge can assess the full trajectory, including tool and MCP calls, while developers run scenarios locally or in CI. The broader platform adds offline and online evaluations, OpenTelemetry-native tracing, prompt versioning, datasets, red teaming, dashboards, and deployment options for cloud, self-managed, or dedicated environments.
How to Use LangWatch
- Start with LangWatch Cloud or deploy the open-source stack in an approved environment.
- Connect an agent through a supported framework adapter, SDK, API, proxy, or OpenTelemetry.
- Capture traces with useful thread, user, version, and environment metadata.
- Define scenarios, simulated user behavior, success criteria, and expected tool calls.
- Run tests locally and in CI, then inspect failed turns and the judge's reasoning.
- Build datasets from traces and add offline experiments or live monitors for important quality signals.
- Configure retention, masking, access control, and alerts before tracing sensitive production traffic.
Core Features
- Agent simulations: Tests multi-turn text and voice conversations with configurable users and success criteria.
- Trajectory evaluation: Judges the whole conversation and verifies tools used during long interactions.
- Red teaming: Exercises jailbreak, unsafe behavior, refusal, and policy-failure scenarios.
- Offline and online evals: Supports batch experiments, production monitors, code checks, and LLM judges.
- Agent observability: Captures nested traces, sessions, tokens, cost, latency, and execution graphs.
- Prompt management: Versions prompts and compares changes across models, quality, cost, and latency.
- Open integrations: Works with popular frameworks, Python and TypeScript SDKs, APIs, and OpenTelemetry.
- Flexible deployment: Runs as managed cloud, self-managed software, or enterprise infrastructure.
Use Cases
- Test a support agent across long, realistic conversations before release.
- Verify that an agent chooses the right tool and respects business policies.
- Inject noise and interruptions into voice-agent simulations.
- Gate a pull request when evaluation scores fall below an agreed threshold.
- Monitor production topics, cost, latency, hallucination, toxicity, or PII risk.
- Reproduce a real production failure as a scenario and prevent recurrence.
Pricing
The Developer plan is free forever with no credit card and currently includes 50,000 events per month, two users, limited scenarios, simulations, and custom evaluations, plus 14-day data access. Growth is listed at €29 per core seat per month with included events and usage charges. Enterprise pricing is custom. Limits and rates can change, so confirm them on the pricing page.
Frequently Asked Questions
Can LangWatch be self-hosted?
Yes. The full platform can run on customer-managed infrastructure for greater data control. The organization remains responsible for securing, updating, and operating that environment.
Is the platform open source?
Yes. LangWatch states that its platform is available under the Apache 2.0 license, including testing, tracing, evaluations, prompts, datasets, and related core capabilities.
Are LLM judges enough for release approval?
No. Calibrate judges against human-reviewed examples, combine them with deterministic checks, and review high-risk failures manually.


