Best LLM Observability Tools in 2026: 12 Platforms for Tracing, Evaluating, and Monitoring Production LLM Apps

Turn this article into takeaways for your work.

Each assistant summarizes the article only for you and suggests best practices for your work.

If your LLM app runs on LangChain, LangSmith covers the most ground with the least setup. If you want the full feature set for free and no vendor lock-in, Langfuse self-hosts at no cost and is now backed by ClickHouse's infrastructure. If traces need to sit next to the APM and infra dashboards you already pay for, Datadog LLM Observability is built for exactly that. And if evaluation, not a pretty dashboard, is the actual job you're hiring for, Braintrust and Galileo lead on testing an application's output systematically instead of eyeballing a log after a customer complains. This guide compares 12 platforms built to trace, evaluate, and monitor LLM applications in production, evaluated against each vendor's own pricing and documentation pages on October 2, 2026.

The thing that makes this category hard to shop is that three different jobs hide under one label. Most vendors claim to do all three and are genuinely strong at one. So instead of one overall ranking, this guide scores each tool per job, prices every option at a real volume instead of the free-tier headline, and treats "speaks OpenTelemetry" as the lock-in question it actually is rather than a checkbox.

Key Facts

  • The LLM observability platform market is on pace to grow from $1.97 billion in 2025 to $2.69 billion in 2026, a 36.3% jump, and reach $9.26 billion by 2030 at a 36.2% CAGR, per a Business Research Company report distributed through ResearchAndMarkets.
  • 2% of all LLM-call spans logged in production returned an error as of March 2026, and 60% of February 2026's logged errors came from exceeded rate limits rather than model failures, per Datadog's State of AI Engineering report.
  • Only 28% of LLM call spans show any cached-read input tokens even though most frontier models support prompt caching, per the same Datadog report, which is exactly the kind of waste a real trace catches and a transcript never will.
  • As of July 17, 2026, not one gen_ai.* span, event, metric, or attribute in OpenTelemetry's dedicated GenAI semantic-conventions repository carries a Stable badge, so every vendor in this guide claiming full OpenTelemetry support is building on a convention that is still moving, per an independent audit of the OTel registry.
  • Two of this category's best-known names changed owners in 2026 alone: ClickHouse acquired Langfuse in January alongside a $400 million Series D, confirmed on ClickHouse's own blog, and Dynatrace closed a roughly $915 million acquisition of Arize on October 1, confirmed on Dynatrace's announcement.

What "LLM Observability" Actually Covers

Three separate jobs get sold under this one label, and almost nobody is equally good at all three:

  1. Tracing and debugging. Capturing every prompt, completion, retrieval step, and tool call in a single run so you can see exactly where something went wrong, not just that it did.
  2. Evaluation and scoring. Running a test set against a new prompt or model version before you ship it (offline), or scoring live production traffic after it ships (online), usually with an LLM-as-judge or a custom scorer.
  3. Production monitoring and alerting. Watching latency, cost, error rate, and quality drift over time and paging someone when a number moves the wrong way.

A tool built for the first job can look thin on the third, and vice versa. Datadog's whole pitch is job three: correlating LLM spans with the APM and infrastructure alerts a platform team already watches. Braintrust's whole pitch is job two: treating evaluation as the release gate rather than a dashboard you check after the fact. Buying on "it has observability" without asking which of the three you're actually buying is how teams end up with a clean-looking dashboard and no idea whether the app is right.

This guide also covers single-call and RAG-style LLM applications, not just autonomous agents. General AI observability is the broadest version of this idea (logs, metrics, traces, and evals for any AI system). AI agent observability is the narrower case: tracing a multi-step loop where an agent calls tools and decides its own next action across five, ten, or fifty steps. If that loop, not a single prompt-response pair, is what you're instrumenting, read our dedicated agent observability roundup alongside this one. The vendor lists overlap heavily, but an agent buyer cares most about step-level tool-call tracing and evaluating multi-step behavior, while an LLM-app buyer (a chatbot, a RAG pipeline, a classification endpoint) cares more about raw trace economics at volume.

Quick Comparison Table

Tool Best For Starting Price Key Strength Key Limitation
LangSmith Teams already building on LangChain or LangGraph Free (1 seat); Plus $39/seat/mo Native tracing plus a real OTLP endpoint if you need to leave later Overage bills in opaque compute/storage units, not a flat per-trace rate
Langfuse Teams that want the full product free, self-hosted Free, self-hosted; Cloud from $29/mo MIT-licensed core, OpenTelemetry-native on every tier Enterprise features (SCIM, audit logs) jump straight to $2,499/mo
Arize (Phoenix and AX) Teams wanting open source now with an enterprise upgrade path Phoenix free (self-hosted); AX from $50/mo OpenTelemetry-native core, unlimited evals on paid entry tiers Phoenix's license is source-available (Elastic License 2.0), not permissive open source
Braintrust Teams that want evaluation and CI-gated releases as the job Free; Pro $249/mo Evaluation-first design: scorers, experiments, autonomous eval iteration Monthly fee doubles as model credit, which takes budgeting to understand
Datadog LLM Observability Teams already paying for Datadog APM Free to 40K LLM spans/mo; Pro $160/mo Only LLM-provider-call spans are billed; tool and workflow spans are free SaaS only, no self-hosted option at any tier
W&B Weave (CoreWeave Forge) ML teams already tracking experiments in Weights and Biases Free; Pro from $60/mo Familiar home for teams already living in W&B's experiment registry Now sold under CoreWeave Forge branding after a 2025 acquisition; OTel support undocumented
Helicone Teams that want the fastest possible setup Free; Pro $79/mo One-line proxy swap, Sessions group calls into one agent-style view No published flat overage rate; "usage-based" beyond the free tier isn't a number you can budget
HoneyHive Teams that want OpenTelemetry auto-instrumentation out of the box Free; Enterprise custom Natively built on OpenTelemetry, both automated and human evals on free tier No self-serve tier between Free and a sales-gated Enterprise
Portkey Teams that want an AI gateway and observability in one bill Free; Production $49/mo OpenTelemetry-compliant, bundles routing and fallbacks with the logs Evaluation is thin; guardrails exist but there's no dedicated scoring product
Traceloop (OpenLLMetry) Teams that want the instrumentation layer, not just a dashboard Free (OpenLLMetry SDK); Traceloop free to 50K spans/mo Apache 2.0 SDK that works with any OTel-compatible backend, including competitors here No self-serve paid tier; outgrowing 50K spans/mo means a sales call
Comet Opik Budget-conscious teams that still want a real paid tier Free (open source or cloud); Pro $19/mo Apache 2.0, full OTel support and LLM-as-judge even on the cheapest paid tier Smaller mindshare than LangSmith or Langfuse in most stacks
Galileo Enterprise teams wanting purpose-built judge models Free; Pro $100/mo (billed annually) Luna evaluator family plus real-time guardrails at the enterprise tier Month-to-month Pro rate and OpenTelemetry support are both undocumented publicly

Score by Job: Tracing, Evaluation, or Monitoring

Rate each candidate on the three jobs separately before you shortlist on a single "best overall" score, because that score almost never exists in this category.

Tool Tracing and Debugging Evaluation and Scoring Production Monitoring and Alerting
LangSmith Strong Strong Good
Langfuse Strong Strong Good
Arize (Phoenix and AX) Strong Strong Good
Braintrust Good Strong Limited
Datadog LLM Observability Good Good Strong
W&B Weave (CoreWeave Forge) Good Good Limited
Helicone Good Limited Strong
HoneyHive Strong Strong Good
Portkey Good Limited Strong
Traceloop (OpenLLMetry) Strong Good Good
Comet Opik Strong Strong Good
Galileo Good Strong Strong

A "Good" in monitoring usually means the dashboards exist but alerting isn't the product's reason for being. A "Limited" in evaluation usually means you'll pair the tool with a specialist (Braintrust or Galileo) rather than trust its built-in scoring to gate a release.

The Real Cost at 1 Million Traces a Month

Free-tier limits are where most of this category's marketing lives, and they're close to useless for budgeting because nobody runs a free-tier volume in production. Here's what a genuinely real month, 1 million billable units, actually costs on each vendor's own published overage math. The billed unit differs by vendor (traces, spans, observations, logs, or GB of data), so read the unit column before comparing the dollar figure.

Tool Billed Unit Cost at ~1M Units/Month How It's Calculated
LangSmith Traces, then opaque compute/storage units Not computable from published rates Overage bills in LangChain Standard Units ($1.00/LSU) blended with per-feature metering, not a flat per-trace number; budget conservatively above the $39/seat base
Langfuse Observations ("units") About $101/mo Core plan $29/mo includes 100K units; the remaining 900K fall in the $8/100K band, 9 x $8 = $72; $29 + $72 = $101
Arize AX Spans About $810/mo (reported) Pro $50/mo includes 50K spans; 950K overage spans at a reported $0.0008/span = $760; $50 + $760 = $810
Braintrust Scores and GB processed, not traces Not directly comparable Braintrust doesn't meter on trace count at all; budgeting needs a real score volume and data-processed estimate, not a trace figure
Datadog LLM Observability LLM-provider-call spans only About $475/mo (overage rate reported) Pro $160/mo includes 100K LLM-call spans; 900K overage at a reported $3.50/10K (annual) = $315; $160 + $315 = $475
W&B Weave (CoreWeave Forge) GB of Weave data ingested, not trace count Not directly comparable Billed on ingested data volume ($0.10/MB overage on Pro); 1M traces could be a few GB or dozens depending on payload size, so no single number is honest here
Helicone Requests Not computable from published rates Pro's overage beyond the included 10K requests is "usage-based" with no published flat rate; get a quote before budgeting a 1M-request month
HoneyHive Events Enterprise quote required Free caps at 10K events/mo with no published mid-tier; 1M events forces a sales conversation
Portkey Recorded logs About $130/mo Production $49/mo includes 100K logs; 900K overage at $9/100K = $81; $49 + $81 = $130
Traceloop (OpenLLMetry) Spans Enterprise quote required Free caps at 50K spans/mo with no published self-serve paid tier; 1M spans forces a sales conversation
Comet Opik Spans About $64/mo Pro Cloud $19/mo includes 100K spans; 900K overage at $5/100K = $45; $19 + $45 = $64
Galileo Traces Custom quote required Pro's $100/mo (annual) includes 50K traces; beyond that, pricing "scales based on number of traces" with no published flat rate

The flip worth sitting with: the two cheapest headline free tiers in this list (Opik and Langfuse) are also the two cheapest at real volume, because their overage math is public and linear. Datadog and Arize AX, both attractive on the free tier, get expensive fast once volume is real, and three vendors (Helicone, HoneyHive, Traceloop) simply don't publish a number past their free tier at all, which is itself a data point about who they're built for.

OpenTelemetry, Self-Hosting, and Licensing

OpenTelemetry support is the actual lock-in question in this category: a tool that ingests standard OTLP traces lets you swap the backend later without re-instrumenting a single line of application code. A tool that only works through its own proprietary SDK is cheaper to adopt today and more expensive to leave. Self-hosting matters here for a reason that's easy to forget: an LLM trace routinely contains the full prompt and completion, which means customer PII, health information, or anything else that can't leave your VPC travels with every single trace, not just with your database.

Tool OpenTelemetry Support Self-Host Option License (where open-core)
LangSmith Yes, OTLP endpoint with GenAI semantic-convention mapping Enterprise only Closed source
Langfuse Yes, native, every tier Yes, free, any scale MIT (core); separate license for the ee/ enterprise modules
Arize Phoenix Yes, built on OpenTelemetry/OpenInference Yes, free Elastic License 2.0 (source-available, not OSI-approved open source)
Arize AX Yes, same OTel core as Phoenix Enterprise only Closed source
Braintrust Listed as a supported integration path Enterprise only (on-prem/hosted-private) Closed source
Datadog LLM Observability Yes, natively supports the OTel GenAI semantic conventions No, SaaS only at any tier Closed source
W&B Weave (CoreWeave Forge) Not documented on CoreWeave's Forge pricing or product pages Personal tier only (single user, local) Closed source
Helicone Partial, via an async OpenLLMetry-based integration Enterprise only Apache License 2.0
HoneyHive Yes, natively built on OpenTelemetry Enterprise (self-hosted, hybrid, single-tenant) Closed source
Portkey Yes, OTel-compliant observability plus a dedicated OTel export feature Yes, free (open-source gateway tier) Not confirmed from pages this fetch could reach
Traceloop (OpenLLMetry) Yes, OpenTelemetry-native by design, this is the product's identity Enterprise (on-prem, including air-gapped) Apache 2.0 (the OpenLLMetry SDK itself)
Comet Opik Yes, native, every tier, more languages than most competitors Yes, free (open source) Apache License 2.0
Galileo Not confirmed from pages this fetch could reach Enterprise (hosted, VPC, or on-prem) Closed source

The licensing line is worth slowing down on. Three of the four open-core names here (Langfuse, Opik, Helicone) are genuinely permissive open source. The fourth, Arize Phoenix, is not: its LICENSE file names the Elastic License 2.0, a source-available license that lets you read, run, and modify the code but restricts offering Phoenix itself as a competing hosted service. That's a different legal position than Apache or MIT, so a procurement or legal team should check the actual LICENSE file, not a marketing page that just says "open source."

Two Acquisitions and a Shutdown Worth Knowing About

This category moved fast enough in 2026 that pricing isn't the only thing worth re-verifying before you buy.

Langfuse is now owned by ClickHouse. ClickHouse announced the acquisition on January 16, 2026, alongside a $400 million Series D round. ClickHouse has stated plainly that Langfuse "remains 100% open-source" under its existing MIT license, that self-hosted deployments need no changes, and that Langfuse Cloud continues as an independent product with its existing SLAs, per ClickHouse's announcement.

Arize is now owned by Dynatrace. The deal was valued at roughly $915 million (about $815 million in cash plus replacement equity) and closed October 1, 2026, one day before this article was fetched. Dynatrace's own announcement confirms "Arize AX and the Dynatrace platform remain available as standalone offerings, and Phoenix remains available as an open-source project," with a longer-term plan to connect Arize's workflows into Dynatrace's broader platform. Read the full statement on Dynatrace's blog. Pricing continuity wasn't addressed, so re-verify AX pricing before committing budget past a pilot.

Literal AI is dead. If you've seen it recommended in an older roundup, it shut down. Its own migration guide states the hosted service was "fully operational until October 31st 2025" and that "all data will be permanently deleted" after that date; the main literalai.com domain now 404s. That guide pointed departing users toward LangSmith or Langfuse, two of the stronger picks on this list anyway. Check for this kind of dead-product risk before you shortlist anything from a 2024 or 2025 source.

Self-Hosting and the Compliance Angle

An LLM trace is not a generic log line. It usually carries the full prompt sent to the model and the full completion that came back, so any customer name, contract clause, or health detail that passed through your app sits inside the observability vendor's database too, not just your own. For a team handling regulated data, that turns "which tool is cheapest" into a secondary question behind "which of these can legally run inside our own VPC."

Three platforms here answer that cleanly with a genuinely free, genuinely open self-hosted path: Langfuse, Arize Phoenix, and Comet Opik. Traceloop's OpenLLMetry SDK adds a fourth option at the instrumentation layer, since the SDK itself is free and portable even though Traceloop's own hosted dashboard gates self-hosting behind Enterprise. Everything else here (LangSmith, Braintrust, Datadog, W&B Weave, Helicone, HoneyHive, Portkey, Galileo) either requires an Enterprise contract to self-host or doesn't offer it at any price. If your AI agent's data privacy posture rules out sending prompt content to a third party's SaaS by policy, confirm that with legal before, not after, a pilot generates the first trace.

1. LangSmith: Deepest Tracing for LangChain and LangGraph Teams

LangSmith is LangChain's own tracing and evaluation platform, and the setup reflects it: point an environment variable at a LangChain or LangGraph app and every chain, tool call, and retry shows up as a run in a full execution tree with no extra instrumentation. It's not LangChain-only traffic, though. LangSmith exposes a real OTLP endpoint and maps OpenTelemetry's GenAI attributes (gen_ai.system, gen_ai.prompt, and related fields) to its own schema, so a mixed stack can still land everything in one place. Evaluation is built in rather than bolted on: online and offline evals, dataset creation, and human annotation queues ship on every tier. LangChain vs LlamaIndex works through LangSmith's pricing for a five-person team at 50,000 traces a month.

What you get What you don't
Deep native tracing for LangChain/LangGraph plus a real OTLP endpoint Developer tier caps at one seat
Online and offline evals, datasets, and annotation queues on every tier Overage bills in opaque LSU compute/storage units, not a flat per-trace rate
SaaS, hybrid, and fully self-hosted deployment at Enterprise Self-hosting isn't available below Enterprise

Pricing: Developer $0/seat (up to 5,000 base traces/month, 1 seat max). Plus $39/seat/month, unlimited seats (up to 10,000 base traces/month, one free small serverless deployment). Enterprise custom, with SaaS, hybrid, or self-hosted deployment. Overage beyond base traces bills in LangChain Standard Units at $1.00/LSU, blended with feature-specific metering. Source: langchain.com/pricing.

Best for: Teams already building on LangChain or LangGraph who want tracing that understands the framework with zero setup. Not ideal for: teams that want a predictable flat per-trace bill, since LSU-based overage is genuinely hard to model in advance.

2. Langfuse: The Open-Source Standard, Now Backed by ClickHouse

Langfuse remains the open-source default in this category: the full product (tracing, evaluation, prompt management, datasets) is free to self-host under the MIT license, with Docker Compose for local setup and Kubernetes templates for production. ClickHouse's January 2026 acquisition didn't change that; ClickHouse committed in writing to keeping Langfuse MIT-licensed with no changes required to existing self-hosted deployments. OpenTelemetry support runs across every tier, cloud or self-hosted, and evaluation (LLM-as-judge, datasets, experiments) ships even on the free Hobby plan rather than gated behind a paid tier.

What you get What you don't
Free, fully open-source self-hosting backed by a well-funded parent company Enterprise features (SCIM, audit logs, uptime SLA) start at $2,499/mo
OpenTelemetry-native ingestion on every plan Cloud Pro's 3-year data retention and Teams add-on stack up to $499/mo
LLM-as-judge, datasets, and experiments included on the free Hobby plan The ee/ directory (enterprise-only code) is not MIT, despite the rest of the repo

Pricing: Hobby free (50,000 units/month, 2 users, 30-day data access). Core $29/month or $348/year (100,000 units/month, unlimited users, 90-day access). Pro $199/month or $2,388/year (100,000 units/month, 3-year access, unlimited annotation queues; +$300/month Teams add-on). Enterprise $2,499/month (audit logs, SCIM, custom rate limits). Overage runs $8 per 100,000 units from 100K to 1M, stepping down to $6 per 100,000 at the highest volume tier. Self-hosting is free at any scale. Source: langfuse.com/pricing.

Best for: Teams that want the complete feature set for free by self-hosting, with a real exit path if they ever want to leave. Not ideal for: teams that specifically want to avoid being acquired-and-absorbed risk in a vendor, though ClickHouse's public commitments here are unusually concrete.

3. Arize (Phoenix and AX): Open Source Core, Now Under Dynatrace

Arize covers both ends of this category under one brand, and that brand now sits inside Dynatrace after an October 1, 2026 acquisition. Phoenix is the self-hosted half: built on OpenTelemetry and Arize's own OpenInference instrumentation, it runs on Docker or Kubernetes for free, though its license is the Elastic License 2.0, source-available rather than permissive open source, so check that distinction before assuming it behaves like an Apache-licensed project. Arize AX is the commercial evolution, and both AX and Phoenix remain standalone offerings per Dynatrace's own acquisition announcement. AX Free and Pro both include unlimited evaluations, unusually generous for a paid entry tier in this category.

What you get What you don't
Free, self-hosted Phoenix built on OpenTelemetry and OpenInference Phoenix's Elastic License 2.0 restricts offering it as a competing hosted service
Unlimited evals on both AX Free and AX Pro Self-hosting the commercial AX product needs Enterprise
Confirmed standalone status post-acquisition, in writing, from Dynatrace Future AX pricing and roadmap under a new parent company aren't addressed yet

Pricing: Phoenix free and open source, self-hosted, under the Elastic License 2.0. AX Free $0/month (reported: 25,000 spans/month, 1GB storage, 15-day retention). AX Pro $50/month (reported: 50,000 spans/month, 10GB storage, 30-day retention, overage around $0.0008/span). AX Enterprise custom. AX figures above are labeled (reported), from third-party aggregation; Phoenix's license is published in Arize-ai/phoenix's GitHub LICENSE file. Acquisition details confirmed on Dynatrace's announcement.

Best for: Teams that want a free, OpenTelemetry-native open-source starting point with a real enterprise upgrade path. Not ideal for: teams that assumed Phoenix was Apache or MIT licensed; verify the Elastic License 2.0's restrictions apply acceptably to your use case first.

4. Braintrust: Built Around Evaluation, Not Dashboards

Braintrust leads with evaluation as the primary product rather than a feature bolted onto a tracing tool. Every plan includes unlimited projects, datasets, playgrounds, and experiments, and its "Loop" feature runs autonomous evaluation iteration, generating test cases and refining scorers without a human writing every rubric from scratch. That makes it a natural fit for a CI-gated release process: run the eval suite against a new prompt or model version before it ships, not after a customer notices something's off.

What you get What you don't
Evaluation-first design: LLM-as-judge, custom code scorers, autonomous eval iteration Tracing is present but secondary; not the deepest trace view in this list
Unlimited projects, datasets, and experiments even on the free Starter plan On-prem or hosted-private deployment requires Enterprise
100+ pre-built integrations across model providers and frameworks Pricing meters on scores and GB processed, not traces, so it doesn't map to a per-trace budget

Pricing: Starter free ($10 in model credits/month, 1GB processed data/month, 10,000 scores/month, 14-day retention). Pro $249/month ($100 in model credits/month, 5GB processed data/month, 50,000 scores/month, 30-day retention). Enterprise custom, with on-prem or hosted-private deployment. Overage: processed data $4/GB (Starter) or $3/GB (Pro); scores $2.50 per 1,000 (Starter) or $1.50 per 1,000 (Pro). Source: braintrust.dev/pricing.

Best for: Teams that want evaluation, not monitoring, to be the center of how they ship changes. Not ideal for: a team whose primary need is deep step-level tracing of a running agent; pair Braintrust with a tracing-first tool if that's also a requirement.

5. Datadog LLM Observability: Agent Traces Next to the Infra You Already Watch

Datadog's entry (branded LLM Observability, increasingly folded into a broader "Agent Observability" line) wins through consolidation, not depth: if your infrastructure, APM, and log data already live in Datadog, LLM traces show up correlated with the same services and hosts you already monitor, instead of a sixth tab to check separately. Tracing covers the full execution path (prompts, retrieval, tool calls, agent decisions) with latency, token usage, and errors captured at each step, and it natively supports the OpenTelemetry GenAI semantic conventions. Billing is scoped tightly: only spans that hit an LLM provider count against the metered limit, so tool, workflow, and retrieval spans trace for free.

What you get What you don't
Traces correlated with existing APM, infrastructure, and RUM data SaaS only, no self-hosted or on-prem option at any tier
Only LLM-provider-call spans are billed; everything else traces free Needs an existing or new Datadog relationship to get full value
Built-in evaluators for hallucination, prompt injection, and PII exposure At real volume, overage adds up faster than the $160 headline suggests

Pricing: Free up to 40,000 LLM spans/month, 15-day retention. Pro $160/month for up to 100,000 LLM spans/month; additional on-demand usage billed beyond that, reported around $3.50 per 10,000 spans on an annual commitment. Retention add-ons available at 30, 60, or 90 days. Source: datadoghq.com/product/llm-observability.

Best for: Teams already running Datadog for APM and infrastructure who want LLM traces in the same place rather than standing up a new vendor. Not ideal for: teams without an existing Datadog relationship, or anyone who needs a self-hosted option for compliance reasons.

6. W&B Weave: Now Living Inside CoreWeave Forge

Weights and Biases, the experiment-tracking platform many ML teams already use for training runs, was acquired by CoreWeave in May 2025 for roughly $1.7 billion. Weave, its LLM and agent observability product, now ships under the "CoreWeave Forge" brand rather than wandb.ai, though the underlying pricing and feature set carried over largely unchanged. The pitch is still the same: if your ML team already lives in Weave or W&B for training history, agent traces and evaluations land in the same registry instead of a separate tool with a separate login. It traces at the operation level and includes LLM-as-a-judge metrics and PII redaction on every tier, including free.

What you get What you don't
Familiar home for ML teams already using W&B for training runs Now rebranded to CoreWeave Forge, which some existing users won't expect
LLM-as-a-judge metrics and PII redaction on every tier, including free OpenTelemetry support isn't documented anywhere on CoreWeave's Forge pages
A free single-user Personal tier for local experimentation Billed on GB of ingested trace data, not trace count, so budgeting needs a payload-size estimate

Pricing: Free $0/month (up to 5 model seats, 5GB storage/month, 1GB Weave data ingestion/month). Pro from $60/month (up to 10 seats, 100GB storage/month, 1.5GB ingestion/month; overage $0.03/GB storage, $0.10/MB ingestion). Enterprise custom (HIPAA, SSO, audit logs; ingestion overage $0.20/MB). An Academic tier offers 200GB storage and 25GB/month ingestion for free. Source: coreweave.com/forge-pricing (the legacy wandb.ai/site/pricing URL redirects there).

Best for: ML teams already running experiments in Weights and Biases who want agent and LLM tracing in the same platform. Not ideal for: teams that specifically need OpenTelemetry portability; it isn't documented as a supported path here.

7. Helicone: The Fastest Proxy-Based Setup

Helicone's original design is a proxy: point API calls at Helicone instead of directly at OpenAI or Anthropic, and logging happens automatically with essentially no code change. That's still its fastest path to value for teams that mostly want cost, latency, and request-level visibility without touching instrumentation. Multi-step visibility comes through Sessions, which groups related calls (LLM calls, vector queries, tool calls) into one view of a full run. An async OpenLLMetry-based integration exists for teams that want OTel-style logging without routing traffic through the proxy, though that's a partial path, not full-stack OTel support.

What you get What you don't
One-line proxy integration, among the fastest setups in this category Native evaluation and LLM-as-judge tooling is thin next to specialists
Sessions groups tool calls and vector queries into a full run view No published flat overage rate once you outgrow the free tier
Unlimited seats even on the $79/mo Pro tier On-prem deployment requires Enterprise

Pricing: Hobby free (10,000 requests/month, 1GB storage, 1 seat, 7-day retention). Pro $79/month (unlimited seats, usage-based overage on requests and storage, 1-month retention). Team $799/month (5 organizations, 3-month retention, SOC 2 and HIPAA compliance). Enterprise custom, with SAML SSO and on-prem deployment. Source: helicone.ai/pricing; the page doesn't publish a flat per-request overage rate.

Best for: Teams that want request-level cost and latency visibility live in minutes, with evaluation handled by a separate tool. Not ideal for: anyone trying to budget a specific monthly number past the free tier without first getting a quote.

8. HoneyHive: OpenTelemetry-Native With No Self-Serve Middle Tier

HoneyHive's technical foundation is genuinely strong: it's natively built on OpenTelemetry, which the company describes as making it "fully agnostic across models, frameworks, and clouds," with native support for LangChain, CrewAI, and several agent frameworks out of the box. Both evaluation types ship even on the free tier: code-based checks and LLM-as-judge scoring for automated evaluation, plus human annotation queues and offline experiments.

The catch is the pricing ladder itself: HoneyHive's public pricing page lists only a Free developer tier (10,000 events/month, up to 5 users, 30-day retention) and a custom-quoted Enterprise tier. There's no visible self-serve paid plan in between, so a team that outgrows the free allotment goes straight into a sales conversation rather than an upgrade button.

What you get What you don't
Natively built on OpenTelemetry, genuinely agnostic across frameworks No public self-serve tier between Free and Enterprise
Both automated and human evaluation on the free tier Growing past 10,000 events/month requires a sales call
Enterprise offers self-hosted, hybrid, or single-tenant deployment choices Overage rates for the free tier aren't published because there isn't an overage path

Pricing: Free/Developer $0/month (10,000 events/month, up to 5 users, 30-day retention, multi-tenant SaaS only). Enterprise custom (unlimited users, customizable retention, SAML/SSO, multi-tenant, single-tenant, hybrid, or self-hosted deployment). Source: honeyhive.ai/pricing.

Best for: Teams that want auto-instrumented OpenTelemetry tracing out of the box and are comfortable moving straight to a sales conversation once they outgrow the free tier. Not ideal for: a self-serve-only buying process; there isn't one past 10,000 events a month.

9. Portkey: Observability Bundled Into an AI Gateway

Portkey's core product is an AI gateway (routing, fallbacks, caching, and retries across model providers), and observability rides along with that gateway traffic rather than existing as a separate purchase. The docs describe the suite as OpenTelemetry-compliant with a dedicated OTel export feature, and there's a genuinely free, unlimited-request open-source self-hosted tier for teams that want the gateway without a vendor relationship at all. Guardrails exist for policy enforcement, but there's no dedicated evaluation or scoring product the way Braintrust or Galileo offer one.

What you get What you don't
Observability bundled with gateway routing, fallbacks, and caching in one bill No dedicated evaluation/scoring product; guardrails cover policy, not systematic scoring
A genuinely free, unlimited-request open-source self-hosted tier License for the self-hosted tier isn't named on the pricing page itself
OTel-compliant with a dedicated export feature, confirmed in the docs Production tier's log retention (30 days) is short next to Langfuse or Opik

Pricing: Open Source (self-hosted) free, unlimited requests. Developer free forever (10,000 recorded logs/month, 3-day log retention, 30-day metrics retention). Production $49/month (100,000 recorded logs/month, 30-day log/90-day metrics retention; overage $9 per additional 100,000 requests, up to 3M). Enterprise custom (10M+ requests, VPC hosting, SSO, SOC2/GDPR/HIPAA). Source: portkey.ai/pricing.

Best for: Teams that want an AI gateway and its observability in a single tool and a single bill. Not ideal for: teams that need deep, dedicated evaluation tooling; pair Portkey with a specialist if systematic scoring is a hard requirement.

10. Traceloop (OpenLLMetry): The OpenTelemetry Instrumentation Layer Itself

Traceloop built OpenLLMetry, an Apache 2.0-licensed, fully open-source instrumentation library that extends OpenTelemetry specifically for LLM and agent applications. Where most tools in this list added OTel support to an existing product, OpenLLMetry's semantic conventions for LLM calls were built alongside, and helped shape, the broader OTel GenAI semantic-convention effort the rest of the category is converging toward. The SDK itself works with any OTel-compatible backend, including several competitors on this list, which makes it the purest "instrumentation, not destination" option here.

What you get What you don't
OpenTelemetry-native from the ground up; the SDK works with any OTel backend Free tier data retention is only 24 hours
Apache 2.0-licensed OpenLLMetry SDK, genuinely free with no feature gate No visible self-serve tier between Free and Enterprise
On-prem deployment across AWS, GCP, Azure, and Kubernetes, including air-gapped setups Evaluation depth is thinner than dedicated eval-first tools

Pricing: OpenLLMetry SDK free and open source (Apache 2.0), usable with any OTel-compatible backend. Traceloop Free Forever $0/month (up to 50,000 spans/month, up to 5 seats, 24-hour retention). Enterprise custom (unlimited seats, custom retention, SOC 2, on-prem deployment). Source: traceloop.com/pricing.

Best for: Teams that want the OpenTelemetry instrumentation itself to be portable, viewing traces in Traceloop or pointing them anywhere else that speaks OTel. Not ideal for: a team that wants a self-serve paid tier between free and Enterprise; that gap doesn't exist here.

11. Comet Opik: Open Source Plus the Cheapest Real Paid Tier

Comet, an established ML experiment-tracking platform and a direct competitor to Weights and Biases, built Opik as its answer to LLM observability and open-sourced it under Apache 2.0. That gives Opik the same "free forever if you self-host" option as Langfuse, plus a hosted Free Cloud tier and a genuinely cheap Pro Cloud tier at $19/month, the lowest paid entry price in this comparison by a wide margin, and (per the cost table above) still the cheapest at a real 1M-span month.

OpenTelemetry support is native across every tier, including languages beyond Python and JavaScript, and LLM-as-judge evaluation, custom metrics, and test suites ship at every paid level, not gated to Enterprise.

What you get What you don't
$19/month Pro Cloud tier with full OTel support and LLM-as-judge included Smaller ecosystem and mindshare than LangSmith or Langfuse
Free, open-source self-hosted option in addition to the hosted Free Cloud tier Extended data retention beyond 60 days costs extra even on Pro
Cheapest option in this guide's 1M-unit cost comparison AI-powered debugging assistant is Enterprise-only

Pricing: Open Source free, self-hosted, Apache 2.0. Free Cloud $0/month (up to 10 team members, 25,000 spans/month, 60-day retention). Pro Cloud $19/month (up to 50 members, 100,000 spans/month, 60-day retention; overage $5 per 100,000 spans). Enterprise custom, flexible deployment including on-premises. Source: comet.com/site/pricing; license confirmed on comet-ml/opik's GitHub LICENSE file.

Best for: Budget-conscious teams that still want real OpenTelemetry support and LLM-as-judge evaluation rather than a stripped-down free tier. Not ideal for: teams that weight brand recognition or ecosystem size heavily in a vendor decision.

12. Galileo: Purpose-Built Judge Models for Enterprise Evaluation

Galileo positions itself as an observability, evaluation, and guardrail platform, and its evaluation approach is the clearest differentiator: rather than relying only on a general-purpose LLM prompted to act as a judge, Galileo built its own Luna family of smaller, purpose-tuned evaluator models aimed at faster, cheaper, more consistent scoring than a frontier-model judge call every time. Evaluations are unlimited even on the free tier, unusual generosity for a category where most competitors gate LLM-as-judge behind a paid plan. Enterprise adds real-time guardrails and dedicated low-latency inference, aimed at catching a bad output before it reaches a customer rather than logging it afterward.

What you get What you don't
Unlimited custom evaluations even on the free tier OpenTelemetry support isn't documented on pages this fetch could reach
Purpose-built Luna evaluator models instead of only general LLM-as-judge prompts Month-to-month Pro pricing isn't stated, only the annual-billed rate
Real-time guardrails and dedicated inference infrastructure at Enterprise Deployment flexibility (VPC, on-prem) is Enterprise-only

Pricing: Free $0/month (5,000 traces/month, unlimited users, unlimited custom evals). Pro $100/month billed annually, described as a 33% discount versus month-to-month (50,000 traces/month, standard RBAC, advanced analytics); overage "scales based on number of traces" with no published flat rate. Enterprise custom, hosted, VPC, or on-premises deployment. Source: galileo.ai/pricing.

Best for: Enterprise teams that want purpose-built judge models and real-time guardrails rather than a general LLM prompted to grade its own category. Not ideal for: a team that needs OpenTelemetry portability confirmed before buying; that information isn't public here.

Stage and Sizing Fit

Tool Best Team Size Typical Buyer
LangSmith Any size already on LangChain/LangGraph Eng lead already inside that framework
Langfuse Solo founder through enterprise (self-host scales both directions) Technical founder or platform team
Arize (Phoenix and AX) Mid-size through enterprise ML/AI platform team
Braintrust Small to mid-size, eval-focused Eng lead shipping prompt or model changes often
Datadog LLM Observability Mid-size through enterprise already on Datadog Platform or SRE team
W&B Weave (CoreWeave Forge) ML research teams of any size ML team already in W&B
Helicone Solo founder to small team Builder who wants cost/latency visibility in minutes
HoneyHive Small team comfortable with a sales call past free Eng lead auto-instrumenting an OTel stack
Portkey Solo to mid-size, gateway-first teams Eng lead who also wants routing and fallbacks
Traceloop (OpenLLMetry) Teams standardizing on portable instrumentation Platform engineer who wants to own the SDK layer
Comet Opik Budget-conscious small to mid-size teams Eng lead wanting a real paid tier cheap
Galileo Mid-size through enterprise, compliance-aware Eng leader needing purpose-built judge models

Buying Mistakes to Avoid

  • Buying tracing and calling it done. A clean dashboard with no failed run in months usually means nobody's grading correctness, not that nothing's wrong. Budget for evaluation from day one, not as a phase-two project.
  • Pricing off the free tier. The cost table above shows why: two vendors with attractive free tiers (Datadog, Arize AX) get expensive fast at real volume, while two others (Opik, Langfuse) stay cheap because their overage math is public and linear.
  • Assuming "OpenTelemetry support" means full portability. Check whether evaluation and datasets travel with the traces or stay locked to the vendor; ingestion-only OTel support is common and easy to miss.
  • Trusting a vendor's "open source" claim without opening the LICENSE file. Arize Phoenix is the clearest example here: it's source-available under the Elastic License 2.0, not Apache or MIT, and that distinction matters for anyone building on top of it commercially.
  • Not re-checking who owns the vendor. Langfuse (ClickHouse) and Arize (Dynatrace) both changed owners in 2026. Neither changed the product's standalone availability, per their own statements, but a procurement team should still ask the question at renewal.
  • Treating LLM-as-judge as ground truth. It scales further than human review alone, but a model grading another model's work still needs a human-reviewed rubric and periodic spot checks, especially before a release that actually matters.

How to Choose: Decision Framework

If you need... Pick... Why
Deepest tracing already on LangChain or LangGraph LangSmith Native run-tree tracing, plus a real OTLP endpoint if you need to move later
To self-host for free with no license surprises Langfuse MIT-licensed core, OpenTelemetry-native on every tier, now backed by ClickHouse
Unlimited evaluations on a free or cheap paid tier Arize Phoenix/AX or Galileo Both leave LLM-as-judge unlocked below Enterprise
Traces sitting next to the APM dashboards you already pay for Datadog LLM Observability Correlated with existing infrastructure, APM, and RUM data
Evaluation and CI-gated releases as the primary job Braintrust Built around scorers, experiments, and autonomous eval iteration
The cheapest real paid tier with full OpenTelemetry support Comet Opik $19/month Pro tier, cheapest at 1M-unit volume in this guide
A proxy-based setup live in minutes Helicone One-line swap, no re-instrumentation required
An AI gateway and observability bundled in one bill Portkey Routing, fallbacks, and OTel-compliant logs in a single product
The instrumentation layer itself, portable to any backend Traceloop (OpenLLMetry) Apache 2.0 SDK works with any OTel-compatible tool, competitors included
Regulated data that cannot leave your VPC Langfuse, Phoenix, or Comet Opik All three offer a genuine free self-hosted path, not just a sales conversation

Frequently Asked Questions

What to Do Next

Pick one LLM application you already have running, ideally the one with the highest traffic or the most customer exposure, and instrument it with your shortlisted tool's free tier for two weeks before signing anything. Confirm you can actually see the full prompt and completion for a failed run, not just a success/failure flag, and run at least one offline evaluation against a small set of real cases before you trust the dashboard. Then price it at your real monthly volume using the cost table above, not the free-tier number, since that's where several of these vendors' actual cost picture changes completely. If that application is a RAG pipeline, the RAG tools roundup covers the retrieval stack those traces will be recording.

If you're specifically instrumenting a multi-step agent rather than a single-call or RAG-style app, read our dedicated agent observability roundup next; the buying questions around tool-call tracing and evaluating multi-step behavior are more specific than what a general LLM observability tool is built to answer. And if cost is the open question once you've picked a tool, AI agent cost optimization covers the token-spend side of the same budget this guide's pricing tables are built to inform.

About the author

Camellia

Camellia

Principal Product Marketing Strategist

Camellia is Principal Product Marketing Strategist at Rework, helping B2B buyers pick the right software with confidence. With 6+ years in product marketing and 150+ SaaS tools evaluated across CRM, project management, and sales engagement, Camellia turns competitive intelligence into clear, honest comparisons. Readers get vendor evaluations they can trust to cut through marketing noise and decide faster.