Somewhere in your trace store right now, there's a pattern. A specific tool call that fails for a specific kind of request. A retry loop that eats your token budget before quietly giving up. A path your agent keeps taking that never looks like a problem in any single trace, because no single trace is the problem.
You have the data. You just can't see the shape of it.
We ran into this ourselves building our own agent harness. Any single trace, we could open up and read step by step, exactly what happened, exactly where it broke. What we couldn't do was step back and see it across a thousand traces at once, notice the same failure taking the same shape over and over. We saw the tree, but missed the forest, the pattern in the crowd, not the one conversation.
That's the gap we built Kubit's agent observability tools to close, and today we're opening early access to the first three pieces: Trajectory Flows, Trigger Insights, and Polaris Evaluator, a small language model we built in-house for labeling every span.
The problem with one trace at a time
Most observability tools are genuinely good at showing you a single trace end to end. Request came in, here's every span, here's where it broke. That part isn't the problem.
The problem is a single trace can't tell you whether a failure is a fluke or the start of a trend. It can't tell you that a dozen failed checkouts all share a payment method, or a retry pattern, or a slow upstream step. It can't tell you which commit caused a regression unless you already suspect one and go looking for it. By the time an aggregate error rate moves enough for anyone to notice, the pattern's usually been running for days.
Trajectory Flows: seeing the shape, not just the trace
Trajectory Flows clusters your agents' execution paths, every tool call, every retrieval, every reasoning step, into the patterns that actually predict failure. Instead of a wall of individual traces, you get a picture of where runs diverge, where they loop, where they stall, and how often each pattern resolves versus drags on.
A pattern with a thousand instances behind it is a different kind of problem than one bad conversation. It should look like one. A retry loop quietly burning through your token budget should be visible the moment it starts climbing, not three weeks later in a support escalation.
Trigger Insights: getting to why
Finding the pattern is half the job. Knowing what caused it is the other half, and that's usually where the tooling stops. We didn't want to stop there.
When a cluster starts climbing, Trigger Insights checks four different kinds of evidence. It looks for structural signals first: is some attribute, a payment method, a category, a customer segment, statistically overrepresented in the failure compared to your overall traffic. If that comes back empty, it clusters the actual text of the failures, since two errors with the same error code can mean completely different things once you read what they say. It checks whether something earlier in the same trace predicted the outcome, because a slow retrieval step three hops upstream is a real signal that's invisible if you only look at the span that failed. And when the cause is a code change, it traces the cluster straight to the commit: the diff, the message, and the jump in error rate right after it shipped.
Every one of these comes with a lift ratio and a significance test attached. A cluster that shares an attribute with 500,000 successful traces isn't a finding, it's noise, and we'd rather show nothing than dress up a coincidence as an insight.
Polaris Evaluator: the part that makes this affordable
Running an LLM judge over every span you produce is expensive, and most vendors solve that by sampling: score 5%, maybe 10%, and hope the rest looks similar. It usually doesn't.
Polaris Evaluator is a small language model we built and tuned in-house specifically to avoid that trade-off. It's fast and cheap enough to run on every span you produce, not a sample, with sub-second performance and no per-token bill stacking up behind it: tool_error, loop_detected, escalated_to_human, hallucination, faithfulness, tone, and tool-selection correctness, alongside intent, sentiment, resolution, and persona.
It's not a fixed rulebook either. Label a handful of your own production spans and Polaris learns your team's actual definition of resolved, on-brand, or in-scope, since what counts as a good outcome for a support bot isn't what counts as one for a coding agent. Already running custom scorers in Databricks MLflow? Polaris uses those too. We're not asking anyone to rebuild what already works.
Where it runs
If your traces already land in Unity Catalog through MLflow Tracing, this works today, nothing new to install. Same story on Snowflake, BigQuery, or ClickHouse: bring your own cloud, nothing leaves your account. Not there yet? OpenTelemetry gets you streaming in minutes.
Design partners, not a waitlist
We're working with a small number of design partners running agents at enterprise scale, on real production traffic, to prove this out before we open it up further. If you're running agents in production and you've ever had the trace without the pattern, we'd like to talk to you.
Request early access →
Table of contents
Running agents in production?
We are opening a first cohort of 7 teams. Your traces stay in your warehouse.
Summarize this post
