LLM Observability with Honeycomb - Bring Your Own LangChain LLM

Your evals live in LangSmith. Your traces live in Honeycomb. Nobody should have to re-instrument their app to put those two facts in the same place, so I spent a week finding out what it actually takes.

Grace 17 September 2026 Evals OpenTelemetry

Every team shipping on LLMs is already grading its output. A judge model scores answers, a reviewer works an annotation queue, a regex checks a schema, a user clicks thumbs down. Those grades pile up in LangSmith or Braintrust or, let's be honest, a spreadsheet. And not one of them is where you go at 2am when the thing is misbehaving.

The industry's answer to this is to tell customers to re-instrument against the GenAI semantic conventions first, and then they get value. That's backwards. It puts the entire cost up front and the payoff somewhere over the horizon, which is exactly why these integrations stall at the proof-of-concept stage and quietly die.

So I tried it the other way round: take the grades in whatever shape they already have, and add only the two things that make them usable. Which operation the grade is about, and what the score means. Here's what broke, what fixed it, and the two problems I couldn't solve.

Three words do a lot of work below.

Operation
One run of the thing you're watching. Here, one agent answering one question. Its record is a span.
Grade
One opinion about that answer: a name (correctness), a score or label (true), and usually an explanation. The grader might be a judge model, a person, a code check or an annoyed user. LangSmith stores all four and calls them feedback.
Late
The grade shows up after the operation's span has ended and shipped, so it can't be added to that span. Nothing to do with being slow.

Which operation is this grade about?

Solved, and it's the piece I'd keep

The two systems share no identifier. LangSmith has a run UUID, Honeycomb has a trace and span ID, and absolutely nothing connects them. My first working version had me pasting IDs in by hand. Great for a demo, useless for a customer.

The fix is much smaller than the problem sounds. Ask the eval tool for a run ID before the run starts, and set it on the operation span as an attribute when the span opens. Now both systems agree on one value, and one filter gets you the operation and every grade about it. If you're already exporting LangSmith traces over OpenTelemetry you don't even need that much: metadata you set on a run turns up in Honeycomb as langsmith.metadata.<key>, and the key rides along for free.

What I checked

Stamping byoe.langsmith.run_id works, offline and in Honeycomb. The metadata path is offline-tested only. Child runs inherit the metadata, so a query wants the root span.

Here's what you must not do. LangSmith derives an eight-byte span ID from the front of its run UUID, and those UUIDs are time-ordered. Five thousand runs created back to back gave me four distinct IDs. Four. Build your join on that and it will collide silently, in production, on the traffic that matters most.

A span that has ended can't be told anything

Solved, and the spec agrees

Everyone's first instinct is to hang the grade on the operation it judges. You can't. Once a span has ended and shipped, the API won't let you touch it and nothing downstream is going to merge it for you. That's not a Honeycomb limitation, that's how tracing works, and it's the single fact that shapes this entire problem.

So the grade becomes its own record: a span link pointing at the operation, and the operation's trace and span IDs copied onto it as plain attributes. The link is for humans navigating. The attributes are what you query, because a link arrives as its own row with none of the grade's columns on it.

What I checked

OpenTelemetry defines span events as points during a span's own duration, and links for asynchronous follow-on work. This is the documented pattern, not a hack I invented to get around something.

One thing I'd watch: the span events API is deprecated in favour of log-based events, and a log record can carry a finished span's trace context. That's where I'd build next. Whether Honeycomb renders such a record next to the span is the very first thing I'd test, because I don't know.

Your field names are fine. Teach them to the agent.

Solved with the product as it ships today

Here's the part that changed how I think about this whole problem. The pitch for adopting shared conventions is that otherwise nothing downstream can understand your data. That was true right up until the thing reading your data became an agent that can be handed a glossary.

Honeycomb's team telemetry schema does exactly that, and it's been sitting there the whole time. Paste a Weaver-format YAML describing your own attributes into Team Settings, and the MCP tools hand those descriptions to whatever agent is investigating. No data moves. No names change. An agent reading the dataset now knows that byoe.target_mapping=user-supplied means nobody verified it, which is the kind of thing that never survives a migration into a shared schema anyway.

What I checked

I pasted this project's registry. Thirteen byoe.* attributes and the eval.* group come back from the registry as team_custom, right next to OpenTelemetry's own, which already carry gen_ai.evaluation.name, score.value, score.label and explanation at semconv 1.41.0.

It fixes meaning, not shape. Comparing one customer's grades against another's still needs a real mapping, and I'm not going to pretend otherwise.

One column asked to mean two things

Found the hard way, and still live

My judge returned a boolean. LangSmith stored it and handed it back as 1.0. That number went to Honeycomb as a double, into a column typed boolean, and came back as true. Every layer behaved exactly as designed.

On this record nothing was lost. The judge really had answered true, so the boolean column happened to restore the right meaning. That is the uncomfortable part. Once they land, 1.0 and true are indistinguishable, so a faithful row and a flattened one look identical. A score of 0.5 is the one that would not survive, and it has never been sent.

This is a semantic bug wearing a formatting bug's clothes, and it's the strongest argument I have for declaring what a score means instead of assuming it. The fix is written and not yet emitted: the score goes in a field per type with the source type recorded beside it, numbers to score.value and categories to score.label the way the convention defines them. The original payload stays untouched either way, which is the only reason this was catchable at all.

What I checked

Checked 17 September 2026, environment test, dataset byoe-langchain-eval. The judge's own run record holds "score": true for target span ba68aa1035802e37. The imported audit copy of that same judgement holds "score": 1.0. The importer passes the value through on a type(value) in (str, bool, int, float) test with no cast, so a double is what leaves the process. eval.result.score is typed boolean, no eval.result.value column exists yet, and status_code is 0 on all five rows, so nothing here came from an error.

Not checked: whether Honeycomb coerces at ingest or types the column from the data at read. Both explain what I see, and separating them needs a 0.5 written on purpose.

Why this matters past one bug

The convention itself warns that a score of 1 means "relevant" in one system and "not relevant" in another. A number with no declared direction isn't a measurement, it's a rumour. A shared column quietly makes it worse.

What did the failures have in common?

Not solved, and it's the question that actually matters

This is what anyone actually wants. Bad grades in one group, good ones in another, BubbleUp telling you what separates them. My design can't answer it, and I'd rather say that plainly than demo around it.

BubbleUp compares fields on the events you selected. Relational fields reach across one trace, not two. My grades sit in their own trace carrying nothing but a score, so there's nothing to compare. The fix is to copy an allowlist of the operation's context onto the grade as it's written: model, prompt version, retrieved document count, cost. That's a snapshot, and snapshots drift. I'd rather make that trade deliberately than discover it during an incident.

Why it's still unproven

Every grade in my dataset passed. All of them. With nothing failing there's nothing to contrast, and any demo I built on it would be theatre. It needs a batch with real failures in it.

A practical trap while you're here: link rows and span events live in the same dataset, so any count of grades wants meta.annotation_type does-not-exist or your baseline fills up with rows that aren't grades.

What did a correct answer cost?

Half solved

Cost per correct answer is the number that makes any of this worth wiring into observability. It breaks on something far dumber than arithmetic: the model you asked for is not the model that answered.

Across 289 calls my logs asked for moonshotai/Kimi-K3 and were served accounts/fireworks/models/kimi-k3. Only the served name has a price. My price table listed only the requested one, so every experiment had been quietly reporting cost as unknown and I hadn't noticed. Price the served model, freeze the number on the record, and write down which price table and version produced it.

What I checked

The token counts are real, read from each response's own usage metadata across 60 calls with no errors. The money is not: $0.73 is those counts against Fireworks' published price, not a number anyone billed me. Passing a deliberately wrong provider returned no price at all rather than the wrong price, which is the behaviour you want.

That gap is the whole point. A computed cost has to carry its price table and version, and the customer's own cost field beats it every time, because published prices don't know what they negotiated.

The same grade, twice, with different IDs

Designed, not proven

One judgement reached Honeycomb twice: once from the process that ran the judge, once imported from LangSmith. A naive count says two grades. Worse, the two copies don't even agree on the score's type.

Identifiers aren't stable either. LangSmith's migration tool re-mints run IDs, issues new feedback IDs, shifts experiment timestamps, and doesn't migrate feedback attached to production traces at all. So deduplicating on the feedback ID works right up until a customer moves instances, which is a terrible time to find out.

What I'd do

Identify a judgement by run, evaluation name and rationale rather than by its score, and keep the source instance and lineage on the mapping record. A duplicate delivery is then one grade, and a migrated record is still the same grade.

How would you know any of this is trustworthy?

The habit, more than the code

Integration write-ups love to claim the top rung of a ladder they climbed two steps of. There are five separate claims hiding in "it works": the record was exported, Honeycomb received it, you can find the operation from the grade, you can navigate between them, and the grade shows up on the operation itself. Each one needs its own evidence.

A flush returning true proves the first and nothing else. I have the first three, for one target, by querying the records back by ID. Four and five are unchecked. There are a lot of unknowns here, and I'd rather hand you the list than let a screenshot imply I've closed it.

Checked

One filter on the operation's span ID returns both grades and nothing else. The attribute descriptions resolve as team_custom. The offline suite is 59 tests, no network.

Not checked

Clicking through in the UI, whether a late grade renders on the operation, and whether a single customer actually wants this.

What I deliberately didn't build

The four experiments I'd run next

Every one of them is small, and every one can come back negative.

  1. Open an existing trace and look. Does the operation show anything about the grades pointing at it, and can you click from a grade to its target? Costs nothing and settles the navigation question today.
  2. Send one log-based gen_ai.evaluation.result record carrying a finished span's trace context and see whether Honeycomb puts it beside that span. Two spans, one query.
  3. Seed a batch with planted failures and operation context, then run BubbleUp. Pass criteria written first: the planted factor ranked top, rates within five points, and a correct refusal when there's nothing to compare.
  4. Ask three teams who already grade their LLM output whether they start from one bad trace or from a chart of scores. That answer changes what's worth building, and no amount of code substitutes for it.

None of this is exotic. It's the same discipline the rest of observability already runs on: keep the customer's data in the customer's shape, say what you actually verified, and put the meaning somewhere a person or an agent can read it later. Evals are just another kind of telemetry that shows up late and means something specific. We should treat them that way.