LLM Observability with Honeycomb - Bring Your Own LangChain LLM
Your evals live in LangSmith. Your traces live in Honeycomb. Nobody should have to re-instrument their app to put those two facts in the same place, so I spent a week finding out what it actually takes.
Every team shipping on LLMs is already grading its output. A judge model scores answers, a reviewer works an annotation queue, a regex checks a schema, a user clicks thumbs down. Those grades pile up in LangSmith or Braintrust or, let's be honest, a spreadsheet. And not one of them is where you go at 2am when the thing is misbehaving.
The industry's answer to this is to tell customers to re-instrument against the GenAI semantic conventions first, and then they get value. That's backwards. It puts the entire cost up front and the payoff somewhere over the horizon, which is exactly why these integrations stall at the proof-of-concept stage and quietly die.
So I tried it the other way round: take the grades in whatever shape they already have, and add only the two things that make them usable. Which operation the grade is about, and what the score means. Here's what broke, what fixed it, and the two problems I couldn't solve.
Three words do a lot of work below.
- Operation
- One run of the thing you're watching. Here, one agent answering one question. Its record is a span.
- Grade
- One opinion about that answer: a name (
correctness), a score or label (true), and usually an explanation. The grader might be a judge model, a person, a code check or an annoyed user. LangSmith stores all four and calls them feedback. - Late
- The grade shows up after the operation's span has ended and shipped, so it can't be added to that span. Nothing to do with being slow.
Which operation is this grade about?
Solved, and it's the piece I'd keep
The two systems share no identifier. LangSmith has a run UUID, Honeycomb has a trace and span ID, and absolutely nothing connects them. My first working version had me pasting IDs in by hand. Great for a demo, useless for a customer.
The fix is much smaller than the problem sounds. Ask the eval tool for a run ID before the
run starts, and set it on the operation span as an attribute when the span opens. Now both systems
agree on one value, and one filter gets you the operation and every grade about it. If you're already
exporting LangSmith traces over OpenTelemetry you don't even need that much: metadata you set on a
run turns up in Honeycomb as langsmith.metadata.<key>, and the key rides along for
free.
What I checked
Stamping byoe.langsmith.run_id works, offline and in Honeycomb. The metadata path is
offline-tested only. Child runs inherit the metadata, so a query wants the root span.
Here's what you must not do. LangSmith derives an eight-byte span ID from the front of its run UUID, and those UUIDs are time-ordered. Five thousand runs created back to back gave me four distinct IDs. Four. Build your join on that and it will collide silently, in production, on the traffic that matters most.
A span that has ended can't be told anything
Solved, and the spec agrees
Everyone's first instinct is to hang the grade on the operation it judges. You can't. Once a span has ended and shipped, the API won't let you touch it and nothing downstream is going to merge it for you. That's not a Honeycomb limitation, that's how tracing works, and it's the single fact that shapes this entire problem.
So the grade becomes its own record: a span link pointing at the operation, and the operation's trace and span IDs copied onto it as plain attributes. The link is for humans navigating. The attributes are what you query, because a link arrives as its own row with none of the grade's columns on it.
What I checked
OpenTelemetry defines span events as points during a span's own duration, and links for asynchronous follow-on work. This is the documented pattern, not a hack I invented to get around something.
One thing I'd watch: the span events API is deprecated in favour of log-based events, and a log record can carry a finished span's trace context. That's where I'd build next. Whether Honeycomb renders such a record next to the span is the very first thing I'd test, because I don't know.
Your field names are fine. Teach them to the agent.
Solved with the product as it ships today
Here's the part that changed how I think about this whole problem. The pitch for adopting shared conventions is that otherwise nothing downstream can understand your data. That was true right up until the thing reading your data became an agent that can be handed a glossary.
Honeycomb's team telemetry schema does exactly that, and it's been sitting there the whole time.
Paste a Weaver-format YAML describing your own attributes into Team Settings, and the MCP tools hand
those descriptions to whatever agent is investigating. No data moves. No names change. An agent
reading the dataset now knows that byoe.target_mapping=user-supplied means nobody
verified it, which is the kind of thing that never survives a migration into a shared schema anyway.
What I checked
I pasted this project's registry. Thirteen byoe.* attributes and the
eval.* group come back from the registry as team_custom, right next to
OpenTelemetry's own, which already carry gen_ai.evaluation.name,
score.value, score.label and explanation at semconv 1.41.0.
It fixes meaning, not shape. Comparing one customer's grades against another's still needs a real mapping, and I'm not going to pretend otherwise.
One column asked to mean two things
Found the hard way, and still live
My judge returned a boolean. LangSmith stored it and handed it back as 1.0. That number
went to Honeycomb as a double, into a column typed boolean, and came back as true. Every
layer behaved exactly as designed.
On this record nothing was lost. The judge really had answered true, so the boolean
column happened to restore the right meaning. That is the uncomfortable part. Once they land,
1.0 and true are indistinguishable, so a faithful row and a flattened one
look identical. A score of 0.5 is the one that would not survive, and it has never been
sent.
This is a semantic bug wearing a formatting bug's clothes, and it's the strongest argument I have
for declaring what a score means instead of assuming it. The fix is written and not yet emitted: the
score goes in a field per type with the source type recorded beside it, numbers to
score.value and categories to score.label the way the convention defines
them. The original payload stays untouched either way, which is the only reason this was catchable
at all.
What I checked
Checked 17 September 2026, environment test, dataset
byoe-langchain-eval. The judge's own run record holds "score": true for
target span ba68aa1035802e37. The imported audit copy of that same judgement holds
"score": 1.0. The importer passes the value through on a
type(value) in (str, bool, int, float) test with no cast, so a double is what leaves
the process. eval.result.score is typed boolean, no eval.result.value
column exists yet, and status_code is 0 on all five rows, so nothing here
came from an error.
Not checked: whether Honeycomb coerces at ingest or types the column from the data at read. Both
explain what I see, and separating them needs a 0.5 written on purpose.
Why this matters past one bug
The convention itself warns that a score of 1 means "relevant" in one system and "not relevant" in another. A number with no declared direction isn't a measurement, it's a rumour. A shared column quietly makes it worse.
What did the failures have in common?
Not solved, and it's the question that actually matters
This is what anyone actually wants. Bad grades in one group, good ones in another, BubbleUp telling you what separates them. My design can't answer it, and I'd rather say that plainly than demo around it.
BubbleUp compares fields on the events you selected. Relational fields reach across one trace, not two. My grades sit in their own trace carrying nothing but a score, so there's nothing to compare. The fix is to copy an allowlist of the operation's context onto the grade as it's written: model, prompt version, retrieved document count, cost. That's a snapshot, and snapshots drift. I'd rather make that trade deliberately than discover it during an incident.
Why it's still unproven
Every grade in my dataset passed. All of them. With nothing failing there's nothing to contrast, and any demo I built on it would be theatre. It needs a batch with real failures in it.
A practical trap while you're here: link rows and span events live in the same dataset, so any
count of grades wants meta.annotation_type does-not-exist or your baseline fills up
with rows that aren't grades.
What did a correct answer cost?
Half solved
Cost per correct answer is the number that makes any of this worth wiring into observability. It breaks on something far dumber than arithmetic: the model you asked for is not the model that answered.
Across 289 calls my logs asked for moonshotai/Kimi-K3 and were served
accounts/fireworks/models/kimi-k3. Only the served name has a price. My price table
listed only the requested one, so every experiment had been quietly reporting cost as unknown and I
hadn't noticed. Price the served model, freeze the number on the record, and write down which price
table and version produced it.
What I checked
The token counts are real, read from each response's own usage metadata across 60 calls with no errors. The money is not: $0.73 is those counts against Fireworks' published price, not a number anyone billed me. Passing a deliberately wrong provider returned no price at all rather than the wrong price, which is the behaviour you want.
That gap is the whole point. A computed cost has to carry its price table and version, and the customer's own cost field beats it every time, because published prices don't know what they negotiated.
The same grade, twice, with different IDs
Designed, not proven
One judgement reached Honeycomb twice: once from the process that ran the judge, once imported from LangSmith. A naive count says two grades. Worse, the two copies don't even agree on the score's type.
Identifiers aren't stable either. LangSmith's migration tool re-mints run IDs, issues new feedback IDs, shifts experiment timestamps, and doesn't migrate feedback attached to production traces at all. So deduplicating on the feedback ID works right up until a customer moves instances, which is a terrible time to find out.
What I'd do
Identify a judgement by run, evaluation name and rationale rather than by its score, and keep the source instance and lineage on the mapping record. A duplicate delivery is then one grade, and a migrated record is still the same grade.
How would you know any of this is trustworthy?
The habit, more than the code
Integration write-ups love to claim the top rung of a ladder they climbed two steps of. There are five separate claims hiding in "it works": the record was exported, Honeycomb received it, you can find the operation from the grade, you can navigate between them, and the grade shows up on the operation itself. Each one needs its own evidence.
A flush returning true proves the first and nothing else. I have the first three, for one target, by querying the records back by ID. Four and five are unchecked. There are a lot of unknowns here, and I'd rather hand you the list than let a screenshot imply I've closed it.
Checked
One filter on the operation's span ID returns both grades and nothing else. The attribute
descriptions resolve as team_custom. The offline suite is 59 tests, no network.
Not checked
Clicking through in the UI, whether a late grade renders on the operation, and whether a single customer actually wants this.
What I deliberately didn't build
- A rule that customers adopt the GenAI conventions before Honeycomb is useful to them.
- A translation that drops or overwrites the fields they sent.
- A failing grade recorded as a span error. The judge worked. The answer was wrong. Those are different events and conflating them poisons your error rate.
- Field mappings decided by embedding similarity. Opposites sit close together in that space, and the distances move when the model changes.
- An LLM anywhere in the write path. Queries produce the numbers. The model explains them.
- A knowledge graph, until lineage questions are a real workload instead of a possibility.
- A conformance percentage for any library. There's no honest denominator, and it turns "not observed" into "not done".
The four experiments I'd run next
Every one of them is small, and every one can come back negative.
- Open an existing trace and look. Does the operation show anything about the grades pointing at it, and can you click from a grade to its target? Costs nothing and settles the navigation question today.
- Send one log-based
gen_ai.evaluation.resultrecord carrying a finished span's trace context and see whether Honeycomb puts it beside that span. Two spans, one query. - Seed a batch with planted failures and operation context, then run BubbleUp. Pass criteria written first: the planted factor ranked top, rates within five points, and a correct refusal when there's nothing to compare.
- Ask three teams who already grade their LLM output whether they start from one bad trace or from a chart of scores. That answer changes what's worth building, and no amount of code substitutes for it.
None of this is exotic. It's the same discipline the rest of observability already runs on: keep the customer's data in the customer's shape, say what you actually verified, and put the meaning somewhere a person or an agent can read it later. Evals are just another kind of telemetry that shows up late and means something specific. We should treat them that way.