Don’t invent a language for this
The obvious next step after normalizing LLM telemetry is a config format for the mappings, and a standard to submit it to. I tried the format. It was the wrong abstraction. The standard is worth doing — for something other than the mappings.
Last post I argued that GenAI instrumentation libraries each spell the same facts differently, that you should normalize them at the pipeline, and that whatever your normalizer cannot carry should be written onto the span rather than dropped in silence.
Two questions came back, and they are the right two. Should the mappings become a format — JSON, a small DSL, something SPL-shaped — instead of code? And should any of this be standardized, by somebody other than me?
No, and yes. The no is more interesting than it sounds: I built the format before deciding against it, and what killed it was not taste. And the yes turns out not to be about the mappings at all, which is where this ends up and the part I would keep if you only read one section.
Three layers that keep getting conflated
Normalizing telemetry is three separable things, and almost every conversation about it slides between them mid-sentence.
The rules. llm.token_count.prompt and gen_ai.usage.input_tokens are the
same fact. That is data. It is a table.
The execution. Something has to actually rewrite the span, in a pipeline, at volume, without dropping anything.
The claim. An assertion that the two names are the same fact, that translating between them costs nothing — or costs exactly this much — and that the assertion is auditable by somebody who was not there when it was made.
The industry has quietly settled the second layer several times over, privately,
inside each vendor’s ingest. It has never published the third. Arize AX rewrites
inbound gen_ai.* attributes into OpenInference before your spans are stored.
That is a defensible product decision and I am not picking on them; the point is
that you cannot read the mapping, cannot version it, and cannot find out what it
discarded. Every backend is doing some version of this and none of them will
show you the table.
So: which layer would a new language be for?
The rules should be data. That is not the same as a new format.
I went and did this to my own code, which is the only honest way to find out what it costs.
genai-interlingua had two halves that both looked like tables and only one of
which was. The schema half — which attribute name each convention version uses,
which values it admits — is now generated from the registries OpenTelemetry
publishes, rather than transcribed by me. That was straightforwardly correct and
I should have done it first. The derivation turned out to be a rule with one
exception: the attribute key is gen_ai. plus the field name, except for the
token-cache field that v1.41.0 spells cache_creation and the new repository
spells cache_write. Fifty keys and seventy-two keys, reproduced exactly. The
hand-typed tables had been right, which I would not have bet on.
The mapping half is where it gets interesting. I converted five of the six dialects from Go calls into declared tables — 86 mappings state cleanly as data against 38 that do not, so roughly seven in ten. Stating one looks like: read this key, lowercase it, cut it at the first dot, run it through this lookup, write it there.
The ratio is an average over emitters that are nothing like each other, which matters more than the average. LiteLLM is 20 against 4. OpenInference is 16 against 16, because it packs every sampling parameter — temperature, top_p, the penalties, the seed, nine others — into one JSON string, and no table reaches inside a string. The sixth dialect is not migratable at all: it is the fallback for spans nobody designed, and it works by resolving each attribute against the entire semantic-convention registry at runtime, so writing it as data would either duplicate the registry or admit it carries nothing.
The three in ten that do not state are not waiting for a better format.
Consider reassembling gen_ai.prompt.0.tool_calls.1.arguments — two levels of
indexing — into one nested JSON document. Or deciding what traceloop.entity.name
means by looking at whether a sibling attribute says this span is an agent or a
tool. Or refusing to map a Braintrust span carrying three evaluation scores,
because the conventions model one evaluation per span and picking one silently
would be worse than admitting the mismatch.
Those are not transformations of a value. They are readings of a span — they need the whole span in hand, not one attribute, because what they produce depends on what else is there. A format expressive enough to state them has conditionals, loops and a JSON parser, at which point you have written a programming language, and a worse one than the several already available.
And one of those already exists
OTTL — the OpenTelemetry Transformation Language — is a DSL for exactly this, governed by OpenTelemetry, shipped in every Collector distribution, running in production at a scale mine never will. Proposing a new config format for telemetry rewriting in 2026 means proposing to compete with it. That is a bad trade for a mapping table.
The better move is the opposite one: emit OTTL. So genai-interlingua now does.
$ interlingua -emit ottl -dialect litellm -target v1.41.0
Out comes a transform processor you paste into your own Collector. No custom
build, no Go, no dependency on my repository continuing to exist.
A stock Collector accepts it; I checked, in both target schemas, because a generated config that has only ever been compared against a golden file has not been tested, it has been photographed.
Then I checked the thing that actually matters, which is whether it means the same thing. Every captured span goes through the Go processor and through a real Collector running the exported config, and the two outputs are compared attribute by attribute. It found four bugs on its first run, and every one of them had been sitting behind a config that parsed cleanly and looked right:
- A scalar-to-list conversion that wrapped unconditionally, where the Go code wraps only a string. On the one emitter whose value is already an array, the attribute came out as a list containing a list.
- A lookup table that never fired on array values at all. Go maps over the elements; OTTL has no iteration, so the generated statements compare the whole array against each table key and match nothing. That one is not fixable — it is now declared instead, which is the point.
- An enum guard that was more destructive than the processor, deleting a non-conformant value the processor deliberately preserves.
- And a header claiming coverage it did not have.
The third and fourth are the ones I would have shipped. An export that quietly does less than it advertises is the exact failure this post is about, and I wrote it into my own tool while writing the argument against it.
And because the export is a subset, it says so in its own header:
# WHAT THIS CARRIES (20 attributes)
# ...
# WHAT THIS CANNOT CARRY (4 attributes)
#
# gen_ai.input.messages
# gen_ai.output.messages
# gen_ai.response.finish_reasons
# gen_ai.tool.definitions
Three things do not survive the trip. The unstated mappings, above. Per-span
loss accounting, because interlingua.lossy is computed from what a given span
turned out to carry and a static config can only describe itself. And detection:
picking a dialect means scoring every candidate across the whole span and taking
the maximum, and OTTL has neither a loop nor an argmax. An exported config is
pinned to one library and scoped by the spellings only that library uses.
That last one is a genuine limit, not a gap to fill later. It is also the smallest of the three.
The rule I would give anyone else: if your mapping can be stated as data, state it as data, then hand the data to a language somebody else maintains. Keep code for the rules that are readings rather than transformations, and be loud about which is which. My compiler cannot tell those apart, so a test does — it runs every captured span through the parser and fails the build if any field comes out that is neither declared as a rule nor admitted as unstateable. Without it, the next mapping I add as a method quietly disappears from the export while the header still claims completeness.
So what would you standardize?
Not the mapping table. That is data with a shelf life. OpenInference
instrumentations already dual-emit gen_ai.*; the conventions are stabilizing;
in two years the dialect tables are archaeology. A standards body should not be
asked to ratify a compatibility shim for a problem that is actively dissolving.
The durable thing is the third layer — the claim — and it is genuinely missing. Here is what exists today and what each artifact can say:
| can express | cannot express | |
|---|---|---|
| Telemetry Schema File 1.1.0 | rename_attributes, rename_events, rename_metrics, metrics-only split |
value rewriting, unit or type change, flattening, translation between vendors |
Weaver schema-changes |
added / renamed / obsoleted, within one registry’s own lineage | value-level transforms, cross-registry mappings |
| Weaver registries | third-party registries, dependencies, layering | equivalence claims between sibling registries |
| OTTL | all of it, imperatively | a portable, reviewable claim about equivalence and fidelity |
The gap is the same shape in every row. You can rename an attribute. You cannot
say that bedrock and aws.bedrock are the same provider, that milliseconds
and seconds are the same duration, or — the one I care about most — that a
particular translation was not faithful, and in exactly which respect.
I stopped asserting that and measured it. genai-interlingua now emits schema
files too, and the resulting gap table is generated from the rule tables
rather than written by me. The blunt version: for LiteLLM at v1.41.0, a schema
file expresses zero of the mappings. At genai-main it manages one. Not
because the format is bad — because every transformation it has changes a
name, and the two mappings LiteLLM actually needs change a value. Eighteen
of the rest need no rename at all, since LiteLLM already writes the conventions’
own names, and counting those as coverage would credit the format for work
nobody did.
That measurement also cost me half my argument, which is the useful part. Most of what a schema file cannot carry, it cannot carry because the work is a reading of a span — and no declarative format should try to express that. Adding value transforms would not make schema files sufficient for normalizing GenAI telemetry. It would make them able to describe migrations the conventions have already made, which is a narrower claim and the only one the evidence supports.
Then it cost me the correction too, which is the more useful part. I migrated a second dialect and the ratio inverted: the gaps a format change would fix went from a minority to a clear majority, and I rewrote the conclusion to say the fixable share grows as you sample more emitters. Then I migrated three more and it went back.
| sample | fixable by a format change | not expressible in any format | fixable share |
|---|---|---|---|
| 1 dialect | 4 | 8 | 33% |
| 2 dialects | 24 | 16 | 60% |
| 5 dialects | 44 | 76 | 37% |
The middle row was one unusually-shaped emitter at n=2. The Vercel AI SDK writes
the same fact under an ai.* name on the span your code creates and a gen_ai.*
name on the span its provider adapter creates, so nearly every mapping it has
carries two spellings — and twelve of those twenty-four fixable gaps were that
one property of that one library. My correction was worse than the thing it
corrected.
So the first answer was right, and I only know that because the table regenerates. If I had written “a minority” as prose in September I would have had no way to discover in October that I had briefly and confidently believed otherwise. Generate the number that your argument turns on. Not because generated numbers are true — these moved three times — but because a number that moves in front of you is one you can still be wrong about out loud.
Two things are worth writing down:
Cross-registry equivalence with value transforms. Not just “this key became that key” but the lookup, the unit, the coercion.
A fidelity vocabulary. A way to record, on the telemetry itself, which schema it was translated from, which it was translated to, what could not be carried, and — when the source convention was inferred rather than declared — how confident that inference was. I have been emitting a private version of this for months. It has no business being private. It generalizes far past GenAI: every semantic-convention migration, every vendor’s ingest pipeline, the whole HTTP attribute rename wave, has the same unanswerable question sitting under it.
And the ground is explicitly reserved. OTEP 0152, which introduced telemetry schemas in the first place, says the transformation set was held to “the bare minimum … with more types of transformations potentially proposed in the future,” and that changing the file format version must go through the OTEP process. That is not a wall. That is a door with a sign on it.
Who you submit it to
OpenTelemetry. Three different places, depending on which piece:
- The GenAI mapping work, and the missing schema URL, go to
semantic-conventions-genai— whose README still readsSchema URL: TODO, which is a fair summary of the situation. - Cross-registry composition is Weaver’s territory, and Weaver is actively building multi-registry layering right now.
- A change to the schema file format is an OTEP, filed in the specification
repository’s
oteps/directory. Note that the standaloneotepsrepo was archived in November 2025; proposals moved.
Not W3C — they own trace context, the propagation format on the wire, not attribute semantics. Not IETF. Not CNCF directly, which hosts OpenTelemetry but does not review its specs. And not a bilateral deal with the instrumentation vendors, who have entirely rational product reasons to keep their own vocabularies.
One more thing, which is the part nobody puts in the “how to contribute” guide: the order matters more than the document. An OTEP filed cold by someone with no history in the project is a document that gets politely queued. The sequence that works is to ship the thing, get a couple of real users, open an issue describing the problem rather than your solution, show up to the SIG call twice before proposing anything, and file small useful patches first. The spec comes last, if it comes at all. I am at the beginning of that, not the end of it, and I would rather say so than write this post as though the standard were already in flight.
The short version
Don’t build a language; one exists, and half your rules were never going to be data anyway. Do make the statable half statable, and export it into the language somebody else maintains. Do not try to standardize your mapping table.
Do standardize the thing nobody currently can say: this telemetry was translated, here is from what and to what, and here is precisely what did not survive.
Both proposals are drafted in the open and neither is filed, which is the honest state of them. The second one is visibly weaker than the first, and its own draft says so.
A translation layer that drops data silently is worse than no translation layer, because you will trust the result. That is true of my normalizer, it is true of your vendor’s ingest pipeline, and right now neither of us has a standard way to tell you which.