Repository navigation
otel: read the engine's provenance flags instead of guessing from values - #488
Merged
Merged
Conversation
Bumps core 1.8.0 -> 1.10.0 and consumes what SmooAI/smooth-operator-core#177 added, closing a hole in #474 that I introduced. #474 decided "was this measured?" by testing `prompt_tokens > 0`. That inferred provenance from the VALUE, and it only worked on one of the two fabrication sites in core: the streaming one hardcodes `prompt_tokens = 0`, but the non-streaming one estimates it from the outgoing request's JSON length and so produces a plausible non-zero. The heuristic caught the first and silently waved the second through as measured. `record_turn_usage` now reads `AgentEvent::Completed.usage_estimated`. A flag carries the fact; a heuristic guesses it — the same reason `absent` beats `0`. Two new attributes, both previously impossible without the core change: - `gen_ai.usage.cost_source` = `gateway` | `estimated`, set alongside `cost_usd` from `Completed.cost_estimated`. Local `ModelPricing` returns the FREE tier for any model it doesn't recognise, so an estimate can be a wild under-count while looking exact; a billed surface must not render the two identically. - `gen_ai.response.id` from `Completed.response_id`, recorded whenever present. It joins to `LiteLLM_SpendLogs.request_id`, whose row carries the gateway's authoritative dollars AND real token counts — so it matters most exactly when the counts above are missing, and it turns "measured vs estimated" from a claim into something reconcilable. Both had to be declared `tracing::field::Empty` at span creation; `record` silently no-ops on an undeclared field, which is how the first run recorded neither while every test still compiled. `TurnUsage` carries the three engine flags but they are TELEMETRY-only — `eventual_response` still serializes exactly `{costUsd, promptTokens, completionTokens}`, now pinned by an assertion on the object's key count with all three flags set in the fixture. It also loses `Copy`, since `response_id` is an owned `String`. Verified by restoring the old `prompt_tokens > 0` gate: the new test fails with `input_tokens: "372"` — an invented count published as measured. 576 tests pass across 46 suites (+1 new). clippy -D warnings and fmt --check clean. Cargo.lock also unifies socket2 0.5.10 -> 0.6.4 for hyper-util/tokio/tokio-postgres; both versions remain in the lock and it is incidental to the core bump. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01LkCz96UfUxcai5RU4LwPnG
|
This was referenced Aug 18, 2026
brentrager
added a commit
that referenced
this pull request
Aug 18, 2026
#488 merged without one, same as #470, #471 and #474 before it. Four for four on this ticket — the release here is changeset-driven, so a merged fix that carries no changeset publishes nothing and every consumer keeps the old crate. Claude-Session: https://claude.ai/code/session_01LkCz96UfUxcai5RU4LwPnG Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Required follow-up to #474, closing a hole I introduced there. Consumes SmooAI/smooth-operator-core#177 (published as core 1.10.0, verified on crates.io — see below).
The hole in #474
#474 decided "was this measured?" by testing
prompt_tokens > 0. That infers provenance from the value, and there are two fabrication sites in core that differ:prompt_tokens0So the heuristic caught the streaming path (the common one, which is why #474 was still a real improvement) and silently waved the non-streaming path through as measured.
record_turn_usagenow readsAgentEvent::Completed.usage_estimated. A flag carries the fact; a heuristic guesses it — the same reasonabsentbeats0.Proof, by restoring the old gate and re-running the new test:
372is invented. The old gate published it as a measurement.Two new attributes
gen_ai.usage.cost_source=gateway|estimated, set alongsidecost_usdfromCompleted.cost_estimated. LocalModelPricingreturns the free tier for any model it doesn't recognise, so an estimate can be a wild under-count while looking exact. A billed surface must not render the two identically.gen_ai.response.idfromCompleted.response_id, recorded whenever present. Joins toLiteLLM_SpendLogs.request_id, whose row carries the gateway's authoritative dollars and real token counts — so it matters most exactly when the counts are missing, and it turns "measured vs estimated" from a claim into something reconcilable after the fact.Both had to be declared
tracing::field::Emptyat span creation.Span::recordsilently no-ops on an undeclared field, so the first run recorded neither while everything still compiled and only the assertions caught it.Wire compatibility
TurnUsagenow carries the three engine flags, but they are telemetry-only:eventual_responsestill serializes exactly{costUsd, promptTokens, completionTokens}. Pinned by an assertion on the object's key count, with all three flags set to non-default values in the fixture so a leak fails the test.TurnUsagelosesCopy(response_idis an ownedString); the one call site already passed by reference.Verification
cargo test -p smooai-smooth-operator-server -p smooai-smooth-operatorcargo clippy --all-targets -- -D warningscargo fmt --all --checkCore publish verified directly rather than trusting a green workflow — literal
200fromstatic.crates.ioforsmooai-smooth-operator-coreand-temporalat both 1.9.0 and 1.10.0, and the published 1.10.0 tarball was unpacked and confirmed to containpub usage_estimated,pub response_id,ResponseId {andcost_estimated.Cargo.lockalso unifiessocket20.5.10 → 0.6.4 for hyper-util/tokio/tokio-postgres. Both versions remain in the lock; incidental to the core bump, and I diffed the lock rather than assuming.No changeset — the release for this repo is being driven separately.
🤖 Generated with Claude Code
https://claude.ai/code/session_01LkCz96UfUxcai5RU4LwPnG