Skip to content

test(telemetry): cover the pipeline and the backend contract - #1484

Open
pblazej wants to merge 2 commits into
blaze/telemetry-stack/5-pipelinefrom
blaze/telemetry-stack/6-pipeline-tests
Open

pblazej wants to merge 2 commits into
blaze/telemetry-stack/5-pipelinefrom
blaze/telemetry-stack/6-pipeline-tests

Conversation

@pblazej

@pblazej pblazej commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Tests for #1483: the pipeline end to end, and the backend contract one test per row. Every link below points at the stage that adds the test.

Changes

Backend contract: every status code and condition → behaviour → test (29 rows)

A Room hands over its server URL and token; only a parsed *.livekit.cloud host over TLS gets https://<host>/observability/client/{logs,traces}/otlp/v0 with Authorization: Bearer <token>. Every collector answer is classified by the core; behaviour marked custom is not prescribed by the OTLP/HTTP spec.

Condition Behaviour Test Why
Server URL parsed (WHATWG); token only for a TLS scheme, a domain under .livekit.cloud with its own label, default port, no userinfo; endpoint built from that host alone; anything else collects nothing look_alike_server_urls_never_get_the_token
only_livekit_cloud_hosts_get_an_ingest_url
the_ingest_url_and_token_come_from_the_room
self_hosted_servers_get_nothing_and_nothing_is_kept
custom — self-hosted servers have no ingest (CLT-3335); credentials only to the authorized origin
Token handed over again same pair: no-op (a reconnect retries a 404'd ingest); new token: uploads resume at once handing_over_the_same_token_again_is_free
a_reconnect_with_the_same_token_retries_a_404
an_expired_token_holds_uploads_until_a_fresh_one_arrives
RFC 6750 bearer; custom refresh contract
Token expired or missing never sent; batches wait (hard hold). A past, negative, fractional-past or non-numeric exp is expired; only a missing one means no known expiry; a far-future one is clamped to a year an_expired_token_holds_uploads_until_a_fresh_one_arrives
untrusted_expiry_claims_fail_closed
uploads_wait_for_a_destination
hard_holds_have_no_escape_hatch
custom — untrusted claims fail closed
First token lacks the observability grant nothing collected for that project a_room_without_the_observability_grant_sends_nothing
a_token_without_the_grant_is_not_consent
custom — the grant is the customer's consent
Refresh drops the grant (today's SFU) keep the granted token until it expires, then wait a_refresh_that_drops_the_grant_keeps_uploading_with_the_granted_token
a_refresh_that_drops_the_grant_keeps_the_granted_token_until_it_expires
custom — bridge until OCD-3387
Ownership each record (and each RTC window, at open) captures (project, session); credentials keyed by it; the answer is attributed to the project the request went to; a Room switching projects takes nothing along; an unconnected Room never borrows another Room's project; a Room's records from before it connected, cached ones included, are bound on disk to that Room's first project by the exporter pass its connect wakes (at once if the exporter is idle, otherwise once the export pass then running finishes; no single-request bound), so they replay after a restart; a crash once a batch's bound copy is journaled rolls forward to it. A Room that never connects before the process ends — or a process killed after the connect but before that batch's bound copy is journaled — leaves the batch unowned: held until the 24 h expiry, then counted expired answers_to_pre_connect_batches_are_attributed_to_their_own_project
pre_connect_records_replay_after_a_restart
a_crash_while_binding_still_replays_to_the_rooms_project
a_room_switching_projects_takes_nothing_along
an_unconnected_room_never_borrows_another_rooms_project
two_rooms_on_two_projects_never_share_a_token_or_a_destination
ownership_survives_a_room_changing_projects
rooms_never_borrow_each_others_tokens
custom — never borrow a credential (RFC 6750 §5.2)
Credential lifecycle a refusal is recorded by credential identity, shared by every slot (Room, project copy, process route, restart), until the token expires; a session's credentials stay while it is alive (one per project it used) or cached batches need them — pre-connect batches included; project copies only while needed refusals_stick_and_credentials_follow_live_rooms
a_gone_rooms_pre_connect_backlog_keeps_its_credential
custom
Restart with cached data, no token wait for a token of the same project (≤ 24 h); tokens never persisted a_restart_with_cached_data_and_no_token_waits_for_the_same_project custom
2xx batch removed batches_events_into_one_otlp_request OTLP specifies 200 for success; accepting every 2xx is custom
2xx + partial_success rejected records counted, never retried partial_success_counts_the_refused_records_and_never_retries
success_and_partial_success
OTLP: MUST NOT retry partial success
400, 3xx, other 4xx dropped, counted rejected; a delay hint changes nothing a_bad_request_drops_the_batch
client_errors
delay_hints_never_make_a_final_status_retryable
OTLP: not retryable
413 halves replace the batch in a journaled cache transaction (on disk stays on disk), retried at once down to one record, each against the pass budget; a lone oversized record dropped (oversized) payload_too_large_splits_down_to_single_records
a_split_never_moves_a_batch_off_the_disk (added in #1485)
a_split_crashed_at_every_step_loses_and_duplicates_nothing
custom — OTLP only says don't retry as-is; RFC 9110 §15.5.14 (test lands in #1485)
401/403 "data recording is disabled by owner" project silent for the process, its cache purged disabled_project_goes_silent_and_purges_its_cache
disabled_by_owner
custom — exact phrase match, no machine code yet
Other 401/403 kept; that token never sent again; wait for the next unauthorized_holds_the_batch_until_a_new_token
repeated_unauthorized_answers_wait_for_renewal_without_loss
a_refused_token_is_never_sent_again
custom — RFC 9110 §15.5.2
404 project silent until its next token or a reconnect; cache purged not_found_on_the_derived_endpoint_goes_silent
not_found_recovers_with_the_next_token_disabled_does_not
a_reconnect_with_the_same_token_retries_a_404
custom (CLT-3335), scoped to the origin
429 that destination pauses for Retry-After, else RetryInfo, else 60 s; collection continues throttling_honors_retry_after_then_retry_info_then_a_minute
a_real_http_collector_throttles_and_recovers
OTLP: retryable, honor Retry-After; RFC 6585 §4; 60 s default custom
503 with a delay that destination pauses for it throttling_honors_retry_after_then_retry_info_then_a_minute OTLP: retryable; RFC 9110 §10.2.3
502 / 503 / 504 backoff 1 s → 60 s, full jitter; exhaustion pauses, never deletes (24 h retention) a_failing_server_never_costs_a_batch
server_errors
OTLP: retryable with exponential backoff; the constants and retention are custom
500 + RetryInfo retried after the named delay (a header alone does not count) a_retryable_500_waits_for_its_retry_info
delay_hints_never_make_a_final_status_retryable
custom — Cloud sends retryable 500s
Other 5xx dropped, counted rejected other_server_errors_drop_the_batch OTLP: MUST NOT retry
Retry-After / RetryInfo values seconds or HTTP-date; garbage ignored, negative → now, oversized saturates; clamped to 24 h; shutdown never cuts it short untrusted_time_values_are_bounded
retry_after_http_dates
shutdown_never_cuts_a_server_delay_short
RFC 9110 §10.2.3; the clamp is custom
One project failing only that destination pauses; evictions count as throttled only for the paused destination a_failing_project_does_not_pause_the_others custom
Timeout backoff; counted apart timeouts_are_counted_apart_from_failures OTLP: retryable
Connection, DNS, TLS failure backoff 1 s → 60 s full jitter; never a capability decision no_answer_backs_off_exponentially_with_full_jitter
backoff_doubles_with_full_jitter_up_to_a_minute
OTLP: retry with backoff
Foreign transport throws an undeclared exception a retryable failure, not a panic (the Rust conversion hook is tested; no generated Swift/Kotlin transport test) an_undeclared_foreign_exception_is_a_retryable_failure custom policy on the UniFFI callback contract
Invalid request (transport-side) dropped an_invalid_request_is_dropped custom — treated as permanent
Redirect a returned 3xx drops the batch; the livekit-net client strips Authorization when the host or port changes (a scheme-only change on the same explicit port is not tested) redirects_to_another_host_or_port_never_carry_the_token
client_errors
RFC 9110 §15.4; custom
Request shape protobuf, gzip, Priority: u=7; ≤ 512 records and ≤ 1 MiB encoded, session attributes and the self-report included; the self-report is added at most once per pass and never takes the place of real records, even at a batch size of 1; a single record over the limit dropped (oversized) a_failing_self_report_never_starves_real_records (added in #1485)
requests_are_gzipped_and_low_priority
encoded_requests_stay_under_the_byte_limit_for_both_signals (added in #1485)
final_limits_hold_after_decoration_and_caller_strings_are_bounded (added in #1485)
OTLP/HTTP binary + gzip; RFC 9218; the limits and the priority are custom (test lands in #1485)
Local collector LK_TELEMETRY_ENDPOINT: everything there, no Cloud rules, no token the_override_reaches_a_local_collector_without_cloud_rules
the_override_takes_everything_without_a_token
custom, test-only, no platform API
Verification

At 3205dd2a, from a clean checkout (CI's test workflow runs only for PRs into main, so these were run locally; there is no clippy job in CI):

  • cargo fmt -- --check
  • cargo clippy -p livekit-telemetry --all-targets --all-features -- -D warnings
  • cargo check -p livekit-telemetry --all-targets --no-default-features with features [], [net], [uniffi], [net,uniffi]
  • cargo test -p livekit-telemetry: 126 unit, 2 doc; --all-features: 129 unit, 2 doc

cargo doc -D warnings reports exactly what it reports on 6aba1b68 (private-item links, ExportError::from_response).

Drive the pipeline through a scripted transport: batching into one OTLP
request, gzip and priority, retry after backoff, rejection and throttling,
the write-ahead cache replayed on the next start, holds while connecting
and under device pressure, the per-interval budget, the byte cap, flood
guard and log floors, sessions with their own trace and attributes, typed
and subscribe spans, RTC stats windows and the loss counters.
One test per row of the backend contract: every collector status code and
condition, the destination and credential rules, and a mock collector
driven over the real `livekit-net` HTTP stack.
@pblazej pblazej added the internal to tag changes that don't require changelog documentation label Oct 1, 2026
@pblazej
pblazej force-pushed the blaze/telemetry-stack/6-pipeline-tests branch from 59a5570 to 3205dd2 Compare October 1, 2026 13:53
@pblazej
pblazej added this pull request to stack #1486 October 1, 2026 14:23
@pblazej
pblazej marked this pull request as ready for review October 1, 2026 14:33
@pblazej
pblazej requested a review from ladvoc as a code owner October 1, 2026 14:33

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔍 Devin Review: 2 flags

Not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)

Devin Review

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

internal to tag changes that don't require changelog documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant