Conversation
🦋 Changeset detectedLatest commit: 3da1b73 The changes in this PR will be included in the next version bump. This PR includes changesets to release 1 package
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
|
Dependency diff: |
Bridge to livekit-telemetry via UniFFI, with process-wide opt-out, attribute setting, and custom event emission. Pipeline initializes at first Room creation and shuts down at opt-out. Production warnings and errors are forwarded to telemetry.
Instrument Room connect, LocalParticipant publish, and RemoteTrackPublication subscribe with telemetry spans.
Instrument reconnect cycles in RTCEngine, and connect checkpoints with context propagation in SignalClient.
Track OS device signals (thermal state, memory pressure, network connectivity, battery level, audio route changes, app lifecycle) and capture failures for observability.
Periodic WebRTC stats export for connected peer connections.
Expose disableTelemetry on LiveKit object, and emitTelemetryEvent/setTelemetryAttribute on Room.
Robolectric tests exercise full lifecycle against local OTLP collector, verifying span export and attribute propagation.
Add a telemetry E2E step to CI with a host Rust build and a local collector, add the changeset entry, and update the detekt baseline.
aff8d93 to
3da1b73
Compare
There was a problem hiding this comment.
Devin Review found 6 potential issues.
2 flags not posted on this PR by your GitHub settings — view them in Devin Review. (Configure)
| fun disable() { | ||
| synchronized(this) { disabled = true } // after this, ifCollecting starts nothing | ||
| try { | ||
| telemetryDisable() | ||
| } catch (e: Throwable) { // a missing native library included: the opt-out never fails the app | ||
| diagnose(e, "The opt-out did not reach the core; Rooms created from now on still collect nothing.") | ||
| } |
There was a problem hiding this comment.
🟡 Early opt-out leaves cached batches
When disableTelemetry() runs before any Room exists, Telemetry.disable leaves previous-launch batches on disk. Only a later Telemetry.scope deletes that directory, so an app creating no Room retains the batches.
Learn more
The process opt-out is available before the first Room. A previous process can have left unsent telemetry under storageDirectory. In this case no pipeline is installed, and the opt-out only removes that directory on a later scope call. An app that opts out and does not create a Room retains the previous launch's batches.
Example: The app launches after an offline call left livekit-telemetry batches, calls LiveKit.disableTelemetry() in startup, and never creates a Room. The batches remain in its cache directory.
Recommended fix: Retain an application context or provide one to the early opt-out so Telemetry.disable can delete storageDirectory immediately, including when the core has not been configured; keep deletion synchronized with installation.
Was this helpful? React with 👍 or 👎 to provide feedback.
| build() | ||
| } | ||
| participant.signalClient.sendUpdateSubscription(isDesired, participantTracks) | ||
| if (subscribed) guarded { participant.signalClient.rtcTelemetry?.subscribeIntent(this, participant) } |
There was a problem hiding this comment.
🟡 Cancelled subscriptions retain open spans
After setSubscribed(true) starts a pending subscribe, setSubscribed(false) never ends it. The publication remains registered, so cancelled subscriptions appear pending until track removal or disconnection.
Learn more
The manual subscription hook starts a core subscribe span when the app requests a track. Unsubscribing changes isDesired and sends a signal update, but the core only receives trackEnded for track removal events. setSubscribed does not remove the publication, so a pending subscription remains open.
Example: An app requests an unsubscribed camera, then cancels the request before the camera arrives. The lk.subscribe span keeps measuring a request the app no longer wants.
Recommended fix: Notify the RTC telemetry instrument on the transition to false and end/cancel the pending subscribe for that SID. Keep this distinct from a later server-driven track removal.
Was this helpful? React with 👍 or 👎 to provide feedback.
|
|
||
| @Suppress("InjectDispatcher") | ||
| val poller = connection.launch(Dispatchers.Default + failOpen) { | ||
| pollStats(interval = { scope.statsPollIntervalMs().toLong() }, wake = wake) { |
There was a problem hiding this comment.
🟡 Poll interval failure stops RTC reporting
If statsPollIntervalMs() throws, pollStats terminates outside the poller's error handler. That Room never submits another RTC statistics report during the connection.
Learn more
The core supplies a polling interval for the next RTC stats request. The interval is evaluated by pollStats, outside the catch surrounding recordPeerStats. A UniFFI exception ends the poller coroutine; CoroutineExceptionHandler logs it but does not restart the worker. Other telemetry failures in the poll are explicitly swallowed so future samples continue.
Example: A core call fails while reading the next interval after a track appears. Every subsequent RTC window for the connected Room is missing, although the peer connections still work.
Recommended fix: Catch failures from scope.statsPollIntervalMs() inside the polling loop and use a bounded fallback interval, then retry the core on the next iteration.
Was this helpful? React with 👍 or 👎 to provide feedback.
| val state = DeviceState( | ||
| thermal = ThermalState.UNKNOWN, | ||
| lowPowerMode = null, | ||
| appState = AppState.FOREGROUND, |
There was a problem hiding this comment.
🟡 Background sessions report foreground state
When a background service creates the first Room, DeviceTelemetry starts with appState set to foreground. The earlier UI-hidden callback cannot reach its newly registered listener, so the session misreports background activity.
Learn more
The first Room installs the device instrument and registers memoryCallbacks afterwards. That listener only changes app state to background when a UI-hidden trim callback arrives. If an Android service first creates a Room after all activities are already hidden, that callback has already passed, while the initial state is foreground. No subsequent trim is needed to keep the process backgrounded, so telemetry reports foreground until a later lifecycle transition.
Example: The user leaves the app, then a background service creates its first Room for an audio call. Its device records begin with foreground app state and remain there while the app stays hidden.
Recommended fix: Initialize app state from a process-level activity tracker that is active before the first Room, or represent the initial state as unknown until a reliable lifecycle signal arrives.
Was this helpful? React with 👍 or 👎 to provide feedback.
| if (console || Telemetry.captures(loggingLevel)) { | ||
| val text = message() | ||
| if (console) { | ||
| logger?.log(loggingLevel, t, text) | ||
| } | ||
| Telemetry.log(loggingLevel, t, text) |
There was a problem hiding this comment.
🟥 Signaling tokens can enter telemetry logs
When signaling fails, Telemetry.log exports SDK warnings regardless of console level. sendRequestImpl logs the full request, so token-bearing requests can reach telemetry.
Was this helpful? React with 👍 or 👎 to provide feedback.
| Request.Builder() | ||
| .url(request.url) | ||
| .post(request.body.toRequestBody()) | ||
| .apply { request.headers.forEach { (name, value) -> header(name, value) } } | ||
| .build() |
There was a problem hiding this comment.
Client telemetry for Android, on top of the shared Rust core (livekit/rust-sdks#1396). Every Room reports its spans, RTC statistics, SDK warnings/errors and device state to its LiveKit Cloud project (only when the token carries the observability grant), for 1,160 lines of Kotlin and no new public types.
Public API
LiveKit.disableTelemetry()Room.emitTelemetryEvent(name, attributes = emptyMap())Room.setTelemetryAttribute(key, value)nullremovesNothing else is public. Configuration, instruments, transport and the UniFFI types stay internal; every tuning value (60 s export, 60 s windows, stats poll interval) is the core's default.
Platform code
What this platform adds on top of Rust (everything else — destination, token handling, retries, cache, holds, stats mapping, span state — is in the core).
Three files, 822 lines; 1,160 lines in total including the wiring in existing files and the build.
Files and responsibilities
telemetry/Telemetry.ktdisableTelemetry()returns: no stats request or submit starts after it; a Room created afterwards also deletes a previous launch's cache); SDK log records (ambient span, else the ambient Room, else the process); WebRTC errors; OkHttp transport that returns the raw answer and follows no redirect; enum mappingstelemetry/DeviceTelemetry.ktDeviceState, drained from one channel on its own serial dispatcher (a failing change is skipped, never ends the stream); audio route / focus (audio switch: one listener pair per handler, however many Rooms share it, released by the last) and camera / microphone failures → device eventstelemetry/RTCTelemetry.ktlk.subscribe(join-announced tracks reconciled at connect and after a full reconnect); onegetStats()per peer connection everystatsPollIntervalMs()(a track appearing re-reads the interval but keeps the deadline) →recordPeerStats; cancelled at the opt-out; a failing core call is logged and skipped, never the app's crash; report flatteningChanges in existing code
Wiring only: every existing public API, the console logging and per-track statistics behave as on main. Every telemetry call in existing code is wrapped in
guarded {}: a core failure is reported to the console on a best-effort basis (even a throwing app logger is swallowed) and never skips the SDK's own teardown or rollback.Changed files
Roomrelease(); hands the scope the app's URL and token as soon asconnect()accepts the attempt; onelk.connectspan perconnect()(engine, ICE and room-connected checkpoints); starts the RTC instrument before the join;setRoomat join and room update,disconnectedat clean-upRTCEngine(+ detekt baseline)ROOM_MOVEDis a TODO inSignalClient— so there is no move to forward); a server Leave's protocol reason and a reconnect that gave up (reconnect_failed) reachlk.room.disconnected; suspendingpeerStats()for the poller; onelk.reconnectspan per reconnect cycle with its reason, attempts as checkpoints;signal/join_recv/pc_createdcheckpointsreconnect()keeps its signatureSignalClientws_open/offer_sent/answer_sentcheckpoints; carries the Room's scope on its coroutinesLocalParticipantlk.publishspan per publish attempt, nested under a still-running ambient span and ambient itself while the publish runsRemoteTrackPublication.setSubscribedlk.subscribeand wakes the stats pollerLKLogRTCModule,CameraCapturerUtilsenableWebRTCLoggingthe console forward writes to the app's logger directly, so WebRTC lines are captured onceLiveKitdisableTelemetry(), synchronoussettings.gradle-PlivekitUniffiVersion=<version>(orLIVEKIT_UNIFFI_VERSION) resolves livekit-uniffi-android at that version from Maven Local (cargo make android-package-localin rust-sdks/livekit-uniffi publishes0.0.1), else the released artifact fromlibs.versions.tomllivekit-android-testMockPeerConnection.statsReport; JNA's JVM natives and-PlivekitUniffiLibraryPath(orLIVEKIT_UNIFFI_LIBRARY_PATH) to run the real core under Robolectric;RoomTestverifies the network-change reconnect reason and gives its mocked participant a real participant's empty sid.github/workflows/android.ymlTelemetry E2E teststep after the existing build-and-test step: buildslivekit-uniffifor the host at the rust-sdks tag matchinggradle/libs.versions.toml(livekit-uniffi/v<version>), starts otelcol-contrib 0.162.0 (SHA-256 checked) on :4319, and runs the telemetry package's tests (TelemetryMockE2ETestand the platform tests that need the core) withLK_TELEMETRY_ENDPOINT; a collector that never listens, or a skippedTelemetryMockE2ETest, fails the joblibs.versions.toml: until then the earlier build step fails to compileEvents
10 of 19 SPEC signals fully covered, 8 partially (platform limits or pending core capture), 1 skipped (smoke-test only).
Event coverage table
lk.connectspan (+ checkpoints)Room.connect:ws_open·signal·join_recv·pc_created·offer_sent·answer_sent·engine·pc_connected·room_connected(signal/join_recvand the last three are stamped together: the SDK has no separate moment for them)lk.reconnectspanRTCEngine.reconnectcycle; reason from the trigger (signal close, primary/publisher ICE, network change);attempt <n> quick|fullper attemptlk.publishspanLocalParticipantpublish, nested under a still-running ambient span; its warnings point at itlk.subscribespansetSubscribed(true)(wakes the poller); subscribed / failed / unsubscribed / unpublished; first media seen by the corelk.rtc.stats.samplegetStats()per peer connection (publisher + subscriber), paced by the corelk.room.disconnectedagent_errortoo), orreconnect_failedwhen a reconnect cycle gives uplk.telemetry.reportcustom.<name>Room.emitTelemetryEventLKLog) and WebRTC errors. The Rust core's own warnings/errors are not captured: Android does not install the Rust log forwarder (logForwardBootstrap), which is what makes the core copy them itself; it never passes Rust log entries totelemetryLog, so nothing is counted twicelk.device.thermal.changedPowerManagerthermal status on API 29+; reported as unknown below (not exposed by the OS)lk.device.low_power.changedACTION_POWER_SAVE_MODE_CHANGED); unknown without a power servicelk.device.app_state.changedTRIM_MEMORY_UI_HIDDEN), foreground at an activity start; no lifecycle dependencylk.device.memory.changedonTrimMemory/onLowMemory; the OS never signals the end of pressure, so it counts as normal again when an activity starts; API 34+ no longer delivers the running-app levelslk.device.network.changedlk.device.battery.changedACTION_BATTERY_CHANGEDlk.device.audio_route.changedAudioSwitchHandleronly); it names no reasonlk.device.audio.interruptionAudioSwitchHandleronly)lk.device.capture.failedother)lk.pingLocal testing
Run the Android end-to-end test against a local OTel backend (LGTM) and look at a whole call in the local LGTM UI on :3000, in a few minutes.
Commands
The test (
TelemetryMockE2ETest) runs the real Rust core on the SDK's mocks and reads back what its own collector wrote, so the core talks to that collector on :4319 and an overlay config (otelcol-lgtm.yaml) makes it forward everything to LGTM on :4318 as well.Then open http://localhost:3000 → Explore:
{resource.service.name="livekit-client-android"}→ the call'slk.connect,lk.reconnect,lk.publish,lk.subscribespans, with their checkpoints as span events.{service_name="livekit-client-android"}→lk.rtc.stats.samplewindows,lk.device.*events,custom.e2e.checkpoint,lk.room.disconnectedand the warning/error records; filter by the session's trace id to follow one Room.Pointing
LK_TELEMETRY_ENDPOINTstraight athttp://localhost:4318works for a sample app, but the e2e test then skips: it asserts on its own collector's output.