Skip to content

test(cli): harden output capture and signal failure diagnostics #2150

Description

@cristim

Unchanged local suite at PR #2143 head ef3a2dc showed two test-harness reliability gaps on macOS/Go1.26.6.

Observed: a race/short suite timed out at20m in TestCSVCapDropsTruncationBelowMinCount, blocked writing AppLogger output to the pipe in captureAppOutput. cmd/multi_service_stats_test.go:17 runs fn synchronously and starts reading only afterwards, so enough output can deadlock the writer; the read end is also not explicitly closed. GCP signal tests intermittently failed their10s marker waits at cmd/configure_gcp_signal_test.go:39 and:66 without including child stderr or exit diagnostics. This does not establish a production GCP cleanup defect or conclusively establish the cause of the marker timeout.

Reproduction evidence: full race suite timed out in reportMinCountDrops at cmd/multi_service.go:376 through captureAppOutput; unchanged isolated CSV test subsequently passed under race. Independent unchanged signal tests reproduced one SIGINT upload marker failure, then passed three repetitions and a race invocation. A complete unchanged short suite subsequently passed in383.064s with GOGC=10, GOFLAGS=-p=1 and off-host requests blocked by a closed loopback proxy. Resource containment is not causal proof.

Expected: output capture drains while the callback runs and closes both ends reliably, including failure paths. Failed signal assertions report captured child stderr/process state so cold-start, exit and actual hangs can be distinguished. Preserve the deliberate blocked-stdout interrupt scenario.

Proposed fix: repair captureAppOutput in cmd/multi_service_stats_test.go using the existing concurrent capture pattern if available; add a large-output failing-first regression that exceeds pipe capacity. Improve marker-timeout diagnostics in cmd/configure_gcp_signal_test.go without extending timeouts blindly or draining away its intended blocked-stdout condition. Verify repeated full race suites under representative resource pressure. Keep production code unchanged.

Severity: low, internal test reliability. References: PR #2143, existing GCP cleanup work #2119, separate ancillary SDK mock-isolation follow-up #2148. No cloud purchase or live-account acceptance was performed during this investigation.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions