Context
Surfaced during adversarial review of PR LeanerCloud/cloud-commitments-cli#1227 (admin-bootstrap race fix).
PR LeanerCloud/cloud-commitments-cli#1227 added TestIntegration_CreateAdminIfNone_ConcurrentBootstrapOnce in internal/auth/store_postgres_db_test.go, which fires two goroutines × 25 barrier-synced iterations. PR body confirms this reliably fails (2 winners) against the pre-fix code, so it does catch the race.
Gap
Two callers is the production UX (two operators bootstrapping at the same time), but it's the bare minimum for a race that needs serialization. A wider concurrency factor (N=10-100) would:
- Harden the regression net against a future "optimization" that subtly weakens the lock (e.g. someone switching from
pg_advisory_xact_lock to pg_try_advisory_xact_lock and silently dropping a contender).
- Stress the connection pool / pgxpool semaphore behavior under heavy contention on the bootstrap lock.
- Catch any code path that accidentally bypasses the tx (e.g. someone refactoring to use
s.db.Exec instead of tx.Exec for the INSERT).
Proposal
Add a higher-fanout variant (e.g. TestIntegration_CreateAdminIfNone_ConcurrentBootstrapOnce_Stress) with numGoroutines = 50 and iterations = 10, behind the same integration build tag. Use the same barrier-channel + send outcomes via channel pattern (matches feedback_require_in_goroutines.md and feedback_no_sleep_in_tests.md). Keep the existing 2-goroutine test as the cheap fast-path; the stress variant runs the same assertions (exactly 1 winner, exactly 1 admin row).
Acceptance criteria
Priority
Optional — current test catches the race. This would harden against future regressions, not fix a real bug.
Context
Surfaced during adversarial review of PR LeanerCloud/cloud-commitments-cli#1227 (admin-bootstrap race fix).
PR LeanerCloud/cloud-commitments-cli#1227 added
TestIntegration_CreateAdminIfNone_ConcurrentBootstrapOnceininternal/auth/store_postgres_db_test.go, which fires two goroutines × 25 barrier-synced iterations. PR body confirms this reliably fails (2 winners) against the pre-fix code, so it does catch the race.Gap
Two callers is the production UX (two operators bootstrapping at the same time), but it's the bare minimum for a race that needs serialization. A wider concurrency factor (N=10-100) would:
pg_advisory_xact_locktopg_try_advisory_xact_lockand silently dropping a contender).s.db.Execinstead oftx.Execfor the INSERT).Proposal
Add a higher-fanout variant (e.g.
TestIntegration_CreateAdminIfNone_ConcurrentBootstrapOnce_Stress) withnumGoroutines = 50anditerations = 10, behind the sameintegrationbuild tag. Use the samebarrier-channel + send outcomes via channelpattern (matchesfeedback_require_in_goroutines.mdandfeedback_no_sleep_in_tests.md). Keep the existing 2-goroutine test as the cheap fast-path; the stress variant runs the same assertions (exactly 1 winner, exactly 1 admin row).Acceptance criteria
winners == 1andCountGroupMembers(DefaultAdminGroupID) == 1with N≥50 concurrent callers across multiple iterations.pg_advisory_xact_lockline locally.outcome{}over a channel; allrequire.*calls stay on the test goroutine.Priority
Optional — current test catches the race. This would harden against future regressions, not fix a real bug.