Repository navigation
Move over gocached - #3
Merged
Merged
Conversation
Migrated-from: bradfitz/go-tool-cache@4fed81d
Signed-off-by: Brad Fitzpatrick <brad@danga.com> Migrated-from: bradfitz/go-tool-cache@e435233
Migrated-from: bradfitz/go-tool-cache@fa12dc1
Migrated-from: bradfitz/go-tool-cache@a411ed2
Migrated-from: bradfitz/go-tool-cache@ef6c7b1
Signed-off-by: Brad Fitzpatrick <brad@danga.com> Migrated-from: bradfitz/go-tool-cache@1a57bfe
Signed-off-by: Brad Fitzpatrick <brad@danga.com> Migrated-from: bradfitz/go-tool-cache@4c504e6
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@2db3b0f
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@d997e5f
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@e221bff
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@1471a41
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@cf8026f
Zero-sized objects are common and worth optimizing a bit. Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@77ecb3b
... for efficiency when multiple actions map to the same dup object ID. And to start laying the groundwork for namespaces. Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@fd8964a
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@643853b
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@5ffda40
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@13dbb96
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@48dd315
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@1346ba9
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@01466c7
So we can configure it so guest VMs don't have access to debug info. Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@2513dd2
Fixes bradfitz/go-tool-cache#15 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@e2eb857
…ool-cache#18) Running with a cold disk cache for go-cacher and a warm disk cache for go-cacher-server, I was getting no cache hits. This is due to a missing component in the output filename when serving up what go-cacher-server thinks is a cache hit; it was trying to serve from `<dir>/o-<hex>` instead of the real file location of `<dir>/<o2>/o-<hex>` where o2 is the first 2 characters of the hex string. To help with consistency, I consolidated all places that need to calculate an action or output file path into 2 shared functions, but there is only one functional change. Tested via the benchmarks I included in bradfitz/go-tool-cache#17, but it's probably worth investing in some unit tests as well as we start to use this more seriously. Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@c868ae4
…l-cache#21) Some minor fix-ups to avoid unnecessarily wide fs permissions or HTTP error messages that could accidentally expose details that we expect to stay internal. Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@38eb341
Adds the ability to run gocached with a token exchange endpoint that starts a new session each time it receives a valid JWT. Each session is associated with an opaque access token and has session-based cache stats that can be retrieved out of band at the end of the build/test process. JWT issuer and claim support is intentionally narrow initially, only supporting a single issuer and exact-match string claims. The next planned follow-up is to support read/write for _all_ tokens, but only JWTs that satisfy the -global-jwt-claim requirements will be able to open a session that can write to the global namespace 0. All other sessions will write to their own namespace based on some other criteria. Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@f8ee03e
…e#22) Pulls the bulk of gocached into a library package to more easily reuse it with custom config from other main packages. There are still likely some more changes to come to the JWT configuration API surface, but this makes sense as a first chunk of work to land. Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@0124e69
We weren't specifying the NamespaceID so the primary key wasn't being used, as the namespace ID is the prefix of the primary key. Updates tailscale/corp#34696 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@80f0f91
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@9fc622d
One of our cigocached servers' WAL had grown to 46 GB over ~87 days of
uptime, with walFindFrame pegging CPU and stalling go-cacher clients on
the CI runners. A CI vet job that normally runs in ~100s on a laptop
was sitting at 30+ minutes.
SQLite's autocheckpoint runs only PASSIVE checkpoints, which reuse WAL
space in place but never shrink the file on disk.
Fix:
- Set journal_size_limit=1 GiB per connection so even passive
checkpoints will truncate down to that bound when they can.
- Run wal_checkpoint(TRUNCATE) once at startup (this is what shrinks
the existing 46 GB WAL on the running server after deploy) and
periodically (every minute) from a background goroutine, logging
any partial checkpoints that hint at a long-running reader.
- Run a final TRUNCATE checkpoint on Close.
- Bump modernc.org/sqlite v1.45.0 -> v1.51.0 for general fixes
accumulated since February.
Also add two new gauges for ongoing visibility:
gocached_sqlite_data_bytes size of the main .db file
gocached_sqlite_wal_bytes size of the .db-wal file
Updates tailscale/corp#42670
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Migrated-from: bradfitz/go-tool-cache@097527a
The usageStats() LEFT JOIN every 5 minutes didn't scale. At 272M
Actions it took ~7 minutes per call, pinning a reader snapshot ~60% of
the time. The WAL grew to 52 GB because SQLite's PASSIVE
autocheckpoint could never advance past the snapshot; walFindFrame
pegged CPU and go-cacher clients stalled. Startup also blocked on
usageStats + cleanOldObjects for 10+ minutes. cleanCandidates was the
same shape: GROUP BY scaling with total rows.
This redoes the usage stats + LRU cleanup pipeline.
Shard the histogram, persist it, refresh in the background:
* New BlobShardStats(Prefix, ScannedAt, StatsJSON) table.
* 256 shards keyed on a SHA256 hex prefix (--shard-prefix-len=1..4).
* runShardStatsLoop picks the oldest shard, sleeps until it crosses
shardStalenessTarget (10m), then rescans.
* Per-shard scan is one SUM(CASE WHEN…) query scoped by
SHA256 >= ? AND SHA256 < ?. No row streaming.
* On restart, loadShardStats seeds lastUsage from the persisted rows.
Server starts up to usable right away without a blocking step.
Action-LRU cleanup driven by idx_actions_access:
* evictOldestActions walks the access-time index for the N oldest
stale Actions, deletes them, orphan-deletes Blobs.
* INDEXED BY locks the plan; TestEvictionQueryPlan{,_atScale} asserts
via EXPLAIN QUERY PLAN.
* Action-LRU instead of Blob-LRU: a stale Action is evicted even if
its Blob has fresher siblings. Equivalent for the common 1:1 case;
stricter LRU on shared content.
* One short Tx per 200-row batch on a 1s tick. No multi-minute reader
snapshots.
Dead-reckon PUTs and evictions between scans:
* Per-shard mutex-guarded counters track (count, bytes) added or
removed since the shard was last persisted.
* cleanupTick includes the delta in the size pressure check so a
burst of PUTs between scans triggers cleanup immediately.
Observability: /usage shows cohort table, shard freshness histogram,
per-shard scan duration p25/p50/p90; new gauges
gocached_shard_stats_{unscanned,oldest_age_seconds}, pending blob
count/bytes; new histogram gocached_shard_scan_duration_seconds
replayed from persisted data at startup; new counter
gocached_evicted_actions.
Schema migration is additive (schemaVersion stays 4).
Perf test (TestPerfQueries) always runs but defaults to 100 rows so
normal `go test` is sub-second. Operators set --performance-test-rows
larger for real scale; seed reused across runs via
--performance-test-dir. Sample, cumulative across rows:
rows DB scanShard evict legacy GROUP BY ratio usageStats
1K 396K 188µs 5.2ms 648µs 0.1x 321ns
10K 3.4M 248µs 5.5ms 6.4ms 1.2x 227ns
100K 35M 1.8ms 6.1ms 73ms 12x 328ns
1M 349M 22ms 7.5ms 763ms 102x 231ns
2M 698M 51ms 8.2ms 1.6s 196x 261ns
5M 1.8G 162ms 8.7ms 4.1s 475x 224ns
10M 3.5G 349ms 9.0ms 8.3s 922x 249ns
usageStats and evictOldestActions are flat in N. scanShard scales with
per-shard rows. The "legacy GROUP BY" column is the cleanCandidates
query gocached used before this branch (since deleted); it scales
linearly with total rows, the behavior that was wedging prod at
250M.
Updates tailscale/corp#42670
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
Migrated-from: bradfitz/go-tool-cache@d48e363
Remove the special-case concept of "global writes" and instead allow callers to provide a namespace mapping function that makes policy decisions about which clients are globally trusted and which clients should be isolated from each other. As a result, all clients are now able to write, but not necessarily into the shared global namespace. All clients can still read from the global namespace, as well as their own. The Namespaces table already existed in the schema, but we drop the lowercase constraint. However, there are some breaking changes in the package API, WithJWTAuth now takes issuer URLs only and policy moves into the new WithNamespaceMapping option. cmd/gocached implements the spirit of the old API in terms of a namespace mapping function, with the main difference that it now allows writes if you don't have the global claims, but just into your own isolated namespace. Updates tailscale/corp#38092 Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@acf1bfd
Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@6c2916a
…radfitz/go-tool-cache#38) Adds new gocached_blob_size_bytes histogram metric for tracking blob sizes as transmitted on the wire. Adds result label (put, dup, error) to PUTs latency, and storage label (disk, inline, none, error) to GETs latency. Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@d4ff6bc
Our AWS cache instance stores everything on its ephemeral NVMe disk, so the whole cache is lost on every reboot and its size is capped by the local disk. Let the blob directory live on a big slow filesystem (an S3 Files NFS mount) as the source of truth, with a new optional --hot-dir flag naming a fast local tier (NVMe) holding a bounded copy of recently used blobs, capped by --hot-capacity-gb. The new --sqlite-dir flag lets the metadata database live on its own volume (EBS), defaulting to the blob directory as before. PUTs write through to both tiers in one pass; a hot tier failure never fails the PUT since a later GET re-promotes. GETs try the hot tier first and asynchronously promote cold hits into it. The hot tier is tracked by an in-memory LRU index seeded from an mtime-ordered scan at startup, evicting least recently used files down to 95% of capacity; hot files are disposable copies, so eviction and reboot loss are always safe. When tiering is not enabled, behavior is unchanged. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I298711f6ee38efa15304aad7295736197c3de084 Migrated-from: bradfitz/go-tool-cache@180dd42
After deploying storage tiering to the AWS CI cache instance (tailscale/corp#44699), the instance showed a sustained ~11 MB/s and ~1400 IOPS of writes to the EBS volume holding the SQLite database, roughly 200x the logical metadata growth rate, and PUT latency averaged 3.6 seconds with ~520 PUTs in flight against a throughput ceiling of ~115 PUTs/sec. The cause: each PUT held sqliteWriteMu across two separate autocommit statements (the Blobs insert and the Actions insert), paying two WAL commits and two fsyncs per PUT. At a few milliseconds per fsync, that serialized writer capped throughput at almost exactly the observed rate; everything beyond it was queueing. Wrap the two inserts in a single explicit transaction so each PUT pays for one WAL commit instead of two, roughly halving the writer's per-PUT hold time and doubling the metadata throughput ceiling. Also only update the in-memory usage deltas after the transaction commits, so a failed commit can no longer skew the usage accounting. This is a stopgap; the longer term plan is to queue PUT metadata in memory (serving GETs from it) and batch the SQLite writes asynchronously, coalescing hot B-tree pages across many PUTs. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: Id3967d0e5000df638b7b76d203f2e819 Migrated-from: bradfitz/go-tool-cache@a0c534b
…itz/go-tool-cache#40) Most of the time, PUTs don't significantly add to a build's latency, but occasionally we've seen an outsized amount of time spent on doing PUTs for a build. In addition to making BestEffortHTTP ignore errors, also make it run PUTs in a background goroutine with support for timeouts and concurrency control. This better represents best-effort PUTs as an optimisation, not a dependency for correctness. Disk writes still have to happen synchronously though, because cmd/go expects the file to exist and be readable/seekable as soon as the RPC returns. Updates tailscale/corp#45334 Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@0bb331e
Pure refactor in preparation for the async PUT pipeline: move the two PUT metadata inserts into insertPutTx (so a future batch flusher can reuse them per item inside one transaction) and the GET response tail into writeObjectResponse over a new objectSource struct (so a pending, not-yet-committed PUT can be served through the same code path as a row from SQLite). No behavior change. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I13d45d1570e8c5dfeee289097f39dab5 Migrated-from: bradfitz/go-tool-cache@7cbc4ae
The upcoming async PUT pipeline spools uploads into a .put-queue directory on the hot tier's filesystem (so installing a spooled blob into the hot tier is a cheap same-device rename). Spool files are not hot tier entries, so teach the startup scan to skip dot-directories rather than index their contents. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I13648874ec103811e9274a18f66757f3 Migrated-from: bradfitz/go-tool-cache@e2c2f07
Add the data structures for the async PUT pipeline: a pending map of PUTs whose metadata hasn't reached SQLite yet (so GETs can serve them in the meantime), semaphore-based backpressure that reserves a PUT's declared size and an entry slot before any bytes are spooled (blocked PUTs abort with the HTTP request context when the client goes away), and a spooler that writes blob bytes to a .put-queue directory on the hot tier's filesystem, lz4-compressed with the main tier's policy so installing a spooled file elsewhere is a rename or a dumb copy. Startup deletes any spool files left over from a previous process: their metadata lived only in that process's memory, so they're unreachable, and losing them is fine for a cache. Nothing uses the queue yet; a following change adds the background movers and the batch flusher and cuts the PUT path over. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I943286dfcf3133042205c8fce5280100 Migrated-from: bradfitz/go-tool-cache@472e27f
Add the pipeline that retires pending PUTs: mover goroutines copy spooled blobs into the main blob directory (in parallel, since that may be a high-latency network filesystem), and a single flusher commits metadata in batched SQLite transactions of up to 512 entries or one second, whichever comes first. Batching is what removes the ~200x write amplification of per-PUT transactions: dirtied B-tree pages coalesce across a batch and the WAL fsync is paid once per batch instead of once per PUT. After a batch commits, each entry's spool file is renamed into the hot tier (same filesystem, so it's free) or deleted. The ordering is deliberate: the main-directory copy happens before the SQLite commit and the hot install after it, so the hot tier never holds the only copy of a blob, and once an entry leaves the pending map its row is already committed, leaving no window where the object is unreadable. Copy or flush failures are retried a few times and then the pending PUTs are dropped: the client already saw success, but a cache is allowed to forget. drainPendingPuts runs the whole pipeline synchronously for tests (which disable the background loops) and for shutdown. Nothing feeds the queue yet; the PUT handler cutover is next. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I69093dff61bf134df2613da9fbcddfef Migrated-from: bradfitz/go-tool-cache@a959d2f
Cut the PUT path over to the put-queue pipeline. A PUT now reserves queue room (blocking under backpressure, aborting if the client goes away), hashes and spools its body to the local fast disk (or memory, for inline-sized objects), registers a pending entry, and returns. No SQLite transaction, no copy to the (possibly slow, networked) main blob directory, and no hot tier write remain on the request path; the background movers and the batching flusher do all of that. GETs consult the pending map before SQLite, so a just-written object is immediately readable. If a pending entry retires between lookup and open, the handler falls through to the SQL path, where the row is already committed. Duplicate PUTs of a still-pending action are detected synchronously against the pending map (first one wins); duplicates of an action already in SQLite are discovered and counted at flush time instead, since the request path no longer reads the database. On the AWS CI cache instance, per-PUT metadata transactions cost ~200x write amplification on the SQLite EBS volume (every PUT dirtied ~5 random B-tree pages, each written twice via the WAL) and the serialized per-PUT fsync capped throughput at a few hundred PUTs per second, with multi-second queueing latency during bursts. Delete writeDiskBlob, finishHotWrite, and failsafeWriter: blob installs now happen off the request path, and the hot tier copy is a rename of the spool file after the metadata commits rather than a second synchronous write. Close drains the queue so a graceful shutdown loses nothing. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: Ib0087b405a2bf71e5057ea32686b103e Migrated-from: bradfitz/go-tool-cache@0ba34ac
Benchmark the end-to-end server cost of PUTs, including the amortized background pipeline (drained synchronously every 4096 requests since benchmarks run without the background loops). On a dev box, inline PUTs cost ~70us and 8 KB disk PUTs ~86us, versus the multi-millisecond serialized fsync each PUT paid when it ran its own SQLite transaction. Draining without running workers exposed a test-only deadlock: drains retire entries from the pending map but nothing consumed the worker channels, so their buffers eventually filled and enqueue blocked. Make drainPendingPuts empty the channels first; in production the count semaphore bounds channel occupancy, so the workers never see this. Also document the acknowledge-then-flush design on the Server type. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: Icf872ce78d3eca1caac358eb67620686 Migrated-from: bradfitz/go-tool-cache@87f35db
A PUT reserves put-queue quota for its declared size before reading the body, so a client that dies mid-upload (e.g. a CI VM getting terminated) pins that reservation until the read fails. With no server read timeouts, that's Go's default TCP keepalive: some 2-3 minutes on Linux. Large reservations parked that long are wasted queue budget, and a pathologically slow-but-alive sender could hold one indefinitely. Set an explicit per-request read deadline via ResponseController, scaled by the declared Content-Length at putMinUploadRate plus a flat grace. Clients are same-region, same-VPC AWS instances that normally upload orders of magnitude faster, but they're TLS-speaking guest VMs that may be CPU over-subscribed, so the rate is deliberately slack (1 MB/s): the point is some reasonable bound, not a tight one. The deadline is set after the backpressure wait so queueing time doesn't count against the upload, and cleared on return so it doesn't outlive the request on a reused connection. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I07deb508042c57a40852fae012810d0c Migrated-from: bradfitz/go-tool-cache@f8a21aa
A crash between a mover's copy into the main blob directory and the flusher's SQLite commit leaves a blob file with no metadata row. Pending state is memory-only, nothing walks the blob directory comparing it against the database, and shard stats are SQL-driven, so such orphans were invisible: unaccounted space consumed forever unless the same content happened to be PUT again. The same window existed before the async pipeline (writeDiskBlob renamed the blob into place before the insert), just microseconds wide instead of seconds. Close it with a write-ahead cleanup intent: before copying a blob, the mover creates an empty file named .cleanup/<8-hex-unix-seconds>-<blob filename>, a durable note to delete that blob at the encoded time (an hour out) if it's still unaccounted for then. After the metadata commits, the intent is deleted; the flusher hands those deletions back to the movers so it never blocks on an NFS round-trip itself. The intent directory, unlike the put-queue spool, survives restarts: intents that outlive a crash are exactly the point. A minutely sweeper lists the intent directory (the fixed-width hex time makes lexical order chronological, so it stops at the first intent not yet due) and deletes blobs with no Blobs row, skipping any SHA referenced by an in-flight pending PUT: a later PUT of identical content adopts an existing file via the mover's stat-skip without rewriting it, so the pending check is what protects the whole copy-to-commit window, with no locking against the movers or flusher. Intents for accounted-for or in-flight blobs cost nothing but their own eventual deletion. Dropped PUTs deliberately keep their intent so whatever their copy left behind is reaped too. The common case costs each big PUT two extra EFS metadata operations (intent create and delete), both on the movers, off the PUT request path. Updates tailscale/corp#44699 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I07bd041641d469144982481ead1417bf Migrated-from: bradfitz/go-tool-cache@984c711
…bradfitz/go-tool-cache#43) A few minor follow-ups to bradfitz/go-tool-cache#40 to improve the API and observability: * Allow callers to use different timeouts for different object sizes. * Track stats of how many async PUTs are cancelled/timed out. * Add a ctx to Shutdown's args for better separation of timeout vs shutdown behaviour. Updates tailscale/corp#45334 Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@0b25549
…-tool-cache#44) Change WithNamespaceMapping's signature to return a more granular NamespaceGrant struct, enabling read-only sessions. The session's permissions are returned in an OAuth scope so clients can skip writes if they will be rejected, and the mapping function also accepts a ctx so we can gracefully make requests to other services as part of namespace policy decisions. Updates tailscale/corp#45427 Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@0df7f03
Wiping the global namespace previously meant an O(n) delete of every blob and metadata row. Instead, let the operator bump a generation number (--global-generation=N): reads and writes of the global namespace then resolve to a fresh per-generation NamespaceID, so the namespace starts out empty without any deletion work. The previous generation's rows stop being accessed, sink to the bottom of the AccessTime LRU, and age out through the existing max-size/max-age eviction. Blobs are content-addressed and shared, so content rewritten into the new generation reuses existing blobs; only blobs unique to the old generation are eventually orphaned and swept. Generations 0 and 1 both mean the original NamespaceID 0, so existing deployments are unaffected. Later generations live under a synthetic "!global:N" row in the Namespaces table; the "!" prefix is outside the character set allowed for mapping-provided namespaces, so they can't collide. Rolling the flag back to an earlier generation makes that generation's surviving entries visible again. Updates tailscale/corp#45834 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Change-Id: I8f3d92c47a1e5b06d9ce24f8a71b3c5e4d20a9f6 Migrated-from: bradfitz/go-tool-cache@5d673d2
…ache#46) Adds support for ES384 alg and allow issuers to omit "alg" in the JWKS endpoint. Also ensure we fail loudly and early if an issuer is specified but no usable signing keys are found on the JWKS endpoint, as we previously rejected all tokens from AWS STS outbound identity federation and didn't error until clients tried to auth. Updates tailscale/corp#45427 Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@f900c6d
) Our dependabot PRs have been failing their token exchange with gocached since rolling out namespaces because their actor has square brackets. Allow those, add a test, and unconditionally log errors due to namespace validation failure as that always indicates a bug in the namespace mapping function, which should instead return an error if it intends for a token to be rejected. Updates tailscale/corp#46834 Signed-off-by: Tom Proctor <tomhjp@users.noreply.github.com> Migrated-from: bradfitz/go-tool-cache@52b050f
Cache entries can be linked binaries (such as test binaries uploaded by gotst under cmd/go's exact link action IDs) that cmd/go then executes directly out of the cache directory on a link-action cache hit. Output files were created with os.CreateTemp's mode 0600, so any such execution failed with "fork/exec ...: permission denied". Write output files with mode 0755 instead, matching what cmd/go's own DiskCache does for cached executables. Action index files stay 0644. Updates tailscale/corp#46045 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@a1c7321
When the put queue fills, the existing metrics say that it's full but not why: whether the movers copying spooled blobs into the main blob directory are behind, or the SQLite metadata flusher is, and how far pending is from a cap that only the source code knows. Add histograms for the two background stages, the intent-write-plus- copy of each spooled blob into the main blob directory and each metadata batch transaction (including the wait for the SQLite write lock), gauges for the backlog feeding each stage, and a gauge for the pending count cap so utilization can be charted as a ratio. No behavior change. The copy histogram in particular is the input any tuning of mover concurrency has to be based on: it shows the main directory's per-copy latency and how it moves under load. Updates tailscale/corp#47471 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@471b04a
… missing Each spooled blob's copy into the main blob directory did an MkdirAll on its two-hex-char shard directory, one round trip to what may be a network filesystem, to create a directory that has existed since the first blob landed there. Rename first and only MkdirAll and retry when the rename fails with not-exist. The common case costs one operation, and there is no remembered state to go stale: if the shard directories are wiped out from under the server, the next copy recreates them. Updates tailscale/corp#47471 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@f5ba3c2
During CI write bursts the put queue on ci-gocached-aws-1 sits at its caps (8192 entries or 16 GiB) for tens of minutes at a time. Every PUT, including zero-byte ones, then waits FIFO on the same weighted semaphore behind multi-megabyte blobs whose drain is bounded by the movers copying them onto the NFS-mounted main blob directory. cigocacher gives small PUTs only 5s, so CI logs fill with error PUT /<action>/e3b0c442...: context deadline exceeded (the empty-string SHA-256) and those objects never reach the shared cache. Measured during one such burst: 47 PUTs/s blocked, 5/s canceled by the client, mean server time per PUT 2.4s to 2.9s. Split admission into two lanes. PUTs of at most smallObjectSize (inline, including empty) reserve a slot in their own semaphore and never touch the spooled lane's byte or count caps; their bytes are settled in memory as soon as the body is read and they only need the metadata flusher. Larger PUTs use the existing byte and count caps. The reservation becomes a small struct so retire and the handler's early-return path release the right lane. The inline lane's cap is a 64 MiB memory budget divided by smallObjectSize. Split the put_queue_blocked counter by which cap bound (inline, spool bytes, spool count) and add gauges for the inline lane's pending count and cap, so a full queue says which constraint is binding instead of just that it's full. Updates tailscale/corp#47471 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@2550b52
The put queue's caps were two bare constants: 16 GiB of pending spooled bytes and 8192 pending entries. Neither says what it bounds or how to know when it's wrong. The byte cap bounds spool disk usage, a property of the host: the spool shares a disk with the hot tier when tiering is enabled. Make it a server option (WithPutSpoolCapacity, --put-spool-gb) defaulting to one eighth of the spool filesystem's size via statfs, or 16 GiB where that can't be read (Windows). Export it as gocached_put_queue_pending_bytes_cap so pending bytes can be charted as a ratio. The count cap bounds memory: a pending spooled entry holds no blob bytes, only its pendingPut, map slot, and channel buffer slots. Write it as a 32 MiB budget at a 512-byte per-entry allowance, which comes to 65536 entries. The old 8192 was low enough that a burst of small blobs hit it in minutes while pending bytes sat well under their cap. Updates tailscale/corp#47471 Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com> Migrated-from: bradfitz/go-tool-cache@c660171
This merges the history of github.com/bradfitz/go-tool-cache's main branch, through bradfitz/go-tool-cache@c660171c910c, rewritten for its new home in tb. In each rewritten commit: * the gocached server package moved to gocache/gocached, its logger and internal/jwt packages to gocache/logger and gocache/internal/jwt, and the other packages (cachers, cacheproc, wire) to gocache/*. * cmd/gocached stayed at cmd/gocached. * cmd/go-cacher, cmd/go-cacher-server, LICENSE, and the GitHub CI workflow were removed, and commits that only touched them dropped. * import paths and the go.mod module line use github.com/tailscale/tb. * bare issue references were qualified as bradfitz/go-tool-cache#N, and a Migrated-from trailer names the original commit. This merge commit itself only combines go.mod and go.sum. Updates #cleanup Signed-off-by: Brad Fitzpatrick <bradfitz@tailscale.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This moves over gocached from https://github.com/bradfitz/go-tool-cache
OutputIDfor compatibility with Go1.24