Conversation
Contributor
|
@pbhandar2 has imported this pull request. If you are a Meta employee, you can view this in D113115379. |
Squash of PR facebook#470 (3153247, 695f774, 7874531) rebased onto main. Adds optional offload of Navy's data checksumming to Intel DSA through DTO. On the write path the value copy into the region buffer and its CRC are one fused "Memory Copy with CRC Generation" descriptor submitted before the fiber yields; on the read path every verification site (lookup, reclaim, cleanup, reinsertion) can verify on the accelerator. BigHash bucket checksums offload with the bloom-filter rebuild as overlap work. CRC-32C is used so the CPU and DSA polynomials agree. Config: BlockCacheConfig::setChecksumOffload(enable, minSize), NavyConfig plumbing, cachebench navyChecksumOffload / navyChecksumOffloadMinSize. Build: BUILD_WITH_DTO (installed DTO package or a pinned fetch).
…hecks
ChecksumOffload grows plain-copy offloads next to the fused copy+CRC:
copyWithOffload() (one Memory Move, fiber-yielding wait) and
copyLargeWithOffload() (one descriptor or a Batch of `parts` moves;
a refused batch degrades to one descriptor, a device failure to memcpy)
with LargeCopyWait::{kSpinOrYield,kSleep}. The wait loop yields to
runnable fibers first, blocks in DTO when DTO_WAIT_METHOD=aggregator
makes dto_wait_blocks() true, and otherwise pause-polls or nanosleeps.
Self-checks now tell a working accelerator from a silently failing one:
AsyncChecksumOp::completedOnDevice() reports whether the last wait
completed on DSA, and checksumOffloadSelfCheck() / copyOffloadSelfCheck()
require it before checking parity - previously the fallback path
recomputed the right answer and the self-check passed on 100% CPU
fallback (e.g. a work queue without block-on-fault rejecting every
descriptor with INVALID_FLAGS). copyOffloadSelfCheck() no longer demands
the Batch opcode, which DSA 1.0 lacks and the flush path does not use.
Device failures redone on the CPU are counted (CopyLargeWaitStats::
fallbacks, getChecksumWaitStats(), getCopyOutWaitStats()). dtoWaitBlocks()
no longer caches DTO state before DTO initialised.
DTO pin moves to byrnedj/DTO@e683bdc (branch cachelib-navy), which has the
batch/memmove/blocking-wait entry points and only sets the BOF descriptor
flag when the work queue allows it.
Tests: CompletedOnDeviceIsTruthful, CopyLargeDegradesWithoutBatch,
WaitStatsExposeFallbacks; SelfCheck skips without a work queue instead of
passing vacuously.
…cache-control knobs
RegionManager::flushBuffer copies the whole region buffer into a write
buffer before deviceWrite. With BlockCacheConfig::setFlushCopyOffload(true)
that copy is one DSA Memory Move (copyLargeWithOffload, kSleep wait), gated
by copyOffloadSelfCheck() at construction; the write buffers come from a
pool sized to the most flushes that can be in flight (max(2 x workers,
numInMemBuffers)) instead of one 16 MiB allocation per flush.
The in-memory region buffers, and any write buffer the pool has to
allocate while the offload is on, are populated (MADV_POPULATE_WRITE) at
allocation. A DSA descriptor that first-touches an unpopulated page stalls
on an IOMMU page request - a kernel round trip per 4 KiB page that also
holds up the other descriptors on that engine. With 200 in-memory buffers
from fresh mmaps and a pool that overflowed under flush bursts, a 2.16M-op
BigCache smoke took 1.28M page requests and ~400 us of accelerator wait per
checksum op (152 s user vs 53 s); populated, 17 page requests and 4 us/op
(53.7 s), on stock glibc without huge pages.
On a 12M-op BigCache replay this removes the memmove that was 3.8% of
process CPU: -7% user CPU on top of the checksum offload (-21.8% vs
software CRC, -4.4% vs no checksum at all), insert p50 unchanged at 14 us.
BlockCache: writeEntry writes the entry descriptor and key before
submitting the DSA copy of the value into the same slot, so the CPU and
the device no longer write the slot's last cache line concurrently.
setChecksumOffloadReadMinSize() gives read-side verification its own gate
(0 = follow the write gate, UINT32_MAX = verify on the CPU) and
setChecksumOffloadCacheControl() controls the cache-control hint on the
fused write. Measured on the same replay: read-side offload is worth 5.7%
user CPU and cache control is neutral, so both defaults stay.
Counters (RegionManager::getCounters): navy_bc_csum_wait_*,
navy_bc_csum_device_fallbacks, navy_bc_copyout_*,
navy_bc_flush_copy_{offloaded,fallbacks,device_fallbacks,...},
navy_bc_flush_writebuf_pooled. A non-zero *_device_fallbacks with the
offload "active" means descriptors are being rejected.
NvmCache::Config::copyOutOffload (+ copyOutOffloadMinSize, default 64 KiB) copies a flash hit's payload from the Navy read buffer into the DRAM item with a DSA Memory Move (navy::copyWithOffload) on the paths where Navy has already verified the value; construction runs copyOffloadSelfCheck() and falls back to memcpy when DSA is unusable. Counters nvm.copy_out_offloaded, nvm.copy_out_offloaded_bytes, nvm.copy_out_fallbacks. Used by a CDN cache tier serving large objects; lookup latency is unchanged, the saving is CPU.
Config: navyBlockCacheFlushCopyOffload, navyBlockCacheDirectFlush, navyChecksumOffloadReadMinSize, navyChecksumOffloadCacheControl, copy-out offload plumbing. run_dsa_cachebench.sh: DTO_LIB_DIR ahead of the system libdto, DTO_WAIT_METHOD, fd soft limit for the Navy thread counts, watchdog knobs, memory as well as CPU bound to the DSA socket (numactl -N 0 -m 0), printNvmCounters on so the offload counters reach the log, and a checksum-error grep that does not count the navy_bc_*_checksum_errors counter names as errors.
Submitting and polling a DSA descriptor costs a few microseconds whatever the size, about what the CPU needs to CRC 16-32 KiB. Measured on cachebench's cdn hit-ratio workload (values p50 13 KB / p90 78 KB, 3 reps, user CPU vs software CRC): 4 KiB gate -17.6%, 16 KiB -21.5%, 32 KiB -23.0% (within rep noise of 16 KiB), 64 KiB -19.3%. Values under ~16 KiB cost more to offload than to checksum. Raise the default from 4096 to 16384 in BlockCache::Config, NavyConfig, cachebench and the harness; a config that sets the gate explicitly is unchanged.
byrnedj
force-pushed
the
crc_upstream
branch
from
September 25, 2026 16:23
7874531 to
454f95b
Compare
|
@byrnedj has updated the pull request. You must reimport the pull request before landing. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds optional offload of Navy's data checksumming to Intel DSA (Data
Streaming Accelerator), using the DTO library.
On the write path, the value copy into the region buffer and its CRC are
fused into a single DSA
Memory Copy with CRC Generationdescriptor thatthe calling fiber/thread submits before yielding; on the read path, all
verification sites (lookup, reclaim, cleanup, reinsertion, random-alloc)
can verify on the accelerator. BigHash bucket checksums offload with the
bloom-filter rebuild as the overlap work. Note, we need to use CRC32C for
compatible CRC polynomial for DSA.
Also adds DTO transparent usage to cache insertions in cachebench.
We use navyChecksumOffload
(+checksumOffloadMinSize, default 4096) and aBUILD_WITH_DTO` cmake option, with a runtimeDSA-vs-CPU parity self-check at engine creation that falls back to
software on mismatch.
Benchmarks
BigCache production trace replay (12M ops, ~48KB avg objects,
BlockCache-only, fiber scheduler):
ucache_bench (DCPerf, hybrid DRAM+Navy, small objects): CPU
−2.9% (fibers) / +0.7% (thread pool), P99.9 −14% / −21%, equal QPS.
How to test