Skip to content

CRC / Copy offloading to Intel Data Streaming Accelerator - #470

Open
byrnedj wants to merge 6 commits into
facebook:mainfrom
byrnedj:crc_upstream
Open

byrnedj wants to merge 6 commits into
facebook:mainfrom
byrnedj:crc_upstream

Conversation

@byrnedj

@byrnedj byrnedj commented Jul 15, 2026 •

Copy link
Copy Markdown
Contributor

Adds optional offload of Navy's data checksumming to Intel DSA (Data
Streaming Accelerator), using the DTO library.
On the write path, the value copy into the region buffer and its CRC are
fused into a single DSA Memory Copy with CRC Generation descriptor that
the calling fiber/thread submits before yielding; on the read path, all
verification sites (lookup, reclaim, cleanup, reinsertion, random-alloc)
can verify on the accelerator. BigHash bucket checksums offload with the
bloom-filter rebuild as the overlap work. Note, we need to use CRC32C for
compatible CRC polynomial for DSA.

Also adds DTO transparent usage to cache insertions in cachebench.

We use navyChecksumOffload(+checksumOffloadMinSize, default 4096) and a BUILD_WITH_DTO` cmake option, with a runtime
DSA-vs-CPU parity self-check at engine creation that falls back to
software on mismatch.

Benchmarks

BigCache production trace replay (12M ops, ~48KB avg objects,
BlockCache-only, fiber scheduler):

no checksum software CRC DSA offload
Insert p50 / p99.9 17 / 262 µs 41 / 628 µs 16 / 451 µs
Lookup p99 9.3 ms 10.6 ms 8.6 ms

ucache_bench (DCPerf, hybrid DRAM+Navy, small objects): CPU
−2.9% (fibers) / +0.7% (thread pool), P99.9 −14% / −21%, equal QPS.

How to test

# unit tests (needs an enabled DSA shared WQ; falls back cleanly without)
build-cachelib/navy/navy-test-ChecksumOffloadTest

# end-to-end harness: builds with DTO, configures DSA WQs, runs on socket 0
./run_dsa_cachebench.sh --workload bigcache --offload off --fibers on   # B: software
./run_dsa_cachebench.sh --workload bigcache --offload on  --fibers on   # C: DSA
./run_dsa_cachebench.sh --workload cdn --offload on --intercept on      # + transparent memcpy intercept

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Jul 15, 2026
@meta-codesync

meta-codesync Bot commented Jul 21, 2026

Copy link
Copy Markdown
Contributor

@pbhandar2 has imported this pull request. If you are a Meta employee, you can view this in D113115379.

Squash of PR facebook#470 (3153247, 695f774, 7874531) rebased onto main.

Adds optional offload of Navy's data checksumming to Intel DSA through
DTO. On the write path the value copy into the region buffer and its CRC
are one fused "Memory Copy with CRC Generation" descriptor submitted before
the fiber yields; on the read path every verification site (lookup,
reclaim, cleanup, reinsertion) can verify on the accelerator. BigHash
bucket checksums offload with the bloom-filter rebuild as overlap work.
CRC-32C is used so the CPU and DSA polynomials agree.

Config: BlockCacheConfig::setChecksumOffload(enable, minSize), NavyConfig
plumbing, cachebench navyChecksumOffload / navyChecksumOffloadMinSize.
Build: BUILD_WITH_DTO (installed DTO package or a pinned fetch).
…hecks

ChecksumOffload grows plain-copy offloads next to the fused copy+CRC:
copyWithOffload() (one Memory Move, fiber-yielding wait) and
copyLargeWithOffload() (one descriptor or a Batch of `parts` moves;
a refused batch degrades to one descriptor, a device failure to memcpy)
with LargeCopyWait::{kSpinOrYield,kSleep}. The wait loop yields to
runnable fibers first, blocks in DTO when DTO_WAIT_METHOD=aggregator
makes dto_wait_blocks() true, and otherwise pause-polls or nanosleeps.

Self-checks now tell a working accelerator from a silently failing one:
AsyncChecksumOp::completedOnDevice() reports whether the last wait
completed on DSA, and checksumOffloadSelfCheck() / copyOffloadSelfCheck()
require it before checking parity - previously the fallback path
recomputed the right answer and the self-check passed on 100% CPU
fallback (e.g. a work queue without block-on-fault rejecting every
descriptor with INVALID_FLAGS). copyOffloadSelfCheck() no longer demands
the Batch opcode, which DSA 1.0 lacks and the flush path does not use.

Device failures redone on the CPU are counted (CopyLargeWaitStats::
fallbacks, getChecksumWaitStats(), getCopyOutWaitStats()). dtoWaitBlocks()
no longer caches DTO state before DTO initialised.

DTO pin moves to byrnedj/DTO@e683bdc (branch cachelib-navy), which has the
batch/memmove/blocking-wait entry points and only sets the BOF descriptor
flag when the work queue allows it.

Tests: CompletedOnDeviceIsTruthful, CopyLargeDegradesWithoutBatch,
WaitStatsExposeFallbacks; SelfCheck skips without a work queue instead of
passing vacuously.
…cache-control knobs

RegionManager::flushBuffer copies the whole region buffer into a write
buffer before deviceWrite. With BlockCacheConfig::setFlushCopyOffload(true)
that copy is one DSA Memory Move (copyLargeWithOffload, kSleep wait), gated
by copyOffloadSelfCheck() at construction; the write buffers come from a
pool sized to the most flushes that can be in flight (max(2 x workers,
numInMemBuffers)) instead of one 16 MiB allocation per flush.

The in-memory region buffers, and any write buffer the pool has to
allocate while the offload is on, are populated (MADV_POPULATE_WRITE) at
allocation. A DSA descriptor that first-touches an unpopulated page stalls
on an IOMMU page request - a kernel round trip per 4 KiB page that also
holds up the other descriptors on that engine. With 200 in-memory buffers
from fresh mmaps and a pool that overflowed under flush bursts, a 2.16M-op
BigCache smoke took 1.28M page requests and ~400 us of accelerator wait per
checksum op (152 s user vs 53 s); populated, 17 page requests and 4 us/op
(53.7 s), on stock glibc without huge pages.
On a 12M-op BigCache replay this removes the memmove that was 3.8% of
process CPU: -7% user CPU on top of the checksum offload (-21.8% vs
software CRC, -4.4% vs no checksum at all), insert p50 unchanged at 14 us.

BlockCache: writeEntry writes the entry descriptor and key before
submitting the DSA copy of the value into the same slot, so the CPU and
the device no longer write the slot's last cache line concurrently.
setChecksumOffloadReadMinSize() gives read-side verification its own gate
(0 = follow the write gate, UINT32_MAX = verify on the CPU) and
setChecksumOffloadCacheControl() controls the cache-control hint on the
fused write. Measured on the same replay: read-side offload is worth 5.7%
user CPU and cache control is neutral, so both defaults stay.

Counters (RegionManager::getCounters): navy_bc_csum_wait_*,
navy_bc_csum_device_fallbacks, navy_bc_copyout_*,
navy_bc_flush_copy_{offloaded,fallbacks,device_fallbacks,...},
navy_bc_flush_writebuf_pooled. A non-zero *_device_fallbacks with the
offload "active" means descriptors are being rejected.
NvmCache::Config::copyOutOffload (+ copyOutOffloadMinSize, default 64 KiB)
copies a flash hit's payload from the Navy read buffer into the DRAM item
with a DSA Memory Move (navy::copyWithOffload) on the paths where Navy has
already verified the value; construction runs copyOffloadSelfCheck() and
falls back to memcpy when DSA is unusable. Counters nvm.copy_out_offloaded,
nvm.copy_out_offloaded_bytes, nvm.copy_out_fallbacks. Used by a CDN cache
tier serving large objects; lookup latency is unchanged, the saving is CPU.
Config: navyBlockCacheFlushCopyOffload, navyBlockCacheDirectFlush,
navyChecksumOffloadReadMinSize, navyChecksumOffloadCacheControl,
copy-out offload plumbing. run_dsa_cachebench.sh: DTO_LIB_DIR ahead of the
system libdto, DTO_WAIT_METHOD, fd soft limit for the Navy thread counts,
watchdog knobs, memory as well as CPU bound to the DSA socket
(numactl -N 0 -m 0), printNvmCounters on so the offload counters reach
the log, and a checksum-error grep that does not count the
navy_bc_*_checksum_errors counter names as errors.
Submitting and polling a DSA descriptor costs a few microseconds whatever
the size, about what the CPU needs to CRC 16-32 KiB. Measured on
cachebench's cdn hit-ratio workload (values p50 13 KB / p90 78 KB, 3 reps,
user CPU vs software CRC): 4 KiB gate -17.6%, 16 KiB -21.5%, 32 KiB -23.0%
(within rep noise of 16 KiB), 64 KiB -19.3%. Values under ~16 KiB cost
more to offload than to checksum. Raise the default from 4096 to 16384 in
BlockCache::Config, NavyConfig, cachebench and the harness; a config that
sets the gate explicitly is unchanged.
@facebook-github-tools

Copy link
Copy Markdown

@byrnedj has updated the pull request. You must reimport the pull request before landing.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant