Skip to content

Cache backend FakeTensor dispatches during ExportPass replay (#21700) - #21700

Open
apullin wants to merge 2 commits into
pytorch:mainfrom
apullin:export-D115374097
Open

Cache backend FakeTensor dispatches during ExportPass replay (#21700)#21700
apullin wants to merge 2 commits into
pytorch:mainfrom
apullin:export-D115374097

Conversation

@apullin

@apullin apullin commented Aug 9, 2026

Copy link
Copy Markdown
Contributor

Summary:

Split from D97528110 so the FakeTensor cache extension can be reviewed independently from cold-node fast-copy.

FakeTensorMode normally caches only aten, prim, and prims operations. ExportPass replay repeatedly dispatches deterministic backend operators from quantized_decomposed, tosa, and cortex_m, but PyTorch stores non-builtin negative cache entries for them. This change temporarily extends torch._library.utils.is_builtin while ExportPass runs, evicts only matching non-builtin bypass entries from the class-global and active ShapeEnv caches, and restores the exact original predicate in finally. Positive entries remain available across successive passes, which is the source of the speedup.

This intentionally remains a narrow process-global monkeypatch. Current FakeTensorMode has no per-instance cache-policy hook, and its non-symbolic cache is class-global. The reentrant lock serializes participating ExportPass contexts, but unrelated threads can observe the extended predicate while the context is active. Fully isolated behavior requires an upstream policy hook plus cache-policy identity or per-mode cache storage; keeping this in a separate diff makes that tradeoff explicit.

A/B benchmark against the fast-copy-only parent on the same host, each side run twice:

  • CombinedControl U55 lowering: 124.343s / 126.491s with this change vs 148.465s / 149.726s in the parent. Warm speedup: 15.5%; two-run mean speedup: 15.9%.
  • Synthetic U55 suite: 80.915s / 80.905s vs 81.238s / 82.371s. The 1.1% mean difference is within run-to-run noise; CombinedControl is the representative workload.

Differential Revision: D115374097

@pytorch-bot

pytorch-bot Bot commented Aug 9, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21700

Note: Links to docs will display an error until the docs builds have been completed.

❌ 7 New Failures

As of commit 173c7bd with merge base fb5eedc (image):

NEW FAILURES - The following jobs have failed:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 9, 2026
@github-actions github-actions Bot added ciflow/trunk module: arm Issues related to arm backend labels Aug 9, 2026
@meta-codesync

meta-codesync Bot commented Aug 9, 2026

Copy link
Copy Markdown
Contributor

@apullin has exported this pull request. If you are a Meta employee, you can view the originating Diff in D115374097.

@github-actions

github-actions Bot commented Aug 9, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@meta-codesync meta-codesync Bot changed the title Cache backend FakeTensor dispatches during ExportPass replay Cache backend FakeTensor dispatches during ExportPass replay (#21700) Aug 9, 2026
apullin added a commit to apullin/executorch that referenced this pull request Aug 9, 2026
…#21700)

Summary:

Split from D97528110 so the FakeTensor cache extension can be reviewed independently from cold-node fast-copy.

`FakeTensorMode` normally caches only `aten`, `prim`, and `prims` operations. ExportPass replay repeatedly dispatches deterministic backend operators from `quantized_decomposed`, `tosa`, and `cortex_m`, but PyTorch stores `non-builtin` negative cache entries for them. This change temporarily extends `torch._library.utils.is_builtin` while ExportPass runs, evicts only matching `non-builtin` bypass entries from the class-global and active ShapeEnv caches, and restores the exact original predicate in `finally`. Positive entries remain available across successive passes, which is the source of the speedup.

This intentionally remains a narrow process-global monkeypatch. Current `FakeTensorMode` has no per-instance cache-policy hook, and its non-symbolic cache is class-global. The reentrant lock serializes participating ExportPass contexts, but unrelated threads can observe the extended predicate while the context is active. Fully isolated behavior requires an upstream policy hook plus cache-policy identity or per-mode cache storage; keeping this in a separate diff makes that tradeoff explicit.

A/B benchmark against the fast-copy-only parent on the same host, each side run twice:
- CombinedControl U55 lowering: 124.343s / 126.491s with this change vs 148.465s / 149.726s in the parent. Warm speedup: 15.5%; two-run mean speedup: 15.9%.
- Synthetic U55 suite: 80.915s / 80.905s vs 81.238s / 82.371s. The 1.1% mean difference is within run-to-run noise; CombinedControl is the representative workload.

Differential Revision: D115374097
@apullin
apullin force-pushed the export-D115374097 branch from 4e3427a to 639de3e Compare August 9, 2026 23:53
apullin added 2 commits August 9, 2026 17:28
Summary:
After ARM pass-skipping and targeted-op ownership landed separately in D106781989, this diff optimizes the remaining `ExportPass` replay cost for passes that declare `target_ops` or `targeted_ops`.

Cold operator nodes are copied with `graph.node_copy` instead of being re-dispatched through FakeTensor. Old-to-new node remapping preserves dependencies and `get_attr` values. Fast-copy is disabled when `call()` is overridden, for exact convolution or linear targets, or after a hot node changes nested tensor metadata.

The fast path preflights every input before mutating the new graph or module tree, so a remapping fallback cannot leave orphaned nodes, attributes, or remap entries. An explicitly empty `targeted_ops` remains authoritative instead of falling back to legacy `target_ops`.

Nested ARM control-flow submodules use the established `ArmPass.should_run_pass()` contract. This diff does not duplicate pass auto-discovery or generic skipping logic owned by D106781989.

Per review from `jrstevens`, the process-global FakeTensor cache extension is now isolated in child D115374097 so its monkeypatching design and incremental performance can be reviewed independently.

A/B benchmark against the parent revision on the same host, each side run twice:
- CombinedControl U55 lowering: 148.465s / 149.726s with fast-copy vs 165.886s / 171.181s before it. Warm speedup: 12.5%; two-run mean speedup: 11.5%.
- Synthetic U55 suite: 81.238s / 82.371s vs 83.816s / 80.802s. The difference is within run-to-run noise; the large model is the representative workload.

Differential Revision: D97528110
…#21700)

Summary:

Split from D97528110 so the FakeTensor cache extension can be reviewed independently from cold-node fast-copy.

`FakeTensorMode` normally caches only `aten`, `prim`, and `prims` operations. ExportPass replay repeatedly dispatches deterministic backend operators from `quantized_decomposed`, `tosa`, and `cortex_m`, but PyTorch stores `non-builtin` negative cache entries for them. This change temporarily extends `torch._library.utils.is_builtin` while ExportPass runs, evicts only matching `non-builtin` bypass entries from the class-global and active ShapeEnv caches, and restores the exact original predicate in `finally`. Positive entries remain available across successive passes, which is the source of the speedup.

This intentionally remains a narrow process-global monkeypatch. Current `FakeTensorMode` has no per-instance cache-policy hook, and its non-symbolic cache is class-global. The reentrant lock serializes participating ExportPass contexts, but unrelated threads can observe the extended predicate while the context is active. Fully isolated behavior requires an upstream policy hook plus cache-policy identity or per-mode cache storage; keeping this in a separate diff makes that tradeoff explicit.

A/B benchmark against the fast-copy-only parent on the same host, each side run twice:
- CombinedControl U55 lowering: 124.343s / 126.491s with this change vs 148.465s / 149.726s in the parent. Warm speedup: 15.5%; two-run mean speedup: 15.9%.
- Synthetic U55 suite: 80.915s / 80.905s vs 81.238s / 82.371s. The 1.1% mean difference is within run-to-run noise; CombinedControl is the representative workload.

Differential Revision: D115374097
@apullin
apullin force-pushed the export-D115374097 branch from 639de3e to 173c7bd Compare August 10, 2026 00:28
@apullin

apullin commented Aug 10, 2026

Copy link
Copy Markdown
Contributor Author

Hold on any reviewing here. There iis a way to do this without any monkeypatching, but it requires an extension to torch itself. That diff is written, will be submitted RFC to torch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ciflow/trunk CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. meta-exported module: arm Issues related to arm backend

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant