Skip to content

Fix test-webgpu-native: export the missing rope_hf_dynamic_sequence fixture - #21690

Open
shoumikhin wants to merge 1 commit into
mainfrom
shoumikhin/fix-webgpu-rope-sequence-fixture
Open

Fix test-webgpu-native: export the missing rope_hf_dynamic_sequence fixture#21690
shoumikhin wants to merge 1 commit into
mainfrom
shoumikhin/fix-webgpu-rope-sequence-fixture

Conversation

@shoumikhin

@shoumikhin shoumikhin commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

What is broken

test-webgpu-native is red on main. The failing test is WebGPUNative.RopeHfDynamicSequenceReusedGraph and it fails with Error::AccessFailed, which is what ExecuTorch returns when it cannot open a .pte file.

Why it is broken

The C++ test loads a model file named rope_hf_dynamic_sequence.pte out of the directory given by the WEBGPU_TEST_ROPE_HF_DIR environment variable.

The CI script that prepares the test fixtures, backends/webgpu/scripts/test_webgpu_native_ci.sh, only called export_rope_hf_dynamic(...). It never called export_rope_hf_dynamic_sequence(...), so rope_hf_dynamic_sequence.pte was never written. The test then tried to open a file that did not exist.

The Python helper export_rope_hf_dynamic_sequence already exists in backends/webgpu/test/ops/test_rope_hf.py. It was just never called from CI.

The contract test that is supposed to catch exactly this kind of gap only checked for the other fixture, so nothing failed at lint time either.

The fix

  1. Call export_rope_hf_dynamic_sequence('${ROPE_HF_DIR}') in the CI script, right next to the existing export_rope_hf_dynamic call.
  2. Add a require_file check for each of the two .pte files. If a fixture ever goes missing again, the script now stops early with a readable message instead of failing deep inside a C++ test with a numeric error code.
  3. Extend test_native_ci_contract.py so it asserts that both exports and both require_file checks are present.

No test logic and no backend code changed. This only produces a file the test always expected to find.

How this was verified

test-webgpu-native was dispatched on this branch. In that run
WebGPUNative.RopeHfDynamicSequenceReusedGraph passes and the script prints
=== WebGPU native tests on Dawn: all run targets passed ===, so the fixture is
now produced and the test that is red on main is green.

On its own this branch does not make the whole job green: the script then reaches
a separate op-test stage that fails on an unrelated problem in the op-test
generator, which #21697 fixes. The two were therefore also tested together, on a
branch holding both changes, and that run passes end to end:

[       OK ] WebGPUNative.RopeHfDynamicSequenceReusedGraph (29 ms)
=== WebGPU native tests on Dawn: all run targets passed ===
Generated 390 cases -> /tmp/webgpu_op_tests/manifest.json
[==========] 391 tests from 88 test suites ran. (37413 ms total)
[  PASSED  ] 391 tests.
=== WebGPU op-test framework on Dawn: passed ===

So this PR plus #21697 turn test-webgpu-native green. Either one alone is not
enough, and they are separate problems, so they are separate changes.

Copilot AI lite review requested due to automatic review settings August 8, 2026 23:01
@pytorch-bot

pytorch-bot Bot commented Aug 8, 2026

Copy link
Copy Markdown

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/21690

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 Cancelled Job

As of commit 2df108d with merge base fb5eedc (image):

CANCELLED JOB - The following job was cancelled. Please retry:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed. label Aug 8, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@github-actions

github-actions Bot commented Aug 8, 2026

Copy link
Copy Markdown

This PR needs a release notes: label

If your change should be included in the release notes (i.e. would users of this library care about this change?), please use a label starting with release notes:. This helps us keep track and include your important work in the next release notes.

To add a label, you can comment to pytorchbot, for example
@pytorchbot label "release notes: none"

For more information, see
https://github.com/pytorch/pytorch/wiki/PyTorch-AutoLabel-Bot#why-categorize-for-release-notes-and-how-does-it-work.

@shoumikhin

Copy link
Copy Markdown
Contributor Author

How this was tested

test-webgpu-native.yml has no pull_request trigger, so it never runs on a pull request. I ran it by hand against this branch instead: https://github.com/pytorch/executorch/actions/runs/31283613850

  • On main today, WebGPUNative.RopeHfDynamicSequenceReusedGraph fails in 0 ms with Error::AccessFailed, which is what Module::load_forward() returns when it cannot open the .pte file.
  • With this change, that same test reports [ OK ] WebGPUNative.RopeHfDynamicSequenceReusedGraph (28 ms), and every gtest suite in the job passes.

The job is still red, for an unrelated reason

Once the C++ tests pass, the job continues into op test generation and stops there:

[Vulkan Partitioner] Due to [op args not supported], skipping
  aten.convolution.default([2, 8, 16, 16] -> [2, 16, 16, 16])
WARNING: No Vulkan subgraphs can be partitioned!
RuntimeError: conv2d/gemm_batched produced NO VulkanBackend delegate

This pull request does not cause that and does not claim to fix it. The gemm_batched case and the rule that a missing delegate is a hard error both live in backends/webgpu/test/op_tests/ on main, and neither file is touched here. The only reason nobody has seen this before is that the job used to die at the C++ tests first, so it never got this far. Repairing the fixture is what makes the second problem visible.

In other words this is the first of two fixes. It stands on its own and is worth landing on its own, and the batch of 2 conv2d case needs a separate follow up from someone who owns the Vulkan partitioner.

@shoumikhin
shoumikhin force-pushed the shoumikhin/fix-webgpu-rope-sequence-fixture branch from 3f99b6c to 2df108d Compare August 10, 2026 05:51
Copilot AI review requested due to automatic review settings August 10, 2026 05:51

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Facebook bot. Authors need to sign the CLA before a PR can be reviewed.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants