Commits affected: 97cf80f1a8fcf6a43e1bdc7132f5bfbc61de2d3b
Run: https://github.com/pytorch/ao/actions/runs/35123695166
Summary
The test-cpu-ops (macos-14) job fails at the "build shared_kernels with ExecuTorch" step because the pip-installed executorch package no longer ships the executorch/extension/threadpool/threadpool.h header in its include dirs, so the ExecuTorch parallel backend fails to compile. This is an external-dependency/environment issue, not a regression from the triggering commit.
Failure details
- Failed job:
test-cpu-ops (macos-14) in workflow "Run Regression Tests (aarch64)".
- Failing step:
torchao/csrc/cpu - build shared_kernels with ExecuTorch (sh build_shared_kernels.sh executorch).
- The sibling
test-cpu-ops (linux.arm64.2xlarge) job was cancelled (fail-fast), not independently failed.
- Representative error:
torchao/csrc/cpu/shared_kernels/internal/parallel-executorch-impl.h:9:10: fatal error:
'executorch/extension/threadpool/threadpool.h' file not found
#include <executorch/extension/threadpool/threadpool.h>
3 warnings and 1 error generated.
- ExecuTorch itself was located successfully during CMake configure:
-- EXECUTORCH_INCLUDE_DIRS: .../site-packages/executorch/lib/cmake/executorch/../../../include; .../include/executorch/runtime/core/portable_type/c10
-- EXECUTORCH_LIBRARIES: executorch::runtime;executorch::threadpool;executorch::kernels_optimized;...
i.e. the package is present, but the extension/threadpool/threadpool.h header is absent from those include directories.
Root cause
torchao/csrc/cpu/shared_kernels/internal/parallel-executorch-impl.h unconditionally #includes <executorch/extension/threadpool/threadpool.h> when building the ExecuTorch parallel backend. The aarch64 workflow installs ExecuTorch via an unpinned pip install executorch (see .github/workflows/regression_test_aarch64.yml). A recent executorch wheel release evidently moved or stopped packaging the extension/threadpool headers under the installed include tree, so the compile fails even though find_package(ExecuTorch) succeeds and executorch::threadpool is listed among the libraries.
Evidence this is an infra/dependency issue rather than a code regression:
- The triggering commit
97cf80f ("Enable JSON serialization/deserialization of pybind11-backed enums") is a Python-only change and touches nothing in torchao/csrc/cpu or any C++/build path.
- The last green run of this workflow on
main was 2026-09-11 (9e5ea7fde7); the two most recent runs (97cf80f and parent-range cfeff3cd5f) both fail with the identical threadpool.h error at the same step. Deterministic across consecutive commits ⇒ not a per-run flake.
- The failing header/include predates this commit; no torchao source changed to cause it. The variable is the externally-pulled executorch wheel.
Classification
infra failure — the build breaks because the externally-installed (unpinned) executorch package no longer provides executorch/extension/threadpool/threadpool.h in its include path, an environment/dependency-packaging change outside torchao. Deterministic (fails identically on two consecutive main commits), so not a flaky test; unrelated to the pybind11-enum triggering commit, so not a real regression. Distinct from open issues #4864 (float8 TP compile aten.abs), #4708 (FSDP2 fp8 NaN parity), and #4663 (ROCm gfx950).
Suggested fix
- Pin the executorch version in
.github/workflows/regression_test_aarch64.yml (pip install executorch==<known-good>) so the build uses a wheel that still ships the threadpool extension headers, and bump deliberately once the include path is confirmed.
- Alternatively, update the ExecuTorch parallel backend include to match the new packaging layout (locate the current header path in the installed executorch wheel and adjust
parallel-executorch-impl.h / the include dirs added in shared_kernels/Utils.cmake accordingly), or add the executorch extension/threadpool include directory explicitly.
- Longer term, verify at configure time that the required threadpool header exists (fail early with a clear message) instead of only checking that the ExecuTorch package/libraries are present.
Commits affected:
97cf80f1a8fcf6a43e1bdc7132f5bfbc61de2d3bRun: https://github.com/pytorch/ao/actions/runs/35123695166
Summary
The
test-cpu-ops (macos-14)job fails at the "build shared_kernels with ExecuTorch" step because the pip-installedexecutorchpackage no longer ships theexecutorch/extension/threadpool/threadpool.hheader in its include dirs, so the ExecuTorch parallel backend fails to compile. This is an external-dependency/environment issue, not a regression from the triggering commit.Failure details
test-cpu-ops (macos-14)in workflow "Run Regression Tests (aarch64)".torchao/csrc/cpu - build shared_kernels with ExecuTorch(sh build_shared_kernels.sh executorch).test-cpu-ops (linux.arm64.2xlarge)job wascancelled(fail-fast), not independently failed.extension/threadpool/threadpool.hheader is absent from those include directories.Root cause
torchao/csrc/cpu/shared_kernels/internal/parallel-executorch-impl.hunconditionally#includes<executorch/extension/threadpool/threadpool.h>when building the ExecuTorch parallel backend. The aarch64 workflow installs ExecuTorch via an unpinnedpip install executorch(see.github/workflows/regression_test_aarch64.yml). A recent executorch wheel release evidently moved or stopped packaging theextension/threadpoolheaders under the installed include tree, so the compile fails even thoughfind_package(ExecuTorch)succeeds andexecutorch::threadpoolis listed among the libraries.Evidence this is an infra/dependency issue rather than a code regression:
97cf80f("Enable JSON serialization/deserialization of pybind11-backed enums") is a Python-only change and touches nothing intorchao/csrc/cpuor any C++/build path.mainwas 2026-09-11 (9e5ea7fde7); the two most recent runs (97cf80fand parent-rangecfeff3cd5f) both fail with the identicalthreadpool.herror at the same step. Deterministic across consecutive commits ⇒ not a per-run flake.Classification
infra failure — the build breaks because the externally-installed (unpinned)
executorchpackage no longer providesexecutorch/extension/threadpool/threadpool.hin its include path, an environment/dependency-packaging change outside torchao. Deterministic (fails identically on two consecutivemaincommits), so not a flaky test; unrelated to the pybind11-enum triggering commit, so not a real regression. Distinct from open issues #4864 (float8 TP compile aten.abs), #4708 (FSDP2 fp8 NaN parity), and #4663 (ROCm gfx950).Suggested fix
.github/workflows/regression_test_aarch64.yml(pip install executorch==<known-good>) so the build uses a wheel that still ships the threadpool extension headers, and bump deliberately once the include path is confirmed.parallel-executorch-impl.h/ the include dirs added inshared_kernels/Utils.cmakeaccordingly), or add the executorchextension/threadpoolinclude directory explicitly.