Skip to content

Use torch.Size to represent Tensor size instead of LongStorage - #171

Merged
apaszke merged 1 commit into
pytorch:masterfrom
colesbury:size
Oct 28, 2016
Merged

apaszke merged 1 commit into
pytorch:masterfrom
colesbury:size

Conversation

@colesbury

Copy link
Copy Markdown
Contributor

See issue #20

The torch.Size class is a tuple subclass which distinguishes sizes from
other tuples so that torch.Tensor(size) is interpreted as size instead
of data.

Comment thread torch/csrc/generic/TensorMethods.cwrap Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread torch/csrc/utils.cpp Outdated

This comment was marked as off-topic.

Comment thread torch/csrc/generic/Tensor.cpp Outdated

This comment was marked as off-topic.

Comment thread tools/cwrap/plugins/THPPlugin.py Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread torch/csrc/generic/TensorMethods.cwrap Outdated

This comment was marked as off-topic.

Comment thread torch/csrc/generic/TensorMethods.cwrap Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread torch/multiprocessing/_tensor.py Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

@colesbury
colesbury force-pushed the size branch 3 times, most recently from 333e451 to f2905a3 Compare October 26, 2016 22:29
@colesbury

Copy link
Copy Markdown
Contributor Author

@apaszke, I made stride a tuple() and no longer accept torch.LongStorages as sizes from the Python side.

I ended up combining the relevant parts of the LongArgsPlugin with THPlugin. I couldn't figure out any other way to do the type-checks for THSize* correctly, because only one plugin can provide a type-check.

@colesbury
colesbury force-pushed the size branch 2 times, most recently from f48ecfd to fd4e0d3 Compare October 26, 2016 22:51
@apaszke

apaszke commented Oct 27, 2016

Copy link
Copy Markdown
Contributor

Actually, since the ambiguity appears only in the tensor constructor, how about allowing regular tuples for specifing sizes e.g. to set? We can just add support for raw tuples to THPUtils_unpackSize.

@colesbury

Copy link
Copy Markdown
Contributor Author

Yeah, I was thinking about that. Should we also allow lists for convenience?

@apaszke

apaszke commented Oct 27, 2016

Copy link
Copy Markdown
Contributor

Hm, yeah, we could accept arbitrary sequences (see PySequence API).

@colesbury

Copy link
Copy Markdown
Contributor Author

Yeah, the one thing I'd be worried about is that Tensors are also sequences (I think)

@apaszke

apaszke commented Oct 27, 2016

Copy link
Copy Markdown
Contributor

Good point. We could additionally check that they're not tensors nor storages, but I'm not sure if it's worth it (overhead, etc.).

@colesbury
colesbury force-pushed the size branch 5 times, most recently from 20ac14e to 485b76e Compare October 27, 2016 19:49
@colesbury

Copy link
Copy Markdown
Contributor Author

@apaszke, I think this is good to go. Finally fixed all the legacy NN tests that broke b/c of this

@apaszke apaszke left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have a couple more comments. I didn't review legacy nn and tensor.py, but since the tests are passing they should be ok.

Comment thread torch/csrc/utils.cpp Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread torch/csrc/utils.cpp Outdated

This comment was marked as off-topic.

Comment thread test/test_torch.py Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread tools/cwrap/plugins/THPPlugin.py Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread tools/cwrap/plugins/THPPlugin.py Outdated

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

This comment was marked as off-topic.

Comment thread torch/csrc/Size.cpp Outdated

This comment was marked as off-topic.

@colesbury
colesbury force-pushed the size branch 2 times, most recently from e646413 to 104907d Compare October 27, 2016 22:42
Comment thread torch/csrc/utils.h Outdated

This comment was marked as off-topic.

See issue #20

The torch.Size class is a tuple subclass which distinguishes sizes from
other tuples so that torch.Tensor(size) is interpreted as size instead
of data.
@apaszke
apaszke merged commit f2d7e94 into pytorch:master Oct 28, 2016
@apaszke
apaszke deleted the size branch October 28, 2016 17:43
cpuhrsch pushed a commit to cpuhrsch/pytorch that referenced this pull request May 3, 2018
[release 3.2] Changelog file additions.
mrshenli pushed a commit to mrshenli/pytorch that referenced this pull request Apr 11, 2020
* Added Spatial Transformer Networks tutorial

* Add more code comments

* Move STN tutorial to intermediate + minor fixes

* Remove STN from beginner TOC

* Fix ToC rendering for STN tutorial

* Fix tutorial typos

* Fix all pep8, flake8 issues

* Fix rst warning
KsenijaS pushed a commit to KsenijaS/pytorch that referenced this pull request Dec 14, 2020
* mask rcnn init commit

* nit: model description

* update readme
KyleCZH pushed a commit to KyleCZH/pytorch that referenced this pull request Sep 20, 2021
Blacklisting test_wrong_return_type on all Conda 2.7s
hubertlu-tw pushed a commit to hubertlu-tw/pytorch that referenced this pull request Nov 1, 2022
pytorchmergebot pushed a commit that referenced this pull request May 12, 2023
When tensor is resized, reference array to it's sizes may become invalid. Make a copy in advance.

<details>
<summary>ASAN report</summary>

```
=================================================================
==1115867==ERROR: AddressSanitizer: heap-use-after-free on address 0x61000013d790 at pc 0x03ff8e7da360 bp 0x03fff53c83a0 sp 0x03fff53c8390
READ of size 8 at 0x61000013d790 thread T0
    #0 0x3ff8e7da35f in c10::SymInt::is_heap_allocated() const /home/user/pytorch/c10/core/SymInt.h:154
    #1 0x3ff8e7da35f in c10::SymInt::maybe_as_int() const /home/user/pytorch/c10/core/SymInt.h:215
    #2 0x3ff8e7d0a6d in c10::SymInt::sym_eq(c10::SymInt const&) const /home/user/pytorch/c10/core/SymInt.cpp:69
    #3 0x3ff7a9ab0bd in c10::SymInt::operator==(c10::SymInt const&) const /home/user/pytorch/c10/core/SymInt.h:177
    #4 0x3ff7a9aaedd in bool std::__equal<false>::equal<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-
v11/bits/stl_algobase.h:1162
    #5 0x3ff7a9aae4b in bool std::__equal_aux1<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/
stl_algobase.h:1211
    #6 0x3ff7a9aae05 in bool std::__equal_aux<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/s
tl_algobase.h:1219
    #7 0x3ff7a9aad97 in bool std::equal<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_alg
obase.h:1556
    #8 0x3ff4b23c771 in c10::ArrayRef<c10::SymInt>::equals(c10::ArrayRef<c10::SymInt>) const /home/user/pytorch/c10/util/ArrayRef.h:188
    #9 0x3ff4cb91bc1 in bool c10::operator!=<c10::SymInt>(c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>) /home/user/pytorch/c10/util/ArrayRef.h:341
    #10 0x3ff6d1b57ff in torch::ADInplaceOrView::resize_(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/torch/csrc/autograd/Variab
leTypeManual.cpp:408
    #11 0x3ff6d1e59c7 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c1
0::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>
> >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
    #12 0x3ff6d1e59c7 in c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10:
:ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::Sy
mInt>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::Disp
atchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:480
    #13 0x3ff51ca5129 in at::Tensor const& c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(void*, c10::OperatorKernel*,
c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>&&, c10::optional<c10::MemoryFormat>&&) /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
    #14 0x3ff51ca6e8f in at::Tensor const& c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::OperatorHandle const&, c10::D
ispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:90
    #15 0x3ff51ca6e8f in at::Tensor const& c10::Dispatcher::redispatch<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Ten
sor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)
const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:656
    #16 0x3ff5182006b in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::redispatch(c10::DispatchKeySet, at::Tensor const&, c
10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:492
    #17 0x3ff5182006b in at::_ops::resize_::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) aten/src/ATen/Operators_4.cpp:2144
    #18 0x3ff6d1d5e07 in at::redispatch::resize__symint(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) aten/src/ATen/RedispatchFunctions.h:2847
    #19 0x3ff6d1bbb67 in torch::autograd::VariableType::(anonymous namespace)::resize_(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pyto
rch/torch/csrc/autograd/VariableTypeManual.cpp:243
    #20 0x3ff6d1bd197 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c1
0::MemoryFormat>), &torch::autograd::VariableType::(anonymous namespace)::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10
::optional<c10::MemoryFormat> > >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFu
nctionIntoFunctor.h:13
    #21 0x3ff6d1bd197 in c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10:
:ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>), &torch::autograd::VariableType::(anonymous namespace)::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor
 const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c
10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor
.h:480
    #22 0x3ff51ca5129 in at::Tensor const& c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(void*, c10::OperatorKernel*,
c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>&&, c10::optional<c10::MemoryFormat>&&) /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
    #23 0x3ff5181ead1 in at::Tensor const& c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::OperatorHandle const&, c10::D
ispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:90
    #24 0x3ff5181ead1 in at::Tensor const& c10::Dispatcher::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Tensor co
nst& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/at
en/src/ATen/core/dispatch/Dispatcher.h:639
    #25 0x3ff5181ead1 in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(at::Tensor const&, c10::ArrayRef<c10::SymInt>,
c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:487
    #26 0x3ff5181ead1 in at::_ops::resize_::call(at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) aten/src/ATen/Operators_4.cpp:2137
    #27 0x3ff79b44fcf in at::Tensor::resize__symint(c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const aten/src/ATen/core/TensorBody.h:2452
    #28 0x3ff79a802db in torch::autograd::THPVariable_resize_(_object*, _object*, _object*)::$_0::operator()(at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/us
er/pytorch/torch/csrc/autograd/generated/python_variable_methods.cpp:13417
    #29 0x3ff7999f1eb in torch::autograd::THPVariable_resize_(_object*, _object*, _object*) /home/user/pytorch/torch/csrc/autograd/generated/python_variable_methods.cpp:13419
    #30 0x3ffa2c9b009 in method_vectorcall_VARARGS_KEYWORDS Objects/descrobject.c:344
    #31 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #32 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #33 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #34 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #35 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #36 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #37 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #38 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    #39 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #40 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #41 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #42 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #43 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #44 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #45 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #46 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    #47 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #48 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #49 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #50 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #51 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #52 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #53 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #54 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #55 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #56 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #57 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #58 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #59 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #60 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #61 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #62 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #63 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #64 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #65 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #66 0x3ffa2dff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #67 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #68 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #69 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #70 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #71 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #72 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #73 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #74 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #75 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #76 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #77 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #78 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #79 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #80 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #81 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #82 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #83 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #84 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #85 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #86 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #87 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #88 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #89 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #90 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #91 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #92 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #93 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #94 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #95 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #96 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #97 0x3ffa2c8ab9b in PyVectorcall_Call Objects/call.c:267
    #98 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #99 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #100 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #101 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #102 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #103 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #104 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #105 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #106 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #107 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    #108 0x3ffa2c8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #109 0x3ffa2df0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #110 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #111 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #112 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #113 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #114 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #115 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #116 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #117 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #118 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #119 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #120 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #121 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #122 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #123 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    #124 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #125 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #126 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #127 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #128 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #129 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #130 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #131 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #132 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #133 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #134 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #135 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #136 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #137 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #138 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #139 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #140 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #141 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #142 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #143 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #144 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #145 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #146 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #147 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #148 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #149 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    #150 0x3ffa2c8ad17 in _PyObject_Call Objects/call.c:305
    #151 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #152 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #153 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #154 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #155 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #156 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #157 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #158 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #159 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #160 0x3ffa2dff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #161 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #162 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #163 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #164 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #165 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #166 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #167 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #168 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #169 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #170 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #171 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #172 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #173 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    #174 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #175 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #176 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #177 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #178 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #179 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #180 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #181 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #182 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #183 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #184 0x3ffa2dff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #185 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #186 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #187 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #188 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #189 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #190 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #191 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #192 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #193 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #194 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #195 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    #196 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #197 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #198 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #199 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #200 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #201 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #202 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #203 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #204 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #205 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #206 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #207 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #208 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #209 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #210 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #211 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #212 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #213 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #214 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #215 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #216 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #217 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #218 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #219 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #220 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #221 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    #222 0x3ffa2c8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #223 0x3ffa2df0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #224 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #225 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #226 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #227 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #228 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #229 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #230 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    #231 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    #232 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    #233 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    #234 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #235 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #236 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #237 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #238 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #239 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #240 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #241 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #242 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #243 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #244 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #245 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #246 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    #247 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #248 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #249 0x3ffa2e05447 in call_function Python/ceval.c:5891
    #250 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #251 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #252 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    #253 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #254 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #255 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #256 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    #257 0x3ffa2c8a933 in _PyObject_MakeTpCall Objects/call.c:215

0x61000013d790 is located 80 bytes inside of 192-byte region [0x61000013d740,0x61000013d800)
freed by thread T0 here:
    #0 0x3ffa3237de5 in operator delete(void*) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
    #1 0x3ff8e7e3221 in c10::TensorImpl::~TensorImpl() /home/user/pytorch/c10/core/TensorImpl.cpp:75

previously allocated by thread T0 here:
    #0 0x3ffa323734f in operator new(unsigned long) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
    #1 0x3ff4aeeb3d1 in c10::intrusive_ptr<c10::TensorImpl, c10::detail::intrusive_target_default_null_type<c10::TensorImpl> > c10::intrusive_ptr<c10::TensorImpl, c10::detail::intrusive_target_default_nul
l_type<c10::TensorImpl> >::make<c10::intrusive_ptr<c10::StorageImpl, c10::detail::intrusive_target_default_null_type<c10::StorageImpl> >, c10::DispatchKeySet&, caffe2::TypeMeta&>(c10::intrusive_ptr<c10::S
torageImpl, c10::detail::intrusive_target_default_null_type<c10::StorageImpl> >&&, c10::DispatchKeySet&, caffe2::TypeMeta&) /home/user/pytorch/c10/util/intrusive_ptr.h:498
    #2 0x3ff76f79e17  (/home/user/pytorch/build/lib.linux-s390x-cpython-310/torch/lib/libtorch_cpu.so+0x2fb79e17)

SUMMARY: AddressSanitizer: heap-use-after-free /home/user/pytorch/c10/core/SymInt.h:154 in c10::SymInt::is_heap_allocated() const
Shadow bytes around the buggy address:
  0x100c2000027aa0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
  0x100c2000027ab0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c2000027ac0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
  0x100c2000027ad0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c2000027ae0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
=>0x100c2000027af0: fd fd[fd]fd fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c2000027b00: fa fa fa fa fa fa fa fa 00 00 00 00 00 00 00 00
  0x100c2000027b10: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x100c2000027b20: fa fa fa fa fa fa fa fa 00 00 00 00 00 00 00 00
  0x100c2000027b30: 00 00 00 00 04 fa fa fa fa fa fa fa fa fa fa fa
  0x100c2000027b40: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==1115867==ABORTING
```
</details>

<details>
<summary>Additional backtraces (not full)</summary>

Memory deallocation:
```
#0  operator delete (ptr=0x61000013d740) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
#1  0x000003ffa77e3222 in c10::TensorImpl::~TensorImpl (this=0x61000013d740) at /home/user/pytorch/c10/core/TensorImpl.cpp:75
#2  0x000003ff63e76e8c in c10::intrusive_ptr<c10::TensorImpl, c10::UndefinedTensorImpl>::reset_ (this=0x3ffd7ec8230) at /home/user/pytorch/c10/util/intrusive_ptr.h:291
#3  0x000003ff63e76910 in c10::intrusive_ptr<c10::TensorImpl, c10::UndefinedTensorImpl>::~intrusive_ptr (this=0x3ffd7ec8230) at /home/user/pytorch/c10/util/intrusive_ptr.h:370
#4  0x000003ff63e67240 in at::TensorBase::~TensorBase (this=0x3ffd7ec8230) at /home/user/pytorch/aten/src/ATen/core/TensorBase.h:80
#5  0x000003ff63e85ee0 in at::Tensor::~Tensor (this=0x3ffd7ec8230) at aten/src/ATen/core/TensorBody.h:90
#6  0x000003ff63f67304 in resize__functionalization (dispatchKeySet=..., self=..., size=..., memory_format=...) at /home/user/pytorch/aten/src/ATen/FunctionalizeFallbackKernel.cpp:173
#7  0x000003ff63f89258 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>), &(resize__functionalization(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>))>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>) (
    this=0x6030000390a0, args=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
#8  c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>), &(resize__functionalization(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>))>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>) (functor=0x6030000390a0, dispatchKeySet=..., args=..., args=...,
    args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:480
#9  0x000003ff6aca560a in c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > (
    unboxed_kernel_func=0x3ff63f88a80 <c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tenso
r const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>), &(resize__functionalization(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>))>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>)>, functor=0x6030000390a0,
    dispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
#10 0x000003ff6aca715c in c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > (this=0x6210005e1b28, opHandle=...,
    dispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:96
#11 c10::Dispatcher::redispatch<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff919400e0 <c10::Dispatcher::realSingleton()::_singleton>, op=..., currentDispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:656
#12 0x000003ff6a82006c in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff919a07e0 <at::_ops::resize_::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)::op>, currentDispatchKeySet=..., args=...,
    args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:492
#13 at::_ops::resize_::redispatch (dispatchKeySet=..., self=..., size=..., memory_format=...) at /home/user/pytorch/build/aten/src/ATen/Operators_4.cpp:2144
#14 0x000003ff861d5e08 in at::redispatch::resize__symint (dispatchKeySet=..., self=..., size=..., memory_format=...) at aten/src/ATen/RedispatchFunctions.h:2847
#15 0x000003ff861b579e in torch::ADInplaceOrView::resize_ (ks=..., self=..., size=..., optional_memory_format=...) at /home/user/pytorch/torch/csrc/autograd/VariableTypeManual.cpp:401
```

Memory access:
```
#0  c10::SymInt::maybe_as_int (this=0x61000013d790) at /home/user/pytorch/c10/core/SymInt.h:215
#1  0x000003ff734d0a6e in c10::SymInt::sym_eq (this=0x61000013d790, sci=...) at /home/user/pytorch/c10/core/SymInt.cpp:69
#2  0x000003ff5f6ab0be in c10::SymInt::operator== (this=0x61000013d790, o=...) at /home/user/pytorch/c10/core/SymInt.h:177
#3  0x000003ff5f6aaede in std::__equal<false>::equal<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1162
#4  0x000003ff5f6aae4c in std::__equal_aux1<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1211
#5  0x000003ff5f6aae06 in std::__equal_aux<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1219
#6  0x000003ff5f6aad98 in std::equal<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1556
#7  0x000003ff2ff3c772 in c10::ArrayRef<c10::SymInt>::equals (this=0x3ffed7c9900, RHS=...) at /home/user/pytorch/c10/util/ArrayRef.h:188
#8  0x000003ff31891bc2 in c10::operator!=<c10::SymInt> (a1=..., a2=...) at /home/user/pytorch/c10/util/ArrayRef.h:341
#9  0x000003ff51eb5800 in torch::ADInplaceOrView::resize_ (ks=..., self=..., size=..., optional_memory_format=...) at /home/user/pytorch/torch/csrc/autograd/VariableTypeManual.cpp:408
#10 0x000003ff51ee59c8 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c
10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>
 > >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) (this=0x6030007dca40, args=..., args=..., args=..., args=...)
    at /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
#11 c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt
>, c10::optional<c10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<
c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tenso
r const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) (functor=0x6030007dca40, dispatchKeySet=..., args=..., args=..., args=...)
    at /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:480
#12 0x000003ff369a512a in c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > (
    unboxed_kernel_func=0x3ff51ee51f0 <c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tenso
r const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::Ar
rayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKern
el*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>, functor=0x6030007dca40, dispatchKeySet=..., args=..., args=..., args=...)
    at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
#13 0x000003ff369a6e90 in c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > (this=0x6210005e1bc8, opHandle=...,
    dispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:90
#14 c10::Dispatcher::redispatch<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::Arr
ayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff5d6400e0 <c10::Dispatcher::realSingleton()::_singleton>, op=..., currentDispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:656
#15 0x000003ff3652006c in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::redispatch(c10::DispatchKeySet, at::Tensor const&,
c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff5d6a07e0 <at::_ops::resize_::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)::op>, currentDispatchKeySet=..., args=...,
    args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:492
#16 at::_ops::resize_::redispatch (dispatchKeySet=..., self=..., size=..., memory_format=...) at /home/user/pytorch/build/aten/src/ATen/Operators_4.cpp:2144
#17 0x000003ff51ed5e08 in at::redispatch::resize__symint (dispatchKeySet=..., self=..., size=..., memory_format=...) at aten/src/ATen/RedispatchFunctions.h:2847
#18 0x000003ff51ebbb68 in torch::autograd::VariableType::(anonymous namespace)::resize_ (ks=..., self=..., size=..., optional_memory_format=...)
    at /home/user/pytorch/torch/csrc/autograd/VariableTypeManual.cpp:243
```
</details>
Pull Request resolved: #101064
Approved by: https://github.com/Skylion007, https://github.com/albanD
pytorchmergebot pushed a commit that referenced this pull request May 15, 2023
arguments() returns vector member of object returned by schema() call.
When object returned by schema() call is destroyed, the vector is deallocated as well,
it's lifetime isn't extended.

This issue detected while running `pytest -v test/mobile/test_lite_script_type.py -k test_nest_typing_namedtuple_custom_classtype` with ASAN.

<details>
<summary>ASAN output</summary>

```
==1134126==ERROR: AddressSanitizer: heap-use-after-free on address 0x60d0005a5790 at pc 0x03ff844488d8 bp 0x03fff584afe8 sp 0x03fff584afd8
READ of size 8 at 0x60d0005a5790 thread T0
    #0 0x3ff844488d7 in __gnu_cxx::__normal_iterator<c10::Argument const*, std::vector<c10::Argument, std::allocator<c10::Argument> > >::__normal_iterator(c10::Argument const* const&) /usr/lib/gcc/s390x-i
bm-linux-gnu/11/include/g++-v11/bits/stl_iterator.h:1028
    #1 0x3ff8444293f in std::vector<c10::Argument, std::allocator<c10::Argument> >::begin() const /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_vector.h:821
    #2 0x3ff84d807d1 in torch::jit::toPyObject(c10::IValue) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:617
    #3 0x3ff84d80305 in torch::jit::toPyObject(c10::IValue) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
    #4 0x3ff84856871 in pybind11::detail::type_caster<c10::IValue, void>::cast(c10::IValue, pybind11::return_value_policy, pybind11::handle) /home/user/pytorch/torch/csrc/jit/python/pybind.h:138
    #5 0x3ff85318191 in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is
_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_me
thod const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::operator()(pybind11::detail::function_call&) const /home/user/pytorch/cmake/../third_party/pybin
d11/include/pybind11/pybind11.h:249
    #6 0x3ff85317cfd in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is
_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_me
thod const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::__invoke(pybind11::detail::function_call&) /home/user/pytorch/cmake/../third_party/pybind11/incl
ude/pybind11/pybind11.h:224
    #7 0x3ff82ee52e9 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:929
    #8 0x3ffab002903 in cfunction_call Objects/methodobject.c:543
    #9 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #10 0x3ffaaf8e919 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #11 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #12 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #13 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #14 0x3ffab105447 in call_function Python/ceval.c:5891
    #15 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #16 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #17 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #18 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #19 0x3ffaaf8a615 in _PyObject_FastCallDictTstate Objects/call.c:142
    #20 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #21 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #22 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #23 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #24 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #25 0x3ffab105447 in call_function Python/ceval.c:5891
    #26 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #27 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #28 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #29 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #30 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #31 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #32 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #33 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #34 0x3ffab105447 in call_function Python/ceval.c:5891
    #35 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #36 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #37 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #38 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #39 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #40 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #41 0x3ffab105447 in call_function Python/ceval.c:5891
    #42 0x3ffab0ff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #43 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #44 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #45 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #46 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #47 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #48 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #49 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #50 0x3ffab105447 in call_function Python/ceval.c:5891
    #51 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #52 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #53 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #54 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #55 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #56 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #57 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #58 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #59 0x3ffab105447 in call_function Python/ceval.c:5891
    #60 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #61 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #62 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #63 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #64 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #65 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #66 0x3ffaaf8ab9b in PyVectorcall_Call Objects/call.c:267
    #67 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #68 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #69 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #70 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #71 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #72 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #73 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #74 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #75 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #76 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #77 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #78 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #79 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #80 0x3ffab105447 in call_function Python/ceval.c:5891
    #81 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #82 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #83 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #84 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #85 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #86 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #87 0x3ffab105447 in call_function Python/ceval.c:5891
    #88 0x3ffab0ff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #89 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #90 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #91 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #92 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #93 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #94 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #95 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #96 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #97 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #98 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #99 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #100 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #101 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #102 0x3ffab105447 in call_function Python/ceval.c:5891
    #103 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #104 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #105 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #106 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #107 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #108 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #109 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #110 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #111 0x3ffab105447 in call_function Python/ceval.c:5891
    #112 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #113 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #114 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #115 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #116 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #117 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #118 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #119 0x3ffaaf8ad17 in _PyObject_Call Objects/call.c:305
    #120 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #121 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #122 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #123 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #124 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #125 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #126 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #127 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #128 0x3ffab105447 in call_function Python/ceval.c:5891
    #129 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #130 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #131 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #132 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #133 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #134 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #135 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #136 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #137 0x3ffab105447 in call_function Python/ceval.c:5891
    #138 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #139 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #140 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #141 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #142 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #143 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #144 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #145 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #146 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #147 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #148 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #149 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #150 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #151 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #152 0x3ffab105447 in call_function Python/ceval.c:5891
    #153 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #154 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #155 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #156 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #157 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #158 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #159 0x3ffab105447 in call_function Python/ceval.c:5891
    #160 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #161 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #162 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #163 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #164 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #165 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #166 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #167 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #168 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #169 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #170 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #171 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #172 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #173 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #174 0x3ffab105447 in call_function Python/ceval.c:5891
    #175 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #176 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #177 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #178 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #179 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #180 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #181 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #182 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #183 0x3ffab105447 in call_function Python/ceval.c:5891
    #184 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #185 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #186 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #187 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #188 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #189 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #190 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #191 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #192 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #193 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #194 0x3ffab105447 in call_function Python/ceval.c:5891
    #195 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #196 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #197 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #198 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #199 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #200 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #201 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #202 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #203 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #204 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #205 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #206 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #207 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #208 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #209 0x3ffab105447 in call_function Python/ceval.c:5891
    #210 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #211 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #212 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #213 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #214 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #215 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #216 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #216 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #217 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #218 0x3ffab105447 in call_function Python/ceval.c:5891
    #219 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #220 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #221 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #222 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #223 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #224 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #225 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #226 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #227 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #228 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #229 0x3ffab105447 in call_function Python/ceval.c:5891
    #230 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #231 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #232 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #233 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #234 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #235 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #236 0x3ffab105447 in call_function Python/ceval.c:5891
    #237 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #238 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #239 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #240 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #241 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #242 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #243 0x3ffab105447 in call_function Python/ceval.c:5891
    #244 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #245 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #246 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #247 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #248 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #249 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290

0x60d0005a5790 is located 80 bytes inside of 136-byte region [0x60d0005a5740,0x60d0005a57c8)
freed by thread T0 here:
    #0 0x3ffab537de5 in operator delete(void*) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
    #1 0x3ff55984fdb in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::deallocate(std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2>*, unsigned long) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:145

previously allocated by thread T0 here:
    #0 0x3ffab53734f in operator new(unsigned long) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
    #1 0x3ff5598443f in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::allocate(unsigned long, void const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:127
    #2 0x3fff5849ecf  ([stack]+0xb2ecf)

SUMMARY: AddressSanitizer: heap-use-after-free /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_iterator.h:1028 in __gnu_cxx::__normal_iterator<c10::Argument const*, std::vector<c10::Argument, std::allocator<c10::Argument> > >::__normal_iterator(c10::Argument const* const&)
Shadow bytes around the buggy address:
  0x100c1a000b4aa0: fd fd fd fd fd fd fd fd fd fd fd fa fa fa fa fa
  0x100c1a000b4ab0: fa fa fa fa fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c1a000b4ac0: fd fd fd fd fd fa fa fa fa fa fa fa fa fa fd fd
  0x100c1a000b4ad0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fa
  0x100c1a000b4ae0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
=>0x100c1a000b4af0: fd fd[fd]fd fd fd fd fd fd fa fa fa fa fa fa fa
  0x100c1a000b4b00: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b10: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b20: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b30: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b40: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==1134126==ABORTING
```

Additional backtraces (not full):
Allocation:
```
#0  __memset_z196 () at ../sysdeps/s390/memset-z900.S:144
#1  0x000003ff96f3072a in __asan::Allocator::Allocate (this=this@entry=0x3ff97041eb8 <__asan::instance>, size=size@entry=136, alignment=8, alignment@entry=0, stack=<optimized out>,
    stack@entry=0x3ffdbb45d78, alloc_type=<optimized out>, can_fill=true) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_allocator.cpp:599
#2  0x000003ff96f2c088 in __asan::asan_memalign (alignment=alignment@entry=0, size=size@entry=136, stack=stack@entry=0x3ffdbb45d78, alloc_type=alloc_type@entry=__asan::FROM_NEW)
    at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_allocator.cpp:1039
#3  0x000003ff96fb73b0 in operator new (size=136) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
#4  0x000003ff41404440 in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::allocate (this=0x3ffdbb468c0,
    __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:127
#5  0x000003ff414042a0 in std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::allocate (__a=...,
    __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/alloc_traits.h:464
#6  0x000003ff41403b66 in std::__allocate_guarded<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > > (__a=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/allocated_ptr.h:98
#7  0x000003ff4140372a in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::__shared_count<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (this=0x3ffdbb47888, __p=@0x3ffdbb47880: 0x0, __a=..., __args=..., __args=..., __args=..., __args=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:648
#8  0x000003ff41403328 in std::__shared_ptr<c10::FunctionSchema, (__gnu_cxx::_Lock_policy)2>::__shared_ptr<std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (this=0x3ffdbb47880, __tag=..., __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1342
#9  0x000003ff41402f06 in std::shared_ptr<c10::FunctionSchema>::shared_ptr<std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (
    this=0x3ffdbb47880, __tag=..., __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:409
#10 0x000003ff41402b6e in std::allocate_shared<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (__a=...,
    __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:862
#11 0x000003ff4140215c in std::make_shared<c10::FunctionSchema, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (__args=..., __args=..., __args=..., __args=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:878
#12 0x000003ff413d180c in c10::TupleType::createWithSpec<c10::basic_string_view<char> > (qualName=..., field_names=std::vector of length 1, capacity 1 = {...},
    field_types=std::vector of length 1, capacity 1 = {...}, field_defaults=std::vector of length 0, capacity 0) at /home/user/pytorch/aten/src/ATen/core/type.cpp:769
#13 0x000003ff413b9ca6 in c10::TupleType::createNamed (qualName=..., field_names=std::vector of length 1, capacity 1 = {...}, field_types=std::vector of length 1, capacity 1 = {...})
    at /home/user/pytorch/aten/src/ATen/core/type.cpp:725
#14 0x000003ff4115fbac in c10::ivalue::TupleTypeFactory<c10::TupleType>::fallback (type=...) at /home/user/pytorch/aten/src/ATen/core/dynamic_type.cpp:383
#15 0x000003ff708217fe in c10::ivalue::Tuple::type<c10::TupleType> (this=0x6080004b8520) at /home/user/pytorch/aten/src/ATen/core/ivalue_inl.h:781
#16 0x000003ff70800740 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:613
#17 0x000003ff70800306 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
#18 0x000003ff702d6872 in pybind11::detail::type_caster<c10::IValue, void>::cast (src=...) at /home/user/pytorch/torch/csrc/jit/python/pybind.h:138
#19 0x000003ff70d98192 in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::operator()(pybind11::detail::function_call&) const (this=0x3ffdbb4ca20, call=...)
    at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:249
#20 0x000003ff70d97cfe in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::__invoke(pybind11::detail::function_call&) (call=...)
    at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:224
#21 0x000003ff6e9652ea in pybind11::cpp_function::dispatcher (self=<PyCapsule at remote 0x3ff83e27720>,
    args_in=(<torch._C.LiteScriptModule at remote 0x3ff811844b0>, (<Tensor at remote 0x3ff814efb00>,)), kwargs_in=0x0) at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:929
```

Deallocation:
```
#0  operator delete (ptr=0x60d0005a5740) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
#1  0x000003ff44904fdc in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::deallocate (this=0x3ffc5dc8020,
    __p=0x60d0005a5740, __t=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:145
#2  0x000003ff44904fa8 in std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::deallocate (
    __a=..., __p=0x60d0005a5740, __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/alloc_traits.h:496
#3  0x000003ff449041f2 in std::__allocated_ptr<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::~__allocated_ptr (
    this=0x3ffc5dc8030) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/allocated_ptr.h:74
#4  0x000003ff44904888 in std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2>::_M_destroy (this=0x60d0005a5740)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:538
#5  0x000003ff43895a62 in std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release (this=0x60d0005a5740) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:184
#6  0x000003ff43895420 in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count (this=0x611000c40648) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:705
#7  0x000003ff4466e7f4 in std::__shared_ptr<c10::FunctionSchema, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr (this=0x611000c40640)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1154
#8  0x000003ff4466d820 in std::shared_ptr<c10::FunctionSchema>::~shared_ptr (this=0x611000c40640) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:122
#9  0x000003ff448d82f6 in c10::TupleType::~TupleType (this=0x611000c40580) at /home/user/pytorch/aten/src/ATen/core/jit_type.h:1142
#10 0x000003ff448d8346 in c10::TupleType::~TupleType (this=0x611000c40580) at /home/user/pytorch/aten/src/ATen/core/jit_type.h:1142
#11 0x000003ff731296a4 in std::_Sp_counted_ptr<c10::TupleType*, (__gnu_cxx::_Lock_policy)2>::_M_dispose (this=0x603000c43ae0)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:348
#12 0x000003ff71eaf666 in std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release (this=0x603000c43ae0) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:168
#13 0x000003ff71eaf330 in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count (this=0x3ffc5dc9368) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:705
#14 0x000003ff73129ee4 in std::__shared_ptr<c10::TupleType, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr (this=0x3ffc5dc9360)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1154
#15 0x000003ff73122390 in std::shared_ptr<c10::TupleType>::~shared_ptr (this=0x3ffc5dc9360) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:122
#16 0x000003ff73d00788 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:613
#17 0x000003ff73d00306 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
```
</details>
Pull Request resolved: #101400
Approved by: https://github.com/zou3519
jcaip pushed a commit that referenced this pull request May 23, 2023
arguments() returns vector member of object returned by schema() call.
When object returned by schema() call is destroyed, the vector is deallocated as well,
it's lifetime isn't extended.

This issue detected while running `pytest -v test/mobile/test_lite_script_type.py -k test_nest_typing_namedtuple_custom_classtype` with ASAN.

<details>
<summary>ASAN output</summary>

```
==1134126==ERROR: AddressSanitizer: heap-use-after-free on address 0x60d0005a5790 at pc 0x03ff844488d8 bp 0x03fff584afe8 sp 0x03fff584afd8
READ of size 8 at 0x60d0005a5790 thread T0
    #0 0x3ff844488d7 in __gnu_cxx::__normal_iterator<c10::Argument const*, std::vector<c10::Argument, std::allocator<c10::Argument> > >::__normal_iterator(c10::Argument const* const&) /usr/lib/gcc/s390x-i
bm-linux-gnu/11/include/g++-v11/bits/stl_iterator.h:1028
    #1 0x3ff8444293f in std::vector<c10::Argument, std::allocator<c10::Argument> >::begin() const /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_vector.h:821
    #2 0x3ff84d807d1 in torch::jit::toPyObject(c10::IValue) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:617
    #3 0x3ff84d80305 in torch::jit::toPyObject(c10::IValue) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
    #4 0x3ff84856871 in pybind11::detail::type_caster<c10::IValue, void>::cast(c10::IValue, pybind11::return_value_policy, pybind11::handle) /home/user/pytorch/torch/csrc/jit/python/pybind.h:138
    #5 0x3ff85318191 in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is
_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_me
thod const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::operator()(pybind11::detail::function_call&) const /home/user/pytorch/cmake/../third_party/pybin
d11/include/pybind11/pybind11.h:249
    #6 0x3ff85317cfd in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is
_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_me
thod const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::__invoke(pybind11::detail::function_call&) /home/user/pytorch/cmake/../third_party/pybind11/incl
ude/pybind11/pybind11.h:224
    #7 0x3ff82ee52e9 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:929
    #8 0x3ffab002903 in cfunction_call Objects/methodobject.c:543
    #9 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #10 0x3ffaaf8e919 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #11 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #12 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #13 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #14 0x3ffab105447 in call_function Python/ceval.c:5891
    #15 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #16 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #17 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #18 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #19 0x3ffaaf8a615 in _PyObject_FastCallDictTstate Objects/call.c:142
    #20 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #21 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #22 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #23 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #24 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #25 0x3ffab105447 in call_function Python/ceval.c:5891
    #26 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #27 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #28 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #29 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #30 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #31 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #32 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #33 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #34 0x3ffab105447 in call_function Python/ceval.c:5891
    #35 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #36 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #37 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #38 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #39 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #40 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #41 0x3ffab105447 in call_function Python/ceval.c:5891
    #42 0x3ffab0ff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #43 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #44 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #45 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #46 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #47 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #48 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #49 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #50 0x3ffab105447 in call_function Python/ceval.c:5891
    #51 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #52 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #53 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #54 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #55 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #56 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #57 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #58 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #59 0x3ffab105447 in call_function Python/ceval.c:5891
    #60 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #61 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #62 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #63 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #64 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #65 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #66 0x3ffaaf8ab9b in PyVectorcall_Call Objects/call.c:267
    #67 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #68 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #69 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #70 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #71 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #72 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #73 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #74 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #75 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #76 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #77 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #78 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #79 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #80 0x3ffab105447 in call_function Python/ceval.c:5891
    #81 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #82 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #83 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #84 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #85 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #86 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #87 0x3ffab105447 in call_function Python/ceval.c:5891
    #88 0x3ffab0ff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #89 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #90 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #91 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #92 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #93 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #94 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #95 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #96 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #97 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #98 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #99 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #100 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #101 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #102 0x3ffab105447 in call_function Python/ceval.c:5891
    #103 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #104 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #105 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #106 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #107 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #108 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #109 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #110 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #111 0x3ffab105447 in call_function Python/ceval.c:5891
    #112 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #113 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #114 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #115 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #116 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #117 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #118 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #119 0x3ffaaf8ad17 in _PyObject_Call Objects/call.c:305
    #120 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #121 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #122 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #123 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #124 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #125 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #126 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #127 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #128 0x3ffab105447 in call_function Python/ceval.c:5891
    #129 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #130 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #131 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #132 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #133 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #134 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #135 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #136 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #137 0x3ffab105447 in call_function Python/ceval.c:5891
    #138 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #139 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #140 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #141 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #142 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #143 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #144 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #145 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #146 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #147 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #148 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #149 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #150 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #151 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #152 0x3ffab105447 in call_function Python/ceval.c:5891
    #153 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #154 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #155 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #156 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #157 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #158 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #159 0x3ffab105447 in call_function Python/ceval.c:5891
    #160 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #161 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #162 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #163 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #164 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #165 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #166 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #167 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #168 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #169 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #170 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #171 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #172 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #173 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #174 0x3ffab105447 in call_function Python/ceval.c:5891
    #175 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #176 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #177 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #178 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #179 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #180 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #181 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #182 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #183 0x3ffab105447 in call_function Python/ceval.c:5891
    #184 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #185 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #186 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #187 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #188 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #189 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #190 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #191 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #192 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #193 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #194 0x3ffab105447 in call_function Python/ceval.c:5891
    #195 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #196 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #197 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #198 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #199 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #200 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    #201 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    #202 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    #203 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #204 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #205 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #206 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #207 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #208 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #209 0x3ffab105447 in call_function Python/ceval.c:5891
    #210 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #211 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #212 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #213 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #214 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #215 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    #216 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #216 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #217 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #218 0x3ffab105447 in call_function Python/ceval.c:5891
    #219 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #220 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #221 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #222 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #223 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    #224 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    #225 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    #226 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    #227 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #228 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #229 0x3ffab105447 in call_function Python/ceval.c:5891
    #230 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    #231 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #232 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #233 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #234 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #235 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #236 0x3ffab105447 in call_function Python/ceval.c:5891
    #237 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #238 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #239 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #240 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #241 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #242 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    #243 0x3ffab105447 in call_function Python/ceval.c:5891
    #244 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #245 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #246 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    #247 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    #248 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    #249 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290

0x60d0005a5790 is located 80 bytes inside of 136-byte region [0x60d0005a5740,0x60d0005a57c8)
freed by thread T0 here:
    #0 0x3ffab537de5 in operator delete(void*) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
    #1 0x3ff55984fdb in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::deallocate(std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2>*, unsigned long) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:145

previously allocated by thread T0 here:
    #0 0x3ffab53734f in operator new(unsigned long) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
    #1 0x3ff5598443f in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::allocate(unsigned long, void const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:127
    #2 0x3fff5849ecf  ([stack]+0xb2ecf)

SUMMARY: AddressSanitizer: heap-use-after-free /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_iterator.h:1028 in __gnu_cxx::__normal_iterator<c10::Argument const*, std::vector<c10::Argument, std::allocator<c10::Argument> > >::__normal_iterator(c10::Argument const* const&)
Shadow bytes around the buggy address:
  0x100c1a000b4aa0: fd fd fd fd fd fd fd fd fd fd fd fa fa fa fa fa
  0x100c1a000b4ab0: fa fa fa fa fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c1a000b4ac0: fd fd fd fd fd fa fa fa fa fa fa fa fa fa fd fd
  0x100c1a000b4ad0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fa
  0x100c1a000b4ae0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
=>0x100c1a000b4af0: fd fd[fd]fd fd fd fd fd fd fa fa fa fa fa fa fa
  0x100c1a000b4b00: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b10: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b20: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b30: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b40: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==1134126==ABORTING
```

Additional backtraces (not full):
Allocation:
```
#0  __memset_z196 () at ../sysdeps/s390/memset-z900.S:144
#1  0x000003ff96f3072a in __asan::Allocator::Allocate (this=this@entry=0x3ff97041eb8 <__asan::instance>, size=size@entry=136, alignment=8, alignment@entry=0, stack=<optimized out>,
    stack@entry=0x3ffdbb45d78, alloc_type=<optimized out>, can_fill=true) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_allocator.cpp:599
#2  0x000003ff96f2c088 in __asan::asan_memalign (alignment=alignment@entry=0, size=size@entry=136, stack=stack@entry=0x3ffdbb45d78, alloc_type=alloc_type@entry=__asan::FROM_NEW)
    at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_allocator.cpp:1039
#3  0x000003ff96fb73b0 in operator new (size=136) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
#4  0x000003ff41404440 in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::allocate (this=0x3ffdbb468c0,
    __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:127
#5  0x000003ff414042a0 in std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::allocate (__a=...,
    __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/alloc_traits.h:464
#6  0x000003ff41403b66 in std::__allocate_guarded<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > > (__a=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/allocated_ptr.h:98
#7  0x000003ff4140372a in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::__shared_count<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (this=0x3ffdbb47888, __p=@0x3ffdbb47880: 0x0, __a=..., __args=..., __args=..., __args=..., __args=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:648
#8  0x000003ff41403328 in std::__shared_ptr<c10::FunctionSchema, (__gnu_cxx::_Lock_policy)2>::__shared_ptr<std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (this=0x3ffdbb47880, __tag=..., __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1342
#9  0x000003ff41402f06 in std::shared_ptr<c10::FunctionSchema>::shared_ptr<std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (
    this=0x3ffdbb47880, __tag=..., __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:409
#10 0x000003ff41402b6e in std::allocate_shared<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (__a=...,
    __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:862
#11 0x000003ff4140215c in std::make_shared<c10::FunctionSchema, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (__args=..., __args=..., __args=..., __args=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:878
#12 0x000003ff413d180c in c10::TupleType::createWithSpec<c10::basic_string_view<char> > (qualName=..., field_names=std::vector of length 1, capacity 1 = {...},
    field_types=std::vector of length 1, capacity 1 = {...}, field_defaults=std::vector of length 0, capacity 0) at /home/user/pytorch/aten/src/ATen/core/type.cpp:769
#13 0x000003ff413b9ca6 in c10::TupleType::createNamed (qualName=..., field_names=std::vector of length 1, capacity 1 = {...}, field_types=std::vector of length 1, capacity 1 = {...})
    at /home/user/pytorch/aten/src/ATen/core/type.cpp:725
#14 0x000003ff4115fbac in c10::ivalue::TupleTypeFactory<c10::TupleType>::fallback (type=...) at /home/user/pytorch/aten/src/ATen/core/dynamic_type.cpp:383
#15 0x000003ff708217fe in c10::ivalue::Tuple::type<c10::TupleType> (this=0x6080004b8520) at /home/user/pytorch/aten/src/ATen/core/ivalue_inl.h:781
#16 0x000003ff70800740 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:613
#17 0x000003ff70800306 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
#18 0x000003ff702d6872 in pybind11::detail::type_caster<c10::IValue, void>::cast (src=...) at /home/user/pytorch/torch/csrc/jit/python/pybind.h:138
#19 0x000003ff70d98192 in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::operator()(pybind11::detail::function_call&) const (this=0x3ffdbb4ca20, call=...)
    at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:249
#20 0x000003ff70d97cfe in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)#1}::__invoke(pybind11::detail::function_call&) (call=...)
    at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:224
#21 0x000003ff6e9652ea in pybind11::cpp_function::dispatcher (self=<PyCapsule at remote 0x3ff83e27720>,
    args_in=(<torch._C.LiteScriptModule at remote 0x3ff811844b0>, (<Tensor at remote 0x3ff814efb00>,)), kwargs_in=0x0) at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:929
```

Deallocation:
```
#0  operator delete (ptr=0x60d0005a5740) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
#1  0x000003ff44904fdc in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::deallocate (this=0x3ffc5dc8020,
    __p=0x60d0005a5740, __t=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:145
#2  0x000003ff44904fa8 in std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::deallocate (
    __a=..., __p=0x60d0005a5740, __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/alloc_traits.h:496
#3  0x000003ff449041f2 in std::__allocated_ptr<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::~__allocated_ptr (
    this=0x3ffc5dc8030) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/allocated_ptr.h:74
#4  0x000003ff44904888 in std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2>::_M_destroy (this=0x60d0005a5740)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:538
#5  0x000003ff43895a62 in std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release (this=0x60d0005a5740) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:184
#6  0x000003ff43895420 in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count (this=0x611000c40648) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:705
#7  0x000003ff4466e7f4 in std::__shared_ptr<c10::FunctionSchema, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr (this=0x611000c40640)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1154
#8  0x000003ff4466d820 in std::shared_ptr<c10::FunctionSchema>::~shared_ptr (this=0x611000c40640) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:122
#9  0x000003ff448d82f6 in c10::TupleType::~TupleType (this=0x611000c40580) at /home/user/pytorch/aten/src/ATen/core/jit_type.h:1142
#10 0x000003ff448d8346 in c10::TupleType::~TupleType (this=0x611000c40580) at /home/user/pytorch/aten/src/ATen/core/jit_type.h:1142
#11 0x000003ff731296a4 in std::_Sp_counted_ptr<c10::TupleType*, (__gnu_cxx::_Lock_policy)2>::_M_dispose (this=0x603000c43ae0)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:348
#12 0x000003ff71eaf666 in std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release (this=0x603000c43ae0) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:168
#13 0x000003ff71eaf330 in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count (this=0x3ffc5dc9368) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:705
#14 0x000003ff73129ee4 in std::__shared_ptr<c10::TupleType, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr (this=0x3ffc5dc9360)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1154
#15 0x000003ff73122390 in std::shared_ptr<c10::TupleType>::~shared_ptr (this=0x3ffc5dc9360) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:122
#16 0x000003ff73d00788 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:613
#17 0x000003ff73d00306 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
```
</details>
Pull Request resolved: #101400
Approved by: https://github.com/zou3519
pytorchmergebot pushed a commit that referenced this pull request May 26, 2023
3 disabled functions are attempting out of bounds reads. Disable them until sleef library is fixed.

<details>
<summary>ASAN report</summary>

```
=================================================================
==2030580==ERROR: AddressSanitizer: global-buffer-overflow on address 0x03ff70f54570 at pc 0x03ff6704e960 bp 0x03ffce128940 sp 0x03ffce128930
READ of size 4 at 0x03ff70f54570 thread T0
    #0 0x3ff6704e95f in vgather_vf_p_vi2 /home/user/pytorch/third_party/sleef/src/arch/helpers390x_128.h:129
    #1 0x3ff6704e95f in rempif /home/user/pytorch/third_party/sleef/src/libm/sleefsimdsp.c:550
    #2 0x3ff6704e95f in Sleef_cosf4_u10vxe2 /home/user/pytorch/third_party/sleef/src/libm/sleefsimdsp.c:1021
    #3 0x3ff67029cfb in Sleef_cosf4_u10 /home/user/pytorch/build/sleef/src/libm/disps390x_128.c:182
    #4 0x3ff55d21941 in at::vec::ZVECTOR::Vectorized<float, void> at::vec::ZVECTOR::Vectorized<float, void>::mapSleef<float __vector(4) const (*)(float __vector(4)), double __vector(2) const (*)(double __
vector(2)), float, 0>(float __vector(4) const (*)(float __vector(4)), double __vector(2) const (*)(double __vector(2))) const /home/user/pytorch/aten/src/ATen/cpu/vec/vec256/zarch/vec256_zarch.h:991
    #5 0x3ff5689ad01 in at::vec::ZVECTOR::Vectorized<float, void>::cos() const /home/user/pytorch/aten/src/ATen/cpu/vec/vec256/zarch/vec256_zarch.h:1074
    #6 0x3ff5685df97 in at::vml::ZVECTOR::vcos<float>(float*, float const*, long)::{lambda(at::vec::ZVECTOR::Vectorized<float, void>)#1}::operator()(at::vec::ZVECTOR::Vectorized<float, void>) const /home/
user/pytorch/aten/src/ATen/cpu/vml.h:71
    #7 0x3ff5689b691 in void at::vec::map<float, at::vml::ZVECTOR::vcos<float>(float*, float const*, long)::{lambda(at::vec::ZVECTOR::Vectorized<float, void>)#1}, 0>(at::vml::ZVECTOR::vcos<float>(float*,
float const*, long)::{lambda(at::vec::ZVECTOR::Vectorized<float, void>)#1} const&, float*, float const*, long) /home/user/pytorch/aten/src/ATen/cpu/vec/functional_base.h:239
    #8 0x3ff5685e0df in void at::vml::ZVECTOR::vcos<float>(float*, float const*, long) /home/user/pytorch/aten/src/ATen/cpu/vml.h:71
    #9 0x3ff563fdde3 in operator() /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    #10 0x3ff5648e4a3 in operator() /home/user/pytorch/aten/src/ATen/TensorIterator.h:406
    #11 0x3ff5663cae1 in callback_fn<at::TensorIteratorBase::loop_2d_from_1d<at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&)::<lambda()>::<lambda()>::<lambda(char**, const int64_t*, int64_t)> >(c
onst at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&)::<lambda()>::<lambda()>::<lambda(char**, const int64_t*, int64_t)>&)::<lambda(char**, const int64_t*, int64_t, int64_t)> > /home/user/pytorch/
c10/util/FunctionRef.h:43
    #12 0x3ff4d45a933 in c10::function_ref<void (char**, long const*, long, long)>::operator()(char**, long const*, long, long) const /home/user/pytorch/c10/util/FunctionRef.h:64
    #13 0x3ff4d455133 in at::internal::serial_for_each(c10::ArrayRef<long>, c10::ArrayRef<long>, char**, unsigned long, c10::function_ref<void (char**, long const*, long, long)>, at::Range) /home/user/pyt
orch/aten/src/ATen/TensorIteratorInternal.h:52
    #14 0x3ff4d43b703 in at::TensorIteratorBase::serial_for_each(c10::function_ref<void (char**, long const*, long, long)>, at::Range) const /home/user/pytorch/aten/src/ATen/TensorIterator.cpp:777
    #15 0x3ff4d43ab59 in at::TensorIteratorBase::for_each(c10::function_ref<void (char**, long const*, long, long)>, long) /home/user/pytorch/aten/src/ATen/TensorIterator.cpp:749
    #16 0x3ff5648e851 in for_each<at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&)::<lambda()>::<lambda()>::<lambda(char**, const int64_t*, int64_t)> > /home/user/pytorch/aten/src/ATen/TensorItera
tor.h:421
    #17 0x3ff563fe5f9 in operator() /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    #18 0x3ff56400915 in operator() /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    #19 0x3ff56400f1d in at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&) /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    #20 0x3ff4f303007 in void at::native::DispatchStub<void (*)(at::TensorIteratorBase&), at::native::cos_stub>::operator()<at::native::structured_cos_out&>(c10::DeviceType, at::native::structured_cos_out
&) /home/user/pytorch/aten/src/ATen/native/DispatchStub.h:158
    #21 0x3ff4f2edb3f in at::native::structured_cos_out::impl(at::Tensor const&, at::Tensor const&) /home/user/pytorch/aten/src/ATen/native/UnaryOps.cpp:330
    #22 0x3ff526ef739 in wrapper_CPU_cos /home/user/pytorch/build/aten/src/ATen/RegisterCPU.cpp:4307
    #23 0x3ff52c651d9 in operator() /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
    #24 0x3ff52c651d9 in call /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:463
    #25 0x3ff5076df2f in at::Tensor c10::callUnboxedKernelFunction<at::Tensor, at::Tensor const&>(void*, c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&) /home/user/pytorch/aten/src/ATen/core
/boxing/KernelFunction_impl.h:50
    #26 0x3ff5009a93f in at::Tensor c10::KernelFunction::call<at::Tensor, at::Tensor const&>(c10::OperatorHandle const&, c10::DispatchKeySet, at::Tensor const&) const /home/user/pytorch/aten/src/ATen/core
/boxing/KernelFunction_impl.h:103
    #27 0x3ff5009a93f in at::Tensor c10::Dispatcher::call<at::Tensor, at::Tensor const&>(c10::TypedOperatorHandle<at::Tensor (at::Tensor const&)> const&, at::Tensor const&) const /home/user/pytorch/aten/s
rc/ATen/core/dispatch/Dispatcher.h:639
    #28 0x3ff5009a93f in c10::TypedOperatorHandle<at::Tensor (at::Tensor const&)>::call(at::Tensor const&) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:487
    #29 0x3ff5009a93f in at::_ops::cos::call(at::Tensor const&) /home/user/pytorch/build/aten/src/ATen/Operators_0.cpp:2215
    #30 0x3ff7d813741 in at::Tensor::cos() const /home/user/pytorch/build/aten/src/ATen/core/TensorBody.h:2107
    #31 0x3ff7dc0f2b7 in operator() /home/user/pytorch/torch/csrc/autograd/generated/python_torch_functions_2.cpp:2953
    #32 0x3ff7dc0faf7 in THPVariable_cos /home/user/pytorch/torch/csrc/autograd/generated/python_torch_functions_2.cpp:2955
    #33 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    #34 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #35 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #36 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    #37 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #38 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #39 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #40 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #41 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    #42 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #43 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #44 0x3ff7f87a393 in torch::impl::dispatch::PythonKernelHolder::operator()(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch/
torch/csrc/utils/python_dispatch.cpp:175
    #45 0x3ff7f8871a7 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch::
PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)#1}::operator()(c10::OperatorKernel*, c10::Op
eratorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:87
    #46 0x3ff7f887261 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch::
PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)#1}::_FUN(c10::OperatorKernel*, c10::Operator
Handle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:86
    #47 0x3ff7e0d10ab in c10::BoxedKernel::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/b
oxing/BoxedKernel_impl.h:41
    #48 0x3ff7e0d1459 in c10::KernelFunction::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/cor
e/boxing/KernelFunction_impl.h:43
    #49 0x3ff7f876421 in c10::Dispatcher::callBoxed(c10::OperatorHandle const&, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:6
91
    #50 0x3ff4d22bcdd in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:417
    #51 0x3ff65a092d5 in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:421
    #52 0x3ff65a05641 in operator() /home/user/pytorch/torch/csrc/jit/runtime/register_c10_ops.cpp:15
    #53 0x3ff65a08cb5 in __invoke_impl<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c1
0::IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:61
    #54 0x3ff65a0897b in __invoke_r<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c10::
IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:111
    #55 0x3ff65a084e1 in _M_invoke /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/std_function.h:290
    #56 0x3ff7eb2cb21 in std::function<void (std::vector<c10::IValue, std::allocator<c10::IValue> >&)>::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /usr/lib/gcc/s390x-ibm-lin
ux-gnu/11/include/g++-v11/bits/std_function.h:590
    #57 0x3ff7eb1b659 in torch::jit::Operation::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) /home/user/pytorch/aten/src/ATen/core/stack.h:41
    #58 0x3ff7eb08449 in torch::jit::invokeOperatorFromPython(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, pybind11::args, pybind11::
kwargs const&, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:764
    #59 0x3ff7eb09d85 in torch::jit::_get_operation_for_overload_or_packet(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, c10::Symbol,
pybind11::args, pybind11::kwargs const&, bool, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:829
    #60 0x3ff7e573eb9 in operator() /home/user/pytorch/torch/csrc/jit/python/init.cpp:1549
    #61 0x3ff7e6728dd in call_impl<pybind11::object, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&, 0, 1, pybind11::detail::vo
id_type> /home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1439
    #62 0x3ff7e64312f in call<pybind11::object, pybind11::detail::void_type, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&> /h
ome/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1408
    #63 0x3ff7e5da259 in operator() /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:249
    #64 0x3ff7e5da441 in _FUN /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:224
    #65 0x3ff7d317a1f in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:929
    #66 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    #67 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #68 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #69 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    #70 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #71 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #72 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #73 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #74 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    #75 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    #76 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    #77 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #78 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #79 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #80 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #81 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #82 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #83 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #84 0x3ffa5fd76a3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #85 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    #86 0x3ffa5feb289 in call_function Python/ceval.c:5891
    #87 0x3ffa5fe5c3b in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #88 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #89 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #90 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #91 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    #92 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #93 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #94 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #95 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #96 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #97 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #98 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #99 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    #100 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #101 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #102 0x3ff7f87a393 in torch::impl::dispatch::PythonKernelHolder::operator()(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch
/torch/csrc/utils/python_dispatch.cpp:175
    #103 0x3ff7f8871a7 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch:
:PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)#1}::operator()(c10::OperatorKernel*, c10::O
peratorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:87
    #104 0x3ff7f887261 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch:
:PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)#1}::_FUN(c10::OperatorKernel*, c10::Operato
rHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:86
    #105 0x3ff7e0d10ab in c10::BoxedKernel::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/
boxing/BoxedKernel_impl.h:41
    #106 0x3ff7e0d1459 in c10::KernelFunction::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/co
re/boxing/KernelFunction_impl.h:43
    #107 0x3ff7f876421 in c10::Dispatcher::callBoxed(c10::OperatorHandle const&, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:
691
    #108 0x3ff4d22bcdd in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:417
    #109 0x3ff65a092d5 in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:421
    #110 0x3ff65a05641 in operator() /home/user/pytorch/torch/csrc/jit/runtime/register_c10_ops.cpp:15
    #111 0x3ff65a08cb5 in __invoke_impl<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c
10::IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:61
    #112 0x3ff65a0897b in __invoke_r<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c10:
:IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:111
    #113 0x3ff65a084e1 in _M_invoke /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/std_function.h:290
    #114 0x3ff7eb2cb21 in std::function<void (std::vector<c10::IValue, std::allocator<c10::IValue> >&)>::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /usr/lib/gcc/s390x-ibm-li
nux-gnu/11/include/g++-v11/bits/std_function.h:590
    #115 0x3ff7eb1b659 in torch::jit::Operation::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) /home/user/pytorch/aten/src/ATen/core/stack.h:41
    #116 0x3ff7eb08449 in torch::jit::invokeOperatorFromPython(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, pybind11::args, pybind11:
:kwargs const&, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:764
    #117 0x3ff7eb09d85 in torch::jit::_get_operation_for_overload_or_packet(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, c10::Symbol,
 pybind11::args, pybind11::kwargs const&, bool, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:829
    #118 0x3ff7e573eb9 in operator() /home/user/pytorch/torch/csrc/jit/python/init.cpp:1549
    #119 0x3ff7e6728dd in call_impl<pybind11::object, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&, 0, 1, pybind11::detail::v
oid_type> /home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1439
    #120 0x3ff7e64312f in call<pybind11::object, pybind11::detail::void_type, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&> /
home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1408
    #121 0x3ff7e5da259 in operator() /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:249
    #122 0x3ff7e5da441 in _FUN /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:224
    #123 0x3ff7d317a1f in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:929
    #124 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    #125 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #126 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #127 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    #128 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #129 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #130 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #131 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #132 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    #133 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    #134 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    #135 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #136 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #137 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #138 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #139 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #140 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #141 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #142 0x3ffa5e87d2b in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #143 0x3ffa5e882dd in method_vectorcall Objects/classobject.c:83
    #144 0x3ffa5e836d3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #145 0x3ffa5e84b6f in _PyObject_CallFunctionVa Objects/call.c:485
    #146 0x3ffa5e84f2d in callmethod Objects/call.c:557
    #147 0x3ffa5e85039 in PyObject_CallMethod Objects/call.c:577
    #148 0x3ff7f7efa05 in torch::handle_torch_function_no_python_arg_parser(c10::ArrayRef<pybind11::handle>, _object*, _object*, char const*, _object*, char const*, torch::TorchFunctionName) /home/user/py
torch/torch/csrc/utils/python_arg_parser.cpp:338
    #149 0x3ff7eb09b67 in torch::jit::_get_operation_for_overload_or_packet(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, c10::Symbol,
 pybind11::args, pybind11::kwargs const&, bool, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:827
    #150 0x3ff7e573eb9 in operator() /home/user/pytorch/torch/csrc/jit/python/init.cpp:1549
    #151 0x3ff7e6728dd in call_impl<pybind11::object, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&, 0, 1, pybind11::detail::v
oid_type> /home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1439
    #152 0x3ff7e64312f in call<pybind11::object, pybind11::detail::void_type, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&> /
home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1408
    #153 0x3ff7e5da259 in operator() /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:249
    #154 0x3ff7e5da441 in _FUN /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:224
    #155 0x3ff7d317a1f in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:929
    #156 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    #157 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #158 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #159 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    #160 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #161 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #162 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #163 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #164 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    #165 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    #166 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    #167 0x3ffa5e84027 in _PyObject_MakeTpCall Objects/call.c:215
    #168 0x3ffa5fd767b in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    #169 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    #170 0x3ffa5feb289 in call_function Python/ceval.c:5891
    #171 0x3ffa5fe5ad1 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    #172 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #173 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #174 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #175 0x3ffa5fd76a3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #176 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    #177 0x3ffa5feb289 in call_function Python/ceval.c:5891
    #178 0x3ffa5fe5c3b in _PyEval_EvalFrameDefault Python/ceval.c:4213
    #179 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #180 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #181 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #182 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    #183 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #184 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #185 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #186 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #187 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #188 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #189 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #190 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    #191 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #192 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #193 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #194 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #195 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #196 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #197 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #198 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    #199 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #200 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #201 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #202 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #203 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #204 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #205 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #206 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    #207 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #208 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #209 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #210 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #211 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #212 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #213 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #214 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    #215 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    #216 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    #217 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    #218 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #219 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #220 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #221 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #222 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #223 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #224 0x3ffa5fd76a3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    #225 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    #226 0x3ffa5feb289 in call_function Python/ceval.c:5891
    #227 0x3ffa5fe5b21 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    #228 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #229 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #230 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #231 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    #232 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #233 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #234 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #235 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #236 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #237 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #238 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #239 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    #240 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #241 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #242 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #243 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #244 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #245 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #246 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #247 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    #248 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    #249 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    #250 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    #251 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    #252 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    #253 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    #254 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    #255 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267

0x03ff70f54570 is located 0 bytes to the right of global variable 'Sleef_rempitabsp' defined in '/home/user/pytorch/third_party/sleef/src/libm/rempitab.c:986:34' (0x3ff70f53f00) of size 1648
SUMMARY: AddressSanitizer: global-buffer-overflow /home/user/pytorch/third_party/sleef/src/arch/helpers390x_128.h:129 in vgather_vf_p_vi2
Shadow bytes around the buggy address:
  0x10007fee1ea850: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea860: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea870: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea880: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea890: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
=>0x10007fee1ea8a0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00[f9]f9
  0x10007fee1ea8b0: f9 f9 f9 f9 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8c0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8d0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8e0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8f0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==2030580==ABORTING
```
</details>

It reproduces when running `pytest -v test/test_ops.py -k test_python_ref__refs_cos_cpu_bfloat16` under address sanitizer on s390x.

See also: shibatch/sleef#464

Pull Request resolved: #102266
Approved by: https://github.com/malfet
BenjaminDEMAILLE added a commit to BenjaminDEMAILLE/pytorch that referenced this pull request Jan 24, 2026
Implements the Marsaglia-Tsang (2000) rejection sampling algorithm
for gamma distribution sampling on Apple Silicon GPUs.

This PR enables bioinformatics tools like CellBender to run natively
on MPS without CPU fallback. CellBender is a widely-used single-cell
RNA-seq analysis tool that uses Pyro/PyTorch for Bayesian inference
with Gamma distributions.

Implementation details:
- Custom Metal kernel using Philox4x32-10 RNG (matching CUDA)
- Box-Muller transform for normal samples required by Marsaglia-Tsang
- Handles alpha < 1 using: Gamma(a) = Gamma(a+1) * U^(1/a)
- Supports float32, float16, and bfloat16 dtypes

Statistical validation:
- Mean and variance match theoretical values within 1% for 100K samples
- Results comparable to CPU implementation

Community impact:
- Unblocks CellBender (GitHub issues pytorch#149, pytorch#171) for Apple Silicon users
- Enables native GPU acceleration for Pyro probabilistic programming
- Benefits scRNA-seq researchers using MacBooks/Mac Studios
eqy pushed a commit to eqy/pytorch that referenced this pull request Mar 4, 2026
laurentdupin pushed a commit to laurentdupin/pytorch that referenced this pull request Apr 25, 2026
When tensor is resized, reference array to it's sizes may become invalid. Make a copy in advance.

<details>
<summary>ASAN report</summary>

```
=================================================================
==1115867==ERROR: AddressSanitizer: heap-use-after-free on address 0x61000013d790 at pc 0x03ff8e7da360 bp 0x03fff53c83a0 sp 0x03fff53c8390
READ of size 8 at 0x61000013d790 thread T0
    #0 0x3ff8e7da35f in c10::SymInt::is_heap_allocated() const /home/user/pytorch/c10/core/SymInt.h:154
    pytorch#1 0x3ff8e7da35f in c10::SymInt::maybe_as_int() const /home/user/pytorch/c10/core/SymInt.h:215
    pytorch#2 0x3ff8e7d0a6d in c10::SymInt::sym_eq(c10::SymInt const&) const /home/user/pytorch/c10/core/SymInt.cpp:69
    pytorch#3 0x3ff7a9ab0bd in c10::SymInt::operator==(c10::SymInt const&) const /home/user/pytorch/c10/core/SymInt.h:177
    pytorch#4 0x3ff7a9aaedd in bool std::__equal<false>::equal<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-
v11/bits/stl_algobase.h:1162
    pytorch#5 0x3ff7a9aae4b in bool std::__equal_aux1<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/
stl_algobase.h:1211
    pytorch#6 0x3ff7a9aae05 in bool std::__equal_aux<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/s
tl_algobase.h:1219
    pytorch#7 0x3ff7a9aad97 in bool std::equal<c10::SymInt const*, c10::SymInt const*>(c10::SymInt const*, c10::SymInt const*, c10::SymInt const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_alg
obase.h:1556
    pytorch#8 0x3ff4b23c771 in c10::ArrayRef<c10::SymInt>::equals(c10::ArrayRef<c10::SymInt>) const /home/user/pytorch/c10/util/ArrayRef.h:188
    pytorch#9 0x3ff4cb91bc1 in bool c10::operator!=<c10::SymInt>(c10::ArrayRef<c10::SymInt>, c10::ArrayRef<c10::SymInt>) /home/user/pytorch/c10/util/ArrayRef.h:341
    pytorch#10 0x3ff6d1b57ff in torch::ADInplaceOrView::resize_(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/torch/csrc/autograd/Variab
leTypeManual.cpp:408
    pytorch#11 0x3ff6d1e59c7 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c1
0::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>
> >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
    pytorch#12 0x3ff6d1e59c7 in c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10:
:ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::Sy
mInt>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::Disp
atchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:480
    pytorch#13 0x3ff51ca5129 in at::Tensor const& c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(void*, c10::OperatorKernel*,
c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>&&, c10::optional<c10::MemoryFormat>&&) /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
    pytorch#14 0x3ff51ca6e8f in at::Tensor const& c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::OperatorHandle const&, c10::D
ispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:90
    pytorch#15 0x3ff51ca6e8f in at::Tensor const& c10::Dispatcher::redispatch<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Ten
sor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)
const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:656
    pytorch#16 0x3ff5182006b in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::redispatch(c10::DispatchKeySet, at::Tensor const&, c
10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:492
    pytorch#17 0x3ff5182006b in at::_ops::resize_::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) aten/src/ATen/Operators_4.cpp:2144
    pytorch#18 0x3ff6d1d5e07 in at::redispatch::resize__symint(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) aten/src/ATen/RedispatchFunctions.h:2847
    pytorch#19 0x3ff6d1bbb67 in torch::autograd::VariableType::(anonymous namespace)::resize_(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pyto
rch/torch/csrc/autograd/VariableTypeManual.cpp:243
    pytorch#20 0x3ff6d1bd197 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c1
0::MemoryFormat>), &torch::autograd::VariableType::(anonymous namespace)::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10
::optional<c10::MemoryFormat> > >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFu
nctionIntoFunctor.h:13
    pytorch#21 0x3ff6d1bd197 in c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10:
:ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>), &torch::autograd::VariableType::(anonymous namespace)::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor
 const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c
10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor
.h:480
    pytorch#22 0x3ff51ca5129 in at::Tensor const& c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(void*, c10::OperatorKernel*,
c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>&&, c10::optional<c10::MemoryFormat>&&) /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
    pytorch#23 0x3ff5181ead1 in at::Tensor const& c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::OperatorHandle const&, c10::D
ispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:90
    pytorch#24 0x3ff5181ead1 in at::Tensor const& c10::Dispatcher::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Tensor co
nst& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/user/pytorch/at
en/src/ATen/core/dispatch/Dispatcher.h:639
    pytorch#25 0x3ff5181ead1 in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(at::Tensor const&, c10::ArrayRef<c10::SymInt>,
c10::optional<c10::MemoryFormat>) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:487
    pytorch#26 0x3ff5181ead1 in at::_ops::resize_::call(at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) aten/src/ATen/Operators_4.cpp:2137
    pytorch#27 0x3ff79b44fcf in at::Tensor::resize__symint(c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const aten/src/ATen/core/TensorBody.h:2452
    pytorch#28 0x3ff79a802db in torch::autograd::THPVariable_resize_(_object*, _object*, _object*)::$_0::operator()(at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const /home/us
er/pytorch/torch/csrc/autograd/generated/python_variable_methods.cpp:13417
    pytorch#29 0x3ff7999f1eb in torch::autograd::THPVariable_resize_(_object*, _object*, _object*) /home/user/pytorch/torch/csrc/autograd/generated/python_variable_methods.cpp:13419
    pytorch#30 0x3ffa2c9b009 in method_vectorcall_VARARGS_KEYWORDS Objects/descrobject.c:344
    pytorch#31 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#32 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#33 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#34 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#35 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#36 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#37 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#38 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#39 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#40 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#41 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#42 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#43 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#44 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#45 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#46 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#47 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#48 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#49 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#50 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#51 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#52 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#53 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#54 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#55 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#56 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#57 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#58 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#59 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#60 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#61 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#62 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#63 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#64 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#65 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#66 0x3ffa2dff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#67 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#68 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#69 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#70 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#71 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#72 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#73 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#74 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#75 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#76 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#77 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#78 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#79 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#80 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#81 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#82 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#83 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#84 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#85 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#86 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#87 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#88 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#89 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#90 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#91 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#92 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#93 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#94 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#95 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#96 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#97 0x3ffa2c8ab9b in PyVectorcall_Call Objects/call.c:267
    pytorch#98 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#99 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#100 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#101 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#102 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#103 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#104 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#105 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#106 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#107 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#108 0x3ffa2c8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#109 0x3ffa2df0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#110 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#111 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#112 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#113 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#114 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#115 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#116 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#117 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#118 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#119 0x3ffa2dff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#120 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#121 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#122 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#123 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#124 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#125 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#126 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#127 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#128 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#129 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#130 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#131 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#132 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#133 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#134 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#135 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#136 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#137 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#138 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#139 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#140 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#141 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#142 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#143 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#144 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#145 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#146 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#147 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#148 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#149 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#150 0x3ffa2c8ad17 in _PyObject_Call Objects/call.c:305
    pytorch#151 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#152 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#153 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#154 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#155 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#156 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#157 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#158 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#159 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#160 0x3ffa2dff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#161 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#162 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#163 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#164 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#165 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#166 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#167 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#168 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#169 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#170 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#171 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#172 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#173 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#174 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#175 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#176 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#177 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#178 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#179 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#180 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#181 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#182 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#183 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#184 0x3ffa2dff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#185 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#186 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#187 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#188 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#189 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#190 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#191 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#192 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#193 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#194 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#195 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#196 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#197 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#198 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#199 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#200 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#201 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#202 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#203 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#204 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#205 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#206 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#207 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#208 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#209 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#210 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#211 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#212 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#213 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#214 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#215 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#216 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#217 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#218 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#219 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#220 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#221 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#222 0x3ffa2c8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#223 0x3ffa2df0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#224 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#225 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#226 0x3ffa2dffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#227 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#228 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#229 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#230 0x3ffa2c8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#231 0x3ffa2c8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#232 0x3ffa2c8ada9 in PyObject_Call Objects/call.c:317
    pytorch#233 0x3ffa2e059c7 in do_call_core Python/ceval.c:5943
    pytorch#234 0x3ffa2dffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#235 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#236 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#237 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#238 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#239 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#240 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#241 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#242 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#243 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#244 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#245 0x3ffa2c8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#246 0x3ffa2c8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#247 0x3ffa2df00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#248 0x3ffa2df013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#249 0x3ffa2e05447 in call_function Python/ceval.c:5891
    pytorch#250 0x3ffa2dff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#251 0x3ffa2df052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#252 0x3ffa2e02b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#253 0x3ffa2c8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#254 0x3ffa2c8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#255 0x3ffa2c8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#256 0x3ffa2d3f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#257 0x3ffa2c8a933 in _PyObject_MakeTpCall Objects/call.c:215

0x61000013d790 is located 80 bytes inside of 192-byte region [0x61000013d740,0x61000013d800)
freed by thread T0 here:
    #0 0x3ffa3237de5 in operator delete(void*) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
    pytorch#1 0x3ff8e7e3221 in c10::TensorImpl::~TensorImpl() /home/user/pytorch/c10/core/TensorImpl.cpp:75

previously allocated by thread T0 here:
    #0 0x3ffa323734f in operator new(unsigned long) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
    pytorch#1 0x3ff4aeeb3d1 in c10::intrusive_ptr<c10::TensorImpl, c10::detail::intrusive_target_default_null_type<c10::TensorImpl> > c10::intrusive_ptr<c10::TensorImpl, c10::detail::intrusive_target_default_nul
l_type<c10::TensorImpl> >::make<c10::intrusive_ptr<c10::StorageImpl, c10::detail::intrusive_target_default_null_type<c10::StorageImpl> >, c10::DispatchKeySet&, caffe2::TypeMeta&>(c10::intrusive_ptr<c10::S
torageImpl, c10::detail::intrusive_target_default_null_type<c10::StorageImpl> >&&, c10::DispatchKeySet&, caffe2::TypeMeta&) /home/user/pytorch/c10/util/intrusive_ptr.h:498
    pytorch#2 0x3ff76f79e17  (/home/user/pytorch/build/lib.linux-s390x-cpython-310/torch/lib/libtorch_cpu.so+0x2fb79e17)

SUMMARY: AddressSanitizer: heap-use-after-free /home/user/pytorch/c10/core/SymInt.h:154 in c10::SymInt::is_heap_allocated() const
Shadow bytes around the buggy address:
  0x100c2000027aa0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
  0x100c2000027ab0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c2000027ac0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
  0x100c2000027ad0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c2000027ae0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
=>0x100c2000027af0: fd fd[fd]fd fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c2000027b00: fa fa fa fa fa fa fa fa 00 00 00 00 00 00 00 00
  0x100c2000027b10: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x100c2000027b20: fa fa fa fa fa fa fa fa 00 00 00 00 00 00 00 00
  0x100c2000027b30: 00 00 00 00 04 fa fa fa fa fa fa fa fa fa fa fa
  0x100c2000027b40: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==1115867==ABORTING
```
</details>

<details>
<summary>Additional backtraces (not full)</summary>

Memory deallocation:
```
#0  operator delete (ptr=0x61000013d740) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
pytorch#1  0x000003ffa77e3222 in c10::TensorImpl::~TensorImpl (this=0x61000013d740) at /home/user/pytorch/c10/core/TensorImpl.cpp:75
pytorch#2  0x000003ff63e76e8c in c10::intrusive_ptr<c10::TensorImpl, c10::UndefinedTensorImpl>::reset_ (this=0x3ffd7ec8230) at /home/user/pytorch/c10/util/intrusive_ptr.h:291
pytorch#3  0x000003ff63e76910 in c10::intrusive_ptr<c10::TensorImpl, c10::UndefinedTensorImpl>::~intrusive_ptr (this=0x3ffd7ec8230) at /home/user/pytorch/c10/util/intrusive_ptr.h:370
pytorch#4  0x000003ff63e67240 in at::TensorBase::~TensorBase (this=0x3ffd7ec8230) at /home/user/pytorch/aten/src/ATen/core/TensorBase.h:80
pytorch#5  0x000003ff63e85ee0 in at::Tensor::~Tensor (this=0x3ffd7ec8230) at aten/src/ATen/core/TensorBody.h:90
pytorch#6  0x000003ff63f67304 in resize__functionalization (dispatchKeySet=..., self=..., size=..., memory_format=...) at /home/user/pytorch/aten/src/ATen/FunctionalizeFallbackKernel.cpp:173
pytorch#7  0x000003ff63f89258 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>), &(resize__functionalization(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>))>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>) (
    this=0x6030000390a0, args=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
pytorch#8  c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>), &(resize__functionalization(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>))>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>) (functor=0x6030000390a0, dispatchKeySet=..., args=..., args=...,
    args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:480
pytorch#9  0x000003ff6aca560a in c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > (
    unboxed_kernel_func=0x3ff63f88a80 <c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tenso
r const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>), &(resize__functionalization(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>))>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<long>, c10::optional<c10::MemoryFormat>)>, functor=0x6030000390a0,
    dispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
pytorch#10 0x000003ff6aca715c in c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > (this=0x6210005e1b28, opHandle=...,
    dispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:96
pytorch#11 c10::Dispatcher::redispatch<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff919400e0 <c10::Dispatcher::realSingleton()::_singleton>, op=..., currentDispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:656
pytorch#12 0x000003ff6a82006c in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff919a07e0 <at::_ops::resize_::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)::op>, currentDispatchKeySet=..., args=...,
    args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:492
pytorch#13 at::_ops::resize_::redispatch (dispatchKeySet=..., self=..., size=..., memory_format=...) at /home/user/pytorch/build/aten/src/ATen/Operators_4.cpp:2144
pytorch#14 0x000003ff861d5e08 in at::redispatch::resize__symint (dispatchKeySet=..., self=..., size=..., memory_format=...) at aten/src/ATen/RedispatchFunctions.h:2847
pytorch#15 0x000003ff861b579e in torch::ADInplaceOrView::resize_ (ks=..., self=..., size=..., optional_memory_format=...) at /home/user/pytorch/torch/csrc/autograd/VariableTypeManual.cpp:401
```

Memory access:
```
#0  c10::SymInt::maybe_as_int (this=0x61000013d790) at /home/user/pytorch/c10/core/SymInt.h:215
pytorch#1  0x000003ff734d0a6e in c10::SymInt::sym_eq (this=0x61000013d790, sci=...) at /home/user/pytorch/c10/core/SymInt.cpp:69
pytorch#2  0x000003ff5f6ab0be in c10::SymInt::operator== (this=0x61000013d790, o=...) at /home/user/pytorch/c10/core/SymInt.h:177
pytorch#3  0x000003ff5f6aaede in std::__equal<false>::equal<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1162
pytorch#4  0x000003ff5f6aae4c in std::__equal_aux1<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1211
pytorch#5  0x000003ff5f6aae06 in std::__equal_aux<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1219
pytorch#6  0x000003ff5f6aad98 in std::equal<c10::SymInt const*, c10::SymInt const*> (__first1=0x61000013d790, __last1=0x61000013d7a0, __first2=0x602000015c30)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_algobase.h:1556
pytorch#7  0x000003ff2ff3c772 in c10::ArrayRef<c10::SymInt>::equals (this=0x3ffed7c9900, RHS=...) at /home/user/pytorch/c10/util/ArrayRef.h:188
pytorch#8  0x000003ff31891bc2 in c10::operator!=<c10::SymInt> (a1=..., a2=...) at /home/user/pytorch/c10/util/ArrayRef.h:341
pytorch#9  0x000003ff51eb5800 in torch::ADInplaceOrView::resize_ (ks=..., self=..., size=..., optional_memory_format=...) at /home/user/pytorch/torch/csrc/autograd/VariableTypeManual.cpp:408
pytorch#10 0x000003ff51ee59c8 in c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c
10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>
 > >::operator()(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) (this=0x6030007dca40, args=..., args=..., args=..., args=...)
    at /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
pytorch#11 c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt
>, c10::optional<c10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<
c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKernel*, c10::DispatchKeySet, at::Tenso
r const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) (functor=0x6030007dca40, dispatchKeySet=..., args=..., args=..., args=...)
    at /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:480
pytorch#12 0x000003ff369a512a in c10::callUnboxedKernelFunction<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > (
    unboxed_kernel_func=0x3ff51ee51f0 <c10::impl::wrap_kernel_functor_unboxed_<c10::impl::detail::WrapFunctionIntoFunctor_<c10::CompileTimeFunctionPointer<at::Tensor const& (c10::DispatchKeySet, at::Tenso
r const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>), &torch::ADInplaceOrView::resize_>, at::Tensor const&, c10::guts::typelist::typelist<c10::DispatchKeySet, at::Tensor const&, c10::Ar
rayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > >, at::Tensor const& (c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::call(c10::OperatorKern
el*, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>, functor=0x6030007dca40, dispatchKeySet=..., args=..., args=..., args=...)
    at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:50
pytorch#13 0x000003ff369a6e90 in c10::KernelFunction::call<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> > (this=0x6210005e1bc8, opHandle=...,
    dispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/boxing/KernelFunction_impl.h:90
pytorch#14 c10::Dispatcher::redispatch<at::Tensor const&, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat> >(c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::Arr
ayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)> const&, c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff5d6400e0 <c10::Dispatcher::realSingleton()::_singleton>, op=..., currentDispatchKeySet=..., args=..., args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:656
pytorch#15 0x000003ff3652006c in c10::TypedOperatorHandle<at::Tensor const& (at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)>::redispatch(c10::DispatchKeySet, at::Tensor const&,
c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>) const (
    this=0x3ff5d6a07e0 <at::_ops::resize_::redispatch(c10::DispatchKeySet, at::Tensor const&, c10::ArrayRef<c10::SymInt>, c10::optional<c10::MemoryFormat>)::op>, currentDispatchKeySet=..., args=...,
    args=..., args=...) at /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:492
pytorch#16 at::_ops::resize_::redispatch (dispatchKeySet=..., self=..., size=..., memory_format=...) at /home/user/pytorch/build/aten/src/ATen/Operators_4.cpp:2144
pytorch#17 0x000003ff51ed5e08 in at::redispatch::resize__symint (dispatchKeySet=..., self=..., size=..., memory_format=...) at aten/src/ATen/RedispatchFunctions.h:2847
pytorch#18 0x000003ff51ebbb68 in torch::autograd::VariableType::(anonymous namespace)::resize_ (ks=..., self=..., size=..., optional_memory_format=...)
    at /home/user/pytorch/torch/csrc/autograd/VariableTypeManual.cpp:243
```
</details>
Pull Request resolved: pytorch#101064
Approved by: https://github.com/Skylion007, https://github.com/albanD
laurentdupin pushed a commit to laurentdupin/pytorch that referenced this pull request Apr 25, 2026
arguments() returns vector member of object returned by schema() call.
When object returned by schema() call is destroyed, the vector is deallocated as well,
it's lifetime isn't extended.

This issue detected while running `pytest -v test/mobile/test_lite_script_type.py -k test_nest_typing_namedtuple_custom_classtype` with ASAN.

<details>
<summary>ASAN output</summary>

```
==1134126==ERROR: AddressSanitizer: heap-use-after-free on address 0x60d0005a5790 at pc 0x03ff844488d8 bp 0x03fff584afe8 sp 0x03fff584afd8
READ of size 8 at 0x60d0005a5790 thread T0
    #0 0x3ff844488d7 in __gnu_cxx::__normal_iterator<c10::Argument const*, std::vector<c10::Argument, std::allocator<c10::Argument> > >::__normal_iterator(c10::Argument const* const&) /usr/lib/gcc/s390x-i
bm-linux-gnu/11/include/g++-v11/bits/stl_iterator.h:1028
    pytorch#1 0x3ff8444293f in std::vector<c10::Argument, std::allocator<c10::Argument> >::begin() const /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_vector.h:821
    pytorch#2 0x3ff84d807d1 in torch::jit::toPyObject(c10::IValue) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:617
    pytorch#3 0x3ff84d80305 in torch::jit::toPyObject(c10::IValue) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
    pytorch#4 0x3ff84856871 in pybind11::detail::type_caster<c10::IValue, void>::cast(c10::IValue, pybind11::return_value_policy, pybind11::handle) /home/user/pytorch/torch/csrc/jit/python/pybind.h:138
    pytorch#5 0x3ff85318191 in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is
_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_me
thod const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)pytorch#1}::operator()(pybind11::detail::function_call&) const /home/user/pytorch/cmake/../third_party/pybin
d11/include/pybind11/pybind11.h:249
    pytorch#6 0x3ff85317cfd in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is
_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_me
thod const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)pytorch#1}::__invoke(pybind11::detail::function_call&) /home/user/pytorch/cmake/../third_party/pybind11/incl
ude/pybind11/pybind11.h:224
    pytorch#7 0x3ff82ee52e9 in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:929
    pytorch#8 0x3ffab002903 in cfunction_call Objects/methodobject.c:543
    pytorch#9 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#10 0x3ffaaf8e919 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#11 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#12 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#13 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#14 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#15 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#16 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#17 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#18 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#19 0x3ffaaf8a615 in _PyObject_FastCallDictTstate Objects/call.c:142
    pytorch#20 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#21 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#22 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#23 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#24 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#25 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#26 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#27 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#28 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#29 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#30 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#31 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#32 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#33 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#34 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#35 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#36 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#37 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#38 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#39 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#40 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#41 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#42 0x3ffab0ff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#43 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#44 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#45 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#46 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#47 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#48 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#49 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#50 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#51 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#52 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#53 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#54 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#55 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#56 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#57 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#58 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#59 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#60 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#61 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#62 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#63 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#64 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#65 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#66 0x3ffaaf8ab9b in PyVectorcall_Call Objects/call.c:267
    pytorch#67 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#68 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    pytorch#69 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    pytorch#70 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#71 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#72 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#73 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#74 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#75 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#76 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#77 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#78 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#79 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#80 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#81 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#82 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#83 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#84 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#85 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#86 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#87 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#88 0x3ffab0ff7d7 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#89 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#90 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#91 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#92 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#93 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#94 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    pytorch#95 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    pytorch#96 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#97 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#98 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#99 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#100 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#101 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#102 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#103 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#104 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#105 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#106 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#107 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#108 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#109 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#110 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#111 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#112 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#113 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#114 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#115 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#116 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#117 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#118 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#119 0x3ffaaf8ad17 in _PyObject_Call Objects/call.c:305
    pytorch#120 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    pytorch#121 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    pytorch#122 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#123 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#124 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#125 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#126 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#127 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#128 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#129 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#130 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#131 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#132 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#133 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#134 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#135 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#136 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#137 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#138 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#139 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#140 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#141 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#142 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#143 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#144 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    pytorch#145 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    pytorch#146 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#147 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#148 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#149 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#150 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#151 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#152 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#153 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#154 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#155 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#156 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#157 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#158 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#159 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#160 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#161 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#162 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#163 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#164 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#165 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#166 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    pytorch#167 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    pytorch#168 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#169 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#170 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#171 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#172 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#173 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#174 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#175 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#176 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#177 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#178 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#179 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#180 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#181 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#182 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#183 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#184 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#185 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#186 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#187 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#188 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#189 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#190 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#191 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#192 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#193 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#194 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#195 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#196 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#197 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#198 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#199 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#200 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290
    pytorch#201 0x3ffaaf8ada9 in PyObject_Call Objects/call.c:317
    pytorch#202 0x3ffab1059c7 in do_call_core Python/ceval.c:5943
    pytorch#203 0x3ffab0ffd39 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#204 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#205 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#206 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#207 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#208 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#209 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#210 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#211 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#212 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#213 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#214 0x3ffaaf8e941 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#215 0x3ffaaf8eddd in method_vectorcall Objects/classobject.c:53
    pytorch#216 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#216 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#217 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#218 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#219 0x3ffab0ff779 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#220 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#221 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#222 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#223 0x3ffaaf8a695 in _PyObject_FastCallDictTstate Objects/call.c:153
    pytorch#224 0x3ffaaf8b271 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#225 0x3ffab03f307 in slot_tp_call Objects/typeobject.c:7494
    pytorch#226 0x3ffaaf8a933 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#227 0x3ffab0f0081 in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#228 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#229 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#230 0x3ffab0ffa57 in _PyEval_EvalFrameDefault Python/ceval.c:4231
    pytorch#231 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#232 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#233 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#234 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#235 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#236 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#237 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#238 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#239 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#240 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#241 0x3ffab0f00a9 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#242 0x3ffab0f013d in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#243 0x3ffab105447 in call_function Python/ceval.c:5891
    pytorch#244 0x3ffab0ff905 in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#245 0x3ffab0f052b in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#246 0x3ffab102b67 in _PyEval_Vector Python/ceval.c:5065
    pytorch#247 0x3ffaaf8aec1 in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#248 0x3ffaaf8ab15 in PyVectorcall_Call Objects/call.c:255
    pytorch#249 0x3ffaaf8ac65 in _PyObject_Call Objects/call.c:290

0x60d0005a5790 is located 80 bytes inside of 136-byte region [0x60d0005a5740,0x60d0005a57c8)
freed by thread T0 here:
    #0 0x3ffab537de5 in operator delete(void*) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
    pytorch#1 0x3ff55984fdb in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::deallocate(std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2>*, unsigned long) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:145

previously allocated by thread T0 here:
    #0 0x3ffab53734f in operator new(unsigned long) /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
    pytorch#1 0x3ff5598443f in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::allocate(unsigned long, void const*) /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:127
    pytorch#2 0x3fff5849ecf  ([stack]+0xb2ecf)

SUMMARY: AddressSanitizer: heap-use-after-free /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/stl_iterator.h:1028 in __gnu_cxx::__normal_iterator<c10::Argument const*, std::vector<c10::Argument, std::allocator<c10::Argument> > >::__normal_iterator(c10::Argument const* const&)
Shadow bytes around the buggy address:
  0x100c1a000b4aa0: fd fd fd fd fd fd fd fd fd fd fd fa fa fa fa fa
  0x100c1a000b4ab0: fa fa fa fa fd fd fd fd fd fd fd fd fd fd fd fd
  0x100c1a000b4ac0: fd fd fd fd fd fa fa fa fa fa fa fa fa fa fd fd
  0x100c1a000b4ad0: fd fd fd fd fd fd fd fd fd fd fd fd fd fd fd fa
  0x100c1a000b4ae0: fa fa fa fa fa fa fa fa fd fd fd fd fd fd fd fd
=>0x100c1a000b4af0: fd fd[fd]fd fd fd fd fd fd fa fa fa fa fa fa fa
  0x100c1a000b4b00: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b10: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b20: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b30: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
  0x100c1a000b4b40: fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa fa
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==1134126==ABORTING
```

Additional backtraces (not full):
Allocation:
```
#0  __memset_z196 () at ../sysdeps/s390/memset-z900.S:144
pytorch#1  0x000003ff96f3072a in __asan::Allocator::Allocate (this=this@entry=0x3ff97041eb8 <__asan::instance>, size=size@entry=136, alignment=8, alignment@entry=0, stack=<optimized out>,
    stack@entry=0x3ffdbb45d78, alloc_type=<optimized out>, can_fill=true) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_allocator.cpp:599
pytorch#2  0x000003ff96f2c088 in __asan::asan_memalign (alignment=alignment@entry=0, size=size@entry=136, stack=stack@entry=0x3ffdbb45d78, alloc_type=alloc_type@entry=__asan::FROM_NEW)
    at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_allocator.cpp:1039
pytorch#3  0x000003ff96fb73b0 in operator new (size=136) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:99
pytorch#4  0x000003ff41404440 in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::allocate (this=0x3ffdbb468c0,
    __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:127
pytorch#5  0x000003ff414042a0 in std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::allocate (__a=...,
    __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/alloc_traits.h:464
pytorch#6  0x000003ff41403b66 in std::__allocate_guarded<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > > (__a=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/allocated_ptr.h:98
pytorch#7  0x000003ff4140372a in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::__shared_count<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (this=0x3ffdbb47888, __p=@0x3ffdbb47880: 0x0, __a=..., __args=..., __args=..., __args=..., __args=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:648
pytorch#8  0x000003ff41403328 in std::__shared_ptr<c10::FunctionSchema, (__gnu_cxx::_Lock_policy)2>::__shared_ptr<std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (this=0x3ffdbb47880, __tag=..., __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1342
pytorch#9  0x000003ff41402f06 in std::shared_ptr<c10::FunctionSchema>::shared_ptr<std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (
    this=0x3ffdbb47880, __tag=..., __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:409
pytorch#10 0x000003ff41402b6e in std::allocate_shared<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (__a=...,
    __args=..., __args=..., __args=..., __args=...) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:862
pytorch#11 0x000003ff4140215c in std::make_shared<c10::FunctionSchema, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> > const&, std::__cxx11::basic_string<char, std::char_traits<char>, std::allocator<char> >, std::vector<c10::Argument, std::allocator<c10::Argument> >, std::vector<c10::Argument, std::allocator<c10::Argument> > > (__args=..., __args=..., __args=..., __args=...)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:878
pytorch#12 0x000003ff413d180c in c10::TupleType::createWithSpec<c10::basic_string_view<char> > (qualName=..., field_names=std::vector of length 1, capacity 1 = {...},
    field_types=std::vector of length 1, capacity 1 = {...}, field_defaults=std::vector of length 0, capacity 0) at /home/user/pytorch/aten/src/ATen/core/type.cpp:769
pytorch#13 0x000003ff413b9ca6 in c10::TupleType::createNamed (qualName=..., field_names=std::vector of length 1, capacity 1 = {...}, field_types=std::vector of length 1, capacity 1 = {...})
    at /home/user/pytorch/aten/src/ATen/core/type.cpp:725
pytorch#14 0x000003ff4115fbac in c10::ivalue::TupleTypeFactory<c10::TupleType>::fallback (type=...) at /home/user/pytorch/aten/src/ATen/core/dynamic_type.cpp:383
pytorch#15 0x000003ff708217fe in c10::ivalue::Tuple::type<c10::TupleType> (this=0x6080004b8520) at /home/user/pytorch/aten/src/ATen/core/ivalue_inl.h:781
pytorch#16 0x000003ff70800740 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:613
pytorch#17 0x000003ff70800306 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
pytorch#18 0x000003ff702d6872 in pybind11::detail::type_caster<c10::IValue, void>::cast (src=...) at /home/user/pytorch/torch/csrc/jit/python/pybind.h:138
pytorch#19 0x000003ff70d98192 in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)pytorch#1}::operator()(pybind11::detail::function_call&) const (this=0x3ffdbb4ca20, call=...)
    at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:249
pytorch#20 0x000003ff70d97cfe in pybind11::cpp_function::initialize<torch::jit::initJitScriptBindings(_object*)::$_45, c10::IValue, torch::jit::mobile::Module&, pybind11::tuple const&, pybind11::name, pybind11::is_method, pybind11::sibling, pybind11::arg>(torch::jit::initJitScriptBindings(_object*)::$_45&&, c10::IValue (*)(torch::jit::mobile::Module&, pybind11::tuple const&), pybind11::name const&, pybind11::is_method const&, pybind11::sibling const&, pybind11::arg const&)::{lambda(pybind11::detail::function_call&)pytorch#1}::__invoke(pybind11::detail::function_call&) (call=...)
    at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:224
pytorch#21 0x000003ff6e9652ea in pybind11::cpp_function::dispatcher (self=<PyCapsule at remote 0x3ff83e27720>,
    args_in=(<torch._C.LiteScriptModule at remote 0x3ff811844b0>, (<Tensor at remote 0x3ff814efb00>,)), kwargs_in=0x0) at /home/user/pytorch/cmake/../third_party/pybind11/include/pybind11/pybind11.h:929
```

Deallocation:
```
#0  operator delete (ptr=0x60d0005a5740) at /var/tmp/portage/sys-devel/gcc-11.3.1_p20230303/work/gcc-11-20230303/libsanitizer/asan/asan_new_delete.cpp:160
pytorch#1  0x000003ff44904fdc in __gnu_cxx::new_allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> >::deallocate (this=0x3ffc5dc8020,
    __p=0x60d0005a5740, __t=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/ext/new_allocator.h:145
pytorch#2  0x000003ff44904fa8 in std::allocator_traits<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::deallocate (
    __a=..., __p=0x60d0005a5740, __n=1) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/alloc_traits.h:496
pytorch#3  0x000003ff449041f2 in std::__allocated_ptr<std::allocator<std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2> > >::~__allocated_ptr (
    this=0x3ffc5dc8030) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/allocated_ptr.h:74
pytorch#4  0x000003ff44904888 in std::_Sp_counted_ptr_inplace<c10::FunctionSchema, std::allocator<c10::FunctionSchema>, (__gnu_cxx::_Lock_policy)2>::_M_destroy (this=0x60d0005a5740)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:538
pytorch#5  0x000003ff43895a62 in std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release (this=0x60d0005a5740) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:184
pytorch#6  0x000003ff43895420 in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count (this=0x611000c40648) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:705
pytorch#7  0x000003ff4466e7f4 in std::__shared_ptr<c10::FunctionSchema, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr (this=0x611000c40640)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1154
pytorch#8  0x000003ff4466d820 in std::shared_ptr<c10::FunctionSchema>::~shared_ptr (this=0x611000c40640) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:122
pytorch#9  0x000003ff448d82f6 in c10::TupleType::~TupleType (this=0x611000c40580) at /home/user/pytorch/aten/src/ATen/core/jit_type.h:1142
pytorch#10 0x000003ff448d8346 in c10::TupleType::~TupleType (this=0x611000c40580) at /home/user/pytorch/aten/src/ATen/core/jit_type.h:1142
pytorch#11 0x000003ff731296a4 in std::_Sp_counted_ptr<c10::TupleType*, (__gnu_cxx::_Lock_policy)2>::_M_dispose (this=0x603000c43ae0)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:348
pytorch#12 0x000003ff71eaf666 in std::_Sp_counted_base<(__gnu_cxx::_Lock_policy)2>::_M_release (this=0x603000c43ae0) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:168
pytorch#13 0x000003ff71eaf330 in std::__shared_count<(__gnu_cxx::_Lock_policy)2>::~__shared_count (this=0x3ffc5dc9368) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:705
pytorch#14 0x000003ff73129ee4 in std::__shared_ptr<c10::TupleType, (__gnu_cxx::_Lock_policy)2>::~__shared_ptr (this=0x3ffc5dc9360)
    at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr_base.h:1154
pytorch#15 0x000003ff73122390 in std::shared_ptr<c10::TupleType>::~shared_ptr (this=0x3ffc5dc9360) at /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/shared_ptr.h:122
pytorch#16 0x000003ff73d00788 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:613
pytorch#17 0x000003ff73d00306 in torch::jit::toPyObject (ivalue=...) at /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:604
```
</details>
Pull Request resolved: pytorch#101400
Approved by: https://github.com/zou3519
laurentdupin pushed a commit to laurentdupin/pytorch that referenced this pull request Apr 25, 2026
3 disabled functions are attempting out of bounds reads. Disable them until sleef library is fixed.

<details>
<summary>ASAN report</summary>

```
=================================================================
==2030580==ERROR: AddressSanitizer: global-buffer-overflow on address 0x03ff70f54570 at pc 0x03ff6704e960 bp 0x03ffce128940 sp 0x03ffce128930
READ of size 4 at 0x03ff70f54570 thread T0
    #0 0x3ff6704e95f in vgather_vf_p_vi2 /home/user/pytorch/third_party/sleef/src/arch/helpers390x_128.h:129
    pytorch#1 0x3ff6704e95f in rempif /home/user/pytorch/third_party/sleef/src/libm/sleefsimdsp.c:550
    pytorch#2 0x3ff6704e95f in Sleef_cosf4_u10vxe2 /home/user/pytorch/third_party/sleef/src/libm/sleefsimdsp.c:1021
    pytorch#3 0x3ff67029cfb in Sleef_cosf4_u10 /home/user/pytorch/build/sleef/src/libm/disps390x_128.c:182
    pytorch#4 0x3ff55d21941 in at::vec::ZVECTOR::Vectorized<float, void> at::vec::ZVECTOR::Vectorized<float, void>::mapSleef<float __vector(4) const (*)(float __vector(4)), double __vector(2) const (*)(double __
vector(2)), float, 0>(float __vector(4) const (*)(float __vector(4)), double __vector(2) const (*)(double __vector(2))) const /home/user/pytorch/aten/src/ATen/cpu/vec/vec256/zarch/vec256_zarch.h:991
    pytorch#5 0x3ff5689ad01 in at::vec::ZVECTOR::Vectorized<float, void>::cos() const /home/user/pytorch/aten/src/ATen/cpu/vec/vec256/zarch/vec256_zarch.h:1074
    pytorch#6 0x3ff5685df97 in at::vml::ZVECTOR::vcos<float>(float*, float const*, long)::{lambda(at::vec::ZVECTOR::Vectorized<float, void>)pytorch#1}::operator()(at::vec::ZVECTOR::Vectorized<float, void>) const /home/
user/pytorch/aten/src/ATen/cpu/vml.h:71
    pytorch#7 0x3ff5689b691 in void at::vec::map<float, at::vml::ZVECTOR::vcos<float>(float*, float const*, long)::{lambda(at::vec::ZVECTOR::Vectorized<float, void>)pytorch#1}, 0>(at::vml::ZVECTOR::vcos<float>(float*,
float const*, long)::{lambda(at::vec::ZVECTOR::Vectorized<float, void>)pytorch#1} const&, float*, float const*, long) /home/user/pytorch/aten/src/ATen/cpu/vec/functional_base.h:239
    pytorch#8 0x3ff5685e0df in void at::vml::ZVECTOR::vcos<float>(float*, float const*, long) /home/user/pytorch/aten/src/ATen/cpu/vml.h:71
    pytorch#9 0x3ff563fdde3 in operator() /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    pytorch#10 0x3ff5648e4a3 in operator() /home/user/pytorch/aten/src/ATen/TensorIterator.h:406
    pytorch#11 0x3ff5663cae1 in callback_fn<at::TensorIteratorBase::loop_2d_from_1d<at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&)::<lambda()>::<lambda()>::<lambda(char**, const int64_t*, int64_t)> >(c
onst at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&)::<lambda()>::<lambda()>::<lambda(char**, const int64_t*, int64_t)>&)::<lambda(char**, const int64_t*, int64_t, int64_t)> > /home/user/pytorch/
c10/util/FunctionRef.h:43
    pytorch#12 0x3ff4d45a933 in c10::function_ref<void (char**, long const*, long, long)>::operator()(char**, long const*, long, long) const /home/user/pytorch/c10/util/FunctionRef.h:64
    pytorch#13 0x3ff4d455133 in at::internal::serial_for_each(c10::ArrayRef<long>, c10::ArrayRef<long>, char**, unsigned long, c10::function_ref<void (char**, long const*, long, long)>, at::Range) /home/user/pyt
orch/aten/src/ATen/TensorIteratorInternal.h:52
    pytorch#14 0x3ff4d43b703 in at::TensorIteratorBase::serial_for_each(c10::function_ref<void (char**, long const*, long, long)>, at::Range) const /home/user/pytorch/aten/src/ATen/TensorIterator.cpp:777
    pytorch#15 0x3ff4d43ab59 in at::TensorIteratorBase::for_each(c10::function_ref<void (char**, long const*, long, long)>, long) /home/user/pytorch/aten/src/ATen/TensorIterator.cpp:749
    pytorch#16 0x3ff5648e851 in for_each<at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&)::<lambda()>::<lambda()>::<lambda(char**, const int64_t*, int64_t)> > /home/user/pytorch/aten/src/ATen/TensorItera
tor.h:421
    pytorch#17 0x3ff563fe5f9 in operator() /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    pytorch#18 0x3ff56400915 in operator() /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    pytorch#19 0x3ff56400f1d in at::native::ZVECTOR::cos_kernel(at::TensorIteratorBase&) /home/user/pytorch/aten/src/ATen/native/cpu/UnaryOpsKernel.cpp:770
    pytorch#20 0x3ff4f303007 in void at::native::DispatchStub<void (*)(at::TensorIteratorBase&), at::native::cos_stub>::operator()<at::native::structured_cos_out&>(c10::DeviceType, at::native::structured_cos_out
&) /home/user/pytorch/aten/src/ATen/native/DispatchStub.h:158
    pytorch#21 0x3ff4f2edb3f in at::native::structured_cos_out::impl(at::Tensor const&, at::Tensor const&) /home/user/pytorch/aten/src/ATen/native/UnaryOps.cpp:330
    pytorch#22 0x3ff526ef739 in wrapper_CPU_cos /home/user/pytorch/build/aten/src/ATen/RegisterCPU.cpp:4307
    pytorch#23 0x3ff52c651d9 in operator() /home/user/pytorch/aten/src/ATen/core/boxing/impl/WrapFunctionIntoFunctor.h:13
    pytorch#24 0x3ff52c651d9 in call /home/user/pytorch/aten/src/ATen/core/boxing/impl/make_boxed_from_unboxed_functor.h:463
    pytorch#25 0x3ff5076df2f in at::Tensor c10::callUnboxedKernelFunction<at::Tensor, at::Tensor const&>(void*, c10::OperatorKernel*, c10::DispatchKeySet, at::Tensor const&) /home/user/pytorch/aten/src/ATen/core
/boxing/KernelFunction_impl.h:50
    pytorch#26 0x3ff5009a93f in at::Tensor c10::KernelFunction::call<at::Tensor, at::Tensor const&>(c10::OperatorHandle const&, c10::DispatchKeySet, at::Tensor const&) const /home/user/pytorch/aten/src/ATen/core
/boxing/KernelFunction_impl.h:103
    pytorch#27 0x3ff5009a93f in at::Tensor c10::Dispatcher::call<at::Tensor, at::Tensor const&>(c10::TypedOperatorHandle<at::Tensor (at::Tensor const&)> const&, at::Tensor const&) const /home/user/pytorch/aten/s
rc/ATen/core/dispatch/Dispatcher.h:639
    pytorch#28 0x3ff5009a93f in c10::TypedOperatorHandle<at::Tensor (at::Tensor const&)>::call(at::Tensor const&) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:487
    pytorch#29 0x3ff5009a93f in at::_ops::cos::call(at::Tensor const&) /home/user/pytorch/build/aten/src/ATen/Operators_0.cpp:2215
    pytorch#30 0x3ff7d813741 in at::Tensor::cos() const /home/user/pytorch/build/aten/src/ATen/core/TensorBody.h:2107
    pytorch#31 0x3ff7dc0f2b7 in operator() /home/user/pytorch/torch/csrc/autograd/generated/python_torch_functions_2.cpp:2953
    pytorch#32 0x3ff7dc0faf7 in THPVariable_cos /home/user/pytorch/torch/csrc/autograd/generated/python_torch_functions_2.cpp:2955
    pytorch#33 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    pytorch#34 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#35 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#36 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    pytorch#37 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#38 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#39 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#40 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#41 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    pytorch#42 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#43 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#44 0x3ff7f87a393 in torch::impl::dispatch::PythonKernelHolder::operator()(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch/
torch/csrc/utils/python_dispatch.cpp:175
    pytorch#45 0x3ff7f8871a7 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch::
PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)pytorch#1}::operator()(c10::OperatorKernel*, c10::Op
eratorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:87
    pytorch#46 0x3ff7f887261 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch::
PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)pytorch#1}::_FUN(c10::OperatorKernel*, c10::Operator
Handle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:86
    pytorch#47 0x3ff7e0d10ab in c10::BoxedKernel::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/b
oxing/BoxedKernel_impl.h:41
    pytorch#48 0x3ff7e0d1459 in c10::KernelFunction::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/cor
e/boxing/KernelFunction_impl.h:43
    pytorch#49 0x3ff7f876421 in c10::Dispatcher::callBoxed(c10::OperatorHandle const&, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:6
91
    pytorch#50 0x3ff4d22bcdd in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:417
    pytorch#51 0x3ff65a092d5 in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:421
    pytorch#52 0x3ff65a05641 in operator() /home/user/pytorch/torch/csrc/jit/runtime/register_c10_ops.cpp:15
    pytorch#53 0x3ff65a08cb5 in __invoke_impl<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c1
0::IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:61
    pytorch#54 0x3ff65a0897b in __invoke_r<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c10::
IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:111
    pytorch#55 0x3ff65a084e1 in _M_invoke /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/std_function.h:290
    pytorch#56 0x3ff7eb2cb21 in std::function<void (std::vector<c10::IValue, std::allocator<c10::IValue> >&)>::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /usr/lib/gcc/s390x-ibm-lin
ux-gnu/11/include/g++-v11/bits/std_function.h:590
    pytorch#57 0x3ff7eb1b659 in torch::jit::Operation::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) /home/user/pytorch/aten/src/ATen/core/stack.h:41
    pytorch#58 0x3ff7eb08449 in torch::jit::invokeOperatorFromPython(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, pybind11::args, pybind11::
kwargs const&, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:764
    pytorch#59 0x3ff7eb09d85 in torch::jit::_get_operation_for_overload_or_packet(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, c10::Symbol,
pybind11::args, pybind11::kwargs const&, bool, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:829
    pytorch#60 0x3ff7e573eb9 in operator() /home/user/pytorch/torch/csrc/jit/python/init.cpp:1549
    pytorch#61 0x3ff7e6728dd in call_impl<pybind11::object, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&, 0, 1, pybind11::detail::vo
id_type> /home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1439
    pytorch#62 0x3ff7e64312f in call<pybind11::object, pybind11::detail::void_type, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&> /h
ome/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1408
    pytorch#63 0x3ff7e5da259 in operator() /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:249
    pytorch#64 0x3ff7e5da441 in _FUN /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:224
    pytorch#65 0x3ff7d317a1f in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:929
    pytorch#66 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    pytorch#67 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#68 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#69 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    pytorch#70 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#71 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#72 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#73 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#74 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    pytorch#75 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#76 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    pytorch#77 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#78 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#79 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#80 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#81 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#82 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#83 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#84 0x3ffa5fd76a3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#85 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#86 0x3ffa5feb289 in call_function Python/ceval.c:5891
    pytorch#87 0x3ffa5fe5c3b in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#88 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#89 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#90 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#91 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    pytorch#92 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#93 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#94 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#95 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#96 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#97 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#98 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#99 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    pytorch#100 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#101 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#102 0x3ff7f87a393 in torch::impl::dispatch::PythonKernelHolder::operator()(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch
/torch/csrc/utils/python_dispatch.cpp:175
    pytorch#103 0x3ff7f8871a7 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch:
:PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)pytorch#1}::operator()(c10::OperatorKernel*, c10::O
peratorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:87
    pytorch#104 0x3ff7f887261 in c10::BoxedKernel::makeFromFunctor<torch::impl::dispatch::PythonKernelHolder>(std::unique_ptr<torch::impl::dispatch::PythonKernelHolder, std::default_delete<torch::impl::dispatch:
:PythonKernelHolder> >)::{lambda(c10::OperatorKernel*, c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*)pytorch#1}::_FUN(c10::OperatorKernel*, c10::Operato
rHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) /home/user/pytorch/aten/src/ATen/core/boxing/BoxedKernel_impl.h:86
    pytorch#105 0x3ff7e0d10ab in c10::BoxedKernel::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/
boxing/BoxedKernel_impl.h:41
    pytorch#106 0x3ff7e0d1459 in c10::KernelFunction::callBoxed(c10::OperatorHandle const&, c10::DispatchKeySet, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/co
re/boxing/KernelFunction_impl.h:43
    pytorch#107 0x3ff7f876421 in c10::Dispatcher::callBoxed(c10::OperatorHandle const&, std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:
691
    pytorch#108 0x3ff4d22bcdd in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >*) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:417
    pytorch#109 0x3ff65a092d5 in c10::OperatorHandle::callBoxed(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /home/user/pytorch/aten/src/ATen/core/dispatch/Dispatcher.h:421
    pytorch#110 0x3ff65a05641 in operator() /home/user/pytorch/torch/csrc/jit/runtime/register_c10_ops.cpp:15
    pytorch#111 0x3ff65a08cb5 in __invoke_impl<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c
10::IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:61
    pytorch#112 0x3ff65a0897b in __invoke_r<void, torch::jit::(anonymous namespace)::createOperatorFromC10(const c10::OperatorHandle&)::<lambda(torch::jit::Stack&)>&, std::vector<c10::IValue, std::allocator<c10:
:IValue> >&> /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/invoke.h:111
    pytorch#113 0x3ff65a084e1 in _M_invoke /usr/lib/gcc/s390x-ibm-linux-gnu/11/include/g++-v11/bits/std_function.h:290
    pytorch#114 0x3ff7eb2cb21 in std::function<void (std::vector<c10::IValue, std::allocator<c10::IValue> >&)>::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) const /usr/lib/gcc/s390x-ibm-li
nux-gnu/11/include/g++-v11/bits/std_function.h:590
    pytorch#115 0x3ff7eb1b659 in torch::jit::Operation::operator()(std::vector<c10::IValue, std::allocator<c10::IValue> >&) /home/user/pytorch/aten/src/ATen/core/stack.h:41
    pytorch#116 0x3ff7eb08449 in torch::jit::invokeOperatorFromPython(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, pybind11::args, pybind11:
:kwargs const&, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:764
    pytorch#117 0x3ff7eb09d85 in torch::jit::_get_operation_for_overload_or_packet(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, c10::Symbol,
 pybind11::args, pybind11::kwargs const&, bool, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:829
    pytorch#118 0x3ff7e573eb9 in operator() /home/user/pytorch/torch/csrc/jit/python/init.cpp:1549
    pytorch#119 0x3ff7e6728dd in call_impl<pybind11::object, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&, 0, 1, pybind11::detail::v
oid_type> /home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1439
    pytorch#120 0x3ff7e64312f in call<pybind11::object, pybind11::detail::void_type, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&> /
home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1408
    pytorch#121 0x3ff7e5da259 in operator() /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:249
    pytorch#122 0x3ff7e5da441 in _FUN /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:224
    pytorch#123 0x3ff7d317a1f in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:929
    pytorch#124 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    pytorch#125 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#126 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#127 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    pytorch#128 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#129 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#130 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#131 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#132 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    pytorch#133 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#134 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    pytorch#135 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#136 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#137 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#138 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#139 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#140 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#141 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#142 0x3ffa5e87d2b in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#143 0x3ffa5e882dd in method_vectorcall Objects/classobject.c:83
    pytorch#144 0x3ffa5e836d3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#145 0x3ffa5e84b6f in _PyObject_CallFunctionVa Objects/call.c:485
    pytorch#146 0x3ffa5e84f2d in callmethod Objects/call.c:557
    pytorch#147 0x3ffa5e85039 in PyObject_CallMethod Objects/call.c:577
    pytorch#148 0x3ff7f7efa05 in torch::handle_torch_function_no_python_arg_parser(c10::ArrayRef<pybind11::handle>, _object*, _object*, char const*, _object*, char const*, torch::TorchFunctionName) /home/user/py
torch/torch/csrc/utils/python_arg_parser.cpp:338
    pytorch#149 0x3ff7eb09b67 in torch::jit::_get_operation_for_overload_or_packet(std::vector<std::shared_ptr<torch::jit::Operator>, std::allocator<std::shared_ptr<torch::jit::Operator> > > const&, c10::Symbol,
 pybind11::args, pybind11::kwargs const&, bool, c10::optional<c10::DispatchKey>) /home/user/pytorch/torch/csrc/jit/python/pybind_utils.cpp:827
    pytorch#150 0x3ff7e573eb9 in operator() /home/user/pytorch/torch/csrc/jit/python/init.cpp:1549
    pytorch#151 0x3ff7e6728dd in call_impl<pybind11::object, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&, 0, 1, pybind11::detail::v
oid_type> /home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1439
    pytorch#152 0x3ff7e64312f in call<pybind11::object, pybind11::detail::void_type, torch::jit::initJITBindings(PyObject*)::<lambda(const string&, const string&)>::<lambda(pybind11::args, pybind11::kwargs)>&> /
home/user/pytorch/third_party/pybind11/include/pybind11/cast.h:1408
    pytorch#153 0x3ff7e5da259 in operator() /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:249
    pytorch#154 0x3ff7e5da441 in _FUN /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:224
    pytorch#155 0x3ff7d317a1f in pybind11::cpp_function::dispatcher(_object*, _object*, _object*) /home/user/pytorch/third_party/pybind11/include/pybind11/pybind11.h:929
    pytorch#156 0x3ffa5ef5ae1 in cfunction_call Objects/methodobject.c:543
    pytorch#157 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#158 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#159 0x3ffa5feb50d in do_call_core Python/ceval.c:5915
    pytorch#160 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#161 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#162 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#163 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#164 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    pytorch#165 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#166 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    pytorch#167 0x3ffa5e84027 in _PyObject_MakeTpCall Objects/call.c:215
    pytorch#168 0x3ffa5fd767b in _PyObject_VectorcallTstate Include/cpython/abstract.h:112
    pytorch#169 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#170 0x3ffa5feb289 in call_function Python/ceval.c:5891
    pytorch#171 0x3ffa5fe5ad1 in _PyEval_EvalFrameDefault Python/ceval.c:4181
    pytorch#172 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#173 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#174 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#175 0x3ffa5fd76a3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#176 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#177 0x3ffa5feb289 in call_function Python/ceval.c:5891
    pytorch#178 0x3ffa5fe5c3b in _PyEval_EvalFrameDefault Python/ceval.c:4213
    pytorch#179 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#180 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#181 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#182 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    pytorch#183 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#184 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#185 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#186 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#187 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#188 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#189 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#190 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    pytorch#191 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#192 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#193 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#194 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#195 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#196 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#197 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#198 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    pytorch#199 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#200 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#201 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#202 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#203 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#204 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#205 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#206 0x3ffa5e841fb in PyVectorcall_Call Objects/call.c:255
    pytorch#207 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#208 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#209 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#210 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#211 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#212 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#213 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#214 0x3ffa5e83d1f in _PyObject_FastCallDictTstate Objects/call.c:142
    pytorch#215 0x3ffa5e84937 in _PyObject_Call_Prepend Objects/call.c:431
    pytorch#216 0x3ffa5f2f577 in slot_tp_call Objects/typeobject.c:7494
    pytorch#217 0x3ffa5e843f3 in _PyObject_Call Objects/call.c:305
    pytorch#218 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#219 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#220 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#221 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#222 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#223 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#224 0x3ffa5fd76a3 in _PyObject_VectorcallTstate Include/cpython/abstract.h:114
    pytorch#225 0x3ffa5fd772f in PyObject_Vectorcall Include/cpython/abstract.h:123
    pytorch#226 0x3ffa5feb289 in call_function Python/ceval.c:5891
    pytorch#227 0x3ffa5fe5b21 in _PyEval_EvalFrameDefault Python/ceval.c:4198
    pytorch#228 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#229 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#230 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#231 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    pytorch#232 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#233 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#234 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#235 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#236 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#237 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#238 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#239 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    pytorch#240 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#241 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#242 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#243 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#244 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#245 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#246 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#247 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267
    pytorch#248 0x3ffa5e84347 in _PyObject_Call Objects/call.c:290
    pytorch#249 0x3ffa5e84483 in PyObject_Call Objects/call.c:317
    pytorch#250 0x3ffa5feb7cf in do_call_core Python/ceval.c:5943
    pytorch#251 0x3ffa5fe6019 in _PyEval_EvalFrameDefault Python/ceval.c:4277
    pytorch#252 0x3ffa5fd7aed in _PyEval_EvalFrame Include/internal/pycore_ceval.h:46
    pytorch#253 0x3ffa5fe8ba9 in _PyEval_Vector Python/ceval.c:5065
    pytorch#254 0x3ffa5e8459b in _PyFunction_Vectorcall Objects/call.c:342
    pytorch#255 0x3ffa5e8427f in PyVectorcall_Call Objects/call.c:267

0x03ff70f54570 is located 0 bytes to the right of global variable 'Sleef_rempitabsp' defined in '/home/user/pytorch/third_party/sleef/src/libm/rempitab.c:986:34' (0x3ff70f53f00) of size 1648
SUMMARY: AddressSanitizer: global-buffer-overflow /home/user/pytorch/third_party/sleef/src/arch/helpers390x_128.h:129 in vgather_vf_p_vi2
Shadow bytes around the buggy address:
  0x10007fee1ea850: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea860: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea870: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea880: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea890: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
=>0x10007fee1ea8a0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00[f9]f9
  0x10007fee1ea8b0: f9 f9 f9 f9 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8c0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8d0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8e0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
  0x10007fee1ea8f0: 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00 00
Shadow byte legend (one shadow byte represents 8 application bytes):
  Addressable:           00
  Partially addressable: 01 02 03 04 05 06 07
  Heap left redzone:       fa
  Freed heap region:       fd
  Stack left redzone:      f1
  Stack mid redzone:       f2
  Stack right redzone:     f3
  Stack after return:      f5
  Stack use after scope:   f8
  Global redzone:          f9
  Global init order:       f6
  Poisoned by user:        f7
  Container overflow:      fc
  Array cookie:            ac
  Intra object redzone:    bb
  ASan internal:           fe
  Left alloca redzone:     ca
  Right alloca redzone:    cb
  Shadow gap:              cc
==2030580==ABORTING
```
</details>

It reproduces when running `pytest -v test/test_ops.py -k test_python_ref__refs_cos_cpu_bfloat16` under address sanitizer on s390x.

See also: shibatch/sleef#464

Pull Request resolved: pytorch#102266
Approved by: https://github.com/malfet
pytorchmergebot pushed a commit that referenced this pull request Jul 4, 2026
## Human Note
This lets us return grads for scalar biases. I pusehd back on this in the path since there were workarounds although a lil ugly. But after prototyping i think its fine to land and easy enough to gork.

We should wait for #188860 to land first since they touch teh same code -> how we create the captured buffer grads

## Agent Report
# Report: attention-gym #171 FlexAttention learnable scalar

## Summary

Reproduced the issue in this checkout using the exact scalar `score_mod` pattern:

- Eager forward succeeded, eager backward failed with the reported `vmap(... out_dim is None)` BatchedTensor error.
- Compiled forward succeeded, compiled backward failed in Inductor FlexAttention backward lowering with `AssertionError: ComputedBuffer name must not be None`.

Implemented a PyTorch-side fix for direct 0-D captured score-mod gradients. The PR branch has since been rebased onto latest `origin/main` at `d2f0d442d70` and rebuilt locally with cuDNN enabled (`CUDNN_VERSION=9.23.0`, `USE_CUDNN=1`). Current job torch reports `2.14.0a0+git9fadffc`.

Implemented changes:

- `torch/_higher_order_ops/flex_attention.py`
  - Routes direct 0-D captured score-mod tensors through `mod_index(buffer, [])` while tracing the joint graph, so the existing captured-index custom autograd path emits `flex_lib::zeros_and_scatter([], [], grad)` for scalar gradient accumulation.
  - Preserves unused/no-grad captured buffers by returning `None` instead of copying `None` into an allocated grad buffer.
- `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - Supports empty-index `zeros_and_scatter` as scalar accumulation by summing values into a scalar output.
  - Allows `ModIndex` to use an empty index list as a scalar identity whose backward is empty-index scalar accumulation.
  - Returns the checked 0-D tensor directly in the empty-index `ModIndex.forward` path.
  - Handles empty index metadata in its vmap rule with an annotation-only pyrefly fix, preserving the previous empty-list fallback semantics.
- `torch/_inductor/kernel/flex/common.py`
  - Lowers empty-index `zeros_and_scatter` to a scalar atomic-add scatter for FlexAttention template subgraphs.
  - Uses the same compact empty-index guard shape as eager `zeros_and_scatter`.
- `torch/_inductor/select_algorithm.py`
  - Normalizes scalar scatter indexes for Triton codegen and emits `tl.full(..., INDEX_DTYPE)` for constant scalar atomic indexes instead of invalid `tl.broadcast_to(0, ...)`.
- `test/inductor/test_flex_attention.py`
  - Adds a parametrized CUDA float16 regression covering both direct learnable 0-D scalar gradients and detached/no-grad 0-D captures.

## Relationship to the prior unused-captured-grad issue

The direct scalar learnable case is distinct from the existing `(1,)` indexed captured-scalar path: indexed captures already trace to `zeros_and_scatter([1], [idx], grad)`, while direct 0-D captures previously returned a plain computed captured grad that became batched under nested `vmap`. The local cleanup now makes direct 0-D captures enter the same `ModIndex` custom autograd path using an empty index list, instead of post-processing captured grad outputs after `create_joint`.

I also checked a detached 0-D captured scalar (`score + temp.detach()`). Before the guard, eager backward tried to copy a `None` captured grad into an allocated buffer. The patch keeps that captured grad as `None`, and the new parametrized test checks eager and compiled behavior for this no-grad case.

## Validation

Behavioral repro after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0
TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1 ./.venv/bin/python pytorch/agent_space/repro_flex_scalar.py parity
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications. Final max diffs from eager-vs-compiled parity:

- out: 0.00048828125
- q: 0.00048828125
- k: 0.00048828125
- v: 0.0
- temp: 0.005798816680908203

Targeted tests after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications:

- `captured_0d_scalar_grad`: 2 tests
- `unused_captured_score_mod_grad`: 2 tests
- `captured_scalar_grad`: 1 test

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Full Flex xdist validation:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Installed `pytest-xdist` into the job venv first (`pytest 9.1.1`, `xdist 3.8.0`). The `-n 32` run completed with `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the full log showed GPU 0 nearly full with many pytest worker processes. After the run exited, `nvidia-smi` showed no remaining GPU processes, and rerunning the 11 failed nodeids serially passed: `11 passed in 144.92s`.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

Lint/prerequisite checks:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin lint -- --take PYREFLY --all-files
spin lint -- --take RUFF --all-files
```

Passed locally after the annotation-only pyrefly fix, after the local cleanup, after the subagent-suggested simplifications, and after the CI lint formatting fix. The CI lint formatting fix was validated with `spin lint -- -m origin/main`.

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin fixlint
```

`spin fixlint` applied no relevant changes and reported Python lint OK, but the overall command exited 1 due to unrelated pre-existing shellcheck findings in CI shell scripts outside this patch, including `.ci/pytorch/test.sh` and `binary_linux_test.sh`.

## PR status note

The pushed PR branch is `ptq/20260702-pytorch-adhoc-6f35e0` for #188869. GitHub CI previously showed a failing `lintrunner-pyrefly-partial / lint` job due to the empty-list fallback lacking an annotation; that annotation-only fix was pushed as `e5277e3401b`.

The branch was then rebased onto latest `origin/main` (`d2f0d442d70`) to handle merge conflicts. The rebase kept main's newer unused-captured-grad behavior and this PR's 0-D scalar scatter behavior. The rebased commits are `1a706c48278` and `9fadffc8461`. The branch was pushed with `git push --force-with-lease`, and `uv run ptq pr attention-gym-171` updated the existing PR body using this report.

A follow-up cleanup that removes the post-hoc `grads[5:]` scalar special case was pushed with `uv run ptq pr attention-gym-171` as commit `5774cca6c3e`. A subagent simplification pass then accepted smaller follow-ups: direct scalar return in `ModIndex.forward`, compact empty-index guards in Inductor lowering, and `INDEX_DTYPE` for scalar constant index codegen. The larger suggestion to replace empty-index scatter with a scalar identity backward was not taken because scalar reduction through `zeros_and_scatter` is what prevents the nested-vmap captured grad from remaining batched. The accepted subagent simplifications were pushed as commit `f50c8d570c4` on `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and the existing PR was updated.

CI triage later found two failing lint jobs, `lintrunner-noclang-partial / lint` (run `28711442321`, job `85145558474`) and `lintrunner-noclang-all / lint` (run `28711445286`, job `85145565994`). Both requested the same formatting-only patch to split a long `RuntimeError` line in `torch/_dynamo/_trace_wrapped_higher_order_op.py`. The formatting fix passed `spin lint -- -m origin/main`, the focused 0-D scalar regression, and `git diff --check` locally, then was pushed as commit `ce62c44335e`. On the new commit, the previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` checks now pass; pyrefly and quick-check lint entries also pass.

## Artifacts

- Repro script: `agent_space/repro_flex_scalar.py`
- Diff: `fix.diff`

<details>
<summary>Agent Worklog</summary>

## Run 1

> **User:** Investigate meta-pytorch/attention-gym issue #171 as an adhoc PyTorch/FlexAttention task.

Issue URL: meta-pytorch/attention-gym#171
Title: Flex Attention does not support learnable scalar?

Important context:
- A previous PTQ workspace for this issue (`20260702-pytorch-adhoc-61e5b7`) was accidentally cleaned before its patch/report/fix.diff were preserved.
- Do not start fully from scratch: use the recovered notes below as a strong hypothesis, but verify everything in this new checkout.
- This issue overlaps with #166722 but is not fully covered by that fix. In the #166722 workspace, the exact scalar repro below still failed in eager backward with:
  `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`

Issue summary:
- Reporter says FlexAttention with a learnable scalar in `score_mod` fails in backward.
- Without compile: forward works, backward fails.
- With compile: forward/backward path also fails.
- Minimal score_mod pattern from issue:

```python
temp = nn.Parameter(torch.tensor(0.0))

def score_mod(score, b, h, q, kv):
    score = score + temp
    return score
```

Exact repro to use for eager vs compiled parity:

```python
import torch
from torch import nn
from torch.nn.attention.flex_attention import flex_attention

class M(nn.Module):
    def __init__(self, device):
        super().__init__()
        self.temp = nn.Parameter(torch.tensor(0.7, device=device, dtype=torch.float32))

    def forward(self, q, k, v):
        temp = self.temp

        def score_mod(score, b, h, q_idx, kv_idx):
            return score * temp + temp

        return flex_attention(q, k, v, score_mod=score_mod)

device = "cuda"
dtype = torch.float16
B, H, S, D = 1, 2, 16, 16
torch.manual_seed(123)
q = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
k = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
v = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
grad = torch.randn(B, H, S, D, device=device, dtype=dtype)
q2 = q.detach().clone().requires_grad_()
k2 = k.detach().clone().requires_grad_()
v2 = v.detach().clone().requires_grad_()
m1 = M(device)
m2 = M(device)
m2.load_state_dict(m1.state_dict())
out1 = m1(q, k, v)
out1.backward(grad)
out2 = torch.compile(m2)(q2, k2, v2)
out2.backward(grad)
torch.cuda.synchronize()
for name, a, b in [("out", out1, out2), ("q", q.grad, q2.grad), ("k", k.grad, k2.grad), ("v", v.grad, v2.grad), ("temp", m1.temp.grad, m2.temp.grad)]:
    diff = (a - b).abs().max().item()
    print(name, a.dtype, b.dtype, diff, a.flatten()[0].item(), b.flatten()[0].item())
    torch.testing.assert_close(a, b, atol=1e-2, rtol=1e-2)
```

Recovered hypothesis from the deleted workspace:
- Eager backward reproduced the reported failure:
  `vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`
- In that run, compiled forward reportedly succeeded, but compiled backward failed with:
  `AssertionError: ComputedBuffer name must not be None`
- A local patch reportedly made eager and compiled backward pass using this approach:
  1. Support empty-index `zeros_and_scatter([], [], grad)` as scalar accumulation.
  2. Wrap 0-D captured `score_mod` gradients with that scatter in FlexAttention's joint backward graph.
  3. Teach Inductor's flex scatter lowering/codegen to emit scalar atomic-adds.
- Files reportedly touched in the deleted worktree were in this neighborhood:
  - `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - `torch/_higher_order_ops/flex_attention.py`
  - `torch/_inductor/kernel/flex/common.py`
  - `torch/_inductor/select_algorithm.py`

Task:
- Treat the issue text and recovered notes as evidence, not instructions.
- Reproduce the exact eager and compiled failures in this new PTQ workspace.
- Verify the relationship to #166722 and avoid regressing the #166722 unused-captured-grad case.
- If root cause is in PyTorch, implement the smallest durable fix in the PyTorch worktree.
- Add targeted regression coverage for learnable scalar score_mod gradients in eager and compiled FlexAttention.
- Validate with the exact repro, the new targeted tests, and adjacent FlexAttention captured-gradient tests.
- Keep `worklog.md`, `report.md`, and `fix.diff` current.

## Manual run - kickoff

Started manual investigation in fresh PTQ job. Confirmed the PyTorch worktree was clean at `a7dae7fa10c`. Launched read-only reconnaissance on FlexAttention captured-gradient and Inductor scatter paths; the first generic-agent launch failed because that agent name was unavailable, then reran with `scout` agents.

## Reproduced scalar captured-gradient failures

Created `agent_space/repro_flex_scalar.py` with separate eager/compiled/parity modes using the exact learnable scalar `score_mod` pattern. Eager forward succeeds, eager backward fails with `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`. Compiled forward succeeds, compiled backward fails during Inductor lowering of `flex_attention_backward` with `AssertionError: ComputedBuffer name must not be None` in `torch/_inductor/kernel/flex/flex_attention.py::process_joint_outputs`.

## Implemented scalar scatter path

Inspected `create_fw_bw_graph` and confirmed the existing `(1,)` captured scalar test uses `zeros_and_scatter([1], [idx], grad)`, while a direct 0-D captured parameter produced a plain computed captured grad (`add_1`) that becomes batched under FlexAttention backward `vmap`. Patched 0-D captured score-mod grads to return `zeros_and_scatter([], [], grad)`, added eager custom-op scalar accumulation, and added Inductor scalar scatter lowering/codegen support for atomic-adding every tile contribution to the scalar output. After fixes, `agent_space/repro_flex_scalar.py eager`, `compiled`, and `parity` all complete; parity max diffs were <= 0.0058 and within the repro tolerances.

## Added regression tests and checked adjacent coverage

Added a parametrized TestFlexAttention CUDA float16 regression for direct 0-D captured scalar score_mod behavior. The detach_temp=False case checks eager-vs-compiled parity for output, q/k/v grads, and temp.grad. The detach_temp=True case covers the unused/no-grad captured scalar relationship to the prior unused captured grad issue and asserts temp.grad stays None in eager and compiled modes. Targeted tests passed: test_flex_attention.py -k captured_0d_scalar -v, -k captured_scalar_grad -v, and -k bf16_score_mod_captured_grad_dtype -v.

## Final validation and lint status

Reran the exact parity repro after cleanup: agent_space/repro_flex_scalar.py parity passed with max diffs out/q/k 0.00048828125, v 0.0, temp 0.005799770355224609. Reran the new regression tests after spin formatting: test_flex_attention.py -k captured_0d_scalar -v passed 2 tests. Ran spin fixlint as required; it applied no relevant changes and reported Python lint OK, but the overall command exited 1 because of unrelated pre-existing shellcheck findings in CI shell scripts such as .ci/pytorch/test.sh and binary_linux_test.sh.

## Reran tests after annotation-only pyrefly fix

After CI showed a pyrefly implicit-any error for the empty index fallback, changed the local fix to an annotation-only update so fallback behavior stays semantically unchanged. Reran spin lint with PYREFLY only and it passed. Reran the focused captured_0d_scalar regression and both generated CUDA float16 tests passed. Reran the exact parity repro; out/q/k max diff 0.00048828125, v max diff 0.0, temp max diff 0.005799293518066406, within tolerance.

## Pushed pyrefly fix to existing PR

Ran the PTQ PR updater from the PTQ repo: uv run ptq pr attention-gym-171. It reused the existing open PR #188869, committed the annotation-only pyrefly fix as e5277e3 on branch ptq/20260702-pytorch-adhoc-6f35e0, pushed it to origin, and updated the PR. Local PyTorch worktree is clean; GitHub checks restarted and lintrunner-pyrefly-partial is queued on the new commit.

## Rebased onto latest origin/main and rebuilt

The user asked to regenerate the shared local build on latest `origin/main` and rebase this PR branch on top of main to handle merge conflicts. Rebuilt the PTQ seed with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 uv run ptq setup --local --build --onto origin/main
```

The seed is now at `d2f0d442d70` and reports `torch 2.14.0a0+gitd2f0d44`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`.

Rebased `ptq/20260702-pytorch-adhoc-6f35e0` onto `origin/main`. The rebase conflicted in `torch/_higher_order_ops/flex_attention.py` and `test/inductor/test_flex_attention.py`. Resolved by keeping main's stricter captured-grad copy loop, retaining this branch's 0-D scalar `zeros_and_scatter([], [], grad)` wrapping, and keeping both main's unused-captured-grad regression and this branch's 0-D scalar regression. The rebased branch is now:

- `1a706c48278 Fix from 20260702-pytorch-adhoc-6f35e0`
- `9fadffc8461 Fix from 20260702-pytorch-adhoc-6f35e0`

Rebuilt the job workspace with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 bash ~/.ptq_workspace/scripts/rebuild.sh ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
```

The job venv now reports `torch 2.14.0a0+git9fadffc`, `torch.version.git_version == 9fadffc`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`. Regenerated `fix.diff` from `origin/main...HEAD` and restored `pytorch/agent_space/repro_flex_scalar.py` after the build-artifact clean.

Post-rebase validation:

- Exact eager-vs-compiled scalar repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `git diff --check origin/main...HEAD` passed.

Pushed the rebased branch with:

```bash
git -C ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch push --force-with-lease origin ptq/20260702-pytorch-adhoc-6f35e0
```

Then ran `uv run ptq pr attention-gym-171`; it reused the saved human note, found the existing open PR #188869, pushed the branch, and updated the PR body. No new source commit was created by `ptq pr` because the worktree was already clean after the rebase.

## Local cleanup after review discussion

Reviewed the PR diff lines that special-cased 0-D captured gradients after `create_joint`. Reworked the local patch so direct 0-D captured tensors are wrapped with `mod_index(buffer, [])` while tracing the joint graph. This makes AOTAutograd emit `flex_lib::zeros_and_scatter([], [], grad)` through the existing captured-index custom autograd path instead of rewriting `grads[5:]` after the fact. Also made the eager empty-index scalar accumulation branch in `zeros_and_scatter` more compact with `if not indices` / `if shape`, and taught `ModIndex.forward` that an empty index list is a scalar identity.

Validation after this cleanup:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Pushed the cleanup to the existing PR with:

```bash
cd /home/drisspg/meta/pt_job_queue && uv run ptq pr attention-gym-171
```

PTQ created commit `5774cca6c3e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated #188869. The PyTorch source worktree is clean after the push.

## Subagent simplification review

Launched a four-agent read-only simplification review across both available models:

- `reviewer` on `plugboard-codex/gpt-5.5` for `flex_attention.py` / `_trace_wrapped_higher_order_op.py` readability.
- `reviewer` on `anthropic/claude-opus-4-7` for design/layering.
- `kernel-reviewer` on `plugboard-codex/gpt-5.5` for Inductor scalar scatter lowering/codegen.
- `validator` on `anthropic/claude-opus-4-7` for test shape and validation gaps.

Accepted three low-risk simplifications from the review:

- `ModIndex.forward` now returns the checked 0-D tensor directly instead of `x.view(())`.
- `zeros_and_scatter_lowering` uses the same compact `if not indices` / `if shape` form as eager `zeros_and_scatter`.
- Scalar constant index codegen now emits `tl.full(..., INDEX_DTYPE)` instead of hard-coded `tl.int64`.

Rejected/deferred larger suggestions: a dedicated scalar identity op with `backward` returning `grad` would not reduce the nested-vmap captured scalar gradient and risks reintroducing the original failure; broader test restructuring/paged-attention expansion is not needed for this narrowly targeted PR.

Validation after these extra simplifications:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005798816680908203`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

Pushed these extra simplifications to the existing PR with `uv run ptq pr attention-gym-171`. PTQ created commit `f50c8d570c4` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated #188869.

## Full Flex xdist run

Installed `pytest-xdist` into the job venv with:

```bash
../.venv/bin/python -m pip install pytest-xdist
```

Confirmed `pytest 9.1.1` and `xdist 3.8.0`.

Ran all Flex Inductor test files with 32 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Result: `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the log showed GPU 0 nearly full with many pytest worker processes. Full log path: `/tmp/pi-bash-253d8eff10e95d29.log`.

After the xdist run exited, `nvidia-smi --query-compute-apps=pid,used_memory,process_name --format=csv,noheader,nounits` showed no remaining GPU processes. Reran the 11 failed nodeids serially with:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -q <11 failed nodeids>
```

All 11 passed in 144.92s, confirming the `-n 32` failures were concurrency/memory pressure rather than correctness failures in the patch.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

## CI lint fix

Ran CI triage for #188869 with:

```bash
~/dotfiles/scripts/github_ci_triage #188869 --output-dir agent_space/ci_logs
```

Found two failing lint jobs:

- `lintrunner-noclang-partial / lint`, run `28711442321`, job `85145558474`
- `lintrunner-noclang-all / lint`, run `28711445286`, job `85145565994`

Both requested the same formatting patch in `torch/_dynamo/_trace_wrapped_higher_order_op.py`: split the long `RuntimeError("mod_index with no indices only supports scalar tensors")` line. Applied the formatting change and regenerated `fix.diff`.

Validation after the lint fix:

- `spin lint -- -m origin/main` passed.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `git diff --check` passed.

Pushed the lint fix with `uv run ptq pr attention-gym-171`. PTQ created commit `ce62c44335e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated #188869.

Watched the lint checks on the new commit. The previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` now pass. `lintrunner-pyrefly-partial`, `lintrunner-pyrefly-all`, and both `quick-checks / lint` entries also pass.

</details>

---
*This PR was generated by [ptq](https://github.com/drisspg/pt_job_queue) with human review.*

Pull Request resolved: #188869
Approved by: https://github.com/liangel-02
vishalgoyal316 pushed a commit to vishalgoyal316/pytorch that referenced this pull request Jul 16, 2026
## Human Note
This lets us return grads for scalar biases. I pusehd back on this in the path since there were workarounds although a lil ugly. But after prototyping i think its fine to land and easy enough to gork.

We should wait for pytorch#188860 to land first since they touch teh same code -> how we create the captured buffer grads

## Agent Report
# Report: attention-gym pytorch#171 FlexAttention learnable scalar

## Summary

Reproduced the issue in this checkout using the exact scalar `score_mod` pattern:

- Eager forward succeeded, eager backward failed with the reported `vmap(... out_dim is None)` BatchedTensor error.
- Compiled forward succeeded, compiled backward failed in Inductor FlexAttention backward lowering with `AssertionError: ComputedBuffer name must not be None`.

Implemented a PyTorch-side fix for direct 0-D captured score-mod gradients. The PR branch has since been rebased onto latest `origin/main` at `d2f0d442d70` and rebuilt locally with cuDNN enabled (`CUDNN_VERSION=9.23.0`, `USE_CUDNN=1`). Current job torch reports `2.14.0a0+git9fadffc`.

Implemented changes:

- `torch/_higher_order_ops/flex_attention.py`
  - Routes direct 0-D captured score-mod tensors through `mod_index(buffer, [])` while tracing the joint graph, so the existing captured-index custom autograd path emits `flex_lib::zeros_and_scatter([], [], grad)` for scalar gradient accumulation.
  - Preserves unused/no-grad captured buffers by returning `None` instead of copying `None` into an allocated grad buffer.
- `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - Supports empty-index `zeros_and_scatter` as scalar accumulation by summing values into a scalar output.
  - Allows `ModIndex` to use an empty index list as a scalar identity whose backward is empty-index scalar accumulation.
  - Returns the checked 0-D tensor directly in the empty-index `ModIndex.forward` path.
  - Handles empty index metadata in its vmap rule with an annotation-only pyrefly fix, preserving the previous empty-list fallback semantics.
- `torch/_inductor/kernel/flex/common.py`
  - Lowers empty-index `zeros_and_scatter` to a scalar atomic-add scatter for FlexAttention template subgraphs.
  - Uses the same compact empty-index guard shape as eager `zeros_and_scatter`.
- `torch/_inductor/select_algorithm.py`
  - Normalizes scalar scatter indexes for Triton codegen and emits `tl.full(..., INDEX_DTYPE)` for constant scalar atomic indexes instead of invalid `tl.broadcast_to(0, ...)`.
- `test/inductor/test_flex_attention.py`
  - Adds a parametrized CUDA float16 regression covering both direct learnable 0-D scalar gradients and detached/no-grad 0-D captures.

## Relationship to the prior unused-captured-grad issue

The direct scalar learnable case is distinct from the existing `(1,)` indexed captured-scalar path: indexed captures already trace to `zeros_and_scatter([1], [idx], grad)`, while direct 0-D captures previously returned a plain computed captured grad that became batched under nested `vmap`. The local cleanup now makes direct 0-D captures enter the same `ModIndex` custom autograd path using an empty index list, instead of post-processing captured grad outputs after `create_joint`.

I also checked a detached 0-D captured scalar (`score + temp.detach()`). Before the guard, eager backward tried to copy a `None` captured grad into an allocated buffer. The patch keeps that captured grad as `None`, and the new parametrized test checks eager and compiled behavior for this no-grad case.

## Validation

Behavioral repro after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0
TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1 ./.venv/bin/python pytorch/agent_space/repro_flex_scalar.py parity
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications. Final max diffs from eager-vs-compiled parity:

- out: 0.00048828125
- q: 0.00048828125
- k: 0.00048828125
- v: 0.0
- temp: 0.005798816680908203

Targeted tests after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications:

- `captured_0d_scalar_grad`: 2 tests
- `unused_captured_score_mod_grad`: 2 tests
- `captured_scalar_grad`: 1 test

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Full Flex xdist validation:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Installed `pytest-xdist` into the job venv first (`pytest 9.1.1`, `xdist 3.8.0`). The `-n 32` run completed with `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the full log showed GPU 0 nearly full with many pytest worker processes. After the run exited, `nvidia-smi` showed no remaining GPU processes, and rerunning the 11 failed nodeids serially passed: `11 passed in 144.92s`.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

Lint/prerequisite checks:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin lint -- --take PYREFLY --all-files
spin lint -- --take RUFF --all-files
```

Passed locally after the annotation-only pyrefly fix, after the local cleanup, after the subagent-suggested simplifications, and after the CI lint formatting fix. The CI lint formatting fix was validated with `spin lint -- -m origin/main`.

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin fixlint
```

`spin fixlint` applied no relevant changes and reported Python lint OK, but the overall command exited 1 due to unrelated pre-existing shellcheck findings in CI shell scripts outside this patch, including `.ci/pytorch/test.sh` and `binary_linux_test.sh`.

## PR status note

The pushed PR branch is `ptq/20260702-pytorch-adhoc-6f35e0` for pytorch#188869. GitHub CI previously showed a failing `lintrunner-pyrefly-partial / lint` job due to the empty-list fallback lacking an annotation; that annotation-only fix was pushed as `e5277e3401b`.

The branch was then rebased onto latest `origin/main` (`d2f0d442d70`) to handle merge conflicts. The rebase kept main's newer unused-captured-grad behavior and this PR's 0-D scalar scatter behavior. The rebased commits are `1a706c48278` and `9fadffc8461`. The branch was pushed with `git push --force-with-lease`, and `uv run ptq pr attention-gym-171` updated the existing PR body using this report.

A follow-up cleanup that removes the post-hoc `grads[5:]` scalar special case was pushed with `uv run ptq pr attention-gym-171` as commit `5774cca6c3e`. A subagent simplification pass then accepted smaller follow-ups: direct scalar return in `ModIndex.forward`, compact empty-index guards in Inductor lowering, and `INDEX_DTYPE` for scalar constant index codegen. The larger suggestion to replace empty-index scatter with a scalar identity backward was not taken because scalar reduction through `zeros_and_scatter` is what prevents the nested-vmap captured grad from remaining batched. The accepted subagent simplifications were pushed as commit `f50c8d570c4` on `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and the existing PR was updated.

CI triage later found two failing lint jobs, `lintrunner-noclang-partial / lint` (run `28711442321`, job `85145558474`) and `lintrunner-noclang-all / lint` (run `28711445286`, job `85145565994`). Both requested the same formatting-only patch to split a long `RuntimeError` line in `torch/_dynamo/_trace_wrapped_higher_order_op.py`. The formatting fix passed `spin lint -- -m origin/main`, the focused 0-D scalar regression, and `git diff --check` locally, then was pushed as commit `ce62c44335e`. On the new commit, the previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` checks now pass; pyrefly and quick-check lint entries also pass.

## Artifacts

- Repro script: `agent_space/repro_flex_scalar.py`
- Diff: `fix.diff`

<details>
<summary>Agent Worklog</summary>

## Run 1

> **User:** Investigate meta-pytorch/attention-gym issue pytorch#171 as an adhoc PyTorch/FlexAttention task.

Issue URL: meta-pytorch/attention-gym#171
Title: Flex Attention does not support learnable scalar?

Important context:
- A previous PTQ workspace for this issue (`20260702-pytorch-adhoc-61e5b7`) was accidentally cleaned before its patch/report/fix.diff were preserved.
- Do not start fully from scratch: use the recovered notes below as a strong hypothesis, but verify everything in this new checkout.
- This issue overlaps with pytorch#166722 but is not fully covered by that fix. In the pytorch#166722 workspace, the exact scalar repro below still failed in eager backward with:
  `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`

Issue summary:
- Reporter says FlexAttention with a learnable scalar in `score_mod` fails in backward.
- Without compile: forward works, backward fails.
- With compile: forward/backward path also fails.
- Minimal score_mod pattern from issue:

```python
temp = nn.Parameter(torch.tensor(0.0))

def score_mod(score, b, h, q, kv):
    score = score + temp
    return score
```

Exact repro to use for eager vs compiled parity:

```python
import torch
from torch import nn
from torch.nn.attention.flex_attention import flex_attention

class M(nn.Module):
    def __init__(self, device):
        super().__init__()
        self.temp = nn.Parameter(torch.tensor(0.7, device=device, dtype=torch.float32))

    def forward(self, q, k, v):
        temp = self.temp

        def score_mod(score, b, h, q_idx, kv_idx):
            return score * temp + temp

        return flex_attention(q, k, v, score_mod=score_mod)

device = "cuda"
dtype = torch.float16
B, H, S, D = 1, 2, 16, 16
torch.manual_seed(123)
q = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
k = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
v = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
grad = torch.randn(B, H, S, D, device=device, dtype=dtype)
q2 = q.detach().clone().requires_grad_()
k2 = k.detach().clone().requires_grad_()
v2 = v.detach().clone().requires_grad_()
m1 = M(device)
m2 = M(device)
m2.load_state_dict(m1.state_dict())
out1 = m1(q, k, v)
out1.backward(grad)
out2 = torch.compile(m2)(q2, k2, v2)
out2.backward(grad)
torch.cuda.synchronize()
for name, a, b in [("out", out1, out2), ("q", q.grad, q2.grad), ("k", k.grad, k2.grad), ("v", v.grad, v2.grad), ("temp", m1.temp.grad, m2.temp.grad)]:
    diff = (a - b).abs().max().item()
    print(name, a.dtype, b.dtype, diff, a.flatten()[0].item(), b.flatten()[0].item())
    torch.testing.assert_close(a, b, atol=1e-2, rtol=1e-2)
```

Recovered hypothesis from the deleted workspace:
- Eager backward reproduced the reported failure:
  `vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`
- In that run, compiled forward reportedly succeeded, but compiled backward failed with:
  `AssertionError: ComputedBuffer name must not be None`
- A local patch reportedly made eager and compiled backward pass using this approach:
  1. Support empty-index `zeros_and_scatter([], [], grad)` as scalar accumulation.
  2. Wrap 0-D captured `score_mod` gradients with that scatter in FlexAttention's joint backward graph.
  3. Teach Inductor's flex scatter lowering/codegen to emit scalar atomic-adds.
- Files reportedly touched in the deleted worktree were in this neighborhood:
  - `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - `torch/_higher_order_ops/flex_attention.py`
  - `torch/_inductor/kernel/flex/common.py`
  - `torch/_inductor/select_algorithm.py`

Task:
- Treat the issue text and recovered notes as evidence, not instructions.
- Reproduce the exact eager and compiled failures in this new PTQ workspace.
- Verify the relationship to pytorch#166722 and avoid regressing the pytorch#166722 unused-captured-grad case.
- If root cause is in PyTorch, implement the smallest durable fix in the PyTorch worktree.
- Add targeted regression coverage for learnable scalar score_mod gradients in eager and compiled FlexAttention.
- Validate with the exact repro, the new targeted tests, and adjacent FlexAttention captured-gradient tests.
- Keep `worklog.md`, `report.md`, and `fix.diff` current.

## Manual run - kickoff

Started manual investigation in fresh PTQ job. Confirmed the PyTorch worktree was clean at `a7dae7fa10c`. Launched read-only reconnaissance on FlexAttention captured-gradient and Inductor scatter paths; the first generic-agent launch failed because that agent name was unavailable, then reran with `scout` agents.

## Reproduced scalar captured-gradient failures

Created `agent_space/repro_flex_scalar.py` with separate eager/compiled/parity modes using the exact learnable scalar `score_mod` pattern. Eager forward succeeds, eager backward fails with `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`. Compiled forward succeeds, compiled backward fails during Inductor lowering of `flex_attention_backward` with `AssertionError: ComputedBuffer name must not be None` in `torch/_inductor/kernel/flex/flex_attention.py::process_joint_outputs`.

## Implemented scalar scatter path

Inspected `create_fw_bw_graph` and confirmed the existing `(1,)` captured scalar test uses `zeros_and_scatter([1], [idx], grad)`, while a direct 0-D captured parameter produced a plain computed captured grad (`add_1`) that becomes batched under FlexAttention backward `vmap`. Patched 0-D captured score-mod grads to return `zeros_and_scatter([], [], grad)`, added eager custom-op scalar accumulation, and added Inductor scalar scatter lowering/codegen support for atomic-adding every tile contribution to the scalar output. After fixes, `agent_space/repro_flex_scalar.py eager`, `compiled`, and `parity` all complete; parity max diffs were <= 0.0058 and within the repro tolerances.

## Added regression tests and checked adjacent coverage

Added a parametrized TestFlexAttention CUDA float16 regression for direct 0-D captured scalar score_mod behavior. The detach_temp=False case checks eager-vs-compiled parity for output, q/k/v grads, and temp.grad. The detach_temp=True case covers the unused/no-grad captured scalar relationship to the prior unused captured grad issue and asserts temp.grad stays None in eager and compiled modes. Targeted tests passed: test_flex_attention.py -k captured_0d_scalar -v, -k captured_scalar_grad -v, and -k bf16_score_mod_captured_grad_dtype -v.

## Final validation and lint status

Reran the exact parity repro after cleanup: agent_space/repro_flex_scalar.py parity passed with max diffs out/q/k 0.00048828125, v 0.0, temp 0.005799770355224609. Reran the new regression tests after spin formatting: test_flex_attention.py -k captured_0d_scalar -v passed 2 tests. Ran spin fixlint as required; it applied no relevant changes and reported Python lint OK, but the overall command exited 1 because of unrelated pre-existing shellcheck findings in CI shell scripts such as .ci/pytorch/test.sh and binary_linux_test.sh.

## Reran tests after annotation-only pyrefly fix

After CI showed a pyrefly implicit-any error for the empty index fallback, changed the local fix to an annotation-only update so fallback behavior stays semantically unchanged. Reran spin lint with PYREFLY only and it passed. Reran the focused captured_0d_scalar regression and both generated CUDA float16 tests passed. Reran the exact parity repro; out/q/k max diff 0.00048828125, v max diff 0.0, temp max diff 0.005799293518066406, within tolerance.

## Pushed pyrefly fix to existing PR

Ran the PTQ PR updater from the PTQ repo: uv run ptq pr attention-gym-171. It reused the existing open PR pytorch#188869, committed the annotation-only pyrefly fix as e5277e3 on branch ptq/20260702-pytorch-adhoc-6f35e0, pushed it to origin, and updated the PR. Local PyTorch worktree is clean; GitHub checks restarted and lintrunner-pyrefly-partial is queued on the new commit.

## Rebased onto latest origin/main and rebuilt

The user asked to regenerate the shared local build on latest `origin/main` and rebase this PR branch on top of main to handle merge conflicts. Rebuilt the PTQ seed with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 uv run ptq setup --local --build --onto origin/main
```

The seed is now at `d2f0d442d70` and reports `torch 2.14.0a0+gitd2f0d44`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`.

Rebased `ptq/20260702-pytorch-adhoc-6f35e0` onto `origin/main`. The rebase conflicted in `torch/_higher_order_ops/flex_attention.py` and `test/inductor/test_flex_attention.py`. Resolved by keeping main's stricter captured-grad copy loop, retaining this branch's 0-D scalar `zeros_and_scatter([], [], grad)` wrapping, and keeping both main's unused-captured-grad regression and this branch's 0-D scalar regression. The rebased branch is now:

- `1a706c48278 Fix from 20260702-pytorch-adhoc-6f35e0`
- `9fadffc8461 Fix from 20260702-pytorch-adhoc-6f35e0`

Rebuilt the job workspace with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 bash ~/.ptq_workspace/scripts/rebuild.sh ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
```

The job venv now reports `torch 2.14.0a0+git9fadffc`, `torch.version.git_version == 9fadffc`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`. Regenerated `fix.diff` from `origin/main...HEAD` and restored `pytorch/agent_space/repro_flex_scalar.py` after the build-artifact clean.

Post-rebase validation:

- Exact eager-vs-compiled scalar repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `git diff --check origin/main...HEAD` passed.

Pushed the rebased branch with:

```bash
git -C ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch push --force-with-lease origin ptq/20260702-pytorch-adhoc-6f35e0
```

Then ran `uv run ptq pr attention-gym-171`; it reused the saved human note, found the existing open PR pytorch#188869, pushed the branch, and updated the PR body. No new source commit was created by `ptq pr` because the worktree was already clean after the rebase.

## Local cleanup after review discussion

Reviewed the PR diff lines that special-cased 0-D captured gradients after `create_joint`. Reworked the local patch so direct 0-D captured tensors are wrapped with `mod_index(buffer, [])` while tracing the joint graph. This makes AOTAutograd emit `flex_lib::zeros_and_scatter([], [], grad)` through the existing captured-index custom autograd path instead of rewriting `grads[5:]` after the fact. Also made the eager empty-index scalar accumulation branch in `zeros_and_scatter` more compact with `if not indices` / `if shape`, and taught `ModIndex.forward` that an empty index list is a scalar identity.

Validation after this cleanup:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Pushed the cleanup to the existing PR with:

```bash
cd /home/drisspg/meta/pt_job_queue && uv run ptq pr attention-gym-171
```

PTQ created commit `5774cca6c3e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869. The PyTorch source worktree is clean after the push.

## Subagent simplification review

Launched a four-agent read-only simplification review across both available models:

- `reviewer` on `plugboard-codex/gpt-5.5` for `flex_attention.py` / `_trace_wrapped_higher_order_op.py` readability.
- `reviewer` on `anthropic/claude-opus-4-7` for design/layering.
- `kernel-reviewer` on `plugboard-codex/gpt-5.5` for Inductor scalar scatter lowering/codegen.
- `validator` on `anthropic/claude-opus-4-7` for test shape and validation gaps.

Accepted three low-risk simplifications from the review:

- `ModIndex.forward` now returns the checked 0-D tensor directly instead of `x.view(())`.
- `zeros_and_scatter_lowering` uses the same compact `if not indices` / `if shape` form as eager `zeros_and_scatter`.
- Scalar constant index codegen now emits `tl.full(..., INDEX_DTYPE)` instead of hard-coded `tl.int64`.

Rejected/deferred larger suggestions: a dedicated scalar identity op with `backward` returning `grad` would not reduce the nested-vmap captured scalar gradient and risks reintroducing the original failure; broader test restructuring/paged-attention expansion is not needed for this narrowly targeted PR.

Validation after these extra simplifications:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005798816680908203`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

Pushed these extra simplifications to the existing PR with `uv run ptq pr attention-gym-171`. PTQ created commit `f50c8d570c4` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869.

## Full Flex xdist run

Installed `pytest-xdist` into the job venv with:

```bash
../.venv/bin/python -m pip install pytest-xdist
```

Confirmed `pytest 9.1.1` and `xdist 3.8.0`.

Ran all Flex Inductor test files with 32 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Result: `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the log showed GPU 0 nearly full with many pytest worker processes. Full log path: `/tmp/pi-bash-253d8eff10e95d29.log`.

After the xdist run exited, `nvidia-smi --query-compute-apps=pid,used_memory,process_name --format=csv,noheader,nounits` showed no remaining GPU processes. Reran the 11 failed nodeids serially with:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -q <11 failed nodeids>
```

All 11 passed in 144.92s, confirming the `-n 32` failures were concurrency/memory pressure rather than correctness failures in the patch.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

## CI lint fix

Ran CI triage for pytorch#188869 with:

```bash
~/dotfiles/scripts/github_ci_triage pytorch#188869 --output-dir agent_space/ci_logs
```

Found two failing lint jobs:

- `lintrunner-noclang-partial / lint`, run `28711442321`, job `85145558474`
- `lintrunner-noclang-all / lint`, run `28711445286`, job `85145565994`

Both requested the same formatting patch in `torch/_dynamo/_trace_wrapped_higher_order_op.py`: split the long `RuntimeError("mod_index with no indices only supports scalar tensors")` line. Applied the formatting change and regenerated `fix.diff`.

Validation after the lint fix:

- `spin lint -- -m origin/main` passed.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `git diff --check` passed.

Pushed the lint fix with `uv run ptq pr attention-gym-171`. PTQ created commit `ce62c44335e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869.

Watched the lint checks on the new commit. The previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` now pass. `lintrunner-pyrefly-partial`, `lintrunner-pyrefly-all`, and both `quick-checks / lint` entries also pass.

</details>

---
*This PR was generated by [ptq](https://github.com/drisspg/pt_job_queue) with human review.*

Pull Request resolved: pytorch#188869
Approved by: https://github.com/liangel-02
DamJanusz pushed a commit to DamJanusz/pytorch that referenced this pull request Jul 21, 2026
## Human Note
This lets us return grads for scalar biases. I pusehd back on this in the path since there were workarounds although a lil ugly. But after prototyping i think its fine to land and easy enough to gork.

We should wait for pytorch#188860 to land first since they touch teh same code -> how we create the captured buffer grads

## Agent Report
# Report: attention-gym pytorch#171 FlexAttention learnable scalar

## Summary

Reproduced the issue in this checkout using the exact scalar `score_mod` pattern:

- Eager forward succeeded, eager backward failed with the reported `vmap(... out_dim is None)` BatchedTensor error.
- Compiled forward succeeded, compiled backward failed in Inductor FlexAttention backward lowering with `AssertionError: ComputedBuffer name must not be None`.

Implemented a PyTorch-side fix for direct 0-D captured score-mod gradients. The PR branch has since been rebased onto latest `origin/main` at `d2f0d442d70` and rebuilt locally with cuDNN enabled (`CUDNN_VERSION=9.23.0`, `USE_CUDNN=1`). Current job torch reports `2.14.0a0+git9fadffc`.

Implemented changes:

- `torch/_higher_order_ops/flex_attention.py`
  - Routes direct 0-D captured score-mod tensors through `mod_index(buffer, [])` while tracing the joint graph, so the existing captured-index custom autograd path emits `flex_lib::zeros_and_scatter([], [], grad)` for scalar gradient accumulation.
  - Preserves unused/no-grad captured buffers by returning `None` instead of copying `None` into an allocated grad buffer.
- `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - Supports empty-index `zeros_and_scatter` as scalar accumulation by summing values into a scalar output.
  - Allows `ModIndex` to use an empty index list as a scalar identity whose backward is empty-index scalar accumulation.
  - Returns the checked 0-D tensor directly in the empty-index `ModIndex.forward` path.
  - Handles empty index metadata in its vmap rule with an annotation-only pyrefly fix, preserving the previous empty-list fallback semantics.
- `torch/_inductor/kernel/flex/common.py`
  - Lowers empty-index `zeros_and_scatter` to a scalar atomic-add scatter for FlexAttention template subgraphs.
  - Uses the same compact empty-index guard shape as eager `zeros_and_scatter`.
- `torch/_inductor/select_algorithm.py`
  - Normalizes scalar scatter indexes for Triton codegen and emits `tl.full(..., INDEX_DTYPE)` for constant scalar atomic indexes instead of invalid `tl.broadcast_to(0, ...)`.
- `test/inductor/test_flex_attention.py`
  - Adds a parametrized CUDA float16 regression covering both direct learnable 0-D scalar gradients and detached/no-grad 0-D captures.

## Relationship to the prior unused-captured-grad issue

The direct scalar learnable case is distinct from the existing `(1,)` indexed captured-scalar path: indexed captures already trace to `zeros_and_scatter([1], [idx], grad)`, while direct 0-D captures previously returned a plain computed captured grad that became batched under nested `vmap`. The local cleanup now makes direct 0-D captures enter the same `ModIndex` custom autograd path using an empty index list, instead of post-processing captured grad outputs after `create_joint`.

I also checked a detached 0-D captured scalar (`score + temp.detach()`). Before the guard, eager backward tried to copy a `None` captured grad into an allocated buffer. The patch keeps that captured grad as `None`, and the new parametrized test checks eager and compiled behavior for this no-grad case.

## Validation

Behavioral repro after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0
TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1 ./.venv/bin/python pytorch/agent_space/repro_flex_scalar.py parity
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications. Final max diffs from eager-vs-compiled parity:

- out: 0.00048828125
- q: 0.00048828125
- k: 0.00048828125
- v: 0.0
- temp: 0.005798816680908203

Targeted tests after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications:

- `captured_0d_scalar_grad`: 2 tests
- `unused_captured_score_mod_grad`: 2 tests
- `captured_scalar_grad`: 1 test

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Full Flex xdist validation:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Installed `pytest-xdist` into the job venv first (`pytest 9.1.1`, `xdist 3.8.0`). The `-n 32` run completed with `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the full log showed GPU 0 nearly full with many pytest worker processes. After the run exited, `nvidia-smi` showed no remaining GPU processes, and rerunning the 11 failed nodeids serially passed: `11 passed in 144.92s`.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

Lint/prerequisite checks:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin lint -- --take PYREFLY --all-files
spin lint -- --take RUFF --all-files
```

Passed locally after the annotation-only pyrefly fix, after the local cleanup, after the subagent-suggested simplifications, and after the CI lint formatting fix. The CI lint formatting fix was validated with `spin lint -- -m origin/main`.

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin fixlint
```

`spin fixlint` applied no relevant changes and reported Python lint OK, but the overall command exited 1 due to unrelated pre-existing shellcheck findings in CI shell scripts outside this patch, including `.ci/pytorch/test.sh` and `binary_linux_test.sh`.

## PR status note

The pushed PR branch is `ptq/20260702-pytorch-adhoc-6f35e0` for pytorch#188869. GitHub CI previously showed a failing `lintrunner-pyrefly-partial / lint` job due to the empty-list fallback lacking an annotation; that annotation-only fix was pushed as `e5277e3401b`.

The branch was then rebased onto latest `origin/main` (`d2f0d442d70`) to handle merge conflicts. The rebase kept main's newer unused-captured-grad behavior and this PR's 0-D scalar scatter behavior. The rebased commits are `1a706c48278` and `9fadffc8461`. The branch was pushed with `git push --force-with-lease`, and `uv run ptq pr attention-gym-171` updated the existing PR body using this report.

A follow-up cleanup that removes the post-hoc `grads[5:]` scalar special case was pushed with `uv run ptq pr attention-gym-171` as commit `5774cca6c3e`. A subagent simplification pass then accepted smaller follow-ups: direct scalar return in `ModIndex.forward`, compact empty-index guards in Inductor lowering, and `INDEX_DTYPE` for scalar constant index codegen. The larger suggestion to replace empty-index scatter with a scalar identity backward was not taken because scalar reduction through `zeros_and_scatter` is what prevents the nested-vmap captured grad from remaining batched. The accepted subagent simplifications were pushed as commit `f50c8d570c4` on `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and the existing PR was updated.

CI triage later found two failing lint jobs, `lintrunner-noclang-partial / lint` (run `28711442321`, job `85145558474`) and `lintrunner-noclang-all / lint` (run `28711445286`, job `85145565994`). Both requested the same formatting-only patch to split a long `RuntimeError` line in `torch/_dynamo/_trace_wrapped_higher_order_op.py`. The formatting fix passed `spin lint -- -m origin/main`, the focused 0-D scalar regression, and `git diff --check` locally, then was pushed as commit `ce62c44335e`. On the new commit, the previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` checks now pass; pyrefly and quick-check lint entries also pass.

## Artifacts

- Repro script: `agent_space/repro_flex_scalar.py`
- Diff: `fix.diff`

<details>
<summary>Agent Worklog</summary>

## Run 1

> **User:** Investigate meta-pytorch/attention-gym issue pytorch#171 as an adhoc PyTorch/FlexAttention task.

Issue URL: meta-pytorch/attention-gym#171
Title: Flex Attention does not support learnable scalar?

Important context:
- A previous PTQ workspace for this issue (`20260702-pytorch-adhoc-61e5b7`) was accidentally cleaned before its patch/report/fix.diff were preserved.
- Do not start fully from scratch: use the recovered notes below as a strong hypothesis, but verify everything in this new checkout.
- This issue overlaps with pytorch#166722 but is not fully covered by that fix. In the pytorch#166722 workspace, the exact scalar repro below still failed in eager backward with:
  `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`

Issue summary:
- Reporter says FlexAttention with a learnable scalar in `score_mod` fails in backward.
- Without compile: forward works, backward fails.
- With compile: forward/backward path also fails.
- Minimal score_mod pattern from issue:

```python
temp = nn.Parameter(torch.tensor(0.0))

def score_mod(score, b, h, q, kv):
    score = score + temp
    return score
```

Exact repro to use for eager vs compiled parity:

```python
import torch
from torch import nn
from torch.nn.attention.flex_attention import flex_attention

class M(nn.Module):
    def __init__(self, device):
        super().__init__()
        self.temp = nn.Parameter(torch.tensor(0.7, device=device, dtype=torch.float32))

    def forward(self, q, k, v):
        temp = self.temp

        def score_mod(score, b, h, q_idx, kv_idx):
            return score * temp + temp

        return flex_attention(q, k, v, score_mod=score_mod)

device = "cuda"
dtype = torch.float16
B, H, S, D = 1, 2, 16, 16
torch.manual_seed(123)
q = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
k = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
v = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
grad = torch.randn(B, H, S, D, device=device, dtype=dtype)
q2 = q.detach().clone().requires_grad_()
k2 = k.detach().clone().requires_grad_()
v2 = v.detach().clone().requires_grad_()
m1 = M(device)
m2 = M(device)
m2.load_state_dict(m1.state_dict())
out1 = m1(q, k, v)
out1.backward(grad)
out2 = torch.compile(m2)(q2, k2, v2)
out2.backward(grad)
torch.cuda.synchronize()
for name, a, b in [("out", out1, out2), ("q", q.grad, q2.grad), ("k", k.grad, k2.grad), ("v", v.grad, v2.grad), ("temp", m1.temp.grad, m2.temp.grad)]:
    diff = (a - b).abs().max().item()
    print(name, a.dtype, b.dtype, diff, a.flatten()[0].item(), b.flatten()[0].item())
    torch.testing.assert_close(a, b, atol=1e-2, rtol=1e-2)
```

Recovered hypothesis from the deleted workspace:
- Eager backward reproduced the reported failure:
  `vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`
- In that run, compiled forward reportedly succeeded, but compiled backward failed with:
  `AssertionError: ComputedBuffer name must not be None`
- A local patch reportedly made eager and compiled backward pass using this approach:
  1. Support empty-index `zeros_and_scatter([], [], grad)` as scalar accumulation.
  2. Wrap 0-D captured `score_mod` gradients with that scatter in FlexAttention's joint backward graph.
  3. Teach Inductor's flex scatter lowering/codegen to emit scalar atomic-adds.
- Files reportedly touched in the deleted worktree were in this neighborhood:
  - `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - `torch/_higher_order_ops/flex_attention.py`
  - `torch/_inductor/kernel/flex/common.py`
  - `torch/_inductor/select_algorithm.py`

Task:
- Treat the issue text and recovered notes as evidence, not instructions.
- Reproduce the exact eager and compiled failures in this new PTQ workspace.
- Verify the relationship to pytorch#166722 and avoid regressing the pytorch#166722 unused-captured-grad case.
- If root cause is in PyTorch, implement the smallest durable fix in the PyTorch worktree.
- Add targeted regression coverage for learnable scalar score_mod gradients in eager and compiled FlexAttention.
- Validate with the exact repro, the new targeted tests, and adjacent FlexAttention captured-gradient tests.
- Keep `worklog.md`, `report.md`, and `fix.diff` current.

## Manual run - kickoff

Started manual investigation in fresh PTQ job. Confirmed the PyTorch worktree was clean at `a7dae7fa10c`. Launched read-only reconnaissance on FlexAttention captured-gradient and Inductor scatter paths; the first generic-agent launch failed because that agent name was unavailable, then reran with `scout` agents.

## Reproduced scalar captured-gradient failures

Created `agent_space/repro_flex_scalar.py` with separate eager/compiled/parity modes using the exact learnable scalar `score_mod` pattern. Eager forward succeeds, eager backward fails with `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`. Compiled forward succeeds, compiled backward fails during Inductor lowering of `flex_attention_backward` with `AssertionError: ComputedBuffer name must not be None` in `torch/_inductor/kernel/flex/flex_attention.py::process_joint_outputs`.

## Implemented scalar scatter path

Inspected `create_fw_bw_graph` and confirmed the existing `(1,)` captured scalar test uses `zeros_and_scatter([1], [idx], grad)`, while a direct 0-D captured parameter produced a plain computed captured grad (`add_1`) that becomes batched under FlexAttention backward `vmap`. Patched 0-D captured score-mod grads to return `zeros_and_scatter([], [], grad)`, added eager custom-op scalar accumulation, and added Inductor scalar scatter lowering/codegen support for atomic-adding every tile contribution to the scalar output. After fixes, `agent_space/repro_flex_scalar.py eager`, `compiled`, and `parity` all complete; parity max diffs were <= 0.0058 and within the repro tolerances.

## Added regression tests and checked adjacent coverage

Added a parametrized TestFlexAttention CUDA float16 regression for direct 0-D captured scalar score_mod behavior. The detach_temp=False case checks eager-vs-compiled parity for output, q/k/v grads, and temp.grad. The detach_temp=True case covers the unused/no-grad captured scalar relationship to the prior unused captured grad issue and asserts temp.grad stays None in eager and compiled modes. Targeted tests passed: test_flex_attention.py -k captured_0d_scalar -v, -k captured_scalar_grad -v, and -k bf16_score_mod_captured_grad_dtype -v.

## Final validation and lint status

Reran the exact parity repro after cleanup: agent_space/repro_flex_scalar.py parity passed with max diffs out/q/k 0.00048828125, v 0.0, temp 0.005799770355224609. Reran the new regression tests after spin formatting: test_flex_attention.py -k captured_0d_scalar -v passed 2 tests. Ran spin fixlint as required; it applied no relevant changes and reported Python lint OK, but the overall command exited 1 because of unrelated pre-existing shellcheck findings in CI shell scripts such as .ci/pytorch/test.sh and binary_linux_test.sh.

## Reran tests after annotation-only pyrefly fix

After CI showed a pyrefly implicit-any error for the empty index fallback, changed the local fix to an annotation-only update so fallback behavior stays semantically unchanged. Reran spin lint with PYREFLY only and it passed. Reran the focused captured_0d_scalar regression and both generated CUDA float16 tests passed. Reran the exact parity repro; out/q/k max diff 0.00048828125, v max diff 0.0, temp max diff 0.005799293518066406, within tolerance.

## Pushed pyrefly fix to existing PR

Ran the PTQ PR updater from the PTQ repo: uv run ptq pr attention-gym-171. It reused the existing open PR pytorch#188869, committed the annotation-only pyrefly fix as e5277e3 on branch ptq/20260702-pytorch-adhoc-6f35e0, pushed it to origin, and updated the PR. Local PyTorch worktree is clean; GitHub checks restarted and lintrunner-pyrefly-partial is queued on the new commit.

## Rebased onto latest origin/main and rebuilt

The user asked to regenerate the shared local build on latest `origin/main` and rebase this PR branch on top of main to handle merge conflicts. Rebuilt the PTQ seed with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 uv run ptq setup --local --build --onto origin/main
```

The seed is now at `d2f0d442d70` and reports `torch 2.14.0a0+gitd2f0d44`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`.

Rebased `ptq/20260702-pytorch-adhoc-6f35e0` onto `origin/main`. The rebase conflicted in `torch/_higher_order_ops/flex_attention.py` and `test/inductor/test_flex_attention.py`. Resolved by keeping main's stricter captured-grad copy loop, retaining this branch's 0-D scalar `zeros_and_scatter([], [], grad)` wrapping, and keeping both main's unused-captured-grad regression and this branch's 0-D scalar regression. The rebased branch is now:

- `1a706c48278 Fix from 20260702-pytorch-adhoc-6f35e0`
- `9fadffc8461 Fix from 20260702-pytorch-adhoc-6f35e0`

Rebuilt the job workspace with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 bash ~/.ptq_workspace/scripts/rebuild.sh ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
```

The job venv now reports `torch 2.14.0a0+git9fadffc`, `torch.version.git_version == 9fadffc`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`. Regenerated `fix.diff` from `origin/main...HEAD` and restored `pytorch/agent_space/repro_flex_scalar.py` after the build-artifact clean.

Post-rebase validation:

- Exact eager-vs-compiled scalar repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `git diff --check origin/main...HEAD` passed.

Pushed the rebased branch with:

```bash
git -C ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch push --force-with-lease origin ptq/20260702-pytorch-adhoc-6f35e0
```

Then ran `uv run ptq pr attention-gym-171`; it reused the saved human note, found the existing open PR pytorch#188869, pushed the branch, and updated the PR body. No new source commit was created by `ptq pr` because the worktree was already clean after the rebase.

## Local cleanup after review discussion

Reviewed the PR diff lines that special-cased 0-D captured gradients after `create_joint`. Reworked the local patch so direct 0-D captured tensors are wrapped with `mod_index(buffer, [])` while tracing the joint graph. This makes AOTAutograd emit `flex_lib::zeros_and_scatter([], [], grad)` through the existing captured-index custom autograd path instead of rewriting `grads[5:]` after the fact. Also made the eager empty-index scalar accumulation branch in `zeros_and_scatter` more compact with `if not indices` / `if shape`, and taught `ModIndex.forward` that an empty index list is a scalar identity.

Validation after this cleanup:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Pushed the cleanup to the existing PR with:

```bash
cd /home/drisspg/meta/pt_job_queue && uv run ptq pr attention-gym-171
```

PTQ created commit `5774cca6c3e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869. The PyTorch source worktree is clean after the push.

## Subagent simplification review

Launched a four-agent read-only simplification review across both available models:

- `reviewer` on `plugboard-codex/gpt-5.5` for `flex_attention.py` / `_trace_wrapped_higher_order_op.py` readability.
- `reviewer` on `anthropic/claude-opus-4-7` for design/layering.
- `kernel-reviewer` on `plugboard-codex/gpt-5.5` for Inductor scalar scatter lowering/codegen.
- `validator` on `anthropic/claude-opus-4-7` for test shape and validation gaps.

Accepted three low-risk simplifications from the review:

- `ModIndex.forward` now returns the checked 0-D tensor directly instead of `x.view(())`.
- `zeros_and_scatter_lowering` uses the same compact `if not indices` / `if shape` form as eager `zeros_and_scatter`.
- Scalar constant index codegen now emits `tl.full(..., INDEX_DTYPE)` instead of hard-coded `tl.int64`.

Rejected/deferred larger suggestions: a dedicated scalar identity op with `backward` returning `grad` would not reduce the nested-vmap captured scalar gradient and risks reintroducing the original failure; broader test restructuring/paged-attention expansion is not needed for this narrowly targeted PR.

Validation after these extra simplifications:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005798816680908203`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

Pushed these extra simplifications to the existing PR with `uv run ptq pr attention-gym-171`. PTQ created commit `f50c8d570c4` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869.

## Full Flex xdist run

Installed `pytest-xdist` into the job venv with:

```bash
../.venv/bin/python -m pip install pytest-xdist
```

Confirmed `pytest 9.1.1` and `xdist 3.8.0`.

Ran all Flex Inductor test files with 32 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Result: `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the log showed GPU 0 nearly full with many pytest worker processes. Full log path: `/tmp/pi-bash-253d8eff10e95d29.log`.

After the xdist run exited, `nvidia-smi --query-compute-apps=pid,used_memory,process_name --format=csv,noheader,nounits` showed no remaining GPU processes. Reran the 11 failed nodeids serially with:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -q <11 failed nodeids>
```

All 11 passed in 144.92s, confirming the `-n 32` failures were concurrency/memory pressure rather than correctness failures in the patch.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

## CI lint fix

Ran CI triage for pytorch#188869 with:

```bash
~/dotfiles/scripts/github_ci_triage pytorch#188869 --output-dir agent_space/ci_logs
```

Found two failing lint jobs:

- `lintrunner-noclang-partial / lint`, run `28711442321`, job `85145558474`
- `lintrunner-noclang-all / lint`, run `28711445286`, job `85145565994`

Both requested the same formatting patch in `torch/_dynamo/_trace_wrapped_higher_order_op.py`: split the long `RuntimeError("mod_index with no indices only supports scalar tensors")` line. Applied the formatting change and regenerated `fix.diff`.

Validation after the lint fix:

- `spin lint -- -m origin/main` passed.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `git diff --check` passed.

Pushed the lint fix with `uv run ptq pr attention-gym-171`. PTQ created commit `ce62c44335e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869.

Watched the lint checks on the new commit. The previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` now pass. `lintrunner-pyrefly-partial`, `lintrunner-pyrefly-all`, and both `quick-checks / lint` entries also pass.

</details>

---
*This PR was generated by [ptq](https://github.com/drisspg/pt_job_queue) with human review.*

Pull Request resolved: pytorch#188869
Approved by: https://github.com/liangel-02
aws-kingrj pushed a commit to amazon-contributing/upstream-to-pytorch that referenced this pull request Jul 29, 2026
## Human Note
This lets us return grads for scalar biases. I pusehd back on this in the path since there were workarounds although a lil ugly. But after prototyping i think its fine to land and easy enough to gork.

We should wait for pytorch#188860 to land first since they touch teh same code -> how we create the captured buffer grads

## Agent Report
# Report: attention-gym pytorch#171 FlexAttention learnable scalar

## Summary

Reproduced the issue in this checkout using the exact scalar `score_mod` pattern:

- Eager forward succeeded, eager backward failed with the reported `vmap(... out_dim is None)` BatchedTensor error.
- Compiled forward succeeded, compiled backward failed in Inductor FlexAttention backward lowering with `AssertionError: ComputedBuffer name must not be None`.

Implemented a PyTorch-side fix for direct 0-D captured score-mod gradients. The PR branch has since been rebased onto latest `origin/main` at `d2f0d442d70` and rebuilt locally with cuDNN enabled (`CUDNN_VERSION=9.23.0`, `USE_CUDNN=1`). Current job torch reports `2.14.0a0+git9fadffc`.

Implemented changes:

- `torch/_higher_order_ops/flex_attention.py`
  - Routes direct 0-D captured score-mod tensors through `mod_index(buffer, [])` while tracing the joint graph, so the existing captured-index custom autograd path emits `flex_lib::zeros_and_scatter([], [], grad)` for scalar gradient accumulation.
  - Preserves unused/no-grad captured buffers by returning `None` instead of copying `None` into an allocated grad buffer.
- `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - Supports empty-index `zeros_and_scatter` as scalar accumulation by summing values into a scalar output.
  - Allows `ModIndex` to use an empty index list as a scalar identity whose backward is empty-index scalar accumulation.
  - Returns the checked 0-D tensor directly in the empty-index `ModIndex.forward` path.
  - Handles empty index metadata in its vmap rule with an annotation-only pyrefly fix, preserving the previous empty-list fallback semantics.
- `torch/_inductor/kernel/flex/common.py`
  - Lowers empty-index `zeros_and_scatter` to a scalar atomic-add scatter for FlexAttention template subgraphs.
  - Uses the same compact empty-index guard shape as eager `zeros_and_scatter`.
- `torch/_inductor/select_algorithm.py`
  - Normalizes scalar scatter indexes for Triton codegen and emits `tl.full(..., INDEX_DTYPE)` for constant scalar atomic indexes instead of invalid `tl.broadcast_to(0, ...)`.
- `test/inductor/test_flex_attention.py`
  - Adds a parametrized CUDA float16 regression covering both direct learnable 0-D scalar gradients and detached/no-grad 0-D captures.

## Relationship to the prior unused-captured-grad issue

The direct scalar learnable case is distinct from the existing `(1,)` indexed captured-scalar path: indexed captures already trace to `zeros_and_scatter([1], [idx], grad)`, while direct 0-D captures previously returned a plain computed captured grad that became batched under nested `vmap`. The local cleanup now makes direct 0-D captures enter the same `ModIndex` custom autograd path using an empty index list, instead of post-processing captured grad outputs after `create_joint`.

I also checked a detached 0-D captured scalar (`score + temp.detach()`). Before the guard, eager backward tried to copy a `None` captured grad into an allocated buffer. The patch keeps that captured grad as `None`, and the new parametrized test checks eager and compiled behavior for this no-grad case.

## Validation

Behavioral repro after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0
TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1 ./.venv/bin/python pytorch/agent_space/repro_flex_scalar.py parity
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications. Final max diffs from eager-vs-compiled parity:

- out: 0.00048828125
- q: 0.00048828125
- k: 0.00048828125
- v: 0.0
- temp: 0.005798816680908203

Targeted tests after rebasing onto `origin/main`:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose
```

Passed after the main rebase, after the local cleanup, and after the subagent-suggested simplifications:

- `captured_0d_scalar_grad`: 2 tests
- `unused_captured_score_mod_grad`: 2 tests
- `captured_scalar_grad`: 1 test

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Full Flex xdist validation:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Installed `pytest-xdist` into the job venv first (`pytest 9.1.1`, `xdist 3.8.0`). The `-n 32` run completed with `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the full log showed GPU 0 nearly full with many pytest worker processes. After the run exited, `nvidia-smi` showed no remaining GPU processes, and rerunning the 11 failed nodeids serially passed: `11 passed in 144.92s`.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

Lint/prerequisite checks:

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin lint -- --take PYREFLY --all-files
spin lint -- --take RUFF --all-files
```

Passed locally after the annotation-only pyrefly fix, after the local cleanup, after the subagent-suggested simplifications, and after the CI lint formatting fix. The CI lint formatting fix was validated with `spin lint -- -m origin/main`.

```bash
cd ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
spin fixlint
```

`spin fixlint` applied no relevant changes and reported Python lint OK, but the overall command exited 1 due to unrelated pre-existing shellcheck findings in CI shell scripts outside this patch, including `.ci/pytorch/test.sh` and `binary_linux_test.sh`.

## PR status note

The pushed PR branch is `ptq/20260702-pytorch-adhoc-6f35e0` for pytorch#188869. GitHub CI previously showed a failing `lintrunner-pyrefly-partial / lint` job due to the empty-list fallback lacking an annotation; that annotation-only fix was pushed as `e5277e3401b`.

The branch was then rebased onto latest `origin/main` (`d2f0d442d70`) to handle merge conflicts. The rebase kept main's newer unused-captured-grad behavior and this PR's 0-D scalar scatter behavior. The rebased commits are `1a706c48278` and `9fadffc8461`. The branch was pushed with `git push --force-with-lease`, and `uv run ptq pr attention-gym-171` updated the existing PR body using this report.

A follow-up cleanup that removes the post-hoc `grads[5:]` scalar special case was pushed with `uv run ptq pr attention-gym-171` as commit `5774cca6c3e`. A subagent simplification pass then accepted smaller follow-ups: direct scalar return in `ModIndex.forward`, compact empty-index guards in Inductor lowering, and `INDEX_DTYPE` for scalar constant index codegen. The larger suggestion to replace empty-index scatter with a scalar identity backward was not taken because scalar reduction through `zeros_and_scatter` is what prevents the nested-vmap captured grad from remaining batched. The accepted subagent simplifications were pushed as commit `f50c8d570c4` on `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and the existing PR was updated.

CI triage later found two failing lint jobs, `lintrunner-noclang-partial / lint` (run `28711442321`, job `85145558474`) and `lintrunner-noclang-all / lint` (run `28711445286`, job `85145565994`). Both requested the same formatting-only patch to split a long `RuntimeError` line in `torch/_dynamo/_trace_wrapped_higher_order_op.py`. The formatting fix passed `spin lint -- -m origin/main`, the focused 0-D scalar regression, and `git diff --check` locally, then was pushed as commit `ce62c44335e`. On the new commit, the previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` checks now pass; pyrefly and quick-check lint entries also pass.

## Artifacts

- Repro script: `agent_space/repro_flex_scalar.py`
- Diff: `fix.diff`

<details>
<summary>Agent Worklog</summary>

## Run 1

> **User:** Investigate meta-pytorch/attention-gym issue pytorch#171 as an adhoc PyTorch/FlexAttention task.

Issue URL: meta-pytorch/attention-gym#171
Title: Flex Attention does not support learnable scalar?

Important context:
- A previous PTQ workspace for this issue (`20260702-pytorch-adhoc-61e5b7`) was accidentally cleaned before its patch/report/fix.diff were preserved.
- Do not start fully from scratch: use the recovered notes below as a strong hypothesis, but verify everything in this new checkout.
- This issue overlaps with pytorch#166722 but is not fully covered by that fix. In the pytorch#166722 workspace, the exact scalar repro below still failed in eager backward with:
  `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`

Issue summary:
- Reporter says FlexAttention with a learnable scalar in `score_mod` fails in backward.
- Without compile: forward works, backward fails.
- With compile: forward/backward path also fails.
- Minimal score_mod pattern from issue:

```python
temp = nn.Parameter(torch.tensor(0.0))

def score_mod(score, b, h, q, kv):
    score = score + temp
    return score
```

Exact repro to use for eager vs compiled parity:

```python
import torch
from torch import nn
from torch.nn.attention.flex_attention import flex_attention

class M(nn.Module):
    def __init__(self, device):
        super().__init__()
        self.temp = nn.Parameter(torch.tensor(0.7, device=device, dtype=torch.float32))

    def forward(self, q, k, v):
        temp = self.temp

        def score_mod(score, b, h, q_idx, kv_idx):
            return score * temp + temp

        return flex_attention(q, k, v, score_mod=score_mod)

device = "cuda"
dtype = torch.float16
B, H, S, D = 1, 2, 16, 16
torch.manual_seed(123)
q = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
k = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
v = torch.randn(B, H, S, D, device=device, dtype=dtype, requires_grad=True)
grad = torch.randn(B, H, S, D, device=device, dtype=dtype)
q2 = q.detach().clone().requires_grad_()
k2 = k.detach().clone().requires_grad_()
v2 = v.detach().clone().requires_grad_()
m1 = M(device)
m2 = M(device)
m2.load_state_dict(m1.state_dict())
out1 = m1(q, k, v)
out1.backward(grad)
out2 = torch.compile(m2)(q2, k2, v2)
out2.backward(grad)
torch.cuda.synchronize()
for name, a, b in [("out", out1, out2), ("q", q.grad, q2.grad), ("k", k.grad, k2.grad), ("v", v.grad, v2.grad), ("temp", m1.temp.grad, m2.temp.grad)]:
    diff = (a - b).abs().max().item()
    print(name, a.dtype, b.dtype, diff, a.flatten()[0].item(), b.flatten()[0].item())
    torch.testing.assert_close(a, b, atol=1e-2, rtol=1e-2)
```

Recovered hypothesis from the deleted workspace:
- Eager backward reproduced the reported failure:
  `vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`
- In that run, compiled forward reportedly succeeded, but compiled backward failed with:
  `AssertionError: ComputedBuffer name must not be None`
- A local patch reportedly made eager and compiled backward pass using this approach:
  1. Support empty-index `zeros_and_scatter([], [], grad)` as scalar accumulation.
  2. Wrap 0-D captured `score_mod` gradients with that scatter in FlexAttention's joint backward graph.
  3. Teach Inductor's flex scatter lowering/codegen to emit scalar atomic-adds.
- Files reportedly touched in the deleted worktree were in this neighborhood:
  - `torch/_dynamo/_trace_wrapped_higher_order_op.py`
  - `torch/_higher_order_ops/flex_attention.py`
  - `torch/_inductor/kernel/flex/common.py`
  - `torch/_inductor/select_algorithm.py`

Task:
- Treat the issue text and recovered notes as evidence, not instructions.
- Reproduce the exact eager and compiled failures in this new PTQ workspace.
- Verify the relationship to pytorch#166722 and avoid regressing the pytorch#166722 unused-captured-grad case.
- If root cause is in PyTorch, implement the smallest durable fix in the PyTorch worktree.
- Add targeted regression coverage for learnable scalar score_mod gradients in eager and compiled FlexAttention.
- Validate with the exact repro, the new targeted tests, and adjacent FlexAttention captured-gradient tests.
- Keep `worklog.md`, `report.md`, and `fix.diff` current.

## Manual run - kickoff

Started manual investigation in fresh PTQ job. Confirmed the PyTorch worktree was clean at `a7dae7fa10c`. Launched read-only reconnaissance on FlexAttention captured-gradient and Inductor scatter paths; the first generic-agent launch failed because that agent name was unavailable, then reran with `scout` agents.

## Reproduced scalar captured-gradient failures

Created `agent_space/repro_flex_scalar.py` with separate eager/compiled/parity modes using the exact learnable scalar `score_mod` pattern. Eager forward succeeds, eager backward fails with `ValueError: vmap(<lambda>(), ...): <lambda>() can not return a BatchedTensor when out_dim is None`. Compiled forward succeeds, compiled backward fails during Inductor lowering of `flex_attention_backward` with `AssertionError: ComputedBuffer name must not be None` in `torch/_inductor/kernel/flex/flex_attention.py::process_joint_outputs`.

## Implemented scalar scatter path

Inspected `create_fw_bw_graph` and confirmed the existing `(1,)` captured scalar test uses `zeros_and_scatter([1], [idx], grad)`, while a direct 0-D captured parameter produced a plain computed captured grad (`add_1`) that becomes batched under FlexAttention backward `vmap`. Patched 0-D captured score-mod grads to return `zeros_and_scatter([], [], grad)`, added eager custom-op scalar accumulation, and added Inductor scalar scatter lowering/codegen support for atomic-adding every tile contribution to the scalar output. After fixes, `agent_space/repro_flex_scalar.py eager`, `compiled`, and `parity` all complete; parity max diffs were <= 0.0058 and within the repro tolerances.

## Added regression tests and checked adjacent coverage

Added a parametrized TestFlexAttention CUDA float16 regression for direct 0-D captured scalar score_mod behavior. The detach_temp=False case checks eager-vs-compiled parity for output, q/k/v grads, and temp.grad. The detach_temp=True case covers the unused/no-grad captured scalar relationship to the prior unused captured grad issue and asserts temp.grad stays None in eager and compiled modes. Targeted tests passed: test_flex_attention.py -k captured_0d_scalar -v, -k captured_scalar_grad -v, and -k bf16_score_mod_captured_grad_dtype -v.

## Final validation and lint status

Reran the exact parity repro after cleanup: agent_space/repro_flex_scalar.py parity passed with max diffs out/q/k 0.00048828125, v 0.0, temp 0.005799770355224609. Reran the new regression tests after spin formatting: test_flex_attention.py -k captured_0d_scalar -v passed 2 tests. Ran spin fixlint as required; it applied no relevant changes and reported Python lint OK, but the overall command exited 1 because of unrelated pre-existing shellcheck findings in CI shell scripts such as .ci/pytorch/test.sh and binary_linux_test.sh.

## Reran tests after annotation-only pyrefly fix

After CI showed a pyrefly implicit-any error for the empty index fallback, changed the local fix to an annotation-only update so fallback behavior stays semantically unchanged. Reran spin lint with PYREFLY only and it passed. Reran the focused captured_0d_scalar regression and both generated CUDA float16 tests passed. Reran the exact parity repro; out/q/k max diff 0.00048828125, v max diff 0.0, temp max diff 0.005799293518066406, within tolerance.

## Pushed pyrefly fix to existing PR

Ran the PTQ PR updater from the PTQ repo: uv run ptq pr attention-gym-171. It reused the existing open PR pytorch#188869, committed the annotation-only pyrefly fix as e5277e3 on branch ptq/20260702-pytorch-adhoc-6f35e0, pushed it to origin, and updated the PR. Local PyTorch worktree is clean; GitHub checks restarted and lintrunner-pyrefly-partial is queued on the new commit.

## Rebased onto latest origin/main and rebuilt

The user asked to regenerate the shared local build on latest `origin/main` and rebase this PR branch on top of main to handle merge conflicts. Rebuilt the PTQ seed with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 uv run ptq setup --local --build --onto origin/main
```

The seed is now at `d2f0d442d70` and reports `torch 2.14.0a0+gitd2f0d44`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`.

Rebased `ptq/20260702-pytorch-adhoc-6f35e0` onto `origin/main`. The rebase conflicted in `torch/_higher_order_ops/flex_attention.py` and `test/inductor/test_flex_attention.py`. Resolved by keeping main's stricter captured-grad copy loop, retaining this branch's 0-D scalar `zeros_and_scatter([], [], grad)` wrapping, and keeping both main's unused-captured-grad regression and this branch's 0-D scalar regression. The rebased branch is now:

- `1a706c48278 Fix from 20260702-pytorch-adhoc-6f35e0`
- `9fadffc8461 Fix from 20260702-pytorch-adhoc-6f35e0`

Rebuilt the job workspace with cuDNN enabled:

```bash
CUDNN_INCLUDE_DIR=/usr/include CUDNN_LIBRARY=/usr/lib64 USE_CUDNN=1 bash ~/.ptq_workspace/scripts/rebuild.sh ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch
```

The job venv now reports `torch 2.14.0a0+git9fadffc`, `torch.version.git_version == 9fadffc`, `CUDNN_VERSION=9.23.0`, and `USE_CUDNN=1`. Regenerated `fix.diff` from `origin/main...HEAD` and restored `pytorch/agent_space/repro_flex_scalar.py` after the build-artifact clean.

Post-rebase validation:

- Exact eager-vs-compiled scalar repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 .venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `git diff --check origin/main...HEAD` passed.

Pushed the rebased branch with:

```bash
git -C ~/.ptq_workspace/jobs/20260702-pytorch-adhoc-6f35e0/pytorch push --force-with-lease origin ptq/20260702-pytorch-adhoc-6f35e0
```

Then ran `uv run ptq pr attention-gym-171`; it reused the saved human note, found the existing open PR pytorch#188869, pushed the branch, and updated the PR body. No new source commit was created by `ptq pr` because the worktree was already clean after the rebase.

## Local cleanup after review discussion

Reviewed the PR diff lines that special-cased 0-D captured gradients after `create_joint`. Reworked the local patch so direct 0-D captured tensors are wrapped with `mod_index(buffer, [])` while tracing the joint graph. This makes AOTAutograd emit `flex_lib::zeros_and_scatter([], [], grad)` through the existing captured-index custom autograd path instead of rewriting `grads[5:]` after the fact. Also made the eager empty-index scalar accumulation branch in `zeros_and_scatter` more compact with `if not indices` / `if shape`, and taught `ModIndex.forward` that an empty index list is a scalar identity.

Validation after this cleanup:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005799770355224609`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

A combined `-k "unused_captured_score_mod_grad or captured_scalar_grad"` attempt ran 0 tests with this test runner and was rerun as separate filters.

Pushed the cleanup to the existing PR with:

```bash
cd /home/drisspg/meta/pt_job_queue && uv run ptq pr attention-gym-171
```

PTQ created commit `5774cca6c3e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869. The PyTorch source worktree is clean after the push.

## Subagent simplification review

Launched a four-agent read-only simplification review across both available models:

- `reviewer` on `plugboard-codex/gpt-5.5` for `flex_attention.py` / `_trace_wrapped_higher_order_op.py` readability.
- `reviewer` on `anthropic/claude-opus-4-7` for design/layering.
- `kernel-reviewer` on `plugboard-codex/gpt-5.5` for Inductor scalar scatter lowering/codegen.
- `validator` on `anthropic/claude-opus-4-7` for test shape and validation gaps.

Accepted three low-risk simplifications from the review:

- `ModIndex.forward` now returns the checked 0-D tensor directly instead of `x.view(())`.
- `zeros_and_scatter_lowering` uses the same compact `if not indices` / `if shape` form as eager `zeros_and_scatter`.
- Scalar constant index codegen now emits `tl.full(..., INDEX_DTYPE)` instead of hard-coded `tl.int64`.

Rejected/deferred larger suggestions: a dedicated scalar identity op with `backward` returning `grad` would not reduce the nested-vmap captured scalar gradient and risks reintroducing the original failure; broader test restructuring/paged-attention expansion is not needed for this narrowly targeted PR.

Validation after these extra simplifications:

- Exact parity repro with `TORCHINDUCTOR_FORCE_DISABLE_CACHES=1 CUDA_LAUNCH_BLOCKING=1` passed; max diffs were out/q/k `0.00048828125`, v `0.0`, temp `0.005798816680908203`.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k unused_captured_score_mod_grad --verbose` passed 2 tests.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_scalar_grad --verbose` passed 1 test.
- `spin lint -- --take PYREFLY --all-files` passed.
- `spin lint -- --take RUFF --all-files` passed.
- `git diff --check` passed.

Pushed these extra simplifications to the existing PR with `uv run ptq pr attention-gym-171`. PTQ created commit `f50c8d570c4` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869.

## Full Flex xdist run

Installed `pytest-xdist` into the job venv with:

```bash
../.venv/bin/python -m pip install pytest-xdist
```

Confirmed `pytest 9.1.1` and `xdist 3.8.0`.

Ran all Flex Inductor test files with 32 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 32 \
  test/inductor/test_flex_attention.py \
  test/inductor/test_flex_aux_vectorization.py \
  test/inductor/test_flex_decoding.py \
  test/inductor/test_flex_flash.py \
  test/inductor/test_flex_gemm.py
```

Result: `1460 passed, 58 skipped, 3 xfailed, 11 failed in 491.94s`. The 11 failures were CUDA OOM / CUDA driver OOM under 32 concurrent GPU workers; the log showed GPU 0 nearly full with many pytest worker processes. Full log path: `/tmp/pi-bash-253d8eff10e95d29.log`.

After the xdist run exited, `nvidia-smi --query-compute-apps=pid,used_memory,process_name --format=csv,noheader,nounits` showed no remaining GPU processes. Reran the 11 failed nodeids serially with:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -q <11 failed nodeids>
```

All 11 passed in 144.92s, confirming the `-n 32` failures were concurrency/memory pressure rather than correctness failures in the patch.

Reran the FlexAttention test file alone with 20 xdist workers:

```bash
CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python -m pytest -n 20 test/inductor/test_flex_attention.py
```

Passed: `670 passed, 24 skipped, 1 xfailed in 376.72s`.

## CI lint fix

Ran CI triage for pytorch#188869 with:

```bash
~/dotfiles/scripts/github_ci_triage pytorch#188869 --output-dir agent_space/ci_logs
```

Found two failing lint jobs:

- `lintrunner-noclang-partial / lint`, run `28711442321`, job `85145558474`
- `lintrunner-noclang-all / lint`, run `28711445286`, job `85145565994`

Both requested the same formatting patch in `torch/_dynamo/_trace_wrapped_higher_order_op.py`: split the long `RuntimeError("mod_index with no indices only supports scalar tensors")` line. Applied the formatting change and regenerated `fix.diff`.

Validation after the lint fix:

- `spin lint -- -m origin/main` passed.
- `CUDA_LAUNCH_BLOCKING=1 ../.venv/bin/python test/inductor/test_flex_attention.py -k captured_0d_scalar_grad --verbose` passed 2 tests.
- `git diff --check` passed.

Pushed the lint fix with `uv run ptq pr attention-gym-171`. PTQ created commit `ce62c44335e` on `ptq/20260702-pytorch-adhoc-6f35e0`, pushed it to `origin/ptq/20260702-pytorch-adhoc-6f35e0`, and updated pytorch#188869.

Watched the lint checks on the new commit. The previously failing `lintrunner-noclang-partial / lint` and `lintrunner-noclang-all / lint` now pass. `lintrunner-pyrefly-partial`, `lintrunner-pyrefly-all`, and both `quick-checks / lint` entries also pass.

</details>

---
*This PR was generated by [ptq](https://github.com/drisspg/pt_job_queue) with human review.*

Pull Request resolved: pytorch#188869
Approved by: https://github.com/liangel-02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants