Skip to content

Allocate arena_matrix copies once instead of twice - #3431

Open
jachymb wants to merge 1 commit into
stan-dev:developfrom
jachymb:bugfix/arena-matrix-double-allocation
Open

jachymb wants to merge 1 commit into
stan-dev:developfrom
jachymb:bugfix/arena-matrix-double-allocation

Conversation

@jachymb

@jachymb jachymb commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

A one-liner that as has been eating up significant amount of autodiff memory since 2020

Summary

arena_matrix's constructor from an Eigen expression or matrix allocates its buffer twice:

template <typename T, require_eigen_t<T>* = nullptr>
arena_matrix(const T& other)
    : Base::Map(ChainableStack::instance_->memalloc_.alloc_array<Scalar>(
                    other.size()),
                get_rows(other), get_cols(other)) {
  *this = other;
}

Inside the constructor, *this = other resolves to arena_matrix's own templated operator=(const T&), an exact match. That operator places the map onto a second arena allocation of the same size and copies into it, so the first buffer is never used. The pattern dates back to f94d6f1 (#1970).

Every copy onto the arena that goes through this constructor therefore took twice the memory: 16N bytes instead of 8N for double or var. That includes:

  • to_arena of any Eigen argument that is not already an arena_matrix, rvalues included;
  • arena_t<T> x = expr; for any expression, and for any matrix except a plain rvalue of exactly the matrix type, which is moved instead;
  • the operand copies and the zero-filled partials in the distributions' partials edges;
  • var_value<Matrix> constructed from a value that is not an arena_matrix. Value and adjoint together drop from 3N to 2N doubles, and from 4N to 2N when both are given.

Copies of the same arena_matrix type and of exactly Map<MatrixType> were not affected.

The fix assigns into the buffer the constructor has already allocated:

  Base::operator=(other);

This is the same call operator= makes after allocating. It uses the same get_rows/get_cols, so orientation, double → var conversion and empty inputs behave as before. Only the target buffer differs.

Tests

Added a test which fails on current develop. Mostly meant to document the issue.

Side Effects

Pure gain.

Some experiments showing that this is quite significant:

The arena bytes per call are measured between two sentinel allocations. At N = 1024 a gradient takes:

Gradient develop this PR change
normal_lpdf(y, mu, sigma) with a var vector mu 65,712 B 49,328 B −25 %
sum(elt_multiply(a, b)) 139,392 B 106,624 B −24 %
sum(multiply(X, beta)), X is N × 5 data 139,600 B 82,216 B −41 %
dot_product(a, b) 81,992 B 65,608 B −20 %

In normal_lpdf, half of the saving comes from the copy of mu and half from its zero-filled partials.

Speed change within noise.

Release notes

Reduced autodiff memory use.

Checklist

  • Copyright holder: Jáchym Barvínek [email protected]

    The copyright holder is typically you or your assignee, such as a university or company. By submitting this pull request, the copyright holder is agreeing to the license the submitted work under the following licenses:
    - Code: BSD 3-clause (https://opensource.org/licenses/BSD-3-Clause)
    - Documentation: CC-BY 4.0 (https://creativecommons.org/licenses/by/4.0/)

  • the basic tests are passing

    • unit tests pass (to run, use: ./runTests.py test/unit)
    • header checks pass, (make test-headers)
    • dependencies checks pass, (make test-math-dependencies)
    • docs build, (make doxygen)
    • code passes the built in C++ standards checks (make cpplint)
  • the code is written in idiomatic C++ and changes are documented in the doxygen

  • the new changes are tested

The copying constructor allocated its buffer and then called
`*this = other`, which resolves to arena_matrix's own templated
operator= and allocates a second buffer. The first one was never used,
so every arena copy took twice the memory. Assign into the buffer
with Base::operator= instead.
@stan-buildbot

Copy link
Copy Markdown
Contributor
Name Old Result New Result Ratio Performance change( 1 - new / old )
stat_comp_benchmarks/benchmarks/gp_regr/gp_regr.stan 0.23 0.23 0.99 -0.59% slower
stat_comp_benchmarks/benchmarks/gp_regr/gen_gp_data.stan 0.06 0.06 0.99 -0.77% slower
stat_comp_benchmarks/benchmarks/low_dim_corr_gauss/low_dim_corr_gauss.stan 0.02 0.02 0.94 -5.89% slower
stat_comp_benchmarks/benchmarks/arma/arma.stan 0.71 0.65 1.08 7.5% faster
stat_comp_benchmarks/benchmarks/arK/arK.stan 3.2 3.17 1.01 0.89% faster
stat_comp_benchmarks/benchmarks/pkpd/one_comp_mm_elim_abs.stan 42.86 42.53 1.01 0.78% faster
stat_comp_benchmarks/benchmarks/pkpd/sim_one_comp_mm_elim_abs.stan 0.59 0.6 1.0 -0.38% slower
stat_comp_benchmarks/benchmarks/gp_pois_regr/gp_pois_regr.stan 4.54 4.45 1.02 1.85% faster
stat_comp_benchmarks/benchmarks/low_dim_gauss_mix_collapse/low_dim_gauss_mix_collapse.stan 20.98 20.9 1.0 0.38% faster
stat_comp_benchmarks/benchmarks/garch/garch.stan 0.89 0.89 0.99 -0.74% slower
stat_comp_benchmarks/benchmarks/eight_schools/eight_schools.stan 0.11 0.11 1.02 2.32% faster
stat_comp_benchmarks/benchmarks/irt_2pl/irt_2pl.stan 8.46 8.33 1.02 1.61% faster
stat_comp_benchmarks/benchmarks/low_dim_gauss_mix/low_dim_gauss_mix.stan 6.32 6.32 1.0 -0.13% slower
stat_comp_benchmarks/benchmarks/sir/sir.stan 165.72 166.26 1.0 -0.33% slower
performance.compilation 379.28 382.64 0.99 -0.89% slower
Mean result: 1.004479020037531

Jenkins Console Log
Jenkins Build Stages
Commit hash: 993f35f70b6d7c411c87d88bbb003259bfff0fcd

Machine information
Distributor ID:	Ubuntu
Description:	Ubuntu 20.04.3 LTS
Release:	20.04
Codename:	focal

CPU:

Architecture:                            x86_64
CPU op-mode(s):                          32-bit, 64-bit
Byte Order:                              Little Endian
Address sizes:                           43 bits physical, 48 bits virtual
CPU(s):                                  256
On-line CPU(s) list:                     0-255
Thread(s) per core:                      2
Core(s) per socket:                      64
Socket(s):                               2
NUMA node(s):                            2
Vendor ID:                               AuthenticAMD
CPU family:                              23
Model:                                   49
Model name:                              AMD EPYC 7742 64-Core Processor
Stepping:                                0
Frequency boost:                         enabled
CPU MHz:                                 1386.314
CPU max MHz:                             3416.0681
CPU min MHz:                             1500.0000
BogoMIPS:                                4491.56
Virtualization:                          AMD-V
L1d cache:                               4 MiB
L1i cache:                               4 MiB
L2 cache:                                64 MiB
L3 cache:                                512 MiB
NUMA node0 CPU(s):                       0-63,128-191
NUMA node1 CPU(s):                       64-127,192-255
Vulnerability Gather data sampling:      Not affected
Vulnerability Indirect target selection: Not affected
Vulnerability Itlb multihit:             Not affected
Vulnerability L1tf:                      Not affected
Vulnerability Mds:                       Not affected
Vulnerability Meltdown:                  Not affected
Vulnerability Mmio stale data:           Not affected
Vulnerability Old microcode:             Not affected
Vulnerability Reg file data sampling:    Not affected
Vulnerability Retbleed:                  Mitigation; untrained return thunk; SMT enabled with STIBP protection
Vulnerability Spec rstack overflow:      Mitigation; Safe RET
Vulnerability Spec store bypass:         Mitigation; Speculative Store Bypass disabled via prctl
Vulnerability Spectre v1:                Mitigation; usercopy/swapgs barriers and __user pointer sanitization
Vulnerability Spectre v2:                Mitigation; Retpolines; IBPB conditional; STIBP always-on; RSB filling; PBRSB-eIBRS Not affected; BHI Not affected
Vulnerability Srbds:                     Not affected
Vulnerability Tsa:                       Not affected
Vulnerability Tsx async abort:           Not affected
Vulnerability Vmscape:                   Mitigation; IBPB before exit to userspace
Flags:                                   fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ht syscall nx mmxext fxsr_opt pdpe1gb rdtscp lm constant_tsc rep_good nopl xtopology nonstop_tsc cpuid extd_apicid aperfmperf rapl pni pclmulqdq monitor ssse3 fma cx16 sse4_1 sse4_2 x2apic movbe popcnt aes xsave avx f16c rdrand lahf_lm cmp_legacy svm extapic cr8_legacy abm sse4a misalignsse 3dnowprefetch osvw ibs skinit wdt tce topoext perfctr_core perfctr_nb bpext perfctr_llc mwaitx cpb cat_l3 cdp_l3 hw_pstate ssbd mba ibrs ibpb stibp vmmcall fsgsbase bmi1 avx2 smep bmi2 cqm rdt_a rdseed adx smap clflushopt clwb sha_ni xsaveopt xsavec xgetbv1 xsaves cqm_llc cqm_occup_llc cqm_mbm_total cqm_mbm_local clzero irperf xsaveerptr rdpru wbnoinvd amd_ppin arat npt lbrv svm_lock nrip_save tsc_scale vmcb_clean flushbyasid decodeassists pausefilter pfthreshold avic v_vmsave_vmload vgif v_spec_ctrl umip rdpid overflow_recov succor smca sev sev_es

G++:

g++ (Ubuntu 9.4.0-1ubuntu1~20.04) 9.4.0
Copyright (C) 2019 Free Software Foundation, Inc.
This is free software; see the source for copying conditions.  There is NO
warranty; not even for MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.

Clang:

clang version 10.0.0-4ubuntu1 
Target: x86_64-pc-linux-gnu
Thread model: posix
InstalledDir: /usr/bin

@WardBrian
WardBrian requested a review from SteveBronder October 6, 2026 15:41
Comment on lines +263 to +265
char* before = static_cast<char*>(memalloc.alloc(8));
arena_matrix<Eigen::VectorXd> a(x);
char* after = static_cast<char*>(memalloc.alloc(8));

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You can just ask for memalloc.bytes_allocated() before and after the allocation

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

bytes_allocated() only counts whole blocks, so it doesn't change for a mere 40-byte allocation inside the current one. i.e. the test would pass on current develop without this patch.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ah sorry I just looked at the function names. Could you add something like this to the stack allocator and use it in the test

inline size_t approx_bytes_used() const {
  size_t sum = 0;
  for (size_t i = 0; i < cur_block_; ++i) {
    sum += sizes_[i];
  }
  return sum + static_cast<size_t>(next_loc_ - blocks_[cur_block_]);
}

@SteveBronder

Copy link
Copy Markdown
Collaborator

Thanks! Besides the note above to simplify the test a bit I think this is good

@jachymb

jachymb commented Oct 7, 2026

Copy link
Copy Markdown
Contributor Author

@SteveBronder I believe your simplification is not applicable here because of the amount of tested memory being too small, see the comment above.

@WardBrian
WardBrian requested a review from SteveBronder October 8, 2026 02:47

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants