Skip to content

Add fixnums array stride knowledge for frozen arrays - #1081

Draft
AlexaCampusano wants to merge 1 commit into
Shopify:masterfrom
AlexaCampusano:nc/add-array-element-stride
Draft

AlexaCampusano wants to merge 1 commit into
Shopify:masterfrom
AlexaCampusano:nc/add-array-element-stride

Conversation

@AlexaCampusano

@AlexaCampusano AlexaCampusano commented Oct 8, 2026 •

Copy link
Copy Markdown

Related read Storage Strategies for Collectionsin Dynamically Typed Languages

What are we trying to do?

This PR is a proposal to add fixnums stride information into the array object, in order for heap arrays to have more information about the data it's storing and we can optimize it.

In this PR I decided to scope it to only frozen heap arrays that contain fixnums. The reason behind it is, this is a good first step, and we can tackle some of the problem in a safe way.

How does it work?

1. Lifecycle of one array

  flowchart LR
      A["Array created<br/>8 B per element"] --> T{"trigger?"}
      T -- "Array#freeze" --> N
      T -- "[...].freeze literal<br/>(compile time)" --> N
      T -- "neither" --> A
      N{"ary_try_narrow<br/>heap-backed?<br/>all Fixnum?<br/>fits 32 bit?"}
      N -- no --> A
      N -- yes --> P["packed<br/>1 / 2 / 4 B per element"]
      P -- "element reads<br/>[] each sum == include?" --> P
      P -- "write · raw buffer ·<br/>dup / slice / C ext" --> W["rb_ary_widen<br/>back to 8 B, for good"]
Loading

2. e.g.: TABLE = [0, 1, …, 199].freeze (200 elements)

  flowchart LR
      subgraph before["before"]
          S1["slot 40 B<br/>stride VALUE"] --> B1["malloc 1,600 B<br/>200 × VALUE"]
      end
      subgraph after["after narrowing"]
          S2["slot 40 B<br/>stride W8 · unsigned"] --> B2["malloc 200 B<br/>200 × uint8_t"]
      end
      before -- "opt_ary_freeze →<br/>rb_ary_narrow" --> after
Loading

3. Who touches a narrow array, and how

Readers, stay narrow. RARRAY_AREF (internal) decodes elements in place, so [], each, map, sum, include?, ==, index, count, join, pack and Marshal.dump never allocate a wide copy and never change the representation. Internally RARRAY_AREF goes through RARRAY_CONST_PTR, which hands out a VALUE * and would therefore widen; we gave it a stride-aware decode path instead.

GC. Mark: nothing to trace (no object references in a packed buffer). Compaction: nothing to re-point, and rb_ary_embeddable_p refuses narrow arrays so they're never re-embedded. Free: ARY_HEAP_SIZE is stride aware

Ractor. make_shareable skips the element loop (every element is an immediate). rb_ary_widen takes the VM lock and re-checks, so two Ractors widening the same shareable array do it exactly once.

Widen, permanently Raw-buffer readers widen. For C extensions that's unavoidable, RARRAY_PTR promises a VALUE * the caller may hold or write through. For Ruby's own dup/slices it's a deliberate trade because it creates shared copies.

Why am I not tackling embedded arrays?

In some of the stats we collected in https://shopify.slack.com/archives/C0BSQ56F7HD/p1789753762082559, majority of array literals containing fixnums in Core/SFR are very tiny in size (< 3 elements hold, at least for SFR ~93% of all fixnums arrays, meaning they are embedded). Embedded arrays are allocated in Ruby heap slots they can fit, while they fit in there they don't malloc, and we cannot change the size of the slots an object is already in outside of compaction, so basically we can't really do much with them by narrowing VALUE to fit the Fixnum with a smaller stride.

I do have an idea where we COULD, and I have not tested this or how much work would it be, but see if we can reduce the amount of Arrays that end up being malloc'ed because the slots are using VALUE to measure capacity. In theory we could potentially fit bigger arrays of smaller stride in one slot, and prevent that array from allocating C heap memory.

Why only frozen?

I chose to do frozen arrays because it's one layer of protection against widening-churn. Some other reasons I had in mind:

  • We narrow an array, and then it gets modified at runtime and it suddenly has another value. It will have to be widen because it's not homogeneous anymore.
  • In ruby/ruby, there are a lot of places where we obtain heap arrays buffer and use it in many different ways, ways we cannot control. There's even a public access to it through RARRAY_CONST_PTR that expects a collection of VALUE, that can be used by native extensions.
  • Picking the right moment to narrow is worth a discussion, do we do it after boot? do we do it every time we write to an array? and does this put a lot of work in a hot path, check the stride of the element -> if matches continue to narrow -> if not widen the entire thing, only trigger it on arrays that become old gen? Each of those brings risks worth discussing

It also keeps the blast radius of changes minimal, and we can iterate on it slowly to teach all the different parts of Ruby how to be stride aware.

Some numbers!!!

Reading narrowed arrays:

image

ruby-bench

  Memory: master -> stride   (base 8fdf4342b5; stride = e73608f5d7 + latest sync)
  Array/total = ObjectSpace.memsize_of_all after full GC; RSS = mean per-iteration after warmup

  == Narrow path exercised ==============================================

  blurhash
    1 array x 124,848 elems (pixel buffer, make_shareable)
    Array       1042.9 KiB  ->      188.7 KiB   -81.9%
    total          3.2 MiB  ->        2.4 MiB   -26.3%
    RSS           20.8 MiB  ->       20.6 MiB   -1.0%

  lee
    2 arrays x ~2,800 elems (grids)
    Array        187.8 KiB  ->      151.4 KiB   -19.4%
    total          3.9 MiB  ->        4.0 MiB   +2.5%
    RSS           43.7 MiB  ->       44.1 MiB   +1.0%

  sudoku
    324 arrays x 9 elems (lookup tables, make_shareable)
    Array        156.1 KiB  ->      137.6 KiB   -11.9%
    total          2.3 MiB  ->        2.2 MiB   -1.3%
    RSS           21.9 MiB  ->       21.5 MiB   -2.2%

  == Not exercised: large Fixnum arrays present but never frozen ========

  lobsters
    57 unfrozen arrays, ~3.3 MiB
    Array       1925.9 KiB  ->     1926.4 KiB     same
    total         30.3 MiB  ->       30.4 MiB   +0.6%
    RSS          375.7 MiB  ->      368.7 MiB   -1.9%

  railsbench
    38 unfrozen arrays, ~2.7 MiB (mail Ragel tables)
    Array        945.5 KiB  ->      947.7 KiB   +0.2%
    total         12.6 MiB  ->       12.7 MiB   +1.0%
    RSS          124.7 MiB  ->      140.6 MiB   +12.8%   (GC heap-sizing variance on master, not the patch)

  mail
    38 unfrozen arrays, ~2.5 MiB (Ragel tables)
    Array        182.4 KiB  ->      184.5 KiB   +1.2%
    total          5.6 MiB  ->        5.8 MiB   +2.6%
    RSS           74.4 MiB  ->       74.5 MiB   +0.2%

  optcarrot
    16 unfrozen arrays, ~1.7 MiB (lookup tables)
    Array      30641.5 KiB  ->    30640.8 KiB     same
    total         31.1 MiB  ->       31.1 MiB     same
    RSS           62.1 MiB  ->       61.7 MiB   -0.6%

  rubocop
    20 unfrozen arrays, ~600 KiB
    Array       1814.4 KiB  ->     1815.7 KiB   +0.1%
    total         12.4 MiB  ->       12.6 MiB   +1.5%
    RSS          109.3 MiB  ->      113.0 MiB   +3.4%

  chunky-png
    1 unfrozen array, ~300 KiB (pixels)
    Array        139.9 KiB  ->      141.9 KiB   +1.4%
    total          3.9 MiB  ->        4.1 MiB   +3.6%
    RSS           98.7 MiB  ->       98.0 MiB   -0.7%

  == Not exercised: no large Fixnum arrays ==============================

  binarytrees
    Array         66.9 KiB  ->       66.2 KiB   -1.1%
    total          2.3 MiB  ->        2.3 MiB   -1.2%
    RSS           26.3 MiB  ->       23.8 MiB   -9.5%   (GC heap-sizing variance on master, not the patch)

  liquid-render
    Array        397.2 KiB  ->      399.2 KiB   +0.5%
    total          6.2 MiB  ->        6.4 MiB   +2.0%
    RSS           59.8 MiB  ->       60.9 MiB   +1.9%

  nbody
    Array         69.1 KiB  ->       68.4 KiB   -1.0%
    total          3.1 MiB  ->        3.1 MiB   -0.2%
    RSS           21.3 MiB  ->       21.3 MiB   +0.1%

@AlexaCampusano
AlexaCampusano force-pushed the nc/add-array-element-stride branch from e73608f to 99fce74 Compare October 8, 2026 21:00

@tenderlove tenderlove left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This looks really good!

Comment thread internal/array.h
* also safe where RARRAY_CONST_PTR is not (opt_aref, GC callbacks).
*
* Not PURE: it is still a pure function of (ary, i) in the narrow case, but
* the wide case reads through RARRAY_CONST_PTR, which is not pure either. */

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not sure if we need comments about purity. But I think this function doesn't have any side-effects (it is pure). It's true that RARRAY_CONST_PTR could widen the array, but it seems like the early return would prevent that from ever happening.

Comment thread array.c
* this point in parallel through RARRAY_CONST_PTR. Widening swaps and
* frees the buffer, which must happen exactly once. Take the VM lock and
* re-check: whoever loses the race finds the array already widened. */
RB_VM_LOCKING() {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nice, yes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants