Skip to content

fix: emit all allocas in the init block instead of the current block - #1665

Merged
TomerStarkware merged 1 commit into
mainfrom
tomer/fix-alloca-in-loops
Oct 1, 2026
Merged

TomerStarkware merged 1 commit into
mainfrom
tomer/fix-alloca-in-loops

Conversation

@TomerStarkware

@TomerStarkware TomerStarkware commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Several places emitted llvm.alloca into the libfunc's own block rather than the function's init (pre-entry) block. When that block sits inside a loop, the alloca re-executes on every iteration and the stack grows without bound.

All allocas now go through the init block, which runs exactly once per function call:

  • felt252_dict_squash: range-check and gas pointers.
  • squashed_dict into_entries: result array slot.
  • qm31 binary ops: lhs/rhs operand slots.
  • RuntimeBindingsMeta::libfunc_qm31_bin_op: result slot. Now takes the LibfuncHelper instead of a bare &Module (same as dict_get); callers already passed the helper via deref.
  • trace_dump::build_state_snapshot (with-trace-dump feature): per-variable snapshot slots emitted at every statement. Takes an explicit init block; both call sites in compiler.rs pass the pre-entry block.

Test plan

  • cargo check with and without with-trace-dump
  • cargo clippy --all-targets clean, cargo fmt applied
  • cargo test --lib -- qm31 felt252_dict squashed_dict (41 passed)

🤖 Generated with Claude Code


This change is Reviewable

Several libfuncs (felt252_dict_squash, squashed_dict into_entries, qm31
binary ops), the qm31 runtime binding and the trace-dump state snapshot
emitted `llvm.alloca` into the libfunc's own block. When that block is
part of a loop, the alloca re-executes on every iteration and the stack
grows unboundedly. Route every alloca through the function's init block,
which runs exactly once per call.

`libfunc_qm31_bin_op` now takes the `LibfuncHelper` (like `dict_get`) so
it can reach the init block, and `build_state_snapshot` takes an explicit
init block argument.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@TomerStarkware
TomerStarkware requested a review from orizi October 1, 2026 10:21

@orizi orizi left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

:lgtm:

@orizi reviewed 6 files and all commit messages, and made 1 comment.
Reviewable status: :shipit: complete! all files reviewed, all discussions resolved (waiting on TomerStarkware).

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Benchmark results Main vs HEAD.

Base

Command Mean [s] Min [s] Max [s] Relative
base dict_insert.cairo (JIT) 1.491 ± 0.012 1.473 1.513 1.02 ± 0.01
base dict_insert.cairo (AOT) 1.466 ± 0.009 1.452 1.478 1.00

Head

Command Mean [s] Min [s] Max [s] Relative
head dict_insert.cairo (JIT) 1.510 ± 0.060 1.457 1.675 1.02 ± 0.04
head dict_insert.cairo (AOT) 1.485 ± 0.028 1.460 1.558 1.00

Base

Command Mean [s] Min [s] Max [s] Relative
base dict_snapshot.cairo (JIT) 1.334 ± 0.012 1.317 1.351 1.02 ± 0.01
base dict_snapshot.cairo (AOT) 1.313 ± 0.013 1.298 1.337 1.00

Head

Command Mean [s] Min [s] Max [s] Relative
head dict_snapshot.cairo (JIT) 1.324 ± 0.010 1.306 1.344 1.02 ± 0.01
head dict_snapshot.cairo (AOT) 1.293 ± 0.006 1.281 1.301 1.00

Base

Command Mean [s] Min [s] Max [s] Relative
base factorial_2M.cairo (JIT) 1.371 ± 0.013 1.355 1.397 1.01 ± 0.02
base factorial_2M.cairo (AOT) 1.355 ± 0.019 1.334 1.398 1.00

Head

Command Mean [s] Min [s] Max [s] Relative
head factorial_2M.cairo (JIT) 1.373 ± 0.008 1.360 1.386 1.00 ± 0.01
head factorial_2M.cairo (AOT) 1.368 ± 0.013 1.351 1.394 1.00

Base

Command Mean [s] Min [s] Max [s] Relative
base fib_2M.cairo (JIT) 1.324 ± 0.012 1.310 1.355 1.02 ± 0.01
base fib_2M.cairo (AOT) 1.298 ± 0.007 1.282 1.306 1.00

Head

Command Mean [s] Min [s] Max [s] Relative
head fib_2M.cairo (JIT) 1.351 ± 0.034 1.323 1.414 1.04 ± 0.03
head fib_2M.cairo (AOT) 1.304 ± 0.004 1.299 1.308 1.00

Base

Command Mean [s] Min [s] Max [s] Relative
base linear_search.cairo (JIT) 1.332 ± 0.007 1.320 1.348 1.01 ± 0.01
base linear_search.cairo (AOT) 1.314 ± 0.007 1.306 1.326 1.00

Head

Command Mean [s] Min [s] Max [s] Relative
head linear_search.cairo (JIT) 1.338 ± 0.005 1.328 1.344 1.01 ± 0.01
head linear_search.cairo (AOT) 1.319 ± 0.009 1.305 1.335 1.00

Base

Command Mean [s] Min [s] Max [s] Relative
base logistic_map.cairo (JIT) 1.320 ± 0.008 1.300 1.328 1.02 ± 0.01
base logistic_map.cairo (AOT) 1.297 ± 0.007 1.287 1.309 1.00

Head

Command Mean [s] Min [s] Max [s] Relative
head logistic_map.cairo (JIT) 1.333 ± 0.005 1.327 1.343 1.02 ± 0.01
head logistic_map.cairo (AOT) 1.307 ± 0.006 1.298 1.319 1.00

@github-actions

github-actions Bot commented Oct 1, 2026

Copy link
Copy Markdown

Benchmarking results

Benchmark for program dict_insert

Open benchmarks
Command Mean [s] Min [s] Max [s] Relative
Cairo-vm (Rust, Cairo 1) 11.259 ± 0.033 11.222 11.311 6.27 ± 0.11
cairo-native (embedded AOT) 1.796 ± 0.030 1.767 1.871 1.00
cairo-native (embedded JIT using LLVM's ORC Engine) 1.796 ± 0.013 1.777 1.821 1.00 ± 0.02

Benchmark for program dict_snapshot

Open benchmarks
Command Mean [ms] Min [ms] Max [ms] Relative
Cairo-vm (Rust, Cairo 1) 514.2 ± 4.9 508.0 525.3 1.00
cairo-native (embedded AOT) 1618.5 ± 20.7 1594.8 1655.5 3.15 ± 0.05
cairo-native (embedded JIT using LLVM's ORC Engine) 1632.4 ± 15.7 1602.9 1658.6 3.17 ± 0.04

Benchmark for program factorial_2M

Open benchmarks
Command Mean [s] Min [s] Max [s] Relative
Cairo-vm (Rust, Cairo 1) 5.036 ± 0.038 4.993 5.128 2.99 ± 0.03
cairo-native (embedded AOT) 1.683 ± 0.008 1.664 1.695 1.00
cairo-native (embedded JIT using LLVM's ORC Engine) 1.686 ± 0.017 1.658 1.712 1.00 ± 0.01

Benchmark for program fib_2M

Open benchmarks
Command Mean [s] Min [s] Max [s] Relative
Cairo-vm (Rust, Cairo 1) 4.952 ± 0.023 4.912 4.987 3.06 ± 0.02
cairo-native (embedded AOT) 1.616 ± 0.007 1.604 1.625 1.00
cairo-native (embedded JIT using LLVM's ORC Engine) 1.642 ± 0.016 1.616 1.671 1.02 ± 0.01

Benchmark for program linear_search

Open benchmarks
Command Mean [ms] Min [ms] Max [ms] Relative
Cairo-vm (Rust, Cairo 1) 559.2 ± 7.2 552.9 575.4 1.00
cairo-native (embedded AOT) 1641.0 ± 11.7 1627.9 1666.9 2.93 ± 0.04
cairo-native (embedded JIT using LLVM's ORC Engine) 1666.5 ± 9.1 1647.8 1675.4 2.98 ± 0.04

Benchmark for program logistic_map

Open benchmarks
Command Mean [ms] Min [ms] Max [ms] Relative
Cairo-vm (Rust, Cairo 1) 469.6 ± 1.7 466.4 472.6 1.00
cairo-native (embedded AOT) 1618.2 ± 19.7 1602.5 1665.6 3.45 ± 0.04
cairo-native (embedded JIT using LLVM's ORC Engine) 1643.3 ± 13.7 1621.4 1669.4 3.50 ± 0.03

@TomerStarkware
TomerStarkware added this pull request to the merge queue Oct 1, 2026
Merged via the queue into main with commit 60cdf41 Oct 1, 2026
29 checks passed
@TomerStarkware
TomerStarkware deleted the tomer/fix-alloca-in-loops branch October 1, 2026 12:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants