Skip to content

Take package-shaped repeats and classes out of the per-character phase loop - #3

Merged
proggeramlug merged 1 commit into
mainfrom
perf/package-regex-shapes
Sep 27, 2026
Merged

proggeramlug merged 1 commit into
mainfrom
perf/package-regex-shapes

Conversation

@proggeramlug

@proggeramlug proggeramlug commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Perry's package benchmarks (Perry benchmarks/packages/PROFILE.md) attribute about 17% of Perry's excess instructions over Node across eleven npm packages to regular expressions, roughly three quarters of that inside this matcher. The patterns doing it are short subjects with counted, folded and lazy repeats: uuid's validate, jws's JWS_REGEX, dotenv's line pattern, date-format tokens. Each paid several phase round trips per character. This PR moves those shapes out of the per-character phase loop. The VM and its answers are unchanged, and so is the charge for every character on the paths it touches.

Changes

  • Folded classes are closed at compile time (compiler/classes.rs). A folded class of at most 256 characters with no property compiles to its case closure as an unfolded class, sorted and merged: [0-9a-f]/i becomes three ranges, and [a-z]/iu becomes [A-Za-zſK]. Case equivalence is symmetric (equivalents walks one cycle from any member), so the closure accepts exactly the characters that match-time folding accepted. Those classes can now use the inline scan and the byte run.
  • Counted minimums, bounded allowances and lazy minimums as runs (executor/atom.rs). The run loop and the byte run used to serve only a greedy unbounded repeat that had already met its minimum. They now take the whole quota: what the repeat still owes, then a greedy allowance.
  • Lazy repeats skip endpoints their continuation cannot start at (atom_commit → lazy_skip). This is the lazy counterpart of the existing greedy retreat filter, and it uses the same condition (the continuation's first consumed instruction after its SAVEs). Where that condition rejects the next character, extending directly reaches the same state that commit, fail, pop, rollback and extend reached. On ASCII storage the skip walks bytes. The condition test is uncharged, because a commit that asks for more frames is resumed and would otherwise pay for it twice. A quantum that runs out part way resumes the same commit.
  • An unfolded class is decided inside the instruction (class_decide) whenever the decision fits the class phase's batch. It examines the same ranges in the same order and charges them the same way.
  • ATOM_CLASS_RANGES goes from 8 to 16, so \s (ten ranges) takes the inline class test in repeats.

The docs are updated in casefold.md, classes.md, repetition.md and performance.md, which has the measurement table.

Measurements

Instructions per search, perf stat -e instructions:u. Each figure is the difference between 11,000 and 1,000 warm finds divided by 10,000. Both arms use the same driver on one idle AMD EPYC 9254, built against 3fcc37e (0.1.10) and against this branch:

Case 0.1.10 This PR Change
uuid validate regex (/i) 20,211 5,891 −70.9%
same, no i 15,202 5,680 −62.6%
jws JWS_REGEX, 155-char token 89,373 13,491 −84.9%
dotenv line pattern 38,569 36,719 −4.8%
^l-\d{1,2}$/i (node-cron) 2,828 2,054 −27.4%
^\d+$ 2,345 2,059 −12.2%
(\w+)@(\w+)\.com 7,950 6,881 −13.4%
[a-z]+[0-9]+ 10,621 10,116 −4.8%
\w+! over 60 chars 3,185 2,869 −9.9%
(?<=\$)\d+ 2,346 2,058 −12.3%
<(.+?)> 7,653 2,592 −66.1%
^\s+|\s+$ 4,390 2,596 −40.9%
(\w)\1 6,245 5,577 −10.7%
dayjs-style format tokens, g 4,374 3,340 −23.6%
\p{L}+/u 2,941 2,814 −4.3%
%..|. (split piece) 1,342 1,346 +0.3%
needle 1,386 1,389 +0.2%
needle/i 1,754 1,758 +0.2%
cat|dog|bird after 200 chars 4,673 4,678 +0.1%
zzz over 1,000 chars (miss) 2,065 2,070 +0.2%

The five +0.x% rows go through paths this PR does not change. They move by 3–5 instructions per search, and I did not isolate where those come from.

In Perry on the same host, built against this branch through a local [patch.crates-io], with both arms built the same way and no auto-optimize:

  • uuid REGEX.test: 23,500 → 8,225 instructions per call
  • JWS_REGEX.test: 96,711 → 15,394
  • ^\d+$ test: 4,453 → 4,163
  • dotenv's exec loop over a 36-line document: 1,609,585 → 1,559,285

Checks run (all on this branch)

  • cargo fmt --check, cargo check --all-targets, cargo clippy --all-targets -D warnings
  • cargo test --locked and cargo test --locked --release: all pass
  • Every differential harness in the README, run against Node 26.8.1 (the CI pin), all passing: engine (with the two CI allow flags), casefold, names, legacy, admission, repetition, atom-filter / sequences / candidate / end-candidate / classes (each plain and with --quantum 1|17 --relocate --grow), modifiers, unicode-sets, sets, properties, input, reference.mjs --check, reference.test.mjs, and check-test262.mjs against test262 4249661388e5
  • New tests:
    • folded_class_closure_matches_match_time_folding (tests/classes.rs) compares each closed class against the same class plus the Private Use Area, which is too large to close and has no case mappings. It covers plain and negated classes, i and iu, the BMP's cased blocks and astral cased blocks.
    • counted_runs_and_lazy_skips_keep_general_vm_captures_and_errors (tests/atom_repeat.rs) runs 14 patterns × 4 flag sets × 25 subjects × byte/UTF-16 storage against the general repetition instructions, and checks every insufficient work allowance.
    • I sabotage-checked both tests: dropping U+212A from the closure, and letting the lazy skip pass a k continuation, each makes its test fail.

Notes

  • refuses_what_it_cannot_generate no longer lists /[a-z]/i. It now compiles to an unfolded class, which the native emitter accepts.
  • Charges change for folded classes on the closed path, which are now charged like any unfolded class, and for \s-sized classes in repeats, which are now charged per class rather than per range visited. Paused and unpaused searches still reach the same total; the resumption tests and the q1/q17 harness runs cover that.
  • Not run: perex-bench against V8 wall-clock (bench/compare.sh). The figures above are instruction counts, not CPU time.
  • Perry adoption: Perry is on 0.1.9 and adopting this needs a release. 0.1.10's RunError needs a small Perry adaptation, which is ready on Perry branch wip/perf-regex-matcher, and the release then has to clear Perry's 7-day min-publish-age soak.

Summary by CodeRabbit

  • Performance
    • Improved matching speed for character classes, counted repeats, and lazy repeats, particularly in package-shaped searches.
    • Reduced execution overhead for eligible character-class searches while preserving matching results and work limits.
  • Documentation
    • Added guidance on case-insensitive class matching, class execution, repeat optimizations, and performance comparisons.

…phase loop

Perry's package benchmarks put a sixth of its excess over Node across eleven
npm packages in regular expressions, three quarters of that in this matcher,
and the patterns doing it were short subjects with counted, folded and lazy
repeats: uuid's validate, jws's JWS_REGEX, dotenv's line pattern, date format
tokens. Each spent its time in phase round trips per character.

- A folded class of at most 256 characters with no property is compiled to
  its case closure as an unfolded class. Equivalence is symmetric, so this
  answers every character as folding at match time did, and the class can now
  use the inline scan and the byte run.
- A repeat's owed characters, a bounded allowance and a lazy minimum are
  walked by the run loop and the byte run, not one per round trip. Charges
  are unchanged per character.
- A lazy repeat's commit extends directly past endpoints its continuation's
  first consumed instruction rejects, byte-wise over ASCII storage, reaching
  the state commit, fail, rollback and extend reached. The condition's test is
  uncharged so a resumed commit does not pay it twice.
- An unfolded class whose decision fits the phase's batch is decided in the
  instruction, examining and charging the same ranges.
- The inline class test takes up to sixteen ranges, so `\s` (ten) qualifies.

Instructions per search against 0.1.10: uuid validate -70.9%, JWS_REGEX
-84.9%, `<(.+?)>` -66.1%, trim -40.9%, format tokens -23.6%, dotenv's line
-4.8%; searches none of this reaches move by at most +0.3%
(docs/performance.md has the table and the method).

`[a-z]/i` is no longer refused by the native emitter, since it is now an
unfolded class, so it leaves that test's refusal list.
@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: bc28925c-84f3-49cd-8f22-64a68af5592a

📥 Commits

Reviewing files that changed from the base of the PR and between 3fcc37e and c9a541b.

📒 Files selected for processing (11)
  • docs/casefold.md
  • docs/classes.md
  • docs/performance.md
  • docs/repetition.md
  • src/compiler.rs
  • src/compiler/classes.rs
  • src/executor.rs
  • src/executor/atom.rs
  • src/native/emit.rs
  • tests/atom_repeat.rs
  • tests/classes.rs
💤 Files with no reviewable changes (1)
  • src/native/emit.rs

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.


📝 Walkthrough

Walkthrough

The changes add compile-time closure for eligible case-insensitive classes, direct decisions for eligible non-folded classes, and quota-bounded scans and lazy skips for repeated atoms. Tests compare optimized paths with general VM behavior, and documentation reports behavior and instruction-count comparisons.

Changes

Regex matching

Layer / File(s) Summary
Compile-time folded classes
src/compiler/classes.rs, src/compiler.rs, tests/classes.rs, docs/casefold.md, src/native/emit.rs
Eligible case-insensitive classes are replaced with sorted, merged closure ranges while preserving negation. Differential tests compare results with match-time folding. The case-insensitive [a-z] pattern is removed from the unsupported-pattern test cases.
Direct class decisions
src/executor.rs, src/executor/atom.rs, docs/classes.md, docs/repetition.md
Eligible non-folded classes can be decided directly within available work. Plain class scans use a 16-range threshold. The documentation describes decision conditions, range examination order, and work accounting.
Repeated-atom scanning and lazy skips
src/executor.rs, src/executor/atom.rs, tests/atom_repeat.rs, docs/repetition.md, docs/performance.md
Repeated atoms are scanned in quota-bounded runs. For eligible lazy repeats, the VM can extend past endpoints rejected by a continuation probe. Tests compare optimized and general VM results and check insufficient work allowances. Performance documentation reports instruction-count comparisons across package-derived and other workloads.

Priority: ⬇️ Low

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Refactor

Sequence Diagram(s)

sequenceDiagram
  participant Vm
  participant atom_scan
  participant atom_commit
  participant lazy_skip
  participant continuation_probe
  Vm->>atom_scan: scan repeated atom within quota and work limits
  atom_scan-->>Vm: commit when quota is reached
  Vm->>atom_commit: pass available work
  atom_commit->>lazy_skip: attempt lazy-repeat extension
  lazy_skip->>continuation_probe: inspect the continuation condition
  continuation_probe-->>lazy_skip: return probe condition or no supported probe
  lazy_skip-->>atom_commit: return skip outcome
Loading

Merge Risk: ⚪ Minimal · up to c9a54

No actionable merge-blocking risk is established; the change is mergeable after normal checks.

Security Architecture Review

Security architecture risk: 🟡 Moderate · up to c9a54

Some previously compilable case-insensitive patterns can now fail when callers provide tight compilation limits or storage. Matching equivalence has supporting tests, but the effect on applications that depend on successful compilation is not established.

Retained concerns

  • Medium · reliability · observed: Compile-time closure can make a formerly accepted folded class fail under the same caller-supplied work budget or range storage. For callers that require successful compilation of validation patterns, this changes availability at the compilation boundary.
Security review details

Security Blast Radius

  • inferred — If a host compiles caller-controlled patterns with tight shared limits or scratch storage, the changed admission behavior can produce compilation errors on that path. No particular host, tenant, privileged sink, or independently attackable deployment was established.

Security Findings and Attack Paths

  • inferred — No bypass or privilege-gain path was established. The conditional security-relevant outcome is loss of regex compilation availability where an application depends on a bounded-budget or exactly sized range allocation; how such an application handles the error is unknown.

Trust Boundaries and Controls

  • observed — Compilation work remains charged to the caller-supplied budget, with exhaustion returned as WorkLimit. Runtime class decisions charge comparisons, and unsupported direct decisions retain the resumable class path.

Resilience and Maintainability Implications

  • inferred — The inspected atom transitions preserve charged progress across ordinary execution, failure, and resumption; the outstanding containment question is whether callers provision for the new compilation cost and handle its errors safely.

Hardening Proposals

  • proposed — Consider retaining match-time folding when closure cannot fit the caller’s remaining range storage, and document or exercise compilation-budget requirements for bounded callers. Verify native-emission equivalence separately if newly eligible patterns use that backend.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 45.83% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 6 files. (4 skipped: … Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main optimization: moving package-shaped repeats and classes out of the per-character phase loop.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Full details: Docstring Coverage

Explanation

Docstring coverage is 45.83% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 24 functions across 6 files. (4 skipped: 4 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant