Repository navigation
opt-level=0 does not do constant folding #136366
Description
Activity
- addedneeds-triageThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triagingThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triaging
on Jan 31, 2025 Additional info: A google search of the term "constant folding" results in several articles that describe the process that calculates the literal expression and replaces the expression with a single value, the behavour I am advocating here. Some of these articles even state that this is a "compiler optimization". However, gcc, by default, applies this process even when the optimization level is set to none; therefore, it seems to me that the gcc developers considered this to be a desirable feature that doesn't need an option to invoke. gcc does have a means to inhibit constant folding for floating point numbers, in particular, with the option
-frounding-math.The semantics of these values is that they are indeed potentially evaluated at runtime.
Consider instead https://rust.godbolt.org/z/637djG7cd which uses the following code:
#[inline(never)] pub fn whatever() -> i32 { let _blah = const { 1 << (5 + 2) }; _blah }
The use of the
const {}forces the value to be evaluated at compile time.I am not entirely sure why
1 << (5+2)is not folded but1 << 7is, the difference in the MIR is not so great, but the LLVMIR we emit has already had1 << 7computed in it.The expression is correctly folded with overflow checks disabled;
https://rust.godbolt.org/z/6c1x3sae4Looks like overflow checks require
opt-level ⩾ 1to be removed, so it prevents constant folding?
E.g.1 + 2isn't folded as3, andn << mis folded as it can't overflow.Seems like a regression from stable to stable;
https://rust.godbolt.org/z/cj16n7hEz@rustbot label +A-codegen +C-optimization +I-heavy +T-compiler
Reacted by Jubilee- addedA-codegenArea: Code generationArea: Code generationC-optimizationCategory: An issue highlighting optimization opportunities or PRs implementing suchCategory: An issue highlighting optimization opportunities or PRs implementing suchI-heavyIssue: Problems and improvements with respect to binary size of generated code.Issue: Problems and improvements with respect to binary size of generated code.T-compilerRelevant to the compiler team, which will review and decide on the PR/issue.Relevant to the compiler team, which will review and decide on the PR/issue.
on Jan 31, 2025 Oh, thank you for identifying this is a regression! I was wondering.
- addedregression-from-stable-to-stablePerformance or correctness regression from one stable version to another.Performance or correctness regression from one stable version to another.
on Jan 31, 2025 - addedI-prioritizeIssue needs a team member to assess the impact. Will be replaced by P-{low,medium,high,critical}Issue needs a team member to assess the impact. Will be replaced by P-{low,medium,high,critical}
on Jan 31, 2025 searched toolchains nightly-2023-03-04 through nightly-2023-05-27 ******************************************************************************** Regression in nightly-2023-04-16 ******************************************************************************** fetching https://static.rust-lang.org/dist/2023-04-15/channel-rust-nightly-git-commit-hash.txt nightly manifest 2023-04-15: 40 B / 40 B [=============================================================================================================================================================================================================================================================] 100.00 % 2.29 MB/s converted 2023-04-15 to 84dd17b56a931a631a23dfd5ef2018fd3ef49108 fetching https://static.rust-lang.org/dist/2023-04-16/channel-rust-nightly-git-commit-hash.txt nightly manifest 2023-04-16: 40 B / 40 B [=============================================================================================================================================================================================================================================================] 100.00 % 3.09 MB/s converted 2023-04-16 to 5cdb7886a5ece816864fab177f0c266ad4dd5358 looking for regression commit between 2023-04-15 and 2023-04-16 fetching (via remote github) commits from max(84dd17b56a931a631a23dfd5ef2018fd3ef49108, 2023-04-13) to 5cdb7886a5ece816864fab177f0c266ad4dd5358 ending github query because we found starting sha: 84dd17b56a931a631a23dfd5ef2018fd3ef49108 get_commits_between returning commits, len: 10 commit[0] 2023-04-14: Auto merge of #110331 - matthiaskrgr:rollup-9vldvow, r=matthiaskrgr commit[1] 2023-04-14: Auto merge of #110197 - cjgillot:codegen-discr, r=pnkfelix commit[2] 2023-04-15: Auto merge of #110142 - Mark-Simulacrum:reduce-core-counts, r=pietroalbini commit[3] 2023-04-15: Auto merge of #109802 - notriddle:notriddle/rustdoc-search-generics-nested, r=GuillaumeGomez commit[4] 2023-04-15: Auto merge of #110335 - asomers:rust-gdb-freebsd, r=jyn514 commit[5] 2023-04-15: Auto merge of #109900 - cjgillot:disable-const-prop, r=oli-obk commit[6] 2023-04-15: Auto merge of #110323 - lcnr:dropck-uwu, r=compiler-errors commit[7] 2023-04-15: Auto merge of #110349 - rust-lang:pa-bump-1.71.0, r=pietroalbini commit[8] 2023-04-15: Auto merge of #110227 - klensy:bs-win, r=Mark-Simulacrum commit[9] 2023-04-15: Auto merge of #110361 - ehuss:disable-jobserver-error, r=Mark-SimulacrumThis was probably due to the following PR which changed const prop to happen at higher MIR opt levels: #109900
- removedC-bugCategory: This is a bug.Category: This is a bug.needs-triageThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triagingThis issue may need triage. Remove when done. See docs forge.rust-lang.org/release/issue-triaging
on Jan 31, 2025 If this sort of regression is a problem for you, you should definitely be using
opt-level = 1. This use case is not a consideration for opt level 0 (and I don't think it should be).
opt level 0 is mostly guided by compile times. If this optimization benefits compile times we could do it, if it doesn't we won't.
The PR that regressed this had a positive compile time impact, so I will classify this as intentional.Reacted by JubileeReacted by scottmcm- removedregression-from-stable-to-stablePerformance or correctness regression from one stable version to another.Performance or correctness regression from one stable version to another.I-prioritizeIssue needs a team member to assess the impact. Will be replaced by P-{low,medium,high,critical}Issue needs a team member to assess the impact. Will be replaced by P-{low,medium,high,critical}
on Feb 1, 2025 - changed the title
[-]Literal Expressions are Not Resolved at Compile Time[/-][+]opt-level=0 does not do constant folding[/+]on Feb 1, 2025 Closing as not-a-bug. At opt-level=0, the compiler doesn't care about binary size (or runtime performance).
Reacted by noraReacted by scottmcmThis is an unfortunate decision in my opinion; I admit, I know little about actual compiler design but I don't understand how interpreting a constant expression and producing a single value (an int in the example) can result in an increase in compile times compared to the amount of object code generated by the compiler to produce a computation result that is run on the target. This is especially important in the embedded field where debugging on an embedded target with limited flash memory may become more difficult if target code size suddenly blooms when setting opt-level=0 to ensure breakpoints are able to land on the correct line of source code.
The intuition for why such passes can hurt compile times is that they do a lot of extra checks for all the cases where you don't have such a case.
That said, a more minimal constant folding pass could potentially be a compile time win. If someone wants to implement that and it's good for compile times that sounds good.for debugging: does opt-level 1 have an insufficient debugging experience for you?
@Noratrieb: To be honest, I don't know, considering the very limited introduction I've had with this language and compiler. Nevertheless, in my over 25 years experience with a certain commercial embedded compiler used by my employer, I have encountered many instances where where optimization has interfered with useful debugging and disabling optimization for even single compilation units can be enough to cause the link to fail due to insufficient memory.
This is an unfortunate decision in my opinion; I admit, I know little about actual compiler design but I don't understand how interpreting a constant expression and producing a single value (an int in the example) can result in an increase in compile times compared to the amount of object code generated by the compiler to produce a computation result that is run on the target.
Simply put: Merely generating code is very fast. It is optimization and reasoning about code in more complicated ways that takes more time.
Disabling optimization causing linkage to fail sounds very weird, to me.
Due to the Rust compiler having a somewhat... legalistic approach to code generation, at times,
-Copt-level=1is often used with software written in Rust that people intend to debug. In a language with potentially a lot of layers of abstraction in a program, removing only some of them still leaves enough code to reason about what is going on. Indeed, because idiomatic Rust tends to use abstractions that are then optimized away, sometimes it's easier to see what is going on after enabling that low level of optimization.
Consider this code:
When inspecting the assembly (from gdb), I expected to see this happen:
Instead, this happened:
This is the compiler actually generating code that computes the value of 1 << (5 + 2) at run time, along with a whole bunch of stuff that appears to relate to the safety of performing such an operation on variable data at run time.
This is unexpected. The entire expression on the right side of the assignment statement, 1 << (5 + 2), involves no variables. This can be safely resolved at compile time.
Strangely, the following code,
actually produces the expected result. This isn't consistent behaviour. Why is this expression, 1 << 7, resolved at compile time while 1 << (5 + 2) is not?
If I set the optimization level to 1 in the Cargo.toml as follows:
The expected code is produced with either expression (as long as the variable is used later, of course). However, I do not see how this is thought of as an optimization; optimization is supposed to be applied to executed code that involves actual data in variables or references. Since both 1 << (5 + 2) and 1 << 7 are literal expressions that resolve to the same identical value, and they do not involve data, they do not need to be "optimized". Each expression ought to be treated as a single entity and treated the same.
Rationale:
I suppose that this will be viewed as trivial in the context of a PC or server environment. However, in an embedded environment, the additional unnecessary code may result in consumption of valuable resources, if optimization must be disabled in order to perform debugging activites on an embedded system. This might even cause the unoptimized code to not compile at all on a sytem with tight memory constraints. Therefore, it is desirable, and in my opinion, a show stopper, that evaluation of literal expressions ought to be resolved at compile time regardless of optimization level.
Meta
rustc --version --verbose:Backtrace