Skip to content

Improve CUDA/LLVM compatibility - #2181

Open
tgrant-nv wants to merge 5 commits into
AcademySoftwareFoundation:mainfrom
tgrant-nv:fix-cuda-compat
Open

tgrant-nv wants to merge 5 commits into
AcademySoftwareFoundation:mainfrom
tgrant-nv:fix-cuda-compat

Conversation

@tgrant-nv

@tgrant-nv tgrant-nv commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Description

This PR addresses some compatibility issues that crop up when mixing certain LLVM and CUDA versions.

We had discussed not pushing these changes and just accepting gaps in the compatibility matrix. But I recently started experiencing build issues due to conflicting use of __noinline__ between CUDA and libstdc++ that require some kind of fix. The fix we arrived at is conceptually and mechanically similar to the LLVM/CUDA compatibility fixes, so I have decided to submit them as a bundle.

libstdc++ Compatibility Fix

I am actually not sure what changed in my environment to expose this problem, but in the last week or two I started seeing errors like this:

/usr/lib/gcc/x86_64-linux-gnu/13/../../../../include/c++/13/bits/basic_string.h:2544:22: error: type name does not allow function specifier to be specified
/usr/local/cuda-12.9/include/crt/host_defines.h:91:24: note: expanded from macro '__noinline__'
        __attribute__((noinline))
                       ^

These errors are due to a collision between CUDA's __noinline__ macro and the use of the __noinline__ attribute in GNU libstdc++. The solution is to undefine that macro before including the <memory> and <string> headers, and then to restore the definition afterwards.

CUDA Compatibility Fixes

The CUDA compatibility fixes address two issues:

  • The "missing" texture_fetch_functions.h that was removed from CUDA 13 but is unconditionally included by older versions of Clang/LLVM. The fix is to include a stub header when using CUDA 13 with older LLVM versions.
  • The undefined _NV_RSQRT_SPECIFIER macro required by CUDA 13.2+, but left undefined in LLVM versions older than 22.1.2.

With these fixes applied I have been able to build OSL with GPU support using CUDA versions from 12.8 to 13.3, and with LLVM 15.0.7 through 22.1.8.

Tests

No new tests have been added, and no change in test results was observed.

Checklist:

  • I have read the guidelines on contributions and code review procedures.
  • I have read the Policy on AI Coding Assistants
    and if I used AI coding assistants, I have an Assisted-by: TOOL / MODEL
    line in the pull request description above.
  • I have updated the documentation if my PR adds features or changes
    behavior.
  • I am sure that this PR's changes are tested in the testsuite.
  • I have run and passed the testsuite in CI before submitting the
    PR, by pushing the changes to my fork and seeing that the automated CI
    passed there. (Exceptions: If most tests pass and you can't figure out why
    the remaining ones fail, it's ok to submit the PR and ask for help. Or if
    any failures seem entirely unrelated to your change; sometimes things break
    on the GitHub runners.)
  • My code follows the prevailing code style of this project and I
    fixed any problems reported by the clang-format CI test.

Assisted-by: OpenAI Codex / GPT-6

CUDA 13 removed texture_fetch_functions.h, but older Clang CUDA wrappers can still include it while generating device bitcode. Search an OSL compatibility directory first and provide an empty stub because OSL does not use the legacy texture-reference API.

Assisted-by: OpenAI Codex / GPT-6
Signed-off-by: Tim Grant <tgrant@nvidia.com>
CUDA 13.2 added _NV_RSQRT_SPECIFIER, but older Clang CUDA wrappers include math_functions.hpp without defining it. Supply a narrow compatibility wrapper that mirrors the definition in newer Clang, including noexcept(true) on glibc 2.42.

With the header mismatch handled locally, remove the configuration-time rejection of CUDA 13.2+ with LLVM older than 22.1.2.

Assisted-by: OpenAI Codex / GPT-6
Signed-off-by: Tim Grant <tgrant@nvidia.com>
For non-Windows LLVM versions older than 18, force-include an OSL header that temporarily hides the CUDA __noinline__ macro while loading string and memory. Restore the macro afterward so device code can still use it; standard-library include guards prevent the affected declarations from being parsed again.

Track the pre-include header as a bitcode build dependency.

Assisted-by: OpenAI Codex / GPT-6
Signed-off-by: Tim Grant <tgrant@nvidia.com>
The cache declaration accidentally used a literal dollar-prefixed variable name while NVCC_COMPILE reads OSL_EXTRA_NVCC_ARGS. Declare the intended cache variable so builds can inject dependency-specific NVCC flags without hard-coding them in OSL.

Assisted-by: OpenAI Codex / GPT-6
Signed-off-by: Tim Grant <tgrant@nvidia.com>
Signed-off-by: Tim Grant <tgrant@nvidia.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant