Skip to content

Describe a backend by what it can do, not by its name (#151) - #151

Closed
joy-xiaojizhang wants to merge 1 commit into
meta-pytorch:mainfrom
joy-xiaojizhang:export-D118187263
Closed

joy-xiaojizhang wants to merge 1 commit into
meta-pytorch:mainfrom
joy-xiaojizhang:export-D118187263

Conversation

@joy-xiaojizhang

@joy-xiaojizhang joy-xiaojizhang commented Sep 9, 2026 •

Copy link
Copy Markdown

Summary:

The generation path decided device behaviour with a single Jinja conditional:

{% if device_string == "xpu" %}
        if not hasattr(torch, 'xpu') or not torch.xpu.is_available():
{% else %}
        if not torch.cuda.is_available():
{% endif %}

An if/else over two known names is not an abstraction. It silently routes every
third backend down the CUDA path -- generating CUDA availability checks under
another backend's name -- and a backend defined outside this repository cannot
add an arm to it.

PlatformConfig now carries the five capabilities a generated harness actually
needs: availability_check, device_setup, synchronize_call, test_prelude
and default_num_workers. The template reads them; it no longer knows any
device name. register_platform() is the seam that lets a backend defined
elsewhere become selectable without editing anything here -- an explicit call
rather than an import-time decorator, so registration order stays something the
caller controls.

Three smaller corrections fall out of the same change:

  • DEFAULT_PLATFORM was declared and then ignored in favour of hardcoded
    "cuda" literals. It is now the value actually used.
  • TritonKernelAgent passed the raw target_platform to PromptManager while
    passing the normalized one to WorkerManager, so two places independently
    re-derived the same default. The resolved config is now used for both, and it
    is resolved before the worker count, because the backend supplies that default.
  • The two device='cuda' literals in the mock test-generation fallback now
    follow the selected backend.

A fake backend is registered alongside cuda and xpu: no accelerator, a
check that asserts nothing because there is nothing to assert, an empty
synchronize_call, and one worker. It is named for what it is, and its guidance
block says outright that nothing it produces is a performance claim -- the same
reasoning as the existing noop implementations in triton_kernel_agent.platform.
Empty synchronize_call is a real answer, not a gap to be filled with the CUDA
call.

No Meta-internal import enters the generic tree; a test asserts that.

Differential Revision: D118187263

@meta-cla meta-cla Bot added the CLA Signed This label is managed by the Meta Open Source bot. label Sep 9, 2026
@meta-codesync

meta-codesync Bot commented Sep 9, 2026

Copy link
Copy Markdown

@joy-xiaojizhang has exported this pull request. If you are a Meta employee, you can view the originating Diff in D118187263.

Summary:

The generation path decided device behaviour with a single Jinja conditional:

```
{% if device_string == "xpu" %}
        if not hasattr(torch, 'xpu') or not torch.xpu.is_available():
{% else %}
        if not torch.cuda.is_available():
{% endif %}
```

An if/else over two known names is not an abstraction. It silently routes every
third backend down the CUDA path -- generating CUDA availability checks under
another backend's name -- and a backend defined outside this repository cannot
add an arm to it.

`PlatformConfig` now carries the five capabilities a generated harness actually
needs: `availability_check`, `device_setup`, `synchronize_call`, `test_prelude`
and `default_num_workers`. The template reads them; it no longer knows any
device name. `register_platform()` is the seam that lets a backend defined
elsewhere become selectable without editing anything here -- an explicit call
rather than an import-time decorator, so registration order stays something the
caller controls.

Three smaller corrections fall out of the same change:

- `DEFAULT_PLATFORM` was declared and then ignored in favour of hardcoded
  `"cuda"` literals. It is now the value actually used.
- `TritonKernelAgent` passed the *raw* `target_platform` to `PromptManager` while
  passing the *normalized* one to `WorkerManager`, so two places independently
  re-derived the same default. The resolved config is now used for both, and it
  is resolved before the worker count, because the backend supplies that default.
- The two `device='cuda'` literals in the mock test-generation fallback now
  follow the selected backend.

A `fake` backend is registered alongside `cuda` and `xpu`: no accelerator, a
check that asserts nothing because there is nothing to assert, an empty
`synchronize_call`, and one worker. It is named for what it is, and its guidance
block says outright that nothing it produces is a performance claim -- the same
reasoning as the existing `noop` implementations in `triton_kernel_agent.platform`.
Empty `synchronize_call` is a real answer, not a gap to be filled with the CUDA
call.

No Meta-internal import enters the generic tree; a test asserts that.

Differential Revision: D118187263
@meta-codesync meta-codesync Bot changed the title Describe a backend by what it can do, not by its name Describe a backend by what it can do, not by its name (#151) Sep 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CLA Signed This label is managed by the Meta Open Source bot. meta-exported

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant