test(otel): add GPU, Neuron, and EFA DRA-path integration tests - #753
Open
samehkhalil wants to merge 1 commit into
Open
samehkhalil wants to merge 1 commit into
samehkhalil wants to merge 1 commit into
Conversation
samehkhalil
force-pushed
the
test/multi-efa-dra-per-device-correlation
branch
from
September 8, 2026 16:28
077c30e to
f803c04
Compare
samehkhalil
force-pushed
the
test/multi-efa-dra-per-device-correlation
branch
from
September 8, 2026 16:56
f803c04 to
d43ae5d
Compare
Add integration coverage for the awsdevicepodcorrelation processor's DRA (Dynamic Resource Allocation) path, mirroring the device-plugin GPU/Neuron/EFA correlation tests. Each package exposes its devices via a DRA driver (through a ResourceClaimTemplate) instead of the device-plugin resource, and asserts per-device pod correlation. - test/otel/multi_efa_dra: EFA via dranet (driver dra.net); efaburn claims one of two devices, the other stays unclaimed. Guards the per-device correlation collapse (ResourceSlice keying via dra.net/rdmaDevice plus the groupbyattrs split before the resource-level promote). - test/otel/neuron_dra: Neuron via the AWS Neuron DRA driver (DeviceClass neuron.aws.com). Single Trainium device (trn1.2xlarge); the claimed device's two cores attribute to the burn pod and to no other pod. The Neuron DRA driver supports Trainium only, so this targets trn1.2xlarge. - test/otel/gpu_dra: GPU via the NVIDIA DRA driver (DeviceClass gpu.nvidia.com). g4dn.12xlarge (4 GPUs); one claimed GPU correlates to the burn pod, the other three stay uncorrelated. Asserts device count, consecutive indices, and all DCGM metrics per device. New terraform modules under terraform/eks/daemon (otel-multi-efa-dra, otel-neuron-dra, otel-gpu-dra) install the DRA driver in place of the device plugin and apply a ResourceClaimTemplate burn workload. The processor uses the GA resource.k8s.io/v1 DRA API (available since Kubernetes 1.34), so the clusters run k8s 1.35 like the rest of the suite. Wired into the test case generator. Requires a chart carrying the DRA correlation config and resource.k8s.io RBAC.
samehkhalil
force-pushed
the
test/multi-efa-dra-per-device-correlation
branch
from
September 22, 2026 19:32
d43ae5d to
3073b10
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description of the issue
EFA metrics on EKS can be exposed to pods two ways: the EFA device plugin
(
vpc.amazonaws.com/efa) and Dynamic Resource Allocation (DRA), where EFAdevices are allocated via a DRA driver (dranet, driver
dra.net) andResourceClaims. The agent's OTel Container Insights pipeline supports both, andthe
awsdevicepodcorrelationprocessor has a dedicated DRA code path that watchesResourceClaims/ResourceSlicesvia the K8s API and bridges the DRA deviceidentity (a PCI name, e.g.
pci-0000-00-1e-0) to the EFA metric label (e.g.rdmap0s30) via thedra.net/rdmaDeviceResourceSlice attribute.Today there is no integration coverage for the DRA path — only the device-plugin
path is exercised. This adds an end-to-end test so the DRA path is validated on a
real cluster: EFA metrics are emitted per device and correctly correlated to the
pods that claim them.
Description of changes
Add an integration test that provisions a cluster exposing EFA via DRA and
validates per-device EFA metric correlation end to end:
test/otel/multi_efa_dra/— queries the EFA metrics in CloudWatch andasserts the DRA path produces the expected per-device series and pod
attribution: each device is a distinct series with its EFA attributes
(
aws.efa.device, ENI, port), the device a pod claims is correlated to thatpod (name/namespace/container), and devices not claimed by any pod carry no pod
attributes.
terraform/eks/daemon/otel-multi-efa-dra/— provisions the cluster,installs dranet (
eks/aws-dranet) as the DRA driver, deploys theobservability chart and agent, and runs an
efaburnworkload that requests oneEFA through a
ResourceClaimTemplate.helm_chart_repo_urlvariable so the chart can be cloned from a fork whenvalidating chart changes not yet merged upstream (defaults to upstream).
Pinned to k8s 1.34 (the rest of the suite is on 1.35): the processor's DRA
informers watch
resource.k8s.io/v1beta1, which 1.34 still serves alongside theGA
v1. The generator entry is commented with this rationale; it moves to 1.35once the processor's DRA client is bumped to
v1.Dependencies (the EKS lane is green only after these land): the chart on
mainmust render the DRA correlation config (dra_device_typeson thedra.netdriver +dra.net/rdmaDevicekeying) and grant the agent ServiceAccountget/list/watchonresource.k8s.ioresourceclaims/resourceslices, and thereleased agent image must include the DRA processor path. Committed defaults
already point at that merged end-state (public agent image, chart
main, upstreamchart URL), so no follow-up edit is needed once those merge.
License
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.
Tests
Ran the module end to end on a live cluster (EKS 1.34,
c6in.32xlarge, EFA viadranet). All tests passed.
Verified against ground truth:
rdmap0s30andrdmap0s31.efaburn'sResourceClaimTemplateallocatedpci-0000-00-1e-0, which maps viadra.net/rdmaDevicetordmap0s30→ correlated to theefaburnpod(namespace
default, containerefaburn).rdmap0s31was unclaimed → no pod attributes.Each EFA metric (
efa_rx_bytes,efa_tx_bytes,efa_rx_dropped,efa_rdma_read_bytes) reported per device with correct DRA-based pod correlation.Cluster torn down after the run.