Conversation
erofs: instrument warm up cache
The carrier implements the OpenTelemetry "Environment Variables as Context Propagation Carriers" specification, which is how the shim hands the trace context to the processes it spawns. Signed-off-by: Maksym Pavlenko <pavlenko.maksym@gmail.com>
full diff: containerd/go-runc@v1.2.0...v1.2.1 Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
runc and the OCI hooks it spawns are separate processes, so the trace context can not reach them over ttrpc. Inject it into their environment instead, riding on the request context via go-runc's WithExtraEnv, which does not race with concurrent requests the way setting it on the shim process would. The otelttrpc interceptor extracts that context, so chain it first explicitly instead of leaving it to plugin initialization order. Also flush the spans the tracing plugin buffered before the shim exits. Signed-off-by: Maksym Pavlenko <pavlenko.maksym@gmail.com>
Send a known trace context with the API call and assert that the environment of the OCI hook runc executes carries it. The shim propagates the context only when built with the shim_tracing tag, so build the Linux integration job with it. The integration target also has to pass the project build tags to the test compile, like the other build targets already do. Signed-off-by: Maksym Pavlenko <pavlenko.maksym@gmail.com>
update logrus to remove legacy indirect dependencies Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
vendor: github.com/containerd/go-runc v1.2.1
Update CI to include Go 1.27
chore(api): update github.com/sirupsen/logrus v1.10.2
- Add support for creating timers with custom histogram buckets. - Update `github.com/prometheus/client_golang` to v1.20.5. - Update the minimum supported Go version to Go 1.21. - Improve HTTP handler instrumentation, including minor performance improvements and cleanup. - Avoid mutating shared label maps when creating namespaces and metrics. - Encapsulate the Prometheus collector used by `HTTPMetric`. - Improve package and API documentation. - Update `golang.org/x/sys` and other dependencies. full diff: docker/go-metrics@v0.0.1...v0.1.0 Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
Removes some dependencies, and fixes some concurrency issues. full diff: cncf-tags/container-device-interface@05ae4b5...0427870 Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
contains a fix for [GHSA-2v4p-qf9q-27wj] full diff: grpc/grpc-go@v1.83.1...v1.83.2 [GHSA-2v4p-qf9q-27wj]: GHSA-2v4p-qf9q-27wj Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
vendor: github.com/docker/go-metrics v0.1.0
A shim declared the mount types it performs itself through a comma separated string in the generic annotations map of its RuntimeInfo, discovered by invoking the shim binary a second time with -info. The values had to reproduce by hand the "<transform>/*" sentinel the mount manager constructs internally, there was nowhere to put anything but a name, and it was keyed by runtime name rather than shim instance, so it could not reflect anything about the environment a particular shim is running in. Remove it along with loadShimInfo, shimInfo and the shimInfos cache, and the runc.v2/runhcs.v1 shortcut that only existed to avoid paying for it on the default runtimes. This is a breaking change targeted at 2.4; known usage is limited to Nerdbox, which can adopt the shim capability protocol added in the following commits after this merges. getRuntimeInfo itself stays: it is also used by PluginInfo and validateRuntimeFeatures. Signed-off-by: Derek McGowan <derek@mcg.dev>
A shim that performs some mount types or transformations itself, so that containerd's mount manager does not perform them on its behalf, had no way to say so other than a comma separated string in the generic annotations map of its RuntimeInfo. Add a containerd.types.MountCapabilities extension a shim can attach to its BootstrapResult, naming the mount types and transforms it handles. Detail lives in an extension rather than fields of BootstrapResult so that containerd can ignore an extension whose type it does not recognize, letting a shim advertise unconditionally. Presence of the extension is enough: it is not gated behind a Capability enum value, since with no other consumer of that field an enum value here would only ever duplicate what the extension's presence already says, while still permitting states with no single correct interpretation, such as the enum value being set with no extension attached, or the reverse. This is additive: only a new field is added to BootstrapResult, and buf breaking is clean. Signed-off-by: Derek McGowan <derek@mcg.dev>
Pull in the mount capabilities extension. The api module is consumed as a published version, so a local replace is required to build against the unreleased proto changes. The replace is added by 'make protos' whenever api/next.txtpb changes. It must be dropped and go.mod pinned to a released api version before this is merged. This also moves the root module's own github.com/sirupsen/logrus requirement from v1.10.1 to v1.10.2, which is not an intended dependency bump. The published api module (v1.12.0-beta.0) requires logrus v1.9.3, so it does not affect the root's resolution; the local api/go.mod on upstream main already requires v1.10.2, and once the replace substitutes it in, MVS takes the higher of the two. Verified by running 'go mod tidy' on a copy of this tree with the replace directive removed: the root stays at v1.10.1. This reverts on its own once the replace above is dropped. Signed-off-by: Derek McGowan <derek@mcg.dev>
parseStartResponse decodes a JSON bootstrap response into the deprecated three-field client.BootstrapParams, so every other field of BootstrapResult is silently discarded. This is reached when reloading a bundle: writeBootstrapParams stores bootstrap.json with encoding/json, and readBootstrapParams parses it back through here. It matters once BootstrapResult carries extensions, because a container joining an existing sandbox shim recovers that shim's result from bootstrap.json rather than from a fresh shim start, and so would have seen none of them. Decode into BootstrapResult instead. The stored format does not change and neither reader nor writer changes what it produces; encoding/json already round-trips every field, including extensions, whose Any is an opaque type URL and byte slice to it and so needs no type resolution, and including a capability value this version of containerd does not define, which must survive rather than being silently dropped. This also removes the last use of the deprecated client.BootstrapParams. Also fix TestRestoreBootstrapParams, which compared protobuf messages with reflect-based deep equality. That is not sound, as generated types carry internal state that varies with how the message was built. Signed-off-by: Derek McGowan <derek@mcg.dev>
A caller that applies a mount transform itself had to say so by naming the mount type "format/*", a sentinel the mount manager constructs internally by concatenating a transform name with "/*". Both sides had to independently know and agree on that convention, and the resulting string is neither a mount type nor a transform name. Add AllowTransforms, and WithAllowTransform to set it, so a transform can be claimed by its own name instead. The "/*" convention is removed rather than kept as a second accepted spelling: it existed only to be produced by the runtime-allow-mounts annotation, which is removed in an earlier commit in this series, and there is no other producer of it in this tree. Signed-off-by: Derek McGowan <derek@mcg.dev>
Activate's planning phase — deciding, for each mount, which transforms the manager must apply and whether the manager or the caller performs the mount itself — was inline in a bolt transaction spanning hundreds of lines, and every existing test of it goes through the full Activate path, which needs root to mount anything for real. Extract it into planActivation, a pure function of the mount list, activation options, and registered transforms/handlers. No behavior change; this is what makes the next commit's fix testable without root. Signed-off-by: Derek McGowan <derek@mcg.dev>
A claimed transform was recorded but never actually left unapplied: every parsed transformer in a mount's chain was applied in full regardless of what the caller had claimed, so a shim advertising it performs "mkdir" in "format/mkdir/overlay" never got the chance to, since the manager silently applied it anyway. This predates the shim capability protocol — it was reachable before through the "format/*" annotation convention — but this series turns it into a documented, supported way for a shim to claim a transform, so a shim relying on it (at least one is: Kata's EROFS snapshotter shim, per its annotation usage) would find its claim silently ignored. A claimed transform can only be honored as a suffix of the chain: each transform's input is the previous one's output, so an outer, unclaimed transform must still be applied even when an inner one is claimed, since the inner one cannot run without it. planActivation now tracks, per mount, how much of its chain — starting from the beginning — the manager must apply itself: through the last unclaimed transform, which by construction covers every transform before it too. Applying only that prefix and returning the true, contiguous claimed suffix to the caller is what actually makes the claim meaningful. This is only observable for mounts[firstSystemMount], the one mount partially transformed by the manager and returned to the caller to finish; every earlier mount is fully internal to the manager and was, and remains, always fully transformed regardless of any claim. While changing this expression, also guard it against the case where firstSystemMount reaches len(mounts) — every mount consumed inside the manager, with no system mount left to transform — which indexes out of range today. This is a distinct, pre-existing latent panic (for example, reachable via WithTemporary combined with a transformed mount type), tracked separately from the claim-handling fix above; guarding it here only prevents a crash newly reachable by touching this line, it does not address the mount.ActivationInfo.System non-empty invariant that scenario also violates. Signed-off-by: Derek McGowan <derek@mcg.dev>
A shim declared the mount types it performs itself through the annotation removed in an earlier commit, discovered by invoking the shim binary a second time with -info. Read them instead from the MountCapabilities extension a shim advertises in its bootstrap result, which is already available from the shim start that Create() performs anyway. This requires starting the shim before activating the task's mounts. Starting a shim does not need the rootfs; only the task.Create call consumes it. Reordering means a failure after the shim has started must now tear the shim down, so the teardown that was inlined in the task.Create error path becomes a deferred cleanup covering the whole window, reusing the existing cleanupShimTask helper. That defer is registered after the mount deactivation defer so that the shim, which may still be using the mounts, is stopped first. Extract the mount activation logic into a new taskMountController in task_mounts.go, rather than growing it further inline in Create(): it already has its own state transitions (owned vs. reused activation, NotImplemented, AlreadyExists) independent of shim startup, and this keeps that reasoning testable without spawning a shim. Also fix two latent nil pointer dereferences: Create() and Delete() both called the mount.Manager interface directly with no nil check, which panics whenever the mount manager plugin is not registered. taskMountController.Activate and Deactivate treat a nil manager as nothing to do. Signed-off-by: Derek McGowan <derek@mcg.dev>
When a container joins an existing sandbox via a Task API address supplied directly by the sandbox controller (rather than by restoring bootstrap.json), ShimManager.Start synthesized a BootstrapResult carrying only the version, protocol and address. A shim that advertised MountCapabilities at sandbox startup was therefore treated as if it had advertised nothing for every container that joins the sandbox this way, silently ignoring mount handling it explicitly claimed. Recover the extensions from the sandbox's shim instance, which containerd already holds in memory, via the existing shimCapabilities interface. This is best effort: an external sandboxer whose shim instance containerd does not track degrades to no extensions rather than failing sandbox startup, matching today's behavior. Reported independently by fuweid and Copilot on #14002. Signed-off-by: Derek McGowan <derek@mcg.dev>
Kata's EROFS snapshotter shim set containerd.io/runtime-allow-mounts before the MountCapabilities bootstrap extension existed (kata-containers/kata-containers#12763), so removing the annotation outright would break it and anyone else who adopted it early with no migration window. Rather than keep the annotation as a first-class, permanently supported mechanism, confine falling back to it in task_mounts_deprecated.go: when a shim's bootstrap result carries no MountCapabilities extension, its runtime name is checked against the annotation as a migration path, with the result cached per runtime name for the life of the daemon, same as the removed mechanism did. Two runtimes known to never set it, runc.v2 and runhcs.v1, are skipped outright rather than queried. The extension stays authoritative whenever a shim provides it: this fallback only runs in its absence, and is confined to one file so it can be deleted outright once known early adopters have migrated. Signed-off-by: Derek McGowan <derek@mcg.dev>
Add a registry entry for the containerd.types.MountCapabilities extension, and update the mount and bootstrap protocol docs for extensions replacing the reserved-for-future-use capabilities field, and for the removal of the runtime-allow-mounts annotation and the "/*" transform convention. Signed-off-by: Derek McGowan <derek@mcg.dev>
Neither doc previously said how much of a claimed transform chain is actually honored; a reader could assume claiming any transform in a chain was sufficient on its own. State the suffix rule plainly, with the format/mkdir/overlay example the code itself now enforces, and call out that format in particular can only ever be claimed as part of a suffix that covers what follows it, never alone, since it resolves mount points internal to the mount manager. Signed-off-by: Derek McGowan <derek@mcg.dev>
Shim mount handler protocol
Bumps [github.com/prometheus/client_golang](https://github.com/prometheus/client_golang) from 1.24.0 to 1.24.1. - [Release notes](https://github.com/prometheus/client_golang/releases) - [Changelog](https://github.com/prometheus/client_golang/blob/main/CHANGELOG.md) - [Commits](prometheus/client_golang@v1.24.0...v1.24.1) --- updated-dependencies: - dependency-name: github.com/prometheus/client_golang dependency-version: 1.24.1 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com>
To make task.Delete retriable, this patch caches the delete result in memory after the shim successfully deletes the task. If a later shutdown attempt fails, a subsequent Delete retry can complete cleanup and return the cached exit information. If containerd restarts after deleting the task, the in-memory cache is lost, and containerd will publish TaskExit and TaskDelete events again to subscribers. Currently, the exit status returned by shim-delete is always 137. In a follow-up, containerd-shim-runc-v2 should store the actual exit status in the state directory, so that shim-delete can reuse it instead of returning a hard-coded value. Implements: #8981 Signed-off-by: Wei Fu <fuweid89@gmail.com>
This is the second patch release in the 1.5.z release series of runc, which primarily includes a workaround for a Linux kernel bug causing random runc crashes when using cgroup v2, and other fixes. Fixed - `runc exec -p` with a `process.json` lacking env now sets HOME again (a regression in runc 1.3.0). - Worked around a Linux kernel bug (present since kernel v6.17, fixed in v7.2) which caused the kernel to write past the end of the structure provided by userspace (runc). This resulted in memory corruption inside runc (manifesting as random crashes) when configuring device rules on cgroup v2 systems. - `runc exec --cgroup` (and the equivalent libcontainer `Process.SubCgroupPaths` API) no longer accepts a sub-cgroup path that escapes the container's cgroup into a sibling cgroup sharing the same name prefix. Note that using `--cgroup` requires the same privileges as running runc exec itself, so this is a correctness rather than a security fix. - Fixed a missing `O_CLOEXEC` when opening the cgroup v2 directory to set up device rules. - Some long-standing file-descriptor leaks on the eBPF devices cgroups were fixed. - When rootfsPropagation is set to `rslave`, the rootfs parent mount is no longer made private before pivoting into the rootfs, so unmount/remount events on host mountpoints under the rootfs are now propagated to the running container. - runc no longer misdetects a non-initial user namespace as the initial one when that namespace has a full identity ID mapping (0 0 4294967295), as used by systemd >= 260 units with `PrivateUsers=full`. Previously this made runc skip its user namespace code paths, so starting a container in such a unit failed with `bpf_prog_query(BPF_CGROUP_DEVICE) failed: operation not permitted`. - Fixed a runc init panic (`SIGABRT`) on the error path, caused by SELinux labels being reset after the cached libpathrs procfs handle was already closed. This is fixed both by not resetting the labels on the init error path, and by updating to libpathrs v0.2.6, which now handles a closed procfs handle gracefully. - Fixed various issues when the libseccomp version runc is run with differs from the one it was compiled against (e.g. built with libseccomp >= 2.6.0 and run with an older one), by updating to libseccomp-golang v0.12.0. This also supersedes the `SECCOMP_FILTER_FLAG_WAIT_KILLABLE_RECV` workaround added in runc 1.5.1. - The libseccomp library statically linked into release binaries is now built with optimizations enabled (the default `-g -O2 CFLAGS`); previously it was built unoptimized. Changed - Switched to opencontainers/cgroups v0.1.0, which no longer uses the high-level cilium/ebpf API to manage cgroup v2 device rules. As a result, the runc binary shrunk by about 1 MiB (7.5%) on amd64. This also means runc no longer calls the cilium/ebpf code affected by GO-2026-6238. - Updated golang.org/x/net to v0.55.0. - Updated builds to libseccomp v2.6.1. release-notes: https://github.com/opencontainers/runc/releases/tag/v1.5.2 full diff: https://github.com/opencontainers/runc/releases/compare/v1.5.1...v1.5.2 Signed-off-by: Sebastiaan van Stijn <github@gone.nl>
Private forks of containerd are used for validating security patches ahead of public disclosure. Some jobs should not run in that context. Guard pull_request and push jobs with !github.event.repository.private so they continue to run on containerd/containerd and public forks. Because schedule events do not populate github.event.repository, guard scheduled runs in nightly.yml and scorecards.yml with github.repository == 'containerd/containerd'. Assisted-by: Antigravity Signed-off-by: Samuel Karp <samuelkarp@google.com>
Skip the linters steps when running in private repositories while allowing the job itself to succeed so downstream jobs with `needs: [..., linters, ...]` still run. Exclude macos-latest and windows-latest from the linters matrix on private repositories to avoid starting unnecessary runners. Assisted-by: Antigravity Signed-off-by: Samuel Karp <samuelkarp@google.com>
Commit 1921680 removed the separate integration/client Go module and merged its dependencies into the root module and vendor directory. Running `go mod download` before `go test` in integration/client is unnecessary because `go test` uses the root vendor directory. Assisted-by: Antigravity Signed-off-by: Samuel Karp <samuelkarp@google.com>
golangci-lint: enable perfsprint linter
Bumps the docker-actions group with 1 update: [docker/setup-buildx-action](https://github.com/docker/setup-buildx-action). Updates `docker/setup-buildx-action` from 4.3.0 to 4.4.1 - [Release notes](https://github.com/docker/setup-buildx-action/releases) - [Commits](docker/setup-buildx-action@37fe631...f87e599) --- updated-dependencies: - dependency-name: docker/setup-buildx-action dependency-version: 4.4.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: docker-actions ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps the k8s group with 2 updates: [k8s.io/cri-api](https://github.com/kubernetes/cri-api) and [k8s.io/streaming](https://github.com/kubernetes/streaming). Updates `k8s.io/cri-api` from 0.37.0 to 0.37.1 - [Commits](kubernetes/cri-api@v0.37.0...v0.37.1) Updates `k8s.io/streaming` from 0.37.0 to 0.37.1 - [Commits](kubernetes/streaming@v0.37.0...v0.37.1) --- updated-dependencies: - dependency-name: k8s.io/cri-api dependency-version: 0.37.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: k8s - dependency-name: k8s.io/streaming dependency-version: 0.37.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: k8s ... Signed-off-by: dependabot[bot] <support@github.com>
Bumps [google.golang.org/grpc](https://github.com/grpc/grpc-go) from 1.83.2 to 1.84.0. - [Release notes](https://github.com/grpc/grpc-go/releases) - [Commits](grpc/grpc-go@v1.83.2...v1.84.0) --- updated-dependencies: - dependency-name: google.golang.org/grpc dependency-version: 1.84.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com>
Conditional actions
update runc binary to v1.5.2
The default profile allowed socket(2) through open-ended range checks around AF_ALG and AF_VSOCK. Any address family added to Linux later was allowed automatically. List the Linux address families the default profile supports and build the socket rules from that list. AF_ALG and AF_VSOCK stay blocked, and unknown or future domains now get the default EPERM response. A single rule cannot bound both ends of a range, because runc treats repeated comparisons on one argument as OR. The consecutive run starting at AF_UNIX therefore becomes one less-than rule, which is safe because AF_UNSPEC cannot be created. Every other domain gets its own equality rule. This matches the same change in Moby's default seccomp profile. Signed-off-by: Paweł Gronowski <pawel.gronowski@docker.com>
…e.golang.org/grpc-1.84.0 build(deps): bump google.golang.org/grpc from 1.83.2 to 1.84.0
…ocker-actions-b530e7f9d1 build(deps): bump docker/setup-buildx-action from 4.3.0 to 4.4.1 in the docker-actions group
…09a0ec50e build(deps): bump the k8s group with 2 updates
seccomp: Use an allow-list for socket domains
…alls A proxy snapshotter is a separate process, and none of its 10 gRPC methods carry a deadline unless the caller already set one. Snapshot GC calls Walk and Remove under the snapshotter-level lock, so one wedged call blocks every later mutating call on that snapshotter. Add default_timeout to the proxy_plugins config section. When set, each proxy snapshotter call gets that deadline if the incoming context has none; a context that already carries its own deadline is left alone. Empty or unset keeps today's behavior, so nothing changes for existing configs. NewSnapshotter keeps its signature and wraps the new NewSnapshotterWithOpts, so existing callers of the exported constructor are unaffected. Assisted-by: Claude Code Signed-off-by: Nahum Litvin <nahuml@wix.com>
The AF_ALG address family exposes the Linux kernel crypto API to userspace via sockets. This has been a source of container escape vulnerabilities (see https://copy.fail/). Unlike seccomp, which can only filter arguments of the direct socket(2) syscall, AppArmor hooks into the kernel's security_socket_create() LSM callback, which fires regardless of the syscall entry point. This means AppArmor also blocks AF_ALG sockets created via the legacy socketcall(2) multiplexer (used by 32-bit binaries), which seccomp cannot inspect because the address family is behind a userspace pointer that BPF cannot dereference. The "deny network alg," rule is placed right after the blanket "network," allow rule so the deny takes precedence for this specific address family. This mirrors moby/profiles commit 946fe8e08d9debcafe76ee60d69261a0c0990368 ("apparmor: Deny AF_ALG sockets in default container profile"). Signed-off-by: Paweł Gronowski <pawel.gronowski@docker.com>
The default seccomp profile excludes AF_VSOCK for direct socket calls,
but 32-bit x86 binaries can reach it through socketcall because
seccomp cannot inspect the pointer argument.
Deny the family in AppArmor so the socket_create LSM hook closes that
path and prevents default-profile containers from reaching host or
guest VSOCK services.
This mirrors moby/profiles commit b7711bef4ad769a13fd40fe1ba710ad0f796fcfd
("apparmor: Deny AF_VSOCK sockets in default container profile").
Signed-off-by: Paweł Gronowski <pawel.gronowski@docker.com>
…timeout snapshots/proxy: add optional default_timeout for proxy snapshotter calls
contrib/apparmor: Deny AF_ALG and AF_VSOCK sockets in default container profile
Signed-off-by: Samuel Karp <samuelkarp@google.com>
Assisted-by: Claude Signed-off-by: syed abdul khaliq <abdul@bugqore.com>
…otdot-path reference: reject `..` path components that rewrite the host
releases: 1.7 EOL, update the rest to current
Signed-off-by: Maksym Pavlenko <pavlenko.maksym@gmail.com>
Signed-off-by: Akihiro Suda <akihiro.suda.cz@hco.ntt.co.jp>
Corresponds to the minimum version required by Kubernetes 1.38 https://github.com/kubernetes/cri-api/blob/v0.38.0-alpha.1/go.mod#L5 Signed-off-by: Akihiro Suda <akihiro.suda.cz@hco.ntt.co.jp>
Correct erofs docs
go.mod: bump minimum Go version to 1.27.0
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to subscribe to this conversation on GitHub.
Already have an account?
Sign in.
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
See Commits and Changes for more details.
Created by
pull[bot] (v2.0.0-alpha.4)
Can you help keep this open source service alive? 💖 Please sponsor : )