You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
feat: design and implement administrator-controlled CDI device admission #4327
As a gateway administrator, I want to define which CDI devices sandbox workloads may receive, so that I can expose selected accelerators without granting every device available to the container runtime.
Problem Statement
OpenShell has no administrator-controlled admission policy for CDI device selectors. GPU attachments are temporarily exempt from external-resource label admission. allow_driver_config controls whether callers may supply nonempty driver JSON, but does not express which CDI names are permitted; public GPU-count requests also resolve device names automatically.
CDI qualified-name validation rejects malformed selectors and host paths. It does not establish that a valid selector is authorized, belongs to a GPU class, or represents one physical GPU. CDI specifications can apply device nodes, mounts, environment variables, and hooks, so the administrator needs to approve the runtime-host-managed device definition as well as its name. This is a feature/design proposal, not a report of a verified vulnerability.
Impact / Why This Matters
Today, operators must disable caller driver config, curate the runtime host's CDI specifications, or use separate gateways/runtime hosts for different device access requirements. Disabling caller config also removes unrelated legitimate overrides and does not define admission for automatically selected GPUs. Host-wide CDI curation affects other runtime consumers and cannot express OpenShell-specific restrictions when multiple device sets coexist.
For example, an operator may want a gateway to select nvidia.com/gpu=0 while reserving GPU 1 for another workload. A policy should govern both an explicit selector and a public GPU-count request, and an aggregate selector such as nvidia.com/gpu=all must not accidentally grant the excluded device. Similar needs apply to other CDI vendors and classes without implying that OpenShell already supports their execution.
Proposed Design
Required operator workflow and behavior:
An administrator configures a CDI admission policy in gateway/driver configuration. Sandbox callers and saved templates cannot weaken it. Keep the existing allow_driver_config gate independent.
Docker and Podman enforce equivalent rules for the same effective policy. Apply admission to caller-provided selectors and automatically resolved selections, including direct driver calls and subsequent launches/restarts; enforce the final device set before runtime attachment.
An administrator can permit exact qualified CDI names. Consider optional vendor/class grants for broader access. If block rules are supported, explicit denial takes precedence over matching allow rules.
A denied request fails with an actionable error identifying the selector and policy reason. Automatic selection chooses permitted candidates or reports insufficient permitted devices; it does not silently attach a denied device.
The policy refers to the runtime host's inventory and trusted CDI specifications, including remote runtimes. Admission by name does not validate every OCI edit or freeze specification contents. Document this trust boundary and what happens when inventory/specifications change.
Design decisions to resolve before implementation:
Configuration location/schema, exact-name versus vendor/class rule syntax, optional block rules, and distinction between absent and empty policy. Decide compatibility defaults and migration behavior explicitly rather than silently changing existing GPU access.
Whether the initial policy governs the existing GPU attachment workflow only or establishes a broader CDI-device contract; how device classes qualify for GPU selection.
Aggregate and alias semantics: indices, UUIDs, MIG selectors, and all may describe overlapping access. Individual grants must not implicitly authorize aggregates containing excluded devices. Decide whether admission is strictly name-based, uses trusted alias/aggregate membership, or rejects selectors whose scope cannot be established under a restrictive policy. Do not infer physical-device identity or count from arbitrary CDI strings.
Older runtimes, unavailable versus known-empty inventory, metadata lookup failures, and remote inventory. Name-only policies may not require discovery, but any rule relying on membership/identity must define behavior when that information is unavailable. Legacy path discovery must not become an admission bypass.
Policy acknowledgement for external drivers, policy changes, existing sandboxes, and runtime-spec changes: document enforcement timing and whether running-workload revocation/revalidation is included initially or deferred.
Keep CDI syntax parsing and Podman inventory discovery as separate work; this issue consumes those capabilities rather than redesigning them.
Acceptance Criteria
Maintainers resolve the design decisions above, including defaults/migration, aggregate/alias semantics, unsupported inventory behavior, and existing/running workload handling.
Administrators can configure exact CDI-name admission consistently for Docker and Podman; callers cannot override it and allow_driver_config remains independent.
Explicit and automatically selected devices are admitted before attachment on supported launch paths, including direct driver requests and restart/recreation.
Denied requests produce actionable errors; automatic selection uses only admitted candidates and reports insufficient admitted capacity.
Gateway configuration and driver documentation explain defaults, rollout, CDI-spec trust, errors, and enforcement timing; related admission acknowledgement/configuration surfaces are updated where needed.
Alternatives Considered
NVIDIA-only syntax validation: restricts vendor compatibility without expressing operator authorization.
Only gate driver JSON: too coarse and misses automatic device selection.
Only curate CDI specifications on the runtime host: useful defense and still required trust, but lacks per-gateway admission and affects other consumers.
Blocklist alone: new devices or aliases can become accessible without an affirmative operator grant. Exact allowlists provide a clearer starting point; broader grants and exceptions require deliberate semantics.
Docker and Podman each have explicit/default CDI selection paths. Shared GPU inventory currently normalizes NVIDIA names and distinguishes indexed/named selectors from the all fallback; it is not a generic identity or authorization model.
User Story
As a gateway administrator, I want to define which CDI devices sandbox workloads may receive, so that I can expose selected accelerators without granting every device available to the container runtime.
Problem Statement
OpenShell has no administrator-controlled admission policy for CDI device selectors. GPU attachments are temporarily exempt from external-resource label admission.
allow_driver_configcontrols whether callers may supply nonempty driver JSON, but does not express which CDI names are permitted; public GPU-count requests also resolve device names automatically.CDI qualified-name validation rejects malformed selectors and host paths. It does not establish that a valid selector is authorized, belongs to a GPU class, or represents one physical GPU. CDI specifications can apply device nodes, mounts, environment variables, and hooks, so the administrator needs to approve the runtime-host-managed device definition as well as its name. This is a feature/design proposal, not a report of a verified vulnerability.
Impact / Why This Matters
Today, operators must disable caller driver config, curate the runtime host's CDI specifications, or use separate gateways/runtime hosts for different device access requirements. Disabling caller config also removes unrelated legitimate overrides and does not define admission for automatically selected GPUs. Host-wide CDI curation affects other runtime consumers and cannot express OpenShell-specific restrictions when multiple device sets coexist.
For example, an operator may want a gateway to select
nvidia.com/gpu=0while reserving GPU 1 for another workload. A policy should govern both an explicit selector and a public GPU-count request, and an aggregate selector such asnvidia.com/gpu=allmust not accidentally grant the excluded device. Similar needs apply to other CDI vendors and classes without implying that OpenShell already supports their execution.Proposed Design
Required operator workflow and behavior:
allow_driver_configgate independent.Design decisions to resolve before implementation:
allmay describe overlapping access. Individual grants must not implicitly authorize aggregates containing excluded devices. Decide whether admission is strictly name-based, uses trusted alias/aggregate membership, or rejects selectors whose scope cannot be established under a restrictive policy. Do not infer physical-device identity or count from arbitrary CDI strings.Keep CDI syntax parsing and Podman inventory discovery as separate work; this issue consumes those capabilities rather than redesigning them.
Acceptance Criteria
allow_driver_configremains independent.Alternatives Considered
Agent Investigation
allfallback; it is not a generic identity or authorization model.Checklist