Skip to content

service: cgroup.kill on instance teardown breaks restart of the servi… - #45

Open
Aukansh wants to merge 1 commit into
openwrt:mainfrom
Aukansh:main
Open

Aukansh wants to merge 1 commit into
openwrt:mainfrom
Aukansh:main

Conversation

@Aukansh

@Aukansh Aukansh commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Since commit f4d512d93a00a701c9604376e759ec60960d4dfb ("jail: place container init and exec into the target cgroup via clone3"), restarting a service that provides the ubus service object fails, and the service disappears from procd's service tree. On OpenWrt this is easily observable on rpcd.

instance_remove_cgroup() is called from instance_delete() while procd
is still processing a "service delete" ubus call issued by the service's
own init script (e.g. "/etc/init.d/rpcd restart" -> "stop" -> ubus call
service delete). Writing "1" to cgroup.kill at that point kills every
task in the cgroup, including the shell running the init script right
now. The shell dies before it can proceed to the "start" phase, the
service is left torn down, its name disappears from the services tree,
and a subsequent restart fails with:

Command failed: ubus call service delete { "name": "rpcd" } (Not found)

Fix this by simply attempting to remove the cgroup and ignoring EBUSY:
if the cgroup still has tasks, it is left in place and reused by the
next instance_start(); the kernel retires it once its last task exits.

…ce that owns the ubus "service" object

Signed-off-by: Au Ychen <au-ychen@foxmail.com>
@Aukansh
Aukansh marked this pull request as ready for review September 17, 2026 12:02
@dqsq2e2

dqsq2e2 commented Sep 21, 2026

Copy link
Copy Markdown

I reproduced this on three OpenWrt-based systems using the tasks service used by iStore.

On the system with procd 2026.03.13~58eb263d, a task running true reports:

exit_code: 0
 data.exit_code: "0"

On two systems with procd 2026.09.19~c230ad87, the same task reports:

exit_code: 137
 data.exit_code: "0"

The task itself completed successfully. The same result is reproducible without iStore or APK: start a sleep 60 service instance, update only its data field, and the cgroup-enabled procd terminates the live instance with 137. The older procd keeps the same PID running.

This also explains the original user-facing symptom: iStore install and removal completed successfully, but the LuCI task status used the outer procd exit code and displayed failure. The proposed change fixes the lifecycle ownership issue at the source.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants