Skip to content

Cancelled stage requests when performing operations on the Poolmanager #8223

Description

@elenamplanas

Dear all,

We've been seeing drops in stage requests for a while now whenever we perform operations on the PoolManager. This happens whether we run a reload -yes on the PoolManager or set a pool to read-only using psu set pool rdonly.

Image

In the attached graph, we can see a sudden drop in restores, which coincides with an operation we performed on the PoolManager.

In the PoolManager logs there are no errors.
In the billing file for each stage cancelled there are messages like following:

09.16 09:21:48 [door:WebDAV-LHCbT1-door06@webdav-lhcb-https-door06Domain:request] ["/DC=ch/DC=cern/OU=Organic Units/OU=Users/CN=lbprods/CN=693025/CN=Robot: LHCb offline productions":31151:1470:2001:67c:1148:301:0:0:0:834] [00002024DA3F0A0F448191BA91462DAF831E,5363003867] [/pnfs/pic.es/data/lhcb/LHCb-Tape/lhcb/data/2026/RAW/TURCAL/LHCb/COLLISION26/345486/345486_00130000_0481.raw] vo-lhcb.lhcb@cta 180007 0 {10006:"Request to [>PoolManager@poolmanager-dccore06Domain] timed out."}
09.16 09:46:24 [pool:dc090_2@dc090_2Domain:remove] [00002024DA3F0A0F448191BA91462DAF831E,0] [Unknown] vo-lhcb.lhcb@cta {0:"Transfer failed and replica is FROM_STORE"}
09.16 09:18:48 [pool:dc090_2@dc090_2Domain:restore] [00002024DA3F0A0F448191BA91462DAF831E,5363003867] [Unknown] vo-lhcb.lhcb@cta 1655680 0 {10006:"Stage was cancelled.

In the pool similar messges appear:

16 Sep 2026 09:46:24 (dc090_2) [door:WebDAV-LHCbT1-door06@webdav-lhcb-https-door06Domain:AAZblHp6hGg PoolManager] Stage of 00002024DA3F0A0F448191BA91462DAF831E failed with CacheException(rc=10006;msg=Stage was cancelled.).
16 Sep 2026 09:46:24 (dc090_2) [door:WebDAV-LHCbT1-door06@webdav-lhcb-https-door06Domain:AAZblHp6hGg PoolManager] Failed to read file size: java.nio.file.NoSuchFileException: /dcpool2/vpool2/data/00002024DA3F0A0F448191BA91462DAF831E
16 Sep 2026 09:46:24 (dc090_2) [door:WebDAV-LHCbT1-door06@webdav-lhcb-https-door06Domain:AAZblHp6hGg PoolManager] Failed to read file size: java.nio.file.NoSuchFileException: /dcpool2/vpool2/data/00002024DA3F0A0F448191BA91462DAF831E
16 Sep 2026 09:46:24 (dc090_2) [door:WebDAV-LHCbT1-door06@webdav-lhcb-https-door06Domain:AAZblHp6hGg PoolManager] Failed to read file size: java.nio.file.NoSuchFileException: /dcpool2/vpool2/data/00002024DA3F0A0F448191BA91462DAF831E

We don't know why these requests are being canceled, but this behavior recurs every time we interact with pools in the PoolManager.
We can reproduce the issue whenever you'd like, and gather more information by increasing the log level wherever you specify.

If the stage was requested through a bulk operation, the system requests it again because the request remains pending in the PinManager. Otherwise, the client receives an error and no retry occurs.

Thank's in advance

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions