Skip to content

SRM HTTPG keep-alive interoperability issue after Java upgrade: chunked response causes gfal2/srm-ifce SOAP error 6 #8235

Description

@telecast-git

Hello,

we have investigated an SRM failure and reduced it to a small reproducible case.

Tested environment:

dCache: 11.2.7
Java: OpenJDK 21

client:
gfal2 2.23.5
srm-ifce 1.24.8
gSOAP 2.8

The client reports:

Communication error on send
Unknown SOAP error (6)

In some cases this is followed by:

DESTINATION MAKE_PARENT
srmMkdir
SRM_AUTHORIZATION_FAILURE
Permission denied

The srmMkdir error initially looked like a namespace permission problem, but it appears to be secondary.

The target directory already exists and dCache reports the preceding srmLs operation as SRM_SUCCESS.

The problem can be reproduced with two sequential SRM directory operations.

With GFAL SRM keep-alive enabled:

[SRM PLUGIN]
KEEP_ALIVE=true

the first response is valid:

HTTP/1.1 200 OK
Transfer-Encoding: chunked

<complete srmLs SOAP response>

However, the client sends the next request before receiving the terminating HTTP chunk:

0

The sequence is therefore effectively:

SOAP response
next POST
0

and the client fails with:

Communication error on send
Unknown SOAP error (6)

With:

KEEP_ALIVE=false

the same test succeeds.

The response uses:

Connection: close

and repeated srmLs operations complete normally.

For comparison, on an older dCache setup where keep-alive works, the sequence is:

SOAP response
0
next POST

Workaround 1: disable connection reuse

The client-side workaround is:

[SRM PLUGIN]
KEEP_ALIVE=false

Since the client configuration is not always under site control, we also implemented a small server-side compatibility plugin.

The plugin does not modify the dCache package. It imports the normal SRM Spring configuration and adds a BeanPostProcessor which changes the SRM GSI/Jetty connector so the HTTP connection is not reused.

Conceptually it provides the server-side equivalent of:

KEEP_ALIVE=false

An A/B test was performed:

plugin enabled:
100 stat operations
100 directory listing operations
FAILURES=0

plugin disabled:
Unknown SOAP error (6)

This strongly indicates that the failure is related to HTTP persistent connection handling rather than namespace permissions.

The plugin is considered a temporary workaround, not a permanent fix.

Workaround 2: use a dedicated subdirectory in the SURL

We also found that changing the target from a base directory such as:

.../data/ops

to a dedicated subdirectory:

.../data/ops/<monitoring-directory>

can avoid the subsequent srmMkdir failure.

The important point is that the original sequence looks like:

srmLs existing-directory
-> client-side communication failure
-> MAKE_PARENT
-> srmMkdir existing-directory

The parent of the existing directory may not be writable by the client, so trying to create that directory correctly returns SRM_AUTHORIZATION_FAILURE.

Using a subdirectory moves the srmMkdir operation below a writable directory.

We therefore consider this a workaround only. It does not explain why the client enters MAKE_PARENT after dCache has already completed srmLs successfully.

It may be entirely caused by the preceding HTTP communication failure, but it could also expose a separate SRM behaviour difference.

Version observation

The problem first appeared after changing Java while keeping the dCache version unchanged:

dCache 10.0.31 + Java 11:
works

same dCache 10.0.31 + Java 17:
fails

The problem remains with newer dCache releases and Java 21.

This makes a database or namespace regression unlikely. The behaviour appears more closely related to the HTTPG/Jetty/Java transport and its interaction with CGSI/gSOAP.

Could you please advise:

  1. Is this a known interoperability issue between dCache/Jetty/Java and CGSI-gSOAP/srm-ifce?
  2. Is it expected that the terminating chunk can reach the client only after it has started sending the next request?
  3. Is there a supported dCache option to disable persistent HTTP connections for the SRM door?
  4. If not, is forcing connection close at the SRM connector a reasonable temporary workaround?
  5. Is this fixed in another dCache release or patch?
  6. Could the MAKE_PARENT / srmMkdir behaviour be fully explained by the failed srmLs exchange, or could there also be an SRM regression around existing destination directories?

We can provide CGSI traces for:

working keep-alive case
failing keep-alive case
working KEEP_ALIVE=false case

as well as the corresponding dCache access logs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions