Skip to content

Add/update OAuth 2.0 flow documentation in IAP - #10

Merged
mbtaylor merged 14 commits into
mainfrom
bearer-tokens
Sep 11, 2026
Merged

mbtaylor merged 14 commits into
mainfrom
bearer-tokens

Conversation

@jesusjuansalgado

Copy link
Copy Markdown
Contributor

Better support of oAuth and examples for:
OAuth Device Code Flow (RFC 8628) (for command line)
OAuth Authorization Code Flow (RFC 8252) (for rich desktop applications)

I have not added the token/token yet as we still need to discuss if we want to promote it further. I think we probably should, as some VO apps already use this.

@mbtaylor

Copy link
Copy Markdown
Member

There are some details here that I don't fully follow, for instance:

  • X-VO-Auth-Error header (not a scheme parameter) - what are clients supposed to do with this value? Also it's only applicable to the ivoa_bearer scheme but there may be multiple challenges with different schemes. And it's optional which reduces its usefulness.

  • The ivoa://ivoa.net/sso#OAuth standard_id value - is this expected to be used with one of the schemes that defines the standard_id parameter (ivoa_cookie, ivoa_x509)? If so, it's not clear to me how this would work. If the user can get a bearer token it seems likely that they'd use it directly rather than exchange it for a cookie or X.509 certificate. I'd like to see this in action before adding it as an option here.

As far as the details of the exchanges go I'm not qualified to review it, at least before attempting an implementation, so it should also be reviewed by somebody else who understands OAuth2 (@aragilar?).

But in any case I think it's premature to make an OAuth2 update to the document. Before that happens I would like to see at least one prototype interoperating service and client implementation of the proposal. OAuth2 is complicated and there are quite a lot of moving parts here, so I think it would be wise to prove that the suggested scheme is sufficiently well-described to be implemented by services and clients, and that the authentication works in at least one example context, before absorbing it into the draft text of the document.

@jesusjuansalgado

jesusjuansalgado commented Jun 13, 2025 •

Copy link
Copy Markdown
Contributor Author

X-VO-Auth-Error header was introduced to identify the problem related to the authorisation (it could describe that the token is not present but needed, the token is invalid, the token is expired... we can create a vocabulary). I think adding this header could be interesting in general, even for other standard_id (s). I was tempted to add it to the other methods but it is only required for OAuth for the time being.

About the ivoa://ivoa.net/sso#OAuth standard_id , I think OAuth is a different thing from cookies or certificates. It has its own path for development and client-server integration, and this is why it is making use of a different standard_id

About the prototypes, I think the path is usually the opposite. We propose a draft, we try the implementation, and we discover if it is implementable, if there are things fundamentally wrong, or if some feedback needs to be added. I think I can coordinate the implementation at the SRCNet (we will not use either cookies or certificates but bearer tokens), in particular in our astroquery module. We are already implementing RFC 8628, but we just need to add the understanding of the error at our server. The current draft, with cookies and certificates, cannot be used by us (and others)

Yes, we could add @aragilar to review it. I would propose finding engineering experts on OAuth to give feedback. I can identify some engineers at SKAO and other missions to take a look at it, but of course, we need the integration in a draft first (I can generate one from my branch), so they could provide feedback while trying to implement it

Please note that as chair of DSP, I would not be confident to allow the progress of the draft without OAuth and only providing support to cookies or certificates, technologies that are being deprecated. Most of the bigger new observatories need to have more secure systems that use federation, so a standard without this would not be very useful in the short term. We need to include OAuth in one way or another to progress with this standard, and this is why I have made to effort to start the writing of this part

@mbtaylor

Copy link
Copy Markdown
Member

About the ivoa://ivoa.net/sso#OAuth standard_id , I think oAuth is a different thing than cookies or certificates. It has its own path, and this is why it is making use of a different standard_id

Agreed that cookies, certificates and OAuth are quite different. But the use of the standard_id parameter defined in AuthVO is, so far, only for use with the ivoa_cookie and ivoa_x509 schemes. Your description of the ivoa_bearer scheme doesn't use it (which is fine, I don't think it's appropriate there). If ivo://ivoa.net/sso#OAuth is a possible value for standard_id, it means that services can issue a challenge like:

   www-authenticate: ivoa_cookie access_url="https://example.org/login",
                                 standard_id="ivo://ivoa.net/sso#OAuth"

In that case the client has to do a lot of work to authenticate using OAuth, as you say "not directly specified by this identifier" to pick up a cookie or certificate for later use. Is this what you envisage? Unless we actually expect services to do that, I don't want to put it in the document since it adds a large implementation burden for clients. The way I'd expect clients to use OAuth is instead using something like the ivoa_oauth challenge you propose, which doesn't use the standard_id parameter.

About the prototypes, I think the path is usually the opposite. We propose a draft, we try the implementation, and we discover if it is implementable, if there are things fundamentally wrong, or if some feedback needs to be added. ... Current draft, with cookies and certificates, cannot be used by us

Also agreed that the existing (cookie, x509) text is no good for OAuth services. My reluctance to accept the PR at this stage is that there is a lot of untested detail there and I can imagine significant changes before we get it right. I think that it will be easier to manage such changes if we experiment using more informal descriptions of the requirements than PRs to standards text. This is generally how I've managed specifications that I've been involved with before (e.g. SAMP, VOParquet). If you disagree strongly with this approach maybe it could be discussed in the TCG.

@jesusjuansalgado

jesusjuansalgado commented Jun 13, 2025 •

Copy link
Copy Markdown
Contributor Author

I think there is a misunderstanding (?)
Once the invocation of the service fails (e.g. trying to invoke a data access URL coming from a data link as we do in our SRCNet system,), the system returns in the error the information needed to start oAuth, that implies to understand where is the document with all the oAuth configuration to authenticate this service. This is not a login URL but something you need to parse in the client https://auth.example.org/.well-known/openid-configuration (because it is more complex)
Also, you need to understand the reason of the error to know if there is an invalid token or you need to refresh it, for example.
The rest of the example describes the standard oAuth, so client developers could modify their existing clients to add the logic and incorporate oAuth (although most will just import an oAuth library)

As said, as chair I would not be able to accept a standard without bearer tokens as that would imply to just describe protocols that are getting obsolete and leaving major missions outside of the loop

@mbtaylor

Copy link
Copy Markdown
Member

Yes, I'm sure we're talking at cross purposes, apologies if I'm not expressing myself clearly. I certainly agree that OAuth2 has to be supported.

You are proposing an ivoa_bearer challenge, which tells the client where to find the configuration document, which sounds fine. But the standard_id parameter, which describes how to submit a username-password pair, is only for use in ivoa_cookie and ivoa_x509 challenges, so it doesn't need to have an OAuth-related value, unless you are saying that people should make cookie- or certificate-based auhentication accessible by use of OAuth2. I expect that OAuth2-based authentication doesn't go near cookies or certificates, so that
standard_id="ivo://ivoa.net/sso#OAuth" has no use.

@jesusjuansalgado

jesusjuansalgado commented Jun 13, 2025 •

Copy link
Copy Markdown
Contributor Author

Ah, I understand your point now. However, I think the purpose of the standard is to allow the identification by clients of an error accessing a particular service due to authorisation problems. This client should be able to understand that this is an error produced by a service that expects a basic authentication (and the error provides the URL to do the login), a certificate authentication or oAuth (where the information on how to handle the authorisation is in one configuration file provided to the client). I have aligned the parameters for the oAuth part, but I think what is wrong now is to assume that all the client-server negotiations are just passing a username and a password (as said in the text). I think this part needs to be more ambiguous to cover more complex negotiations

I am trying to cover all with a similar approach, so this is why I added a new standard_id. The interpretation is different but parameters are the same and the behaviour of the server due to errors very aligned

@mbtaylor

Copy link
Copy Markdown
Member

Agreed not all client-server negotiations are passing a username and password; the standard_id parameter is only for those protocols for which username+password is appropriate (and where there are multiple options of how to do that). For ivoa_bearer you don't want a username+password, and the nature of the client-server negotiation is implicit in the challenge type, so I don't think either standard_id or a differently-named parameter with similar meaning is required in this case.

@jesusjuansalgado

Copy link
Copy Markdown
Contributor Author

I will continue with this branch to see if I can align it more with a username/password approach and I will let you now when this is ready

@mbtaylor

Copy link
Copy Markdown
Member

I think that the parameters for the ivoa_bearer scheme need to be adjusted - it doesn't work in the same way as ivoa_cookie and ivoa_x509, so it shouldn't try to use the same parameters, which don't fit it very well.

For ivoa_cookie and ivoa_x509 the login protocol has the same requirements for both of them, i.e. send a username+password, so it makes sense to use the same parameter (stanadard_id) in both cases.

But ivoa_bearer has a different way of authenticating (not sending a username+password), so it doesn't make sense to use the same parameter with the same set of options. If one of the options for standard_id is ivo://ivoa.net/sso#OAuth then the client would have to figure out what to do if confronted with

   www-authenticate: ivoa_cookie access_url="https://example.com/login",
                                 standard_id="ivo://ivoa.net/sso#OAuth"

which I don't think has a sensible interpretation.

Moreover, the only value for standard_id that makes sense in your ivoa_bearer scheme is standard_id="ivoa://ivoa.net/sso#OAuth", so specifying that parameter in that scheme is not doing any work.

Instead ivoa_bearer should come up with its own set of parameters that make sense for what it needs to do, not force usage of the parameters already in use for ivoa_cookie and ivoa_x509.

So you could have something like:

   www-authenticate: ivoa_bearer discovery_url="https://auth.example.org/.well-known/openid-configuration"

Using HTTP response headers like X-VO-Auth-Error and X-VO-Auth-User-Action for additional information about the authentication requirements is also problematic, since the server might issue multiple challenges and it's not clear which one(s) the other headers apply to. If that extra information is necessary (since it's optional I'm not yet convinced that it is) it might be better to supply those as additional challenge parameters, e.g.

   www-authenticate: ivoa_bearer
                     discovery_url="https://auth.example.org/.well-known/openid-configuration",
                     error="missing_token",
                     user_action="hello you have done something wrong"

But I'm not sure about that. What is usual practice in OAuth2 for transmitting information like the fact that the token has expired?

@jesusjuansalgado

jesusjuansalgado commented Jun 17, 2025 •

Copy link
Copy Markdown
Contributor Author

I'm not entirely sure. What the text is proposing is:

www-authenticate: ivoa_bearer access_url="https://auth.example.org/.well-known/openid-configuration",  
                                 standard_id="ivo://ivoa.net/sso#OAuth"

So, whenever the response includes an ivoa_bearer header, the access_url is interpreted as the discovery_url. Since they are conceptually the same (discovery_url provides all the access_url(s) and other metadata for the negotiation), I don’t see any contradictions. In a previous proposal, I used discovery_url, so changing it now would imply that access_url should become optional in the text. That’s why I reused it — to avoid turning everything into optional parameters.

Regarding X-VO-Auth-Error, I understand your point, but I think it could be useful across all methods. We just need to define a controlled vocabulary for possible values. Other methods might use different error statuses, and this could provide a unified way to handle them.

In summary, I think trying to reuse the same parameters across all authentication methods, where possible, could be beneficial. Of course, we could treat them differently if needed (adding as extra metadata in the www-authenticate line)— but in that case, no parameters would be universally required, only those mandated by each standard_id. Nothing would be compulsory in AuthVO

@mbtaylor

Copy link
Copy Markdown
Member

It is normal for different authentication schemes to have different parameters, since they have different requirements. The Basic auth scheme (RFC7617) has the parameters "realm" and "charset", Bearer (RFC6750) has "realm" and "scope", and Digest (RFC7616) has parameters "realm", "charset", "domain", "nonce", "opaque", and a bunch of others. Where the meanings are the same they re-use the same ones, but if the meanings are different they use different parameters.

@mbtaylor

Copy link
Copy Markdown
Member

I still see the issue that @aragilar pointed out in #6 (comment): a client_id is required for Device Authorization Grant, but this text doesn't say how to get one. James's proposal does say how to get a client_id as well as, I think, giving clear instructions for how to do the rest of the authentication, although I can't say for sure without attempting a client-side implementation. Is that proposal unsuitable for SRCNet operations?

@jesusjuansalgado

jesusjuansalgado commented Jun 20, 2025 •

Copy link
Copy Markdown
Contributor Author

the preferred way to do it is by preregistered clients. At the SRCNet we use Indigo IAM, and the steps are done by:

Login as IAM admin
Navigate to the Admin Console
Go to Client Management → "Add Client"
Set:
Client ID: e.g., ivoa-pyvo-cli
Client type: Public
Grant types: enable:
urn:ietf:params:oauth:grant-type:device_code
authorization_code (for desktop)
optionally refresh_token
Scopes: openid, email, profile, vo.read, etc.
Redirect URIs:
For Device Code: not required
For native apps: e.g., http://localhost:*
Enable Client credentials only if the client is trusted to act without user
Save and document the client_id for external clients (e.g., for pyvo)

for keycloak (we could migrate to it at certain point) the steps are similar:
In Keycloak:

Go to Clients → Create
Enter:
client_id = ivoa-pyvo-cli
Enable public client (no client secret)
Enable device_authorization_grant
Set scopes (e.g., openid, profile, vo.read)

So, this is done by the admins and it depends on the IAM system.
There is a dynamic way to get a client_id based on Dynamic Client Registration (RFC 7591) but I do not think this is secure enough and it could be not accepted by some providers.
What we can do? We can maintain a list of VO client_ids and we could update our services. Something like:

client_ids for VO Clients

Client / Tool client_id Notes
TOPCAT / STILTS ivoa-stilts Covers both desktop GUI (TOPCAT) and CLI (STILTS) workflows
pyVO (CLI scripts) ivoa-pyvo-cli Headless Python client for SIA, TAP, and other VO protocols
Astroquery ivoa-astroquery Programmatic VO and archive access via Python
Aladin Desktop ivoa-aladin Interactive VO image/data viewer
vo-cli ivoa-vo-cli General VO command-line interface
VO Notebook (Jupyter) ivoa-vo-notebook For VO-enabled workflows in Jupyter or web notebooks
GAVO DaCHS Portal ivoa-dachs-service Internal DaCHS service/client communication
Web-based VO Portals ivoa-web-portal Browser-based VO tools and data explorers
Test Clients ivoa-test-client Used for sandbox, development, or interoperability testing

But in this standard, we should say that the clients should be preregistered and that the list is maintained by the IVOA (maybe, propose some concrete examples (?))

@mbtaylor

Copy link
Copy Markdown
Member

I think you are proposing here a federated IVOA-wide list of client_ids accepted by all OAuth2-based providers in the VO that want to be accessible in this way. From a client point of view that would make things quite straightforward, though it could make it harder for new clients to use the system, unless there's a general-purpose ID like "ivoa-client" suitable for anybody to use.

Do you (or other readers) think that VO resource providers would accept client_id management done like this? My feeling is that RFC7591 dynamic client registration would be more palatable to resource providers, but I don't have server-side experience, so I might be wrong.

@jesusjuansalgado

jesusjuansalgado commented Jun 20, 2025 •

Copy link
Copy Markdown
Contributor Author

I think the resource providers will accept the preregistered approach, but they will be reluctant to a dynamic client_id creation based on RFC7591. RFC 7591 is not a good fit, in my view, for the IVOA.
In the IVOA, different services and identity providers are managed by different organizations, so allowing clients to register without approval creates trust and security problems. It becomes hard to tell which clients are legitimate, and there’s a higher risk of abuse or misconfiguration.
Most IVOA clients—like data tools, scripts, and web apps—are known ahead of time and don’t change often, so they can be registered manually and securely, and it would be easy to coordinate

@mbtaylor

Copy link
Copy Markdown
Member

OK, let's ask around and see what the opinions are of OAuth2 VO service providers.

@jesusjuansalgado

Copy link
Copy Markdown
Contributor Author

BTW, I have updated the text to come back to discovery_url, set the errors as error and error_message in the same WWW-Authenticate: definition (removing other extra keywords) and a short sentence on the client_id that could be extended whenever we have an agreement with providers. BTW, now it is inconsistent the part of "Common Challenge Parameters for VO Schemes" as access_url is not present for oAuth. That could be fixed later whenever we have a common view

@andamian

Copy link
Copy Markdown
Contributor

My 2 cents:

  1. I don't know why we talk about OAuth (Authorization) when we actually want OpenID Connect (authentication). OAuth2 maybe? https://www.okta.com/identity-101/whats-the-difference-between-oauth-openid-connect-and-saml/. For OIDC we would only need the issuer URL since .well-known/openid-configuration is part of the standard (or just .well-known as suggested in Support to bearer tokens #6 (comment)
  2. There are already mechanisms for dealing with authentication errors that the generic tools already understand so I don't see the benefit of X-VO-Auth-Error. It won't be useful for certs, which fail during SSL handshake. Well-behave clients could also check the token for expiry and allowed_origins before sending the token.
  3. I'm not sure what the purpose of the client_id list is. I'm not an OIDC specialist, but clients (dynamic or pre-approved) are created to assign different security policies and limit impersonations (through matching of the redirect URLs in requests). Secure clients are also configured with a client secret but that only works with service clients due to the security nature of the secret. Unless we can provide a reason for those clients existence, I don't think they should be mentioned at all.

@jesusjuansalgado

jesusjuansalgado commented Jun 21, 2025 •

Copy link
Copy Markdown
Contributor Author

Hi Adrian,
1- I understand we are describing authorisation (not only authentication). In all the examples, we are trying to, e.g. access a particular data product, not only authenticating. We could mention oAuth2 instead of oAuth. Anyway, we are making reference to specific RFCs in all the text so there are not ambiguities. About the link to the discovery url, giving the issuer URL could be enough too and let the client to derive the openid configuration URL but if the final reason to provide the issuer URL is to deduce where the openid configuration is located (assuming that the file is going to be in the expected place) is adding a complexity in the client workflow (already complex) that, in my view, is unnecessary. There are at least two files that could need to be discovered (/.well-known/openid-configuration and /.well-known/oauth-authorization-server ) if we do not mention anything but the issuer url. I think we could do it better describing the expected configuration file location and the metadata inside (what creates a contract for client and server providers).
I even described the expected metadata inside the configuration file because many VO client developers would need to adapt them to be able to work with the new clients so, although it is complex, the more simplified description we could do, the better.

2- X-VO-Auth-Error was already removed and errors were integrated in the header (see my comment just before yours and the new text). This is not a problem for the bearer token description but I still think that this could have been a good add for all the methods. oAuth2 errors have a vocabulary defined but this is not true in general for other methods so I think it is a missing opportunity to standardise this at VO level. Yes, I know that there could be more than one challenge but we could have more than one X-VO-Auth-Error in the response describing the reason of the error so the clients could take decisions on the next step

3- Well, client_ids is part of the standard so the original authors of oAuth2 considered it necessary for security reasons (maybe we do not understand it but they did). I think we could use a default client_id and remove it from the text (this is what we are doing in the first version implemented at the SRCNet). However, reading the documentation:

  • In public flows like Device Code or Authorization Code with PKCE, the client_id is still required (these workflows are used by some clients)
  • client_ids (e.g., ivoa-stilts, ivoa-pyvo-cli) allow IdPs to apply security profiles per tool, like:
  • Acceptable scopes
  • Rate limits
  • Federation rules
  • ...
    We could investigate more to understand the reasons but, as said, if this is part of the standard there should be a reason. We are not inventing this.

@aragilar

Copy link
Copy Markdown

RFC 7591 I think is required for the VO to effectively use OAuth2/OIDC. There is no point relying on the client_id remaining secret, and all it does is encourage everyone to reuse an existing client_id they found on the internet (this is a well-known thing that all OAuth2/OIDC providers deal with). Let's rather save everyone that pain and do things right (and it's far better that providers assume they have to deal with hostile clients and take preventative measures than assuming that a client_id really does mean a specific client is being used).

For the VO, I think clients shouldn't care whether it's OAuth 2.0, OAuth 2.1, or OIDC, because clients shouldn't need any of the profile details to function (you don't need the user's full name or institute or email). The authorization and resource server can communicate however they wish (and OIDC makes sense to standardise on from a VO service provider's side), so I haven't really worried about the distinction between the versions.

@jesusjuansalgado

jesusjuansalgado commented Jun 21, 2025 •

Copy link
Copy Markdown
Contributor Author

Hi @aragilar, I completely agree with your first point. Dynamic registration is a more secure (if correctly implemented) and scalable option, but it may introduce too much complexity for the current VO ecosystem. Pre-registering public client_ids can still be useful, for example, to scope or restrict their capabilities — e.g., ensuring that stilts can't request unrelated scopes like email (so, at policy level). Of course, authorisation servers must assume clients could be hostile (not just in the VO, but in OAuth in general). So yes, pre-registered public clients offer limited protection and shouldn’t be treated as trusted but they can still help to limit the damage of impersonation.

Dynamic registration (RFC 7591) would provide better isolation and tracking, but implementing it correctly, particularly with access policies and registration control, could be difficult for many VO service providers. They are loosing some control of the clients registered on-the-fly so, I think, a bad implementation at server side facilitates attacks (from my partial understanding of the problem). Anyway, maybe this is something we could revisit later after talking to server providers. I do not have a strong position on this at all.

On your second point, I fully agree. I would avoid locking the standard to any specific OAuth version. The goal is to describe a discovery-based mechanism that lets clients adapt to the server’s authentication capabilities. The more generic and future-proof the top-level description is, the better. We should treat the examples (e.g., device code, authorisation code, etc.) as extensible, not exhaustive, ideally, avoiding full rewrites of the spec when new flows emerge in the IVOA ecosystem.

@aragilar

Copy link
Copy Markdown

@jesusjuansalgado There is nothing stopping groups from using OAuth2 currently (the ESO archive uses it, and I've written a Python wrapper around it https://dev.aao.org.au/adacs/eso-downloader; Rubin uses it; Data Central uses it both internally and as an IdP for MWA, SkyMapper and CASDA, though we don't advertise for VO client use, primarily because we don't have RFC 7591 configured as needed for the VO yet; and there are probably other I don't know about), and unlike the cookies/client certs (where there is not a standard, hence AuthVO), no-one needs a VO rec to implement it with pre-agreed client IDs, this has already happened, and I see that if we keep doing that we're going to end up with silos and users unhappy they cannot use their preferred tool against a specific VO service.

@aragilar

Copy link
Copy Markdown

On OAuth 2 (and access token) vs OIDC (and ID token), please see https://auth0.com/blog/id-token-access-token-what-is-the-difference/

From the server-side it makes sense to use OIDC (adding in additional claims to include missing information like ORCID), but from the client-side in the VO, as the client does not actually care about the ID token, the difference between OAuth 2 and OIDC is moot. Both the discovery and dynamic registration endpoints are the same, and so it does not make sense to try to distinguish between the two within the VO.

@aragilar

Copy link
Copy Markdown

On error handling, please read https://datatracker.ietf.org/doc/html/rfc6749#section-4.1.2.1 carefully, we really do not want to add more things to the issues listed on https://pilcrowonpaper.com/blog/dear-oauth-providers/

@jesusjuansalgado

Copy link
Copy Markdown
Contributor Author

On error handling, please read https://datatracker.ietf.org/doc/html/rfc6749#section-4.1.2.1 carefully, we really do not want to add more things to the issues listed on https://pilcrowonpaper.com/blog/dear-oauth-providers/

The error listed in the doc are

  • invalid_request: Request is missing required parameters or malformed.
  • unauthorized_client: The client is not allowed to use this grant type.
  • access_denied: The resource owner denied the request.
  • unsupported_response_type: Response type is not supported by the server.
  • invalid_scope: The requested scope is invalid or unknown.
  • invalid_token: The token is expired, malformed, or invalid.
  • expired_token: The token has expired and must be refreshed.

So, apart from unsupported_response_type (and unauthorized_client but we do not know what is the level of support of clients that we need to provide), the rest are justified. However those are standard errors,... we are not defining them, just alerting on possible errors to be received by clients in normal use cases. We should point to the RFC with the full list of errors as other errors not included in the list could be raised on peculiar situations (the oAuth server will behave as a complete oAuth server)

@andamian

andamian commented Jun 27, 2025 •

Copy link
Copy Markdown
Contributor

On OAuth 2 (and access token) vs OIDC (and ID token), please see https://auth0.com/blog/id-token-access-token-what-is-the-difference/

From the server-side it makes sense to use OIDC (adding in additional claims to include missing information like ORCID), but from the client-side in the VO, as the client does not actually care about the ID token, the difference between OAuth 2 and OIDC is moot. Both the discovery and dynamic registration endpoints are the same, and so it does not make sense to try to distinguish between the two within the VO.

Thank you for the article @aragilar. It probably depends what you mean by clients, but in a federated environment I don't see clients not using ID tokens. To launch jobs in a science platform for example, a client app needs to know the identity of the user in order to use appropriate posix details.

Very likely that there are some things that I'm not quite grasping here and given the length of this thread, I suggest it would be more effective to schedule a call and discuss these. What do you think?

@aragilar

Copy link
Copy Markdown

Yeah, I think discussing this over a call makes a lot of sense! Is it worth sending out a doodle pool (or whichever one works well now) to find a good time.

@mbtaylor

Copy link
Copy Markdown
Member

I agree, I think a call would be productive.

@jesusjuansalgado

Copy link
Copy Markdown
Contributor Author

Yes, we already mentioned during the interop that we need to call for a dedicated meeting. Probably us and Sara B. or do you want to invite also server providers into the discussion?

@mbtaylor

Copy link
Copy Markdown
Member

I think it would be good to invite anybody who is interested in exposing OAuth2-based services in this way, which hopefully is many/most/all of the OAuth2-using services in the VO. We want to come up with something that's acceptable to as many service providers as possible, and having broad input at an early stage should help with that.

@rra

rra commented Jul 4, 2025

Copy link
Copy Markdown

Hi folks, I was asked to weigh in on this discussion from Rubin's perspective. Our effort allocation for working on authentication mechanisms is still a few months away, so unfortunately my feedback is mostly a high-level discussion of requirements. I am hesitant to offer any detailed feedback on implementation strategy until we've had a chance to start an implementation.

I have also not yet had a chance to read the underlying document here in detail, and I apologize for that. It's entirely possible that some of the things I say below are already addressed there.

Background

Rubin uses OpenID Connect to authenticate the user, after which browser authentication to individual services is done via a cookie. The user can create a bearer token to use from programs or with clients such as TOPCAT, and in the special case of our Jupyter-based notebook service, creation of that token is handled for the user.

Directionally, where we want to go is towards allowing users to use some type of device registration flow to get a token in a more secure way than cutting and pasting tokens between windows and systems and thus potentially exposing them or sending them to the wrong site.

General reactions

Our immediate problem is probably orthogonal to this proposal: currently, we have to advertise HTTP Basic Auth for some services because IVOA clients do not understand RFC 6750 bearer token authentication and its challenges. This is annoying because clients respond to Basic Auth challenges by asking for a username and password, and our users do not have a password; they only have tokens. We therefore have to document how to respond to that prompt and where to put the token (in the password field, with anything in the username, similar to how GitHub handles Basic Auth).

The most valuable change we could get in the short term would be support for bearer challenges which would cause the client to prompt for a token, something the user has or can get through following our documentation, rather than a password that they don't have and can't get and have to understand is really a token.

More directly on topic, though, I am also in favor of standardizing a way to advertise to IVOA clients that some way to use a more secure device authorization flow to get a token for later use as a bearer token is available. It's not clear to me right now, and probably will not become clear without more research and experimentation, whether the right way to do that is RFC 8628, OpenID Connect Client-Initiated Backchannel Authentication Flow, or some other alternative, so I have no useful commentary on those specific details.

Auth schemes

We use bearer as the auth scheme for our challenges (except where we currently have to use basic for client support), and I am very reluctant to change this to something IVOA-specific. This is a well-understood auth scheme with standard protocol documentation that generic clients know how to support.

I have no architectural objections to also adding ivoa_* auth scheme challenges, although the added complexity is a bit unappealing. We have not tested client behavior with multiple challenges and would want reassurance that adding additional challenges wouldn't break already-working clients.

Challenge metadata

When we have device registration and device code flow, I agree with the need to somehow advertise where the client should go to initiate that flow.

I am less convinced of the point of advertising authentication flow metadata outside of the specific case of device registration, since I'm not sure what the client is expected to do with that information. For browser-facing services, we simply redirect unauthenticated users into an OpenID Connect authentication flow, so no challenge metadata is useful. For clients, outside of dynamic client registration, if the client is not authenticated, there is nothing that the client can do about it and we do not want the client to attempt to do anything about it. The correct action is for the user to read our documentation for how to obtain a token, obtain that token, and provide it to the client. The client isn't involved in that process.

Obviously the hope is that in the long run those use cases will be supplanted by some sort of more-secure device registration, but I would expect that to take a very long time (years), so the bearer token flow will remain with us for quite a long time.

Client IDs

I cannot imagine us implementing a device code flow of any type without dynamic client registration and unique client IDs for every script or client install that wants to obtain a token. We certainly will not be maintaining or consuming a static list of known client IDs.

The point of supporting a device code flow, from our perspective, is to provide a more secure way for the user to obtain a token and associate it with a specific application so that they don't cut and paste and copy and reuse tokens willy-nilly and thus risk losing or exposing them. We therefore want it to be immediately usable for any API client the user wants to use, including the one that they just wrote themselves in Python using some standard library or external tool to obtain tokens for it.

Scopes

We use token scopes to limit which services and APIs a given token has access to, and would love to have some way to round-trip that through the challenge to the client and back to any device code flow to inform the defaults the user will be presented with for scopes when authorizing the device. It would also be great to have some way for a client with an existing token with inadequate scopes to be able to request additional scopes be added to their existing token (or be reissued a token that the client knows to substitute for their existing token, I guess). This is more in the weeds of whatever specific protocol we use for that flow, but I wanted to call it out since it has implications for the challenge metadata.

@mbtaylor

mbtaylor commented Jul 4, 2025

Copy link
Copy Markdown
Member

@rra, thanks for your comments. If Rubin doesn't have effort for finalising authentication mechanisms right now so be it, but your ideas about the way that might eventually work are valuable, and hopefully we can write the document in such a way that future revisions to accommodate emerging practices by you and others will be possible.

We have not tested client behavior with multiple challenges and would want reassurance that adding additional challenges wouldn't break already-working clients.

Standards-compliant clients ought to be able to handle this. RFC 9110 makes clear that multiple WWW-Authenticate challenges in the same response are permitted; sec 11.4 says "When creating their values, the user agent ought to do so by selecting the challenge with what it considers to be the most secure auth-scheme that it understands, obtaining credentials from the user as appropriate." Sec 11.6.1 also says about WWW-Authenticate "Furthermore, the header field itself can occur multiple times."

The most valuable change we could get in the short term would be support for bearer challenges which would cause the client to prompt for a token, something the user has or can get through following our documentation, rather than a password that they don't have and can't get and have to understand is really a token.

The document currently does not, but could, advise clients on encountering a Bearer challenge to prompt the user for a token. The domain of such a token would have to be interpreted in the context of a supplied "realm" parameter (see RFC 6750 section 3); in absence of a supplied realm it might be safe to assume the domain to be equivalent to the origin of the URL supplying the challenge, though as far as I know(?) that interpretation is not explicit anywhere.

I think this client behaviour is not common practice because it's not very common for users to have access to suitable tokens (which are typically hidden inside clients). A well-educated Rubin user, who knows they are talking to a Rubin service, might know enough to find their Rubin token and paste it in there. But IIUC not all VO OAuth2 services provide users with a convenient way to generate a suitable token, so in general this is not a good user experience. But maybe it's better than (for Rubin) being made to enter a token as if it's a password and (for other services) just not being able to authenticate.

Do you want me to prototype this behaviour in topcat with a view to considering it as one way for clients to talk OAuth2? Though I still think we should pursue the other ideas here for more robust ways to handle token generation and exchange.

@mbtaylor

mbtaylor commented Jul 4, 2025

Copy link
Copy Markdown
Member

PS: if you want a really quick and hacky improvement to the user experience for Rubin users exposed to Basic Auth when they should be using tokens, you can just issue a challenge like:

   www-authenticate: Basic realm="Use token for password and any username"

At least in TOPCAT the user would see that realm text when prompted for username and password.

@jesusjuansalgado

Copy link
Copy Markdown
Contributor Author

@rra
Thanks for your comments! Here you have a detailed answer, taking into account the text proposed into the current pull request.

Bearer vs Basic Challenges
I completely agree with your frustration about having to fall back to HTTP Basic Auth purely because clients don’t understand Bearer token challenges (RFC 6750). One of the primary motivations for defining ivoa_bearer (or, potentially, just using Bearer directly) is exactly to avoid this problem.

The goal of this pull request is that VO clients like TOPCAT, PyVO, Astroquery, and others will:

  • Recognize WWW-Authenticate: Bearer or ivoa_bearer challenges
  • Prompt the user explicitly for a token (rather than username/password)
  • Know where to obtain further metadata about supported login flows

I dislike to fake a username and paste tokens into password prompts. This was considered as an optional approach but it is not a preferred standard and it is not included in the current pull request.

Device Authorization Flow
I also think that the device code flow (RFC 8628) or similar mechanisms should be the future for headless clients.

In the current pull request:

  • We do not hard-code the use of device code flow.
  • Instead, the discovery_url points to the server’s well-known metadata.
  • Clients can read that metadata to see which grant types (e.g. device_code, authorization_code, etc.) are available.

Currently, we describe in the example the most popular and easy to implement flow but the pull request aims to stay flexible rather than mandate one particular flow.

Challenge Metadata
On the question of challenge metadata:
“I am less convinced of the point of advertising authentication flow metadata outside of the specific case of device registration…”

I see this similarly. For browser-based flows, we expect the service to simply redirect unauthenticated users to an OIDC login page. No challenge metadata is typically necessary.

However, for non-browser clients, the user is stuck unless they know:

  • where to go to obtain a token
  • which flows are supported (device code, dynamic registration, etc.)

Hence the idea of including:

  • a discovery_url in the WWW-Authenticate header
  • optionally scopes, client registration endpoints, etc.

This is not mandatory for browser use cases, but it’s critical for CLI tools like PyVO or scripts running on HPC clusters.

Client IDs and Dynamic Registration
I recognise that Dynamic registration (RFC 7591) is the modern and secure approach. However, I am not totally sure about the possible problems derived of this approach for the authorisation server administrators. I have the feeling the this could produce a wild environment but we want to ask experts on this. To control the dynamic registration would imply admin operations, I think.
In the VO context, it requires authorisation servers to manage and validate potentially untrusted client registrations in a coordinated way.

The current pull request propose to document some recommended static client IDs for widely deployed VO tools (e.g. TOPCAT/STILTS, PyVO) for convenience (maybe only one like ivoa-client) so the admins could control the scope allowed for these command line clients. This would be easily to maintain. In other case, e.g. TOPCAT should maintain a registration of different secrets per IAM service and use the adequate one or we should allow dynamic registration on all the authorisation servers. Currently, not all the services are allowing that. At the SRCNet we are not requesting any client_id because, in some way, all the clients are sharing the same one (a default "public" one with reduced scopes)

Maybe, the final solution should support both static and dynamic client IDs, depending on the environment but this is something I would like to clarify with admins.

Scopes
The current pull request recommends includes scope details through the discovery metadata. This is why it is not part of the challenge response. I think that could be enough as it is closer to the standard.

@jesusjuansalgado

jesusjuansalgado commented Jul 9, 2025 •

Copy link
Copy Markdown
Contributor Author

BTW, I have been speaking with the experts on the SRCNet, and dynamic registration is indeed allowed. This could simplify the client's development, but there is a possible risk associated. One expert from the SRCNet will attend the DSP meeting

@rra

rra commented Aug 7, 2025

Copy link
Copy Markdown

Hi all,

Sorry to have dropped a huge message and then not followed up further. I'll try to be a bit more regular (and concise) in replying to the discussion!

There are multiple discussion threads here, and I wonder if it may make sense to move some of the broader discussion to a mailing list to hammer out the higher-level goals before coming back to specific language.

Bearer tokens: I think where we arrived at for Rubin's current use case, where the user is able to obtain a bearer token for programmatic and client use but where device code registration is not (yet) supported is that, due to existing use of bearer challenges by other sites with different token policies than Rubin, it doesn't make sense for clients such as TOPCAT to prompt the user for a token when seeing any bearer challenge. However, RFC 6750 explicitly allows additional auth-param attributes ("Other auth-param attributes MAY be used as well"). We could therefore define an ivoa-prompt=true boolean attribute that indicates to clients such as TOPCAT that it's reasonable to prompt the user directly for a token.

I considered defining an attribute that would point the user towards documentation for how to obtain that token, but this introduces a new phishing risk if the client blindly presents that as a clickable URL, and I'm not sure the risk is worth the benefits.

Device registration: I did a small amount of additional research on security risks with device code flows, and the security behavior of unconstrained dynamic device registration initiated by the device is grim. This is widely used in real-world phishing attacks by attackers who initiate the device code flow and then trick users into entering their code, and this appears to be a very easy attack to perform. I'm not yet sure what approach to preventing this has the most promise. One idea that I'm kicking around at the moment is to require the user go to an authenticated web site and say explicitly "I want to get an authentication token for my client/script/whatever" and be issued a new client ID (and password?), which they would need to then cut and paste into the client/script so that it can initiate a device code flow as a registered client. This is awkward but would defeat most of the phishing attacks.

Given what I know right now, I'm opposed to any list of known static client IDs because I can't see any reason why an attacker wouldn't just copy that client ID when starting a phishing attack. It therefore is equivalent to not sending a client ID at all so far as I can tell; it looks like it provides some client identification and security benefit, but this is just an illusion.

If sites want to not require a client ID, that's up to them and their local security policy. Just be aware that approach is very vulnerable to phishing (but of course we're living with HTTP Basic Auth right now, so it's all relative).

Audience restriction of bearer tokens: There has been some discussion elsewhere about how a client such as TOPCAT, after it has obtained a token via device code flow, can know where it is safe to use that token. The worry is that the client could be tricked into sending the user's token to a hostile site, which could then use it to act on behalf of the user.

I poked around a little and, so far as I can tell from the various RFCs, the intended mechanism here is to define, in the bearer token, its intended audience. However, exactly how to do this is left rather undefined, so if we want a standard mechanism to do this, it looks like we're going to have to directly define one in the IVOA standard.

At first glance, I think a reasonable way to do that would be to say that the token returned by the device code flow for an IVOA protocol MUST be a JWT with an aud field that contains a list of HTTP origins for which the token is valid. A client that uses device code flow to obtain an access token then MUST NOT send that token to any HTTP origin not listed in the aud claim of that token. I think this would achieve the desired security properties, and it is a valid and common way of using the aud claim. The main drawback is that sites that are using some third-party OAuth implementation may not already be generating tokens compliant with those requirements, and it's not clear how easy it will be to make third-party OAuth code comply.

jesusjuansalgado and others added 4 commits April 16, 2026 00:14
Use the standard Bearer scheme with the RFC 9728 resource_metadata pointer
instead of a custom ivoa_bearer scheme / discovery_url. Add the mandatory
domain-proxy check (PRM origin == issuer origin; requested URL covered by the
PRM resource) and audience-bound tokens (RFC 8707 resource indicator or RFC 8693
token exchange). Rework the token flow steps and add a sequence diagram.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@jesusjuansalgado
jesusjuansalgado requested a review from pdowler July 4, 2026 19:13
@mbtaylor

mbtaylor commented Sep 4, 2026

Copy link
Copy Markdown
Member

@jesusjuansalgado this looks well-written and I think? secure, and it relies on existing RFCs rather than inventing anything IVOA-specific. Your test IAM-based client/server code at https://github.com/jesusjuansalgado/authvo-clients runs smoothly for me; I've also written some code based on the text, though I haven't managed to test that yet. So far so good.

But I've got at least one concern about it.

RFC 9728 section 3.3 says:

If the protected resource metadata was retrieved from a URL returned by the protected resource via the WWW-Authenticate resource_metadata parameter, then the resource value returned MUST be identical to the URL that the client used to make the request to the resource server. If these values are not identical, the data contained in the response MUST NOT be used.

But AuthVO sec 5.4.3 says:

Gate (b) - resource coverage:
the URL the client originally requested MUST be covered by the resource value, i.e. be equal to it or below it at a path-segment boundary

So AuthVO looks more lenient than RFC9728; AuthVO allows the protected resource metadata resource key to name an ancestor of the resource requested, while RFC9728 requires it to be identical to the resource requested.

The restriction as written in AuthVO would be much more convenient, because we don't want to have to carry out this authentication flow (which requires user interaction) for every single resource, but it looks like we're doing something that contradicts the RFC.

Note even the AuthVO interpretation would require new token acquisition flows for resources that cannot be described by a single ancestor (i.e. ones with different origins) - some services will have these, though not all.

A possible solution is the optional protected_resources key described in RFC9728 section 4. I think that might help, but I haven't 100% convinced myself that resolves this.

So I think there is still a bit of work to do. But having said that: the discussion of how to handle bearer tokens in this PR, in PR #18 and in Issue #6 is now extremely lengthy and hard to digest (there are some comments buried in there relevant to the point I've raised above). I understand that you've already agreed with @aragilar that this is basically the way to go. If none of the other interested parties (@rra, @andamian, @pdowler, others?) are opposed then maybe we should merge this PR and tackle any remaining problems under new Issues/PRs based on the resulting document. Comments from those interested parties agreeing/disagreeing with this suggestion would be very welcome.

@rra

rra commented Sep 8, 2026

Copy link
Copy Markdown

I'm sorry to have taken so long to review the updated document. Thank you for the ping!

This looks like solid progression to me and I wholeheartedly endorse @mbtaylor's plan to merge this and then address any additional concerns with subsequent PRs.

For those future PRs:

  • So far as I can tell, it seems generally safe in an IVOA context to relax the requirement that the resource values match, and it will make deployment much easier. I agree that protected_resources is the truly standard way to do this, but I think it will be common for the resources to not be easily enumerable for REST services. The one case where I could see this relaxing of the security rules not being safe is the case where the resource URL points to a shared hosting facility where untrusted parties can add more URLs below that URL, but in the IVOA case I think we can simply prohibit this. It would be nice, though, to learn from the IETF discussion why the restriction is this tight, since perhaps something came up that we haven't thought of.
  • Device registration is still somewhat vulnerable to phishing because an attacker can perform the whole flow and then prompt the user to confirm and many users will blindly click through regardless of what text is on the confirmation page. Many sites won't care and will consider this adequate security, but it may be worth noting this (in a follow-on PR). I have some vague thoughts about how to mitigate this by requiring users to pre-register their intent to register a client in the cases where one wants to be particularly cautious about phishing.

@pdowler pdowler left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I do intend to review more thoroughly, but would benefit from this being merged so I can see the document directly, so I 2nd (3rd?) @mbtaylor 's suggestion to merge and aim for smaller issues with better focus.

@andamian

andamian commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

+1 for merging.

@mbtaylor

mbtaylor commented Sep 9, 2026

Copy link
Copy Markdown
Member

Thanks all. Unless I get an objection from @aragilar I will merge in a day or two.

@mbtaylor
mbtaylor merged commit 1b8bbca into main Sep 11, 2026
1 check passed
@mbtaylor
mbtaylor deleted the bearer-tokens branch September 11, 2026 09:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants