Skip to content

feat(waterdata): request v1 of the Water Data API, pinnable through api_version - #422

Merged
thodson-usgs merged 5 commits into
DOI-USGS:mainfrom
thodson-usgs:feat/waterdata-api-v1
Oct 2, 2026
Merged

thodson-usgs merged 5 commits into
DOI-USGS:mainfrom
thodson-usgs:feat/waterdata-api-v1

Conversation

@thodson-usgs

@thodson-usgs thodson-usgs commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

TL;DR: The Water Data OGC getters now request v1 of the API, released September 2026; v0 stays online until June 2027, and a new api_version setting can pin v0. get_time_series_metadata sends calls that use a field v1 dropped to v0 with a DeprecationWarning. The maintainer should confirm four choices and decide one open question (field-measurement row order).

Closes #421. Merge order is in #426.

Changes

  • OGC requests go to /ogcapi/v1 (release post), set by OGC_API_VERSION in waterdata/endpoints.py. Statistics and STAC have no v1 and stay on v0.
  • WaterdataConfiguration gains api_version, a setting like base_url (ADR 0010) because it describes the deployment, not the query. A configure() block or the file's [waterdata] table can set it; the environment refuses it. The file now refuses only BLOCK_ONLY_SETTINGS (base_url): a base URL can redirect requests to another host, and a version cannot.
  • get_time_series_metadata: v1 responds with 400 or 500 to begin_utc, end_utc, state_name (from state), and hydrologic_unit_code. A call that names one, as a filter or in properties, goes to v0 with a DeprecationWarning. The version is passed per request (get_ogc_data(api_version=...)) rather than through configure(), so the caller's configured version is unchanged. The getter gains statistics_begin, new in v1.
  • Behavior changes: get_time_series_metadata returns begin and end in UTC with a time zone, without the four dropped columns. get_field_measurements returns time as a date (tz-naive midnight, as in get_daily) and the time of day in time_of_day.
  • The queryables snapshot and time-series-metadata fixture are regenerated from v1; the snapshot adds method_category on continuous, which fix(waterdata): accept the continuous method_category queryable #423 makes a get_continuous parameter.

For the maintainer

  1. Horizon 2027-06-01 is when the service retires v0, not the NEWS date plus one year, because the v0 routing stops working then.
  2. The file accepts api_version. To make it code-only like base_url, add api_version to BLOCK_ONLY_SETTINGS and update the api_version tests in tests/configuration_test.py.
  3. Dropped fields route to v0 instead of using _accept_legacy_kwargs, the ADR 0012 default for a renamed argument: translating begin_utc inside properties would rename the returned column without an error.
  4. Open: get_field_measurements row order within a day. Rows sort by time, then monitoring_location_id; with time a date, same-day rows at one location keep the service's order. Sorting by time, then time_of_day restores the v0 order, but the same key would reorder get_peaks, so a fix would be specific to field measurements.
  5. NEWS is dated 09/22/2026; set it to the merge date.

Verification

  • Offline: 1192 passed, 12 deselected; branch coverage 98.94% (ratchet 98.9). ruff check, ruff format --check, mypy --strict (61 files), lint-imports (8 contracts), xenon, complexipy, and pre-commit pass; a Sphinx build without notebook execution adds no warnings.
  • Live (2026-09-22): a default get_time_series_metadata call raises nothing under -W error::DeprecationWarning; each dropped field warns and returns v0 rows; a caller-pinned v1 still applies after a v0-routed call; the file, top-level, and environment sources behave as documented; the queryables monitor passes.
  • Live differential (2026-09-29, 60 cases across every Water Data getter, main against the feat(waterdata): request v1 of the Water Data API, pinnable through api_version #422–feat(waterdata): name every column the OGC collections return #425 stack): at api_version="v0" every frame matches main; at v1 the only differences are the behavior changes above and item 4.

🤖 Generated with Claude Code

thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Sep 22, 2026
Taken so this branch carries no edit that conflicts with DOI-USGS#422. The
`data_gap_interval` fix this branch made to the v0 column list is
subsumed by DOI-USGS#422's rewrite of that list for v1, so the resolution keeps
DOI-USGS#422's version -- including `statistics_begin` -- and this branch keeps
only the monitor.

`dataretrieval/waterdata/metadata.py` is byte-identical to DOI-USGS#422's copy
after this merge, which is the check worth making: resolving a column
list the wrong way would silently revert part of the other PR while every
offline test still passed.
thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Sep 23, 2026
Comparing each collection's /schema with its getter's signature found 20
returned columns that were reachable only through **queryables, so the
getter documented neither the column nor the filter:

- get_field_measurements: control_condition, day,
  field_measurements_series_id, measurement_rated, month, reading_type,
  time_of_day, year
- get_peaks: qualifier, time_of_day, value
- get_monitoring_locations: revision_created, revision_modified,
  revision_note
- get_combined_metadata: data_gap_interval, reading_type
- get_time_series_metadata: data_gap_interval, parameter_description
- get_field_measurements_metadata: reading_type
- get_channel: channel_location_direction

Each is now a named parameter, described in the service's own words. day,
month and year take the integer annotation get_peaks already uses. The
monitoring-location attributes every collection accepts as filters but
does not return stay in **queryables. Existing calls send the same
request as before.

Stacked: this commit also carries DOI-USGS#422 (Water Data API v1), DOI-USGS#423
(continuous method_category and the API-version monitor) and DOI-USGS#424 (the
documented-columns monitor), which merge first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Sep 23, 2026
get_daily, get_continuous and get_time_series_metadata list their
returned columns in the properties docstring. The list is hand-written
and went stale twice without anything noticing: continuous was missing
method_category (DOI-USGS#423) and time-series-metadata data_gap_interval
(DOI-USGS#422).

Add a live test that compares each list with the collection's /schema
and names the docstring to edit when they differ. id is excluded on both
sides because only time-series-metadata lists it in its schema, though
all three accept it.

Stacked: this commit also carries DOI-USGS#422 (Water Data API v1) and DOI-USGS#423
(continuous method_category and the API-version monitor), which merge
first.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Sep 24, 2026
Comparing each collection's /schema with its getter's signature found 20
returned columns that could be passed only through **queryables, so the
getter documented neither the column nor the filter:

- get_field_measurements: control_condition, day,
  field_measurements_series_id, measurement_rated, month, reading_type,
  time_of_day, year
- get_peaks: qualifier, time_of_day, value
- get_monitoring_locations: revision_created, revision_modified,
  revision_note
- get_combined_metadata: data_gap_interval, reading_type
- get_time_series_metadata: data_gap_interval, parameter_description
- get_field_measurements_metadata: reading_type
- get_channel: channel_location_direction

Each is now a named parameter, described in the service's own words. day,
month and year are typed as integers, as get_peaks already types them.
The monitoring-location attributes that every collection accepts as
filters but does not return stay in **queryables. Existing calls send
the same request as before.

Stacked on DOI-USGS#422, DOI-USGS#423 and DOI-USGS#424, which merge first; review this commit
alone.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Sep 29, 2026
Comparing each collection's /schema with its getter's signature found 20
returned columns that could be passed only through **queryables, so their
getters documented neither the column nor the filter:

- get_field_measurements: control_condition, day,
  field_measurements_series_id, measurement_rated, month, reading_type,
  time_of_day, year
- get_peaks: qualifier, time_of_day, value
- get_monitoring_locations: revision_created, revision_modified,
  revision_note
- get_combined_metadata: data_gap_interval, reading_type
- get_time_series_metadata: data_gap_interval, parameter_description
- get_field_measurements_metadata: reading_type
- get_channel: channel_location_direction

Each is now a named parameter, documented from the description the
service publishes for it. day, month and year are typed as integers, as
get_peaks already types them. The monitoring-location attributes that
every collection accepts as filters but does not return stay in
**queryables. Existing calls send the same request as before.

A parametrized test checks that each one is sent in the request. It
also covers DOI-USGS#422's statistics_begin, which no other test sends.

The properties lists of get_monitoring_locations and get_channel now
name the columns they were missing (the three revision_* columns and
channel_location_direction), and get_monitoring_locations joins the
documented-columns monitor. get_channel stays out of it, because its
list names the output column channel_measurements_id where the schema
has id.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@thodson-usgs
thodson-usgs marked this pull request as ready for review September 29, 2026 15:12

@ehinman ehinman left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The diffs look okay. I tried running this fetch:

test,_ = waterdata.get_combined_metadata(state_name = "Iowa")

And I got a warning:

dataretrieval-python/dataretrieval/ogc/shaping.py:292: UserWarning: Could not infer format, so each element will be parsed individually, falling back to `dateutil`. To ensure parsing is consistent and as-expected, please specify a format.
  df[col] = pd.to_datetime(df[col], errors="coerce")

What am I doing wrong? Do I have a package version that is causing this warning?

only: the file and the environment refuse it. The API key is scoped to the host
that accepts it, so a redirected call sends no key.
``/ogcapi/<api_version>``, ``/samples-data``, ``/statistics/v0`` and
``/stac/v0``. Code only: the file and the environment refuse it. The API key is

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does "the file and the environment refuse it" mean?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we instead say something like "the file and the environment fail loudly if the user tries to specify the version there rather than in the configuration"?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

suggest something more literal like, fail and raise a warning. Follow convention used elsewhere in the package.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, "refuse" didn't say what happens. It means a ConfigurationError is raised. That's an error, not a warning, which matches how the rest of configuration handles a setting in the wrong place (e.g. a known key in the wrong table). Fixed in 7cba029: the base_url docstring now says "Code only: setting it in the configuration file or through an environment variable raises ConfigurationError." Same wording in the ADR 0011 note and the code comments.

Comment thread dataretrieval/waterdata/metadata.py Outdated
_V0_ONLY_FILTERS: dict[str, str] = {
"begin_utc": "'begin'",
"end_utc": "'end'",
"state_name": "get_combined_metadata(state=...)",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm confused by this. Isn't the input for get_combined_metadata still "state_name"?

Suggested change
"state_name": "get_combined_metadata(state=...)",
"state_name": "get_combined_metadata(state_name=...)",

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch. get_combined_metadata takes both state and state_name, so the remedy shouldn't switch a state_name caller over to state. Rather than hard-code either one, 7cba029 repeats whichever the caller passed: state_name= → get_combined_metadata(state_name=...), state= → get_combined_metadata(state=...). A new test (test_time_series_metadata_state_warning_points_to_the_same_argument) covers both, and I confirmed both against the live service.

requested = set(args.get("properties", ()))
legacy = [n for n in _V0_ONLY_FILTERS if n in args or n in requested]
for name in legacy:
spelled = "state" if name == "state_name" and state_given else name

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we push users to still use "state_name" or "state_code", rather than "state"? I get that it's a convenience function, but I'd rather encourage users to actually use the API's documentation...

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll follow up with this point. I elected to sanitize the APIs a bit, since they all follow slightly different conventions, which gets ugly when you're interacting with several APIs at once.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Following up on this. I'd like to keep state as the recommended argument, but I agree it should be easier to map to the API docs.

Why keep it: the native state parameter is different in each collection. The location collections take a full name (state_name), the statistics getters take state_code in US:XX form, and NGWMN providers take a postal code (state). state accepts a name, postal code, or FIPS code and sends each collection the form it expects. If someone is working across several of these APIs at once, that's one spelling to learn instead of three.

state_name and state_code aren't deprecated, and they won't be. They send the value to the API unchanged, which covers values the conversion doesn't handle, such as non-US FIPS codes. Passing state together with one of them raises a ValueError.

Your point stands, though: state doesn't appear in the API documentation, so a reader there has no way to match it up. 7a48e10 makes every state docstring say it's a dataretrieval argument and name the API field it's sent as, e.g. "A dataretrieval argument rather than an API field: it is sent as the API's state_name." Anyone who starts from the API docs can now find the corresponding argument.

Comment thread dataretrieval/waterdata/metadata.py Outdated
state_name : string or iterable of strings, optional
The name of the state or state equivalent in which the monitoring location
is located.
Deprecated; see ``state``. The name of the state or state equivalent in

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again, not seeing "state" anywhere in API, but maybe I'm missing something. Or, is this artificially created in this PR?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

state isn't from this PR and isn't an API field. It's a dataretrieval argument added with the shared OGC engine (#324). It accepts a name, postal code, or FIPS code and is sent as state_name (or state_code, depending on the collection). This PR only marks it deprecated on get_time_series_metadata, because v1 dropped state_name there. The docstring now says this outright (7cba029): "state is a dataretrieval argument rather than an API field … and is sent as state_name." The bigger question of whether to steer people toward state or the native names is in the thread above.

Comment thread dataretrieval/_configuration_core.py Outdated
ADAPTER_ONLY_SETTINGS: tuple[str, ...] = ("base_url",)
ADAPTER_ONLY_SETTINGS: tuple[str, ...] = ("base_url", "api_version")

#: The adapter-only settings the file also refuses, so only a ``configure()``

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again, this language does not make sense to me. What is meant by "the file also refuses"? Can you clarify this language a bit to be useful to both a human and an agent?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reworded in 7cba029 to say what happens: "Adapter-only settings that only a configure() block can set: writing one in the configuration file raises ConfigurationError (ADR 0011)." The neighboring _REFUSED_ENV_VARS comment got the same treatment.

def test_api_version_is_refused_from_the_environment(monkeypatch):
"""Refused, as every adapter-only setting is. The message also points to
the file, which accepts it."""
monkeypatch.setenv("API_USGS_API_VERSION", "v0")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we actually have instructions anywhere that would prompt a user to do this? "API_USGS_API_VERSION" doesn't even make sense as a variable name. This check might be a little overkill, and the documentation about it being refused in the environment is confusing. I appreciate the defensive programming, but I do not think it's necessary.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think the idea was the API should be configurable, and I suspect that the documented method is to use a configuration file. Environment variables are more of a fallback, but fine for a test.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing tells a user to set it, and you're right that documenting the refusal in user-facing text was more confusing than helpful. 7cba029 removes it there: the api_version docstring and the user guide now just say it can be set in code or in the [waterdata] table of the config file, and that it has no environment variable. The check itself adds no api_version-specific code. It comes from one rule that already guarded API_USGS_BASE_URL: no adapter-only setting can come from an environment variable, since a variable applies to every adapter. The test is there to show the rule also covers the new setting. If someone guesses the variable name, they get an error that tells them where to set it instead of having the value silently ignored. So I've kept the test, and the guarantee now only shows up in the internals.

def test_key_excluded_for_lookalike_host(self):
"""The key is not sent to a typosquatting/lookalike domain."""
url = "https://api.waterdata.usgs.gov.evil.com/ogcapi/v0/daily/items"
url = "https://api.waterdata.usgs.gov.evil.com/ogcapi/v1/daily/items"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How'd I miss the "evil" test, lol.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ha. It checks that the API key isn't sent to a lookalike host (api.waterdata.usgs.gov.evil.com). The key is only attached when the host matches exactly, not by prefix. This PR just updated its path from /v0 to /v1.

thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Oct 1, 2026
Review on DOI-USGS#422: "the file and the environment refuse it" did not say
what happens. Each mention now names the error. The user-facing
api_version docs drop the environment entirely, since nothing directs
a caller to set one.

get_time_series_metadata's state deprecation now names the argument
the caller passed (state or state_name) in its get_combined_metadata
remedy, since that getter accepts both.
@thodson-usgs

Copy link
Copy Markdown
Collaborator Author

@ehinman re the UserWarning from get_combined_metadata(state_name="Iowa"): you're not doing anything wrong, and it isn't your pandas version. It also isn't caused by this PR. I get the same warning with api_version="v0", and the code involved is the same on main.

The column causing it is construction_date. The service sends it at three precisions. For Iowa (27,307 rows):

format example rows
YYYYMMDD 19950812 16,946
YYYY 2005 2,040
YYYYMM 199508 1,716

_type_cols passes it to pd.to_datetime(errors="coerce") with no format. pandas can't infer a single format, so it warns and parses each value on its own. Worse, all 1,716 YYYYMM values come back as NaT, so data is lost without any error. That makes this a real bug, but a separate one from v1, so I'd rather fix it in its own PR than add to this one.

thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Oct 1, 2026
The service records construction_date at day, month, or year precision
(19950812, 199508, 2005). Parsing it as a datetime raised pandas'
"Could not infer format" UserWarning and turned every month-precision
value into NaT -- 1,716 of 27,307 Iowa rows. No one parse keeps all
three, and the schema types it as a string, so it is no longer a time
column. The R package also leaves it as character.

Reported in review on DOI-USGS#422.
@thodson-usgs

Copy link
Copy Markdown
Collaborator Author

Follow-up on the construction_date warning: none of the follow-up PRs fixed it, so it's fixed here in 95340a2. The column is no longer parsed as a datetime. It now holds the strings the service sends (19950812, 199508, 2005). That matches the collection's /schema, which types it as a string, and the R package, which also leaves it as character. No single datetime format keeps all three precisions. Your get_combined_metadata(state_name="Iowa") call now runs without the warning and keeps all 27,307 values. NEWS records the change.

thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Oct 1, 2026
Review on DOI-USGS#422: state does not appear in the API documentation, so a
reader cannot map it to what they find there. Each state docstring now
says it is a dataretrieval argument and names the field it is sent as:
state_name (monitoring locations, combined metadata, NGWMN sites),
state_code in US:XX form (statistics), or state as a postal code
(NGWMN providers).
thodson-usgs added a commit to thodson-usgs/dataretrieval-python that referenced this pull request Oct 2, 2026
state and now county are package-named arguments for domain concepts
that ADR 0013 otherwise leaves to each service's spelling. Review of
DOI-USGS#422 read state as hiding the APIs' own parameters, and nothing
recorded why it did not contradict the rule.

ADR 0013 gains a clause stating the four conditions both meet: the
native parameters stay, the conversion is an exact table kept against
the service's reference collection, combining the two raises, and the
docstring names what the argument is sent as. CONTEXT.md defines State
and County as domain terms with each service's spelling, and Unified
argument as a core term.
@ehinman

ehinman commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

We got asked about the switch to v1 in drpy at the public NWIS decommission meetings. People are excited for this!

thodson-usgs and others added 5 commits October 2, 2026 11:56
Water Data OGC requests now go to /ogcapi/v1, released September 2026;
v0 stays online until June 2027. Statistics and STAC have no v1.

WaterdataConfiguration gains api_version, which a configure() block or
the [waterdata] file table can set and the environment refuses. The file
now refuses only BLOCK_ONLY_SETTINGS (base_url): a base URL can redirect
requests to another host, and a version cannot.

get_time_series_metadata sends a call that names a field v1 dropped
(begin_utc, end_utc, state_name, hydrologic_unit_code) to v0 with a
DeprecationWarning, without changing the caller's configured version,
and gains statistics_begin. Behavior changes: its begin and end are UTC
with a time zone and those four columns are gone; get_field_measurements
returns time as a date, with the time of day in time_of_day.

Closes DOI-USGS#421.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Review on DOI-USGS#422: "the file and the environment refuse it" did not say
what happens. Each mention now names the error. The user-facing
api_version docs drop the environment entirely, since nothing directs
a caller to set one.

get_time_series_metadata's state deprecation now names the argument
the caller passed (state or state_name) in its get_combined_metadata
remedy, since that getter accepts both.
The service records construction_date at day, month, or year precision
(19950812, 199508, 2005). Parsing it as a datetime raised pandas'
"Could not infer format" UserWarning and turned every month-precision
value into NaT -- 1,716 of 27,307 Iowa rows. No one parse keeps all
three, and the schema types it as a string, so it is no longer a time
column. The R package also leaves it as character.

Reported in review on DOI-USGS#422.
Review on DOI-USGS#422: state does not appear in the API documentation, so a
reader cannot map it to what they find there. Each state docstring now
says it is a dataretrieval argument and names the field it is sent as:
state_name (monitoring locations, combined metadata, NGWMN sites),
state_code in US:XX form (statistics), or state as a postal code
(NGWMN providers).
The field-measurements fixture still carried v0's full datetimes in
time, so nothing tested the documented v1 behavior: time is a date,
parsed to a tz-naive midnight, with the time of day in time_of_day.
The fixture now matches what v1 sends (checked live), and a test pins
the parse.

Also: the begin_utc/end_utc deprecation remedies now read begin=...
and end=..., in the same call form as their get_combined_metadata
siblings, and ogc_api_url takes api_version keyword-only.
@thodson-usgs
thodson-usgs force-pushed the feat/waterdata-api-v1 branch from 0e69af6 to 955db3d Compare October 2, 2026 17:00
@thodson-usgs
thodson-usgs merged commit 917a1d3 into DOI-USGS:main Oct 2, 2026
11 checks passed
@thodson-usgs
thodson-usgs deleted the feat/waterdata-api-v1 branch October 2, 2026 18:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

v1 of the water data apis has been released

2 participants