Skip to content

FRED questions use revised series values, so resolution can differ from what forecasters saw #283

Description

@arpjw

Hi — we hit this while building a vintage-aware pipeline of our own and thought it was worth reporting. Two related issues in the FRED path, both stemming from FRED serving current values for historical dates unless asked otherwise.

Both verified by reading the source at 2e78c66505a7405d1d76ec726bb23f97f8277cc5 (main, 2026-08-27).

1. _fetch_observations omits the real-time parameters

src/sources/fred.py, _fetch_observations (line 267):

params = {
    "api_key": api_key,
    "file_type": "json",
    "series_id": series_id,
    "limit": 10000,
}

/fred/series/observations without realtime_start / realtime_end returns the latest revised value for every historical date, not the value published at the time. realtime, vintage and alfred do not appear anywhere in the file.

Reproduction, re-run today (2026-08-27) against ALFRED — US unemployment for August 2023:

vintage_date=2023-10-01  ->  3.8   (as published)
current series           ->  3.7   (as revised)

A question asking whether unemployment exceeded 3.75 has opposite answers depending on which is read.

One caution if you fix this: vintage_date is not a parameter of the observations endpoint. The server accepts it, returns HTTP 200, and silently serves the revised value anyway — so a fix using vintage_date looks like it worked and does not. realtime_start / realtime_end are the ones that take effect.

2. The due-date baseline is re-read from the revised series at scoring time

This is the one with a direct effect on scores.

Questions are phrased "Will {series} have increased by {resolution_date} as compared to its value on {forecast_due_date}?", and forecasters are shown freeze_datetime_value, captured at freeze time (src/sources/fred.py, _transform_series).

But resolution reads both endpoints from the resolution file (src/sources/_dataset.py, lines 54-76):

# Get values at forecast_due_date (for imputation)
df_standard = pd.merge(..., left_on=["id", "forecast_due_date"], ...)
df_standard["market_value_on_due_date"] = df_standard["value"]
...
float(row["resolved_to"] > row["market_value_on_due_date"])

and that file is rebuilt from a fresh _fetch_observations call on each run.

So if the due-date value is revised between freeze and resolution, the baseline a forecaster was shown is not the baseline they are scored against. Because the comparison is a strict inequality between two values, a revision smaller than the change being forecast can flip the label.

The exposure is not uniform across series. From our own measurement over 48 quarterly vintages, 2015-2026:

series observation dates revised median abs change
PAYEMS 859/1050 (81.8%) 0.009%
CPIAUCSL 191/953 (20.0%) 0.085%
UNRATE 84/941 (8.9%) 2.17%

UNRATE revises least often and by the most, which is the combination most likely to flip an inequality.

Suggested fix

Pass realtime_start / realtime_end on the observations call, and pin the due-date baseline to the vintage current at forecast_due_date rather than re-reading it at scoring time. ALFRED serves this at no extra cost — https://alfred.stlouisfed.org/graph/alfredgraph.csv also works without a key (note alfredgraph.csv, not fredgraph.csv).

Happy to send a PR if useful.


Reported by Laplace Research. We searched the tracker for existing vintage/revision reports and did not find one; apologies if we missed it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions