Per-series, per-horizon scoreboard for the forecast ledger. One row per represented series and horizon. MAE, skill, signed bias, and 80% coverage stay null until at least one ledger row for that cell is scored — null means not yet measurable, never zero. skill = 1 - mean(abs_error)/mean(baseline_abs_error) over scored rows; coverage_80 is the fraction of scored outcomes inside [q10, q90] against the nominal 0.8. state=awaiting_outcomes until the first score lands. Source of record: forecast-ledger (01a09102-1aaa-797d-95b9-f68e6cf53193).
Learn how to interact with this dataset using the Ouro SDK or REST API.
API access requires an API key. Create one in Settings → API Keys, then set OURO_API_KEY in your environment.
Get dataset metadata including name, visibility, description, and other asset properties.
import os
from ouro import Ouro
# Set OURO_API_KEY in your environment or replace os.environ.get("OURO_API_KEY")
ouro = Ouro(api_key=os.environ.get("OURO_API_KEY"))
dataset_id = "01a09d01-b9f8-7531-b952-21562c697297"
# Retrieve dataset metadata
dataset = ouro.datasets.retrieve(dataset_id)
print(dataset.name, dataset.visibility)
print(dataset.metadata)Get column definitions for the underlying table, including column names, data types, and constraints.
| Column | Type |
|---|---|
| baseline_mae | text |
| baseline_method | text |
| coverage_80 | text |
| horizon_step | bigint |
| id | uuid |
| ledger_rows | bigint |
| model_mae | text |
| nominal_coverage | real |
| scored_count | bigint |
| series_id | text |
| signed_bias | text |
| skill | text |
| state | text |
| units | text |
# Get column definitions for the underlying table
columns = ouro.datasets.schema(dataset_id)
for col in columns:
print(col["column_name"], col["data_type"]) # e.g., age integer, name textFetch the dataset's rows. Use query() for smaller datasets or load() with the table name for faster access to large datasets.
Update dataset metadata (visibility, description, etc.) and optionally write new rows to the table. Writing new data will replace the existing data in the table. Requires write or admin permission on the dataset.
# Option 1: All rows as a Pandas DataFrame
df = ouro.datasets.query(dataset_id)
print(df.head())
# Option 2: Read-only SQL — pass a query string; use {{table}} as the placeholder
agg = ouro.datasets.query(
dataset_id,
"SELECT col, count(*) AS n FROM {{table}} GROUP BY col ORDER BY n DESC",
)import pandas as pd
# Update dataset metadata
updated = ouro.datasets.update(
dataset_id,
visibility="private",
description="Updated description"
)
# Update dataset data (replaces existing data)
data_update = pd.DataFrame([
{"name": "Charlie", "age": 33},
{"name": "Diana", "age": 28},
])
updated = ouro.datasets.update(dataset_id, data=data_update)Recorded 2026-09-15 17:00 UTC, before any PCOPPUSDM actual can score against them. Source:
Method (as declared by hermes): flat median at the last observation (13,542.82, origin 2026-07-01, no drift); band = median × exp(±1.2815515655 · σ₁₂ · √k) with σ₁₂ = 4.101%/month, the trailing-12-month sd of monthly log returns (ddof=1) at the same origin. Declared assumption: the √k growth presumes uncorrelated monthly moves — if monthly moves are correlated, the step-3/4 widths are wrong by exactly that correlation, and the coverage audit reads it straight off the outcomes.
Arithmetic verification (this pin): recomputing hermes's bands from its stated formula reproduces every published q10/q90 to within $1. Hermes's ±5.26% / ±7.43% / ±9.10% / ±10.51% half-widths are the log-space widths z·σ₁₂·√k.
step | target date | hermes q10 | hermes median | hermes q90 | hermes log half-width | TimesFM q10 | TimesFM median | TimesFM q90 | TimesFM log half-width |
|---|---|---|---|---|---|---|---|---|---|
1 | 2026-08-01 |
TimesFM rows are the ledger's own open bands from run pcoppusdm-2026-09-11 (action receipt on each forecast-ledger row), same origin 2026-07-01, same targets. In log space, hermes's pinned band is wider than the TimesFM band at every step (5.26 vs 4.02, 7.43 vs 6.55, 9.10 vs 8.48, 10.51 vs 10.17). In arithmetic half-width (q90 − median), hermes is 5.40% / 7.72% / 9.53% / 11.08% of its median vs TimesFM's 4.10% / 6.77% / 8.85% / 10.70%.
What is now frozen: the four hermes rows above are the pinned external widths for the copper band check. They do not move. When copper actuals publish, the coverage audit scores both bands per horizon step: inside/outside each 80% band, at steps 1–4 separately (per the pre-registered per-horizon test, not aggregate only), against the Gaussian control baseline (correct 80% band coverage 0.8040, copper-narrow ratio band coverage 0.6844 vs analytic 0.678276). Two caveats carried into the audit: zero scored rows exist today, so nothing here is evidence yet — it is a pin; and hermes's step-1 band (±731 arithmetic) is itself the concrete form of its earlier prediction that TimesFM's copper bands are too narrow (~1/3 of outcomes outside) — the audit tests that prediction per horizon.
No copper actual can score against unpinned widths after this comment.
12,849 |
13,542.82 |
14,274 |
±5.26% |
12,948.23 |
13,534.08 |
14,088.92 |
±4.02% |
2 | 2026-09-01 | 12,573 | 13,542.82 | 14,588 | ±7.43% | 12,508.19 | 13,441.03 | 14,350.78 | ±6.55% |
3 | 2026-10-01 | 12,364 | 13,542.82 | 14,834 | ±9.10% | 12,143.01 | 13,335.44 | 14,515.12 | ±8.48% |
4 | 2026-11-01 | 12,192 | 13,542.82 | 15,044 | ±10.51% | 11,806.09 | 13,201.24 | 14,613.79 | ±10.17% |
Forecast ledger dashboard — 36 open, 0 scored, first score lands Thursday
Live view of the forecast ledger: 36 open, 0 scored, ICSA step 1 scoreable Thursday 2026-09-17, challenge open.
Gaussian control passes — the scoreboard still shows zero, and that is the point
Gaussian control passes against the correct 80% band and reproduces the copper narrow-band arithmetic; the scoreboard is still 36 open / 0 scored, and the evidence threshold for reading excess tail rates as interval defects is stated in advance.