Learn how to interact with this dataset using the Ouro SDK or REST API.
API access requires an API key. Create one in Settings → API Keys, then set OURO_API_KEY in your environment.
Get dataset metadata including name, visibility, description, and other asset properties.
Get column definitions for the underlying table, including column names, data types, and constraints.
| Column | Type |
|---|---|
| carbon_and_hydrogen | integer |
| carbon_no_hydrogen | integer |
| ch_share | real |
| id | uuid |
| no_carbon | integer |
| no_carbon_share | real |
| partial_year | integer |
| total_entries | integer |
| year | integer |
Fetch the dataset's rows. Use query() for smaller datasets or load() with the table name for faster access to large datasets.
Update dataset metadata (visibility, description, etc.) and optionally write new rows to the table. Writing new data will replace the existing data in the table. Requires write or admin permission on the dataset.
# Get column definitions for the underlying table
columns = ouro.datasets.schema(dataset_id)
for col in columns:
print(col["column_name"], col["data_type"]) # e.g., age integer, name text# Option 1: All rows as a Pandas DataFrame
df = ouro.datasets.query(dataset_id)
print(df.head())
# Option 2: Read-only SQL — pass a query string; use {{table}} as the placeholder
agg = ouro.datasets.query(
dataset_id,
"SELECT col, count(*) AS n FROM {{table}} GROUP BY col ORDER BY n DESC",
)import pandas as pd
# Update dataset metadata
updated = ouro.datasets.update(
dataset_id,
visibility="private",
description="Updated description"
)
# Update dataset data (replaces existing data)
data_update = pd.DataFrame([
{"name": "Charlie", "age": 33},
{"name": "Diana", "age": 28},
])
updated = ouro.datasets.update(dataset_id, data=data_update)import os
from ouro import Ouro
# Set OURO_API_KEY in your environment or replace os.environ.get("OURO_API_KEY")
ouro = Ouro(api_key=os.environ.get("OURO_API_KEY"))
dataset_id = "019fe3d8-6612-7381-bc5b-2e66e05b836d"
# Retrieve dataset metadata
dataset = ouro.datasets.retrieve(dataset_id)
print(dataset.name, dataset.visibility)
print(dataset.metadata)Per-year composition of the Crystallography Open Database by element presence in the reported formula, publication years 1960-2026. Queried from the COD REST API (crystallography.net/cod/result, format=count) on 2026-08-09 UTC. Classes: no_carbon (no C in formula; the inorganic class), carbon_no_hydrogen (C but no H; carbonates, carbides, cyanides, oxalates...), carbon_and_hydrogen (C and H; organic, organometallic, MOF-like). Shares are of all COD entries with that publication year. Caveats: COD holdings reflect ingestion and backfill history (mineral-collection backfills, journal ingest lag), so absolute recent-year counts understate true publication volume; element presence is a heuristic, not a curated class. 2026 is a partial year.
What year is your training data from? The temporal footprint of the Crystallography Open Database
If you train a model on the Crystallography Open Database today, the inorganic structures in your training set have a mean deposition year of 1996. Half were deposited by 2000. The COD covers ~20% of ICSD and captures <4% of post-2015 inorganic crystallography. What does that mean for ML?
Where did the structures go?
Capstone of the COD series. The open database's inorganic intake collapsed from 2,949 entries (2003) to 316 (2024), but crystallography didn't die: ICSD grew from 100k (2007) to 335k (2026) structures. Publisher-family anatomy of all 46,753 no-carbon COD entries since 1990 shows three staggered exit waves (Elsevier+Wiley 2004-05, mineralogy 2014-15, ACS 2015-16) with only the RSC still feeding the open record (70% of 2025 intake). Open capture fell from ~15% to ~2% of ICSD's annual intake.