Which CoinNudge datasets preserve what was known, observed and ingested at each point in time? · By CoinNudge Research · Method reviewed 2026-09-14 · Guide updated 2026-09-14 · Historical study updated · Observation range: – · Calculation coinnudge-pit-coverage-2.1 · Auto-refresh about every 21600 seconds
Point-in-time crypto data for backtesting
Current answer: As of , using CoinNudge read-only research archive catalog: CoinNudge currently catalogs 14 populated archive families plus one aggregate row for 10 active Market Events releases. Every archive envelope separates source event time, local observation time and archive ingestion time; backfilled history must not be marketed as known at the original event time.
Historical study updated
Subscribe to research data · Check a real sample · Get Telegram alerts
Current source-backed snapshot
A backtest can look accurate while using information that was not available at the simulated decision time. CoinNudge's evidence archive separates source event time, local observation time and archive ingestion time, and it labels historical backfills instead of presenting them as contemporaneously known. This page exposes actual populated data families and first/last coverage so researchers can decide what is PIT-eligible, what is reconstructed and where a study must remain out of scope.
Input documentation: CoinNudge data services · CoinNudge sample · CoinNudge API documentation · CoinNudge methodology
Observation window: Actual catalog coverage by dataset
Calculation cadence: Auto-refresh about every 21600 seconds
Calculation version: coinnudge-pit-coverage-2.1
Download current dataset: JSON · CSV Raw values retain the dataset's published precision. Free fair-use limit: 60 requests per minute per IP, shared across all Research JSON and CSV endpoints.

| Dataset family | Rows | Stored bytes | First event UTC | Last event UTC | Partitions / releases | Backtest availability test | Meaning and boundary |
|---|---|---|---|---|---|---|---|
| personal_signal_summary | 109 | 37391 | 2026-09-05 00:00:00 UTC | 2026-09-14 00:00:00 UTC | 10 | observed_at <= simulated decision time | Daily personal signal counts by kind; no user identities or raw evidence |
| delivery_status_summary | 10 | 42504 | 2026-09-05 00:00:00 UTC | 2026-09-14 00:00:00 UTC | 10 | observed_at <= simulated decision time | Daily latest observed delivery-state counts; not actual delivery timestamps |
| candles | 15693853 | 758431884 | 2021-09-01 00:00:00 UTC | 2026-09-14 18:13:00 UTC | 5132 | observed_at <= simulated decision time | Closed Binance spot OHLCV; historical backfills are known only at ingestion time |
| derivatives | 287647 | 11833448 | 2026-09-09 05:00:00 UTC | 2026-09-14 18:00:00 UTC | 6 | observed_at <= simulated decision time | 15-minute venue OI, price and normalized funding; historical table omits native funding type |
| oi_5m | 168032 | 9406840 | 2026-09-11 11:35:00 UTC | 2026-09-14 18:10:00 UTC | 4 | observed_at <= simulated decision time | Observed Hyperliquid OI samples aggregated to closed five-minute buckets |
| coinbase_flow | 23252 | 1459337 | 2026-09-06 16:28:00 UTC | 2026-09-14 18:14:00 UTC | 9 | observed_at <= simulated decision time | Coinbase BTC/ETH executed aggressive volume in one-minute buckets |
| liquidations | 68567 | 7973089 | 2026-09-06 15:36:35 UTC | 2026-09-14 18:14:06 UTC | 9 | observed_at <= simulated decision time | Observed OKX events; coverage is not all liquidations on the exchange |
| stablecoins | 11165 | 220062 | 2026-09-09 05:00:00 UTC | 2026-09-14 18:10:00 UTC | 6 | observed_at <= simulated decision time | Observed stablecoin price and peg deviation |
| listings | 18035 | 1293474 | 2026-09-05 08:52:22 UTC | 2026-09-14 11:41:49 UTC | 10 | observed_at <= simulated decision time | Announcement revisions preserved as first observed locally |
| signals | 3025 | 1219822 | 2026-09-05 08:52:22 UTC | 2026-09-14 18:10:57 UTC | 10 | observed_at <= simulated decision time | Public signals and original evidence; no user identifiers |
| source_snapshots | 28272 | 164107165 | 2026-09-11 14:20:00 UTC | 2026-09-14 18:15:00 UTC | 77 | observed_at <= simulated decision time | Allowlisted local market snapshots and basket membership captured every five minutes |
| source_health | 25536 | 787360 | 2026-09-11 14:20:00 UTC | 2026-09-14 18:15:00 UTC | 4 | observed_at <= simulated decision time | Collector availability observed at archive time |
| research_snapshots | 3931 | 16391714 | 2026-09-11 14:00:00 UTC | 2026-09-14 18:00:00 UTC | 4 | observed_at <= simulated decision time | Versioned live Research calculations captured hourly, no synthetic past snapshots |
| signal_outcomes | 15120 | 855040 | 2026-09-05 08:52:22 UTC | 2026-09-14 17:45:26 UTC | 10 | observed_at <= simulated decision time | Post-recording spot reference returns and excursions, no fees or execution model |
| market_events_releases | 1805 | 6290681 | 2026-09-05 08:52:29 UTC | 2026-09-14 18:00:00 UTC | 10 | release published_at <= simulated decision time | Immutable customer-data releases; release time, row hashes and retrospective feature flags are preserved. |
Observation evidence
First 30 of 4 observations. JSON and CSV contain all published evidence. CSV dataset_section distinguishes summary and observation rows.
| dataset | record_key | event_time | observed_at | ingested_at | journal_captured_at | journal_kind | symbol | venue | sample_fields | calculation_version |
|---|---|---|---|---|---|---|---|---|---|---|
| candles | TRXUSDT:5m:1789403700.0 | 2026-09-14 16:35:00 UTC | 2026-09-14 16:45:09 UTC | 2026-09-14 16:45:09 UTC | 2026-09-14 16:46:18 UTC | first_journal_observation | TRXUSDT | Binance Spot | {"opened":1789403700,"closed":1789403999.999,"c":0.3412,"qv":317574.35708,"taker_buy_valid":1} | coinnudge-pit-coverage-2.1 |
| derivatives | Coinbase International:MEGA-PERP:1789403400 | 2026-09-14 16:30:00 UTC | 2026-09-14 16:46:19 UTC | 2026-09-14 16:46:19 UTC | 2026-09-14 16:46:20 UTC | first_journal_observation | MEGA | Coinbase International | {"venue":"Coinbase International","contract":"MEGA-PERP","funding_hourly_pct":-0.0063999999999999994,"oi_usd":8773.972271999999,"price":0.035496} | coinnudge-pit-coverage-2.1 |
| listings | mexc_listings:17827791538562:0ccb79fd42fa22c631427e2b6c65e13038e3dc5cbc5f415609879276dd36ad0b | 2026-09-14 11:41:49 UTC | 2026-09-14 16:46:28 UTC | 2026-09-14 16:46:28 UTC | 2026-09-14 16:46:31 UTC | first_journal_observation | BEM | mexc_listings | {"id":"mexc_listings:17827791538562","source":"mexc_listings","symbol":"BEM","status":"scheduled","url":"https://www.mexc.com/announcements/article/first-in-market-17827791538562","fingerprint":"48d76410e6ee4f1c9c49280c75e8b4cdf06ca8a270ae4c4025d93a997ae53ff5"} | coinnudge-pit-coverage-2.1 |
| source_snapshots | research_protocol_economics:1789404300 | 2026-09-14 16:45:00 UTC | 2026-09-14 16:46:32 UTC | 2026-09-14 16:46:32 UTC | 2026-09-14 16:46:32 UTC | first_journal_observation | research_protocol_economics | {"key":"research_protocol_economics"} | coinnudge-pit-coverage-2.1 |
Stored bytes equals the sum of compressed_bytes in the published coverage rows: catalog Parquet file lengths plus listed release files. It excludes revision journals, databases and unlisted archive files and is not disk allocation. Derivative catalog counts include all stored envelopes; the derivatives study separately selects its qualified 15-minute contract observations and can have a different cutoff. Archive row count is not proof of continuous coverage or point-in-time eligibility; backfills retain later ingestion times.
How to read this page
- Which CoinNudge datasets preserve what was known, observed and ingested at each point in time?
- Large row counts and old event dates do not prove continuous coverage, historical universe completeness or point-in-time eligibility.
- Inspect the table before relying on the summary.
- JSON and CSV preserve the published evidence.
What can this page tell you quickly?
- Question answered
- Which CoinNudge datasets preserve what was known, observed and ingested at each point in time?
- Measured scope
- Actual archive catalog coverage by data family
- Update schedule
- Recalculated about every 6 hours from stored source observations.
- Do not infer
- Large row counts and old event dates do not prove continuous coverage, historical universe completeness or point-in-time eligibility.
How does point-in-time data reduce look-ahead bias?
A market value can have an old event date but enter the database later through a backfill. If a backtest uses it as though it were available on the event date, future knowledge leaks into the simulation.
Filtering by feature availability and retaining revisions makes that mistake detectable. It does not remove survivorship bias or guarantee the asset universe was complete.
Why keep three timestamps?
Event time answers when the market or announcement event occurred. Observation time answers when CoinNudge first had the relevant source evidence. Ingestion time answers when the archive envelope was persisted or backfilled.
The three can be equal for live-collected records and very different for historical recovery. Collapsing them into one column destroys the distinction a research review needs.
What does the coverage table prove—and not prove?
It proves that stored partitions contain rows between published minimum and maximum event times. It helps scope an extract and exposes which data families are still young.
It does not prove every expected interval exists. Buyers should request gap counts, historical universe membership and feature-validity distributions for the exact study before purchase.
Which keys and units define one PIT envelope?
The archive record key identifies the source record inside a dataset. event_time, observed_at, ingested_at and journal_captured_at are UTC instants represented as Unix seconds in JSON and ISO 8601 on the page. Payload fields retain dataset-specific units, while the public sample_fields object shows the actual stored names.
Deduplication uses dataset plus record_key for the source observation. A new journal envelope records another capture; it is not automatically a semantic revision. Backfilled rows remain distinguishable because their observation or ingestion time is later than their source event time.
How can I verify and cite this result?
Download the page JSON for the complete snapshot and CSV for tabular evidence. Record the API dataset key `point-in-time-data-coverage`, the returned calculation_version, observation interval and named source alongside any quoted number. The visible table rounds values for readability; the downloads retain the published precision.
The stable URL remains unchanged when the data refreshes. A new calculation time does not prove every input changed, so compare source-observed timestamps and the stated actual archive catalog coverage by data family. Missing rows remain missing rather than becoming zero or a favorable estimate.
What is measured, and what is not?
| Measured claim | Evidence on this page | Boundary |
|---|---|---|
| Which CoinNudge datasets preserve what was known, observed and ingested at each point in time? | The current table and downloadable rows using calculation contract point-in-time-data-coverage. | Large row counts and old event dates do not prove continuous coverage, historical universe completeness or point-in-time eligibility. |
| The answer can be independently inspected. | Read the immutable archive catalog and active Market Events release catalog. For each populated family publish rows, compressed bytes, partitions and event-time bounds. Publish real sanitized archive-envelope samples with record key, event_time, observed_at, ingested_at and journal capture time. A backtest may use a feature only when its availability time is no later than the simulated decision; an old event date alone is never sufficient. Storage card=sum(compressed_bytes) over the displayed rows / 1073741824, rounded to two decimals. Includes listed Parquet and release file lengths; excludes revision journals, databases and unlisted archives. This is neither logical size nor disk allocation. Derivative catalog envelopes and the qualified 15-minute contract study have distinct scopes and cutoffs. | Coverage and missingness constrain every conclusion. |
Method and data boundary
Read the immutable archive catalog and active Market Events release catalog. For each populated family publish rows, compressed bytes, partitions and event-time bounds. Publish real sanitized archive-envelope samples with record key, event_time, observed_at, ingested_at and journal capture time. A backtest may use a feature only when its availability time is no later than the simulated decision; an old event date alone is never sufficient. Storage card=sum(compressed_bytes) over the displayed rows / 1073741824, rounded to two decimals. Includes listed Parquet and release file lengths; excludes revision journals, databases and unlisted archives. This is neither logical size nor disk allocation. Derivative catalog envelopes and the qualified 15-minute contract study have distinct scopes and cutoffs.
Large row counts and old event dates do not prove continuous coverage, historical universe completeness or point-in-time eligibility.
Sources and verification
- CoinNudge data servicesDataset families, status and delivery model.
- CoinNudge sampleEvent, context and outcome field examples.
- CoinNudge API documentationAuthentication, limits and timestamp fields.
- CoinNudge methodologyHistorical reconstruction boundaries.
Frequently asked questions
What is point-in-time crypto data?
It preserves what was available at each simulated decision time, including observation and revision boundaries.
Is every archived row PIT-eligible?
No. Backfilled rows can have old event times but later availability and must be filtered.
Does PIT data eliminate all backtest bias?
No. Universe selection, missing venues, execution assumptions and overfitting remain.
Can I download the whole archive?
Only datasets and scopes explicitly included in a plan or custom agreement are deliverable.
Why publish row counts?
They help screen feasibility, but gap and feature-validity reports are still required for a real study.