← Research library
Research · 研究 · 135 · Methodology

What is point-in-time data? Why what you knew then is everything

21 Sept 20267 min readMethodologyShishin Research

This article explains a data discipline used to test quantitative strategies honestly. It is educational and general, not personalised investment advice, and not a recommendation to buy, sell, or hold any security.

The most dangerous number in any backtest is one that was true later but not yet true on the day the strategy claims to have used it. Point-in-time data is the practice of recording history as it actually stood on each date, including the values that were later revised away. What you knew then is the only honest basis for a decision the backtest says you made then.

The short version

Point-in-time (PIT) data is data as it stood on each historical date. That means the earnings figure as first reported, the index membership as of that day, the price that actually printed. Non-PIT data silently substitutes today’s corrected version, which leaks the future into the past. A PIT store keeps vintaged snapshots with effective dates, so a query answers only with what was visible on that day. It is the single strongest defence against look-ahead bias.

As-reported versus later-restated

Financial data is not fixed once it is published. A company reports quarterly earnings, and months later restates them after an audit, a reclassification, or an acquisition. A vendor corrects a bad price or a missed dividend. An index reshuffles its constituents. Each of these is a legitimate correction, and each one quietly rewrites the historical record so that the version you download today is not the version anyone could have seen at the time.

Point-in-time data preserves both. The as-reported value is what was first published on its effective date. The later-restated value is what the record was subsequently corrected to. A PIT query for a past date returns the as-reported version, because that is the only figure a decision-maker could have acted on. The restatement is real, but it belongs to the date it was filed, not to the date the original number was live. The same logic governs membership and prices: what the universe contained as-of a date, and what a stock traded at as-of a date, rather than as-known today.

How non-PIT data leaks the future

The leak is rarely a blunder. It is what happens by default when a pipeline pulls the cleanest, most current dataset it can find and reads it backwards. A few of the common channels:

  • Revised earnings. A screen ranks companies on a quarter’s reported profit. If the pipeline uses the restated figure, filed months after the quarter, it has ranked the past using information that did not yet exist. The strategy appears to have picked the winners; it merely read the answer key.
  • Delistings and the survivor effect. A universe built from names still listed today has erased the companies that died in between. The test never sees them, so it never loses on them. This is the survivor leak, and it is a data-vintaging failure at heart. It is treated in depth in survivorship bias.
  • Ticker reuse. Symbols are recycled. A ticker that belongs to one company today may have belonged to a different, now-dead company years ago. A pipeline keyed on the symbol rather than the entity as-of the date will splice two unrelated price histories into one, and the join looks seamless.
  • Index reconstitution. A backtest that ranks within an index using today’s membership list has handed itself the outcome: the names that were later added are usually the ones that performed. Membership must be read as-of each date. That specific case is covered in the hindsight universe.

In every case the mechanism is the same, and it has a single root cause: reading a cleaned current dataset backwards. A value that only became true later is presented to the strategy as though it were true on the day. The result is a beautiful curve that no operator could have traded. The failure mode itself, in full, is look-ahead bias; point-in-time data is the practice that closes the door on it.

How a point-in-time store is built

A PIT store does not overwrite. When a new value arrives, it is written as a fresh snapshot stamped with the date it became known, its effective date. The old snapshot is kept. Instead of one row per company that mutates over time, the store holds a stack of vintaged rows, each one a photograph of what the record said on the day it was taken.

A query is then always a query as-of a date. Ask “what was this company’s reported earnings as-of March” and the store returns the most recent snapshot whose effective date is on or before March, ignoring every correction filed afterwards. Ask what the tradable universe contained as-of a date and it returns the names that had begun trading and had not yet delisted. The discipline is simple to state and unforgiving in practice: a backtest may only see snapshots whose effective date precedes the decision it is simulating. Get the effective-date logic wrong by a single day and the leak returns.

Building this correctly is also what makes a backtest reproducible: because the as-of view of any past date is fixed, re-running the same rules over the same history returns the same result, rather than a subtly different one each time the underlying data is silently revised beneath it.

How Shishin stores its own history

Two of the inputs that matter most for a momentum and breakout system are price bars and the tradable universe, and both are kept as-of each date. A day’s bars are recorded with the open, high, low, close, and volume that printed on that day, so the raw print is never silently overwritten. Splits and dividends are a legitimate ongoing adjustment, and the honest way to carry them is as adjustment factors with their own effective dates, so a query can reconstruct either the raw print or the series adjusted as it stood on any past date. The universe is the set of symbols that were tradable on that date: a name that listed later does not appear in an earlier date’s universe, and a name that delisted appears with the prices it actually traded at, up to the session it stopped trading.

Every score the four engines produce is derived from that as-of view and only that view. A ranking for a past date is computed from the bars and universe that were visible on the day, never from a version of the record cleaned up by hindsight. This is the same principle that lets the public log at /verify mean something: a signal published on a date was derived from the data of that date, and both are committed at the time rather than reconstructed after the fact.

The honest limit: not all PIT is equal

Point-in-time discipline is easier to claim than to fully achieve, and the difficulty is not uniform across data types. Price and universe membership are the tractable cases. Prices are printed and timestamped as they trade, and membership changes on known dates, so a store that simply refuses to overwrite and stamps every snapshot can reconstruct the as-of view faithfully. This is the PIT we hold with confidence.

True point-in-time fundamentals are a different and far more expensive problem. Capturing the exact figure a company first reported, before every subsequent restatement, requires a vendor that archived each filing as it landed and never folded corrections back into the original. That data exists, but it is costly and comparatively rare, and much of the freely available fundamental history is restated by construction. The honest position is to be explicit about which PIT you actually hold. A system built on price and universe momentum can be genuinely point-in-time on its core inputs; a claim of fully vintaged fundamentals is a much stronger claim, and one that should be doubted until the vendor and the archival method are named. This is one reason we treat freely available (restated) fundamental data as a caution rather than a foundation, and lean on the inputs we can vintage cleanly.

Point-in-time data does not make a backtest true. It removes the most seductive way for a backtest to be false: pretending you knew then what you only learned later. Everything downstream, the whole argument for why most backtests lie, assumes you got this part right first.

Sources & further reading

  • López de Prado, M. (2018). Advances in Financial Machine Learning. Wiley. On data structuring, leakage, and the discipline of using only information available at the decision point.
  • Nagel, S. (2021). Machine Learning in Asset Pricing. Princeton University Press. On using only predictor data that was available in real time, to avoid look-ahead from restated values.
Related reading
MethodologyLook-ahead bias: the subtlest way a backtest lies6 min readMethodologyHow to vet a stock-signal track record: seven questions7 min readMethodologyInside Byakko: the defensive engine, by the numbers8 min read
Frequently asked

What is point-in-time data?

Point-in-time (PIT) data is data as it stood on each historical date: the earnings figure as first reported, the index membership as of that day, and the price that actually printed, rather than today's later-corrected version.

Why does non-point-in-time data leak the future?

The root cause is reading a cleaned current dataset backwards. When a pipeline uses restated earnings, today's index membership, or a survivor-only universe to score a past date, it hands the strategy information that did not yet exist on that date, producing a curve no operator could have traded.

How is a point-in-time store built?

It never overwrites. Each new value is written as a fresh snapshot stamped with the date it became known (its effective date), and old snapshots are kept. A query is always as-of a date and returns the most recent snapshot whose effective date is on or before that date.

How does Shishin store its own history?

Price bars and the tradable universe are kept as-of each date, and the raw print is never silently overwritten. Splits and dividends are carried as adjustment factors with their own effective dates, so scores for a past date are derived only from what was visible on that day.

Is fundamental data as easy to keep point-in-time as prices?

No. Price and universe membership are the tractable cases. True point-in-time fundamentals require a vendor that archived each filing as it landed and never folded corrections back, which is costly and rare, so much freely available fundamental history is restated by construction.