← Research library
Research · 研究 · 42 · Methodology

How to vet a track record.

1 Sept 20267 min readMethodologyShishin Research

This article is an educational framework for evaluating a track record. It is not personalised investment advice and not a recommendation of any product or security.

Every signal service shows a track record. Almost none give you the means to check it. Before you trust a performance claim, ours or anyone’s, there are seven questions that separate a record worth believing from a marketing artifact. None requires a finance degree; all require the publisher to have done the honest work up front.

The seven questions

  • Is it survivorship-free? Were delisted, merged, and bankrupt names in the test universe, or only the companies that happen to still exist today? Testing on survivors alone inflates everything. See survivorship bias.
  • Is it reproducible? Re-run the locked test, does it land on the same number, to the cent? A record that drifts each time it is recomputed is a moving target, not a record.
  • Does it survive losing its best trades? Strip out the handful of biggest winners; is there still an edge? If the whole result rode on a few lucky names, it won’t repeat, the test is in leave out the winners.
  • Does it show the losers? A credible record publishes every trade, not a curated highlight reel. If you can only see the winners, you can’t see the strategy.
  • Is it point-in-time? Did every decision use only data available before the moment it was made, no look-ahead, no hindsight in the features or the labels? This is the most common silent way a backtest lies.
  • Is it net of costs? Are fees, slippage, and a realistic fill price baked in, or is it a gross, frictionless fantasy? Honest tests fill at prices you could have actually gotten.
  • Is it statistically significant, after the searching? One impressive backtest out of hundreds of variants tried is not evidence. A real result corrects for how many things were tested before one looked good (Harvey, Liu & Zhu, 2016; the deflated Sharpe ratio), the subject of statistical significance.

Backtest is a hypothesis; live is the test

Even a record that passes all seven is a hypothesis until it is run forward, transparently, on data that did not exist when it was built. The strongest evidence a service can offer is a backtest paired with a live, public track that behaves like it, which is the whole point of running what was tested.

Apply it to us

We built Shishin’s record to pass this list: a survivorship-free universe, reproducible to the dollar, every trade published, fills on the close, corrected for multiple testing, and paired with a live paper-traded track. Run the seven questions against the track record, and against anyone else’s.

Sources & further reading

  • Harvey, C. R., Liu, Y. & Zhu, H. (2016). “… and the Cross-Section of Expected Returns.” Review of Financial Studies, 29(1), 5 to 68.
  • Bailey, D. H. & López de Prado, M. (2014). “The Deflated Sharpe Ratio.” Journal of Portfolio Management, 40(5), 94 to 107.
  • Brown, S. J., Goetzmann, W. N., Ibbotson, R. G. & Ross, S. A. (1992). “Survivorship Bias in Performance Studies.” Review of Financial Studies, 5(4), 553 to 580.
Related reading
MethodologySurvivorship bias: what most backtests quietly leave out8 min readMethodologyInside Byakko: the engine that works when the market doesn't, by the numbers8 min readMethodologyInside Seiryū: the recovery engine that fires least and earns most per day7 min read
Frequently asked

How do you vet a stock-signal track record?

Ask seven questions: is it survivorship-free, reproducible to the dollar, does it survive removing its best trades, does it show every loser, is it point-in-time (no look-ahead), is it net of costs, and is it statistically significant after correcting for how many variants were tried?

What are the biggest red flags in a backtest?

Testing only on stocks that still exist (survivorship bias), results that change when re-run, showing only winners, look-ahead bias, ignoring fees and slippage, and one impressive result cherry-picked from hundreds of variants.

Is a backtest enough to trust a signal service?

No. A backtest is a hypothesis until it is run forward, transparently, on data that didn't exist when it was built. The strongest evidence is a backtest paired with a live, public track that behaves like it.