Does technical analysis work backtest audit showing search, costs, and unseen data

Does Technical Analysis Work? A Backtest Audit Before You Trade

📅 Originally Published: · Last Reviewed and Updated:

Does technical analysis work? A rule may earn a paper test, but not trust, until the search that produced it is disclosed, realistic costs are included, and the complete rule survives untouched data. The evidence is mixed in-sample and much weaker after those controls. Use the three gates below to decide whether to reject the rule, paper-test it, or consider a tightly limited allocation.

You have a rule with a smooth equity curve, a strong annualized return, and only a few ugly drawdowns. The platform labels the result profitable. The live decision is harder: should you trust it with money?

The chart cannot answer that question by itself. It does not show how many failed variants were tried before this one, which costs were left at zero, or whether the rule has ever faced a truly unseen period. Those omissions are where a convincing backtest can stop being useful.

This guide covers rules built from daily price or volume data, such as moving average, breakout, momentum, and mean-reversion systems. It does not judge judgment-based chart reading, intraday trading, high-frequency systems, or market-making. Readers testing a specific signal can also use the broader Trading Edge research hub to compare the rule with related tests.

What Does the Research Actually Say?

What do academic tests show? Results can look stronger in-sample and weaken after researchers account for rule selection, trading costs, and unseen data. That pattern does not make every technical rule useless; it means the final equity curve is only the start of the evidence.

Sullivan, Timmermann, and White applied White’s Reality Check to a large technical-rule universe. Their historical sample still contained a best rule with evidence of superior performance after the adjustment, so the paper does not support a blanket “nothing works” conclusion. Its durable lesson is about method: judge the winner against the search that produced it.

Bajgrowicz and Scaillet tested 7,846 daily Dow Jones rules from 1897 through 2011. Their False Discovery Rate test found that future winners were hard to select in advance, while even low trading costs erased much of the reported value. Rink extended the question to 6,406 rules across 23 developed and 18 emerging markets. Many markets showed strong in-sample results, but recent forecasting power and carryover from earlier winning rules were much weaker.

What the three core studies support, and what they do not.
Study Design Defensible takeaway Important limit
Sullivan, Timmermann & White (1999) Reality Check applied to a large Dow Jones rule universe The entire search process must be priced into the statistical test The best rule in their historical sample was not automatically eliminated
Bajgrowicz & Scaillet (2012) 7,846 rules, Dow Jones data, 1897-2011 Rule selection was unreliable, and low costs erased much of the economic performance Results concern simple daily index rules, not every form of technical analysis
Rink (2023) 6,406 rules, 41 equity markets, up to 66 years In-sample significance often failed to persist; positive costs worsened out-of-sample results Index-level daily rules do not cover intraday or market-microstructure strategies

📚 Primary sources: Sullivan, Timmermann & White (1999) · Bajgrowicz & Scaillet (2012) · Rink (2023)

Read the studies as separate tests, not as one pooled experiment. They use different markets, periods, rule families, and methods, so their rule counts cannot be combined into one general survival rate.

Before the three gates, check the data and trade timing. Reject a test that uses future data, drops failed or delisted assets from the sample, mishandles splits or dividends, relies on stale prices, or assumes trades that could not have occurred. The three gates below start only after those basic checks are clean.

Gate 1: Was the Rule Chosen Before the Result?

The first gate asks whether the rule was fixed before its performance was known. If hundreds or thousands of versions were tried, the test must account for those extra chances to find a winner. A rule written down in advance starts with a cleaner evidence burden.

Suppose a researcher tests ten moving-average pairs, five stop-loss levels, four assets, three holding periods, and two start dates. The search now contains twelve hundred combinations, so reporting only the best curve hides the size of the test that produced it.

This is the core data-snooping problem: the winning rule can reflect skill, luck, or both. Looking only at the winner cannot separate them. Reality Check, Superior Predictive Ability tests, False Discovery Rate controls, and rules fixed in advance all address the same issue from different angles by judging the winner against the search that produced it.

Writing down a famous rule before your own test does not erase the older search that may have made it famous. A clean test today can reduce your own search bias, but it cannot prove that the rule first became popular without publication or selection effects.

Evidence that clears Gate 1

  • The entry, exit, sizing, benchmark, and rebalance rules were written before the test.
  • The researcher lists every material version tested, including discarded versions.
  • The analysis adjusts for extra trials when more than one candidate rule was tested.
  • Setting ranges were chosen for a clear market or trading reason instead of tuned around the best result.

Evidence that fails Gate 1

  • The seller shows one optimized curve but will not state how many trials were run.
  • The start date, asset, or setting window changes whenever the result weakens.
  • The rule contains vague conditions that can be changed after each chart is visible.
  • The reported p-value treats the winner as a single test even though a large universe was scanned.

At this stage, the signal name matters less than the disclosed search. A plain moving-average rule with weak results is easier to judge than a proprietary “AI signal” chosen from a hidden universe.

Ask how many rules, settings, assets, and dates were tested, because “this was the best setting” is not enough. The size of the search is part of the result.

Gate 2: Does the Rule Survive Real Trading Costs?

The second gate asks what happens after trading costs. The backtest must include the costs created by its own trades, using inputs that fit the instrument, account, order type, and venue. No single commission, slippage, spread, borrow cost, or tax rate fits every rule.

TradingView’s official Pine strategy docs state that a strategy applies no commission when commission settings are omitted and that the default slippage setting is zero. A zero input therefore models zero commission or zero slippage; whether that assumption is realistic must be justified for the rule, instrument, and venue.

📚 Platform docs: TradingView Pine strategy docs states the no-commission behavior when commission arguments are omitted and the zero default for slippage.

A realistic model does not mean inserting the same 0.05% commission and one-tick slippage into every strategy. The estimate should match the instrument, broker, order type, trade size, time of day, and sample period. A liquid ETF traded monthly has a different cost profile from a small-cap stock traded at the open. Short strategies may also need borrow fees and proof that shares were available to borrow.

Taxes need their own scenario because the result depends on account type, tax location, holding period, loss harvesting, cost basis, and the investor’s other gains and losses. A taxable investor and an IRA holder can run the same rule and face different after-tax outcomes, so one fixed annual drag cannot represent every strategy or investor.

A cost model that can be defended

  1. Count actual entries and exits, including reversals and partial fills.
  2. Apply broker fees and observed spread or slippage by trade.
  3. Include financing and borrow costs when the rule uses leverage or shorts.
  4. Run taxable and tax-advantaged cases side by side when taxes matter.
  5. Show gross and net performance side by side so the cost contribution remains visible.
Use your own trade history when possible. The typical gap between the signal price and your actual fill is more useful than a generic slippage input copied from another market.

For readers using platform defaults, the related guide on TradingView settings for investors shows where zero-cost inputs and inherited signal settings enter the workflow. Its specific dollar model should still be judged on its own assumptions.

Gate 3: Does It Work on Unseen Data?

The third gate tests the complete rule on data that played no role in choosing it. Many polished in-sample results weaken once the rule faces a period it never saw.

A proper holdout is more than a later date range opened after the first test disappoints. The researcher decides the training period, validation method, benchmark, and pass threshold before looking at the holdout result. Once the holdout influences another parameter change, it has become part of the training process.

Rink tested this across many markets. The paper found significant in-sample outperformance in 13 of 23 developed markets and 14 of 18 emerging markets. The forecasting power also declined over time. Almost no developed markets remained predictable in the last two subperiods beginning in 2002, and almost all emerging markets were unpredictable in 2009-2016. Portfolios built from the best rules in the prior three-year window did not clearly beat buy-and-hold at zero costs; with positive trading costs, they lagged most of the time.

That result does not prove a new rule will fail, but it shows why an in-sample chart is incomplete evidence. Walk-forward tests, rolling refits, checks across markets, and a final untouched holdout can expose different weak points.

Questions for the unseen-data test

  • Was the holdout period chosen before the researcher saw its performance?
  • Were all parameters frozen before the holdout was opened?
  • Does the benchmark match the asset, risk, and time in the market?
  • Does performance survive more than one regime, market, or reasonable start date?
  • Were costs updated for the holdout rather than copied from the training sample?

A rule that works only in one market or regime can still be useful when that limit is explicit and the allocation is designed around it. The problem begins when a narrow result is sold as a general answer.

A Reproducible 50/200-Day Test

The SPDR S&P 500 ETF Trust (SPY) test below shows how the three gates change the question. It does not estimate the general value of technical analysis; it checks one fixed moving-average crossover rule on one pinned daily SPY file.

The rule holds SPY when its 50-day simple moving average (SMA) is above its 200-day SMA. The signal is lagged by one trading day to avoid using today’s closing information for today’s position. The pinned file runs from January 4, 2010 through December 30, 2019 and contains 2,515 daily observations.

Results from the pinned SPY crossover test, 2010-2019.
Path Final multiple Annualized return Maximum drawdown
Buy and hold 3.460× 13.24% -19.35%
Crossover, before costs 2.314× 8.76% -19.18%
Crossover, 25 bps per position change 2.262× 8.52% -19.98%
Crossover, plus illustrative 100 bps annual drag while invested (sensitivity case) 2.089× 7.66% -20.40%

The signal changed position nine times and was invested for 79.7% of the observations. In this window, the gross crossover lagged buy-and-hold before transaction costs. Most of the difference came from missed market exposure rather than the 25-basis-point cost assumption. The final row adds an extra drag case, not a tax model. It does not simulate realized gains, holding periods, tax lots, losses, or rates for a specific account.

Line chart comparing growth of one dollar in SPY buy-and-hold with a 50/200-day moving-average crossover before costs, after 25 basis points per position change, and after an additional 100-basis-point annual drag, with the 2010 warm-up period shaded.
Buy-and-hold finished at 3.460× versus 2.314× for the gross crossover. The shaded opening period shows the long-average warm-up: the strategy remained in cash until October 19, 2010 while buy-and-hold was invested from day one. The 100-bp line is an illustrative sensitivity case, not a personal tax model.

Method at a glance: one fixed moving-average crossover, a one-day signal lag, the pinned SPY file, and 25 basis points per position change. Full data, warm-up, and return-field limits are documented below.

📚 Pinned dataset: SPY daily CSV at fixed Git commit. The file hash and all four displayed paths were checked again on August 21, 2026; the current catalog runner also matched the 2.089× sensitivity path on the same file.

This test cannot answer whether a different rule, asset, or market period would succeed. It shows why the broad question must become a testable claim with a named rule, benchmark, period, data treatment, and cost model. For a deeper look at crossover-specific turnover, see SMA versus EMA crossover. For a visual example of selection bias, see trendline survivorship bias.

Test details

Data: pinned daily SPY comma-separated values (CSV) file, filtered to 2010-01-04 through 2019-12-30. File hash: f1682f176f9db69ab654405b0126b5c9840f272ee12c6c5bd5c0fec2345fd6f0. The test uses the source file’s close field as supplied. The data source does not explain that field’s past adjustment treatment well enough to call these multiples a clean raw-price or total-return series. Treat the outputs as tied to this file.

Signal: 1 if SMA(50) > SMA(200), else 0, shifted one trading day.

Warm-up: because the pinned file starts with the test period, the 200-day average first exists on October 18, 2010 and the one-day-lagged signal can first trade on October 19. The fixed runner keeps the strategy in cash during the initial warm-up while buy-and-hold is invested from day one. The table matches that runner, but the two paths do not share the same invested start date.

Costs: 25 basis points on each position change. The separate 100-basis-point annual drag is applied only while invested as a sensitivity case; it is not a personal tax model.

Limits: one ETF proxy, one decade, an unclear source-price adjustment, a warm-up mismatch, no search correction inside this test, no live fill data, and no claim about future performance. See the sitewide calculation methodology.

Decision Path: Reject, Paper-Test, or Allocate?

What you do next depends on which gate failed. A weak backtest does not need another signal layered on top; it needs cleaner evidence or a smaller claim.

Action after the three-gate backtest audit.
Evidence state Action Reason
Search universe hidden, rule changed after results, or no credible benchmark Reject The reported curve cannot distinguish selection skill from luck
Rule is fixed and testable, but costs or holdout evidence are incomplete Paper-test A forward paper test can reveal signal behavior, turnover, and operating problems without risking capital; fills remain simulated
Rule fixed in advance, disclosed search, realistic costs, repeated unseen-data evidence Consider a limited allocation The evidence is stronger, but model and regime risk remain
Rule only works before tax or in a specific account Route by account Trading conditions, not the chart alone, determine whether the edge reaches the investor

FREE WORKSHEET

Backtest Audit Worksheet

Run your own backtest through the data-integrity precheck and the same three gates used in this article. The worksheet does not turn the checks into a score; unresolved critical items narrow the claim or stop the test.


Download the worksheet

1-page PDF · Print-friendly

Paper testing cannot confirm real slippage, market impact, partial fills, queue position, or whether shares will be available to borrow. Those require live trades, and real-money behavior can also differ from a simulated account.

Before a limited allocation

  • Write the maximum position size and acceptable drawdown before trading.
  • Define in advance the exact condition that would retire the rule from use.
  • Keep a benchmark account or shadow portfolio for comparison.
  • Record actual fills and costs instead of continuing to use backtest estimates.
  • Do not use leverage merely to restore the return lost by a low-exposure rule.

A paper test is useful only if the rules remain frozen. Changing the settings after every loss creates a new in-sample search in real time. That may feel adaptive, but it destroys the forward evidence the test was meant to collect.

The broad question becomes useful only after the claim is narrowed to a rule, asset, benchmark, cost model, and unseen period. Without those fields, the evidence is too vague to support an allocation.

Technical Analysis FAQ

Does technical analysis work for every market?

No. The evidence varies by market, period, rule family, benchmark, and cost model. Rink found strong in-sample results in many markets but much weaker recent and out-of-sample carryover. A result from a developed-market index cannot automatically be transferred to a single stock, cryptocurrency, futures market, or intraday strategy.

Does technical analysis work better when commissions are zero?

Zero commissions remove only one cost. Spreads, slippage, missed exposure, market impact, borrow costs, financing, and taxes may remain. A low-turnover rule in a liquid ETF can be inexpensive to execute, while a frequent strategy in a thin market can lose a meaningful share of gross performance.

Does technical analysis work after one profitable out-of-sample test?

It is stronger evidence, but one holdout can still be lucky. Look for frozen rules, multiple regimes or markets, realistic costs, a fair benchmark, and results that are not driven by one brief episode. Once the holdout is used to tune the rule, a new untouched period is needed.

Does technical analysis work with a short out-of-sample test?

There is no fixed five-year rule for every strategy. The window should contain enough independent trades and relevant regimes to test the claim. A slow monthly strategy may need a much longer period than a daily strategy, while a structurally changed market may make older data less representative. Set the test plan before viewing the result, then keep it fixed.

Can technical analysis be used as a risk-control overlay?

It can be tested for that job, but the evidence reviewed here does not establish that technical overlays reduce risk in general. A drawdown-control rule may accept lower long-run return or more tracking error, so compare it with a benchmark and risk measure that match the stated objective. Position sizing or diversification may sometimes address the same risk with less timing dependence.

The Bottom Line

Does technical analysis work? Some rule sets have looked credible in-sample, and some technical methods can be tested for narrow risk or trading objectives. The studies reviewed here do not establish a general risk-control benefit. A polished backtest still does not deserve capital until selection, trading costs, and unseen-data performance have been tested separately.

Reject a rule when the search is hidden. Paper-test it when the rule is fixed but real costs or forward evidence are missing. Consider capital only after the rule survives all three gates and the position size can tolerate model failure. Before risking capital, write down which gate remains open and what evidence would close it.

YOUR TURN

What would make you reject a backtest immediately: undisclosed trials, unrealistic costs, no untouched data, or an unfair benchmark?

Sources, Method & Evidence

  • FOUNDATIONAL Sullivan, Timmermann & White (1999): White’s Reality Check applied to a large technical-rule universe; used for the multiple-testing framework, not as proof that every rule fails. Journal record.
  • FOUNDATIONAL Bajgrowicz & Scaillet (2012): 7,846 daily Dow Jones rules over 1897-2011; supports the claims about unreliable ex ante selection and the effect of low transaction costs. Journal record.
  • CONFIRMATORY Rink (2023): 6,406 rules across 23 developed and 18 emerging markets; supports the recent-period and out-of-sample persistence discussion. Open-access article.
  • SUPPORTING TradingView Pine documentation: supports the narrower statement that omitted commission settings apply no commission and the default slippage argument is zero. Official documentation.

Original test: TheFinSense uses the pinned SPY file and the fixed crossover rule described above. The file hash and the displayed gross, cost, and sensitivity paths were checked against the same file. The initial warm-up and the unclear source-price adjustment are disclosed because both affect how the result should be read.

Interpretation: Published study results stay separate from TheFinSense assumptions. The 100-basis-point annual drag is a sensitivity case, not a measured tax estimate, and no article-created cost or wealth-gap model is presented as an academic finding.

Evidence boundary: The reproduced crossover result describes one rule, one pinned file, and one decade. It is not causal evidence, a forecast, a general risk-control result, or a universal estimate of technical-analysis performance.

Update history

  • v2.2
    2026-08-21
    FACT / METHOD

    Rebound the TradingView default-zero claim to the official Pine strategy documentation; clarified the SPY warm-up and return-field limits; narrowed paper-test and risk-control claims; added a data-integrity precheck; and normalized the bottom trust markup.

  • v2.1
    2026-07-20
    EDITORIAL

    Tightened the research section, consolidated correction history into the method and update sections, and added the revised featured image.

  • v2.0
    2026-07-19
    CORRECTION

    Replaced unsupported article-created friction and wealth-gap models with source-bounded interpretations and a reproducible moving-average test; restructured the article around three decision gates.

  • v1.0
    2026-05-02
    PUBLISH

    Original publication.