Does technical analysis work backtest audit showing search, costs, and unseen data

Does Technical Analysis Work? A Backtest Audit Before You Trade

📅 Originally Published: · Last Reviewed and Updated:

Does technical analysis work? A rule may earn a paper test, but not trust, until the search that produced it is disclosed, realistic costs are included, and the complete rule survives untouched data. The evidence is mixed in-sample and much weaker after those controls. Use the three gates below to decide whether to reject the rule, paper-test it, or consider a tightly limited allocation.

You have a rule with a smooth equity curve, a double-digit annualized return, and only a few ugly drawdowns. The platform labels the result profitable. The live decision is harder: should you trust it with money?

The chart cannot answer that question by itself. It does not show how many failed variants were tried before this one, which costs were left at zero, or whether the rule has ever faced a genuinely unseen period. Those omissions are where a convincing backtest can stop being useful.

This guide covers mechanical rules built from daily price or volume data, such as moving-average, breakout, momentum, and mean-reversion systems. It does not judge discretionary chart reading, intraday execution, high-frequency strategies, or market-making models. Readers evaluating a specific indicator can also use the broader Trading Edge research hub to compare the rule with related tests.

What Does the Research Actually Say?

Does technical analysis work in academic tests? The evidence is mixed in-sample and much weaker after researchers account for selection, implementation costs, and unseen data. That does not make every technical rule useless. It means the final equity curve is only the start of the evidence.

Sullivan, Timmermann, and White applied White’s Reality Check to a large technical-rule universe. Their historical sample still contained a best rule with evidence of superior performance after the adjustment, so the paper does not support a blanket “nothing works” conclusion. Its durable lesson is methodological: judge the winner against the search that produced it.

Bajgrowicz and Scaillet tested 7,846 daily Dow Jones rules from 1897 through 2011. Their False Discovery Rate framework found unreliable ex ante selection, while low transaction costs offset much of the reported economic value. Rink extended the question to 6,406 rules across 23 developed and 18 emerging markets. Many markets showed in-sample significance, but recent predictive ability and prior-window persistence were much weaker.

What the three core studies support, and what they do not.
Study Design Defensible takeaway Important limit
Sullivan, Timmermann & White (1999) Reality Check applied to a large Dow Jones rule universe The entire search process must be priced into the statistical test The best rule in their historical sample was not automatically eliminated
Bajgrowicz & Scaillet (2012) 7,846 rules, Dow Jones data, 1897-2011 Rule selection was unreliable, and low costs erased much of the economic performance Results concern simple daily index rules, not every form of technical analysis
Rink (2023) 6,406 rules, 41 equity markets, up to 66 years In-sample significance often failed to persist; positive costs worsened out-of-sample results Index-level daily rules do not cover intraday or market-microstructure strategies

📚 Primary sources: Sullivan, Timmermann & White (1999) · Bajgrowicz & Scaillet (2012) · Rink (2023)

Read the studies as complementary tests, not as one pooled experiment. They use different markets, periods, rule families, and statistical procedures, so their rule counts cannot be combined into a universal survival rate.

Gate 1: Was the Rule Chosen Before the Result?

The first gate asks whether the rule was specified before the performance was known. A backtest selected from hundreds or thousands of variants needs a multiple-testing correction. A genuinely pre-specified rule has a different evidence burden.

Suppose a researcher tests ten moving-average pairs, five stop-loss levels, four assets, three holding periods, and two start dates. That already creates 1,200 combinations. Reporting the best curve as though it came from one clean hypothesis hides the actual experiment.

This is the central data-snooping problem. The winning rule can be real, lucky, or a mixture of both. Looking only at the winner cannot tell you which. Reality Check, Superior Predictive Ability tests, False Discovery Rate controls, and careful pre-registration attack the same problem from different directions: they compare the reported winner with the search that produced it.

Evidence that clears Gate 1

  • The entry, exit, sizing, benchmark, and rebalance rules were written before the test.
  • The researcher discloses every material variant searched, including abandoned versions.
  • A multiple-testing method is used when more than one candidate rule was evaluated.
  • Parameter ranges were chosen for an economic reason instead of tuned around the best result.

Evidence that fails Gate 1

  • The seller shows one optimized curve and will not disclose the number of trials.
  • The start date, asset, or parameter window changes whenever the result weakens.
  • The rule contains vague conditions that can be interpreted differently after each chart is visible.
  • The reported p-value treats the winner as a single test even though a large universe was scanned.

At this stage, the indicator name matters less than the disclosed search. A transparent moving-average rule with weak results is easier to judge than a proprietary “AI signal” selected from an undisclosed universe.

The practical response is simple. Ask how many rules, parameter sets, assets, and dates were tested. An answer such as “this was the best setting” is not enough. The search count is part of the result.

Gate 2: Does the Rule Survive Real Trading Costs?

Does technical analysis work after costs? Implementation is the second gate. The backtest must include the costs created by its own trades, and those inputs must match the instrument, account, and execution method. There is no universal commission, slippage, spread, borrow cost, or tax rate.

TradingView’s official strategy documentation shows that commission and slippage are explicit strategy properties. A backtest that leaves either input at zero is assuming no commission or no slippage, not describing a real account.

📚 Platform documentation: TradingView Strategy Properties documents user-set commission and slippage, both defaulting to zero.

Realistic does not mean inserting the same 0.05% commission and one-tick slippage into every strategy. The estimate should match the instrument, broker, order type, trade size, time of day, and sample period. A liquid ETF traded monthly has a different cost profile from a small-cap stock traded at the open. Short strategies may also need borrow fees and availability constraints.

Taxes belong in a separate scenario. They depend on account type, jurisdiction, holding period, loss harvesting, cost basis, and the investor’s other gains and losses. A taxable investor and an IRA holder can run the same rule and face different after-tax outcomes, so one fixed annual drag cannot represent every strategy or investor.

A cost model that can be defended

  1. Count actual entries and exits, including reversals and partial fills.
  2. Apply broker fees and observed spread or slippage by trade.
  3. Include financing and borrow costs when the rule uses leverage or shorts.
  4. Run taxable and tax-advantaged scenarios separately when taxes matter.
  5. Show gross and net performance side by side so the cost contribution remains visible.
Use your own execution history when possible. A median difference between the signal price and the actual fill is more useful than a generic slippage assumption copied from another market.

For readers using platform defaults, the related guide on TradingView settings for investors shows where zero-cost assumptions and inherited indicator settings enter the workflow. Its specific dollar model should still be judged on its own assumptions.

Gate 3: Does It Work on Unseen Data?

Does technical analysis work on unseen data? The third gate tests the complete rule on data that played no role in choosing it. This is where many polished in-sample results lose their claim to a repeatable edge.

A proper holdout is more than a later date range opened after the first test disappoints. The researcher decides the training period, validation method, benchmark, and pass threshold before looking at the holdout result. Once the holdout influences another parameter change, it has become part of the training process.

Rink’s multi-market study makes this point concrete. The paper found significant in-sample outperformance in 13 of 23 developed markets and 14 of 18 emerging markets. Yet predictive ability declined over time. Almost no developed markets remained predictable in the last two subperiods beginning in 2002, and almost all emerging markets were unpredictable in 2009-2016. Portfolios built from the best rules in the prior three-year window did not significantly beat buy-and-hold at zero costs; with positive transaction costs, they underperformed most of the time.

That result does not prove a new rule will fail. It shows why the in-sample chart is incomplete evidence. Walk-forward tests, rolling re-estimation, cross-market validation, and a final untouched holdout can each reveal a different form of fragility.

Questions for the unseen-data test

  • Was the holdout period chosen before the researcher saw its performance?
  • Were all parameters frozen before the holdout was opened?
  • Does the benchmark match the asset, risk, and time in the market?
  • Does performance survive more than one regime, market, or reasonable start date?
  • Were costs recalculated for the holdout rather than copied from the training sample?

A rule that works only in one market or regime can still be useful when that limit is explicit and the allocation is designed around it. The problem begins when a narrow result is sold as a general answer.

A Reproducible 50/200-Day Test

A simple reproduction shows how the three gates change the question. This test does not estimate the universal value of technical analysis. It checks one fully specified moving-average rule on one pinned daily SPY dataset.

The rule holds SPY when its 50-day simple moving average is above its 200-day simple moving average. The signal is lagged by one trading day to avoid using today’s closing information for today’s position. The sample runs from January 4, 2010 through December 30, 2019 and contains 2,515 daily observations.

Reproduction of a 50/200-day moving-average rule, 2010-2019.
Path Final multiple Annualized return Maximum drawdown
Buy and hold 3.460× 13.24% -19.35%
Crossover, before costs 2.314× 8.76% -19.18%
Crossover, 25 bps per position change 2.262× 8.52% -19.98%
Crossover, plus illustrative 100 bps annual tax drag while invested 2.089× 7.66% -20.40%

The signal changed position nine times and was invested for 79.7% of the observations. In this window, the gross crossover lagged buy-and-hold before transaction costs. Most of the difference came from missed market exposure rather than the 25-basis-point cost assumption. Adding the illustrative tax scenario widened the gap, but that tax input is a sensitivity assumption, not a measured average for every account.

Reproduction details

Data: pinned daily SPY CSV, filtered to 2010-01-04 through 2019-12-30. File SHA-256: f1682f176f9db69ab654405b0126b5c9840f272ee12c6c5bd5c0fec2345fd6f0.

Signal: 1 if SMA(50) > SMA(200), else 0, shifted one trading day.

Costs: 25 basis points on each position change. The separate tax sensitivity subtracts 100 basis points annually while invested.

Limits: one ETF proxy, one decade, no parameter search correction, no live fill data, and no claim of future performance. See the sitewide calculation methodology.

📚 Pinned dataset: SPY daily CSV at fixed Git commit. The project backtest runner and catalog were executed on July 19, 2026.

This example cannot answer whether a different rule, asset, or regime would succeed. It shows why “Does technical analysis work?” must become a testable question with a named rule, benchmark, period, and cost model. For a deeper look at crossover-specific turnover, see SMA versus EMA crossover. For a visual selection-bias example, see trendline survivorship bias.

Decision Router: Reject, Paper-Test, or Allocate?

The correct action depends on which gate failed. A failed backtest does not need another indicator layered on top. It needs either cleaner evidence or a smaller claim.

Action after the three-gate backtest audit.
Evidence state Action Reason
Search universe hidden, rule changed after results, or no credible benchmark Reject The reported curve cannot distinguish selection skill from luck
Rule is fixed and testable, but costs or holdout evidence are incomplete Paper-test Forward observation can reveal fills, turnover, and behavioral pressure without risking capital
Pre-specified rule, disclosed search, realistic costs, repeated unseen-data evidence Consider a limited allocation The evidence is stronger, but model and regime risk remain
Rule only works before tax or in a specific account Route by account Implementation, not the chart, determines whether the edge reaches the investor

Before a limited allocation

  • Write the maximum position size and acceptable drawdown before trading.
  • Define the exact condition that would retire the rule.
  • Keep a benchmark account or shadow portfolio for comparison.
  • Record actual fills and costs instead of continuing to use backtest estimates.
  • Do not use leverage merely to restore the return lost by a low-exposure rule.

A paper test is useful only if the rules remain frozen. Changing the parameters after every loss creates a new in-sample optimization in real time. The investor may feel adaptive while quietly destroying the evidence they were trying to collect.

The broad question becomes useful only after the claim is narrowed. Which rule? Which asset? Which benchmark? Which costs? Which unseen period? Without those fields, the evidence is too vague to support an allocation.

Technical Analysis FAQ

Does technical analysis work for every market?

No. The evidence varies by market, period, rule family, benchmark, and cost model. Rink found in-sample significance in many markets but much weaker recent and out-of-sample persistence. A result from a developed-market index cannot automatically be transferred to a single stock, cryptocurrency, futures market, or intraday strategy.

Does technical analysis work better when commissions are zero?

Zero commissions remove only one cost. Spreads, slippage, missed exposure, market impact, borrow costs, financing, and taxes may remain. A low-turnover rule in a liquid ETF can be inexpensive to execute, while a frequent strategy in a thin market can lose a meaningful share of gross performance.

Does technical analysis work after one profitable out-of-sample test?

It is stronger evidence, but one holdout can still be lucky. Look for frozen rules, multiple regimes or markets, realistic costs, a fair benchmark, and results that are not driven by one brief episode. Once the holdout is used to tune the rule, a new untouched period is needed.

Does technical analysis work with a short out-of-sample test?

There is no universal five-year rule. The window should contain enough independent trades and relevant regimes to test the claim. A slow monthly strategy may need a much longer period than a daily strategy, while a structurally changed market may make older data less representative. Choose the plan before viewing the result.

Does technical analysis work as a risk-control overlay?

Yes, but evaluate the stated job. A rule designed to reduce drawdown may accept a lower long-run return or more tracking error. Compare it with a benchmark that reflects that objective instead of grading every rule only on raw return. Position sizing and diversification can sometimes achieve the same risk goal with less timing dependence.

The Bottom Line

Does technical analysis work? Some rule sets have looked credible in-sample, and technical methods can serve narrow risk or execution purposes. A polished backtest still does not deserve capital until selection, implementation, and unseen-data performance have been tested separately.

Reject a rule when the search is hidden. Paper-test it when the rule is fixed but real costs or forward evidence are missing. Consider capital only after the rule survives all three gates and the position size can tolerate model failure. That decision standard is stricter than a green equity curve, which is exactly why it is useful.

What would make you reject a backtest immediately? Start with the missing field: undisclosed trials, zero costs, no holdout, or an unfair benchmark.

Sources, Method & Evidence

  • FOUNDATIONAL Sullivan, Timmermann & White (1999): White’s Reality Check applied to a large technical-rule universe; used here for the multiple-testing framework, not as proof that every rule fails.
  • FOUNDATIONAL Bajgrowicz & Scaillet (2012): 7,846 daily Dow Jones rules over 1897-2011; supports the claims about unreliable ex ante selection and the effect of low transaction costs.
  • CONFIRMATORY Rink (2023): 6,406 rules across 41 markets; supports the recent-period and out-of-sample persistence discussion.
  • SUPPORTING TradingView documentation: verifies that commission and slippage are user-set properties with zero defaults.

Original analysis: TheFinSense executed catalog backtest bt-003 with the unchanged project runner against a pinned SPY CSV. An offline adapter supplied the hash-verified local copy because the execution environment blocks direct network access. Gross, transaction-cost, and illustrative tax scenarios are displayed separately.

Interpretation control: Published study results are kept separate from TheFinSense assumptions. No article-created friction or wealth-gap model is presented as an academic finding.

Evidence boundary: The reproduced moving-average result is descriptive for one rule and one decade. It is not causal evidence, a forecast, or a universal estimate of technical-analysis performance.

📋 Update History
  • May 2, 2026: Original article published.
  • July 19, 2026: Replaced unsupported article-created friction and wealth-gap models with source-bounded interpretations and a reproducible moving-average test; restructured the article around three decision gates.
  • July 20, 2026: Tightened the research section, consolidated correction history into the method and update sections, and added the new featured image.

Educational quantitative analysis based on published data. Not investment, tax, or legal advice. Consult a licensed professional before acting on any calculation. About TheFinSense.