Most crypto backtests are worthless. They show astronomical returns because they’re full of hidden bugs that would never survive live trading. Here’s how to do it right.
What Is Backtesting?
Backtesting means testing a trading strategy against historical market data to see how it would have performed. You define rules (when to buy, when to sell, position size, stop-loss) and simulate executing those rules on past price data.
Why it matters: Without backtesting, you’re gambling. With proper backtesting, you have statistical evidence for whether a strategy has an edge.
Why most backtests lie: The gap between a backtest and live trading is enormous. Our own Momentum LONG strategy backtested at +400% but showed negative expectancy after fixing a candle index bug. The difference? Bugs that only show up under rigorous validation.
Types of Backtesting
Not all backtests are created equal. Before diving into the rules, understand the three main approaches:
Walk-Forward Analysis: Split your data into sequential windows. Optimize on window 1, test on window 2. Then optimize on windows 1+2, test on window 3. This mimics real-world conditions where you only have past data to work with. It is the closest thing to simulating how your strategy would evolve over time with periodic re-optimization.
Monte Carlo Simulation: Randomize the order of your trades (or slightly vary parameters) and run thousands of iterations. If your strategy only works with trades in the exact historical order, it is fragile. Monte Carlo tells you the probability distribution of outcomes — not just the single historical path. A strategy with a 95th percentile max drawdown of 40% is very different from one with a 95th percentile max drawdown of 15%, even if their average returns are identical.
Out-of-Sample Testing: The most basic and most important. The split we actually use is 2024 training / 2025 validation / 2026 out-of-sample (the SL/TP optimization guide owns that procedure). Test on a period the model has never seen. If performance degrades significantly, you have overfit. This is non-negotiable — any backtest result that does not include out-of-sample validation is essentially worthless.
Each approach catches different failure modes. Walk-forward catches regime-dependent overfitting. Monte Carlo catches sequence-dependent fragility. Out-of-sample catches plain overfitting. Use all three if you want confidence in your results.
The 5 Critical Rules of Honest Backtesting
Rule 1: Only Use Completed Candles
This is the most common and most dangerous mistake in crypto backtesting.
When your bot runs at 10:01 UTC, the 9:00-10:00 candle is complete (confirmed data). The 10:00-11:00 candle is still forming (unknown data).
Wrong approach: Using the current candle’s volume, close price, or indicators as entry conditions. In a backtest, the “current” candle is already complete — but in live trading, you only have 1 minute of data.
Right approach: All signal conditions must use the previous (completed) candle. Fill at the next candle’s open. The close of the signal candle does not exist yet at decision time, so it cannot be the entry price — our engine fills at the open of the bar after the signal.
# Correct: Use prev candle for signals, curr for entry price
signal = prev_candle['bb_squeeze'] == True and prev_candle['volume_ratio'] > 2.0
entry_price = curr_candle['close']
# Wrong: Using current candle data (look-ahead bias)
signal = curr_candle['bb_squeeze'] == True # This data doesn't exist yet!
Rule 2: Match Backtest Logic to Live Logic Exactly
Your backtest code and live trading code must produce identical signals given the same data. Any difference — even a single index offset — can completely change results.
Our lesson: We once had prev vs prev2 candle comparison in our backtest that differed from live code. The backtest showed -20.6% loss. When we fixed it to match live logic exactly, it showed +$794 profit. One index difference. Completely opposite results.
How to verify:
- Extract the signal function from your live code
- Use that exact function in your backtest
- Run both on the same data and compare signals
Rule 3: Include All Costs
A strategy showing +3% average return per trade becomes unprofitable when you account for:
- Trading fees: 0.02% maker / 0.05% taker on OKX USDT perpetuals — a market-order round trip costs 0.1% (taker both sides)
- Slippage: Market orders don’t fill at the displayed price. Our engine charges three liquidity tiers per fill — 0.05% (top 50), 0.10% (top 200), 0.20% (everything else) — applied to entry and exit, so a round trip costs twice that
- Funding rates: Perpetuals settle funding every 8 hours. Our source is the BTC settlement history we collected (Binance USDT-M perpetual BTCUSDT, 3,002 settlements, 2023-12-31 → 2026-09-27). Over it the average was 0.0064%/8h and the highest single settlement was 0.0881% — 0.1% was never once recorded (the fee guide owns that table)
- Spread: Low-liquidity altcoins can have 0.1-0.5% bid-ask spreads
Minimum cost assumption: 0.20% per round trip on top-50 coins at today’s OKX rate — taker 0.05% × 2 = 0.1%, slippage 0.05% × 2 = 0.10%. (This slot used to say 0.15%: fees were counted both ways but slippage only one way.)
Rule 4: Test on Enough Data
A strategy that works on 6 months of data proves nothing. Markets have regimes — bull, bear, sideways, high volatility, low volatility. Your backtest needs to cover multiple regimes.
Our standard at the time (archived pre-OKX run, data 2023-12 → 2026-02): 535 coins, 2+ years of hourly data (Binance USDT perpetual futures). That gave us 2,898 trades across multiple market conditions. Today the same bar is expressed per coin: the engine backtests each coin over the price history we have collected for it — which, for coins listed before our collection began, starts after their listing.
Out-of-sample validation: Split your data into training (optimize parameters) and testing (validate results). If it works on both, it’s more likely real. If it only works on training data, you’ve overfit.
Rule 5: Use Realistic Position Sizing
Don’t calculate returns as percentages and add them up. Simulate actual capital allocation.
# Wrong: Simple percentage addition (fantasy)
total_return = sum(trade_returns) # Shows +2,090%
# Right: Simulate actual account balance
balance = 10000 # Starting capital
for trade in trades:
position_size = min(200, balance * 0.02) # $200 or 2% of balance
pnl = position_size * trade_return * leverage
balance += pnl
# Shows +103% (reality)
The difference? Simple addition ignores that losing trades reduce your capital for future trades. Realistic simulation shows what your actual P&L would be.
The Thresholds We Actually Enforce
Instead of inventing a “healthy range”, here are the values this site actually judges with. Every row names the code that owns it:
| Metric | Threshold | What it means | Owner |
|---|---|---|---|
| Profit factor | 1.5+ / 1.0+ / below 1.0 | 1.0 is break-even (gross profit = gross loss) — a definition, not a number we picked from a measurement | src/utils/format.ts profitFactorTier |
| Win rate | 55%+ / 50%+ / below 50% | Where the break-even win rate is known, we use that instead — win rate alone does not decide good from bad | same file, winRateTier |
| Sample size | under 30 trades | The result sits in the noise band, so it is not a performance estimate | backend/api/schemas.py build_reliability_checks |
| Sample size | under 100 trades | Rankings carry a “low sample” warning | LOW_SAMPLE_MIN_TRADES |
| Leverage × stop-loss | 90 or above | Most trades end in liquidation | same file, leverage axis |
| Coins used | 1 | That coin’s regime, not a diversified edge | same file, diversification axis |
What we removed, and why. This slot used to hold a “healthy range / red flag” table — win rate 50-70%, profit factor 1.5-3.0, max drawdown 10-30%, Sharpe 1.0-3.0, 500+ trades, with red flags at >80%, >5.0 and >50%. None of those numbers have a source inside our system: they are not in the code, not in any measurement of ours, and we judge nothing by those ranges. Max drawdown and Sharpe ratio have no threshold at all, which is why they have no row above — adding one would mean inventing it again. Unsourced thresholds get quoted more the more table-shaped they are, and verified less.
Case study — BB Squeeze SHORT (since killed): 68.6% win rate, 2.22 profit factor, 2,898 trades across 535 coins (archived run, data 2023-12 → 2026-02). Those numbers survived out-of-sample validation across 2024, 2025, and 2026 data — and a fresh out-of-sample re-test on 2026-06-28 still rejected it, so we killed it (strategy journal). Even an honest backtest is an upper bound, not a promise.
Common Backtesting Mistakes
The five traps below also have a standalone page showing how our engine handles each one: Why Backtests Fail.
Overfitting
Adding more and more conditions to improve backtest results. Each added condition makes the strategy fit historical data better but generalizes worse to new data.
Test: If removing one condition destroys your results, you’re probably overfit. A robust strategy should survive small parameter changes.
Survivorship Bias
Only testing on coins that still exist today. Coins that got delisted (often after crashing 99%) are excluded, making results look better than reality.
Fix: Include delisted coins in your dataset, or at least acknowledge this limitation.
Ignoring Market Regime
A strategy that works in a bear market may fail completely in a bull market. We tested 4 different BTC regime filters to see if adapting to market conditions helps. All 4 failed — the overhead of missed trades outweighed the loss prevention.
Lesson: Sometimes the best filter is no filter. Let the data decide.
Common Backtesting Mistakes (Case Studies)
The mistakes above are abstract. Here are three concrete examples of how they destroy real money.
Case Study 1: Look-Ahead Bias — A Costly Lesson
Our team built a Momentum LONG strategy that backtested at +400% returns over 18 months. The equity curve was smooth, the Sharpe ratio was above 3.0, and the drawdowns were manageable. Capital was allocated for live trading.
The problem: the signal function was reading the current candle’s volume spike to trigger entries. In a backtest, that candle is already complete — the volume spike is confirmed. In live trading, the candle has only been open for 1 minute when the bot runs. The “volume spike” does not exist yet. After fixing the candle index bug and retesting, the strategy showed negative expectancy. Significant capital had already been lost in live trading before catching it. This is why Rule 1 (completed candles only) is non-negotiable. One index offset turned a +400% winner into a money-losing strategy.
Case Study 2: Survivorship Bias — Testing Only Top 50 Coins
A common shortcut is to backtest on today’s top 50 coins by market cap. The logic seems sound — these are the most liquid, most traded assets. But here is the problem: coins in today’s top 50 were selected because they succeeded. You are excluding every coin that crashed 95% and got delisted — LUNA, FTT, dozens of DeFi tokens from 2021.
A strategy that goes long on momentum signals will show inflated returns when tested only on survivors. This is why our BB Squeeze SHORT numbers (win rate 68.6%, profit factor 2.22 in that era’s run) were measured on all 535 coins including delisted ones — the universe keeps its casualties, so survivorship bias cannot quietly inflate the result. If you size positions off survivor-only metrics, you are taking on more risk than the data actually supports.
Case Study 3: Overfitting — The 25-Parameter Strategy
Picture a trader (a composite, not a specific person) building an elaborate mean-reversion strategy with 25 adjustable parameters: RSI period, RSI overbought threshold, RSI oversold threshold, Bollinger Band period, BB standard deviations, MACD fast/slow/signal periods, volume filter window, volume multiplier, ATR period, ATR multiplier for stop-loss, trailing stop activation, trailing stop distance, time-of-day filter start, time-of-day filter end, day-of-week filter, minimum spread filter, maximum position hold time, re-entry cooldown, and seven more.
In this illustration the backtest shows a 92% win rate and a profit factor above 8.0 — and live it loses money in the first week. With 25 parameters and 2 years of hourly data, an optimizer has enough degrees of freedom to fit noise perfectly: the strategy memorizes history instead of capturing an edge. Strip it to 4 core parameters (BB period, BB deviation, SL%, TP%) and the shiny win rate drops — but robustness rises. Fewer parameters, tested on more data, beats a complex model every time. (This scenario is a teaching model, not a logged experiment.)
How PRUVIQ Handles These Issues
Every strategy on PRUVIQ goes through a validation pipeline designed to catch exactly these problems.
Monte Carlo Validation: We do not report a single backtest equity curve. We run 1,000 Monte Carlo iterations that bootstrap the trade sequence (reshuffling trade order) to see how fragile the equity path is. If the result only looks good with the exact historical sequence, that fragility shows up here. You can see this in action on the PRUVIQ Simulator.
Out-of-Sample Testing: Every strategy is trained on one time period and validated on a completely separate period (the same 2024 / 2025 / 2026 split as above). Our BB Squeeze SHORT showed a profit-factor edge in both windows — IS 1.29 / OOS 1.08. The fresh out-of-sample re-test on 2026-06-28 rejected it on three adversarial axes in choppy regimes, so we killed it: that edge was bear-beta, not directional skill. It was re-classified conditional on 2026-08-17, and the owner of its current state is the journal. Passing out-of-sample is necessary, not sufficient. Strategies that degrade — in OOS or live — are marked as killed, not hidden.
Real Fee Modeling: We charge a flat 0.05% per side on entry and exit (the OKX USDT-perpetual base taker rate — the simulator’s coin universe is drawn from OKX’s USDT-SWAP listings) plus tiered slippage (0.05–0.20% by liquidity) and a funding-rate cost assumption. No “zero-cost” fantasies. The fee page shows the full schedule, the date we last checked it against OKX, and what we earn on referrals. (Disclosure: the OKX referral link on that page is an affiliate link — we earn a commission when you sign up through it.)
Killed Strategies on Display: most strategy configurations we test fail validation and are publicly documented with full data (11 killed presets ship on the simulator today). We do not hide failures — they are the proof that our validation process actually works. A platform that only shows winners is a platform that is not testing honestly.
Checklist: Before You Trust a Backtest
Use this 10-point checklist before risking real money on any backtested strategy — yours or someone else’s.
- Completed candles only: All signal conditions use the previous (closed) candle. No current-candle data in entry logic.
- Code parity: Backtest signal function is identical to live trading signal function. Verified by running both on the same dataset and comparing outputs.
- Realistic costs: Fees, slippage, spread, and funding rates are included. Total round-trip cost is at least 0.20% for liquid futures (check that fees and slippage are charged on both the entry and the exit).
- Sufficient data: Minimum 2 years covering bull, bear, and sideways regimes. At least 30 trades in the sample (below that is the noise band), and 100+ to stand in a ranking without a warning — the same values the code above owns.
- Out-of-sample validation: Strategy tested on data it was never optimized on. Performance does not degrade more than 20% versus in-sample.
- Survivorship bias addressed: Dataset includes delisted or crashed coins, or the limitation is explicitly acknowledged.
- Parameter count under control: Fewer than 8 free parameters for a simple strategy. Each parameter has economic justification.
- Monte Carlo survival: Strategy maintains positive expectancy across 1,000+ randomized iterations.
- Realistic position sizing: Simulated with actual capital allocation, not simple percentage addition.
- No cherry-picked timeframe: Results are not dependent on starting or ending on a specific date.
If any of these fail, the backtest is not trustworthy. Go back, fix the issue, and retest.
FAQ
How many trades do I need for a statistically significant backtest?
The thresholds our code actually enforces are 30 trades (below that the result is in the noise band and we do not treat it as a performance estimate) and 100 trades (the line where rankings carry a “low sample” warning) — the same values the table above names. The “at minimum, 500 trades” that used to sit in this answer was not one of our thresholds and had no source, so it is gone. Our BB Squeeze SHORT strategy used 2,898 trades across 535 coins — and it still failed live. Sample size protects against noise, not regime change.
Can I backtest on TradingView?
TradingView’s Strategy Tester is a good starting point for quick validation, but it has significant limitations. It does not model realistic slippage, uses simplified fee structures, and makes it easy to accidentally use current-candle data (look-ahead bias). For serious validation, export your logic to Python and run it with proper cost modeling. Use TradingView for idea generation, not for final validation.
How often should I re-optimize my strategy parameters?
Re-optimization is a double-edged sword. Too frequent (weekly) and you are curve-fitting to recent noise. Too rare (never) and your strategy may drift as market microstructure changes. Our approach: re-validate quarterly using walk-forward analysis. If out-of-sample performance drops below 50% of in-sample performance, investigate. But do not change parameters just because last month was bad — that is how you turn a working strategy into an overfit one.
What is the difference between backtesting and paper trading?
Backtesting runs your strategy on historical data — it is fast (minutes to hours) and covers years of market conditions. Paper trading runs your strategy in real-time with simulated money. Paper trading catches bugs that backtesting misses: API latency, order fill issues, rate limits, exchange downtime. You need both. Backtest first to filter out bad strategies quickly, then paper trade survivors for at least 2-4 weeks before risking real capital. Try the PRUVIQ Simulator to run backtests without writing code.
Getting Started
- Get data: Download OHLCV (Open, High, Low, Close, Volume) candle data from your exchange’s API
- Define rules: Write clear, unambiguous entry and exit conditions
- Simulate: Run through historical data candle by candle, tracking positions and P&L
- Validate: Split data into training/testing periods. Check for look-ahead bias
- Go small: If results survive validation, test with minimum position size on live markets
The gap between backtest and live trading is where most strategies die. Honest backtesting narrows that gap.
At PRUVIQ, strategies that fail stay published. See our killed strategies and version history for the full transparency trail.