A profitable backtest is not a decision. It’s an application form. Something has to review that application before real money is on the line — and if the reviewer is the same person who tuned the parameters, with no fixed bar written down in advance, the review will pass everything you show it.
PRUVIQ solves this for its own presets with a fixed battery: every preset carrying the Verified badge had to be profitable across the full multi-year backtest window (PF ≥ 1.05) and hold up out-of-sample and survive a 3-window walk-forward on the live engine. This post lays that battery out as a checklist you can run on your own strategy — and is honest about which steps the product runs for you and which ones only you can do.
Step 0: costs in, or nothing else counts
Every check below is meaningless on a cost-free backtest, because costs are what kill most edges. Every PRUVIQ simulator run charges fees, slippage, and settlement-event funding automatically — the rates and models are documented on the methodology page, and the funding model has its own article, because most backtests get that part wrong. If your own backtest doesn’t charge realistic costs, fix that before running any item on this list.
Step 1: the bar — profitable over the full window, with margin
The entry requirement is the whole window, not a good stretch: PF ≥ 1.05 across the full backtest window — every candle in the measured range, currently reaching back to late 2023. The bar sits deliberately above 1.0 — 1.0 is breakeven, and a strategy that merely ties after costs has no margin for the ways live trading is worse than simulation.
PRUVIQ does: every preset card shows its full-window result, fees included; losing presets are published too, labeled Killed rather than deleted. You do: write your bar down before you run. A bar chosen after seeing the results is not a bar, it’s a rationalization.
Step 2: out-of-sample — a holdout you never touched
Split the data and keep a holdout out of every tuning decision. The battery holds out the most recent 30% of the window, and its bar is asymmetric on purpose: the out-of-sample PF must clear 1.05 while in-sample must clear 1.0 — the unseen data faces the stricter bar, because in-sample performance is partly a product of tuning and proves less.
PRUVIQ does: the simulator’s results panel has a Run Validation button — for presets and custom-built strategies alike — that reruns your exact configuration with a holdout split (default 30%), shows in-sample and out-of-sample metrics side by side, and reports an overfit-risk rating from the degradation between them. You do: freeze parameters before looking at the holdout. The tool can split the data; it cannot know how many times you already peeked. An honest holdout is a discipline, not a feature.
Step 3: walk-forward — every window must carry itself
An aggregate number can hide a strategy that made all its money in one regime and bled through the rest. The battery cuts the full data range into 3 calendar windows and requires each window to stand alone: PF ≥ 1.0 and at least 30 trades, per window. A window that fails is a fail; a window too thin to judge is also a fail — “not enough data to check” never counts as a pass.
PRUVIQ does: this battery runs server-side against the live engine, and the per-window numbers — dates, PF, trade counts — are published in a verification dossier on each verified strategy’s page, so the badge is never an unaccompanied claim. Honesty note: there is no one-click walk-forward button in the simulator UI. You do: replicate any window yourself with the simulator’s Test Period date controls — set the start and end dates to one segment, run, repeat. Three runs is cheap insurance against a one-regime wonder.
Step 4: sequence — resample the trades
Your equity curve is one draw from the trades you happened to get. The same Run Validation panel also runs a Monte Carlo pass — by default, 1,000 bootstrap reruns that resample your trade PnLs with replacement — and reports the 5th–95th percentile band, the probability of ending positive, and the worst-case drawdown across reruns. (Resampling, not just reshuffling: a pure reorder would keep the final return fixed, and the whole point is to see how different the outcome could have been.) The panel states its own limit: it assumes the historical trade distribution repeats, which real markets are not obligated to do.
You do: read the pessimistic band, not the mean. If the 5th percentile would make you abandon the strategy, you will abandon it live the first time variance runs against you.
Step 5: sample size and selection bias
Two traps survive every check above: ratios built on too few trades, and picking the best of many tries. We covered both — the low-sample flag and the Deflated Sharpe Ratio — in How to Read Sharpe, Sortino, and Max Drawdown; run this checklist with those two lenses on.
Step 6: decide what a fail means — before you run
This is the step most checklists omit. Our battery re-runs on a schedule against the live API, and a bar miss does not quietly update the published evidence: the failing numbers are not written into the dossier, the run turns red, and demotion of the badge on a battery failure is a human decision made looking at the failure. There is also a stricter automatic path: a separate daily job re-measures every card, and if a verified preset’s full-window PF decays below the 1.05 gate, the badge drops automatically. Either way, bad numbers are never dressed as green.
You do: the personal equivalent. Before running the checks, write down what happens on a fail — stop, retune from scratch, or shelve. Deciding after the fail is how a fail becomes a footnote.
What passing means — and what it doesn’t
Passing every item is a reproducibility statement, not a prophecy. Our own definition says it directly: verified is a reproducible result, not a live-deploy verdict. And the methodology page’s disclaimer applies to this entire checklist: past performance does not guarantee future results — backtests are simulations, not predictions. The badge track ends there: passing the battery earns the badge and the published dossier, nothing more. On our separate real-money research track, a passed backtest is only a ticket to paper forward-testing, never straight to size — we wrote up that stage in what we validate before risking a dollar.
The fastest way to internalize the checklist is to run it once on a real strategy: open the simulator, load any preset — verified or killed — and press Run Validation. The gap between the headline number and what survives the battery is the whole lesson.