A backtest that flatters you is worse than no backtest, because it replaces “I do not know” with false confidence and a position size to match. This is a diagnostic checklist. Run your harness through it before you let any number it produces influence what you risk.
These are the checks we run on any harness before its output is allowed to influence position size. They take minutes and they routinely fail.
This is the fastest and most damning test there is. Take your stop, change it materially, and re-run the identical seed. The dollar outcome must change. If it does not, the stop is not reaching the exit logic and every risk figure in that report is decoration.
Run it across a few thousand simulated months and watch the dollar distribution, not the ratios. The failure mode to look for is a report where every dollar figure holds steady — same average balance, same worst month — while the R-denominated numbers improve. That is not a better system. R is defined as the stop distance, so shrinking the stop shrinks the divisor and flatters every ratio while the money stays exactly where it was.
Search the source or the settings for a hit rate, a win probability, or a directional accuracy dial. If one exists and the report also prints a win rate, you are reading your own assumption with extra steps.
The honest arrangement is that direction accuracy is an input — it has to come from somewhere — but the win rate is a result, because getting the direction right and making money are not the same event. A correct call that never gives you an entry is not a win.
Any decent harness can tell you how each trade ended. If almost everything exits at a target and almost nothing stops out, ask whether the price generator can produce the trade that fails. A model calibrated to a gentle average move will rarely manufacture the violent one that takes you out, and the result is a distribution with the left tail quietly missing.
Mid-price fills are the single most common way a backtest invents returns. You do not get the mid. You get the bid when you sell and the ask when you buy, and on a short-dated contract that gap is a meaningful fraction of the premium.
Worse is an asymmetric fill assumption, and it is easy to introduce without noticing. A target is a resting limit order and fills exactly; a stop is a market exit and fills at the bid. Price both at the midpoint and the harness is honest about wins and optimistic about losses. Nothing in the output looks wrong, because the bias is structural and points one way.
A backtest that assumes every signal becomes a position overstates everything. Real entry logic refuses trades: the setup never confirms, the spread is too wide, no contract passes the screen, liquidity never shows. In our own measurements a large share of simulated days produce no trade at all, and any harness that silently converts every signal into a fill is describing a system nobody is running.
If the strategy resolves inside a single session and the data is daily, the harness is interpolating the part that decides the outcome. That is not automatically useless, but it must be labelled. Treat modelled sub-second behaviour as mechanics under assumed dynamics, never as measurement.
Not throw it away. An execution simulator that fails the fill test is still useful for checking that your machinery does not deadlock, double-fill, or leave orders working. Just stop using it to size positions.
The rule worth adopting: use the harness for structure — does the logic behave, is it sensitive to the things it should be sensitive to — and use arithmetic for exposure. How much of the account a single loss costs does not need a simulator, and that is the number that decides whether you survive a bad week.
Overfitting and a mis-wired harness look identical from the outside. Separate them by changing a risk parameter and re-running the same seed. If the dollar result does not move at all, the problem is not overfitting but that the parameter never reaches the exit logic.
The usual cause is fill assumptions. Mid-price fills hand you the full bid-ask spread on every round trip. An asymmetric assumption is worse again: if targets fill exactly as limit orders but stops fill at the market, your wins are priced honestly and your losses are not.
Only after you have confirmed the harness responds to the stop in dollars. R is defined as the stop distance, so tightening the stop shrinks the divisor and inflates every R figure even when the money is unchanged. Check the dollars first, then read the R.
Structure. Whether the logic behaves consistently, whether it is sensitive to the things it should be sensitive to, and whether the mechanics deadlock or misfire. Exposure, meaning what one loss costs the account, is arithmetic and does not need a simulator at all.
Options Sniper is free and ships with no strategy at all. The N‑T PRO Logic Pack is the alternative to building your own: a licensed parameter set plus the daily research signal, resolved each morning and held only in memory.
See the N‑T PRO Logic PackRelated reading: Options backtesting software compared · Why a backtest says the strategy works · The N-T PRO Logic Pack