做量化这几年,最坑我的错误几乎都不报错。代码照常跑完,图照常画出来,只是数字错了。
刚开始做策略时,我的流程很直接:有个想法,写代码,回测,调参数,上实盘,失效,然后从头再来。每个策略都是一份单独的脚本,上一个策略踩过的坑,下一个还会再踩一遍。
后来在 ATB,我把策略拆成六块:入场、出场、风险、仓位、止损止盈、执行时机。每一块单独写、单独测,一个策略就是一份 YAML 配置,写明这六块怎么组合。想换一种止损方式,改配置里的一段就行,别的代码不用碰。
最直接的变化是测得快了。以前一个月认真测几个想法,后来能测上百个。
测得越快,越容易骗自己
测一百个想法,总有几个回测曲线很漂亮。这几个里面哪些是真的,哪些只是刚好贴合了那段历史,光看曲线分不出来。
所以框架里后来最花功夫的,是验证那一层。参数只在一段数据上挑,再拿到它没见过的数据上跑,也就是 walk-forward。时间序列做交叉验证之前,先把训练集和测试集交界处的数据剔掉,不然信息会漏过去。最后用蒙特卡洛打乱交易顺序,看结果是不是只靠某一段走运的行情。
比这些方法更管用的是一条规矩:实验开始前先把通过标准写下来,没通过的版本全部存档,不能因为结果看着不错就临时放宽。有一个月我们提了三十多个改进想法,最后留下的只有几个。
两个不报错的坑
第一个是口径。
以前算年化夏普,我在代码里把 √252 写死了,可有的策略一年根本没有 252 个观测点。还有的策略是几条腿先合成日收益再算波动率,一部分波动就这样被抹平了。单看每一步都说得过去,叠在一起数字就偏高。后来我把记录过的策略按同一套口径全部重算,年化因子改成按实际观测频率算。好几个原来看着很强的策略,重算完都掉了一截。
第二个是默认参数。
有一次我在新的研究库里重跑一个老组合,想和旧结果逐笔对账。第一版里有一个策略只对上了一半多。查了很久才发现,那个策略的脚本默认用的是另一个时间周期,而组合调用它的时候我没显式传参数。它不报错,就安安静静地跑成了另一个版本。把参数补上、把指标热身要用的数据补齐以后,将近两千笔交易对上了 99% 以上。剩下几笔是浮点误差让手数取整不一样,要追平得逐笔复现旧库的资金曲线,我决定不追。
这类坑我在研究库里专门记了一份清单,现在已经有八个。我给自己定了两条规矩:参数一律显式写死,不靠默认值;新结果要和旧结果逐笔对上,对不上就不用。
有了 AI 以后
现在写策略代码、跑实验,很多活都交给 AI 了。我在做的研究系统里,模型每一轮自己写信号代码,系统打分,过了闸才留下。
实验变便宜以后,骗自己也跟着变便宜了。AI 不会替你判断一个结果是不是真的,所以我现在花时间最多的,是把闸定清楚:什么样的结果算数,提前写好,让机器照着执行。
回头看,这些规矩没有一条是灵光一现想出来的,都是被坑出来的。
Over the years I’ve done quant research, the mistakes that cost me the most almost never threw an error. The code ran to the end, the charts came out as usual, and the numbers were simply wrong.
When I started building strategies, my process was simple: get an idea, write the code, backtest, tune the parameters, go live, watch it stop working, and start over. Every strategy was its own script, so whatever trap the last one fell into, the next one fell into again.
Later, at ATB, I split a strategy into six parts: entry, exit, risk, position sizing, stops and take-profits, and execution timing. Each part was written and tested on its own, and a strategy became a YAML config that said how the six parts fit together. To try a different stop-loss, you changed one section of the config and didn’t touch any other code.
The most direct change was speed. I went from properly testing a few ideas a month to testing more than a hundred.
The faster you test, the easier it is to fool yourself
Test a hundred ideas and a few of them will always have beautiful backtest curves. Looking at the curves alone, you can’t tell which of those are real and which just happened to fit that stretch of history.
So the part of the framework that ended up taking the most work was the validation layer. Parameters are chosen on one slice of data and then run on data they’ve never seen, which is walk-forward testing. Before cross-validating a time series, you cut out the data around the boundary between the training and test sets, or information leaks across. Finally, a Monte Carlo run shuffles the order of the trades to see whether the result depends on one lucky stretch of the market.
One rule mattered more than any of these methods: write down the pass criteria before the experiment starts, archive every version that fails, and never loosen the bar on the spot because a result looks good. One month we came up with more than thirty ideas for improvements, and only a few of them survived.
Two traps that never threw an error
The first was how the numbers were defined.
When I annualized Sharpe ratios, I had hard-coded √252, but some strategies don’t have 252 observations in a year. Other strategies combined several legs into a daily return before computing volatility, which smoothed part of the volatility away. Each step was defensible on its own, but stacked together they pushed the number up. Later I recalculated every strategy on record with one consistent method, and changed the annualization factor to follow the actual observation frequency. Several strategies that had looked strong came down a fair bit after the recalculation.
The second was default parameters.
Once I re-ran an old portfolio in a new research library and tried to reconcile it trade by trade against the old results. In the first pass, one strategy matched only a little over half of its trades. It took a long time to find the cause: that strategy’s script defaulted to a different timeframe, and I hadn’t passed the parameter explicitly when the portfolio called it. It threw no error and just silently ran as a different version. After I added the parameter and filled in the data the indicators needed to warm up, more than 99% of nearly two thousand trades matched. The few that were left came from floating-point error rounding the position size differently. Closing that gap would have meant reproducing the old library’s equity curve trade by trade, so I decided not to.
I keep a list of traps like these in the research library, and it’s up to eight now. I also set myself two rules: always pass parameters explicitly instead of relying on defaults, and make new results match the old ones trade by trade, or don’t use them.
Once AI came along
A lot of the work of writing strategy code and running experiments now goes to AI. In the research system I’m building, the model writes the signal code itself each round, the system scores it, and it’s only kept if it passes the gate.
Once experiments got cheap, fooling yourself got cheap too. AI won’t decide for you whether a result is real, so what I spend the most time on now is defining the gates clearly: writing down in advance what kind of result counts, and letting the machine enforce it.
Looking back, none of these rules came from a flash of insight; I learned every one of them by getting burned.