The combo finder — and why you can’t trust it
This searches every slice of the data for the best-looking edge. It’s a hypothesis generator, not a verdict — and here’s the honest math on why.
slices tested850
“significant” by luck alone~42at p<0.05
honest thresholdp<0.00006Bonferroni
slices that survive it0
Testing 850 slices, pure chance throws up ~42 that look “significant” at the usual p<0.05. So a single good p means nothing. Correcting for all the tries (Bonferroni), the bar is p<0.00006, and 0 slices clear it. Every row below is a hypothesis to test on fresh data, not a proven edge.
Most significant slices (min 50 trades)
| slice | trades | hit rate | return/trade | p | survives? |
|---|---|---|---|---|---|
| gpt-4.1-nano + short + Large | 218 | 63.3% | +0.78% | 0.0001 | — |
| gpt-4.1-nano + short + morning + Large | 86 | 70.9% | +1.38% | 0.0001 | — |
| gpt-4.1-nano + short + Large + no spam | 198 | 63.6% | +0.74% | 0.0001 | — |
| gpt-4o-mini + short + afternoon + Large + split | 139 | 33.8% | -1.50% | 0.0001 | — |
| gpt-4.1-nano + short + Large + all agree | 180 | 63.9% | +0.81% | 0.0002 | — |
| gpt-4o-mini + short + Large + all agree | 180 | 63.9% | +0.81% | 0.0002 | — |
| consensus + short + Large + all agree | 180 | 63.9% | +0.81% | 0.0002 | — |
| gpt-4.1-nano + short + morning + Large + all agree | 71 | 71.8% | +1.47% | 0.0002 | — |
| gpt-4o-mini + short + morning + Large + all agree | 71 | 71.8% | +1.47% | 0.0002 | — |
| consensus + short + morning + Large + all agree | 71 | 71.8% | +1.47% | 0.0002 | — |
| gpt-4o-mini + afternoon + Large + split | 333 | 39.9% | -0.76% | 0.0002 | — |
| gpt-4.1-nano + short + morning + Large + no spam | 76 | 71.1% | +1.43% | 0.0002 | — |
| gpt-4o-mini + short + afternoon + Large + split + no spam | 132 | 34.1% | -1.34% | 0.0003 | — |
| consensus + short + afternoon + Large + split | 127 | 33.9% | -1.49% | 0.0003 | — |
| gpt-4.1-nano + short + Large + all agree + no spam | 162 | 64.2% | +0.83% | 0.0003 | — |
Highest return/trade — the microcap noise trap
These look spectacular but are tiny, illiquid slices whose p-values say "noise."
| slice | trades | hit rate | return/trade | p |
|---|---|---|---|---|
| gpt-4o-mini + short + Unknown + no spam | 71 | 53.5% | +2.79% | 0.553 |
| consensus + short + Unknown + no spam | 65 | 52.3% | +2.46% | 0.710 |
| gpt-4o-mini + short + Unknown | 85 | 52.9% | +2.33% | 0.588 |
| gpt-4o-mini + short + Unknown + split | 51 | 54.9% | +2.23% | 0.484 |
| consensus + short + Unknown | 79 | 51.9% | +2.02% | 0.736 |
| gpt-4o-mini + morning + Unknown + split | 50 | 60.0% | +1.89% | 0.157 |
| gpt-4o-mini + Unknown + split + no spam | 109 | 50.5% | +1.83% | 0.924 |
| gpt-4o-mini + Unknown + split | 119 | 50.4% | +1.67% | 0.927 |
| consensus + morning + Unknown + split | 50 | 60.0% | +1.62% | 0.157 |
| gpt-4.1-nano + short + morning + Large + all agree + no spam | 61 | 72.1% | +1.55% | 0.001 |