Green Wick · The Signal Lab
The goal: use AI to surface correlations humans miss, distil them into a reliable, frequently-firing, tradable signal, prove it the disciplined way — backtest → private forward-proof → public track record — and only then charge for it. This page is the live workbench where we narrow an impossibly wide search down to the one thing we actually ship.
Yes — but only the disciplined, narrow version.
The category is proven: quant funds, paid signal groups, alt-data shops and prediction-market models all sell predictions for real money. We don't even need to beat the market for our own book — we need a track record convincing enough that people subscribe. That is a slightly lower bar than running a fund, but it is reputationally unforgiving.
The reasons this is hard are exactly the reasons ours will be different.
We don't search all of finance from zero. We refine the edges our own systems already hint at.
We've attacked "find signal" three times. They're the same hunt at different resolutions — and this one is the refinery.
Many agents collaborate → emergent signal. Wide, decentralized.
Many strategies × dynamic on-chain data → ensemble signal. Live, noisy, partial wins.
Distil ONE legible, proven signal. The refinery for the others' hints.
Define the target before searching, or every pretty backtest looks like a winner.
A signal isn't "it went up after X." It is a claim that clears six gates at once. We write these down first so the search has a finish line.
| Gate | The bar | Why |
|---|---|---|
| Reliability | Positive edge per fire, net of costs, with a hit-rate or expected-value that holds out-of-sample. | Subscribers feel losers fast; the edge must survive honesty. |
| Frequency | Fires often enough to matter — think weekly-ish, not once a year. | Rare fires = no statistical proof and no reason to subscribe. |
| Tradability | The instrument is liquid and executable; edge > fees + slippage + impact. | A "signal" you can't act on net-positive is a chart, not a product. |
| Persistence | Survives walk-forward, multiple regimes, and a year of decay. | One lucky window is the overfit trap wearing a suit. |
| Capacity | Enough room that a group acting on it doesn't kill it. | Subscribers are the crowding risk. |
| Mechanism bonus | A plausible reason it exists (forced flows, behaviour, structure). | A "why" is the single best defence against overfitting. |
The scoreboard we'll rank candidates on
The breadth is enormous. Decompose it into axes so we can narrow deliberately, not randomly.
Every candidate signal is one point across five axes. Naming the axes turns "anything in any market" into a finite map we can prune.
Absolute direction is the hardest and most-mined target. Relative value (A vs B, market-neutral) and volatility are structurally more predictable — we bias here.
| Market | Inefficiency | Our data edge | Verdict |
|---|---|---|---|
| Crypto on-chain / DEX | High (retail, 24/7, manipulated) | Strong — multichain flow infra | hunt here |
| Prediction markets (Polymarket/Kalshi) | High (thin, news-driven) | Strong — realtime infra exists | hunt here |
| Crypto social / sentiment | Medium-high | Strong — alphalens | hunt here |
| Equities / FX / commodities | Low (efficient, covered) | Weak — costly data, no latency | deprioritize |
"Out of crypto" sounds like diversification but is actually harder for us: more efficient markets, less rare data. The asset-agnostic ambition is right; the realistic first hunting ground is crypto-native + prediction markets where our data is rare.
Most-mined, weakest residual edge. Use as a feature, never the thesis.
On-chain flows, social velocity, search/news, funding/basis. Where rare data still pays.
One market reliably leads another (BTC→alts, majors→long-tail).
Order flow, funding dislocations, forced rebalances/expiries.
We cannot win sub-second (no co-location, no HFT stack). We hunt hours → weeks, where the edge is in what we see and how we reason, not how fast we click.
Behavioural (retail over/under-reaction) · Structural (forced flows, expiries, rebalances) · Informational (alt-data others ignore) · Risk-premium (paid to hold a risk). A candidate with no identifiable type is probably noise.
The intersection of three circles — that's where we dig first.
Retail-driven, under-covered, new, manipulated — where edges still exist.
Multichain on-chain flows, mempool, DEX/CEX basis, social sentiment, prediction-market books.
Hours-to-weeks, so we never race infrastructure we don't have.
This isn't a section of the method. It IS the method. Break these and we ship noise.
Eight gates. Most candidates die early and cheaply. That's the point.
Concrete, mechanism-based starting bets. Each is a falsifiable claim, not a vibe.
Seed list — to be tested, killed, or promoted through the funnel. A hypothesis earns a row only if it names a mechanism, the data we already hold, and the exact target.
| # | Hypothesis | Mechanism / why | Data we hold | Target |
|---|---|---|---|---|
| H1 | Stablecoin exchange-inflow surges precede short-horizon majors downside. | Coins moving to CEX = intent to sell (structural/behavioural). | On-chain flow infra | BTC/ETH direction, 6–48h |
| H2 | Smart-money wallet accumulation clusters precede token outperformance. | Informed flow leads price (informational). | Wallet trackers, smart-money poller | Token vs basket, days |
| H3 | Funding/basis dislocations (CEX-perp vs DEX-spot) mean-revert. | Crowded leverage gets paid to unwind (structural). | Multichain price + funding | Spread reversion, hours |
| H4 | Social sentiment-velocity spikes lead short-horizon moves. | Attention precedes flow (behavioural) — must filter manipulation. | alphalens TG/X analyzer | Direction/vol, hours |
| H5 | Polymarket price vs news-velocity diverges, then converges. | Thin books lag breaking information (informational). | Polymarket realtime infra | Event price convergence |
| H6 | BTC regime shift → alt-rotation timing is predictable. | Capital rotates majors→long-tail on a lag (cross-asset lead-lag). | Multichain price feeds | Alt-basket relative, days |
| H7 | Disclosed insider/political trades (Congress, Form 4, 13F whales) — test for RESIDUAL edge after the obvious copy-trade decayed. | Informed actors trade on non-public edge; mandatory disclosure exposes the footprint. Obvious version is crowded (see graveyard) — residual edge, if any, is in faster ingestion or less-covered filings (informational). | Free public filings (STOCK Act, SEC EDGAR) | Named equity, days–weeks |
None of these is "the signal" — they are the first batch to run through §05. Expect most to die at Stage 4. The survivors are what we forward-proof.
Thousands publicly claim signals work; almost none publish an out-of-sample test. We harvest the claim and run our gauntlet — the kills feed the graveyard, and once in a while a real one survives. Full catalog in the lab; the priority queue:
| # | Public claim | Arena | Skeptic prior | Data |
|---|---|---|---|---|
| H8 | Overnight vs intraday — the equity premium accrues overnight, not in the session | equities | robust, widely replicated | have |
| H10 | Pre-FOMC drift — S&P ~+49bps in the 24h before FOMC | equities | decayed post-2015 — a live decay test | have |
| H11 | Funding-rate extremes → mean reversion — crowded leverage unwinds | crypto perps | structurally real; strongest of the set | free (Binance) |
| H12 | MVRV / NUPL bands mark cycle tops & bottoms | crypto | real but fires ~yearly → frequency fail | free |
| H13 | Fear & Greed contrarian — fade extreme greed | crypto | weak, widely known | free |
The product angle: publishing "we tested famous signal X and it's dead" is rarer and more trustworthy than another channel shouting buys.
The honest record. Every screen logged, especially the ones that died.
Engine: signal-lab/ — free daily Binance data (12 assets, ~1000 days), a reusable harness with the §04 discipline baked in (time-ordered train/OOS split, cost-real, separate in-sample vs out-of-sample metrics, a persistent multiple-testing counter). A signal earns a forward-test only by clearing out-of-sample.
Pre-registered: sign of BTC's k-day return predicts the alt basket outperforming BTC the next day (relative-value target, 10bps cost, 60/40 train/OOS). Tested a pre-set family k∈{1,3,7} — no cherry-picking.
| k | In-sample Sharpe | OUT-OF-SAMPLE Sharpe | OOS hit-rate | OOS bps/fire |
|---|---|---|---|---|
| 1 | −0.36 | −0.74 | 0.470 | −6.92 |
| 3 headline | +0.82 | −0.59 | 0.495 | −5.46 |
| 7 | +0.24 | −0.43 | 0.500 | −4.01 |
Instead of a new backtest, we read what the Monad wallet-bot already proved — its own out-of-sample validation (172k in-sample vs 74k OOS buys). Which on-chain feature actually survives OOS?
Run deliberately in parallel with the Monad mine, in a different arena, so we don't lock into one thing. Dollar-neutral long-losers / short-winners on 15 large caps, 5-day lookback, cost-real.
The textbook crypto claim: very positive funding = crowded longs → fade for the reversion. First pass on 66 days looked like a clear winner (OOS Sharpe +1.76). We didn't trust it. Pulling the full 3.5 years (3,150 trades) killed it: Sharpe −0.32, −39%, negative every year since 2023.
Screens 05–49 ran the rest of the gauntlet — funding, seasonality, momentum/low-vol/lottery factors, macro & FOMC, valuation bands, vol structure — and produced zero deployable edges; one real-but-unfit lead (the BTC variance-risk premium) sits parked, not sold. The kills are catalogued in the graveyard. The screen that unlocked the order-flow class, in full:
The 49-screen verdict said free price-only data is mined out — the next frontier needs a new data class. The wall was softer than we thought: Binance klines quietly ship taker-buy volume, the share of each bar initiated by aggressive market BUYS. That is genuine order flow (who crossed the spread), free, back to 2017, daily and hourly. Pre-registered continuation thesis: an abnormally buy-heavy bar marks urgent/informed demand → positive next bar (the academic order-flow result — Chordia; Cont et al.).
| Cut | 2017-19 | 2019-22 | 2022-24 | 2024-26 |
|---|---|---|---|---|
| BTC daily long-only (Sharpe) | +1.08 | −0.23 | −0.85 | −0.54 |
| BTC 1h top-decile → fwd 1h (bps · t) | +1.4 | +4.2 · t 2.5 | +0.1 | −2.6 · t −2.6 |
Sixteen distinct alpha hypotheses, each pre-registered, implemented and backtested in parallel: stablecoin depeg reversion, BTC-dominance rotation, the weekend effect, funding extremes, volume breakouts, OI squeezes, vol-regime filters, BTC/ETH pairs, settlement-hour drift, an SMA control, RSI bounces, gap fades, regime-conditioned momentum, equity lead-lag, and a funding-carry basket. Result: 15 graveyard, 1 promising, 0 deployable — a realistic hit rate. Most died the same way: a real micro-effect smaller than realistic taker cost (the maker-vs-taker trap), or an effect that's noise across years, assets and out-of-sample. The headstones are in the graveyard; the one survivor, in full:
Pre-registered: a big DOWN 1h BTC candle × a ≥2× volume spike (trailing-168h baseline, no lookahead) → next-hour mean-reversion. The interaction is the novel bit: with the volume spike the bounce is +30.2bps (t=3.51) vs +13.2bps without — a real +17bps conditional interaction effect. But the honest ledger: in-sample +49bps degrades to out-of-sample +1.7bps against a 10bps round-trip taker cost.
Round 1 (2026-07-07) — the cost-aware execution study is now done, plus fresh cross-asset + finer-grain cuts. It did not rescue the signal:
Round 2 (2026-07-07) — exactly that test, pre-registered before the data was touched, out-of-sample burned exactly once, verdict read off a pre-declared table:
The meme: BTC’s CME futures close for the weekend, spot keeps moving, and the resulting “gap” gets filled — so fade the weekend move at the Sunday reopen. Pre-registered before touching an outcome: fade every |gap| ≥ 1% at the reopen, exit at the Friday-close level or a 120h time stop, 10bps round trip, split at the house 2022-12-17 boundary. 445 weekends since CME launch, 243 tradeable gaps.
| n | gross bps/event | net bps/event | hit rate | |
|---|---|---|---|---|
| In-sample 2017–2022 | 167 | −39.4 | −49.4 (t −0.88) | 70.7% |
| OUT-OF-SAMPLE 2022–2026 | 76 | −6.8 | −16.8 (t −0.27) | 77.6% |
First screen in the price-level class. The FX literature says stops cluster just beyond round numbers, so a fresh breach should cascade and continue. Pre-registered: first hourly close beyond a round $1,000 BTC level after ≥24h on the other side, then hold 24h in the breach direction. The design’s whole point is the control — the identical test on the $x,500 grid, same spacing, same geometry, zero psychology — so generic breakout momentum cannot masquerade as a round-number effect. Only the round-minus-offset excess counts.
| Grid | OOS n | OOS gross bps | OOS net bps | hit rate |
|---|---|---|---|---|
| Round $1,000 levels | 666 | −0.8 (t −0.08) | −10.8 | 45.9% |
| $x,500 control grid | 657 | +4.8 (t +0.48) | −5.2 | 45.7% |
Most people only want the signals that work. Understanding why the rest don't IS the skill.
A signal dies two fundamentally different ways. Knowing which is which is the whole skill.
Over-analyzed (price & chart patterns — every market, including crypto). The signal is completely real. It's also the most-used input on earth, so the edge is competed away almost instantly. Real pattern, ~zero remaining alpha. This is why our Screen 01 lead-lag died — not because the pattern is fake, but because everyone already trades it.
Out-competed (equity stat-arb, latency-gated signals). The edge is real — and already harvested by insanely-funded quant teams with better data, faster pipes and more capital (Renaissance, Citadel, Two Sigma). Not impossible; effectively not ours. If extracting it needs speed or scale we don't have, it isn't our edge.
Decayed after publicity (the half-life problem). Congressional-trade copying — the "Pelosi tracker" — genuinely worked: disclosed trades from informed actors, copyable for free. Then trackers, headlines and ETFs (NANC/KRUZ) crowded in and compressed it. The signal stayed real; its edge ran out of half-life.
The in-sample mirage (overfitting). A rule that looks gorgeous on history and dies on data it was never fit on — because there was never an edge, only noise that happened to fit the past. Live example: Lab-log Screen 01, in-sample Sharpe +0.82 → out-of-sample −0.59. The ONLY genuinely "fake" category — and the most dangerous, because it's the most seductive.
⌦ The headstones — signals we tested out-of-sample and buried
66 hypotheses tested · 0 deployable edges — public + academic claims, our own proprietary on-chain data, price, factors, macro, order flow — and now a 16-screen mega-cook. Exactly two leads have escaped burial so far, and neither is for sale: the BTC variance-risk premium — real, but fading, and it needs options rails + capital we don't allocate (real ≠ ours) — parked; and the liquidation-cascade bounce (screen 58) — a real bounce whose net edge failed a two-round, pre-registered gauntlet (round 1: execution study; round 2: pooled deep-crash test, out-of-sample burned once on n=304 — gross +32bps real, net +4bps ≈ zero, negative ex-BTC and every year since 2024) — shelved, real-but-sub-cost. And our tester is calibrated: it confirmed a known-true equity anomaly, so the kills are credible — not the artifact of an always-negative test. The graveyard grows as we test more. A record of honest kills is rarer — and more trustworthy — than another channel shouting buy.