Systems ~5 min read

Philosophy and Design Choices

The decisions behind Outlier Detector, and the features I had to cut before I could trust it.

This started with something I kept noticing from trading by hand. Certain hours looked more alive than others, and certain shapes on the chart kept seeming to matter. So I wrote a small script that took a year of my trades, bucketed them by minute of the day, then by hour, then by weekday, and added up the P&L in each bucket. Some buckets came out reliably greener than others. Was that real, or was I just looking at noise? I couldn't tell.

What I could tell was that I didn't trust it enough to put money behind it. Something happening before doesn't mean it happens again, and my paper-trading setup back then wasn't good enough to settle it. That gap, between a pattern that looks interesting and something I'd actually risk money on, is most of what this project ended up being about.

Wanting a real signal

My first momentum signal was barely a signal. Sort coins by recent return and you get more or less the top gainers and losers on CoinMarketCap, which anyone can see for free. It also raised a question I still haven't fully shaken: if a coin has already run a long way, why should it keep going? A lot of the time there's more reason to fade it than chase it. The big liquid coins tend to trend and the small ones tend to snap back, and whatever I built had to take that into account.

So I read a fair bit of Cliff Asness and AQR on momentum and value, and tried to work out how much of it carries over to crypto. Some of it does. The problem is the inputs. Most alts still just follow BTC around, and a good chunk of the volume is fake. I could borrow the ideas, but not the assumption that the data was clean.

The score

Every 15 minutes the detector ranks the whole list against itself, and the score it ranks on is deliberately simple. The base is each coin's own momentum: its log-price change over a window, divided by its recent volatility. The volatility part matters. A coin that moved cleanly scores higher than one that covered the same distance while thrashing around. I also skip the last few bars, so the newest and noisiest candle doesn't set the score.

That number means nothing on its own, so I never compare it to zero. Each cycle it gets turned into a z-score against the rest of the list, and clipped at a few standard deviations so one ridiculous print can't take over the leaderboard. The raw momentum says how a coin is moving. The z-score says whether that's interesting compared to everything else moving at the same time.

Then there's curvature, which is whether a move is speeding up or rolling over. I wanted it to matter more than it does. Weighted heavily, it just chased tops and added noise, so now it has a small weight and a tight clip and only nudges the timing. There's also a small bonus for coins that stay near the top across several closed bars instead of spiking once and disappearing.

What goes into the rank

  • Momentum: log-return over a lagged window, divided by volatility. The core of the score.
  • Relative momentum: that momentum z-scored across the list every cycle, then clipped.
  • Curvature: whether the move is speeding up, kept to a small weight.
  • Hurst: whether a series trends or mean-reverts. A hard filter, not a weight.
  • Persistence: a small bonus for staying near the top across closed bars.

Hurst took me the longest to get right. It estimates whether a series trends or mean-reverts, and my first instinct was to add it to the score as another weighted factor. That was wrong. Strong momentum with a low Hurst is exactly the trap: a sharp move with nothing behind it. Average that into a score and the loud momentum just buys back the points Hurst was trying to take away. So now it's a gate. Below the cutoff a coin is out, however good its momentum looks. Some of the most useful changes I made were taking a feature out of the score and turning it into a filter.

BTC decides a lot

BTC drags everything else around, so two inputs sit above the ranking. One is a BTC regime score from zero to three: a point for price above its long EMA, a point for a rising medium EMA, and a point for calm volatility. The other is dominance. I don't have clean spot dominance data, so I use Binance's BTCDOMUSDT futures as a stand-in and turn it into falling, neutral or rising.

Early on I let dominance veto signals outright, and it quietly killed real ones. Now both inputs just move the score a coin has to clear. A bad backdrop means a signal has to be stronger, not that nothing counts.

Live versus closed candles

The same engine runs twice. One pass ranks on the live candle that's still forming, and it's allowed to be early and noisy. The other only ranks on closed 15-minute bars, and that's the one I trust. The live side doesn't treat every flicker at the top as a breakout. A coin has to keep showing up, with its rank and score improving, before it moves from the watchlist to emerging, and again before it reaches the tier I'd actually act on early.

The live prices never touch the closed-candle history. If a half-formed candle leaked into it, the one number I actually rely on would be wrong and I wouldn't know.

A hand-picked list

The list of coins is one I maintain by hand. That's a step down in sophistication, and I did it on purpose. The chart below is why.

SIREN · previously a top-40 coin by market cap, listed across major exchanges

SIRENUSDT chart showing violent price spikes and collapses.

SIREN wasn't some forgotten micro-cap. It was a top-40 coin listed almost everywhere, and the chart still looks like pure manipulation. Automated universe selection mostly uses liquidity and market cap to decide what's real, and SIREN shows that can be wrong even for big names. I'd rather keep a smaller list I can reason about than hand that decision to a filter that trusts the exact numbers being gamed.

Value got left out for a similar reason. In crypto it's hard to define without stitching together activity, revenue and on-chain data, and I didn't want to force it in just so the model sounded complete. The rule I try to stick to is simple: a feature stays if it makes the detector better, not if it makes it sound clever. Curvature barely survived that. Value hasn't made it in yet. Hurst only stayed because it changed jobs.

Testing

Most of the real work has been testing. I run a few copies on rented VPSs, log every row they score to SQLite, and go back through the output by eye a day or two later. There's a replay tool that runs the actual engine over recent Bybit history, not a separate backtest version, so what I'm checking is the same code that runs live. It's slow and a bit tedious. I still trust it far more than a nice first backtest, which in my experience is usually just a bug I haven't found yet.

I didn't want something that only wakes up when a bar closes, but I didn't want it re-ranking on every WebSocket tick either. So there's a short delay before each cycle and a throttle on the live pass. It's always watching, but it rarely says anything.

Where it stands

Outlier Detector is a detector, and that's on purpose. Execution, sizing, position management and everything else a real trading system needs are out of scope until I trust the signal underneath. It's still a work in progress. Building the part I can check first seemed like the only way to stop the whole thing falling over later.

Back to Ideas