EquationDB

← Back to all articles

Searching for edges without fooling yourself

2 October 2026 · EquationDB team

If you test enough trading rules on the same history, some of them will look excellent by pure chance. This is the oldest trap in quantitative research, and it gets worse as tools get faster: the easier it is to try one more condition, the more conditions get tried, and the more of the "discoveries" are noise.

EquationDB's SEARCH command is built around that problem. It lets the database look for edges on its own, but it applies the guardrails a careful researcher would apply by hand, and it applies them every time, whether you ask or not. This article explains what SEARCH does, what the guardrails are, and how to turn a result into something you can query and watch.

Two kinds of search

SEARCH SETUPS combines conditions from families you choose (momentum, trend, volume, gaps, streaks, ranges, volatility, market regime, events) and keeps the combinations that meet a goal you set. SEARCH EVENTS works the other way round: instead of starting from conditions you might think of, it encodes every bar as a market state and mines the transitions into states that were followed by unusual returns.

Both run on daily bars and both report the same statistics, so the results can be compared side by side.

SEARCH SETUPS: a goal, some families, three periods

Here is a full search:

SEARCH SETUPS IN top500
  GOAL win_rate > 0.55 AND mean_ret(5d) > 0.3%
  MIN n = 300 USING momentum, trend, volume, market MAX CONDITIONS 2
  TRAIN 2005-2016 VALIDATE 2017-2020 TEST 2021-2026

Read it clause by clause:

  • GOAL is what a setup must achieve: here, a win rate above 55% and a mean five-day return above 0.3%. Goals can use win_rate, mean_ret(h), median_ret(h), edge (the return beyond the unconditional baseline), t_stat and n, joined with AND.
  • MIN n sets the minimum number of occurrences on the training period. Small samples produce the most impressive and least reliable statistics, so the default is already 100.
  • USING picks the condition families. A narrow list is faster and easier to interpret.
  • MAX CONDITIONS limits how many conditions are combined per candidate (two by default, three at most). Every extra condition multiplies the number of candidates and the opportunity to overfit.
  • TRAIN, VALIDATE, TEST are year ranges. The defaults are 2005–2016, 2017–2020 and 2021 to the current year.

A shorter search uses the defaults and asks for something different, a 20-day edge over the baseline from gap and streak conditions:

SEARCH SETUPS IN top500 GOAL edge > 0.2% USING gap, streak HORIZON 20d

The guardrails

These are always on. You cannot turn them off, and the result tells you how they were applied.

  1. A three-way split. Candidates are selected on the training years and must hold on the validation years. The test years are touched exactly once, to report how the survivors did on data that played no part in choosing them.
  2. False-discovery control. Every candidate tried goes into a Benjamini–Hochberg procedure, and a candidate must reach a false-discovery-adjusted q_value of 0.10 or less. The number of candidates tried is reported as candidates_tried, so you always know how large the haystack was.
  3. Stability. At least 60% of the training years must beat the baseline, and the setup must have fired on at least 20 symbols. An edge that came from one great year or one stock does not pass.
  4. Honest verdicts. Each surviving candidate is labeled "held out-of-sample" or "failed test". Failures are shown, not hidden.

The second point deserves a word. A plain p-value answers "how surprising is this one result?" When you test a thousand candidates, you expect about fifty to clear p < 0.05 even if none of them work. Benjamini–Hochberg controls the expected share of false discoveries among the candidates you keep, which is the question you actually care about after a search: of the things this returned, how many are likely to be real?

Reading the result

Each row is a candidate:

ColumnsWhat they tell you
id, condition, descriptionthe candidate's id (used to promote it) and the rule in KQL and in words
n_train, win_train, mean_train, edge_train, q_valuehow it did where it was chosen, and its adjusted significance
n_val, mean_val, win_valwhether it held on the validation years
n_test, mean_test, win_test, edge_testthe single, final look at untouched data
years_positive, symbolsstability over time and breadth across stocks
verdict"held out-of-sample" or "failed test"

The most useful habit is to read the test columns last and to be suspicious of large drops between training and test. A candidate whose edge halves out of sample may still be useful; one whose edge flips sign was an artifact of the training years.

SEARCH EVENTS: let the market define the states

SEARCH EVENTS does not start from your condition families. It encodes every bar as a market state, a combination of RSI band, trend, volume, the size of the move and position within the Bollinger Bands, then mines transitions into states and two-bar sequences, and tests what followed each one under the same guardrails.

SEARCH EVENTS IN top500 HORIZON 5d MIN n = 200 LIMIT 20

This is the closest thing to letting the database notice patterns you would not have thought to test, while still holding them to an out-of-sample standard.

From candidate to event

Anything a search finds can be promoted to a named event of your own. After a search, take the id of the row you want and run DEFINE EVENT me.capitulation AS CANDIDATE c_01 (with that row's id). You can also skip the search and define an event straight from a condition:

DEFINE EVENT me.washout AS rsi(2) < 5 AND close > sma(200)

User events fire when their condition turns true, and from then on they behave like built-in events. You can list them:

EVENTS me.washout IN top500 LAST 3mo

You can also subscribe with WATCH EVENTS me.washout IN top500 DELIVER inapp, or measure them with OUTCOMES.

Edges decay, so EquationDB watches

A discovery is a snapshot of the past. Every week, EquationDB compares each promoted event's last-year edge with the statistics it had when it was discovered, and notifies you when the edge has halved or flipped. You find out from the database, not from your P&L.

Practical advice

  • Start narrow. A wide USING list with MAX CONDITIONS 3 tries many thousands of combinations. That is slower, and false-discovery control then needs much stronger evidence before anything passes.
  • Set a realistic MIN n. For daily setups across 500 stocks, a few hundred training occurrences is a sensible floor.
  • Prefer edge to raw return. GOAL edge > … asks for return beyond the baseline, which removes the market's drift from the answer.
  • Remember survivorship. Searches run over current universe members, so long histories favor companies that survived. Treat results as research, not a forecast.

SEARCH is included in the Trial, Desk Plus and Enterprise plans. The full reference is on SEARCH.