EquationDB

← Back to all articles

Why we built EquationDB

2 October 2026 · EquationDB team

We built EquationDB because we were tired of waiting for our own questions.

The questions were not exotic. Which of the 500 largest stocks are oversold but still above their 200-day average? Rank every stock in each sector by three-month momentum. What happened in the 20 days after every gap-down since 2005, and was it any better than just holding those stocks? What was this company's P/E as it was known at the time, not as restated later? Each one is ordinary research. Each one, on the databases we were using, took anywhere from a few seconds to a few minutes, and every follow-up question paid the same price again.

What "slow" looks like in practice

The problem is not that general-purpose databases are badly built. It is that market questions are a bad fit for how they work.

A row store or a general-purpose SQL engine keeps prices as rows and fundamentals as more rows. To answer "which stocks are oversold right now?", it has to read enough history for every symbol to compute RSI, compute it, keep the last value, and then filter. The indicator is not stored anywhere; it is recomputed from raw prices on every query. A cross-sectional rank means doing that for the whole universe and sorting. An event study means scanning decades of bars for every symbol to find the days a condition fired, then joining forward returns back on. A point-in-time join means matching every bar to the fundamentals that were known on that date, not the latest restated figures.

None of these is hard to write. All of them are expensive to run, every time, because the database starts from scratch on every question. Recursive indicators make it worse: RSI, ATR and ADX depend on their own previous values, so they cannot be expressed cleanly as plain SQL window functions at all.

We wanted the opposite: a database where the expensive part happens once, when data arrives, and a question is mostly a lookup.

What we did differently

Live state per symbol, maintained incrementally. Indicators update when a bar arrives, not when you ask. Each symbol carries its own state for every function in use (moving averages, RSI, ATR, Bollinger Bands, about a hundred functions in all), and each new bar advances that state by one step. A screen over 500 stocks reads current values from memory. A function used for the first time, say rsi(9) on 15-minute bars, is registered on demand, seeded from full history, and maintained like the built-in ones from then on. EXPLAIN shows whether each field is already warm:

EXPLAIN GET top500 WHERE rsi(9) < 20

One code path for live and history. History queries replay the same function code over stored bars, so the RSI in a backtest, an event study or a chart from three years ago is exactly what live state would have shown that day. Backtests, alerts and screens cannot disagree with each other because of two implementations of the same indicator.

Events computed once and indexed. More than 60 detectors run on every closed bar. Each occurrence is stored once with its magnitude, its rarity against the same symbol's history and, as the future arrives, its forward returns. "What happened after golden crosses?" then reads an index rather than rescanning history.

Partitioned, columnar history. Bars are stored per symbol, per timeframe and per partition (one year for daily bars, one month or one session for intraday), as compressed columns. One symbol's year of history is one small decode, not a scan through everyone else's rows.

A light path for the long tail. Full live state for every listed instrument would be wasteful. For the long tail of the roughly 70,000 symbols we carry, EquationDB keeps the last two daily bars in memory, so simple universe-wide screens on price, change, gap and volume still answer from memory, while full history stays on disk for when you ask about one of them.

Point-in-time fundamentals. Each fundamental value is stored with the session on which it became known, so a history query or a backtest only sees what was known then. Price ratios such as P/E are computed from EquationDB's own close and the latest per-share figures, so they move with the price every day.

Adjustment at read time. Splits are applied as factors when bars are read, so history on disk is never rewritten. Dividends are a per-bar factor you choose to apply:

GET KO SELECT close LAST 5y ADJ total

One engine, two languages. A compact market language for the questions traders ask, and SQL for grouping, joins and window functions, over the same data with indicators as ordinary columns:

SELECT sector, count(*) AS stocks, avg(pe) AS avg_pe, median(rsi(14)) AS median_rsi, sum(CASE WHEN close > sma(200) THEN 1 ELSE 0 END) * 100.0 / count(*) AS pct_above_200d FROM top500 GROUP BY sector ORDER BY pct_above_200d DESC

Did it work?

In our tests (one shared virtual CPU with 1 GB of RAM running both engines, the same 3.8 million daily bars for the top 500 stocks over their full history, medians of five runs after a warm-up, result cache off), we compared EquationDB with a popular general-purpose embedded analytical SQL engine:

QuestionEquationDBEmbedded SQL engine
Screen now: close above the 200-day and below the 20-day average (500 symbols)3.4 ms2,385 ms
Screen now: new 52-week high2.0 ms2,926 ms
One symbol, one year with 50- and 200-day averages1.9 ms9.9 ms
Forward returns after a condition, scanned over full history5,153 ms2,092 ms

"What is true now" questions were 700 to 1,400 times faster, because the answer was already computed when the data arrived. Single-symbol history was about five times faster. We are just as clear about the last row. When OUTCOMES has to scan an arbitrary condition across 45 years and 500 symbols, a vectorized engine was about 2.5 times faster on that small box, because EquationDB replays exact indicator state bar by bar so that historical values match live ones. For the events EquationDB already indexes, it uses pre-filled outcomes instead of scanning.

A database that can think

Speed was the first goal. The second was a database that does more of the reasoning a careful analyst would otherwise do by hand. Here is what that means today, using features that exist now.

It knows what usually happens next, and compares it with a baseline. OUTCOMES returns forward returns after any event or condition, together with the same stocks' unconditional return over the same period and a t-statistic. An "edge" that is just the market drifting upwards shows up as zero.

OUTCOMES AFTER gap_up IN top500 SINCE 2015-01-01 GROUP BY sector

It looks for edges itself, without fooling itself. SEARCH tries combinations of conditions, selects on training years, checks them on validation years, reports on untouched test years, and applies false-discovery control over every candidate it tried. It tells you how many candidates it tested and which survivors failed out of sample.

SEARCH SETUPS IN top500 GOAL edge > 0.2% USING gap, streak HORIZON 20d

It understands questions in plain English, without letting a model write code that runs. A language model extracts the intent; EquationDB checks every field and event name against its catalog and compiles the query deterministically. You always see the query. In the editor, the AI assistant can write a query from a description, and every query it writes is parsed and, if it only reads data, test-run before it reaches you. If the test run fails, the model gets the real error and tries again. Improve with AI sharpens a vague question, Explain walks through what a query does and its caveats, and Fix reads the last error and proposes a corrected query.

It knows how the market is connected. A nightly relationship graph of correlations, betas and intraday lead-lag answers "who really trades with this stock?", "who moves first?" and "if this falls 5%, what else moves?", and it raises events when a stock's relationship with its sector breaks.

Where it is going

The next step is a database that notices things on its own: one that brings you the unusual, the newly significant and the quietly broken before you think to ask.