The Research Ledger

192 ideas tested. 105 of them didn't work.

This is the complete record of every trading idea Brown Bear Research has put through a formal test — overwhelmingly, the ones that failed. Most research sites publish their winners. The failures are the more useful half, and almost nobody shows them.

105Rejected outright
54Still open
20Method lessons
9Still standing
4Untestable

Looking for what did work?

The Brown Bear Report carries the most promising signals this research has produced, and the names each one is firing on now. The Report is what is still standing and being watched live — not a list of guaranteed winners.

How this works

Every idea here was pre-registered before it was tested: the hypothesis, the exact pass/fail thresholds, and the data window were written down and cryptographically fingerprinted in advance. That makes it impossible to move the goalposts after seeing the answer — the single most common way research fools itself.

A rejection here doesn't mean "this never works for anyone". It means: tested on this universe, over this period, against a bar set in advance, it did not clear the bar. Where a test lacked the statistical power to decide, it's recorded as undecided rather than rejected — those are the 54 still open.

The four corrections that changed the most answers
  • Buy at the next morning's open, not tonight's close. Assuming you can buy at the same closing price that triggered the decision roughly doubles apparent performance.
  • Include the companies that went bust. Testing only on firms that still exist today flatters everything. Our own local price store keeps just 7.9% of delisted names, worth roughly 4 percentage points a year of phantom return.
  • Measure against the S&P 500, not in raw terms. A strategy that made 15% in a year the index made 18% lost money in the only sense that matters.
  • Subtract the cost of trading. Roughly 44 basis points per round trip. Many effects that look real are smaller than the cost of harvesting them.

What failed, by family

The rejections aren't 105 unrelated disappointments. They cluster, and the clusters are the interesting part — whole categories of technique that repeatedly produced nothing.

Chart patterns and candlestick setups

0 for 14

Fourteen consecutive named patterns, tested one at a time, none of which survived: bull flags, breakout-and-retest, pump-then-dip, capitulation engulfing bars, shrinking red trios, MACD divergence, EMA reclaims, RSI shock drops, sharp-drop-tight-base, and more.

The pattern behind the pattern failures

Two things recurred often enough to become predictive. First, these setups usually looked promising in raw returns and evaporated once we corrected for when they fired — they cluster on days the whole market rose, so they were measuring the market, not the stock.

Second, and more damning: when we tested the individual clauses separately, the specific clause the pattern is named for was often the one subtracting value. In sharp-drop-tight-base, the drop earned its keep and the "tight base" was dead weight. In breakout-and-retest, both of the proposed conditions inverted. The decoration wasn't neutral — it was actively harmful.

Knowing when to sell

13 rejections

Four separate pre-registered programmes, roughly 700 rule variations between them: trailing stops, chandelier exits, support breaks, relative-strength decay, rank slippage, bearish reversal detectors, deteriorating fundamentals. Nothing beat simply holding.

The uncomfortable finding underneath

The best-powered study in this family found something we didn't expect: the median held position earns nothing at all. Returns come from a small number of large winners. Any sell rule that trims the losers also trims the future winners, because in advance the two are indistinguishable.

Worse, a position that has already gone against you turns out to be positive-expectancy to keep. Selling it is the mistake.

The conclusion we now work from: exiting isn't about detecting a warning sign. It's about whether the money has somewhere better to be.

Cleverer versions of momentum

12 rejected

Faster momentum, acceleration, beta-neutral and "purified" residual momentum, volume-weighted momentum, Ichimoku as a ranking factor, adaptive position counts, ordinal ranking. Every refinement of plain 12-month momentum either matched it or did worse.

Why orthogonalisation did nothing

A standard academic move is to strip out the parts of momentum explained by market beta, volatility and sector, leaving "pure" stock-specific momentum. We ran it. It was a no-op — the cleaned signal performed the same as the dirty one.

Relatedly: volatility, beta, illiquidity and lottery-likeness turned out to be one axis wearing four names, explaining 62–69% of the variation in outcomes between them.

Fundamentals and valuation

10 rejected

Cyclically-adjusted valuation, mechanical discounted-cash-flow gaps, earnings deterioration as an exit signal, debt paydown, post-earnings drift as a buying lane, and quality filters layered onto momentum.

The one distinction that survived

Almost every accounting-based measure failed. Two cash-flow measures survived — free cash flow and operating cash flow relative to assets — while every measure based on accruals died.

A sharper finding: the level of cash generation carries information, but the moment it turns positive carries none. Dated fundamental events are far weaker than standing fundamental conditions.

One methodological trap worth flagging: ranking companies on fundamentals across the whole market is a sector bet wearing a quality label. Software firms and utilities have structurally different cash profiles. Ranking within sector fixed it — and shrank the edge to about 1.5 percentage points a year before costs.

Market timing, breadth and regime

Closed

Whether market breadth improving predicts a turn, whether a macro regime map should size positions, mechanical trend-following gates to cash, and volatility targeting. All rejected as return signals.

What breadth actually detects

Breadth change — "participation has improved over the last month" — was rejected on a five-way convergence of independent evidence. Notably, where breadth carries information at all, it detects tops, not bottoms, which is the opposite of how it is usually sold.

The macro regime layer failed specifically as a return device: the deeply stressed states carry high average returns and high risk simultaneously. It separates risk well. It cannot tell you when to buy.

Sector rotation

Programme closed

All five constructions we could build, including relative-rotation-graph quadrant narratives and execution-footprint detection. The whole layer was closed rather than iterated on.

Seasonality and calendar effects

Closed on arithmetic

Same-calendar-month seasonality, turn-of-the-month, option expiry, and index-rebalance flow windows.

Why this one didn't even need a backtest

The effects are real and far too small. Once you price in roughly 44 basis points per round trip, the arithmetic closes the question before any statistical test is required.

A separate trap we walked into and documented: a smoothing method can manufacture the appearance of persistence out of pure noise. You have to compute what your own smoother does to random data before reading persistence as evidence of anything.

Short interest and short-sale flow

All 3 layers spent

Crowded-short conditions as an avoid signal, daily off-exchange short volume as a state at deep drawdowns, and short interest as a return signal in its own right.

Stocks having personalities

Family rejected

A long-running intuition: individual stocks have persistent behavioural signatures — a characteristic momentum decay, a characteristic reaction to volume spikes, a habit of over- or under-performing its own score.

The finding that closed the family

These traits are genuinely persistent and genuinely measurable. They just map to risk, not return. That turned out to be a general law here rather than a coincidence: every persistent characteristic we return-tested — days-to-cover, short interest, beta, volatility, dollar volume, index coupling — predicted how much a stock would move, never which direction.

Return, where it existed at all, lived only in temporary conditions, never in permanent characteristics.

Our own data, twice

Self-caught

Two entries in this section aren't failed hypotheses at all — they're corrupted inputs we caught in our own pipeline, recorded here because pretending they didn't happen would defeat the point of keeping a ledger.

The run that passed all six criteria on garbage

One study printed six consecutive PASSes against pre-registered thresholds. The thresholds were met. The data feeding them was corrupt — a vendor price-scale splice that our corruption scrubber could not structurally detect, because the prices either side of the seam were individually valid.

This produced the single most useful lesson in the programme, and it is lesson 18 below.

What we learned about testing

These are the most valuable things this programme produced — not edges, but corrections. Each one changed the answer to questions we had already thought we'd settled, and each is reusable by anyone doing this kind of work.

01

Same-day-close backtests roughly double your apparent skill

If a test assumes you buy at the same closing price that triggered the decision, it is crediting you with information you did not have. Correcting one strategy to buy at the next morning's open dropped its Sharpe ratio from 1.22 to 0.62 — the same strategy, honestly measured, was half as good.

02

You cannot predict the downside — only your exposure to it

Around twelve separate attempts to forecast which stocks would fall, over horizons from one week to three months. All failed, and all failed with the same signature. Manage how much you are exposed; don't try to see it coming.

03

Returns come from a handful of enormous winners

Only 44.5% of stocks beat the S&P 500 over our window. Every classifier we built to sort future winners from losers performed at chance. What worked slightly was predicting how big a move might be, never which way it would go.

04

Selling is a capital-allocation decision, not a warning sign

After thirteen failed attempts to find an exit trigger, the framing that survived is that you don't sell because a stock looks bad. You sell because the money has somewhere better to be.

05

This universe paid for volatility

Among large US stocks over our test window, more volatile names paid better, and low-volatility and market-decoupled names were a drag. Stated as an observation about this universe in this era — not a law. The widely cited "betting against beta" effect ran backwards here.

06

Not trading is a free diversifier

Keeping existing positions rather than constantly swapping them helped on returns — and, discovered later and independently, on risk too. A portfolio that rotates slowly ends up holding stocks bought at different moments, and those move together less than a fresh basket picked all on one day.

07

Splitting your data in half by date hides the thing you're looking for

"First half versus second half" is the standard robustness check and it is close to useless. Market regimes don't align with calendar halves; you have to split on the actual regime.

08

Three corrections that repeatedly reversed the answer

Measure against the index rather than in raw terms. Include the companies that went bust or were taken over. And compare a stock against its own history rather than against other stocks — several effects that looked real across stocks vanished, or flipped sign, when measured within each stock.

09

Momentum dies as a cliff, to whole groups at once

When a stock stops leading, it takes a median of nine trading days. But it happens to entire cohorts of similar stocks simultaneously, not to individual companies — and in advance it is indistinguishable from an ordinary wobble, which is five times more common. About 83% of apparent leadership losses recover.

10

How often you trade changes your results, with no rule involved

Trading frequently looks skilful in choppy markets and foolish in trending ones — regardless of whether the trading rule has any merit. This means you can never compare two strategies that trade at different rates without a control that trades at the same rate randomly.

11

A dataset that looks thin is a bug until proven otherwise

Twice in one day, data that appeared to be missing turned out to be a loading error. The most costly instance: a belief that our price history only reached back five years — repeated as a caveat across dozens of studies — was simply false. It reaches 2018 and spans three bear markets.

12

Permanent characteristics predict risk; temporary conditions predict return

The most useful single distinction we found. If a measurement keeps its rank between stocks over time, it is a characteristic and it will predict volatility, never direction. If it doesn't persist, it is a condition — and conditions are the only place return has ever lived here. You can tell which you have before running any return test at all.

13

We audited all 246 tests at once, and both flagship edges failed

A census of every arm ever tested found real structure: 100% of validated results were sign-stable across eras versus 10% of rejected ones, and six recurring failure modes covered 95% of the 184 rejections. Then we ran a pre-registered test on a decade of data we had never examined. Both of our flagship entry edges failed it. Nothing in the programme cleared a trial-count-adjusted significance test.

14

Smoothing manufactures persistence out of noise

Any rolling, smoothed or differenced construction inherits a long memory from its own arithmetic. Before treating persistence as evidence, run the identical calculation on random data and confirm you can tell the two apart.

15

Check the measurement is sane before testing whether it predicts

A construction screen costs about an hour, touches no outcomes, and therefore cannot contaminate anything. It is the cheapest filter available and it kills most ideas before they consume a full study.

16

Two floors any cross-sectional signal must clear

First: compute your measure on the index itself. If the index scores higher than every individual company, you have found a market-wide phenomenon that cannot possibly rank companies against one another. Second: any measure taken over a limited window carries built-in random error — if the spread between companies is no larger than that error, the entire spread is measurement noise and there is nothing to rank on.

17

Peeking at a subgroup spends the whole experiment

Narrowing to a subgroup before running the wider test contaminates the wider test, because the subgroup sits inside it — you already know part of the answer. Run the wider test first, or declare the subgroup in advance.

18

Pre-registration does not protect you from bad data

The hardest lesson here. A locked-in pass/fail threshold tests a number; it cannot tell you the number was computed on garbage. One of our runs passed all six of its pre-registered criteria on corrupted prices. Locked criteria don't prevent a wrong answer — they make a wrong answer look official. Every study now carries a separate plausibility check against something independently known.

19

Price histories contain invisible multi-year holes

A stored series can have years missing with nothing marking the gap, and every standard "look back N days" calculation reads straight across it as though no time passed. In our data this affected about 0.5% of rows and manufactured roughly a full percentage point of phantom edge. Corruption scanners cannot catch it, because the prices on either side of the hole are individually valid.

20

Rarity buys variance, not edge

Rare signals look spectacular and mostly aren't. The rarest tier of results we examined retained only 15% of its apparent advantage on retest, and 26% of rare strategy variants cleared our cost threshold on noise alone. Concentrated, high-conviction, rarely-firing ideas are where self-deception is easiest.

Where this leaves us

The base rate of a candidate idea becoming a believed return edge is about 1–2%. The base rate of surviving our own adversarial re-testing intact is, so far, zero.

That sentence is from our internal meta-study and it is the honest summary of six months' work. Two ideas made it all the way to "believed true" — buying fallen market leaders, and buying proven winners on an ugly pullback — and then both failed when tested on a decade of data we had deliberately never looked at.

What the programme has produced is a validated map and a validated method: six recurring ways that promising ideas turn out to be illusions, and a set of cheap tests that catch most of them in an hour rather than a fortnight. The failure modes have proved more reusable than any of the edges.

Everything above is measured on US large-cap equities, mostly 2010–2026, with delisted companies included, returns measured against the S&P 500, and trading costs applied at roughly 44 basis points per round trip. Nothing here is investment advice.