Research period August 2026 · James Kennedy, founder

Research notes

Why we started looking harder at behaviour

How much should it matter when a horse does not arrive at the start in the state the form book assumes it is in? Everyone who watches racing has an opinion. We wanted to know whether it could be measured.

The bit that happens before anyone is watching properly

You have seen it. Everyone has. A horse gets warm and busy going down. One plants itself and takes four handlers to load. One rears in the stalls and comes out already behind. A saddle slips. A horse misses the break by two lengths in a race where two lengths is the whole story.

None of that is exotic. It is a normal afternoon's racing and the commentary will usually mention it in passing. What interests me is that by the time the tapes go up, the numbers we have all been staring at for an hour describe a horse that turned up in good order and got a fair run at it.

Sometimes that horse did not turn up at all. Something else did, wearing the same colours.

What the form book quietly assumes

Form is a record of what happened. It is very good at that. But almost every rating built on top of it carries an assumption nobody states, which is that the horse arrived at the start ready to run the race the rating describes.

Most of the time that assumption is fine, which is exactly why it survives without being examined. The question is whether the times it fails are random noise, in which case they cost you nothing over a season, or whether they cluster in ways a patient observer could see coming.

That is a genuinely open question. It is also the sort of question racing answers badly, because everyone has a vivid memory of one horse boiling over and losing, and nobody has counted the ones that boiled over and won anyway.

We priced the answer before we went looking for it

Here is the part I am actually proud of, and it is not a finding.

Before collecting anything, we wrote out the arithmetic of how much evidence it would take to detect a behavioural effect if one were really there. Not after the data came in and disappointed us. Before, when we still had no idea which way it would fall and no reason to flatter ourselves.

The answer was sobering. Even assuming a very large true effect, considerably larger than anything I would honestly expect, you need in the region of 800 races to detect it at roughly 76% power. That is a lot of racing. It is a serious commitment of patient, boring, unglamorous capture on days when nothing appears to be happening.

And it tells you something brutal about the small version. A test at 50 races has essentially no power at all. Which means a null result at that sample is not a null. It is a shrug wearing a lab coat. You would have failed to find the effect whether or not the effect exists, and you would have no way of telling those two worlds apart afterwards.

So we did not run it

One sub-study inside this lane was formally refused. Not shelved because it looked unpromising, not quietly dropped because we lost interest. Refused, in writing, on the grounds that the sample could not support a conclusion in either direction.

I understand how that reads. It is the least exciting sentence a research company can publish. We could have run it, got a number, and written it up. Underpowered findings are the most abundant product in this industry precisely because they are so easy to manufacture, and nothing about the write-up would have looked any different to a reader.

But an answer you cannot trust is worse than no answer, because you will build the next three things on top of it. That is how a research programme rots. Not in one dramatic mistake, but in a small unearned conclusion that nobody revisits until it is load bearing.

Refusing on power grounds is a decision about what the evidence can support, taken before the evidence arrives.

What actually exists today

A grammar, mainly. Before you can count behaviour you have to agree what counts as an event and describe it the same way every single time, otherwise you are just collecting opinions with timestamps on them. So there is a defined vocabulary for what we observe, and around 485 historical events have been coded into it, with more coded as each day's racing goes past.

What that is not: deployed, gated, or eligible as proof of anything. The behavioural lane is a specification and a growing set of observations. It feeds no decision. It has never influenced a read. It is not part of any grading.

So the honest position is this. We are taking behaviour seriously, we have built the machinery to measure it properly, and we have not earned a finding. Behaviour does not predict anything as far as Elite Pass is concerned, because we have not yet done the work that would let us say whether it does.

A company that tells you what it has not yet earned is worth more than one that tells you what it hopes. Ask me again in 800 races.

Related research

Elite Pass publishes research, not guaranteed outcomes. Findings are measured against the market and remain subject to replication and sample size. The evidence room and the public ledger carry the full record.