DataJoint has achieved Built-On status with Databricks.

Where does your organization keep its reasons?

Data survives. Reasoning does not.

Most R&D organizations I speak with do not have a data shortage. Storage is cheap, instruments are prolific, and the measurements are almost always there somewhere.

What is missing is why. Why a target was dropped in 2023. Why an assay protocol changed between two campaigns. Which result convinced a team to move forward, and which one told them to stop.

What a typical R&D record retains, and what it does not.

The pattern repeats across organizations of every size: the compound history is in the system, and the decision history is in a few people’s heads. By the time anyone goes looking, one of them has usually left.

AI makes the gap visible immediately

This was survivable when the person who ran the work was the person who explained it. AI removes that arrangement. Ask a model why a program was killed and you will usually get an answer, fluently, assembled from whatever documents happen to have survived.

A fluent answer is not the same as a true answer.

The output is not obviously wrong, which is the problem. A plausible account built from a partial record arrives with the same confidence as a retrieved fact, and nothing in the answer tells you which one you got.

The negative result is the costliest thing you are not keeping

In discovery, the study that did not work is the least likely to be written up with its full context, the least likely to be reviewed, and the first to be forgotten when the person who ran it moves on.

It is also the result that narrows the search space for everyone who comes after. The principle I keep coming back to:

No result should be wasted, including the ones that failed. 

Every negative result that is not preserved is a study somebody else will run again, at full cost, to reach the same conclusion.

The economics of this are documented. In PLOS Biology, Freedman, Cockburn and Simcoe estimated that more than half of US preclinical research cannot be reproduced, at a cost of roughly $28 billion a year in the United States alone.

Their breakdown of causes:

  • biological reagents and reference materials at 36%
  • study design at 28%
  • data analysis and reporting at 25%
  • laboratory protocols at 11%.

More than a third of the problem traces to how work is analyzed, reported and documented, not to the materials or the design.

The same paper notes that the underlying biases can result in entire studies never being published or reported at all. The negative result does not just lose its context. It disappears.

A recollection is not a record

Some organizations try to close the gap afterward: exit interviews before a scientist leaves, retrospective write-ups, a wiki that nobody maintains past the second quarter. Reasoning captured afterward is a different artifact. It is a recollection, which is useful and is not the same as a record.

A record is made at the moment of the decision, by the people inside it, with the alternatives still on the table. A recollection is what someone believes happened, filtered through everything that came after, including how it turned out.

That is why this does not fix itself later. The reasons get captured as the work happens, or they get approximated afterward. And an approximation is exactly what a model will treat as ground truth.

Where does your organization keep its reasons, and who could find them without asking a person?

The second half of that question is the one that matters. If the answer is a person, the reasoning is not stored. It is borrowed, on terms that end when they do.


Sources

Freedman LP, Cockburn IM, Simcoe TS, “The Economics of Reproducibility in Preclinical Research,” PLOS Biology 13(6), 9 June 2015. journals.plos.org 

Science, “Study claims $28 billion a year spent on irreproducible biomedical research,” 9 June 2015. science.org