A great deal of current work is deep learning and “AI”.
We do both, and we are equally willing not to. A twelve-sample dose–response
study needs a hierarchical model and a realistic confidence interval; a screen
across a million compounds is where a transformer earns its place.
The method therefore follows the data: how much of it there is, how
it was collected, what is being asked, and what the consequences are if the
answer is wrong. Where a simple model and a complex one perform the same, we use
the simple one. Equal performance on the data you have is not equal performance
on the data you do not, and the extra parameters are usually fitting noise.
The same applies to infrastructure. Reproducibility is what keeps an
analysis meaningful long after it was first run, so we build it in from the start
rather than adding it afterwards.
See what that has looked like on real projects.