Projects

Projects and apps we have built for our clients.

Beyond direct data analysis, we create custom scientific software: automated analysis pipelines, bespoke databases, visualization tools, and internal platforms that outlive the project they were written for.

The four platforms below are our own. They began as client work, were generalized, and are now applied to new problems. Each one is illustrated with a diagram of how it works.

DrugREV

molecular engine

DrugRev is a molecular engine we developed to explore the chemical binding activity, pharmacokinetics and side effects of drug-like molecules. It draws on a range of mined public and licensed sources — PDB, DrugBank, ChEMBL and FAERS among them — alongside whatever in-house data a client brings, with 3D visualization and docking on one side and statistical and deep learning models on the other.

Together these have helped our clients predict whether a molecule will engage its intended target in a clean and efficacious fashion. DrugRev also offers generative capabilities: designing small molecules and de novo binders — macrocycles, peptides and mini-proteins — with a specified activity profile against a particular set of targets, including the option to ensure the result does not engage designated off-targets.

docking PDB · ChEMBL · DrugBank · FAERS · in-house ADME/T off-target prediction de novo binder design generative design
DrugRev · docking + de novo design selective · 1 of 14 active
binding site KRAS G12C · switch II 2.9 Å 3.1 Å 3.4 Å O N N de novo macrocycle · 9 residues ΔG bind kcal/mol -9.4 poses 128/128 selectivity pIC50 pred cut 6.0 KRAS hERG CYP2D6 SRC CYP3A4 ADRB2 EGFR LCK MAO-A 5-HT2B ABL1 PXR AURKB D2R 4 6 8
  • on-target
  • off-target flag
  • predicted contact
  • de novo binder

BAYESIC

data integration

Multiple datasets spanning various data types — GWAS, transcriptomics, proteomics, binary screens and other specialized measurements — are common in genomics, and integrating them is challenging. The process often gets reduced to a rigid funnel where only genes with positive results in every dataset are considered.

Bayesic selects an appropriate prior for each data type and converts each contribution into a weighted posterior probability. This enables a more flexible evaluation, where a negative result in one dataset does not automatically disqualify a gene, and you can see exactly how much each dataset moved it.

Bayesian integration priors per data type posterior ranking GWAS · RNA-seq · proteomics
Bayesic · priors → posterior priors set · P 0.19
evidence · point ± 1 s.d. null GWAS p 4.7e-12 w 0.34 RNA-seq q 8.3e-04 w 0.26 proteomics q 2.6e-05 w 0.28 CRISPR screen FDR 0.62 negative w 0.12 effect (z) posterior · P(target | data) prior 90% CrI P(target) 0.19 CrI 0.05–0.78 sd 0.227 rank by P(target | all data) top 9 of 1,842 candidates gene Δrank P(target) GCKR·0.34 TCF7L2·0.31 CDKAL1·0.29 HNF1A·0.26 KCNJ11·0.24 ADCY5·0.22 SLC30A8·0.19 MTNR1B·0.17 GLP1R·0.15 an AND-funnel would drop SLC30A8
  • GWAS
  • RNA-seq
  • proteomics
  • CRISPR · negative evidence
  • posterior · SLC30A8

Targette

target discovery

Target identification and validation is the request we get most often. Targette is the tool we built for it: GWAS, transcriptomics, single-cell data and pathway analysis in one place, arranged so that the chain of evidence from a variant to a phenotype stays visible.

Clients use it to go from a genome-wide scan to a handful of candidates that hold up across indications, then to check what else those candidates sit on — the ancillary pathways and protein–protein interactions that turn a clean target into an awkward one.

GWAS single-cell pathway analysis cross-indication network context
Targette · locus → cell type → pathway scanning · 1.1M variants
A · gwas meta-analysis B · single-cell atlas C · pathway graph
  • locus → population → path
  • genome-wide significance
  • background variants & cells

Grimoire

AI assistant

Every research organization has more written down than anyone can hold in their head: assay results, protocols, internal reports, and the literature around all of it. Grimoire is the assistant we build over that corpus — retrieval-augmented, running on a client-specific cloud deployment, and inheriting the same data access privileges as the account asking the question.

The design constraint that matters is provenance. Every answer names the chunks it came from, so a claim can be traced back to the report it came out of rather than taken on trust.

RAG LangChain private corpus per-account access control cited answers
Grimoire · retrieval, with sources idle · 8,431 chunks
corpus 8,431 chunks embedding n = 164 shown context 5 chunks · 3,140 tok assay data3,182 protocols1,046 reports2,417 literature1,786 8,431 chunks · 214 docs scoped to your account UMAP-1 UMAP-2 k = 5 · cosine 768-d · umap query why did IC50 shift in rep 3? answer 4 of 5 cited · 0 unsupported mmr λ 0.5 · 3,140 / 8,192 tok
  • assay data
  • protocols
  • reports
  • literature
  • retrieved
02Stack

The tools we work with.

Our expertise encompasses the computational tools and best practices used by leading biopharmaceutical companies as well as top tech-bio and software organizations, so that whatever we hand over can be picked up and run by your own team.

Languages and pipelines

  • Python
  • R
  • C++
  • SQL
  • Nextflow
  • Snakemake
  • Git

Modeling and ML

  • PyTorch
  • JAX
  • scikit-learn
  • Stan
  • PyMC
  • XGBoost
  • LangChain

Biology and chemistry

  • Scanpy
  • Seurat
  • Bioconductor
  • RDKit
  • OpenMM
  • GROMACS
  • AutoDock

Data and delivery

  • PostgreSQL
  • DuckDB
  • Docker
  • AWS · GCP
  • Plotly Dash
  • R Shiny
  • React

A representative selection rather than a complete list.

How to get started with MatGen

Much of what we build is smaller than the platforms above: an analysis pipeline, a database replacing a folder of spreadsheets, or a dashboard a few people use every day. Navigate here to learn how the engagement process with MatGen works.