Source-only snapshot of the cluster branch for WES execution. Large reference files (HPA/MANE/ensemble FASTA, model weights, ~597 MB) are omitted: they are baked into the container images at build time and mounted from the dreamdock-data PVC at runtime, and exceed the Gitea request size limit. Pipeline entry point is main.nf, which orchestrates the biotransformer, conplex and tissue modules as a single workflow. Ligand inputs are read from the eureka workspace; protein_zarr and chembl_db come from the dreamdock-data PVC.
6.3 KiB
CLAUDE.md
This file provides guidance to Claude Code (claude.ai/code) when working with code in this repository.
What This Is
Digital Trials is a Nextflow pipeline that simulates drug-protein interactions for digital clinical trials. Given a drug (as SMILES), it predicts metabolites, screens the drug and its metabolites against a proteome, performs pathway enrichment, and maps tissue distribution of interacting proteins.
Running the Pipeline
# Local test run (uses params defined in main.nf)
nextflow run main.nf
# With parameter file for cluster runs
nextflow run main.nf -params-file params/input_to_run.json
# Kubernetes via Tower
export TOWER_ACCESS_TOKEN=68358d90995ae27fe78b4c1818f9f0097d834b00
nextflow run main.nf -profile k8s -with-tower https://tower.stg.cluster.omic.ai/api -params-file params/input_to_run.json
Building Containers
# Build and tag individual containers via docker-compose
docker-compose build tissue
docker-compose build chembl
docker-compose build network
# Images push to harbor.cluster.omic.ai/omic/digitaltrials/
Each pipeline step has its own Dockerfile: Dockerfile_biotransformer, Dockerfile_conplex, Dockerfile_tissue, Dockerfile_chembl, Dockerfile_network.
Pipeline Architecture
The pipeline is orchestrated in main.nf (entry point and parameter definitions) and composed from three module files:
Stage 1: Metabolite Prediction (main_biotransformer.nf)
- HUMAN_TRANSFORMER — Runs BioTransformer (Java) to predict human metabolites from input SMILES. Only
HUMANmode is fully implemented. - GET_FINAL_METABOLITES_STATIC — Filters to terminal metabolites (not precursors), checks against local ChEMBL SQLite DB and PubChem, outputs canonical SMILES list.
Stage 2: Proteome Screening (main_conplex.nf)
- PREPROCESS_PROTEIN — Merges patient FASTA files and projects them into a Zarr vector store (GPU, uses
/app/project.pyinside container). - CONPLEX — Screens drug + metabolite SMILES against protein Zarr DBs using ConPLex model. Runs
/app/convert.py,/app/screen.py,/app/get_round_2.pyinside themetabolite-screencontainer. Outputs drug scores and significant interactions TSVs. - NETWORK_ENRICHMENT — Queries STRING DB API for pathway enrichment and protein-protein interactions from significant hits.
Stage 3: Tissue & Biological Properties (main_tissue.nf)
- TISSUE_DISTRIBUTION — Maps interacting proteins to tissue expression using HPA data (
drug_tissue_distribution.py). - BIO_METRICS — Aggregates biological properties across metabolites, drug scores, and interactions (
digital_patient_extract_metrics.py).
Data Flow
Input CSVs (SMILES+TARGET) → HUMAN_TRANSFORMER → GET_FINAL_METABOLITES_STATIC
↓
Patient FASTA → PREPROCESS_PROTEIN → Zarr ──→ CONPLEX (drug + metabolites vs proteome)
Pre-built Zarr ─────────────────────────↗ ↓ ↓
NETWORK_ENRICHMENT TISSUE_DISTRIBUTION
↓
BIO_METRICS
Channels are joined by ligand ID (CSV simpleName), ensuring each drug's metabolites are matched to the correct ConPLex screening run.
Key Parameters (in main.nf)
| Parameter | Purpose | Default |
|---|---|---|
params.mode |
BioTransformer mode — use HUMAN only |
HUMAN |
params.threshold |
ConPLex binding score cutoff for significant interactions | 0.65 |
params.protein_network_threshold |
STRING DB enrichment cutoff | 0.65 |
params.keep_enst |
Keep per-protein Zarr results from ConPLex | false |
params.ligands |
Input directory of per-compound CSV files | — |
params.mutated_protein_fasta |
Directory of patient-specific FASTA files (or blank) |
— |
params.protein_zarr |
Pre-built protein sequence Zarr DB | — |
params.chembl_db |
Path to ChEMBL SQLite database file | — |
params.outdir |
Output root directory | — |
params.bt_initial_memory |
Biotransformer starting memory (GB) | 5 |
params.bt_growth_memory |
Biotransformer additional memory per retry (GB) | 15 |
params.bt_max_retries |
Biotransformer max retry attempts | 10 |
params.bt_fail_action |
Biotransformer failure strategy: 'terminate' or 'ignore' |
'terminate' |
params.conplex_initial_memory |
ConPLex starting memory (GB) | 5 |
params.conplex_growth_memory |
ConPLex additional memory per retry (GB) | 15 |
params.conplex_max_retries |
ConPLex max retry attempts | 1 |
params.conplex_fail_action |
ConPLex failure strategy: 'terminate' or 'ignore' |
'ignore' |
Nextflow Gotchas
- Params loaded from JSON are strings — use
.toInteger()when used in resource directives (e.g.,memory,maxRetries) main.nfdefines all default param values; params files only need to override what they change
Sub-pipeline: Metabolite Screen Adaptive
nf_metabol_screen_adaptive/ is a standalone Nextflow pipeline for ConPLex-based metabolite screening. It has its own main.nf, test.nf, Dockerfile, and nextflow.config. Used for development/testing of the screening step independently.
Container Registry
All images are at harbor.cluster.omic.ai/omic/digitaltrials/. Container refs in main.nf use sha256 digests for reproducibility (except tissue which uses version tags). When updating containers, pin to a digest or explicit version tag.
Input Format
Per-compound .csv files named by identifier (e.g., InChIKey). 4-line key-value format:
SMILES
<SMILES string>
TARGET
<ENST transcript ID or empty>
Input directories are under /data/digital-trials-data/ (e.g., input_to_run/, input_knowen_target/, input_new/input/).
See also /data/runs/docs/digital-trials-input-format.md for format details.
Output Structure
Results are organized under {outdir}/{project_name}/:
1_biotransformer/— Raw metabolite predictions1b_final_metabolites/— Filtered terminal metabolites2_conplex/— Drug scores and significant interaction TSVs3_string/— Network enrichment and interaction TSVs4_tissue_distribution/— Tissue distribution and biological properties TSVs