Evolution-informed methods
EvoClin’s frameworks for myelodysplastic syndromes and colorectal cancer are built on a set of methodological components designed to capture cancer evolution and translate it into interpretable risk models. This page provides an overview of these components.
From genomes to evolutionary trajectories
Many genomic analyses treat mutations as static features: present or absent at a given time. EvoClin’s methods are different by design. They explicitly model evolutionary dynamics, asking how mutations tend to appear over time and how they interact to shape disease progression.
Building on previous work in cancer evolution, EvoClin’s approach typically starts from cross-sectional or longitudinal genomic data and reconstructs cohort-level trajectories, identifying early, intermediate and late events and their directional relationships.
Core methodological components
While the exact implementation is tailored to each indication, EvoClin’s frameworks share a number of core components:
- Data preprocessing and harmonisation – variant calling and copy number profiles are curated, filtered and harmonised across cohorts and centres, often starting from real-world sequencing panels.
- Cancer cell fraction–based ordering – when available, cancer cell fractions or related quantities are used to infer likely temporal orderings between events at the patient or cohort level.
- Inference of evolutionary trajectories – probabilistic or graph-based models are employed to reconstruct predominant evolutionary routes and identify directional relationships between alterations.
- Evolution-informed feature engineering – from these trajectories, specific patterns (e.g. ordered mutation pairs, co-occurring events, “early” vs “late” status of key genes) are turned into candidate features.
- Interpretable risk modelling – regularised Cox proportional hazards models and logistic regression are then used to select a limited subset of these features and to build risk scores for endpoints such as survival and metastasis.
Interpretability by design
EvoClin deliberately favours models that remain interpretable. Each feature, whether a single alteration or an evolutionary pattern, has a clear coefficient and effect size, making its contribution to risk explicit. This is crucial for:
- clinical interpretability and trust;
- hypothesis generation and mechanistic follow-up studies;
- potential integration with existing guidelines and risk scores.
From methods to frameworks
These methodological components are instantiated in indication-specific frameworks:
- In ProgEvo and IPSS-M-Evo, evolutionary modelling in myelodysplastic syndromes is used to derive a small set of evolution-based variables that extend IPSS-M and improve risk stratification for leukemia-free and overall survival.
- In ProgMet (PROGMET framework), evolution-aware representation and feature selection are applied to large colorectal cancer cohorts to identify prognostic and metastatic biomarkers and to derive dual risk scores and patient groups.
As EvoClin expands to additional indications and datasets, the same methodological principles will be adapted and extended, always with the goal of keeping evolution at the centre and models interpretable.