Psoriatic Arthritis Biomarker Discovery
Use Case
Biomedical Research · Started April 2026 · Completed May 2026
Project Overview
An exhaustive distributed search across ~426 million gene expression pairs to identify 2-gene classifiers that distinguish psoriatic arthritis (PsA) from psoriasis (Ps), because earlier diagnosis means earlier treatment, before irreversible joint damage occurs.
Statistics
426M total gene pairs to evaluate
100% of search space covered to date
0.927 best AUC found · 2-gene LR classifier
71.2% of top pairs contain KLF5
Search Progress
This search has actively run and completed on the DCP network. What would have taken a single computer 25 years was completed in 3 weeks. Workers evaluated gene pairs in real time and returned results to a central leaderboard. The numbers below reflect the final results.
Summary
- Search space coverage: 100% complete
- 426M total pairs
- 6 Machine Learning classifiers
- 10-fold CV per pair
- Leaderboard: top 1,000 retained
Background
Psoriasis (Ps) is a chronic inflammatory skin condition affecting approximately 2–3% of the global population. In up to 30% of cases, patients progress to psoriatic arthritis (PsA), a systemic disease involving joint inflammation and destruction that significantly reduces quality of life and is associated with cardiovascular and metabolic comorbidities.
Distinguishing PsA from Ps at the molecular level is clinically important: earlier identification of patients at high risk of joint progression could enable preventive intervention before irreversible joint damage occurs. The two conditions share overlapping clinical features, making classification challenging by conventional means.
Gene expression microarray data offers a window into the transcriptomic differences between these phenotypes. The GSE57383 dataset (NCBI Gene Expression Omnibus) contains expression profiles for 66 patients, 30 Ps and 36 PsA, across 29,201 probesets from the Affymetrix Human Genome U133 Plus 2.0 array.
Computational Approach
Exhaustive testing of all 2-gene combinations from 29,201 probesets produces approximately 426 million candidate classifiers. Each candidate requires 10-fold cross-validation. Training and testing a classifier 10 times on the 66-patient cohort, making this problem computationally intractable on a single machine.
Pipeline
Input Data: 29,201 probesets · 66 patients · GSE57383
DCP Search: MultiRangeObject · classifier × gene_i × gene_j
Worker Fleet: 10-fold CV · up to 5,000 pairs per slice
Leaderboard: Top 1,000 pairs retained
Analysis: Gene frequency · heatmap · federation
Classifiers Evaluated
Six scikit-learn classifiers were evaluated across the full pair space: Logistic Regression (LR), L1-regularized Logistic Regression (L1), Bagging Classifier (BA), Random Forest (RF), Support Vector Machine (SVM), and Decision Tree (DT). Workers received a (classifier, gene_i, gene_j_batch) tuple via MultiRangeObject and returned only their top-10 pairs by AUC. All six classifiers ran sequentially to completion.
Results
The following results reflect 100% of the total search space, evaluated across all six classifiers. The leaderboard retains the top 1,000 pairs by AUC across all runs.
| # | AUC | ACC | CLF | Gene 1 | Gene 2 |
|---|---|---|---|---|---|
| 1 | 0.927 | 0.833 | LR | LINC00189 | KLF5 |
| 2 | 0.918 | 0.818 | LR | KLF5 | LINC02218 |
| 3 | 0.917 | 0.864 | LR | KLF5 | L3MBTL4 |
| 4 | 0.916 | 0.833 | LR | STK4 | KLF5 |
| 5 | 0.914 | 0.818 | LR | KLF5 | ZNF589 |
| 6 | 0.913 | 0.788 | LR | PDXDC1 | KLF5 |
| 7 | 0.912 | 0.894 | LR | KLF5 | CYP4V2 |
| 8 | 0.910 | 0.803 | LR | KLF5 | PDPK1 |
| 9 | 0.909 | 0.773 | LR | KLF5 | PRKCB |
| 10 | 0.909 | 0.864 | LR | SNX29P2 | PDXDC1 |
AUC = area under ROC curve · ACC = accuracy · 10-fold CV · n=66 · all 6 classifiers
Key Findings
Across the complete search space evaluated with all six classifiers, three genes emerge with striking consistency. Two are entirely uncharacterized in the PsA literature.
1. KLF5
- 712 / 1,000 top pairs · best AUC 0.927
- Krüppel-like Factor 5 is a transcription factor governing sphingolipid metabolism and skin barrier function.
2. LINC00189
- Top-ranked pair · AUC 0.927 with KLF5
- A long intergenic non-coding RNA with essentially no published functional characterization.
3. TREML3P
- 26 / 1,000 top pairs · rank 2 partner to LINC00189
- A pseudogene in the TREM gene cluster on chromosome 6, adjacent to TREM1 and TREM2.
Next Steps
Replicate on GSE57405 and GSE61281: Run the same exhaustive pair search on two independent Ps vs PsA datasets.
Single-gene baseline evaluation: Evaluate all 29,201 individual genes as solo classifiers to establish a baseline AUC floor.
Guided triple-gene search: Using the top genes stable across all three datasets as seeds, exhaustively test 3-gene combinations.
Treatment response prediction: Extend the classifier framework to predict which patients respond to which biologic class.
Federated multi-site validation: Deploy the search across independent hospital cohorts.
LINC00189 and TREML3P functional characterization: If cross-dataset stability is confirmed, these warrant dedicated functional investigation to understand whether they play a causal role or are markers of an underlying regulatory mechanism.
Project Video
Decentralized computing unlocks Precision Medicine - YouTube