Psoriatic Arthritis Biomarker Discovery

Use Case

Biomedical Research · Started April 2026 · Completed May 2026

Project Overview

An exhaustive distributed search across ~426 million gene expression pairs to identify 2-gene classifiers that distinguish psoriatic arthritis (PsA) from psoriasis (Ps), because earlier diagnosis means earlier treatment, before irreversible joint damage occurs.

Statistics
426M total gene pairs to evaluate
100% of search space covered to date
0.927 best AUC found · 2-gene LR classifier
71.2% of top pairs contain KLF5

Search Progress

This search has actively run and completed on the DCP network. What would have taken a single computer 25 years was completed in 3 weeks. Workers evaluated gene pairs in real time and returned results to a central leaderboard. The numbers below reflect the final results.

Summary

  • Search space coverage: 100% complete
  • 426M total pairs
  • 6 Machine Learning classifiers
  • 10-fold CV per pair
  • Leaderboard: top 1,000 retained

Background

Psoriasis (Ps) is a chronic inflammatory skin condition affecting approximately 2–3% of the global population. In up to 30% of cases, patients progress to psoriatic arthritis (PsA), a systemic disease involving joint inflammation and destruction that significantly reduces quality of life and is associated with cardiovascular and metabolic comorbidities.

Distinguishing PsA from Ps at the molecular level is clinically important: earlier identification of patients at high risk of joint progression could enable preventive intervention before irreversible joint damage occurs. The two conditions share overlapping clinical features, making classification challenging by conventional means.

Gene expression microarray data offers a window into the transcriptomic differences between these phenotypes. The GSE57383 dataset (NCBI Gene Expression Omnibus) contains expression profiles for 66 patients, 30 Ps and 36 PsA, across 29,201 probesets from the Affymetrix Human Genome U133 Plus 2.0 array.

Computational Approach

Exhaustive testing of all 2-gene combinations from 29,201 probesets produces approximately 426 million candidate classifiers. Each candidate requires 10-fold cross-validation. Training and testing a classifier 10 times on the 66-patient cohort, making this problem computationally intractable on a single machine.

Pipeline

Input Data: 29,201 probesets · 66 patients · GSE57383
DCP Search: MultiRangeObject · classifier × gene_i × gene_j
Worker Fleet: 10-fold CV · up to 5,000 pairs per slice
Leaderboard: Top 1,000 pairs retained
Analysis: Gene frequency · heatmap · federation

Classifiers Evaluated

Six scikit-learn classifiers were evaluated across the full pair space: Logistic Regression (LR), L1-regularized Logistic Regression (L1), Bagging Classifier (BA), Random Forest (RF), Support Vector Machine (SVM), and Decision Tree (DT). Workers received a (classifier, gene_i, gene_j_batch) tuple via MultiRangeObject and returned only their top-10 pairs by AUC. All six classifiers ran sequentially to completion.

Results

The following results reflect 100% of the total search space, evaluated across all six classifiers. The leaderboard retains the top 1,000 pairs by AUC across all runs.

# AUC ACC CLF Gene 1 Gene 2
1 0.927 0.833 LR LINC00189 KLF5
2 0.918 0.818 LR KLF5 LINC02218
3 0.917 0.864 LR KLF5 L3MBTL4
4 0.916 0.833 LR STK4 KLF5
5 0.914 0.818 LR KLF5 ZNF589
6 0.913 0.788 LR PDXDC1 KLF5
7 0.912 0.894 LR KLF5 CYP4V2
8 0.910 0.803 LR KLF5 PDPK1
9 0.909 0.773 LR KLF5 PRKCB
10 0.909 0.864 LR SNX29P2 PDXDC1

AUC = area under ROC curve · ACC = accuracy · 10-fold CV · n=66 · all 6 classifiers

Key Findings

Across the complete search space evaluated with all six classifiers, three genes emerge with striking consistency. Two are entirely uncharacterized in the PsA literature.

1. KLF5

  • 712 / 1,000 top pairs · best AUC 0.927
  • Krüppel-like Factor 5 is a transcription factor governing sphingolipid metabolism and skin barrier function.

2. LINC00189

  • Top-ranked pair · AUC 0.927 with KLF5
  • A long intergenic non-coding RNA with essentially no published functional characterization.

3. TREML3P

  • 26 / 1,000 top pairs · rank 2 partner to LINC00189
  • A pseudogene in the TREM gene cluster on chromosome 6, adjacent to TREM1 and TREM2.

Next Steps

  1. Replicate on GSE57405 and GSE61281: Run the same exhaustive pair search on two independent Ps vs PsA datasets.

  2. Single-gene baseline evaluation: Evaluate all 29,201 individual genes as solo classifiers to establish a baseline AUC floor.

  3. Guided triple-gene search: Using the top genes stable across all three datasets as seeds, exhaustively test 3-gene combinations.

  4. Treatment response prediction: Extend the classifier framework to predict which patients respond to which biologic class.

  5. Federated multi-site validation: Deploy the search across independent hospital cohorts.

  6. LINC00189 and TREML3P functional characterization: If cross-dataset stability is confirmed, these warrant dedicated functional investigation to understand whether they play a causal role or are markers of an underlying regulatory mechanism.

Project Video

Decentralized computing unlocks Precision Medicine - YouTube