# ProFormer: A Transformer-Based Deep Learning Pipeline for Robust Classification of Proteomic Samples
## Introduction
Proteomic datasets present unique challenges for machine learning applications. Modern mass spectrometry-based proteomics routinely identifies thousands of peptide ions per sample, generating high-dimensional data tables that are inherently prone to overfitting when fed into deep neural networks. Traditional approaches have relied on image-based representations of proteomic data—converting peptide signals into heatmaps through fixed grid binning—but this strategy discards potentially valuable information and limits the number of measurable features to just three dimensions.
A transformative solution has emerged in the form of a transformer-based deep learning pipeline that operates directly on tabular peptide ion data. This approach leverages a learnable patch merger to compress high-dimensional proteomic information into compact data tokens, which are then processed through a series of transformer encoder blocks before being classified by a terminal single-layer perceptron. The architecture processes four MS1-derived features simultaneously—mass-to-charge ratio (m/z), retention time (RT), indexed ion mobility (iIM), and intensity—making it uniquely versatile across different data acquisition modes.
Unlike image-based pipelines that impose rigid spatial structures on proteomic signals, this tabular transformer framework preserves the full dimensionality of raw peptide ion profiles while dynamically reducing them through a weighted aggregation mechanism. The result is a classification model that demonstrates remarkable robustness against batch effects, generalizes across unseen biological conditions, and achieves state-of-the-art performance in diverse proteomic applications ranging from bulk cell samples to single-cell analyses and clinical plasma diagnostics.
## Architecture and Methodology
### The Patch Merger: Dynamic Dimensionality Reduction
At the core of the pipeline lies a learnable patch merger module that fundamentally reimagines how high-dimensional proteomic data is prepared for deep learning. Rather than binning peptide signals into a fixed grid structure—as is done when converting proteomic tables into image-like heatmaps—the patch merger divides the full dataset into subsets called patches, each of which summarizes the information content of its constituent peptide ions.
The patch merger operates by computing dot products between input signal vectors and a set of learnable weight vectors, then applying a softmax function to generate attention weights. These weights determine how each peptide ion contributes to each output patch, effectively creating a weighted aggregation of features localized in relationship to dynamically determined reference points in the m/z, retention time, ion mobility, and intensity feature space. Because these weights are optimized during model training against the classification loss function, the patch merger learns to focus on peptide ion regions that are most discriminative for the task at hand.
The output of the patch merger takes the form of data tokens—compact representations of peptide ion subsets—that preserve the information content of the original high-dimensional table. Importantly, this process does not require positional encoding of tokens, since the tabular format inherently preserves the independence of each sample’s features from their input ordering. This design choice allows the downstream transformer layers to identify class-specific relationships without being influenced by artificial sequence constraints.
### Transformer Encoder Architecture
Once the patch merger generates data tokens, a concatenated classification (CLS) token is appended to serve as a readout anchor for the final classification decision. The combined token set is then passed through multiple transformer encoder blocks, each consisting of a multi-head attention module followed by a two-layer perceptron, with interspersed dropout and layer normalization for regularization and training stability.
These encoder blocks iteratively refine the peptide ion-derived tokens by modulating inter-signal dependencies through self-attention mechanisms, allowing the model to capture complex, global relationships across the entire proteomic profile. The final representation of the CLS token is then fed into a single-layer perceptron that outputs a binary prediction score between 0 and 1, indicating the predicted class membership of the sample.
The number of patches generated by the merger and the number of transformer layers are both learnable hyperparameters, enabling the architecture to adapt its complexity to the specific characteristics of each dataset.
## Batch-Robust Classification and Dose-Dependent Generalization
### Benchmarking Against Traditional and Deep Learning Approaches
The classification performance of this transformer-based pipeline has been rigorously evaluated against a comprehensive suite of alternative approaches. Traditional machine learning algorithms—including logistic regression, random forest, and XGBoost—were tested on the same proteomic data, both with and without deterministic dimensionality reduction methods such as wavelet transforms and K-means clustering. While wavelet transforms consistently outperformed K-means clustering across all tested ML algorithms, the best logistic regression model achieved only approximately 65% area under the curve (AUC), demonstrating the inherent limitations of classical approaches when confronted with the complexity of proteomic data.
Among deep learning alternatives, the transformer-based pipeline achieved the highest AUC of approximately 84%, substantially outperforming all competing architectures. This result was particularly notable given that the evaluation was conducted on temporally separated test data collected six months after the training set—a scenario designed to simulate real-world batch effects that frequently undermine model reproducibility in proteomics.
### Robustness to Technical Domain Shifts
Conventional convolutional neural network (CNN) architectures, including residual networks and efficient networks, showed dramatic performance degradation when evaluated on test data from different acquisition batches. Some CNN models exhibited drops of 20–28 percentage points in accuracy between validation and test sets, raising serious concerns about overfitting in previously reported results. Vision transformers fared somewhat better but still showed discrepancies of 6–23 percentage points between validation and test performance.
In stark contrast, the transformer-based pipeline with its patch merger component achieved essentially equivalent results on both validation and test datasets. This finding demonstrates that the combination of tabular input processing with transformer-based attention mechanisms produces models that are inherently more robust to technical variations in measurement conditions, instrument performance, and acquisition timing.
### Multi-Source Domain Generalization
The pipeline’s generalization capabilities extend beyond simple batch effects to encompass more complex domain shifts. In experiments designed to test generalization across unseen biological conditions—different cell cultivation protocols, varying operator techniques, and distinct biological passages—the transformer-based model maintained AUC values as high as 97%, far exceeding the performance of all competing architectures.
Perhaps most impressively, the model demonstrated dose-response generalization when confronted with interferon gamma (IFN-γ) concentrations not present in the training data. Only the transformer-based pipeline exhibited a positive monotonic relationship between treatment concentration and predicted output, with a Spearman correlation of +0.72 and a coefficient of determination (R²) of 0.77 for the dose-response curve. The model successfully recognized IFN-γ treatment effects at concentrations as low as 9.44 ng/mL—approximately one order of magnitude below the lowest concentration used during training—achieving a limit of detection well beyond what was explicitly taught during the training phase.
Other architectures showed either no correlation or negative trends in their dose-response predictions, indicating fundamental failures in their ability to generalize beyond their binary training objective. These results confirm that the tabular transformer approach not only excels at precise binary classification but also captures continuous biological gradients in a biologically interpretable manner.
## The Learnable Patch Merger in Detail
### Information-Preserving Data Aggregation
Analysis of the attention weights generated by the patch merger reveals how this component achieves its remarkable performance. The learned weights organize peptide ion signals into distinct clusters that correspond to functionally relevant regions of the proteomic feature space. Rather than representing random subsets of all detected peptide ions, each patch aggregates signals that share regional similarities in their m/z, retention time, and intensity characteristics.
These attention weight clusters exhibit consistent formation across different test samples regardless of their class membership, suggesting that the patch merger identifies a stable and biologically meaningful decomposition of the proteomic data space. The patches operate with variable borders and numbers, unlike the rigid grid constraints of image-based approaches, and can simultaneously focus on multiple distinct regions of the feature space within a single sample.
### Enhanced Utility for Downstream Machine Learning
An important practical finding is that the output tokens generated by the patch merger, when used as input to traditional machine learning algorithms, substantially improve classification performance compared to using raw wavelet-transformed or K-means-clustered features. A random forest trained on patch merger outputs achieved up to 80% AUC, surpassing the best wavelet-transform-based ML approach and demonstrating that the learned dimensionality reduction captures more discriminative information than any fixed transformation function.
This cross-algorithm benefit is significant because it means the patch merger is not exclusively valuable within the transformer architecture but serves as a general-purpose feature engineering tool for proteomic data analysis. The dynamic, learnable nature of this aggregation step represents a meaningful advance over static dimensionality reduction techniques that have been standard in the proteomics field.
## Single-Cell Proteomics Classification
### Broad Applicability Across Cell Types and States
The transformer-based pipeline has been successfully applied to single-cell proteomics (SCP) data, extending its utility beyond traditional bulk proteomics experiments. Single-cell proteomics presents unique challenges, including lower peptide ion counts per cell, higher rates of missing values, and greater technical variability—all factors that can degrade model performance.
Despite these challenges, the pipeline achieved AUC values ranging from 73.6% to 100% across diverse single-cell classification tasks, including the discrimination of different cell types, cell cycle stages, and differentiation states. The highest accuracies were observed for tasks involving highly divergent cell populations—such as distinguishing completely different cell types or distinct cell cycle phases—where the biological differences between conditions are most pronounced.
### Performance on Subtle Biological Differences
When confronted with more subtle biological distinctions, such as differentiating between migratory and stationary cells or classifying pluripotent states of mouse embryonic stem cells, the pipeline still achieved respectable accuracies of 82.1% and 80.8%, respectively. These results are particularly meaningful given that discriminating altered states of the same cell type is inherently more challenging than contrasting fundamentally different cell types.
Analysis of classification difficulty using t-distributed stochastic neighbor embedding (tSNE) revealed a clear pattern: predictions aligned closely with true labels when cell populations were well-separated in feature space, but became more uncertain in regions of biological overlap. This finding underscores the importance of deep learning classifiers for proteomic data, as simple dimensionality reduction techniques like t-SNE or UMAP alone cannot resolve these overlapping distributions.
### Optimizing Data Quality and Training Set Size
Systematic evaluation of how training data composition affects model performance revealed several practical strategies for improving classification accuracy. Training on low-abundance peptide ions alone produced inferior results compared to models that also incorporated higher-abundance signals. However, for single-cell data specifically, the optimal strategy involved including up to the 60% highest-abundance peptide ions, beyond which additional low-intensity signals degraded performance—likely due to their higher rates of missing values and quantitative inconsistencies.
Remarkably, the pipeline demonstrated consistent performance even with very small training sets. Training on as few as ten samples still achieved AUC values of approximately 85% for bulk experiments and 75% for single-cell data. While larger training sets gradually improved single-cell performance—reaching 80.8% AUC when all available samples were used—the ability to achieve strong results with limited training data makes the pipeline particularly valuable for studies where sample availability is constrained.
### Resilience to Data Noise
The pipeline demonstrated exceptional robustness to artificial noise injection. When Gaussian noise was systematically added to test set features, the transformer-based model maintained its classification accuracy even at noise levels of 10%, while a competing CNN architecture degraded to random classification performance at the same noise level. In the bulk experiment setting, even a minimal noise ratio of just 0.01% caused a dramatic drop in CNN performance from 79.9% to below 60% AUC, whereas the transformer-based pipeline remained completely unaffected.
This noise resilience has important implications for real-world proteomic data analysis, where technical variability and measurement uncertainty are inherent features of the data. The demonstrated ability to maintain performance in the presence of noise validates the pipeline’s potential for classifying biological outliers and analyzing samples that deviate from standard acquisition protocols.
## Clinical Diagnostics: Plasma Proteome Classification
### COVID-19 Severity Discrimination
The pipeline’s applicability to clinical diagnostics was demonstrated using plasma proteome data from the COMBAT consortium, a large-scale multi-omics study comprising 490 participants. The task involved binary classification of severe COVID-19 patients against healthy volunteers and individuals with other severe conditions, including mild COVID-19, sepsis, and severe influenza.
Analyzing only the intensity feature produced initial AUC values of 72.7% on the validation set. However, a critical preprocessing innovation—dividing the m/z feature into its separate mass and charge components—dramatically improved performance to an AUC of 86.3% on the test set, with an F1 score of 76.7% at the optimized classification threshold.
### Diagnostic Performance Characteristics
By setting the prediction cutoff slightly below 0.5, the optimized model achieved 100% sensitivity for severe COVID-19 detection, correctly identifying every positive case while ruling out the disease in 46% of cases. This high recall profile is precisely what would be desired in a screening diagnostic, where the cost of false negatives is clinically unacceptable.
When evaluating isolated discriminations between severe COVID-19 and individual alternative conditions, the model achieved AUC values of 80.6% for mild COVID-19, 87.5% for sepsis, and 98.3% for other severe infections. These results demonstrate that the model can effectively distinguish the proteomic signature of severe COVID-19 from a diverse panel of conditions that share overlapping clinical presentations.
### Explaining Model Decisions with SHAP Values
To build confidence in the clinical applicability of the model, explainability analysis was performed using SHAP (SHapley Additive exPlanation) values. This analysis revealed that all four implemented features—retention time, mass, charge, and intensity—contributed significantly to model predictions, with the mass feature exhibiting the highest individual impact, likely reflecting the high precision with which mass measurements are obtained in mass spectrometry.
At the peptide ion level, SHAP value aggregation identified specific signals that most strongly influenced the model’s decisions. A subset of top-contributing peptide ions achieved equivalent classification accuracy when used to train a simple random forest model, demonstrating that the transformer architecture successfully identified the most diagnostically informative features from among thousands of candidates.
Biological validation of the SHAP-derived insights confirmed that the most impactful peptide ions originated from proteins known to be involved in COVID-19 pathology, including complement system proteins (C3, C4A), coagulation factors (SERPINC1, fibrinogen chains FGA, FGB, FGG), and markers of tissue injury. Notably, the analysis identified a specific C3 peptide (residues 567–592) that marked the beginning of the C3-beta-c fragment, suggesting proteolytic cleavage through complement activation—a biological process that would have been missed by conventional protein-level differential expression analysis.
## Frequently Asked Questions (FAQ)
**Q1: What makes the ProFormer pipeline different from traditional deep learning approaches for proteomic data?**
A1: The key innovation lies in how proteomic data is preprocessed and fed into the model. Traditional approaches convert peptide ion information into image-like heatmaps using fixed grid binning, which is inherently limited to three features and imposes artificial spatial constraints. The ProFormer pipeline operates directly on tabular data using a learnable patch merger that dynamically reduces dimensionality while preserving information across all available features. This allows the model to incorporate four or more MS1-derived features simultaneously and adapt its aggregation strategy to the specific characteristics of each dataset.
**Q2: Why is the patch merger considered learnable, and how does this benefit classification?**
A2: The patch merger uses weight vectors that are optimized during model training, meaning the algorithm learns which peptide ion signals are most relevant for the classification task. Unlike fixed transformations such as wavelet transforms or K-means clustering, the patch merger adjusts its aggregation strategy to minimize the model’s classification loss. This learnability enables the merger to focus on biologically discriminative regions of the proteomic feature space, producing compact data tokens that are highly informative for downstream classification.
**Q3: How does the model handle batch effects between training and test data?**
A3: The combination of tabular input processing, patch-based dimensionality reduction, and transformer-based self-attention mechanisms produces a model that is inherently robust to batch effects. In benchmark evaluations, the pipeline achieved equivalent performance on validation and test sets collected months apart, while CNN architectures showed accuracy drops of 20–28 percentage points. The transformer’s global attention mechanism allows it to identify class-specific peptide ion relationships independent of technical covariates such as acquisition time or instrument variability.
**Q4: Can this pipeline be applied to datasets with very few samples?**
A4: Yes. One of the pipeline’s most notable practical advantages is its ability to achieve strong classification performance with relatively small training sets. Experiments demonstrated that training on as few as ten samples still yielded AUC values above 85% for bulk proteomics data, making the pipeline suitable for studies with limited sample availability, such as rare clinical cohorts or expensive single-cell experiments.
**Q5: What explains the model’s ability to detect IFN-γ effects at concentrations far below those seen during training?**
A5: The transformer’s self-attention mechanism captures complex, non-linear relationships between peptide ion profiles and treatment conditions. Rather than learning a simple threshold-based response, the model learns the continuous biological gradient of the IFN-γ dose-response relationship. This allows it to interpolate and extrapolate treatment effects to concentrations that were never explicitly represented in the training data, achieving a limit of detection approximately one order of magnitude below the lowest training concentration.
**Q6: How reliable are the SHAP-based explanations of the model’s predictions?**
A6: The SHAP analysis in this context has been validated through biological correspondence. Peptide ions identified as highly influential by SHAP values mapped directly to proteins and biological pathways known to be involved in the classification task—for example, complement system activation in COVID-19 severity prediction. Furthermore, a random forest classifier trained exclusively on SHAP-identified top-contributing peptide ions achieved performance equivalent to the full deep learning model, confirming that the explanations identify genuinely informative signals rather than spurious correlations.
**Q7: Is this approach limited to binary classification tasks?**
A7: While the core architecture demonstrated in these evaluations uses a binary classifier with output scores ranging from 0 to 1, the underlying transformer framework is inherently extensible to multi-class classification tasks. The patch merger and transformer encoder components are agnostic to the number of output classes, and modifications to the terminal classifier module would enable application to scenarios with three or more sample categories.
**Q8: How does the model perform when peptide ion counts vary widely between datasets?**
A8: The pipeline demonstrates consistent performance across datasets with dramatically different peptide ion counts, ranging from as few as 763 peptide ions in some single-cell studies to approximately 145,000 in bulk experiments. The learnable patch merger is designed to accommodate variable input dimensions, and the model’s performance scales with the biological divergence of the classes being distinguished rather than being strictly dependent on data dimensionality.
**Q9: What preprocessing steps are required before feeding data into the pipeline?**
A9: Raw proteomic data must first be processed through standard identification and quantification software (such as DIA-NN or MaxQuant) to extract MS1 features including m/z, retention time, ion mobility (if available), and intensity values per peptide ion per sample. Missing values should be addressed through appropriate normalization strategies—for single-cell data, median normalization followed by Z-score standardization has been shown to be effective. The masked peptide ion list (without prior biological interpretation) is then formatted as a tabular input for the pipeline.
**Q10: Can the pipeline be used for regression tasks rather than classification?**
A10: The current architecture is designed for binary classification with a sigmoid output. Extending it to regression tasks—such as predicting continuous biomarker levels or disease severity scores—would require modifying the terminal classification module to output continuous values rather than probabilities, along with appropriate loss function adjustments. The core patch merger and transformer components would remain applicable for such extensions.
## Conclusion
The integration of transformer architectures with learnable patch merging represents a significant advance in deep learning for proteomic data analysis. By processing tabular peptide ion data directly, the pipeline overcomes the dimensionality limitations of image-based approaches while simultaneously achieving superior robustness against batch effects and technical domain shifts. The model’s ability to generalize across unseen biological conditions, recognize dose-dependent treatment effects, and maintain high performance with limited training data establishes it as a practical and powerful tool for diverse proteomic applications.
The successful extension of this approach to single-cell proteomics and clinical plasma diagnostics demonstrates its versatility across sample types and biological contexts. The incorporation of SHAP-based explainability provides a crucial pathway toward transparent and trustworthy deep learning in proteomics, enabling researchers to understand not just what the model predicts but why it makes those predictions—and to connect those predictions back to underlying biological mechanisms.
As proteomic technologies continue to evolve and generate increasingly complex datasets, methods that can effectively extract biologically meaningful patterns while remaining robust to technical variability will be essential. The tabular transformer approach with learnable patch merging offers a compelling solution to these challenges, and its demonstrated performance across bulk, single-cell, and clinical proteomics suggests that similar architectures will find broad adoption in the proteomics and broader biomedical research communities.
Thank you for reading



