The tumor microbiome has emerged as a potential contributor to cancer biology, yet its clinical relevance in prostate adenocarcinoma (PRAD) remains incompletely understood. In this study, we performed an integrative analysis of publicly available microbiome, transcriptomic, and clinical datasets to identify tumor-associated bacterial genera and characterize host molecular alterations associated with bacterial abundance. Microbial abundance profiles were obtained from the Bacteria in Cancer (BIC) database, while transcriptomic and clinical data were collected from The Cancer Genome Atlas (TCGA)-PRAD cohort. Differential abundance analysis identified 15 bacterial genera that differed significantly between 493 primary tumors and 52 adjacent normal tissues (FDR < 0.05).
Among these, Paenibacillus exhibited the greatest enrichment in tumor tissues (log₂FC = 3.04, FDR = 2.54 × 10⁻⁷) and was the only differentially abundant genus significantly associated with both Gleason score (Spearman's ρ = 0.163, P = 9.06 × 10⁻⁴) and pathological stage (P = 9.85 × 10⁻³). Transcriptomic analysis identified 17,555 differentially expressed genes between tumor and normal tissues, including 1,215 genes with |log₂FC| > 1. Stratification of tumors by Paenibacillus abundance identified 1,079 differentially expressed genes (FDR < 0.05 and |log₂FC| > 0.5). Functional enrichment analysis demonstrated that tumors with higher Paenibacillus abundance were characterized by enrichment of MYC target and oxidative phosphorylation signatures, whereas inflammatory response, TNFα signaling via NF-κB, IL6/JAK/STAT3 signaling, hypoxia, TGF-β signaling, and epithelial–mesenchymal transition were relatively enriched in tumors with lower Paenibacillus abundance.
Collectively, these findings identify Paenibacillus as a candidate tumor-associated bacterial genus associated with clinicopathological characteristics and distinct host transcriptomic programs in PRAD. More broadly, this study demonstrates the utility of integrating microbiome, transcriptomic, and clinical data to systematically prioritize candidate tumor-associated microorganisms using publicly available multi-omics datasets and provides a foundation for future experimental validation.