Herrera-Rocha, Fabio; Medina-Ortiz, David; Wyrzykala, Desiree; Sudha, Tharun Srinivasan; Davari, Mehdi D. Multitask Bayesian Neural Networks for Multiparameter Protein Engineering [Preprint] Journal Article In: arXiv, 2026. Abstract | Links | BibTeX | Tags: Artificial neural network, Bayesian probability, Benchmark (surveying), Curse of dimensionality, Dimensionality reduction, DiP-BioCasNavi, DiP-LeFos, Feature engineering, Ranking (information retrieval), Set (abstract data type) @article{herrerarocha2026multitask,Simultaneously engineering multiple protein properties remains a major challenge. Existing machine learning-based pipelines for protein engineering often model properties separately, failing to capture their dependencies and trade-offs. Here, we systematically evaluate how Bayesian parameterization on Multitask Neural Networks can enable robust simultaneous protein engineering under scarce, noisy experimental data. We curated a comprehensive set of 27 multiparameter protein datasets. Then, we compared three algorithm architectures spanning low to full Bayesian parameterization across 16 sequence representations and dimensionality reduction (2,592 models). Bayesian Last Layer models delivered the strongest overall accuracy, generalization, and calibration, ranking as the top-performing model on 70% of benchmark datasets. Dimensionality reduction improved predictive performance by up to 42% and enhanced calibration up to 57% across architectures. Notably, simple One-Hot encoding achieved top performance on 25% of benchmark datasets, particularly with larger datasets. These results establish practical design principles for reliable and data-efficient multiparameter protein engineering. |
Fernández, Diego; García-Vinuesa, Julián Alfonso; Álvarez-Saravia, Diego; Soto-García, Michelle; Medina-Franco, José L.; Sepúlveda-Yáñez, Julieta; Cadet, Xavier F.; Cadet, Frédéric; Davari, Mehdi D.; Uribe-Paredes, Roberto; Herrera-Rocha, Fabio; Medina-Ortiz, David SilkRoute: A Descriptor-Driven Framework for Reproducible Multi-Source Biomolecular Data Acquisition [Preprint] Journal Article In: bioRxiv, 2026. Abstract | Links | BibTeX | Tags: chEMBL, Data acquisition, Data integration, Data retrieval, DiP-BioCasNavi, DiP-LeFos, Identifier, Interoperability, Python (programming language), UniProt, Workflow @article{fernandez2026silkroute,Abstract Background Biomolecular dataset construction often requires coordinated retrieval from heterogeneous repositories, identifier mapping, cross-reference enrichment, source-specific parsing, and provenance recording. These operations are frequently implemented through project-specific scripts, making acquisition procedures difficult to inspect, reproduce, or adapt across studies. We present SilkRoute, an open-source Python framework that formalizes biomolecular data acquisition as descriptor-defined, source-aware, and provenance-tracked workflows, providing a reproducible foundation for multi-source biomolecular dataset construction. Results SilkRoute uses machine-readable YAML descriptors to specify dataset intent, biomolecular modality, workflow mode, query logic, enrichment resources, execution parameters, and export settings. These descriptors drive a common execution model that coordinates primary retrieval and downstream enrichment while preserving source-specific outputs, interaction evidence when available, the original workflow configuration, metadata, and run summaries. We evaluated this model through three representative acquisition scenarios spanning proteins, compounds, and molecular interactions. In the protein-centered workflow, SilkRoute retrieved 2,444 reviewed antimicrobial protein records from UniProt and generated complementary outputs from AlphaFold DB, InterPro, Pathway Commons, and the Protein Data Bank. In the compound-centered workflow, a ChEMBL IC 50 query produced 1,445,939 activity records organized into query-defined potency ranges. In the interaction-centered workflow, 2,253 UniProt protein records were expanded with 902,713 BioGRID interaction records and 5,702 STRING interaction-partner records. Across these scenarios, the framework successfully applied the same descriptor-defined acquisition model to distinct biomolecular entity types, retrieval strategies, enrichment paths, and output structures. Conclusions SilkRoute extends beyond sequence retrieval by providing a reusable acquisition layer for constructing multi-source biomolecular datasets. By separating primary retrieval from enrichment and preserving source-aware outputs together with workflow descriptors and execution metadata, the framework makes acquisition procedures easier to inspect, reproduce, archive, and adapt. SilkRoute does not replace biological curation, label validation, deduplication, partitioning, or benchmarking, but provides structured and traceable acquisition packages that support these downstream processes. |
Herrera-Rocha, Fabio; Medina-Ortiz, David; Mauz, Fabian; Pleiss, Juergen; Davari, Mehdi D. Best Practices for Machine Learning-Assisted Protein Engineering Journal Article In: JOURNAL OF CHEMICAL INFORMATION AND MODELING, vol. 65, no. 23, pp. 12655-12667, 2025, ISSN: 1549-9596. Links | BibTeX | Tags: DiP-LeFos, Machine learning, Optimization, Protein engineering @article{WOS:001616552400001, |
MDR, Sachsen-Anhalt heute Nachhaltige Forschung: Wie man Zuckerrüben noch besser verwerten kann Miscellaneous 2025. Abstract | Links | BibTeX | Tags: DiP-LeFos, Presse @misc{Video002,Nur 10 Prozent der Rübe wird für die Zuckergewinnung verwendet, den Rest nutzen wir gar nicht zur Ernährung. Dabei könnten wir es. Sven Stephan hat es sich von einer Lebensmittelchemikerin der Uni Halle erklären lassen. |