The Consortium of Molecular Design at BYU provides cutting edge interdisciplinary research opportunities for students to push the envelope for protein engineering and drug discovery.
We use close collaboration between laboratories at BYU in Physics, Chemistry, Computer Science, LifeSciences, and Engineering to tackle these challenging topics from all angles.
We actively seek industrial collaboration and support for our efforts and are excited to explore mutually beneficial application of all state-of-the-art technologies to revolutionize molecular design.
News and Events
Selected Publications
Despite significant advances in artificial intelligence (AI) algorithms for prostate cancer detection from whole slide images, the clinical applicability of these models remains limited. Variability in inter- and intrapathological grading, low generalizability of training datasets, and insufficient annotation precision restrict the performance of downstream models. This article introduces a novel Bayesian framework that addresses these challenges by generating pixel-wise posterior distributions, thereby providing a probabilistic output that enables the simulation of a panel consensus and enabling seamless integration of new data and models as they become available. The framework is demonstrated by integrating a Bayesian prior with a trained AI model to produce a per-pixel distribution of Gleason patterns. It is shown that using this distribution of Gleason patterns rather than a ground-truth label can improve model applicability, mitigate errors, and highlight areas of interest for pathologists. Furthermore, we present a high-quality, hand-curated dataset of prostate histopathological images annotated at the gland level by trained premedical students and verified by an expert pathologist. We highlight the potential of this adaptive and uncertainty-aware framework for developing clinically deployable AI tools that can support pathologists in accurate prostate cancer grading, improve diagnostic accuracy, and create positive patient outcomes. This work is presented as an early-stage, proof-of-concept study; the framework has not been validated for clinical use and is not intended for diagnostic deployment in its current form.
Artificial intelligence foundation models are increasingly deployed for prostate cancer Gleason grading, where GP3/GP4 distinction directly impacts treatment decisions (active surveillance vs. intervention). However, these models may achieve high validation accuracy by learning specimen-specific artifacts rather than generalizable biological features, limiting real-world clinical utility. We introduce PANDA-PLUS-Bench, a curated benchmark dataset derived from expertly annotated prostate biopsies designed specifically to quantify this failure mode. The benchmark comprises nine carefully selected whole slide images from nine unique patients containing diverse Gleason patterns, with non-overlapping tissue patches extracted at both 512 × 512 and 224 × 224-pixel resolutions across eight augmentation conditions. Using this benchmark, we evaluate seven foundation models (Virchow, Virchow2, UNI, UNI2, Phikon, Phikon-v2, and HistoEncoder) on their ability to separate biological signals from slide-level confounders. Our results reveal substantial variation in robustness across models: the Virchow models achieved the lowest slide-level encoding among large-scale models (slide ID accuracy: 80.7–81.0%), yet Virchow2 exhibited the lowest cross-slide accuracy (47.2%). HistoEncoder, trained specifically on prostate tissue, demonstrated the highest cross-slide accuracy (59.7%) and the strongest slide-level encoding (slide ID accuracy: 90.3%), suggesting tissue-specific training may enhance both biological feature capture and slide-specific signatures. All models exhibited measurable within-slide vs. cross-slide accuracy gaps, though the magnitude varied from 19.9 percentage points (HistoEncoder) to 26.9 percentage points (Phikon). We provide an open-source Google Colab notebook enabling researchers to evaluate additional foundation models against our benchmark using standardized metrics. PANDA-PLUS-Bench addresses a critical gap in foundation model evaluation by providing a purpose-built resource for robustness assessment in the clinically important context of Gleason grading.
Background
Recent personalized nutrition research has reported large inter-individual differences in postprandial glucose responses to identical foods, raising questions about whether these differences reflect food-specific personal effects or normal day-to-day variability in glucose tolerance.
Objectives
To quantify the relative contributions of measurement variability vs person-specific effects to inter-individual glycemic variation, and to define substitution thresholds for when glycemic index (GI) differences produce distinct physiological effects.
Methods
In this secondary analysis with simulated validation, data from 382 healthy adults (1,022 glucose reference tests, 1,116 food tests across 9 carbohydrate-rich foods) were analyzed using a direct comparison scaling model, in which an individual's food response equals their glucose reference response scaled by the food's average GI. Sensitivity analyses included single-reference predictions, restriction to participants with ≥3 reference tests, and exclusion of a protocol-deviating food.
Results
Predicted errors did not exceed the observed glucose reference test-retest variability (mean root mean square deviation [RMSD]: 0.78 vs. 1.02 mmol/L; Cohen's d = 0.54 [0.45, 0.63]), with ∼90% of predictions falling within each participant's own test-retest range. Bland-Altman analysis confirmed negligible systematic bias (-0.01 mmol/L). Synthetic datasets generated from glucose variability and average GI values reproduced observed response distributions without person-specific parameters. GI differences of ≥15 units produced reliably distinguishable responses in a given individual. All sensitivity analyses yielded equal or stronger effect sizes.
Conclusions
In healthy adults under standardized conditions, inter-individual variation in glycemic responses is predominantly accounted for by variability in day-to-day glucose tolerance, propagating through the GI ratio. The GI concept performs within the reproducibility limits of input data.