The Consortium of Molecular Design at BYU provides cutting edge interdisciplinary research opportunities for students to push the envelope for protein engineering and drug discovery.
We use close collaboration between laboratories at BYU in Physics, Chemistry, Computer Science, LifeSciences, and Engineering to tackle these challenging topics from all angles.
We actively seek industrial collaboration and support for our efforts and are excited to explore mutually beneficial application of all state-of-the-art technologies to revolutionize molecular design.
News and Events
Selected Publications
Artificial intelligence (AI)-based prostate cancer detection through whole slide images (WSIs) offers promising potential to address the global pathologist shortage while improving clinical consistency. Digital slides and improving image analysis methods encourage the creation of tools to aid in WSI classification. Despite promising advances, these tools are still limited by available training data. Current publicly available datasets, such as Kaggle's PANDA Challenge, while large in scale, rely on slide-level labels that may introduce noise and limit model reliability. Others contain detailed annotations, but are smaller in size due to manual processing efforts. In this work, we introduce PANDA-PLUS, a 546-image dataset derived from PANDA images with improved pixel-level annotations, as well as an accompanying annotation pipeline that reduces pathologists' time commitment. We present a detailed comparative analysis between PANDA-PLUS and PANDA using Gleason score and ISUP grade, supported by agreement values, κ, and PABAK under multiple weighting schemes. The results demonstrate consistently lower grading in PANDA-PLUS, with disagreement patterns especially pronounced at higher grades. We also demonstrate through single rater grading of various annotation granularities how slide- and patch-level labels may distort grading proportions and alter image scores. PANDA-PLUS not only improves annotation granularity and reduces label noise but also exposes potential grading errors in the original PANDA dataset. We present PANDA-PLUS's annotations as an improved alternative to the PANDA labels and conclude that it represents a step forward in the development of higher-quality public datasets for clinical AI applications in prostate cancer pathology.
Background: Understanding how different modeling strategies affect associations in nutritional epidemiology is critical, especially given the temporal complexity of dietary and health data.
Objective: To compare how different modeling frameworks—including isotemporal versus time-lagged designs and frequentist versus Bayesian inference—affect estimated associations between carbohydrate subtypes and adiposity.
Methods: Longitudinal data of 415 adults from the NoHoW Study were used to investigate associations between four carbohydrate predictors (free sugars, intrinsic sugars, starch, and dietary fiber) and three indices of adiposity (body fat percentage, BMI, and waist circumference) as outcomes. Four statistical approaches were used contrasting frequentist and Bayesian methods across both isotemporal (concurrent measurement) and time-lagged (6-month temporal shift) frameworks. To specifically evaluate change in adiposity outcomes over time, we implemented additional baseline-adjusted longitudinal models.
Results: Isotemporal and time-lagged models showed directional agreement for nearly all associations; in all but one case, the models either aligned in the direction of the association or differed only in relation to the null. However, time-lagged models identified statistically significant associations and produced larger effect sizes for body fat outcomes and for starch and fiber predictors. Other associations, including intrinsic and free sugars, were weaker and varied with model specification, losing statistical support under time-lagged models. Frequentist models exhibited greater variation across temporal frameworks, including one directional shift among significant associations. Effect estimates were substantially attenuated after adjustment for baseline adiposity.
Discussion: Time-lagged modeling shifted associations between carbohydrate intake and anthropometric outcomes, with increased effect sizes and additional significant associations for starch and fiber, and fewer statistically significant associations for intrinsic and extrinsic sugars. In contrast to frequentist models, Bayesian models yielded more stable and consistent estimates across time-lagged and isotemporal frameworks, showing no differences in the directions of associations across temporal frameworks. Models unadjusted for baseline adiposity overstate dietary impacts; including baseline adiposity is essential to isolate true diet-change effects from initial weight.
Conclusion: Our findings suggest that incorporating temporal structure, especially through Bayesian models, can uncover relevant relationships that concurrent models may overlook. This study demonstrates that model specification, both in temporal framework and statistical approach, meaningfully influences both the detection and interpretations of associations in nutritional epidemiology.
Protein function emerges from dynamic conformational changes, yet structure prediction methods provide only static snapshots. While AlphaFold3 (AF3) predicts protein structures, the potential for extracting dynamic information from its ensemble predictions has remained underexplored. Here, we demonstrate that AF3 structural ensembles contain substantial dynamic information that correlates remarkably well with molecular dynamics simulations (MD). We developed ChronoSort, a novel algorithm that organizes static structure predictions into temporally coherent trajectories by minimizing structural differences between neighboring frames. Through systematic analysis of four diverse protein targets, we show that root-mean-square fluctuations derived from AF3 ensembles can correlate strongly with those from MD (r = 0.53 to 0.84). Principal component analysis reveals that AF3 predictions capture the same collective motion patterns observed in molecular dynamics trajectories, with eigenvector similarities significantly exceeding random distributions. ChronoSort trajectories exhibit structural evolution profiles comparable to MD. These findings suggest that modern AI-based structure prediction tools encode conformational flexibility information that can be systematically extracted without expensive MD. We provide ChronoSort as open-source software to enable broad community adoption. This work offers a novel approach to extracting functional insights from structure prediction tools in minutes, with significant implications for synthetic biology, protein engineering, drug discovery, and structure–function studies.