Research Computing

Research Computing Symposium 2026

Research Computing is hosting a half-day of presentations to highlight the many different ways our research community is using the HPC.  Please plan to come to the University Center, room 300, the morning of September 3rd, 8:30a – noon-ish. Feel free to pop in and out as your schedule permits, but make sure you at least stop by first thing in the morning for a coffee, snack, and a chat with the research computing team and other researchers.

Schedule

8:30am - Hang out, drink coffee, and eat snacks

9:00am - Opening remarks

9:15am - Christodoulos Kyriakopoulos (CERI) - "HPC in Earthquake Modeling"

9:45am - Brian Wentzloff (Research Computing)  - R Studio on Big Blue demo

10:00am - Daniel Frimpong (Biology) - "Supercomputing Centrapalus: Assembling a Complex Plant Genome"

10:30am - Imran Sharker (CAESAR) - "The Impact of HPC on Advanced CFD Simulations and Large-Scale Data Processing"

11:00am - Eric Spangler (Research Computing) - llama.cpp on Big Blue demo

11:30am - Maziar Ganji (Biology) - "HPC-Accelerated Random Forest Optimization for PTSD Subtype Classification and Biomarker Discovery in High-Dimensional Data"

Noon - Closing remarks


Daniel Frimpong (Biology) - "Supercomputing Centrapalus: Assembling a Complex Plant Genome"

The genus Centrapalus represents one of the most ecologically and taxonomically significant lineages within the largest family of flowering plants, Asteraceae. Initially classified into the Vernonia genus ( e.g., Centrapalus pauciflorus, syn. Vernonia pauciflora) and immensely grown for their seed oil composition and other natural epoxidized fatty acids (Carlsson et al., 2011; Perdue et al., 1989), the species within the lineage are rich in bioactive phytochemicals and other secondary metabolites (e.g., sesquiterpenes, flavonoids, and vernolic acid) that are of greater value inbiomedical and agricultural discovery (Funk et al., 2009; Robinson & Washington, 1999).

The genomes of plants, especially Asteraceae, pose one of the most formidable hurdles in genomic studies due to the high levels of repeat DNA regions, complex heterozygosity, and other structural variations. Being able to decode and assemble a high-quality reference genome for Centrapalus will provide the biological blueprint needed to pinpoint specific gene families and the associated regulatory pathways driving these high-value secondary metabolites, ultimately enabling the annotation of certain functional characteristics and in-depth evolutionary studies across the Asteraceae. Nevertheless, reconstructing chromosome-scale assemblies from raw sequencing data to a fully assembled chromosomal scaffold requires huge computational power and resources that goes way beyond the regular capabilities of a standard desktop computer.

Here, we present a bioinformatics assembly pipeline and a quality-control workflow for Centrapalus, highlighting how High-Performance Computing (HPC) resources can be used to overcome traditional computational hurdles. High-accuracy PacBio HiFi long-read sequencing datawere employed to create primary contig graphs using the de novo assembler hifiasm. To expand the assembly to chromosome-level contigs, long-range scaffolds were generated using the Yet Another Hi-C Scaffolder (YaHS) software. Quality evaluation was done through Benchmarking Universal Single-Copy Orthologs (BUSCO) using the eudicots database. By leveraging the HPC cluster, high-memory partitions, and containerized bioinformatics pipelines (i.e., Conda and Slurm), this study successfully transformed the raw sequencing reads into a highly continuous reference genome for Centrapalus, ultimately enabling precise gene expression profiling, evolutionary comparative genomics, and the investigation of functional traits across related plant taxa.


Imran Sharker (CAESAR) - "The Impact of HPC on Advanced CFD Simulations and Large-Scale Data Processing"

High-fidelity fluid dynamics research is fundamentally limited by computational power. Running Large Eddy Simulations (LES) and analyzing the massive, complex datasets they produce can take a local desktop machine several months to complete. This presentation highlights how High-Performance Computing (HPC) overcomes that limitation. We will walk through the critical impact of HPC resources on the entire CFD pipeline—from scalable grid computations to the rapid post-processing of massive flow fields. By shifting from local machines to advanced computing clusters, researchers can accelerate their workflows by orders of magnitude, making resource-intensive techniques like advanced spectral analysis and vortex dynamics modeling highly efficient and accessible.


Maziar Ganji (Biology) - "HPC-Accelerated Random Forest Optimization for PTSD Subtype Classification and Biomarker Discovery in High-Dimensional Data"

Biomarker identification can involve many different analytical approaches. In this project, we approach biomarker discovery as a supervised machine-learning problem, using binary classification to identify molecular features that can distinguish between PTSD clinical subtypes. With thousands of candidate molecular features and a comparatively small number of samples, the primary challenge is not simply model fitting, but achieving a robust estimation of predictive performance while identifying features that consistently contribute to discrimination between the subtypes. 
 
This talk will demonstrate how Random Forest (RF) algorithm is used in our biomarker identification pipeline and how High-Performance Computing (HPC) enables rigorous model evaluation at scale. I will discuss hyperparameter tuning, nested cross-validation, and feature-importance analysis, and show how these computationally intensive steps can be distributed across an HPC cluster using parallel jobs. The goal is to show how HPC can make a reproducible machine-learning workflow for biomarker identification substantially more efficient and scalable.