Publications

Newest first, including theses and conference abstracts. Lab members are in bold.

2026

Docking-score landscapes shape active-learning performance across Vina, Glide, and SILCS

Joseph Chung, Aashish Bhatt, Jacob Ede Levine, Yi‐Chun Lin, Mingtian Zhao, Alexander D. MacKerell, Sunhwan Jo, Sai Chandra Kosaraju, Yun Luo

Journal of Computer-Aided Molecular Design

Compares four active-learning virtual screening workflows built on Vina, Glide, and SILCS docking across several protein targets and library sizes. Latent-embedding analyses suggest the docking-score landscape strongly shapes how well active learning performs.

Two PIEZO2 pore renders: left cartoon in blue/red/grey subunits inside a yellow lipid band with boxed pore; right cylind

PINT: Pathway-pathway interactions for predicting interpretable clinical outcomes from gene expression

Sai Phani Parsa, Beomsu Baek, Euiseong Ko, Sai Chandra Kosaraju, Mingon Kang

bioRxiv

PINT models interactions among biological pathways with self-attention over gene expression data. Across five TCGA cancer datasets it outperformed benchmark models in survival analysis and recovered known links between RAS signaling and other regulatory pathways.

Pathway-pathway attention heatmap (5x5, viridis purple-teal-yellow) beside a small orange node graph showing cytokine an

PIMO: pathway-based interpretable multiomics interactions for multiomics integration

Sai Phani Parsa, Sai Chandra Kosaraju, Euiseong Ko, Beomsu Baek, Tesfaye B Marsha, Mingon Kang

Bioinformatics

PIMO is an interpretable, pathway-based deep learning model that captures regulatory effects of DNA methylation and copy number alterations on gene expression. On TCGA cancer datasets it improved the survival C-index by up to 13% over baselines.

Radial wheel: central 'PIMO' circle with eleven pathway names, surrounded by spokes of colored gene-name labels (green,

Fairness-aware supervised hierarchical contrastive semantic learning for sexual dimorphism analysis

Euiseong Ko, Sai Phani Parsa, Sai Chandra Kosaraju, Tesfaye B. Mersha, Mingon Kang

Bioinformatics

FairHICON is a fairness-aware contrastive learning method that finds sex-common and sex-specific predictive features in transcriptomic data. On cancer and asthma datasets it improved prediction by up to 9% while narrowing the performance gap between males and females.

FairHICON schematic: expression matrix into sex-common and male/female encoders, hatched purple/blue/red representation

Slide-Omics: An Interpretable Bi-Directional Attention Framework for Integrating Multi-Omics and Pathology in Cancer Survival Analysis

Carson Green, Ali Momennsab, Maxwell Nguyen, Ethan Reidel, Siddhi Narayan, Sai Phani Parsa, Sai Chandra Kosaraju

Proceedings of the 17th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics

Slide-Omics predicts cancer survival from whole-slide images and pathway-structured multi-omics data using bidirectional cross-attention between tissue morphology and KEGG pathways. It gives competitive survival prediction across four TCGA cohorts along with pathway-level importance scores.

Variational Generative Modeling for Forecasting Protein Diffusion from Molecular Dynamics

Phuong Cam Thai, M. Hidalgo-Soria, Jacob Ede Levine, Chen Yun Wen, Yun Luo, Sai Chandra Kosaraju

Proceedings of the 17th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics

Presents variational generative models that learn membrane protein diffusion from molecular dynamics trajectories and forecast future motion. The base model captures short-term dynamics well; extensions reduce, but do not fully remove, a long-term super-diffusive bias.

Deep learning models with a hierarchical method using RGB images from UAV and proximal sensor data for real-time detection of strawberry plant health

Anupama Singh, Subodh Bhandari, Amar Raheja, Sai Chandra Kosaraju

Autonomous Air and Ground Sensing Systems for Agricultural Optimization and Phenotyping XI (Proc. SPIE 14045)

Classifies strawberry plant health from UAV RGB images with a stage-wise hierarchy of CNN classifiers, checked against field chlorophyll measurements and live UAV video. ResNet18 reached 87.75% accuracy for healthy plant classification.

CholBindNet as an interpretable neural network for cholesterol-binding site classification

Alexis Hernamdez, Aashish Bhatt, Ivan Revilla, Jacob Ede Levine, Sai Chandra Kosaraju, Yun Lyna Luo

Communications Chemistry

CholBindNet is an interpretable atom-based graph neural network, trained with a positive-unlabeled strategy on over 800 cholesterol-bound transmembrane protein structures, that classifies cholesterol-binding sites. It outperformed AlphaFold3, P2Rank, and DiffDock on this task.

Ribbon render of the PIEZO2 channel (tan, lilac, pale blue subunits) with three magenta space-filling cholesterol molecu

Ensemble Based Real Time Facial Emotion Recognition with Adaptive Feedback for Mental Health Support

Khadeeja Hussain, Sai Phani Parsa, Fateme Jamshidi, Sai Chandra Kosaraju

2026 IEEE International Conference on Electro Information Technology (eIT)

Describes a real-time virtual assistant that recognizes facial emotions with an ensemble of three CNNs and responds with personalized multimedia feedback. Trained on FER2013, the ensemble reached 70.33% accuracy, higher than any single backbone.

Prediction of bacterial protein–compound interactions with only positive samples

Ki‐Hwa Kim, Avinash Yaganapu, Sai Chandra Kosaraju, Aashish Bhatt, Yun Lyna Luo, Sai Phani Parsa, Juyeon Park, Hyun Lee, Jun Hyuck Lee, Tae‐Jin Oh, M. Kang

Bioinformatics

BIN-PU is a positive-unlabeled learning framework that trains bacterial compound-protein interaction models using only known positive interactions. It outperformed existing PU models on bacterial cytochrome P450 data, and predictions on uncurated data were tested experimentally.

Two docking poses: salmon and cyan steroid sticks hovering over a grey heme with dashed distance lines (4.49 Å, 4.69 Å),

TSSR: Two-Stage Swap-Reward-Driven Reinforcement Learning for Character-Level SMILES Generation

Jacob Ede Levine, Yun Lyan Luo, Sai Chandra Kosaraju

arXiv

TSSR is a two-stage reinforcement learning framework for character-level SMILES generation that first rewards token swaps fixing syntax, then rewards fewer RDKit chemistry errors. On the MOSES benchmark it improved syntactic and chemical validity without reducing diversity.

Two rows of 2D molecule drawings showing an invalid SMILES repaired step by step, problem atoms circled in green, under

2025

Evidential deep learning-based ALK-expression screening using H&E-stained histopathological images

Sai Chandra Kosaraju, Sai Phani Parsa, Dae Hyun Song, Hyo Jung An, Yoon‐La Choi, Joungho Han, Jung Wook Yang, Mingon Kang

npj Digital Medicine

Develops an interpretable, evidence-based deep learning model that screens for ALK alterations in non-small cell lung cancer directly from H&E-stained pathology images. It reached over 95% accuracy on both resection and biopsy datasets.

ALK screeningnpj Digit. Med.
Low-power H&E scan of a mucinous lung tumor

Race-Specific Impact of Telehealth Advance Care Planning on Cost of Dementia: A Cost Prediction Study

Peter S. Reed, Yonsu Kim, Jay J. Shen, Sai Kosaraju, Mingon Kang, Jennifer Carson, Iulia Ioanitoaia Chaudhry, Sarah Kim, Connor Jeong, Yena Hwang, Ji Won Yoo

International Journal of Environmental Research and Public Health

Uses electronic health records of dual-eligible Medicare/Medicaid patients with dementia, with regression and machine learning, to estimate how telehealth advance care planning relates to hospitalization costs by race. Receiving advance care planning was associated with USD 23,928.84 lower costs.

Error distributions by race for the cost prediction models

Prediction of ALK Expression on H&E-Stained Pathology Slide Images Using a Deep-Learning Model: the Further Validation

Jung Wook Yang, Sai Chandra Kosaraju, Wookjae Jung, Dae Hyun Song, Hyo Jung An, Mee Sook Roh, Ahrong Kim, Dong‐Hoon Shin, Mingon Kang

Laboratory Investigation

Conference abstract on further validation of a deep learning model that predicts ALK expression from H&E-stained pathology slide images.

ALK validationLab. Invest.
High-power H&E of mucinous adenocarcinoma

2024

Interpretable and Evidential Deep Learning for Medical Image Analysis

Sai Chandra Kosaraju

University of Nevada, Las Vegas (PhD dissertation)

Dissertation developing three interpretable and evidential deep learning models for whole-slide histopathology images: Deep-Hipo for multi-scale patches, HipoMap for slide-level prediction, and Deep-PATHO for detecting ALK rearrangement morphology.

Ph.D. dissertationUNLV
Whole-slide H&E scan of lung adenocarcinoma with annotated regions

ActivePCA: A Novel Framework Integrating PCA and Active Machine Learning for Efficient Dimension Reduction

Priyanka Bhyregowda, Mohammad Masum, Lohuwa Mamudu, Mohammed Chowdhury, Sai Chandra Kosaraju, Hossain Shahriar

IEEE Annual Computers, Software, and Applications Conference (COMPSAC)

In medical data analysis, addressing challenges from high-dimensional datasets is crucial due to issues related to computational complexity, resource utilization, and model interpretability. Principal Component Analysis (PCA), a prevalent dimension reduction technique, aims to tackle these challenges by transforming high-dimensional data into a lower-dimensional representation while preserving ...

CMI: Cluster-Centric Missing Value Imputation with Feature Consistency

Megha Gupta, Shripal Shah, Mohammad Masum, Sai Chandra Kosaraju

2024 IEEE 14th Annual Computing and Communication Workshop and Conference (CCWC)

CMI imputes missing numerical values from clusters of similar data points and uses SHAP values to interpret feature importance after imputation. On liver and kidney disease datasets it outperformed mean, KNN, and MICE imputation.

2023

Evidential deep learning for trustworthy prediction of enzyme commission number

So‐Ra Han, Mingyu Park, Sai Chandra Kosaraju, JeungMin Lee, Hyun Lee, Jun Hyuck Lee, Tae‐Jin Oh, Mingon Kang

Briefings in Bioinformatics

ECPICK is an evidential deep learning model for trustworthy prediction of enzyme commission numbers from protein sequences. It identifies the amino acids behind each prediction without sequence alignment, which helped reveal potential new motif sites in microorganisms.

ECPICK overview: organism icons (fungi, human, bacteria, plant, animal, virus) feeding a DNA strand into a neural networ

2022

Deep learning-based framework for slide-based histopathological image analysis

Sai Chandra Kosaraju, Jeongyeon Park, Hyun Jung Lee, Jung Wook Yang, Mingon Kang

Scientific Reports

HipoMap converts a whole-slide image of any shape or size into a structured, image-like representation map for slide-level analysis with CNNs. It reached an AUC of 0.96 for lung cancer classification and also improved survival analysis.

HipoMapSci. Rep.
Whole-slide H&E scan of small cell lung cancer with annotated region

2020

Deep-Hipo: Multi-scale receptive field deep learning for histopathological image analysis

Sai Chandra Kosaraju, Jie Hao, Hyun Min Koh, Mingon Kang

Methods

Deep-Hipo analyzes histopathology image patches at high and low magnification together, capturing morphology in both small and large receptive fields. It was assessed on gastric cancer whole-slide images and on TCGA stomach and colon adenocarcinoma slides.

Deep-HipoMethods
Normal gastric gland at 20x magnification, H&E

2019

PAGE-Net: Interpretable and Integrative Deep Learning for Survival Analysis Using Histopathological Images and Genomic Data

Jie Hao, Sai Chandra Kosaraju, Nelson Zange Tsaku, Dae Hyun Song, Mingon Kang

Pacific Symposium on Biocomputing 2020

PAGE-Net combines histopathological images, gene expression, and demographic data in an interpretable deep learning model for survival prediction. On TCGA glioblastoma data it reached a C-index of 0.702 and identified prognostic image and genomic factors.

PAGE-NetPSB
Glioblastoma H&E patches with survival-relevant regions highlighted

Texture-based Deep Learning for Effective Histopathological Cancer Image Classification

Nelson Zange Tsaku, Sai Chandra Kosaraju, Tasmia Aqila, Mohammad Masum, Dae Hyun Song, Ananda Mohan Mondal, Hyun Min Koh, Mingon Kang

2019 IEEE International Conference on Bioinformatics and Biomedicine (BIBM)

CAT-Net is a texture-based deep neural network that uses dilated convolutions to learn morphological patterns for classifying cancer in histopathological whole-slide images. It outperformed benchmark methods on accuracy, precision, recall, and F1 score with lower model complexity.

DoT-Net: Document Layout Classification Using Texture-Based CNN

Sai Chandra Kosaraju, Mohammad Masum, Nelson Zange Tsaku, Pritesh Patel, Tanju Bayramoglu, Girish Modgil, Mingon Kang

2019 International Conference on Document Analysis and Recognition (ICDAR)

DoT-Net is a texture-based CNN that classifies scanned document blocks as text, image, table, mathematical expression, or line-diagram, going beyond text versus non-text. It outperformed existing document layout classifiers on accuracy, F1 score, and AUC.

Table of Contents Recognition in OCR Documents using Image-based Machine Learning

Sai Chandra Kosaraju, Nelson Zange Tsaku, Pritesh Patel, Tanju Bayramoglu, Girish Modgil, Mingon Kang

Proceedings of the 2019 ACM Southeast Conference

Recognizes table of contents pages in OCR documents from page images rather than text, classifying pages with one-dimensional horizontal projections. Tested on thesis and dissertation PDFs, it outperformed other methods.

Boosting Recommendation Systems through an Offline Machine Learning Evaluation Approach

Nelson Zange Tsaku, Sai Chandra Kosaraju

Proceedings of the 2019 ACM Southeast Conference

Proposes a recommendation strategy that uses domain knowledge for new users and machine learning for existing users, plus an offline/online evaluation strategy to reduce customer churn. Each system was implemented on the movie Tweetings dataset.

Document Layout Analysis and Recognition Systems

Sai Chandra Kosaraju

Kennesaw State University (MS thesis)

Thesis on a knowledge extraction framework for OCR documents that rebuilds document structure (sections, table of contents, block types) and then answers domain-specific questions by retrieving the most relevant sentences or paragraphs.

2018

Automatic knowledge extraction from OCR documents using hierarchical document analysis

Mohammad Masum, Sai Chandra Kosaraju, Tanju Bayramoglu, Girish Modgil, Mingon Kang

Proceedings of the 2018 Conference on Research in Adaptive and Convergent Systems

Presents a framework that takes keyword queries and extracts the most relevant knowledge from OCR documents using text mining, with relevance ranking. It was tested on GE Power documents of more than a hundred pages each.