Institutional Publications
Permanent URI for this collectionhttps://ndkr-library.nipgr.ac.in/handle/123456789/11
Browse
2 results
Search Results
Item PFGPred: A stack ensemble classifier for the identification of fusion genes in plants(Oxford University Press, 2026) Hamid, Fiza; Mukherjee, Kanka; Chaudhary, Sakshi; Kaushik, Love; Kumar, ShaileshFusion genes play crucial roles in plant biological processes but remain far less explored than their human counterparts, largely due to limited validated datasets and the absence of plant-specific prediction tools. Existing approaches often produce high false-positive rates, restricting reliable discovery. To address this gap, we developed Plant Fusion Gene Predictor (PFGPred), an ensemble machine learning framework that integrates Random Forest, XGBoost, and long short-term memory (LSTM) models into a meta-classifier for accurate identification of true and false fusion genes from RNA-Seq data. PFGPred was trained on a high-confidence dataset of fusion genes validated by both RNA-Seq and whole-genome sequencing from Arabidopsis thaliana, Oryza sativa, Triticum aestivum, and Zea mays, to predict and rank candidate fusion genes for future functional validation. It outperformed individual baseline models, achieving accuracies of 0.97 on training data and 0.77 on independent test data. When evaluated on human datasets, it achieved 0.71 accuracy with lower sensitivity, reflecting biological differences between plant and human fusion events. Comparative analyses confirmed that PFGPred reliably identifies validated fusions, demonstrating its utility as a cost-effective, plant-specific prediction tool for high-throughput fusion gene screening and functional genomics research. It is freely available as a web server at http://www.nipgr.ac.in/PFGPred.Item A protocol for the detection of fusion transcripts using RNA-sequencing data(Springer Nature Publishing AG, 2024) Hamid, Fiza; Arora, Simran; Chitkara, Pragya; Kumar, ShaileshFusion transcripts are formed when two genes or their mRNAs fuse to produce a novel gene or chimeric transcript. Fusion genes are well-known cancer biomarkers used for cancer diagnosis and as therapeutic targets. Gene fusions are also found in normal physiology and lead to the evolution of novel genes that contribute to better survival and adaptation for an organism. Various in vitro approaches, such as FISH, PCR, RT-PCR, and chromosome banding techniques, have been used to detect gene fusion. However, all these approaches have low resolution and throughput. Due to the development of high-throughput next-generation sequencing technologies, the detection of fusion transcript becomes feasible using whole genome sequencing, RNA-Seq data, and bioinformatics tools. This chapter will overview the general computational protocol for fusion transcript detection from RNA-sequencing datasets.
