Publications of NIPGR Scientists
Permanent URI for this communityhttps://ndkr-library.nipgr.ac.in/handle/123456789/1
Browse
10 results
Search Results
Item Validation of plant fusion peptides using proteomics data(Springer Nature Publishing AG, 2026) Hamid, Fiza; Aftab, Sahrish; Shree, Tanu; Kumar, ShaileshFusion transcripts and their fused protein products are emerging as exciting entities in molecular biology, offering potential applications in diagnostics and therapeutics. These fusion proteins, derived from the translation of fusion transcripts, hold promise as unique biomarkers and targets for intervention. While numerous algorithms exist to identify fusion RNAs, the detection and validation of their protein counterparts through proteomics remains a growing area of research. This challenge is particularly intriguing in plant biology, where fusion events may affect stress responses, development, and adaptation. This chapter provides an accessible and practical workflow for validating plant fusion peptides using publicly available proteomics datasets.Item PFGPred: A stack ensemble classifier for the identification of fusion genes in plants(Oxford University Press, 2026) Hamid, Fiza; Mukherjee, Kanka; Chaudhary, Sakshi; Kaushik, Love; Kumar, ShaileshFusion genes play crucial roles in plant biological processes but remain far less explored than their human counterparts, largely due to limited validated datasets and the absence of plant-specific prediction tools. Existing approaches often produce high false-positive rates, restricting reliable discovery. To address this gap, we developed Plant Fusion Gene Predictor (PFGPred), an ensemble machine learning framework that integrates Random Forest, XGBoost, and long short-term memory (LSTM) models into a meta-classifier for accurate identification of true and false fusion genes from RNA-Seq data. PFGPred was trained on a high-confidence dataset of fusion genes validated by both RNA-Seq and whole-genome sequencing from Arabidopsis thaliana, Oryza sativa, Triticum aestivum, and Zea mays, to predict and rank candidate fusion genes for future functional validation. It outperformed individual baseline models, achieving accuracies of 0.97 on training data and 0.77 on independent test data. When evaluated on human datasets, it achieved 0.71 accuracy with lower sensitivity, reflecting biological differences between plant and human fusion events. Comparative analyses confirmed that PFGPred reliably identifies validated fusions, demonstrating its utility as a cost-effective, plant-specific prediction tool for high-throughput fusion gene screening and functional genomics research. It is freely available as a web server at http://www.nipgr.ac.in/PFGPred.Item The intersection of AI and genomics in health and disease: Advancements and applications(Elsevier B.V., 2026) Kaushik, Love; Vivek, A T; Arora, Simran; Hamid, Fiza; Mukherjee, Kanka; Bisht, Niyati; Chaudhary, Sakshi; Shukla, Jagriti; Nawani, Sakshi; Kumar, ShaileshAI and genomics are revolutionizing precision medicine by using machine learning (ML) to analyze large-scale next-generation sequencing (NGS) data, identifying genetic mutations and biomarkers for personalized therapies. In practice, this accelerates drug discovery and enhances variant detection, while in cancer genomics, AI enables early detection via liquid biopsies and refines treatment by integrating multi-omics data to improve therapeutic precision. However, challenges such as data biases in underrepresented populations, limited model interpretability, and ethical concerns regarding privacy and algorithmic inequity hinder clinical adoption and demand robust governance. Efforts to diversify datasets also face standardization hurdles, although explainable AI and federated learning provide promising solutions for improving transparency and privacy. In this chapter, we discuss the role of AI in advancing genomics from diagnostics to novel therapies and emphasize the need for equitable frameworks to ensure responsible implementation, thereby paving the way for breakthroughs in personalized medicine.Item Breaking and making genes: the genesis of novel traits in plants(John Wiley & Sons, 2026) Hamid, Fiza; Arora, Simran; Kumar, ShaileshUnderstanding the mechanisms by which plants adapt, evolve, and acquire new traits is crucial for enhancing agricultural resilience and productivity in the face of global challenges. Among the various mechanisms that drive new gene evolution, gene fusion has emerged as a significant yet relatively understudied contributor. It can arise through chromosomal rearrangements or RNA processing mechanisms, merging segments from different genes to produce novel fusion transcripts. In plants, these fusion events have been associated with key biological functions, including the regulation of specialized metabolism, stress responses, and developmental changes. While fusion genes have been extensively studied in humans, mainly due to their oncogenic potential, their prevalence and functional relevance in plants remain relatively underexplored. This review offers a detailed overview of the molecular mechanisms underlying gene fusion formation, highlighting their participation in gene evolution, functional diversification, and plant adaptation. In addition, we discuss current methodologies for detecting and validating fusion events, including high-throughput sequencing technologies and emerging single-cell sequencing platforms, and outline promising directions for future research aimed at elucidating their biological significance. Collectively, these insights emphasize the expanding importance of gene fusions in plant biology and underscore the need for further investigation into their regulatory and evolutionary roles.Item Molecular and expression analyses indicate the role of fusion transcripts in mediating abiotic stress responses in chickpea(Frontiers Media S.A., 2025) Hamid, Fiza; Zahra, Shafaque; Kumar, ShaileshUnderstanding the transcriptome diversity is essential for deciphering the transcriptional level regulation. High-throughput sequencing technologies have facilitated the detection of fusion transcripts (FTs), which are chimeric mRNA molecules derived from gene fusions due to chromosomal rearrangements or via the splicing machinery at the RNA level. In this study, we investigated the transcriptome complexity in Cicer arietinum resulting from fusion events using high-throughput RNA-Seq datasets from five tissues, i.e., stem, leaves, buds, flowers, and pods, and two abiotic stress conditions, i.e., drought and salinity. Of the 328 unique FTs identified, 69% exhibited the presence of canonical splice sites at their junction, indicating their generation via trans-splicing. Functional annotation and enrichment analyses of fusion partners suggested that these transcripts may expand functional diversity. A total of 10 FTs were validated via RT-PCR followed by Sanger sequencing, which are the first FTs described in the important legume chickpea. Expression analysis of fusion transcripts across various tissues and under abiotic stress conditions revealed evidence of context-dependent regulation. Furthermore, 120 fusion gene pairs were found to be conserved across 17 chickpea genotypes, highlighting their potential biological significance and stability within the species. Overall, these findings suggest that fusion transcripts may contribute to regulatory mechanisms underlying abiotic stress responses in chickpea.Item Fusion transcripts in plants: hidden layer of transcriptome complexity(Elsevier B.V., 2025) Arora, Simran; Hamid, Fiza; Kumar, ShaileshIn the realm of genetic information, fusion transcripts contribute to the intricate complexity of the transcriptome across various organisms. Recently, Cong et al. investigated these RNAs in rice, maize, soybean, and arabidopsis (Arabidopsis thaliana), revealing conserved characteristics. These findings enhance our understanding of the functional roles and evolutionary significance of these fusion transcripts.Item The landscape of fusion transcripts in plants: a new insight into genome complexity(BioMed Central Ltd, 2024) Chitkara, Pragya; Singh, Ajeet; Gangwar, Rashmi; Bhardwaj, Rohan; Zahra, Shafaque; Arora, Simran; Hamid, Fiza; Arya, Ajay; Sahu, Namrata; Chakraborty, Srija; Ramesh, Madhulika; Kumar, ShaileshBackground Fusion transcripts (FTs), generated by the fusion of genes at the DNA level or RNA-level splicing events significantly contribute to transcriptome diversity. FTs are usually considered unique features of neoplasia and serve as biomarkers and therapeutic targets for multiple cancers. The latest findings show the presence of FTs in normal human physiology. Several discrete reports mentioned the presence of fusion transcripts in planta, has important roles in stress responses, morphological alterations, or traits (e.g. seed size, etc.). Results In this study, we identified 169,197 fusion transcripts in 2795 transcriptome datasets of Arabidopsis thaliana, Cicer arietinum, and Oryza sativa by using a combination of tools, and confirmed the translational activity of 150 fusion transcripts through proteomic datasets. Analysis of the FT junction sequences and their association with epigenetic factors, as revealed by ChIP-Seq datasets, demonstrated an organised process of fusion formation at the DNA level. We investigated the possible impact of three-dimensional chromatin conformation on intra-chromosomal fusion events by leveraging the Hi-C datasets with the incidence of fusion transcripts. We further utilised the longread RNA-Seq datasets to validate the most reoccurring fusion transcripts in each plant species followed by further authentication through RT-PCR and Sanger sequencing. Conclusions Our findings suggest that a significant portion of fusion events may be attributed to alternative splicing during transcription, accounting for numerous fusion events without a proportional increase in the number of RNA pairs. Even non-nuclear DNA transcripts from mitochondria and chloroplasts can participate in intra- and inter-chromosomal fusion formation. Genes in close spatial proximity are more prone to undergoing fusion formation, especially in intra-chromosomal FTs. Most of the fusion transcripts may not undergo translation and serve as long non-coding RNAs. The low validation rate of FTs in plants indicated that the fusion transcripts are expressed at very low levels, like in the case of humans. FTs often originate from parental genes involved in essential biological processes, suggesting their relevance across diverse tissues and stress conditions. This study presents a comprehensive repository of fusion transcripts, offering valuable insights into their roles in vital physiological processes and stress responses.Item PFusionDB: a comprehensive database of plant-specific fusion transcripts(Springer Nature Publishing AG, 2024) Arya, Ajay; Arora, Simran; Hamid, Fiza; Kumar, ShaileshFusion transcripts (FTs) are well known cancer biomarkers, relatively understudied in plants. Here, we developed PFusionDB (www.nipgr.ac.in/PFusionDB), a novel plant-specific fusion-transcript database. It is a comprehensive repository of 80,170, 39,108, 83,330, and 11,500 unique fusions detected in 1280, 637, 697, and 181 RNA-Seq samples of Arabidopsis thaliana, Oryza sativa japonica, Oryza sativa indica, and Cicer arietinum respectively. Here, a total of 76,599 (Arabidopsis thaliana), 35,480 (Oryza sativa japonica), 72,099 (Oryza sativa indica), and 9524 (Cicer arietinum) fusion transcripts are non-recurrent i.e., only found in one sample. Identification of FTs was performed by using a total of five tools viz. EricScript-Plants, STAR-Fusion, TrinityFusion, SQUID, and MapSplice. At PFusionDB, available fundamental details of fusion events includes the information of parental genes, junction sequence, expression levels of fusion transcripts, breakpoint coordinates, strand information, tissue type, treatment information, fusion type, PFusionDB ID, and Sequence Read Archive (SRA) ID. Further, two search modules: ‘Simple Search’ and ‘Advanced Search’, along with a ‘Browse’ option to data download, are present for the ease of users. Three distinct modules viz. ‘BLASTN’, ‘SW Align’, and ‘Mapping’ are also available for efficient query sequence mapping and alignment to FTs. PFusionDB serves as a crucial resource for delving into the intricate world of fusion transcript in plants, providing researchers with a foundation for further exploration and analysis. Database URL: www.nipgr.ac.in/PFusionDB.Item A protocol for the detection of fusion transcripts using RNA-sequencing data(Springer Nature Publishing AG, 2024) Hamid, Fiza; Arora, Simran; Chitkara, Pragya; Kumar, ShaileshFusion transcripts are formed when two genes or their mRNAs fuse to produce a novel gene or chimeric transcript. Fusion genes are well-known cancer biomarkers used for cancer diagnosis and as therapeutic targets. Gene fusions are also found in normal physiology and lead to the evolution of novel genes that contribute to better survival and adaptation for an organism. Various in vitro approaches, such as FISH, PCR, RT-PCR, and chromosome banding techniques, have been used to detect gene fusion. However, all these approaches have low resolution and throughput. Due to the development of high-throughput next-generation sequencing technologies, the detection of fusion transcript becomes feasible using whole genome sequencing, RNA-Seq data, and bioinformatics tools. This chapter will overview the general computational protocol for fusion transcript detection from RNA-sequencing datasets.Item In silico identification of tRNA fragments, novel candidates for cancer biomarkers, and therapeutic targets(Springer Nature Publishing AG, 2024) Singh, Ankita; Zahra, Shafaque; Arora, Simran; Hamid, Fiza; Kumar, ShaileshThe identification of a wide variety of RNA molecules using high-throughput sequencing techniques in the transcriptome pool of living organisms has revealed hidden regulatory insights in the cell. The class of non-coding RNA fragments produced from transfer RNA, or tRFs, is one such example. They are heterogeneously sized molecules with lengths ranging between 15 and 50 nt. They have a history of being dysregulated in human malignancies and other illnesses. The detection of these molecules has been made easier by a variety of bioinformatics techniques. The various types of tRFs and how they relate to cancer are covered in this chapter. It also provides a summary of the biological significance of tRFs reported in human cancer. Additionally, it emphasizes the utilities of databases and computational tools that have been created by different research teams for the investigation of tRFs. This will further aid the exploration and analysis of tRFs in cancer research and will support future advancement and a better comprehension of these molecules.
